Skip to content
Judith Solomon
Logistic RegressionNLPClassificationPython

SMS Spam Classification

Binary text classification with logistic regression — and learning what the metrics actually mean.

Role
Learning project
When
August 2026
Tools
Python, pandas, NumPy, scikit-learn, Matplotlib, Jupyter Notebook

Description

A binary text classifier that separates spam SMS messages from normal ones, built on the UCI SMS Spam Collection dataset.

TF-IDF turns each message into numbers, and logistic regression does the classifying. The more useful half of the project was everything after that: understanding why accuracy alone is a misleading score when one class is much rarer than the other.

What I worked through

  • TF-IDF feature extraction to turn SMS text into numerical features.
  • Logistic regression for binary classification.
  • Train/test splitting and evaluation.
  • Confusion matrices — and reading them properly.
  • Precision, recall and F1, and when each one is the score that matters.
  • Grid search for hyperparameter tuning, including on a larger search space.
  • Extending the same ideas to multi-class and multi-label problems.

Results

The held-out test set. 1,206 ordinary messages kept and not one of them wrongly binned; of the spam, 148 caught and 39 let through. Erring that way round is the right way round for a spam filter — a lost real message costs more than one that slips into the inbox.
The same classifier as a ROC curve — AUC 0.99.