Skip to content
Judith Solomon
Decision TreesClassificationscikit-learnPython

Internet Advertisements — Decision Trees

Classifying images as adverts or not, across 1,558 variables.

Role
Learning project
When
August 2026
Tools
Python, pandas, scikit-learn, Jupyter Notebook

Description

Practice with decision trees using the Internet Advertisements dataset from the UCI Machine Learning Repository — deciding whether an image on a web page is an advertisement.

The dataset is deliberately awkward: 3,279 observations described by 1,558 explanatory variables, and heavily imbalanced with 459 advertisements against 2,820 non-advertisements. That imbalance is the point — it is what makes the naive approach look good and be wrong.

Dataset

  • 3,279 observations
  • 1,558 explanatory variables
  • 459 advertisements, 2,820 non-advertisements

Results

Drawn from the classification report the notebook prints: 96% accuracy over 656 pages. The number that matters is the one that is lowest — the model finds 80% of the adverts, and is right 94% of the time it calls one. Only 92 of those 656 pages are adverts, and a lopsided class is exactly where a single accuracy figure stops being informative.