Skip to content
Judith Solomon
Feature EngineeringPreprocessingComputer VisionNLP

Feature Extraction & Preprocessing

Turning text and images into numbers a model can learn from.

Role
Learning project
When
September 2026
Tools
Python, pandas, NumPy, scikit-learn, Jupyter Notebook

Description

A working set of notebooks on the step before modelling: getting raw text and images into a numerical form that a model can use at all.

This is the part of the pipeline I find most interesting, because it is where domain knowledge enters. The algorithm cannot recover information that the features threw away.

What it covers

  • Bag-of-words representations for text.
  • One-hot encoding for categorical variables.
  • Standardisation and scaling.
  • Pixel intensities as image features, and their limits.
  • Points of interest in images.
  • SIFT — scale-invariant feature transform.
  • Extracting text from images.

What the notebooks produce

SIFT keypoints on her own avatar illustration — each circle is a feature the detector found, sized by its scale and turned by its orientation.
The same image in greyscale, with the detected points of interest marked.