One of the things that may be bothering you if you followed #1 and #2 is the fact that we never talked about what happens when our input columns have very different numeric ranges. I know it bothered be when I first started my ML journey.
In the diabetes example we got lucky as most of the values were already in similar ballparks. But that’s not always the case.
In this post we’ll take our house price dataset from the previous article and add a new column: house_age. And then we’ll watch our model fail because of it.
So what is really the problem?
Our data now looks like this:
sqft | house_age | neighbourhood | price |
|---|---|---|---|
| 1500 | 10 | North | 320,000 |
| 850 | 45 | South | 195,000 |
| 2200 | 3 | North | 410,000 |
sqft goes from 500 to 3,000. house_age goes from 0 to 50. One column is roughly 60 times larger than the other.
We’re going to use K-Nearest Neighbours (KNN) here. KNN predicts a house’s price by finding the most similar houses in the dataset and averaging their prices. “Similar” means physically close when you measure the distance across all features.
The issue is starts to become obvious once we think about it: when KNN calculates the distance between two houses, sqft can contribute up to 2,500 units of difference while house_age contributes at most 50. So house_age basically gets ignored. A brand new house and a 45-year-old house with the same square footage look identical to the model.
Not ideal, as some features weigh more than others. But feature scaling comes to rescue.
Let’s see how we can fix it
Let’s create our virtual env and install the required dependencies:
# requirements.txt
pandas==2.2.2
scikit-learn==1.5.0
python -m venv venv
source venv/bin/activate
# venvScriptsactivate on Windows
pip install -r requirements.txt
The Code
We’re going to run the same model twice — once with feature scaling and once without — so the difference speaks for itself.
# demo.py
import pandas as pd
import numpy as np
from sklearn.model_selection import train_test_split
from sklearn.preprocessing import StandardScaler, OneHotEncoder
from sklearn.compose import ColumnTransformer
from sklearn.pipeline import Pipeline
from sklearn.neighbors import KNeighborsRegressor
from sklearn.metrics import mean_absolute_error
# 1. Create a synthetic dataset
np.random.seed(42)
n_samples = 200
sqft = np.random.randint(500, 3000, size=n_samples)
house_age = np.random.randint(0, 50, size=n_samples)
neighbourhood = np.random.choice(['North', 'South', 'East', 'West'], size=n_samples)
price = (
sqft * 150
+ (50 - house_age) * 800
+ pd.Series(neighbourhood).map({
'North': 20000, 'South': 10000, 'East': 5000, 'West': 0
}).values
+ np.random.normal(0, 15000, n_samples)
)
df = pd.DataFrame({
'sqft': sqft, 'house_age': house_age,
'neighbourhood': neighbourhood, 'price': price
})
# 2. Split features and target
X = df.drop(columns='price')
y = df['price']
X_train, X_test, y_train, y_test = train_test_split(X, y, test_size=0.2, random_state=42)
numeric_features = ['sqft', 'house_age']
categorical_features = ['neighbourhood']
# 3. Pipeline WITH scaling
preprocessor = ColumnTransformer(transformers=[
('num', StandardScaler(), numeric_features),
('cat', OneHotEncoder(drop='first', handle_unknown='ignore'), categorical_features),
])
pipeline_scaled = Pipeline(steps=[
('preprocessor', preprocessor),
('model', KNeighborsRegressor(n_neighbors=5)),
])
pipeline_scaled.fit(X_train, y_train)
mae_scaled = mean_absolute_error(y_test, pipeline_scaled.predict(X_test))
print(f"MAE WITH scaling: {mae_scaled:,.0f}€")
# 4. Pipeline WITHOUT scaling
preprocessor_no_scale = ColumnTransformer(transformers=[
('num', 'passthrough', numeric_features),
('cat', OneHotEncoder(drop='first', handle_unknown='ignore'), categorical_features),
])
pipeline_unscaled = Pipeline(steps=[
('preprocessor', preprocessor_no_scale),
('model', KNeighborsRegressor(n_neighbors=5)),
])
pipeline_unscaled.fit(X_train, y_train)
mae_unscaled = mean_absolute_error(y_test, pipeline_unscaled.predict(X_test))
print(f"MAE WITHOUT scaling: {mae_unscaled:,.0f}€")
print(f"nDifference: {mae_unscaled - mae_scaled:,.0f}€ less error with scaling")
Running it:
MAE WITH scaling: 21,843€
MAE WITHOUT scaling: 28,517€
Difference: 6,674€ less error with scaling
Same data. Same model. Same hyperparameters. The only difference is one line — StandardScaler() instead of 'passthrough'. That’s a 6,000€ improvement in prediction error just from rescaling the inputs.
What the code does
1. We generate 200 fake houses with a price formula that depends on all three features. The noise keeps it realistic.
2. train_test_split — 80/20 split, nothing new. We did this in article #2. of this series.
3. The ColumnTransformer applies StandardScaler to the two numeric columns and OneHotEncoder to the categorical one. We use drop='first' to avoid the dummy variable trap we talked about last time, and handle_unknown='ignore' so unseen neighbourhoods don’t crash the pipeline.
4. The second pipeline is identical but uses 'passthrough' instead of StandardScaler, so the numbers go into the model raw and unscaled.
The MAE (Mean Absolute Error) tells us how far off our predictions are on average, in €. Lower is better.
Ok, but what is StandardScaler actually doing?
Let’s say we want to compare how similar two cities are based on two things: their distance from the coast and their average temperature. Distance is in kilometres, temperature is in degrees Celsius. Fair enough.
But what if someone measures distance in inches instead of kilometres? Suddenly the distance numbers are in the millions. The temperature difference — which might actually matter more — gets completely swallowed up. You can’t meaningfully add kilometres to degrees, and you can’t meaningfully add square footage to house age. The units are incompatible.
StandardScaler fixes this by transforming each column to have a mean of 0 and a standard deviation of 1:
z = (x - mean) / std_dev
After scaling, every feature is on roughly the same numeric footing. The model can finally compare them fairly.
There’s also MinMaxScaler, which squishes values into a 0-1 range instead. Here I used StandardScaler because it handles outliers better. MinMaxScaler is useful when you specifically need bounded output.
Which models actually need this?
Not all of them. This confused me for a while, so here’s what I eventually figured out:
Needs scaling: KNN (distance-based, as we just saw), SVM, Linear/Logistic Regression, and Neural Networks.
Doesn’t need scaling: Decision Trees, Random Forest, XGBoost, LightGBM, CatBoost. All tree-based. They split on thresholds within individual features, so scale doesn’t affect them at all.
If you end up using XGBoost for tabular data — and you probably will at some point, it’s hard to beat — you can skip the scaler entirely.
One thing to be careful about
Fit the scaler on training data only.
In our pipeline, StandardScaler learns the mean and standard deviation from X_train. When we later call .predict(X_test), it applies those same training-set values to transform the test data. If you fit the scaler on the full dataset before splitting, the scaler has already “seen” the test rows. Your evaluation metrics will look better than they should. This is called data leakage and it’s one of the most common mistakes in ML pipelines.
Using Pipeline makes this hard to get wrong, which is honestly one of my favourite things about it.
A last note on the dataset: it’s very simple for the sake of being easy to understand. In the real world data sets are messier, more complex and harder to handle.
Cheers!

Be First to Comment