šŸ¤–QGISAdvancedā±ļø 4 mins read

Supervised vs Unsupervised Classification

Published by GISTECHNEWS Editorial Team • Peer-Reviewed & Verified on QGIS 3.34+ LTR & Python 3.10+

Image classification algorithms in remote sensing group continuous multi-spectral pixel arrays into meaningful land-use categories. These algorithms fall into two fundamentally contrasting paradigms: Supervised Classification and Unsupervised Classification. Understanding the theoretical, statistical, and operational differences between these two methodologies is essential for designing defensible environmental monitoring programs and choosing the right tool for your specific satellite imagery dataset.

šŸ“‹ Prerequisites

  • Basic understanding of multispectral band reflectance.
  • QGIS installed with Processing Toolbox enabled.
  • Knowledge of statistical clustering and machine learning principles.

šŸ› ļø Technical Environment

Required Software: QGIS / SAGA GIS / scikit-learn (Recommended: 3.34+ LTR / Python 3.10+)

Practice Dataset: Landsat 8 OLI 30m Multispectral Bands

Source Portal: USGS EarthExplorer

CRS / Format: UTM (GeoTIFF)

Step-by-Step Workflow & Methodological Execution

1

Module 1: Unsupervised Classification Mechanics (K-Means & ISODATA)

In Unsupervised Classification, the computer clusters pixels based purely on natural spectral groupings in multi-dimensional feature space without any prior training labels or human intervention: • K-Means Algorithm: The analyst specifies an arbitrary number of clusters $K$ (e.g., 10). The algorithm randomly assigns $K$ cluster centers, allocates each pixel to its nearest centroid in spectral Euclidean distance, recalculates the mean centroid, and iteratively repeats until cluster assignments stabilize. • ISODATA (Iterative Self-Organizing Data Analysis Technique): A dynamic upgrade over K-Means that automatically splits clusters with high standard deviations and merges clusters whose centroids fall below a minimum distance threshold. • The Post-Classification Task: Once clusters are formed (Cluster 1, 2, 3...), the GIS analyst must inspect the imagery and assign real-world meanings (e.g., Cluster 1 = Deep Water, Cluster 2 = Turbid Water, Cluster 3 = Coniferous Forest).

2

Module 2: Supervised Classification Mechanics (Parametric vs Machine Learning)

In Supervised Classification, the analyst controls the process by providing known training samples for predefined classes, and the computer establishes statistical boundaries based on those samples: • Maximum Likelihood Classifier (MLC): A traditional parametric classifier that assumes training samples follow a multivariate normal Gaussian distribution. It calculates the statistical probability that a given pixel belongs to each class and assigns it to the most probable class. • Random Forest (RF): A non-parametric ensemble learning technique that constructs hundreds of de-correlated decision trees and outputs the majority mode class. Handles multi-modal distributions, non-linear relationships, and ancillary elevation/slope data without distribution assumptions. • Support Vector Machines (SVM): Finds optimal separating hyperplanes that maximize the margin between classes in high-dimensional feature space, performing exceptionally well even with small training sample sets.

3

Module 3: Accuracy Assessment & Confusion Matrix Evaluation

A classification without an accuracy assessment is scientifically invalid. To validate results, generate an independent validation sample set (at least 50 random stratified points per class) and build an Error Matrix (Confusion Matrix): • Overall Accuracy (OA): Total correctly classified validation pixels divided by total sample count. • Producer's Accuracy (100% - Omission Error): Percentage of actual ground features that were correctly categorized by the classifier. • User's Accuracy (100% - Commission Error): Probability that a pixel classified on the map truly represents that category on the ground (reliability from a map reader's perspective). • Cohen's Kappa Coefficient ($\hat{K}$): Measures accuracy agreement above chance expectation. Values > 0.80 indicate strong, reliable classification performance.

āš ļø Common Errors & Troubleshooting

āŒ K-Means clusters merge urban and bare soil

šŸ’” Resolution: Increase cluster count parameter (k) from 5 to 12 during unsupervised clustering, then manually merge clusters post-classification.

āŒ Overall Accuracy is high but Kappa Coefficient is low

šŸ’” Resolution: This indicates severe class imbalance (e.g. 90% of the scene is water). Rely on F1-Score and Producer's Accuracy instead.

šŸ’” Expert Tips & Best Practices

  • Unsupervised classification (K-Means/ISODATA) is ideal when no ground-truth training data exists for an unknown geographic area.
  • Use Random Forest over Maximum Likelihood for complex multimodal spectral distributions.

šŸ Python Scikit-Learn K-Means Unsupervised Spectral Clustering

import rasterio
import numpy as np
from sklearn.cluster import KMeans

# Open 4-band optical satellite image
with rasterio.open("satellite_4band.tif") as src:
    img = src.read() # Shape: (4, rows, cols)
    profile = src.profile

bands, rows, cols = img.shape
# Flatten image array into 2D feature matrix (N_pixels, N_bands)
X = img.reshape(bands, rows * cols).T

# Initialize K-Means clustering into 6 natural spectral clusters
print("Executing unsupervised K-Means clustering...")
kmeans = KMeans(n_clusters=6, random_state=42, n_init=10)
cluster_labels = kmeans.fit_predict(X)

# Reshape back to 2D image dimensions
clustered_map = cluster_labels.reshape(rows, cols)

# Save unsupervised cluster raster
profile.update(dtype=rasterio.uint8, count=1, nodata=255)
with rasterio.open("unsupervised_clusters.tif", "w", **profile) as dst:
    dst.write(clustered_map.astype(rasterio.uint8), 1)

print("K-Means clustering complete. Classes 0-5 ready for analyst labeling.")

šŸ”— Related Tutorials & Practical Workflows