Experimental findings overview
RGB leads in the current experimental setup.
In this specific baseline configuration, RGB transfer learning achieved 94.96% test accuracy versus 88.12% for the 13-band multispectral model — a gap of 6.84 percentage points.
94.96%
RGB test accuracy
88.12%
Multispectral test accuracy
+6.84%
RGB configuration gap
Scientific interpretation note
This gap reflects the current baseline configuration — only the final classification layer was fine-tuned, and the multispectral model's 13-band input layer was adapted without extensive pretraining. It does not establish a universal superiority of RGB over multispectral satellite data.
Training run on a Colab T4 GPU, September 17, 2026. Split 70/15/15, seed 42. Numbers may shift slightly between runs; an earlier September 16 run is kept in the research log for comparison.
Overall test performance
| Model | Test accuracy | Weighted precision | Weighted recall | Weighted F1 | Best val accuracy |
|---|---|---|---|---|---|
| RGB ResNet50 (3 bands) | 94.96% | 94.98% | 94.96% | 94.95% | 94.27% |
| Multispectral ResNet50 (13 bands) | 88.12% | 88.11% | 88.12% | 87.87% | 88.00% |
Training curves

RGB ResNet50 — loss and accuracy over 10 epochs.
Confusion matrices

RGB ResNet50 test-set confusion matrix.

Multispectral ResNet50 test-set confusion matrix.
Per-class F1
Measured on this run's test set. Harder classes here reflect this data and setup, not fixed properties of those land-cover types.
Error analysis (Grad-CAM)
Grad-CAM on the multispectral model surfaced an overconfident mistake: a Residential patch predicted as SeaLake at roughly 95% confidence. It is kept as an error-analysis example, not a success case. Grad-CAM shows where the model focused; it does not prove why it was wrong.

Grad-CAM on a correctly classified RGB example.

Grad-CAM on the Residential → SeaLake error case.
Coming next
Live inference in the Dataset Library and Upload & Predict needs the trained RGB checkpoint wired up behind the FastAPI inference backend.