Methodology
Exact preprocessing, split, and training configuration used for both baselines.
Data split
70% train / 15% validation / 15% test, seed 42 — 18,900 / 4,050 / 4,050 images. The same split is reused for both the RGB and multispectral datasets.
RGB model
- Pretrained ImageNet ResNet50, backbone frozen, final FC layer replaced (10-class head)
- Resize 224×224 → tensor → ImageNet normalization (mean [0.485, 0.456, 0.406], std [0.229, 0.224, 0.225])
- CrossEntropyLoss, Adam, lr = 0.001, batch size 32, 10 epochs
- 23,528,522 total parameters / 20,490 trainable
Multispectral model
- ResNet50 first convolution expanded to 13 channels
- Channels 0–2 initialized from pretrained RGB conv weights; channels 3–12 from their average
- Input scaled by /10000; same split, same training hyperparameters as the RGB model
- Only the final classification layer (plus the necessarily-retrained first conv) trained
Evaluation
Accuracy, weighted precision/recall/F1, confusion matrix, per-class performance, training/ validation loss and accuracy curves, and misclassified-example inspection — run identically for both models. Grad-CAM is applied to both correct predictions and interesting misclassifications for interpretation, not as proof of causation.
Full training and evaluation code: research/experiments/ in the repository.