Methodology

Exact preprocessing, split, and training configuration used for both baselines.

Data split

70% train / 15% validation / 15% test, seed 42 — 18,900 / 4,050 / 4,050 images. The same split is reused for both the RGB and multispectral datasets.

RGB model

  • Pretrained ImageNet ResNet50, backbone frozen, final FC layer replaced (10-class head)
  • Resize 224×224 → tensor → ImageNet normalization (mean [0.485, 0.456, 0.406], std [0.229, 0.224, 0.225])
  • CrossEntropyLoss, Adam, lr = 0.001, batch size 32, 10 epochs
  • 23,528,522 total parameters / 20,490 trainable

Multispectral model

  • ResNet50 first convolution expanded to 13 channels
  • Channels 0–2 initialized from pretrained RGB conv weights; channels 3–12 from their average
  • Input scaled by /10000; same split, same training hyperparameters as the RGB model
  • Only the final classification layer (plus the necessarily-retrained first conv) trained

Evaluation

Accuracy, weighted precision/recall/F1, confusion matrix, per-class performance, training/ validation loss and accuracy curves, and misclassified-example inspection — run identically for both models. Grad-CAM is applied to both correct predictions and interesting misclassifications for interpretation, not as proof of causation.

Full training and evaluation code: research/experiments/ in the repository.