1. Evolution from CNNs to Vision Transformers
While classical Convolutional Neural Networks (ResNet, EfficientNet) process local pixel neighborhoods through convolutional kernels, modern Vision Transformers (ViTs) divide dermoscopic photographs into non-overlapping patches and use self-attention mechanisms to evaluate global structural relationships across the entire lesion surface.
2. Training Pipelines & Feature Extraction
Models are trained on hundreds of thousands of biopsy-verified clinical and dermatoscopic images (ISIC, HAM10000). Features analyzed include pigment network regularity, border termination abruptness, vascular morphology (dotted, linear, comma-shaped vessels), and multi-spectral color variegation.