A DUAL-PATH CONTRASTIVE CNN-TRANSFORMER FRAMEWORK FOR ROBUST AND DISCRIMINATIVE MELANOMA CLASSIFICATION FROM DERMOSCOPIC IMAGES

ICTACT Journal on Image and Video Processing ( Volume: 17 , Issue: 1 )

Abstract

Melanoma is an aggressive form of skin cancer in which early and accurate diagnosis substantially influences clinical management and patient survival. Dermoscopic imaging provides detailed visual information about lesion structures, pigmentation, borders, and internal patterns, making it an important source for computer-aided melanoma assessment. Recent deep learning methods have demonstrated strong diagnostic capability; however, their performance remains sensitive to image variability, class imbalance, subtle inter-class similarities, and limited representation of global lesion context. Conventional convolutional neural networks (CNNs) effectively capture local texture and morphological patterns but may inadequately model long-range relationships, whereas Transformer-based models capture global dependencies but can be sensitive to data availability and computational requirements. Existing hybrid models also frequently rely on simple feature concatenation without explicitly encouraging complementary and discriminative representations. This study proposes a novel Dual-Path Contrastive CNN-Transformer Network (DCCT-Net) for melanoma classification. The framework employs parallel CNN and Transformer pathways to learn complementary local and global lesion representations. A contrastive representation alignment module is introduced to reduce intra-class feature variation and increase inter-class separability, while an adaptive feature fusion mechanism combines the two pathways before classification. Augmentation-aware contrastive learning further improves representation robustness against changes in scale, illumination, rotation, and lesion appearance. The proposed DCCT-Net demonstrates superior performance compared with the DL-DNN, Vision Transformer, and Hybrid U-Net-Inception-ResNet-ViT approaches. At the final training epoch, DCCT-Net achieves 98.4% accuracy, 98.1% precision, 97.9% recall, 98.0% F1-score, and 99.0% specificity. Compared with the three existing methods, the proposed model improves accuracy by 2.2, 2.3, and 1.4%, respectively. The recall reaches 97.9%, representing improvements of 3.1, 3.4, and 1.9%, demonstrating stronger melanoma detection capability. The F1-score of 98.0% confirms a balanced improvement in precision and recall, while the specificity of 99.0% indicates effective discrimination of non-melanoma lesions. These results demonstrate that combining complementary CNN and Transformer representations with contrastive learning and adaptive fusion improves the discriminative capability of automated melanoma classification.

Authors

Sreedevi Kadiyala1, Chandra Srinivas Potluri2
Guru Nanak Institutions Technical Campus, India1, Siddhartha Institute of Engineering and Technology, India2

Keywords

Melanoma Classification, Dermoscopic Images, Contrastive Learning, CNN-Transformer, Skin Lesion Analysis

Published By
ICTACT
Published In
ICTACT Journal on Image and Video Processing
( Volume: 17 , Issue: 1 )
Date of Publication
August 2026
Pages
3998 - 4007
Page Views
35
Full Text Views
5