Artificial Intelligence and Machine Learning–Driven Interpretation of Non-Coding and Structural Variants in Cancer Genomics
Hasti Hosseini,1,*
1. Tehran University of Medical Sciences
Introduction: Over 98% of the human genome is noncoding, yet most cancer genetics has focused on protein-coding mutations. Noncoding variants in promoters, enhancers, silencers, and other regulatory elements, together with structural variants (SVs: deletions, duplications, inversions, translocations), are increasingly recognized as drivers of oncogenesis, risk modifiers, and determinants of therapy response. However, interpreting these variants remains a major bottleneck: they are abundant, context-dependent, and often lack clear functional annotation. Artificial intelligence (AI) and deep learning (DL) have emerged as key tools to decode this “dark genome” by learning complex sequence and regulatory patterns beyond human-designed rules.
Methods: This article is a focused review and conceptual synthesis of AI-driven approaches for interpreting noncoding and structural variants in cancer. We systematically examined recent reviews and primary studies (2015–2026) on: (i) deep learning models for noncoding variant effect prediction (e.g., DeepSEA, Enformer, GET, promoter- and splice-focused models); (ii) machine learning frameworks for SV pathogenicity scoring and filtering (e.g., SVFX, attention-based models such as PhenoSV, ML filters for long-read SV calling); and (iii) integrative pipelines that combine sequence, epigenomics, 3D genome, and expression data to link variants to genes and pathways. We compared models according to architecture (CNNs, transformers, graph neural networks), data requirements (sequence-only vs. multi-omics), cell type specificity, interpretability, and demonstrated cancer applications. We then distilled common design principles into a conceptual framework for a cancer-genetics–oriented AI pipeline: from whole-genome or whole-exome data, through variant calling, functional scoring of noncoding mutations and SVs, integration with regulatory annotations and expression, to prioritization of candidate drivers and clinically relevant alterations.
Results: The literature shows a rapid expansion of AI methods tailored to noncoding and structural variation in cancer. Systematic reviews of 78 deep learning studies (2015–2024) report that convolutional and graph-based architectures achieve state-of-the-art performance in variant calling and tumor stratification, reducing false-negative rates by 30%–40% compared with traditional pipelines and prioritizing pathogenic variants with up to 92% accuracy. For noncoding variants, sequence-to-function models (e.g., DeepSEA, Enformer, GET) predict chromatin states, transcription factor binding, and gene expression from DNA sequence, enabling quantitative estimates of how a single-nucleotide change alters regulatory activity. Newer transformer-based and tissue-specific transfer-learning models improve prediction in relevant cell types and can generalize to unseen contexts, addressing a key limitation of earlier tools. For SVs, machine learning frameworks such as SVFX and attention-based models integrate genomic context, regulatory overlap, conservation, and copy-number impact to assign pathogenicity scores, helping distinguish driver rearrangements (e.g., enhancer hijacking, gene fusions) from passengers. Multi-omics integrative models further link noncoding and structural variants to expression programs, chromatin accessibility, and 3D contacts, revealing mechanisms such as promoter activation, enhancer rewiring, and silencer disruption in specific cancer types. Collectively, these approaches enable: (i) prioritization of functional noncoding mutations (e.g., in TERT promoter, MYC enhancers, super-enhancers); (ii) systematic scoring of SVs affecting coding and non-coding regions; and (iii) generation of testable hypotheses about variant-to-gene-to-pathway relationships that can guide functional validation and clinical interpretation.
Conclusion: AI has transformed the interpretation of noncoding and structural variants from a largely manual, annotation-driven task into a data-driven, predictive science. Deep learning models now provide quantitative, context-aware estimates of regulatory impact and pathogenicity, substantially improving the resolution of cancer genome analysis beyond coding mutations. For cancer genetics, this means more accurate identification of candidate drivers, better understanding of germline risk variants, and richer mechanistic insights into how noncoding and structural alterations rewire gene regulation. Future work should focus on improving interpretability, expanding diverse training data, and integrating these AI scores into routine cancer genomics pipelines alongside ctDNA and single-cell/spatial assays. For a 5–12 minute oral presentation and poster, this topic offers a clear narrative: (1) the problem of the “dark genome” in cancer; (2) how AI models work at a high level; (3) concrete examples of discoveries and applications; and (4) a forward-looking view of AI as a core component of precision cancer genetics.
Keywords: AI; Machine learning; Non-coding variants; Structural variants; Cancer genomics
به خانواده بزرگ کنسر ژنتیکس و ژنومیکس سرطان بپیوندید!