Learning & Workshops

CS&STA & IACR AI/ML Seminar: Deep statistical and contrastive modelling of nanopore sequencing translocation times reveals latent non-B DNA structures

Speaker: Derek Aguiar (Associate Professor in the School of Computing at the University of Connecticut)

Date/Time/Location: Nov 13, 1pm, Tyler Hall RM. 106.

Title: Deep statistical and contrastive modelling of nanopore sequencing translocation times reveals latent non-B DNA structures

Abstract: Non-canonical (or non-B) DNA are genomic regions whose three-dimensional conformation deviates from the canonical double helix. Non-B DNA plays an important role in cellular processes and is associated with genomic instability, gene regulation, and oncogenesis. Experimental methods are low-throughput and detect only a limited set of non-B structures, while computational methods typically rely on sequence motifs that indicate the potential to form non-B DNA but do not establish that a structure is present. I will present our group's work on predicting non-B DNA structure from Oxford Nanopore sequencing by modeling the time required for DNA to translocate through a nanopore. We first formulated non-B structure prediction as a novelty-detection problem and developed a deep statistical model combining representation learning with goodness-of-fit testing. In subsequent work, we addressed sensitivity to noisy motif-based labels using noise-resilient contrastive learning. Both approaches use large-scale hypothesis testing with false discovery rate control to identify significant deviations from canonical B-DNA. Across simulated and whole-genome nanopore sequencing data, our results demonstrate that translocation-time signals contain information about DNA conformation, with predicted non-B structures exhibiting distinct SNP-frequency patterns and predicted G-quadruplexes showing greater G4-forming potential. Together, these studies illustrate how deep representation learning and statistical hypothesis testing can be combined to recover latent biological structure from noisy and imperfectly labeled sequencing data.

Sign in to report a problem