
CS&STA & IACR AI/ML Seminar: Deep statistical and contrastive modelling of nanopore sequencing translocation times reveals latent non-B DNA structures
Speaker: Derek Aguiar (Associate Professor in the School of Computing at the University of Connecticut)
Date/Time/Location: Nov 13, 1pm, Tyler Hall RM. 106.
Title: Deep statistical and contrastive modelling of nanopore sequencing translocation times reveals latent non-B DNA structures
Abstract: Non-canonical (or non-B) DNA are genomic regions whose three-dimensional conformation deviates from the canonical double helix. Non-B DNA plays an important role in cellular processes and is associated with genomic instability, gene regulation, and oncogenesis. Experimental methods are low-throughput and detect only a limited set of non-B structures, while computational methods typically rely on sequence motifs that indicate the potential to form non-B DNA but do not establish that a structure is present. I will present our group's work on predicting non-B DNA structure from Oxford Nanopore sequencing by modeling the time required for DNA to translocate through a nanopore. We first formulated non-B structure prediction as a novelty-detection problem and developed a deep statistical model combining representation learning with goodness-of-fit testing. In subsequent work, we addressed sensitivity to noisy motif-based labels using noise-resilient contrastive learning. Both approaches use large-scale hypothesis testing with false discovery rate control to identify significant deviations from canonical B-DNA. Across simulated and whole-genome nanopore sequencing data, our results demonstrate that translocation-time signals contain information about DNA conformation, with predicted non-B structures exhibiting distinct SNP-frequency patterns and predicted G-quadruplexes showing greater G4-forming potential. Together, these studies illustrate how deep representation learning and statistical hypothesis testing can be combined to recover latent biological structure from noisy and imperfectly labeled sequencing data.