Research area
Statistical Learning & Biological Language Processing
This area grew out of the lab’s earlier work in deep genomics, deep proteomics, and data-driven biological language processing. It now serves as the umbrella for statistical learning methods that connect sequence data, microbiome data, and mechanistic scientific questions.
Biological sequences as language
The legacy premise of this program was that biological sequences can be treated as languages: structured strings generated by nontrivial rules and constraints. From that perspective, machine learning can be used to learn useful subsequence patterns, contextual representations, and predictive models directly from protein, genome, and microbial data.
The lab approached this through both subsequence-based models and language-model-style embeddings, asking how sequence statistics and learned representations could improve structural and functional annotation.
Protein and genome informatics
One major thread focused on protein informatics and representation learning, including learned embeddings for protein sequences, motif discovery, and structure-function annotation from large sequence corpora. This work treated proteins as learnable sequence objects and used statistical representations to bridge raw sequence data with biological interpretation.
Microbial informatics
Another thread focused on microbial communities and 16S rRNA data, including phenotype prediction, biomarker discovery, and reference-light analysis of microbiome sequence data. These projects connected biological language-processing methods directly to host phenotype questions and microbiome interpretation.
Data-driven discovery in biology and medicine
The lab also explored adjacent machine-learning applications in clinical and biomedical settings, especially where data-rich measurements could be connected to mechanistic questions or practical scientific inference. In the current site taxonomy, those efforts are preserved as part of a broader program in statistical learning for biological discovery.
Current framing
Today this area emphasizes statistical learning, biological language processing, and model-guided discovery across proteins, genomes, microbiomes, and related data modalities. Representative outputs are collected in the publications archive.