Research area

Statistical Learning & Biological Language Processing

This area grew out of the lab’s earlier work in deep genomics, deep proteomics, and data-driven biological language processing. It now serves as the umbrella for statistical learning methods that connect sequence data, microbiome data, and mechanistic scientific questions.

Biological sequences as language

The legacy premise of this program was that biological sequences can be treated as languages: structured strings generated by nontrivial rules and constraints. From that perspective, machine learning can be used to learn useful subsequence patterns, contextual representations, and predictive models directly from protein, genome, and microbial data.

The lab approached this through both subsequence-based models and language-model-style embeddings, asking how sequence statistics and learned representations could improve structural and functional annotation.

Protein and genome informatics

One major thread focused on protein informatics and representation learning, including learned embeddings for protein sequences, motif discovery, and structure-function annotation from large sequence corpora. This work treated proteins as learnable sequence objects and used statistical representations to bridge raw sequence data with biological interpretation.

Microbial informatics

Another thread focused on microbial communities and 16S rRNA data, including phenotype prediction, biomarker discovery, and reference-light analysis of microbiome sequence data. These projects connected biological language-processing methods directly to host phenotype questions and microbiome interpretation.

Data-driven discovery in biology and medicine

The lab also explored adjacent machine-learning applications in clinical and biomedical settings, especially where data-rich measurements could be connected to mechanistic questions or practical scientific inference. In the current site taxonomy, those efforts are preserved as part of a broader program in statistical learning for biological discovery.

Current framing

Today this area emphasizes statistical learning, biological language processing, and model-guided discovery across proteins, genomes, microbiomes, and related data modalities. Representative outputs are collected in the publications archive.