An expert in computational genomics and the creator of STAR and STARsolo — the widely used RNA-seq aligner and its single-cell counterpart — who now builds algorithms and infrastructure for Arc Institute's Virtual Cell and Alzheimer's Disease Initiatives in collaboration with the research groups at NVIDIA and Roche, following sixteen years at Cold Spring Harbor Laboratory.
As Director of Bioinformatics at Arc Institute, Dobin built and directed a high-performing bioinformatics team from the ground up, setting the computational strategy for Arc's Virtual Cell and Alzheimer's Disease Initiatives by bridging computational and experimental biology.
His team develops production-scale infrastructure — quality-control frameworks and fully automated cloud pipelines — that delivered a 30-fold speedup and enabled the generation of 100 million CRISPRi-perturbed cells in three months while cutting compute costs by $20,000 a month. He co-led scBaseCount, the world's largest freely accessible single-cell repository (502M+ cells across 27 organisms), co-organized the inaugural Virtual Cell Challenge — 5,000+ participants from 114 countries — and built the evaluation frameworks used to benchmark Arc's STATE virtual-cell model.
This work extends a career-long focus that took shape over sixteen years at Cold Spring Harbor Laboratory, where he rose from computational science developer to Assistant Professor and Head of the Bioinformatics Core. There he created and maintained STAR under a five-year NIH R01 grant, led the ENCODE consortium's RNA-seq pipelines and novel statistical methods for validating splice junctions, transcripts, and fusions, and — in collaboration with David Tuveson's lab — applied machine learning to pancreatic-cancer genomics, inferring gene-regulatory networks across tumors and their microenvironment.
Dobin's career reads differently in hindsight than it must have felt at the time. He earned a Ph.D. in condensed matter physics, then spent years writing simulations of the magnetic materials used in computer hard drives. From there he moved into biology, teaching himself bioinformatics without ever taking a formal course in it — and shortly after making that switch, built STAR and STARsolo, now two of the most widely used tools in RNA-seq analysis.
He credits the physics background directly, saying it taught him "how to look for patterns and logic in the data" — and how to tell the signal from the noise. That same computational instinct now carries over into modeling cellular dynamics at Arc.
Built to align the ENCODE project's transcriptome RNA-seq data — more than 80 billion reads — STAR uses a sequential maximum-mappable-seed search over uncompressed suffix arrays to map spliced reads with high accuracy and speed. Since its 2012 publication in Bioinformatics it has been cited in more than 50,000 papers and become one of the standard aligners in RNA-seq pipelines worldwide, later extended with STARsolo — whose algorithms improved single-cell quantification accuracy while boosting processing speed roughly 30-fold. Development at Cold Spring Harbor was supported by a five-year NIH R01 grant.