After the Genome: The Race to Sequence Proteins
DNA sequencing became one of the steepest cost declines in the history of technology, and a single company captured the majority of the value. Reading a protein the way we now read a gene is the harder problem; with it comes a large prize, and up-and-coming platforms are vying for it.
Written by Sarah Rodriguez, PhD and Jennifer Kan, PhD
Twenty years ago, reading a human genome was a national undertaking that cost billions and took years; today it costs a few hundred dollars and takes about a day. That cost collapse reshaped biology, built one of the most durable franchises in the life sciences, and rewarded the investors who saw where the curve was heading. We think the next chapter is beginning, and this time the molecule is protein.
Genes are instructions, but proteins are what a cell actually builds and uses, the machinery behind nearly every biological function and nearly every disease. Most approved drugs act on proteins, not genes, which is why the genome describes only what could happen while the proteome describes what is happening in a given cell on a given day. Yet although we can read a genome from end to end for a few hundred dollars, reading a protein the same way, amino acid by amino acid, remains far harder and far less mature. The first single-molecule protein sequencers have only just reached the market, but they still identify only a subset of the twenty amino acids (a few specific technologies showcased here and here) and carry error rates DNA sequencing left behind years ago. None can yet read every amino acid in sequence the way we read DNA bases. Closing that gap, from first instruments to end-to-end sequencing, is the frontier this essay is about.
Why We Need Protein Sequencing
Despite decades of progress, the tools we use to study proteins still can’t read one directly. Proteomics, the study of the proteins in a sample, is largely performed by mass spectrometry, which identifies proteins by matching them to a reference database, closer to fingerprinting than to reading, and affinity assays, which find only the proteins a researcher already knows to look for. The shortcomings of mass spectrometry and affinity assays as tools boil down to bias.
For example, in blood plasma a handful of abundant proteins, such as albumin and immunoglobulins alone make up roughly 90% of total protein mass, dominate any bulk measurement and bury everything else. The proteins that matter most for disease are often the rarest: plasma protein concentrations span about ten orders of magnitude, and more than 90% of FDA-approved protein biomarkers sit in the lowest 1% of that mass, far below the abundant proteins that swamp the signal. A tool that recognizes only what it expects will miss what is unexpected: a sequence not in any reference, a disease-specific proteoform (a modified form of a known protein, e.g., altered by chemical changes such as phosphorylation or glycosylation), or a low-abundance signaling protein no panel was built to detect. These are what are measured poorly and often not measured at all by today’s tools. Protein sequencing, which reads the actual order of amino acids and modifications on each individual molecule the way a genome is read base by base, is the tool that is required to identify the low-abundance, previously uncharacterized variant.
Why Proteins Resist Being Read
Reading a protein is harder than reading DNA, and the difficulties compound. First, the alphabet is larger and messier. Twenty amino acids against DNA’s four, several so chemically alike they are hard to tell apart, and a residue can read differently depending on its neighbors. Worse, each amino acid can be “decorated” with chemical modifications such as sugars, phosphates, acetyl groups, multiplying what a reader must distinguish.
The molecule is also physically uncooperative. It folds into a rigid three-dimensional shape, and unlike DNA’s uniformly negative backbone, a protein’s charge varies residue by residue. This makes it difficult to adopt existing DNA sequencing technologies (such as nanopore) to proteins.
Finally, protein sequencing has to work one molecule at a time. Proteins can’t be amplified like DNA, so a faint signal can’t be copied louder; everything must come from the molecules in hand. Protein sequencing requires single-molecule reading to count discrete events so that we can understand nuances, such as whether two modifications sit on the same protein.
Why Now
Protein sequencing is a hard problem, but exciting and credible companies are forming now because several things have converged at once. Single-molecule detection matured on the back of semiconductor chips, engineered nanopores, and optics borrowed from the sequencing and chip industries. Machine learning and AI can decode the noisy readouts these methods produce, so that reading a molecule can now smooth a signal enough to identify an amino acid residue. Protein engineering supplied the recognizer proteins, pores, and enzymes that many of these emerging technologies depend on. And the genome boom left behind experienced operators who have built a measurement platform before and investors trained to recognize the pattern.
The Approaches in Competition
So far, no company has yet achieved from-scratch sequencing at scale, and the contenders differ chiefly in how directly they read. Reported efforts cluster into a few families: mass spectrometry, the incumbent of bulk proteomics that infers sequence from peptide fragment masses; Edman-style chemistry read out by fluorescence, or fluorosequencing; direct nanopore reading of an intact chain; and more indirect routes that reconstruct sequence from DNA-based molecular recording or single-molecule Raman spectroscopy, each trading directness of read against scale (dive in more here and here).
The most direct approaches read residues in order. Quantum-Si sells a single-molecule sequencer that uses enzymes to shave amino acids off a peptide’s end one at a time while fluorescent binding proteins identify each newly exposed residue, though its throughput remains far below whole-proteome depth. The fluorosequencing techniques pursued by Erisyon and Prisma labels a few reactive residue types and reads the pattern left behind as the peptide is stripped, enough to match a protein to a reference database. Nanopore sequencing, a strategy pursued by Oxford Nanopore, demonstrates the most faithful to reading an intact chain and the hardest, threads a protein through a pore and reads the current it disturbs, contending with the molecule’s uneven charge and rigid folds through unfolding, molecular motors that ratchet it along, and enzymes that cleave residues at the pore’s mouth. Others read the molecule’s own physics: Pumpkinseed uses a nanophotonic chip to capture each amino acid’s Raman vibrational signature, identifying residues by their intrinsic spectra rather than by tags or fragmentation.
Against all of these stand the proxies that true sequencing means to surpass: mass spectrometry, and the affinity fingerprinting of Nautilus, Olink, and SomaLogic, which establish which known proteins are present but never read an unknown sequence from scratch.
Engineering is only half the race. The differentiator is increasingly software: machine learning and AI turn noisy single-molecule readouts into accurate calls, compensating for what hardware and wet-lab chemistry can’t perfect. Models that denoise and decode these signals push accuracy and sensitivity past what the physics alone delivers, so the winning platform may be decided as much by its data engine as by its chemistry.
What remains unsettled is which approach will be broadly commercialized first, which will ultimately reach the throughput, parallelization, sensitivity, and cost to become the standard, and whether the winning product is full sequencing, partial fingerprinting, or something in between. No obvious winners have emerged, and exciting breakthrough platforms are emerging. In our opinion, this is a great time to be paying attention.
The Size of the Prize, and Who Won Last Time
The proteomics market is estimated at 42 billion dollars (2025), growing toward roughly one hundred billion by the middle of the next decade, and one analyst has sized the next-generation opportunity at about 75 billion. The more instructive question is whether any single company won DNA sequencing. One did. In 2007, Illumina acquired Solexa’s sequencing-by-synthesis chemistry for about six hundred million dollars in stock, and by 2014 it held roughly seventy percent of the market and had won the thousand-dollar genome. While the acquisition price isn’t eye-popping, it was the platform that helped carry Illumina to a peak of roughly seventy-five billion dollars in market value by 2021, more than a hundredfold above what it paid. It won because its chemistry was best on the metric that came to define the market with cost per accurate base at scale, and locked that lead in with a razor-and-blade model whose switching costs became a moat.
One Technology, Many Frontiers
The stakes of protein sequencing reach well beyond life sciences tools and medicine. In agriculture, it can reveal how crops respond to drought or disease and help breed hardier varieties. In food and nutrition, it can verify what’s actually in a product, spotting allergens, and confirming a product is what the label claims. In environmental monitoring, it can identify what living organisms are doing to our soil and water, offering a readout of ecosystem health that DNA alone can’t give. It matters for security and forensics, too, where a protein signature can flag a biological threat or place a person or species at a scene.
We are looking closely at the latest advances in protein sequencing, from detection chemistry and device physics to the software that turns a noisy signal into an answer. If you are working on reading proteins, we want to hear from you!


