Chemicals & industrial biotechnology
Strain engineering runs on a loop — design, build, test, learn. The loop only moves as fast as the data coming back from it.
We have built research data platforms inside some of the world's largest chemical companies, including the infrastructure behind strain optimization for biotechnological production.
// the loop only runs as fast as the data coming back
One strain. Four systems. No single record.
A strain has a genotype — constructs, edits, a lineage going back several rounds. It has a set of fermentation runs at different scales. Those runs have process traces: feed, pH, dissolved oxygen, temperature, over time. And they have outcomes: titer, rate, yield. Each of those four things lives in a different system, and none of them is wrong.
So the question that matters — which genetic change actually improved yield, and does it still hold at 200 litres — becomes an assembly job. Someone joins four sources by hand, matching strain identifiers that are written differently in each one. It takes days, so it gets done for a subset, and the comparison is made across the runs somebody remembered.
What we build: one record per strain — genotype and lineage, every run it appeared in, the conditions of those runs, and the measured outcomes. A genotype-to-phenotype question becomes a query instead of a project.
What it is worth: more of each round's data informs the next one; fewer runs repeated because the earlier result was findable only in principle; and scale-up decisions made on all the evidence rather than the part that was easy to reach.
Problems we see
- Strain lineage and construct history in one system, fermentation results in another, process traces in a third — and no single record for the strain itself.
- The same strain under three different names, so comparing across rounds is manual and partial.
- Bioreactor time-series exported to CSV, analyzed once for one question, and never read again.
- Genotype-to-phenotype questions that take weeks to assemble and get answered on a subset.
- Negative results repeated, because the earlier run was findable in principle and not in practice.
- Expertise concentrated in a handful of long-tenured people, all of whom have a retirement date.
One record per strain, and a way to ask it questions
One record per strain, assembled from the systems that each hold a piece of it — connected data foundations — with search and assistants that understand strain identifiers, construct names and run vocabulary: answers on your data.
The modeling is the easy half. Deciding what a strain, a run or a batch actually is across four systems takes molecular biologists and process engineers in the room, not a data team working from a spec.
Your strains do not leave your infrastructure
Three constraints that are real in this industry, and that shape the architecture rather than the paperwork.
IP first
Strain and process data is the asset. Deployment inside your own environment is a requirement here rather than a preference, and it rules out sending anything to a third-party model.
REACH and product stewardship
On the chemicals side, substance and composition data decides whether a product can legally be sold in a market. That makes it a traceability problem rather than a reporting one.
CSRD-driven footprint data
Product carbon footprint, life-cycle and scope 3 figures now have to be auditable — traceable to a source, rather than assembled in a spreadsheet each quarter.
The same problem, one industry over
The pattern is not specific to strains. Formulation and experiment history spread across an ELN, a LIMS, reports and legacy databases. Substance and composition data reassembled by hand every time a market asks whether a product is still sellable. Plant parameters that exist in three places and disagree.
Same reassembly problem, different entity. Our depth is in bio-based production, and the approach transfers — which is a different claim from having done it everywhere.
How much of last round's data made it into this round?
Not rhetorical — we would like to know. Tell us what your teams can and cannot get to.
Get in touch