BioConsortia News

Harnessing Genomics and Machine Learning to Unlock Microbial Potential

Written by BioConsortia Scientists | Jul 23, 2026

Part 2 of Decoding Microbial Performance, an ongoing series where BioConsortia's scientists explain, in their own words, how we discover and de-risk the microbes behind our products.

Written by Cris Vera, Scientist III, Bioinformatics, and Don Gibson, Scientist III, Microbiology

Sequencing a microbial genome, once a significant undertaking, is now a standard laboratory procedure. Determining which genes within it are functionally relevant, particularly those lacking clear homology to previously characterized proteins, is the harder problem. Machine learning is beginning to close that gap.

At BioConsortia, our scientists have developed proprietary machine learning algorithms and models that accelerate our findings in genomics and plant phenotyping. These pipelines are embedded in the development cycle for both wild type and gene edited microbial biologics for nitrogen fixation, biostimulant, and biocontrol applications. The use of modern bioinformatic tools enhances our selection pipeline, introduced in 'From Soil to Seed: How Microbial Discovery Is Evolving', to speed up product development and deepen our understanding of our microbes, ensuring our best-performing strains are selected for field trials all over the world.

Genomics and Machine Learning

In agricultural biotech, one of the most important questions we face is deceptively simple: what makes high-performing microbes perform so well? We address this question through systematic diversity screening of microbial populations. By leveraging our large and continually expanding collection of microbial strains derived from our patented AMS™ process, we can explore performance through both traditional microbiology and modern computational approaches. Today, genomics sits at the center of that exploration.

Next-generation sequencing has transformed how we gather insights into microbial potential. Long-read sequencing technologies are pushing the field forward. Machine learning derived models have greatly improved base-calling accuracy and even detect epigenetic modifications directly from sequencing data. Advances in base calling allow us to move beyond static DNA sequences and begin understanding regulatory layers that control plant microbe interactions.

Beyond Standard Annotation

Raw sequence data is only the beginning. Genomic mining tools, many powered by machine learning, enable us to annotate genes and gene clusters at scale and with accuracy that would have been unprecedented a decade ago. Through these tools, we can identify key functional traits such as secondary metabolite production, nitrogen fixation potential, and the presence or absence of toxins or phage-related sequences. This functional annotation is critical for linking genotypes to phenotypes.

We analyze phylogenetic relationships across multiple strains, identifying patterns that distinguish high performers from low performers. These patterns help generate hypotheses about the mode of action and guide future wild type and gene edited strain development.

Thanks to machine learning, we are also beginning to overcome one of genomics’ long-standing limitations: imperfect annotation. Advances in models of conserved protein domains allow us to move beyond simple amino acid similarity when considering homologs in related strains. Machine learning approaches for unsupervised clustering of proteins allow us to group gene products based on amino acid chemistry and structural features, rather than relying solely on existing labels. We can now discover functional relationships that traditional annotation pipelines miss.

The convergence of genomics and machine learning is accelerating in agricultural biotech. As these tools continue to evolve, so too does our ability to ask deeper questions, move faster from discovery to application, and build more effective products for growers.

How it works

At BioConsortia, we leverage vast quantities of screening data with machine learning algorithms and have improved our translation of greenhouse leads to field trial success. Historically, using lab and greenhouse data to choose the right leads for field trials has been risky. The greenhouse is not the field, and traits found in the greenhouse are often at different stages of development than those in a field trial that comes to yield. Our scientists, with over a decade of experience screening biostimulants and nitrogen-fixing microbes have unlocked new plant phenotypes to study.

Figure 1:  Corn under greenhouse screening at BioConsortia, where leaf nutrient sampling, hand-based phenotyping, and plant imaging generate the data behind our machine learning models. 

 

Plant-associated nitrogen-fixing microbes can provide crops with exogenous nitrogen, reducing the need for synthetic fertilizer and increasing yields. These contributions can be subtle and not easy to detect. Our greenhouse data collection includes sampling leaf nutrients, hand-based phenotyping, and advanced plant imaging techniques. With millions of measurements and data points, we use machine learning to bring new insights into which phenotypes are best correlated with field success.

Figure 2:  BioConsortia's nitrogen prediction model, which forecasts nitrogen content in corn leaf tissue, using machine learning applied to greenhouse and field data. 

 

Why it matters

Using this machine learning approach, we developed a nitrogen prediction model that accurately forecasts nitrogen content in corn leaf tissue, with r2 > 0.9. In addition to improving our plant nitrogen content estimate accuracy, our machine learning approach uncovered a correlation not previously described in peer reviewed research articles with a phenotype we were collecting related to nitrogen content. We reviewed our historical data in the greenhouse and field and found that this new phenotype had the highest correlation between greenhouse success and field trial success to date.

Conclusion

Machine learning can help address pressing challenges in agricultural biotechnology. When applied to a large library of genomic sequences from soil bacteria, it can identify protein sequence domains associated with desired bioactivity, aiding in the selection of lead strains for greenhouse experiments. Large datasets of greenhouse and field trial results helped identify the phenotypes in the greenhouse that are most predictive of field trial success. BioConsortia has successfully leveraged these applications of machine learning to advance our most promising leads from our collection to the greenhouse, and from the greenhouse to the field, all the while growing our datasets to prepare us to apply the next generations of machine learning technologies to the next generation of agricultural solutions.

 

 

ABOUT BIOCONSORTIA

BioConsortia, Inc. develops transformative microbial products in agriculture, pioneering the use of directed selection and advanced genomics techniques within microbial communities. BioConsortia’s microbial products deliver superior efficacy, higher consistency, and easier grower adoption because of their focus on microbes that deliver extended shelf- and on-seed life. The company’s rich biological pipeline includes nitrogen fixation microbes to replace or reduce synthetic nitrogen fertilizers; nutrient use efficiency and biostimulants to increase crop yields; and bionematicides and biofungicides to protect crops from pests and diseases. BioConsortia’s patented Advanced Microbial Selection (AMS) process and cutting-edge GenePro platform enable the company to predict, design, and unleash the natural power of microbes.

For further information, please contact info@bioconsortia.com.