Fungal Genomics Machine Learning for Novel Therapeutic Discovery

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current methods for discovering new therapeutics from fungi rely on bioactivity-guided approaches that often rediscover known compounds, lacking integration of big data analytics, 'omics' biology, and artificial intelligence, and are not scalable for unstudied fungal species.

Innovation Solution

A platform combining genomics, metabolomics, and machine learning to analyze fungal genomic and metabolomic data, identifying relationships between biosynthetic gene clusters and mass spectrometric features of metabolites without DNA manipulations, using custom Hidden Markov Models and random forest classifiers for precision and scalability.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If bioactivity-guided approaches are used for compound discovery, then known compounds can be identified, but new compounds are rarely discovered and the approach is not scalable

Engineering Contradiction:
Improvecompound identification reliabilityVSAvoidnew compound discovery rate
Core Design Contradiction:
ReliabilityVSProductivity

Solution Approach 1:

The patent replaces traditional bioactivity-guided fractionation (mechanical/physical separation methods) with a computational approach using machine learning models, random forest classifiers, and genomic-metabolomic data integration. This substitution enables systematic prediction of biosynthetic gene cluster functions and metabolite structures, dramatically increasing new compound discovery while maintaining reliability through validated predictive models

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Solution Approach 2:

The patent changes the fundamental parameters of compound discovery by shifting from phenotype-based screening to genotype-prediction-based identification. By using genomic sequences, metabolomic profiles, and machine learning parameters, the system can predict biosynthetic pathways and metabolite structures in silico, enabling scalable discovery of novel compounds without relying on traditional bioactivity screening

Inventive Principle:
Principle #35Parameter changes

2Measurement precision

If synthetic biology approaches with DNA manipulations are used, then biosynthetic pathways can be studied, but the process becomes expensive and not scalable to unstudied fungal species

Engineering Contradiction:
Improvebiosynthetic pathway characterization precisionVSAvoidDNA manipulation complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent creates computational copies and models of biosynthetic pathways through machine learning. Instead of physically manipulating DNA, the system uses genomic sequence data to create in silico models that predict pathway function and metabolite production. This copying approach maintains measurement precision by using validated predictive models while eliminating the complexity and cost of physical DNA manipulations

Inventive Principle:
Principle #26Copying

Solution Approach 2:

The patent introduces machine learning models and bioinformatics pipelines as intermediaries between genomic data and biosynthetic pathway characterization. These computational intermediaries translate raw genomic and metabolomic data into predictions about pathway function without requiring direct DNA manipulation, thereby reducing complexity while maintaining precision through systematic data analysis

Inventive Principle:
Principle #24Intermediary (Mediator)

3Loss of information

If extensive DNA manipulations are performed to study biosynthetic pathways, then pathway function can be determined, but the process becomes expensive and challenging to implement

Engineering Contradiction:
Improvebiosynthetic information completenessVSAvoiddiscovery process ease
Core Design Contradiction:
Loss of informationVSEase of manufacture

Solution Approach 1:

The patent substitutes mechanical DNA manipulation procedures with computational analysis of genomic and metabolomic data. Machine learning models predict biosynthetic pathway function and metabolite structures directly from sequence data, maintaining information completeness through integrated multi-omics analysis while dramatically improving ease of manufacture by eliminating complex wet-lab procedures

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Solution Approach 2:

The patent performs preliminary computational analysis of genomic sequences to predict biosynthetic pathways before any experimental work. By using machine learning to pre-identify potential pathways and prioritize targets, the system maintains complete biosynthetic information while simplifying subsequent experimental validation and improving overall process ease

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentUS20230035690A1Machine learning tools and a process to discover new natural products by linking genomes and metabolomes in fungi
Publication Date: 2023.02.02 NORTHWESTERN UNIV
  • US20230035690A1 patent drawing
  • US20230035690A1 patent drawing
  • US20230035690A1 patent drawing

AI summary

Provided herein are method of analyzing genomic and metabolomic data from fungi to identify relationships between biosynthetic gene clusters and mass spectrometric features of metabolites.