Machine Learning System for Biological Feature Identification

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current methods fail to effectively and efficiently analyze and predict the biological and genetic resources with practical impact, particularly in identifying biologically active features for pharmaceutical, chemical, and industrial applications, while also considering the ethical and regulatory frameworks like the Nagoya Protocol.

Innovation Solution

A system and method utilizing a resource database and computing components to analyze and identify biological features by creating profiles for retrieving compatible datasets, employing machine learning algorithms to determine similarity and compatibility, and automatically gathering and analyzing datasets for pharmaceutical, chemical, and industrial applications, with the ability to adapt and improve over time.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Productivity

If traditional methods are used to analyze genetic resources, then the process is simple and straightforward, but the efficiency and effectiveness of identifying biologically active features is insufficient

Engineering Contradiction:
Improveefficiency of identifying biologically active featuresVSAvoidcomplexity of the analysis system
Core Design Contradiction:
ProductivityVSDevice complexity

Solution Approach 1:

The system segments the complex task of genetic resource analysis into distinct functional modules: data collection module, data processing module, machine learning prediction module, and validation module. Each module handles specific aspects of the analysis process, enabling efficient identification of biologically active features while maintaining manageable system complexity through modular architecture

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent introduces machine learning algorithms as an intermediary between raw genetic data and biological activity identification. The ML models process and interpret complex genetic sequences, bridging the gap between data collection and actionable insights, thereby significantly improving identification efficiency without requiring direct complex manual analysis

Inventive Principle:
Principle #24Intermediary (Mediator)

2Measurement precision

If machine learning algorithms are employed to predict biological activity, then the accuracy and prediction quality improve, but the computational resources and time required increase

Engineering Contradiction:
Improveaccuracy of biological activity predictionVSAvoidcomputational energy consumption
Core Design Contradiction:
Measurement precisionVSUse of energy by moving object

Solution Approach 1:

The system performs preliminary data processing and feature extraction before feeding data to machine learning models. By pre-processing genetic sequences to identify and extract relevant features in advance, the system reduces the computational burden on ML algorithms during prediction, thereby improving accuracy while lowering energy consumption

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent dynamically adjusts model parameters and complexity levels based on the specific analysis requirements and available computational resources. The system can switch between different ML models or adjust prediction depth based on energy constraints, optimizing the balance between prediction accuracy and energy consumption for different scenarios

Inventive Principle:
Principle #35Parameter changes

3Quantity of substance

If comprehensive data collection from multiple sources is performed, then the quantity and quality of biological data increase, but the time and effort for data processing and validation increase

Engineering Contradiction:
Improvequantity of biological dataVSAvoidtime for data processing and validation
Core Design Contradiction:
Quantity of substanceVSLoss of time

Solution Approach 1:

The system incorporates automated data validation and quality control mechanisms that self-correct and filter data without extensive manual intervention. The machine learning models automatically assess data quality and flag anomalies, enabling the system to process and validate large quantities of data efficiently, reducing time loss while maintaining high data quality standards

Inventive Principle:
Principle #25Self-service

Data Source

PatentUS20240296911A1System and method for the identification of biological compounds from the genetic information in existing biological resources
Publication Date: 2024.09.05 SENCKENBERG GESELLSCHAFT FUR NATURFORSCHUNG
  • US20240296911A1 patent drawing
  • US20240296911A1 patent drawing

AI summary

The present invention is directed to an identification of biological feature in a biological resource. An analyzing component (10,15) can be configured for testing and/or analyzing the datasets retrieved for the same and/or a compatible and/or a similar biological activity of the features known. A similarity of the datasets to the biologically active features can be preset or pre-defined by a minimum threshold value of similarity. This can be a fixed or a dynamic threshold value. The similarity of the datasets to the biologically active features can be determined by a sliding minimum threshold value of similarity. The sliding minimum threshold value of similarity can set by a machine learning algorithm. This also comprises a change of the threshold value over time as a result of a machine learning training algorithm trained by results of testing and/or analyzing the datasets identified for the biological activity of the features known.