ML Data Debiasing Framework for Drug Discovery Prediction Pipelines

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

The rapid increase in biomedical data due to advancements in genomics, proteomics, and medical imaging has made data processing technologies inefficient, requiring substantial computing resources and time, especially in preprocessing and feature selection for predictive models, which are often costly and time-consuming.

Innovation Solution

A machine learning (ML) framework utilizing a parallel computing network to automate and optimize preprocessing algorithms, integrating multiple datasets, and selecting relevant features, reducing computational resource consumption.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If multiple preprocessing algorithms and feature selection methods are tested to optimize predictive model accuracy, then prediction accuracy is improved, but computational time and resource consumption increase substantially

Engineering Contradiction:
Improveprediction accuracyVSAvoidcomputational time
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The system performs preliminary action by pre-evaluating and ranking multiple preprocessing algorithms and feature selection methods before actual predictive modeling. The ML algorithm assesses all candidate methods in advance, establishing an optimized pipeline configuration that can be directly applied without extensive testing during production, thus reducing computational time while maintaining prediction accuracy.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The system implements self-service through automated ML-driven selection and optimization of preprocessing and feature selection methods. The ML algorithm autonomously evaluates different approaches, selects the optimal combination, and configures the predictive model without requiring manual intervention or extensive human expertise in data preprocessing techniques, thereby reducing both time and resource consumption.

Inventive Principle:
Principle #25Self-service

2Measurement precision

If multiple preprocessing algorithms and feature selection methods are tested to optimize predictive model accuracy, then prediction accuracy is improved, but computational resource consumption increases substantially

Engineering Contradiction:
Improveprediction accuracyVSAvoidcomputational resource consumption
Core Design Contradiction:
Measurement precisionVSUse of energy by moving object

Solution Approach 1:

The system implements self-service through automated ML-driven selection and optimization of preprocessing and feature selection methods. The ML algorithm autonomously evaluates different approaches, selects the optimal combination, and configures the predictive model without requiring manual intervention or extensive human expertise in data preprocessing techniques, thereby reducing both time and resource consumption.

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The system extracts and isolates the most critical preprocessing and feature selection methods from the large pool of candidates through ML-based evaluation. By identifying and extracting only the essential methods that contribute significantly to prediction accuracy, the system avoids wasting computational resources on evaluating less effective approaches, thus reducing overall resource consumption while maintaining accuracy.

Inventive Principle:
Principle #2Taking out (Extraction)

3Reliability

If expert biomedical scientists manually perform feature selection and preprocessing optimization, then domain expertise is utilized, but the process becomes time-consuming and resource-intensive

Engineering Contradiction:
Improvedomain expertise utilizationVSAvoidprocessing time
Core Design Contradiction:
ReliabilityVSLoss of time

Solution Approach 1:

The system replaces the mechanical process of manual expert analysis with an automated ML-based system. The ML algorithm incorporates domain knowledge through trained models that can evaluate preprocessing methods and feature selection strategies automatically, substituting human manual work with computational processes that are faster and more scalable while maintaining reliability through the integration of domain-specific training data and expertise.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Data Source

PatentUS12555041B2Methods, systems, and frameworks for debiasing data in drug discovery predictions
Publication Date: 2026.02.17 BIOSYMETRICS INC
  • US12555041B2 patent drawing
  • US12555041B2 patent drawing
  • US12555041B2 patent drawing

AI summary

Some embodiments relate to methods, systems, and frameworks for data analytics using machine learning, such as methods and systems for preprocessing of biomedical data, using machine learning, for input to a predictive model. The method may include receiving data from a data source, using at least one machine learning (ML) algorithm from a plurality of ML algorithms to obtain at least one combination of preprocessing steps, and computing an accuracy score for each of the at least one combination based on accuracy of prediction of the predictive model. The method may further include using at least one ML algorithm to optimize the feature selection of the predictive model, combining a plurality of datasets into a single dataset, and using a parallel computing network to provide a framework for executing such predictive model.