Drug-Target Interaction Data Cleaning via Graph Adjacency Matrix

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current methods for modeling drug-target interactions face challenges due to differences in chemical space between target proteins and drug molecules, leading to over-fitting, poor generalization, and inaccurate predictions.

Innovation Solution

A data cleaning method and system based on the adjacency matrix of a graph data structure are used to quantify and compare the structure and function of drug target proteins, allowing for proper classification, standardization, and improved model accuracy by filtering and processing drug-target interaction data.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Quantity of substance

If drug-target interaction data from different sources are directly mixed for modeling, then the data volume increases, but the model accuracy deteriorates due to differences in chemical space between target proteins

Engineering Contradiction:
Improvedata volumeVSAvoidmodel accuracy
Core Design Contradiction:
Quantity of substanceVSMeasurement precision

Solution Approach 1:

The patent segments the heterogeneous drug-target interaction data into distinct subsets based on target protein types or data sources. Each subset is processed and modeled separately to preserve the specific chemical space characteristics of each target type, avoiding the degradation of model accuracy that would result from direct mixing of diverse data.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent applies local quality by tailoring the modeling approach to each specific target protein type or data source. Each data subset is handled with specialized processing that accounts for its unique chemical space properties, ensuring that local characteristics are preserved while still achieving comprehensive coverage across all targets.

Inventive Principle:
Principle #3Local quality

2Adaptability or versatility

If data from multiple target proteins are combined, then the model coverage increases, but the generalization ability deteriorates due to over-fitting on specific target mechanisms

Engineering Contradiction:
Improvemodel coverageVSAvoidgeneralization ability
Core Design Contradiction:
Adaptability or versatilityVSReliability

Solution Approach 1:

The patent divides the combined data into segmented groups based on target protein classifications. By training separate models or using separate data subsets for different target types, the approach maintains specialized knowledge for each target class while preventing over-fitting to any single target's specific mechanism, thereby improving generalization ability across diverse targets.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent adjusts modeling parameters such as data subset selection, feature extraction methods, and model hyperparameters based on the specific characteristics of each target protein group. This dynamic parameter adjustment allows the model to adapt to different target types while maintaining robust generalization performance across the entire dataset.

Inventive Principle:
Principle #35Parameter changes

Data Source

PatentUS20240047082A1Method and device for cleaning drug-target interaction data
Publication Date: 2024.02.08 AINNOCENCE TECH LLC
  • US20240047082A1 patent drawing
  • US20240047082A1 patent drawing

AI summary

The present invention provides a cleaning method for drug-target interaction data, which comprises the following steps: provide an original collection of drug-target interaction data; screen and filter the original drug-target interaction data set according to a predetermined cleaning rule to obtain a drug-target interaction data set to be studied; wherein, the predetermined cleaning rule is based on the data structure of the adjacency matrix of the graph. The present invention further provides a cleaning system for drug-target interaction data. The present invention provides a method and system for properly describing and comparing the structure, function, etc. of drug target proteins, so as to the differences in target proteins can be quantified.