Drug-Target Interaction Data Cleaning via Graph Adjacency Matrix
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current methods for modeling drug-target interactions face challenges due to differences in chemical space between target proteins and drug molecules, leading to over-fitting, poor generalization, and inaccurate predictions.
Innovation Solution
A data cleaning method and system based on the adjacency matrix of a graph data structure are used to quantify and compare the structure and function of drug target proteins, allowing for proper classification, standardization, and improved model accuracy by filtering and processing drug-target interaction data.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Quantity of substance
If drug-target interaction data from different sources are directly mixed for modeling, then the data volume increases, but the model accuracy deteriorates due to differences in chemical space between target proteins
Solution Approach 1:
The patent segments the heterogeneous drug-target interaction data into distinct subsets based on target protein types or data sources. Each subset is processed and modeled separately to preserve the specific chemical space characteristics of each target type, avoiding the degradation of model accuracy that would result from direct mixing of diverse data.
Solution Approach 2:
The patent applies local quality by tailoring the modeling approach to each specific target protein type or data source. Each data subset is handled with specialized processing that accounts for its unique chemical space properties, ensuring that local characteristics are preserved while still achieving comprehensive coverage across all targets.
2Adaptability or versatility
If data from multiple target proteins are combined, then the model coverage increases, but the generalization ability deteriorates due to over-fitting on specific target mechanisms
Solution Approach 1:
The patent divides the combined data into segmented groups based on target protein classifications. By training separate models or using separate data subsets for different target types, the approach maintains specialized knowledge for each target class while preventing over-fitting to any single target's specific mechanism, thereby improving generalization ability across diverse targets.
Solution Approach 2:
The patent adjusts modeling parameters such as data subset selection, feature extraction methods, and model hyperparameters based on the specific characteristics of each target protein group. This dynamic parameter adjustment allows the model to adapt to different target types while maintaining robust generalization performance across the entire dataset.
Data Source
AI summary
The present invention provides a cleaning method for drug-target interaction data, which comprises the following steps: provide an original collection of drug-target interaction data; screen and filter the original drug-target interaction data set according to a predetermined cleaning rule to obtain a drug-target interaction data set to be studied; wherein, the predetermined cleaning rule is based on the data structure of the adjacency matrix of the graph. The present invention further provides a cleaning system for drug-target interaction data. The present invention provides a method and system for properly describing and comparing the structure, function, etc. of drug target proteins, so as to the differences in target proteins can be quantified.

