Drug interaction prediction method based on multi-dimensional features

By cleaning and quality detection of drug data, a high-quality data sample set is generated, and multi-dimensional feature extraction is used to use a pre-trained drug interaction prediction model to solve the problem of insufficient accuracy of drug interaction prediction in the prior art, and more efficient and reliable prediction results are achieved.

CN120048548AInactive Publication Date: 2025-05-27WEIFANG MEDICAL UNIV
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
CN202510502325.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-22
Publication Date
2025-05-27
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

The prior art fails to fully utilize the multidimensional characteristics of drugs in drug interaction prediction, resulting in insufficient accuracy and reliability of the prediction results, and lacks effective quality detection and optimization methods in the data processing stage, resulting in noise and incomplete data affecting the prediction results.

Method used

By collecting critical data samples from the drug dataset, performing data cleaning and quality detection, high-quality critical data sample sets are generated, and multi-dimensional feature extraction and prediction are used to utilize pre-trained drug interaction prediction models.

Benefits of technology

It improves the accuracy and efficiency of drug interaction prediction, ensures the reliability of input data, reduces prediction errors, and provides more accurate decision-making support for drug research and development and clinical use.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120048548A_ABST
    Figure CN120048548A_ABST
Patent Text Reader

Abstract

The invention provides a multi-dimensional feature-based drug interaction prediction method. The method comprises the steps of collecting key data samples in a drug data set and generating a sample set; sending the sample set to a feature extraction module of a pre-training model, and extracting multi-dimensional feature data; sending the feature data to a server, and generating a drug interaction prediction result by a prediction module; and displaying the result on a computing device interface. Optionally, the method further comprises the steps of data cleaning, quality detection, activity detection, tracking analysis and the like, and the data reliability is ensured. The prediction model is generated through a training data set, comprises a feature extraction module and a prediction module, and is used for processing drug features and outputting prediction labels. According to the method, the result reliability can be improved, and medicine research and development and clinical medication decision making can be assisted.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical fields of pharmaceutical information processing and artificial intelligence. More specifically, the present invention relates to a method for predicting drug interactions based on multi-dimensional features. Background Art

[0002] In today's medical field, the prediction of drug interactions is a crucial research topic. With the continuous increase in the types of drugs and the increasing prevalence of polypharmacy in patients, accurately predicting drug interactions is of great significance for ensuring the safety of patients' medication, improving the therapeutic effect, and reducing the incidence of adverse drug reactions. Traditional methods for predicting drug interactions mainly rely on clinical experience and experimental research. Although these methods can provide valuable information to a certain extent, they have many limitations. Clinical experience is often based on doctors' personal knowledge and past cases, which is difficult to cover all possible drug combinations and is easily affected by subjective factors. Experimental research requires a large amount of time and resources, and it is often difficult to conduct timely research on the interactions of new drugs or rare drug combinations.

[0003] In recent years, with the development of computing technology, data-driven prediction models have gradually become a research hotspot. These models analyze a large amount of drug data and use machine learning or deep learning algorithms to predict drug interactions. However, most of the existing data-driven prediction methods only focus on single features of drugs, such as chemical structure or pharmacological effects, while ignoring the multi-dimensional characteristics of drugs. The interaction of drugs is a complex biochemical process involving the combined action of multiple factors. Relying solely on single features for prediction often leads to insufficient accuracy and reliability of the prediction results. In addition, existing methods also have deficiencies in data processing and fail to fully consider the quality and integrity of the data, which may introduce noise and bias, further affecting the prediction effect.

[0004] In the process of implementing the embodiments of the present invention, the inventors found that there are at least the following problems or defects in the prior art: In the prediction of drug interactions, the prior art fails to fully utilize the multi-dimensional features of drugs for comprehensive analysis, resulting in limited accuracy and reliability of the prediction results; in the data processing stage, there is a lack of effective means for detecting and optimizing data quality, so that noise data and incomplete data have a negative impact on the prediction results; in addition, in the process of model training, the existing methods do not handle relevant and irrelevant data finely enough and fail to fully utilize global and local feature information for prediction, thus affecting the performance and generalization ability of the model. Summary of the Invention

[0005] The present invention provides a method for predicting drug interactions based on multi-dimensional features, including: Collect key data samples from the drug dataset, and generate a key data sample set from the key data samples; send the key data sample set to a feature extraction module in a pre-trained drug interaction prediction model for feature extraction to obtain multi-dimensional feature data, where the feature extraction module of the drug interaction prediction model is pre-deployed on a computing device; send the multi-dimensional feature data to a server for a prediction module corresponding to the drug interaction prediction model in the server to predict drug interactions and obtain a prediction result of drug interactions; in response to receiving the prediction result of drug interactions returned by the server, display the result on the interface of the computing device.

[0006] Further, before sending the key data sample set to the pre-trained drug interaction prediction model, perform the following steps: perform data cleaning processing on the key data samples to remove abnormal data information therein.

[0007] Further, the method further includes: performing data quality detection on each group of data in the drug dataset to generate a data quality detection result set; determining the data quality detection results in the data quality detection result set that meet preset conditions, and determining the drug data corresponding to the data quality detection results that meet the preset conditions as key data samples.

[0008] Further, the method further includes: performing activity detection on the drug dataset to determine target activity data in the drug dataset; performing tracking analysis on the detected target activity data, where tracking the marked target activity data to obtain a tracking identification set; determining the data range of the detected target activity data in the drug dataset to obtain an activity data range set; based on the tracking identification set and the activity data range set, performing data quality detection on the drug data corresponding to each detected target activity data to generate a data quality detection result set.

[0009] Further, the method further includes: in response to determining that there is no data quality detection result in the data quality detection result set that meets the preset conditions, selecting key data samples from the drug data collected by the data acquisition device again.

[0010] Further, based on the tracking identifier set and the active data range set, performing data quality detection on the drug data corresponding to each detected target active data to generate a data quality detection result set, including: determining an integrity index set and a data accuracy index set of the drug data corresponding to each tracking identifier in the tracking identifier set according to the active data range set; performing drug property analysis on the drug data corresponding to each active data range in the active data range set to generate a data property information set, where each data property information in the data property information set corresponds to a drug property, and each data property information includes a classification label of the drug property in the key data sample, and the classification label includes at least one of the following: a chemical structure label, a pharmacological action label, or a metabolic pathway label; determining a data quality detection result set corresponding to each tracking identifier according to the integrity index set, the data accuracy index set, and the data property information set.

[0011] Further, the method further includes: determining an integrity index of the drug data corresponding to each tracking identifier in the tracking identifier set to obtain an integrity index set; determining an accuracy index of the drug data corresponding to each tracking identifier in the tracking identifier set to obtain a data accuracy index set.

[0012] Further, the drug interaction prediction model is generated by a preset model generation method; wherein, the preset model generation method includes: obtaining a training data set, where the training samples in the training data set include relevant data instances, irrelevant data instances, and corresponding relevant labels and irrelevant labels; selecting training samples from the training data set and inputting them into the feature extraction module of the initial drug interaction prediction model to generate relevant feature data and irrelevant feature data, where the model parameters of the feature extraction module in the initial drug interaction prediction model are fixed, and the prediction module in the initial drug interaction prediction model includes a global prediction sub-module and a local prediction sub-module, and the prediction module corresponds to a single drug combination; inputting the relevant feature data into the prediction module of the initial drug interaction prediction model to output relevant prediction labels, where the local prediction sub-module in the prediction module is used to perform prediction processing on the drug local features in the relevant feature data; inputting the irrelevant feature data into the prediction module of the initial drug interaction prediction model to output irrelevant prediction labels, where the global prediction sub-module in the prediction module is used to perform prediction processing on the drug overall features in the relevant feature data; generating a sample error metric according to the obtained relevant prediction labels, irrelevant prediction labels, relevant labels, and irrelevant labels; adjusting the network parameters in the prediction module according to the sample error metric.

[0013] Further, the training data set is generated through the following steps: receiving a removal operation of the user on the drug data in the displayed drug interaction data set; determining the drug data corresponding to the removal operation as irrelevant data instances to obtain a group of irrelevant data instances, where the irrelevant data instance is the irrelevant data removed by the user; determining the tracking identifier corresponding to each irrelevant data in the group of irrelevant data instances as an irrelevant label to obtain a group of irrelevant labels; receiving an association operation of the user on the drug data in the displayed drug interaction data set; adjusting the tracking identifiers corresponding to the association operation to the same identifier, and determining each drug data in the drug interaction data set as relevant data instances to obtain a group of relevant data instances; determining the tracking identifier corresponding to each relevant data instance in the group of relevant data instances as a relevant label to obtain a group of relevant labels; combining the group of relevant data instances, the group of relevant labels, the group of irrelevant data instances and the group of irrelevant labels to generate a training data set.

[0014] Further, the feature extraction is used to extract multiple characteristic features of the drug, including at least one of the following: chemical structure feature, pharmacological action feature, metabolic pathway feature and side effect feature, for multi-feature fusion prediction.

[0015] According to the above embodiments of the present invention, there are at least the following beneficial effects: The present invention can improve the accuracy and efficiency of drug interaction prediction. By extracting multi-dimensional features and analyzing drug data with a pre-trained model, it can comprehensively capture key features such as the chemical structure, pharmacological action and metabolic pathway of the drug. Combining the data cleaning and quality detection mechanisms can ensure the reliability of the input data, thereby reducing prediction errors and providing more accurate decision-making support for drug research and development and clinical medication.

[0016] This method can optimize the drug data analysis process. By detecting activity, tracking analysis and dynamically screening key samples, it can automatically identify high-quality data and eliminate abnormal information. Based on the synergistic effect of the global and local prediction sub-modules, it can take into account the relevance of the overall characteristics and local features of the drug at the same time, further improving the generalization ability of the model, and finally providing an intelligent solution for pharmaceutical research and rational drug use. BRIEF DESCRIPTION OF THE DRAWINGS

[0017] By reading the following detailed description with reference to the accompanying drawings, the above and other objects, features and advantages of the exemplary embodiments of the present invention will become readily understood. In the drawings, several embodiments of the present invention are shown in an exemplary rather than restrictive manner, wherein: Figure 1 It is a schematic flow chart of a drug interaction prediction method based on multi-dimensional features provided by an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0018] The principles and spirit of the present invention will be described below with reference to several exemplary embodiments. It should be understood that these embodiments are provided only to enable those skilled in the art to better understand and then implement the present invention, rather than limiting the scope of the present invention in any way. On the contrary, these embodiments are provided to make the present invention more thorough and complete, and to be able to fully convey the scope of the present invention to those skilled in the art.

[0019] Those skilled in the art know that the embodiments of the present invention can be implemented as a system, device, equipment, method, or computer program product. Therefore, the present invention can be specifically implemented in the following forms, namely: completely hardware, completely software (including firmware, resident software, microcode, etc.), or a combination of hardware and software.

[0020] It should be noted that any number of elements in the drawings is for illustration rather than limitation, and any naming is only for distinction and does not have any limiting meaning.

[0021] The following reference Figure 1 , Figure 1 is a schematic flowchart of a method for predicting drug interactions based on multi-dimensional features provided for an embodiment of the present invention. As Figure 1 shown, a method for predicting drug interactions based on multi-dimensional features includes: S1 Collect key data samples in a drug dataset and generate a key data sample set from the key data samples; S2 Send the key data sample set to a feature extraction module in a pre-trained drug interaction prediction model for feature extraction to obtain multi-dimensional feature data, wherein the feature extraction module of the drug interaction prediction model is pre-deployed on a computing device; S3 Send the multi-dimensional feature data to a server for a prediction module corresponding to the drug interaction prediction model in the server to predict drug interactions and obtain a prediction result of drug interactions; S4 In response to receiving the prediction result of drug interactions returned by the server, display the result on the interface of the computing device.

[0022] It should be noted that the present invention relates to a method for predicting drug interactions based on multi-dimensional features. This method is applied to a computing device. Its core lies in collecting key data samples in a drug dataset and generating a key data sample set, and then using a pre-trained drug interaction prediction model for feature extraction and prediction. Here, the key data samples refer to the representative and important data screened from the drug dataset, which can reflect the basic characteristics of drugs and the potential information of drug interactions. The drug interaction prediction model is a model constructed through machine learning or deep learning techniques, used to analyze the interaction relationships between drugs and output prediction results. The feature extraction module of this model is pre-deployed on the computing device, which means that part of the functions of the model have been integrated into the computing device to facilitate rapid data processing and feature extraction.

[0023] Specifically, a drug dataset usually contains data in multiple aspects such as the chemical structure information, pharmacological effects, metabolic pathways, side effects of drugs, etc. The process of generating the key data sample set is to screen and sort the drug dataset to extract the data related to drug interaction prediction. For example, drug data with a specific chemical structure can be extracted from the drug dataset, or drug data showing an interaction tendency under a specific pharmacological effect can be extracted. Before sending the key data sample set to the drug interaction prediction model, data cleaning processing is required to remove the abnormal data information therein. Data cleaning is an important preprocessing step, which can help improve the quality and reliability of the data. The abnormal data information may include incorrect chemical structure data, unreasonable pharmacological effect descriptions, or incomplete metabolic pathway information, etc. By data cleaning, it can be ensured that the data input into the model is accurate and complete, thereby improving the accuracy of the prediction results.

[0024] Preferably, when generating the key data sample set, data quality detection can be further performed on each group of data in the drug dataset to generate a data quality detection result set. Data quality detection can be achieved through a series of algorithms and rules. For example, check the integrity, accuracy, and consistency of the data. For chemical structure data, it can be detected whether it conforms to known chemical rules; for pharmacological effect data, it can be verified whether it conforms to existing pharmacological knowledge. After determining the data quality detection results in the data quality detection result set that meet the preset conditions, the corresponding drug data is determined as the key data sample. The preset conditions can be set according to actual needs. For example, it is required that the integrity index of the data reaches a certain threshold, or the accuracy index of the data is within an acceptable range. In addition, activity detection can also be performed on the drug dataset to determine the target activity data therein and conduct tracking analysis. Activity detection can help identify the drug components or characteristics that may play a key role in drug interactions, thereby further improving the representativeness of the key data samples.

[0025] In some embodiments, before sending the critical data sample set to a pre-trained drug interaction prediction model, the following steps are performed: performing data cleaning on the critical data samples to remove abnormal data information therein.

[0026] It should be noted that during the process of processing critical data samples, data cleaning is an important preprocessing step. The purpose of data cleaning is to remove abnormal data information therein to ensure that the data input into the drug interaction prediction model is accurate and reliable. Here, the abnormal data information refers to information that does not conform to drug characteristics, has incorrect data formats, or is significantly inconsistent with other data. For example, there may be parts in chemical structure data that do not conform to known chemical rules, or there may be content in pharmacological effect data that is contrary to known drug action mechanisms. Through data cleaning, the interference of such abnormal data on the prediction results can be effectively reduced, and the accuracy and robustness of the model can be improved.

[0027] Specifically, data cleaning can be achieved through various methods. For example, for chemical structure data, chemical informatics tools can be used to verify its rationality and check whether the molecular structure conforms to known chemical bond rules and stereochemical requirements. For pharmacological effect data, it can be compared with existing pharmacological databases to identify data that does not conform to known pharmacological effects.

[0028] Furthermore, some thresholds can be set to judge the rationality of data. For example, for the metabolic rate data of a drug, if a certain value exceeds the known physiological range, it can be regarded as abnormal data and cleaned. In actual operation, the specific parameter settings for data cleaning can be determined according to the characteristics of the drug data set and the requirements of the prediction task. For example, for a data set containing the chemical structures of multiple drugs, a reasonable molecular weight range can be set, and the data outside this range can be marked as abnormal and processed.

[0029] Preferably, during data cleaning, the cleaning strategy can be further refined. For example, for some data that may have errors but are not completely wrong, the method of data correction can be adopted instead of direct deletion. For example, if there are partial missing values in the metabolic pathway data of a certain drug, the metabolic pathway information of similar drugs can be referred to for supplementation and correction. In addition, machine learning algorithms can be introduced to assist data cleaning. For example, a classification model can be trained to identify abnormal data. This model can learn based on labeled normal and abnormal data samples, and thus automatically identify and mark abnormal data in new data. This machine learning-based data cleaning method can not only improve the cleaning efficiency but also improve the cleaning accuracy.

[0030] In some embodiments, the method further includes: performing data quality detection on each group of data in the drug dataset to generate a data quality detection result set; determining the data quality detection results in the data quality detection result set that meet the preset conditions, and determining the drug data corresponding to the data quality detection results that meet the preset conditions as key data samples.

[0031] It should be noted that in the process of generating the key data sample set of the present invention, not only key data samples are directly collected, but also the step of data quality detection is added. The purpose of data quality detection is to comprehensively evaluate each group of data in the drug dataset and generate a data quality detection result set. Here, the data quality detection result set refers to a set containing the quality evaluation results of each group of data, and these results reflect indicators such as the integrity, accuracy, and consistency of the data. By determining the data quality detection results in the data quality detection result set that meet the preset conditions, the corresponding drug data can be determined as key data samples, thereby ensuring the quality and reliability of the key data samples and providing a more accurate data basis for subsequent drug interaction prediction.

[0032] Specifically, data quality detection involves multiple aspects. First, integrity detection refers to checking whether the data lacks key information. For example, for the chemical structure data of drugs, it is necessary to ensure that all necessary parts of the molecular structure have been completely recorded; for the pharmacological effect data, it is necessary to check whether the main action mechanism and target information of the drug are included. Accuracy detection is to verify whether the data is correct. For example, by comparing with a known drug database, it can be checked whether the chemical structure of the drug is consistent with the standard structure and whether the pharmacological effect conforms to known pharmacological knowledge. Consistency detection is to ensure whether the data is logically coherent. For example, whether the metabolic pathway of the drug matches its chemical structure and pharmacological effect. In actual operation, specific thresholds can be set for each detection index. For example, the integrity index can be set that the data missing rate is lower than 10%, and the accuracy index can be set that the data error rate is lower than 5%. These parameters can be adjusted according to the characteristics of the actual dataset and the requirements of the prediction task.

[0033] Preferably, in the process of data quality detection, the detection steps can be further refined. For example, in integrity detection, for the case of missing some data, data imputation methods can be used for supplementation instead of directly discarding the data. The data imputation method can be estimated based on the data of similar drugs or filled using statistical methods. In accuracy detection, in addition to comparing with the known database, expert knowledge can be introduced for manual review, especially for some complex drug characteristic data.

[0034] Furthermore, a multi-dimensional detection method can also be adopted, such as combining the similarity analysis of chemical structures and the correlation analysis of pharmacological effects to comprehensively evaluate the quality of the data. This multi-dimensional detection method can more comprehensively evaluate the data quality and improve the screening accuracy of key data samples.

[0035] In some embodiments, the method further includes: performing activity detection on the drug data set to determine the target activity data in the drug data set; performing a tracking analysis on the detected target activity data, wherein tracking the labeled target activity data to obtain a tracking identification set; determining the data range of the detected target activity data in the drug data set to obtain an activity data range set; and based on the tracking identification set and the activity data range set, performing data quality detection on the drug data corresponding to each detected target activity data to generate a data quality detection result set.

[0036] It should be noted that in the process of generating the key data sample set of the present invention, an activity detection step is further introduced. The purpose of the activity detection is to determine the target activity data in the drug data set, which usually refers to the drug properties or components that have a significant impact in drug interactions. By tracking and analyzing these target activity data, the data characteristics related to drug interactions can be more accurately identified. Specifically, tracking the labeled target activity data to obtain a tracking identification set helps to record and trace the source and changes of the activity data. At the same time, determining the data range of the target activity data in the drug data set to obtain an activity data range set helps to clarify the distribution and coverage of the activity data. Based on the tracking identification set and the activity data range set, performing data quality detection on the drug data corresponding to each target activity data to generate a data quality detection result set further ensures the quality and reliability of the key data samples.

[0037] Specifically, activity detection involves identifying and analyzing active ingredients or characteristics in a drug dataset. For example, for a drug, its active ingredients may be certain specific chemical groups or metabolites, and these ingredients may play a key role in drug interactions. When performing activity detection, specific parameters can be set to determine the target activity data. For example, an activity threshold can be set, and when a certain characteristic of the drug (such as metabolic rate, pharmacological action intensity, etc.) exceeds this threshold, it is regarded as the target activity data. Trace analysis can be achieved by recording the data source, change process, and other associated data. For example, for the metabolic pathway data of a drug, its changes at different time points and its interactions with other drug components can be tracked. When determining the activity data range, factors such as the chemical structure similarity and pharmacological action relevance of the drug can be considered, and data with similar characteristics can be grouped into the same range. Data quality detection can evaluate the integrity, accuracy, and consistency of relevant drug data based on the activity data range set and the trace identification set.

[0038] Preferably, during the activity detection and data quality detection processes, the operation steps can be further refined. For example, in activity detection, a combination of multiple detection methods can be used, such as chemical analysis methods and bioactivity testing methods, to improve the accuracy of detection. For trace analysis, data mining techniques can be introduced to automatically identify and record the change rules and association patterns of activity data. When determining the activity data range, machine learning algorithms can be used to learn based on existing activity data samples and automatically divide the data range. In data quality detection, in addition to evaluating integrity, accuracy, and consistency, data timeliness detection can also be introduced to ensure the update and effectiveness of the data. For example, for the clinical trial data of some drugs, it can be checked whether it is the latest research result. In addition, an expert system can be introduced for manual review, especially for some complex or uncertain data, to further improve the reliability of data quality detection.

[0039] In some embodiments, the method further includes: in response to determining that there is no data quality detection result in the data quality detection result set that meets the preset conditions, selecting key data samples from the drug data collected by the data acquisition device again.

[0040] It should be noted that a feedback mechanism is introduced in the data quality detection process of the present invention. When there is no data quality detection result in the data quality detection result set that meets the preset conditions, that is, all detection results do not reach the expected quality standard, the system will trigger an operation to reselect key data samples. This mechanism ensures the quality of key data samples and avoids inaccurate drug interaction prediction results caused by data quality problems. The preset conditions here refer to the data quality standards set according to the requirements of the drug interaction prediction task, such as the minimum thresholds of indicators such as data integrity and accuracy.

[0041] Specifically, the data quality detection result set is generated by evaluating the quality of each group of data in the drug dataset. During the detection process, integrity indicators, accuracy indicators, etc. of each group of data will be calculated, and these indicators reflect the quality level of the data. For example, the integrity indicator can be evaluated by checking whether there are missing values in the data, and the accuracy indicator can be determined by comparing with known standard data. The preset conditions can be set according to actual needs. For example, the threshold of the integrity indicator is set to 90%, and the threshold of the accuracy indicator is set to 95%. If none of the groups of data in the data quality detection result set meet these preset conditions, it means that the data quality in the current dataset is generally low and cannot be directly used for subsequent prediction model training or prediction tasks.

[0042] Preferably, when reselecting key data samples, the operation steps can be further refined. For example, the data acquisition device can be checked to determine whether there are data acquisition errors or omissions. If problems are found in the data acquisition device, the data acquisition parameters can be repaired or adjusted in time to improve data quality.

[0043] Furthermore, the scope and frequency of data acquisition can also be adjusted. For example, the scope of data acquisition can be expanded to increase the diversity of samples, or the frequency of data acquisition can be increased to obtain more comprehensive and accurate data. After reselecting key data samples, data quality detection can be performed again to ensure that the newly selected data samples meet the preset conditions. If data samples that meet the conditions still cannot be obtained after multiple detections, external data sources can be considered, or the existing data can be further cleaned and preprocessed to improve data quality.

[0044] In some embodiments, based on the tracking identifier set and the active data range set, data quality detection is performed on the drug data corresponding to each detected target active data to generate a data quality detection result set, including: determining an integrity index set and a data accuracy index set of the drug data corresponding to each tracking identifier in the tracking identifier set according to the active data range set; performing drug characteristic analysis on the drug data corresponding to each active data range in the active data range set to generate a data characteristic information set, wherein each data characteristic information in the data characteristic information set corresponds to a drug characteristic, and each data characteristic information includes a classification label of the drug characteristic in the key data sample, and the classification label includes at least one of the following: chemical structure label, pharmacological action label, or metabolic pathway label; determining a data quality detection result set corresponding to each tracking identifier according to the integrity index set, the data accuracy index set, and the data characteristic information set.

[0045] It should be noted that in the process of data quality detection of the present invention, the detection method is further refined. By combining the tracking identifier set and the active data range set, multi-dimensional quality detection is performed on the drug data corresponding to the target active data. Here, the tracking identifier set refers to a set of identifiers generated during the active detection process, used to record and trace the source and changes of the target active data; the active data range set refers to the range covered by the target active data in the drug dataset. Through these two sets, the integrity, accuracy, and characteristic information of the drug data can be more accurately evaluated, and a data quality detection result set is generated. This process not only considers the basic quality indicators of the data but also combines the characteristic information of the drug, such as chemical structure, pharmacological action, and metabolic pathway, etc., thus providing a higher-quality data basis for subsequent drug interaction prediction.

[0046] Specifically, the integrity index set and the data accuracy index set are generated by evaluating the drug data corresponding to each tracking identifier in the tracking identifier set. The integrity index can be calculated by checking whether there are missing values in the data and whether the data fields are complete. For example, for the chemical structure data of a drug, the integrity index can check whether information such as its molecular formula and functional groups is complete. The accuracy index can be determined by comparing with known standard data or authoritative databases. For example, for the pharmacological action data of a drug, it can be compared with the standard description in the pharmacological database to calculate its accuracy. The data characteristic information set is generated by performing drug characteristic analysis on the drug data corresponding to each active data range in the active data range set. For example, the chemical structure label can identify the molecular structure type of the drug, the pharmacological action label can identify the main action mechanism of the drug, and the metabolic pathway label can identify the metabolic process of the drug in the body. These classification labels of the characteristic information provide a richer dimension for subsequent data quality evaluation.

[0047] Preferably, during the process of generating the data quality detection result set, the operation steps can be further refined. For example, when determining the integrity index, a weight mechanism can be introduced, and different weights can be assigned according to the importance of different data fields. For key fields (such as the main chemical components or main pharmacological effects of drugs), higher weights can be given to more accurately reflect the integrity of the data. When determining the accuracy index, a multi-source data comparison method can be adopted, and the information of multiple authoritative databases can be combined to improve the reliability of the accuracy assessment. For the generation of the data characteristic information set, machine learning algorithms can be introduced to automatically identify and classify the characteristic information of drugs. For example, by training a classification model, based on the labeled drug data samples, the chemical structure labels, pharmacological effect labels, and metabolic pathway labels of new drug data can be automatically identified. In addition, an expert review mechanism can be introduced, and for some complex or uncertain drug characteristic information, professional personnel can conduct final review and confirmation to ensure the accuracy of the data characteristic information.

[0048] In some embodiments, the method further includes: determining the integrity index of the drug data corresponding to each tracking identifier in the tracking identifier set to obtain an integrity index set; determining the accuracy index of the drug data corresponding to each tracking identifier in the tracking identifier set to obtain a data accuracy index set.

[0049] It should be noted that in the data quality detection process of the present invention, the methods for determining the integrity index and the data accuracy index are further clarified. The integrity index set refers to a set of indexes generated by evaluating the integrity of the drug data corresponding to each tracking identifier in the tracking identifier set, and these indexes reflect the missing situation of the data in each dimension. The data accuracy index set refers to a set of indexes generated by evaluating the accuracy of the drug data, and these indexes reflect the correctness and reliability of the data. By separately determining these indexes, the quality of the drug data can be evaluated more meticulously, providing more accurate data support for subsequent drug interaction prediction.

[0050] Specifically, the determination of the integrity index can be achieved by checking the missing situation of data fields. For example, for the chemical structure data of drugs, it can be checked whether key information such as molecular formula and functional groups is complete; for the pharmacological action data, it can be checked whether information such as action mechanism and targets is complete. The determination of the data accuracy index can be achieved by comparing with authoritative databases or known standard data. For example, for the drug metabolism pathway data, it can be compared with the known metabolism pathway database to check the accuracy of the data. In actual operation, weights can be set for each data field, and the integrity index and accuracy index can be calculated according to the importance of the field and the missing or error situation. For example, for the main chemical composition field of drugs, a higher weight can be assigned because this information is crucial for predicting drug interactions.

[0051] Preferably, when determining the integrity index and data accuracy index, the operation steps can be further refined. For example, in the calculation of the integrity index, a fault tolerance mechanism can be introduced to allow a certain degree of data missing, but the missing part is marked and recorded. For the data accuracy index, a multi-source data verification method can be adopted, combining the information of multiple authoritative databases to improve the reliability of accuracy evaluation.

[0052] Furthermore, automated tools and scripts can also be introduced to assist in data quality detection, improving the detection efficiency and accuracy. For example, using data cleaning tools to automatically detect and repair errors in the data, or using scripts to automatically compare the differences between the data and the standard database. These methods can effectively improve the efficiency and accuracy of data quality detection, ensuring the quality of key data samples.

[0053] In some embodiments, the drug interaction prediction model is trained and generated by a preset model generation method; wherein, the preset model generation method includes: obtaining a training data set, wherein the training samples in the training data set include relevant data instances, irrelevant data instances, and corresponding relevant labels and irrelevant labels; selecting training samples from the training data set and inputting them into the feature extraction module of the initial drug interaction prediction model to generate relevant feature data and irrelevant feature data, wherein the model parameters of the feature extraction module in the initial drug interaction prediction model are fixed, and the prediction module in the initial drug interaction prediction model includes a global prediction sub-module and a local prediction sub-module, and the prediction module corresponds to a single drug combination; inputting the relevant feature data into the prediction module of the initial drug interaction prediction model to output relevant prediction labels, wherein the local prediction sub-module in the prediction module is used to perform prediction processing on the drug local features in the relevant feature data; inputting the irrelevant feature data into the prediction module of the initial drug interaction prediction model to output irrelevant prediction labels, wherein the global prediction sub-module in the prediction module is used to perform prediction processing on the drug overall features in the relevant feature data; generating a sample error metric based on the obtained relevant prediction labels, irrelevant prediction labels, relevant labels, and irrelevant labels; and adjusting the network parameters in the prediction module according to the sample error metric.

[0054] It should be noted that the drug interaction prediction model mentioned in the present invention is trained and generated by a preset model generation method, and the core of this method lies in using the training data set to train and optimize the model. The training samples in the training data set include relevant data instances, irrelevant data instances, and corresponding relevant labels and irrelevant labels, and these data are used to train the model to distinguish the interactions between drugs. The initial drug interaction prediction model includes a feature extraction module and a prediction module, wherein the feature extraction module is used to extract multi-dimensional feature data of drugs, and the prediction module is used to predict the drug interactions based on these feature data. The prediction module is further divided into a global prediction sub-module and a local prediction sub-module. The global prediction sub-module is responsible for performing prediction processing on the overall features of drugs, while the local prediction sub-module focuses on the local features of drugs. In this way, the model can analyze the interactions between drugs more comprehensively.

[0055] Specifically, the construction of the training dataset is one of the key steps in model generation. Relevant data instances refer to those drug combinations known to have interactions, while irrelevant data instances refer to those drug combinations known not to have interactions. Relevant labels and irrelevant labels are used to mark the categories of these data instances respectively. During the training process, training samples are selected from the training dataset and input into the feature extraction module of the initial drug interaction prediction model to generate relevant feature data and irrelevant feature data. These feature data are then input into the prediction module, where the local prediction sub-module performs prediction processing on the local features of the drugs, and the global prediction sub-module performs prediction processing on the overall features of the drugs. By comparing the relevant prediction labels and irrelevant prediction labels output by the prediction module with the actual relevant labels and irrelevant labels, a sample error metric can be generated. Based on this sample error metric, the network parameters in the prediction module are further adjusted to optimize the performance of the model.

[0056] Preferably, during the model training process, the operation steps can be further refined. For example, when selecting training samples, a stratified sampling method can be adopted to ensure that the training samples have good representativeness between relevant and irrelevant data instances. In the feature extraction module, various feature extraction algorithms can be introduced, such as feature extraction algorithms based on chemical structure, feature extraction algorithms based on pharmacological effects, etc., to more comprehensively extract the multi-dimensional features of the drugs. In the prediction module, deep learning algorithms, such as neural networks, can be used to improve the prediction ability of the model. In addition, techniques such as cross-validation can be introduced to evaluate and optimize the performance of the model. For example, by dividing the training dataset into multiple subsets and performing multiple trainings and validations, the generalization ability and accuracy of the model can be ensured.

[0057] In some embodiments, the training dataset is generated through the following steps: receiving a user's removal operation on the drug data in the displayed drug interaction dataset; determining the drug data corresponding to the removal operation as irrelevant data instances to obtain a group of irrelevant data instances, where the irrelevant data instances are the irrelevant data removed by the user; determining the tracking identifier corresponding to each irrelevant data in the group of irrelevant data instances as an irrelevant label to obtain a group of irrelevant labels; receiving a user's association operation on the drug data in the displayed drug interaction dataset; adjusting the tracking identifiers corresponding to the association operation to the same identifier, and determining each drug data in the drug interaction dataset as relevant data instances to obtain a group of relevant data instances; determining the tracking identifier corresponding to each relevant data instance in the group of relevant data instances as a relevant label to obtain a group of relevant labels; combining the group of relevant data instances, the group of relevant labels, the group of irrelevant data instances, and the group of irrelevant labels to generate a training dataset; The drug interaction dataset is constructed by integrating multi-source heterogeneous data. The data sources of the drug interaction dataset include publicly available medical databases (such as DrugBank, PubChem, ChEMBL), published clinical trial data, pharmacological research literature, and drug labels from drug regulatory agencies (such as FDA, EMA). It can also include experimental measurement results, such as in vitro activity tests and in vivo pharmacokinetic studies. The drug interaction data can include the chemical structure characteristics, pharmacological action characteristics, metabolic pathway characteristics, and side effect characteristics of drugs.

[0058] It should be noted that the generation process of the training dataset in the present invention is completed through the user's operations on the drug interaction dataset. The user performs a removal operation on the drug data, and determines the removed drug data as irrelevant data instances. These irrelevant data instances form an irrelevant data instance group, and the tracking identifier corresponding to each irrelevant data instance is determined as an irrelevant label, forming an irrelevant label group. Similarly, the user performs an association operation on the drug data, and determines the associated drug data as relevant data instances. These relevant data instances form a relevant data instance group, and the tracking identifier corresponding to each relevant data instance is determined as a relevant label, forming a relevant label group. Finally, the relevant data instance group, relevant label group, irrelevant data instance group, and irrelevant label group are combined to generate the training dataset. This process makes full use of the user's operation behavior, enabling the training dataset to better conform to the actual drug interaction situation, thereby improving the training effect and prediction accuracy of the model.

[0059] Specifically, the user's operations on the drug interaction dataset are the basis for generating the training dataset. The removal operation refers to the user removing from the dataset those drug data that are known to have no interactions, and these data are marked as irrelevant data instances. The association operation refers to the user associating those drug data that are known to have interactions, and these data are marked as relevant data instances. The tracking identifier is an identifier used to uniquely identify each drug data, and through the tracking identifier, the data instance can be associated with the corresponding label. When generating the training dataset, some parameters can be set to control the quality and scale of the dataset. For example, a minimum sample number threshold can be set to ensure that the number of data instances in each category (relevant and irrelevant) reaches a certain amount to guarantee the sufficiency of model training. At the same time, the data can be preprocessed, such as removing duplicate data, standardizing the data format, etc., to improve the quality of the data.

[0060] Preferably, during the process of generating the training data set, the operation steps can be further refined. For example, for the irrelevant data instances removed by the user, it can be further verified whether they are indeed irrelevant to drug interactions to avoid data errors caused by misoperations. The expert review mechanism can be introduced to review the data operated by the user to ensure the accuracy of the data. When combining the training data set, data augmentation techniques can be adopted, such as amplifying relevant data instances, to solve the problem of data imbalance.

[0061] Furthermore, a data cleaning step can be introduced to further clean and screen the generated training data set, removing noise data and abnormal data to improve the quality of the training data set. For example, statistical methods can be used to detect and remove outliers in the data, or machine learning algorithms can be used to identify and remove possible error data.

[0062] In some embodiments, the feature extraction is used to extract multiple characteristic features of the drug, including at least one of the following: chemical structure feature, pharmacological action feature, metabolic pathway feature, and side effect feature, for multi-feature fusion prediction.

[0063] It should be noted that the feature extraction module mentioned in the present invention is used to extract multiple characteristic features of the drug, and these features include chemical structure features, pharmacological action features, metabolic pathway features, side effect features, etc. The extraction of these features is for multi-feature fusion prediction to more comprehensively analyze the interactions between drugs. By fusing multiple features, the accuracy and reliability of drug interaction prediction can be improved because the interaction between drugs is a complex biochemical process involving multiple factors. This multi-feature fusion method can better capture the potential interaction relationships between drugs.

[0064] Specifically, the chemical structure feature refers to the chemical composition and structure information of the drug molecule, such as molecular formula, functional group, molecular weight, etc. The pharmacological action feature refers to the mechanism of action of the drug in the organism, such as the action target and pharmacological effect of the drug. The metabolic pathway feature refers to the metabolic process of the drug in the organism, such as metabolic enzymes and metabolites. The side effect feature refers to the adverse reactions that the drug may cause, such as allergic reactions and toxic reactions. In the feature extraction process, specific parameters and extraction methods can be set for each feature. For example, for chemical structure features, chemical informatics tools can be used to extract the topological structure and chemical bond information of the molecule; for pharmacological action features, pharmacological databases can be referred to obtain the mechanism of action and target information of the drug; for metabolic pathway features, metabolomics data can be used to analyze the metabolic process of the drug; for side effect features, relevant information can be extracted by combining clinical data and drug instructions. These feature extraction methods can be selected and adjusted according to specific application scenarios and data sources.

[0065] Preferably, during the feature extraction process, the operation steps can be further refined. For example, for the extraction of chemical structure features, molecular fingerprint technology can be introduced to convert complex chemical structures into digital fingerprints that are easy to process, thereby improving the efficiency and accuracy of feature extraction. For pharmacological effect features, machine learning algorithms can be combined to automatically identify and extract pharmacological effect features related to drug interactions. For metabolic pathway features, a dynamic metabolic model can be introduced to simulate the metabolic process of drugs in the body and extract more comprehensive metabolic features. For side effect features, text mining technology can be introduced to extract side effect information from a large number of clinical reports and literature.

[0066] Furthermore, a feature selection algorithm can be introduced to select the most valuable features for predicting drug interactions from the extracted multiple features, so as to improve the performance and efficiency of the model. For example, a feature selection method based on information gain can be used to select those features that have a significant impact on the prediction results.

[0067] The above-mentioned various embodiments of the present invention have the following beneficial effects: This technical solution can improve the accuracy of predicting drug interactions. By combining multi-dimensional feature extraction with a pre-trained model, it is possible to comprehensively analyze key dimensions such as the chemical structure, pharmacological properties, and metabolic pathways of drugs. The data cleaning and quality detection mechanism can automatically screen high-quality samples and eliminate abnormal data to ensure the reliability of the input information, thereby reducing prediction errors and providing a more scientific decision-making basis for drug research and development and clinical medication.

[0068] Through activity detection and dynamic tracking analysis, this method can intelligently identify target activity data and determine its effective range, further improving the data screening efficiency. The collaborative effect of the global and local prediction sub-modules can simultaneously capture the correlation between the overall characteristics and local features of drugs, enhancing the generalization ability of the model. In addition, the automated data quality assessment and feedback mechanism can continuously optimize the sample set to ensure the stability of the prediction results, and ultimately provide intelligent support for medical research and personalized medication.

[0069] Further, the storage medium of the embodiments of the present application stores program instructions capable of implementing all the above methods. Among them, the program instructions can be stored in the above storage medium in the form of a software product, including several instructions to enable a computer device (which can be a personal computer, a server, or a network device, etc.) or a processor to execute all or part of the steps of the methods described in various embodiments of the present application. The aforementioned storage medium includes: various media that can store program codes such as USB flash drives, mobile hard disks, read-only memories (ROMs), random access memories (RAMs), magnetic disks, or optical discs, or terminal devices such as computers, servers, mobile phones, and tablets.

[0070] The above description is only some preferred embodiments of the present invention and an explanation of the applied technical principles. Those skilled in the art should understand that the scope of the invention involved in the embodiments of the present invention is not limited to the technical solutions formed by the specific combination of the above technical features, but should also cover other technical solutions formed by any combination of the above technical features or their equivalent features without departing from the above inventive concept. For example, the technical solutions formed by mutually replacing the above features with the technical features (but not limited to) having similar functions disclosed in the embodiments of the present invention.

Claims

1. A drug interaction prediction method based on multidimensional features, applied to a computing device, characterized in that: The method comprises: Collecting key data samples in the drug data set, and generating a key data sample set from the key data samples; Sending the key data sample set to a feature extraction module in a pre-trained drug interaction prediction model to extract features and obtain multi-dimensional feature data, wherein the feature extraction module of the drug interaction prediction model is pre-deployed on a computing device; The multidimensional feature data is sent to a server, so that a prediction module corresponding to the drug interaction prediction model in the server can predict drug interactions and obtain drug interaction prediction results; In response to receiving the drug interaction prediction result returned by the server, the result is displayed on the interface of the computing device.

2. The method according to claim 1, characterized in that Before sending the key data sample set to the pre-trained drug interaction prediction model, the following steps are performed: data cleaning processing is performed on the key data samples to remove abnormal data information therein.

3. The method according to claim 1, characterized in that The method also includes: performing data quality detection on each group of data in the drug data set to generate a data quality detection result set; determining the data quality detection results in the data quality detection result set that meet preset conditions, and determining the drug data corresponding to the data quality detection results that meet the preset conditions as key data samples.

4. The method according to claim 1 or 3, characterized in that: The method also includes: performing activity detection on the drug data set to determine target activity data in the drug data set; tracking and analyzing the detected target activity data, wherein the identified target activity data is tracked to obtain a tracking identification set; determining the data range of the detected target activity data in the drug data set to obtain an activity data range set; and based on the tracking identification set and the activity data range set, performing data quality detection on the drug data corresponding to each detected target activity data to generate a data quality detection result set.

5. The method according to claim 3, characterized in that: The method further includes: in response to determining that the data quality detection result set does not contain a data quality detection result that meets a preset condition, again selecting key data samples for the drug data collected by the data acquisition device.

6. The method according to claim 4, characterized in that The method of performing data quality detection on the drug data corresponding to each detected target active data based on the tracking identifier set and the active data range set to generate a data quality detection result set includes: determining, according to the active data range set, an integrity indicator set and a data accuracy indicator set of the drug data corresponding to each tracking identifier in the tracking identifier set; performing drug property analysis on the drug data corresponding to each active data range in the active data range set to generate a data property information set, wherein each data property information in the data property information set corresponds to a drug property, each data property information includes a classification label of the drug property in the key data sample, and the classification label includes at least one of the following: a chemical structure label, a pharmacological action label or a metabolic pathway label; and determining, according to the integrity indicator set, the data accuracy indicator set and the data property information set, a data quality detection result set corresponding to each tracking identifier.

7. The method according to claim 6, characterized in that The method further includes: determining the integrity index of the drug data corresponding to each tracking identifier in the tracking identifier set to obtain an integrity index set; and determining the accuracy index of the drug data corresponding to each tracking identifier in the tracking identifier set to obtain a data accuracy index set.

8. The method according to claim 1, characterized in that The drug interaction prediction model is trained and generated by a preset model generation method; wherein the preset model generation method comprises: obtaining a training data set, wherein the training samples in the training data set comprise relevant data instances, irrelevant data instances and corresponding relevant labels and irrelevant labels; selecting training samples from the training data set and inputting them into a feature extraction module of an initial drug interaction prediction model to generate relevant feature data and irrelevant feature data, wherein the model parameters of the feature extraction module in the initial drug interaction prediction model are fixed, and the prediction module in the initial drug interaction prediction model comprises a global prediction submodule and a local prediction submodule, and the prediction module corresponds to a single drug combination; The relevant feature data are input into the prediction module of the initial drug interaction prediction model, and relevant prediction labels are output, wherein the local prediction submodule in the prediction module is used to predict the local features of the drugs in the relevant feature data; the irrelevant feature data are input into the prediction module of the initial drug interaction prediction model, and irrelevant prediction labels are output, wherein the global prediction submodule in the prediction module is used to predict the overall features of the drugs in the relevant feature data; a sample error metric is generated based on the obtained relevant prediction labels, the irrelevant prediction labels, the relevant labels and the irrelevant labels; and the network parameters in the prediction module are adjusted based on the sample error metric.

9. The method according to claim 8, characterized in that The training data set is generated by the following steps: receiving a user's removal operation on drug data in a displayed drug interaction data set; determining the drug data corresponding to the removal operation as an irrelevant data instance to obtain an irrelevant data instance group, wherein the irrelevant data instance is irrelevant data removed by the user; determining the tracking identifier corresponding to each irrelevant data in the irrelevant data instance group as an irrelevant label to obtain an irrelevant label group; receiving a user's association operation on drug data in a displayed drug interaction data set; adjusting the tracking identifier corresponding to the association operation to the same identifier, and determining each drug data in the drug interaction data set as a related data instance to obtain a related data instance group; determining the tracking identifier corresponding to each related data instance in the related data instance group as a related label to obtain a related label group; combining the related data instance group, the related label group, the irrelevant data instance group and the irrelevant label group to generate a training data set.

10. The method according to claim 1, characterized in that The feature extraction is used to extract multiple characteristic features of the drug, including at least one of the following: chemical structure features, pharmacological action features, metabolic pathway features and side effect features, so as to perform multi-feature fusion prediction.

Citation Information

Patent Citations

  • Drug interaction effect prediction method based on pre-training model and molecular map

    CN114882970A

  • Drug interaction prediction method based on multi-dimensional features

    CN119230128A

  • Anesthesia assessment big data intelligent supervision system

    CN119673473A

  • AI-powered system for predicting drug interactions for personalized medication management

    DE202024104425U1

  • Machine learning-based system for predicting pharmacokinetics and pharmacodynamics

    DE202024104426U1