Artificial intelligence-based drug efficacy prediction method and related device

By employing an AI-based drug efficacy prediction method that utilizes time-series data and signaling pathway analysis, the challenge of predicting the responsiveness of TNF antagonists has been solved, enabling the development of precise treatment plans for autoimmune diseases and improving the accuracy of drug efficacy prediction.

CN116524995BActive Publication Date: 2026-06-02PING AN TECH (SHENZHEN) CO LTD

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
PING AN TECH (SHENZHEN) CO LTD
Filing Date
2023-03-24
Publication Date
2026-06-02

AI Technical Summary

Technical Problem

Current technology cannot effectively predict the responsiveness of biologics such as TNF antagonists, which makes it impossible for doctors to develop accurate medical plans for patients with autoimmune diseases, delaying diagnosis and leading to the loss of treatment intervention.

Method used

An AI-based drug efficacy prediction method is adopted. By acquiring time series data from multiple time points of the user, feature extraction and signal pathway analysis are performed. The time series prediction model is used to predict drug efficacy, and graph neural networks and heterogeneous medical knowledge graphs are combined for prediction.

Benefits of technology

It improves the accuracy of drug efficacy prediction, avoids the loss of diagnostic and treatment intervention periods, and achieves accurate prediction of autoimmune diseases.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116524995B_ABST
    Figure CN116524995B_ABST
Patent Text Reader

Abstract

The application relates to the field of digital medical technology, and provides a drug efficacy prediction method based on artificial intelligence and related equipment.The method comprises the following steps: performing feature extraction on an original data set to obtain a first target feature set; performing analysis on a second target feature set and a third target feature set of each time node respectively to obtain an analysis result of a first signal path of the corresponding time node; and inputting the analysis result of the first signal path of each time node, the second target feature set, the third target feature set and a corresponding drug efficacy result into a pre-trained time sequence prediction model to obtain a drug efficacy prediction result of each time node.The drug efficacy at a subsequent medication time node is predicted in advance, so that doctors can more reasonably formulate a medical scheme for patients, and the working efficiency of the doctors is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of digital medical technology, specifically to a method and related equipment for predicting drug efficacy based on artificial intelligence. Background Technology

[0002] Patients with autoimmune diseases respond significantly to early treatment with biologics. However, these diseases are often diagnosed years after onset, leading to many patients missing the optimal treatment window. Currently, there is no effective cure for autoimmune diseases, although biologics such as TNF antagonists are effective in reducing inflammation and pain.

[0003] However, before or in the early stages of treatment with biologics such as TNF antagonists, no drugs can be found to predict the responsiveness of these biologics, making it difficult for doctors to accurately prescribe treatment plans for patients. Summary of the Invention

[0004] In view of the above, it is necessary to propose a drug efficacy prediction method and related equipment based on artificial intelligence, which can predict the efficacy of drugs at subsequent medication time points in advance, so that doctors can formulate more reasonable medical plans for patients and improve doctors' work efficiency.

[0005] A first aspect of the present invention provides an artificial intelligence-based method for predicting drug efficacy, the method comprising:

[0006] Obtain the user's original dataset, which includes multiple time-series datasets at multiple time points from before medication to the medication process;

[0007] Feature extraction is performed on the original dataset to obtain a first target feature set, wherein the first target feature set includes a first common component feature set and a first key component feature set for each time node;

[0008] A second target feature set is extracted from the first common component feature set of each time node, and a third target feature set is extracted from the first key component feature set of each time node;

[0009] The second target feature set and the third target feature set at each time node are analyzed respectively to obtain the analysis results of the first signal path at the corresponding time node;

[0010] The analysis results of the first signal pathway, the second target feature set, the third target feature set, and the corresponding drug efficacy results at each time point are input into a pre-trained time series prediction model to obtain the drug efficacy prediction results at each time point.

[0011] Optionally, the analysis of the second target feature set and the third target feature set at each time point to obtain the analysis results of the first signal path at the corresponding time point includes:

[0012] Enrichment analysis is performed on the second target feature set and the third target feature set to obtain the enrichment analysis results;

[0013] Pathway analysis was performed on the enrichment analysis results to obtain the first signal pathway;

[0014] Cluster analysis was performed on the first signal path to obtain the analysis results of the first signal path at the corresponding time node.

[0015] Optionally, the original dataset includes: gene expression data, miRNA expression data, DNA methylation data, proteomics data, protein modification omics data, metabolomics data, and gut microbiota 16sRNA data.

[0016] Optionally, before inputting the analysis results of the first signal pathway, the second target feature set, the third target feature set, and the corresponding drug efficacy results at each time point into a pre-trained time series prediction model to obtain the drug efficacy prediction results at each time point, the method further includes:

[0017] Obtain historical sequence datasets from multiple users, wherein each user's historical sequence dataset includes historical sequence data from multiple time points from before medication to the medication process;

[0018] Feature extraction is performed on the historical sequence dataset to obtain a first historical feature set, wherein the first historical feature set includes a common component feature set and a key component feature set for each time node;

[0019] A second historical feature set is extracted from the common component feature set of each time node, and a third historical feature set is extracted from the key component feature set of each time node;

[0020] The second historical feature set and the third historical feature set for each time point are analyzed respectively to obtain the analysis results of the historical signal path;

[0021] The analysis results of multiple historical signal pathways at multiple time points of the multiple users, multiple second historical feature sets, multiple third historical feature sets, and corresponding drug efficacy results are used as training sets.

[0022] A preset graph neural network is trained based on the training set to obtain a time series prediction model.

[0023] Optionally, the step of extracting features from the original dataset to obtain the first target feature set includes:

[0024] The original data in the original dataset is standardized to obtain a standardized matrix;

[0025] Calculate the average value of the sample data for each sample in the normalization matrix;

[0026] The target matrix is ​​obtained by subtracting the corresponding mean from the sample data of each sample.

[0027] Calculate the covariance matrix of the target matrix, and the eigenvalues ​​and eigenvectors of the covariance matrix;

[0028] The eigenvalues ​​are sorted in descending order, and the top-ranked eigenvalues ​​are selected from the sorting results. The eigenvectors of the eigenvalues ​​are then used as row vectors to form a new eigenvector matrix.

[0029] Projecting the target matrix onto the new feature vector matrix yields the first target feature set.

[0030] Optionally, the method further includes:

[0031] Extract the user's dataset prior to the onset of illness from the original dataset;

[0032] Feature extraction is performed on the dataset to obtain a fourth target feature set, wherein the fourth target feature set includes a second common component feature set and a second key component feature set;

[0033] Extract the fifth target feature set from the fourth target feature set;

[0034] Cluster analysis is performed on the fifth target feature set to obtain the second signal pathway;

[0035] The dataset and the second signaling pathway are input into a pre-trained heterogeneous graph neural network model to obtain the user's disease prediction results.

[0036] Optionally, before inputting the dataset and the second signaling pathway into a pre-trained heterogeneous graph neural network model to obtain the user's disease prediction result, the method further includes:

[0037] Obtain data from multiple users, including each user's historical data prior to the onset of illness;

[0038] Obtain the heterogeneous graph corresponding to the heterogeneous medical knowledge graph of the historical data;

[0039] Obtain the signal pathways that interact with the heterogeneous graph;

[0040] The training set is composed of multiple historical data and signal paths of the multiple users.

[0041] The pre-trained heterogeneous graph neural network model is trained using the training set to obtain the heterogeneous graph neural network model.

[0042] A second aspect of the present invention provides an artificial intelligence-based drug efficacy prediction device, the device comprising:

[0043] The acquisition module is used to acquire the user's original dataset, which includes multiple time series datasets at multiple time points from before the user takes medication to the medication process.

[0044] The first extraction module is used to extract features from the original dataset to obtain a first target feature set, wherein the first target feature set includes a first common component feature set and a first key component feature set for each time node;

[0045] The second extraction module is used to extract a second target feature set from the first common component feature set of each time node, and to extract a third target feature set from the first key component feature set of each time node;

[0046] The analysis module is used to analyze the second target feature set and the third target feature set of each time node respectively, and obtain the analysis results of the first signal path of the corresponding time node;

[0047] The input module is used to input the analysis results of the first signal pathway, the second target feature set, the third target feature set, and the corresponding drug efficacy results of each time node into the pre-trained time series prediction model to obtain the drug efficacy prediction results of each time node.

[0048] A third aspect of the present invention provides an electronic device comprising a processor and a memory, wherein the processor is configured to execute a computer program stored in the memory to implement the aforementioned artificial intelligence-based drug efficacy prediction method.

[0049] A fourth aspect of the present invention provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the aforementioned artificial intelligence-based drug efficacy prediction method.

[0050] In summary, the AI-based drug efficacy prediction method and related equipment described in this invention improves the accuracy of drug efficacy prediction by extracting a second target feature set from the first common component feature set at each time point and a third target feature set from the first key component feature set at each time point. This considers the second and third target feature sets, which have a significant impact on drug efficacy prediction. The second and third target feature sets at each time point are analyzed separately to obtain the analysis results of the first signaling pathway at the corresponding time point. The analysis results of the first signaling pathway, the second and third target feature sets, and the corresponding drug efficacy results at each time point are input into a pre-trained time series prediction model to obtain the drug efficacy prediction result for each time point. By considering multiple dimensions in the process of obtaining the drug efficacy prediction result, the accuracy of the drug efficacy prediction result is improved. Attached Figure Description

[0051] Figure 1 This is a flowchart of the artificial intelligence-based drug efficacy prediction method provided in Embodiment 1 of the present invention.

[0052] Figure 2 This is a structural diagram of the artificial intelligence-based drug efficacy prediction device provided in Embodiment 2 of the present invention.

[0053] Figure 3 This is a schematic diagram of the structure of the electronic device provided in Embodiment 3 of the present invention. Detailed Implementation

[0054] To better understand the above-mentioned objects, features, and advantages of the present invention, the present invention will be described in detail below with reference to the accompanying drawings and specific embodiments. It should be noted that, unless otherwise specified, the embodiments of the present invention and the features thereof can be combined with each other.

[0055] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this invention pertains. The terminology used herein in the description of the invention is for the purpose of describing particular embodiments only and is not intended to be limiting of the invention.

[0056] Example 1

[0057] Figure 1 This is a flowchart of the artificial intelligence-based drug efficacy prediction method provided in Embodiment 1 of the present invention.

[0058] In this embodiment, the AI-based drug efficacy prediction method can be applied to electronic devices. For electronic devices that need to perform AI-based drug efficacy prediction, the AI-based drug efficacy prediction function provided by the method of this invention can be directly integrated into the electronic device, or it can run in the electronic device in the form of a software development kit (SDK).

[0059] The embodiments of this invention can acquire and process relevant data based on artificial intelligence technology. Artificial intelligence (AI) refers to the theories, methods, technologies, and application systems that utilize digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use that knowledge to obtain optimal results.

[0060] Foundational technologies for artificial intelligence generally include sensors, dedicated AI chips, cloud computing, distributed storage, big data processing, operating / interactive systems, and mechatronics. AI software technologies mainly encompass computer vision, robotics, biometrics, speech processing, natural language processing, as well as machine learning and deep learning.

[0061] like Figure 1 As shown, the artificial intelligence-based drug efficacy prediction method specifically includes the following steps. Depending on different needs, the order of the steps in this flowchart can be changed, and some steps can be omitted.

[0062] 101. Obtain the user's original dataset, which includes multiple time series datasets for multiple time points from before the user takes medication to the medication process.

[0063] In this embodiment, within the field of digital healthcare, patients with autoimmune diseases typically receive a diagnosis 6-10 years after the onset of the disease. This delayed diagnosis can lead to patients missing the optimal treatment intervention period and resulting in irreversible bone damage. When an autoimmune disease occurs, the efficacy of medication can be predicted in advance based on multiple time-series datasets from various time points before and during medication administration. This avoids missing the optimal treatment intervention period and achieves accurate prediction of autoimmune diseases.

[0064] In this embodiment, a drug efficacy prediction request sent by a user terminal is received, the received drug efficacy prediction request is parsed, a request message is obtained, a patient identification code is obtained from the request message, and the user's original dataset is obtained from a preset data source based on the patient identification code.

[0065] Specifically, the original dataset includes: gene expression data, miRNA expression data, DNA methylation data, proteomics data, protein modification omics data, metabolomics data, and gut microbiota 16sRNA data.

[0066] In this embodiment, the patient identification code is used to uniquely identify the patient's identity, and the preset data source can be a platform that records the patient's medication and medical treatment process.

[0067] 102. Perform feature extraction on the original dataset to obtain a first target feature set, wherein the first target feature set includes a first common component feature set and a first key component feature set for each time node.

[0068] In this embodiment, the first common component refers to the characteristics of the same components in the original dataset, and the first key component refers to the characteristics of the unique components in the original dataset.

[0069] In this embodiment, Principal Component Analysis (PCA) can be used to analyze the principal components in the original dataset to obtain the first target feature set. PCA refers to sequentially finding a set of mutually orthogonal coordinate axes in the original space, mapping n-dimensional features onto k dimensions. These k dimensions are entirely new orthogonal features, also known as principal components, which are reconstructed based on the original n-dimensional features. These k-dimensional features are then determined as the first target feature set.

[0070] In an optional embodiment, the step of extracting features from the original dataset to obtain the first target feature set includes:

[0071] The original data in the original dataset is standardized to obtain a standardized matrix;

[0072] Calculate the average value of the sample data for each sample in the normalization matrix;

[0073] The target matrix is ​​obtained by subtracting the corresponding mean from the sample data of each sample.

[0074] Calculate the covariance matrix of the target matrix, and the eigenvalues ​​and eigenvectors of the covariance matrix;

[0075] The eigenvalues ​​are sorted in descending order, and the top-ranked eigenvalues ​​are selected from the sorting results. The eigenvectors of the eigenvalues ​​are then used as row vectors to form a new eigenvector matrix.

[0076] Projecting the target matrix onto the new feature vector matrix yields the first target feature set.

[0077] In this embodiment, since the original dataset contains multiple omics data, the covariance matrix of the target matrix obtained by calculation satisfies two-dimensional features.

[0078] In this embodiment, the eigenvalues ​​and eigenvectors of the covariance matrix can be calculated using the eigenvalue decomposition method.

[0079] In this embodiment, the first common component feature set and the first key component feature set for each time node are extracted from the original dataset to ensure that the components are orthogonal, that is, there is no information redundancy between the components. The feature set is characterized to the greatest extent with as few components as possible, reducing the number of feature sets. At the same time, the first common component feature set and the first key component feature set are considered when predicting drug efficacy in the future, thereby improving the efficiency of drug efficacy prediction.

[0080] 103. Extract a second target feature set from the first common component feature set of each time node, and extract a third target feature set from the first key component feature set of each time node.

[0081] In this embodiment, features that have a significant impact on the efficacy of drugs can be pre-defined for each autoimmune disease. The second target feature set and the third target feature set refer to features that have a significant impact on the efficacy of drugs.

[0082] In this embodiment, features that have a significant impact on drug efficacy prediction are extracted from the first common component feature set and the first key component feature set at each time point. When predicting drug efficacy, the second target feature set and the third target feature set that have a significant impact on drug efficacy prediction are considered, thereby improving the accuracy of drug efficacy prediction.

[0083] 104. The second target feature set and the third target feature set of each time node are analyzed respectively to obtain the analysis results of the first signal path of the corresponding time node.

[0084] In this embodiment, the signaling pathway refers to an intracellular signal transduction pathway closely related to cytokines, which participates in many important physiological processes such as cell proliferation, differentiation, apoptosis, and immune regulation. If the first signaling pathway is abnormally activated, it may lead to viral infection, decreased antibacterial immune function, and aggravated inflammation in patients. By analyzing the signaling pathways of the second and third target feature sets at each time point, it is possible to predict in advance whether the first signaling pathway is abnormally activated, and to avoid problems such as viral infection, decreased antibacterial immune function, and aggravated inflammation in patients caused by the abnormal activation of the first signaling pathway.

[0085] In an optional embodiment, the analysis of the second target feature set and the third target feature set at each time point to obtain the analysis results of the first signal path at the corresponding time point includes:

[0086] Enrichment analysis is performed on the second target feature set and the third target feature set to obtain the enrichment analysis results;

[0087] Pathway analysis was performed on the enrichment analysis results to obtain the first signal pathway;

[0088] Cluster analysis was performed on the first signal path to obtain the analysis results of the first signal path at the corresponding time node.

[0089] In this embodiment, enrichment refers to the process of classifying the target features in the second target feature set and the third target feature set according to the genomic annotation information corresponding to the target features. After classification, the commonalities of the found target features can be identified, where the commonalities can be the functions, compositions, etc. among the target features.

[0090] In this embodiment, before performing enrichment analysis on the second target feature set and the third target feature set, a gene annotation database corresponding to the target features is pre-constructed. Using a preset algorithm, the second target feature set and the third target feature set are classified according to the annotations in the gene annotation database. The classification results are clustered, and redundant results are removed to obtain enrichment analysis results. The enrichment classification results include differential gene expression analysis, differential genes are screened out, pathway analysis is performed on the differential genes to obtain the first signaling pathway, and cluster analysis is performed on the first signaling pathway.

[0091] In this embodiment, cluster analysis refers to the analysis process of grouping a set of physical or abstract objects into multiple classes composed of similar objects. By performing cluster analysis on the first signal pathway, the analysis results of the first signal pathway at the corresponding time node are obtained, that is, the cluster analysis results of the signal pathway obtained by integrating multi-omics data. The analysis results include the abnormal activation probability value of the first signal pathway.

[0092] In this embodiment, the gene annotation database refers to the database obtained by annotating gene functions from the perspective of different biological models and storing the data.

[0093] In this embodiment, by analyzing the second and third target feature sets at different time points, the abnormal activation probability values ​​of the signal pathways at the corresponding time points are obtained. This avoids the problem of inaccurate drug prediction caused by using a feature set of a single time point for drug efficacy prediction, and improves the accuracy of drug efficacy prediction.

[0094] 105. The analysis results of the first signal pathway, the second target feature set, the third target feature set, and the corresponding drug efficacy results at each time point are input into the pre-trained time series prediction model to obtain the drug efficacy prediction results at each time point.

[0095] In this embodiment, a time series prediction model can be pre-trained, which can be used to predict the efficacy of the drug at each time point.

[0096] For example, if a user is taking a TNF biologic, multiple time-series datasets are obtained at various time points during the user's medication process, such as datasets for the first and second treatment cycles. These datasets are then adjusted, extracted, and analyzed. The analysis results of the first signaling pathway, the second target feature set, the third target feature set, and the corresponding drug efficacy results for each time point are input into a pre-trained time-series prediction model to obtain the drug efficacy prediction result for each time point. This allows for the prediction of the TNF biologic's efficacy; for example, it can be predicted that completing the second course of TNF biologic will effectively reduce inflammation and pain.

[0097] In an optional embodiment, before inputting the analysis results of the first signaling pathway, the second target feature set, the third target feature set, and the corresponding drug efficacy results at each time point into a pre-trained time series prediction model to obtain the drug efficacy prediction results at each time point, the method further includes:

[0098] Obtain historical sequence datasets from multiple users, wherein each user's historical sequence dataset includes historical sequence data from multiple time points from before medication to the medication process;

[0099] Feature extraction is performed on the historical sequence dataset to obtain a first historical feature set, wherein the first historical feature set includes a common component feature set and a key component feature set for each time node;

[0100] A second historical feature set is extracted from the common component feature set of each time node, and a third historical feature set is extracted from the key component feature set of each time node;

[0101] The second historical feature set and the third historical feature set for each time point are analyzed respectively to obtain the analysis results of the historical signal path;

[0102] The analysis results of multiple historical signal pathways at multiple time points of the multiple users, multiple second historical feature sets, multiple third historical feature sets, and corresponding drug efficacy results are used as training sets.

[0103] A preset graph neural network is trained based on the training set to obtain a time series prediction model.

[0104] In this embodiment, the historical sequence dataset contains historical sequence data from multiple time points from before the user takes medication to the medication process. For example, the historical dataset includes gene expression data, miRNA expression data, DNA methylation data, proteomics data, protein modification proteomics data, metabolomics data, and gut microbiota 16sRNA data. Regarding DNA methylation data, the first course of treatment data in the medication process is the methylation level of CpG sites, which is M = log2(M / U), where M = 2.4 and U = 2.5.

[0105] In this embodiment, after obtaining the historical sequence dataset, features are extracted from the historical dataset, and a graph neural network is trained based on the feature extraction results to obtain a time series prediction model.

[0106] In this embodiment, the graph neural network is a type of deep learning method capable of predicting nodes, edges, or graphs. By training a pre-defined graph neural network based on a training set, it is possible to predict the differences in clinical phenotypes at different time points, and based on these differences, predict the efficacy of drugs at subsequent medication time points.

[0107] In this embodiment, the analysis results of the first signal pathway, the second target feature set, the third target feature set, and the corresponding drug efficacy results at each time point are input into a pre-trained time series prediction model to obtain the drug efficacy prediction results at each time point. In the process of obtaining the drug efficacy prediction results, multiple dimensions are considered, which improves the accuracy of the drug efficacy prediction results.

[0108] In other alternative embodiments, the method further includes:

[0109] Extract the user's dataset prior to the onset of illness from the original dataset;

[0110] Feature extraction is performed on the dataset to obtain a fourth target feature set, wherein the fourth target feature set includes a second common component feature set and a second key component feature set;

[0111] Extract the fifth target feature set from the fourth target feature set;

[0112] Cluster analysis is performed on the fifth target feature set to obtain the second signal pathway;

[0113] The dataset and the second signaling pathway are input into a pre-trained heterogeneous graph neural network model to obtain the user's disease prediction results.

[0114] In this embodiment, the fifth target feature set includes a feature set extracted from the second common component feature set of the fourth target feature set and a feature set extracted from the second key component feature set of the fourth target feature set. The features in the fifth target feature set represent features that have a significant impact on the pathogenesis. By performing cluster analysis on these features, multiple classes composed of similar features are analyzed to obtain the second signaling pathway.

[0115] In other alternative embodiments, a heterogeneous graph neural network model can be pre-trained to predict the risk of developing autoimmune diseases. After obtaining the user's original dataset before the onset of the disease, the user's disease prediction result is obtained by inputting the dataset into the heterogeneous graph neural network model. The user's disease prediction result includes the predicted probability of developing the disease.

[0116] In an optional embodiment, before inputting the dataset and the second signaling pathway into a pre-trained heterogeneous graph neural network model to obtain the user's disease prediction result, the method further includes:

[0117] Obtain data from multiple users, including each user's historical data prior to the onset of illness;

[0118] Obtain the heterogeneous graph corresponding to the heterogeneous medical knowledge graph of the historical data;

[0119] Obtain the signal pathways that interact with the heterogeneous graph;

[0120] The training set is composed of multiple historical data and signal paths of the multiple users.

[0121] The pre-trained heterogeneous graph neural network model is trained using the training set to obtain the heterogeneous graph neural network model.

[0122] In this embodiment, a heterogeneous medical knowledge graph method is used to predict the risk of autoimmune diseases. Specifically, the known signaling pathway knowledge graph consists of N+1 heterogeneous graphs, where N represents the signaling pathways that interact with the heterogeneous graphs. For example, the heterogeneous graphs can include one or more of the following combinations: directed graphs of interactions between proteins; heterogeneous graphs of relationships between small molecule metabolites; heterogeneous graphs of relationships between gene expression; heterogeneous graphs of relationships between protein modifications; heterogeneous graphs of relationships between gut microbiota; and directed graphs of mutual regulation relationships between biological signaling pathways. Interactions exist between the N+1 heterogeneous graphs, meaning there are interactions between proteins, small molecule metabolites, genes, protein modification sites, gut microbiota, and biological signaling pathways. Each sub-graph in this heterogeneous medical knowledge graph is connected by edges to autoimmune diseases and the efficacy of TNF antagonists. Based on the user's pre-illness dataset (e.g., proteins, small molecules, genes, etc.) extracted from the original dataset and the features of the second signaling pathway, a heterogeneous graph neural network model is trained and data at different time points are distinguished. This enables the prediction of the risk of autoimmune disease and the efficacy of TNF antagonists at different time points through multi-omics data input.

[0123] In this embodiment, by using heterogeneous graphs corresponding to heterogeneous medical knowledge graphs for disease risk prediction, the performance of disease prediction is enhanced, the impact of insufficient data and data bias is compensated, the matching degree between the disease prediction results and clinical knowledge is improved, and the accuracy of disease prediction results is increased.

[0124] In summary, the AI-based drug efficacy prediction method described in this embodiment improves the accuracy of drug efficacy prediction by extracting a second target feature set from the first common component feature set at each time point and a third target feature set from the first key component feature set at each time point. This considers the second and third target feature sets, which have a significant impact on drug efficacy prediction. The second and third target feature sets at each time point are analyzed separately to obtain the analysis results of the first signaling pathway at the corresponding time point. The analysis results of the first signaling pathway, the second and third target feature sets, and the corresponding drug efficacy results at each time point are input into a pre-trained time series prediction model to obtain the drug efficacy prediction result for each time point. In obtaining the drug efficacy prediction result, multiple dimensions are considered, thus improving the accuracy of the drug efficacy prediction result.

[0125] Example 2

[0126] Figure 2 This is a structural diagram of the artificial intelligence-based drug efficacy prediction device provided in Embodiment 2 of the present invention.

[0127] In some embodiments, the AI-based drug efficacy prediction device 20 may include multiple functional modules composed of program code segments. The program code of each program segment in the AI-based drug efficacy prediction device 20 may be stored in the memory of an electronic device and executed by the at least one processor to perform (see details). Figure 1 (Description) Functionality for predicting drug efficacy based on artificial intelligence.

[0128] In this embodiment, the AI-based drug efficacy prediction device 20 can be divided into multiple functional modules according to its functions. These functional modules may include: an acquisition module 201, a first extraction module 202, a second extraction module 203, an analysis module 204, and an input module 205. The term "module" in this invention refers to a series of computer-readable instruction segments that can be executed by at least one processor and perform a fixed function, stored in memory. In this embodiment, the functions of each module will be detailed in subsequent embodiments.

[0129] The acquisition module 201 is used to acquire the user's original dataset, which includes multiple time series datasets at multiple time points from before the user takes medication to the medication process.

[0130] The first extraction module 202 is used to extract features from the original dataset to obtain a first target feature set, wherein the first target feature set includes a first common component feature set and a first key component feature set for each time node.

[0131] The second extraction module 203 is used to extract a second target feature set from the first common component feature set of each time node, and to extract a third target feature set from the first key component feature set of each time node.

[0132] The analysis module 204 is used to analyze the second target feature set and the third target feature set of each time node respectively, and obtain the analysis results of the first signal path of the corresponding time node.

[0133] The input module 205 is used to input the analysis results of the first signal pathway, the second target feature set, the third target feature set and the corresponding drug efficacy results of each time node into the pre-trained time series prediction model to obtain the drug efficacy prediction results of each time node.

[0134] In an optional embodiment, the first extraction module 202 is configured to: standardize the original data in the original dataset to obtain a standardized matrix; calculate the mean of the sample data for each sample in the standardized matrix; subtract the corresponding mean from the sample data for each sample to obtain a target matrix; calculate the covariance matrix of the target matrix, and the eigenvalues ​​and eigenvectors of the covariance matrix; sort the eigenvalues ​​in descending order, select the top-ranked eigenvalues ​​from the sorting results, and use the eigenvectors of the eigenvalues ​​as row vectors to form a new eigenvector matrix; project the target matrix onto the new eigenvector matrix to obtain a first target feature set.

[0135] In this embodiment, the first common component feature set and the first key component feature set for each time node are extracted from the original dataset to ensure that the components are orthogonal, that is, there is no information redundancy between the components. The feature set is characterized to the greatest extent with as few components as possible, reducing the number of feature sets. At the same time, the first common component feature set and the first key component feature set are considered when predicting drug efficacy in the future, thereby improving the efficiency of drug efficacy prediction.

[0136] In an optional embodiment, the analysis module 204 is configured to: perform enrichment analysis on the second target feature set and the third target feature set to obtain enrichment analysis results; perform path analysis on the enrichment analysis results to obtain a first signal path; and perform cluster analysis on the first signal path to obtain the analysis results of the first signal path at the corresponding time node.

[0137] In an optional embodiment, before inputting the analysis results of the first signal pathway, the second target feature set, the third target feature set, and the corresponding drug efficacy results of each time point into a pre-trained time series prediction model to obtain the drug efficacy prediction results of each time point, a historical sequence dataset of multiple users is obtained. Each user's historical sequence dataset includes historical sequence data from multiple time points from before medication to the medication process. Feature extraction is performed on the historical sequence dataset to obtain a first historical feature set, which includes a common component feature set and a key component feature set for each time point. A second historical feature set is extracted from the common component feature set of each time point, and a third historical feature set is extracted from the key component feature set of each time point. The second historical feature set and the third historical feature set of each time point are analyzed to obtain the analysis results of the historical signal pathway. The analysis results of the multiple historical signal pathways of multiple time points of multiple users, the multiple second historical feature sets, the multiple third historical feature sets, and the corresponding drug efficacy results are used as a training set. A pre-set graph neural network is trained based on the training set to obtain the time series prediction model.

[0138] In this embodiment, the graph neural network is a type of deep learning method capable of predicting nodes, edges, or graphs. By training a pre-defined graph neural network based on a training set, it is possible to predict the differences in clinical phenotypes at different time points, and based on these differences, predict the efficacy of drugs at subsequent medication time points.

[0139] In this embodiment, the analysis results of the first signal pathway, the second target feature set, the third target feature set, and the corresponding drug efficacy results at each time point are input into a pre-trained time series prediction model to obtain the drug efficacy prediction results at each time point. In the process of obtaining the drug efficacy prediction results, multiple dimensions are considered, which improves the accuracy of the drug efficacy prediction results.

[0140] In other optional embodiments, the user's pre-illness dataset is extracted from the original dataset; features are extracted from the dataset to obtain a fourth target feature set, wherein the fourth target feature set includes a second common component feature set and a second key component feature set; a fifth target feature set is extracted from the fourth target feature set; cluster analysis is performed on the fifth target feature set to obtain a second signaling pathway; the dataset and the second signaling pathway are input into a pre-trained heterogeneous graph neural network model to obtain the user's onset prediction result.

[0141] In other alternative embodiments, a heterogeneous graph neural network model can be pre-trained. After obtaining the original dataset of the user before the onset of the disease, the dataset can be input into the heterogeneous graph neural network model to obtain the prediction result of the user's onset of the disease.

[0142] In an optional embodiment, before inputting the dataset and the second signaling pathway into a pre-trained heterogeneous graph neural network model to obtain the user's disease prediction result, multiple users and their historical data before the onset of disease are acquired; a heterogeneous graph corresponding to the heterogeneous medical knowledge graph of the historical data is acquired; signaling pathways that interact with the heterogeneous graph are acquired; multiple historical data and signaling pathways of the multiple users are used as a training set; and the pre-trained heterogeneous graph neural network model is trained based on the training set to obtain the heterogeneous graph neural network model.

[0143] In this embodiment, by using heterogeneous graphs corresponding to heterogeneous medical knowledge graphs for disease risk prediction, the performance of disease prediction is enhanced, the impact of insufficient data and data bias is compensated, the matching degree between the disease prediction results and clinical knowledge is improved, and the accuracy of disease prediction results is increased.

[0144] In summary, the AI-based drug efficacy prediction device described in this embodiment improves the accuracy of drug efficacy prediction by extracting a second target feature set from the first common component feature set at each time point and a third target feature set from the first key component feature set at each time point. This considers the second and third target feature sets, which have a significant impact on drug efficacy prediction. The second and third target feature sets at each time point are analyzed separately to obtain the analysis results of the first signaling pathway at the corresponding time point. The analysis results of the first signaling pathway, the second and third target feature sets, and the corresponding drug efficacy results at each time point are input into a pre-trained time series prediction model to obtain the drug efficacy prediction result for each time point. In obtaining the drug efficacy prediction result, multiple dimensions are considered, thus improving the accuracy of the drug efficacy prediction result.

[0145] Example 3

[0146] See Figure 3 The diagram shown is a structural schematic of an electronic device provided in Embodiment 3 of the present invention. In a preferred embodiment of the present invention, the electronic device 3 includes a memory 31, at least one processor 32, at least one communication bus 33, and a transceiver 34.

[0147] Those skilled in the art should understand that Figure 3The structure of the electronic device shown does not constitute a limitation of the embodiments of the present invention. It can be a bus structure or a star structure. The electronic device 3 may also include more or fewer other hardware or software than shown, or different component arrangements.

[0148] In some embodiments, the electronic device 3 is an electronic device capable of automatically performing numerical calculations and / or information processing according to pre-set or stored instructions. Its hardware includes, but is not limited to, microprocessors, application-specific integrated circuits (ASICs), programmable gate arrays (FPGAs), digital processors, and embedded devices. The electronic device 3 may also include client devices, including, but not limited to, any electronic product capable of human-computer interaction with a client via a keyboard, mouse, remote control, touchpad, or voice control device, such as personal computers, tablet computers, smartphones, and digital cameras.

[0149] It should be noted that the electronic device 3 is merely an example. Other existing or future electronic products that are suitable for this invention should also be included within the scope of protection of this invention and are incorporated herein by reference.

[0150] In some embodiments, the memory 31 is used to store program code and various data, such as an AI-based drug efficacy prediction device 20 installed in the electronic device 3, and to achieve high-speed and automatic access to programs or data during the operation of the electronic device 3. The memory 31 includes read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), one-time programmable read-only memory (OTPROM), electrically erasable programmable read-only memory (EEPROM), compact disc read-only memory (CD-ROM) or other optical disc storage, disk storage, magnetic tape storage, or any other computer-readable medium capable of carrying or storing data.

[0151] In some embodiments, the at least one processor 32 may be composed of integrated circuits, such as a single packaged integrated circuit or multiple integrated circuits packaged with the same or different functions, including combinations of one or more central processing units (CPUs), microprocessors, digital processing chips, graphics processors, and various control chips. The at least one processor 32 is the control unit of the electronic device 3, connecting various components of the entire electronic device 3 via various interfaces and lines. It executes programs or modules stored in the memory 31 and calls data stored in the memory 31 to perform various functions and process data of the electronic device 3.

[0152] In some embodiments, the at least one communication bus 33 is configured to enable communication between the memory 31 and the at least one processor 32, etc.

[0153] Although not shown, the electronic device 3 may also include a power supply (such as a battery) to power the various components. Optionally, the power supply may be logically connected to the at least one processor 32 via a power management device, thereby enabling functions such as charging, discharging, and power consumption management. The power supply may also include one or more DC or AC power supplies, recharging devices, power fault detection circuits, power converters or inverters, power status indicators, and other arbitrary components. The electronic device 3 may also include various sensors, Bluetooth modules, Wi-Fi modules, etc., which will not be described in detail here.

[0154] It should be understood that the embodiments described are for illustrative purposes only and are not limited to this structure in the scope of the patent application.

[0155] The integrated unit implemented as a software functional module described above can be stored in a computer-readable storage medium. This software functional module, stored in a storage medium, includes several instructions to cause a computer device (which may be a personal computer, electronic device, or network device, etc.) or processor to execute portions of the methods described in the various embodiments of the present invention.

[0156] In a further embodiment, combined with Figure 2 The at least one processor 32 can execute the operating device of the electronic device 3 and various installed applications (such as the artificial intelligence-based drug efficacy prediction device 20), program code, etc., for example, the various modules mentioned above.

[0157] The memory 31 stores program code, and the at least one processor 32 can call the program code stored in the memory 31 to execute related functions. For example, Figure 2 The modules described herein are program codes stored in the memory 31 and executed by the at least one processor 32, thereby realizing the functions of the modules to achieve the purpose of predicting drug efficacy based on artificial intelligence.

[0158] For example, the program code can be divided into one or more modules / units, which are stored in the memory 31 and executed by the processor 32 to complete this application. The one or more modules / units can be a series of computer-readable instruction segments capable of performing a specific function, which describe the execution process of the program code in the electronic device 3. For example, the program code can be divided into an acquisition module 201, a first extraction module 202, a second extraction module 203, an analysis module 204, and an input module 205.

[0159] In one embodiment of the present invention, the memory 31 stores a plurality of computer-readable instructions, which are executed by the at least one processor 32 to achieve the function of predicting drug efficacy based on artificial intelligence.

[0160] Specifically, the specific implementation method of the above instructions by the at least one processor 32 can be referred to Figure 1 The descriptions of the relevant steps in the corresponding embodiments are not repeated here.

[0161] In the several embodiments provided by this invention, it should be understood that the disclosed apparatus and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of modules is only a logical functional division, and other division methods may be used in actual implementation.

[0162] The modules described as separate components may or may not be physically separate. The components shown as modules may or may not be physical units; they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs.

[0163] Furthermore, the functional modules in the various embodiments of the present invention can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or in the form of hardware plus software functional modules.

[0164] It will be apparent to those skilled in the art that the present invention is not limited to the details of the exemplary embodiments described above, and that the invention can be implemented in other specific forms without departing from the spirit or essential characteristics of the invention. Therefore, the embodiments should be considered illustrative and non-limiting in all respects, and the scope of the invention is defined by the appended claims rather than the foregoing description. Thus, all variations falling within the meaning and scope of equivalents of the claims are intended to be embraced within the present invention. No reference numerals in the claims should be construed as limiting the scope of the claims. Furthermore, it is clear that the word "comprising" does not exclude other elements, and the singular does not exclude the plural. Multiple elements or devices recited in the present invention may also be implemented by a single element or device in software or hardware. The terms "first," "second," etc., are used to denote names and do not indicate any particular order.

[0165] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit it. Although the present invention has been described in detail with reference to preferred embodiments, those skilled in the art should understand that modifications or equivalent substitutions can be made to the technical solutions of the present invention without departing from the spirit and scope of the technical solutions of the present invention.

Claims

1. A method for predicting drug efficacy based on artificial intelligence, characterized in that, The method includes: Obtain the user's original dataset, which includes multiple time-series datasets at multiple time points from before medication to the medication process; Feature extraction is performed on the original dataset to obtain a first target feature set, wherein the first target feature set includes a first common component feature set and a first key component feature set for each time node; A second target feature set is extracted from the first common component feature set of each time node, and a third target feature set is extracted from the first key component feature set of each time node; The second target feature set and the third target feature set at each time node are analyzed respectively to obtain the analysis results of the first signal path at the corresponding time node; The analysis results of the first signal pathway, the second target feature set, the third target feature set, and the corresponding drug efficacy results at each time point are input into a pre-trained time series prediction model to obtain the drug efficacy prediction result at each time point. The pre-training process of the time series prediction model includes: acquiring historical sequence datasets of multiple users, wherein each user's historical sequence dataset includes historical sequence data of multiple time points from before medication to the medication process; extracting features from the historical sequence dataset to obtain a first historical feature set, wherein the first historical feature set includes a common component feature set and a key component feature set for each time point; extracting a second historical feature set from the common component feature set for each time point, and extracting a third historical feature set from the key component feature set for each time point; analyzing the second historical feature set and the third historical feature set for each time point to obtain the analysis results of the historical signal pathway; using the analysis results of multiple historical signal pathways, multiple second historical feature sets, multiple third historical feature sets, and the corresponding drug efficacy results of multiple time points of multiple users as a training set; and training a pre-set graph neural network based on the training set to obtain the time series prediction model.

2. The artificial intelligence-based drug efficacy prediction method as described in claim 1, characterized in that, The analysis of the second target feature set and the third target feature set at each time node to obtain the analysis results of the first signal path at the corresponding time node includes: Enrichment analysis is performed on the second target feature set and the third target feature set to obtain the enrichment analysis results; Pathway analysis was performed on the enrichment analysis results to obtain the first signal pathway; Cluster analysis was performed on the first signal path to obtain the analysis results of the first signal path at the corresponding time node.

3. The drug efficacy prediction method based on artificial intelligence as described in claim 1, characterized in that, The original dataset includes: gene expression data, miRNA expression data, DNA methylation data, proteomics data, protein modification omics data, metabolomics data, and gut microbiota 16sRNA data.

4. The artificial intelligence-based drug efficacy prediction method as described in claim 1, characterized in that, The step of extracting features from the original dataset to obtain the first target feature set includes: The original data in the original dataset is standardized to obtain a standardized matrix; Calculate the average value of the sample data for each sample in the normalization matrix; The target matrix is ​​obtained by subtracting the corresponding mean from the sample data of each sample. Calculate the covariance matrix of the target matrix, and the eigenvalues ​​and eigenvectors of the covariance matrix; The eigenvalues ​​are sorted in descending order, and the top-ranked eigenvalues ​​are selected from the sorting results. The eigenvectors of the eigenvalues ​​are then used as row vectors to form a new eigenvector matrix. Projecting the target matrix onto the new feature vector matrix yields the first target feature set.

5. The artificial intelligence-based drug efficacy prediction method as described in claim 1, characterized in that, The method further includes: Extract the user's dataset prior to the onset of illness from the original dataset; Feature extraction is performed on the dataset to obtain a fourth target feature set, wherein the fourth target feature set includes a second common component feature set and a second key component feature set; Extract the fifth target feature set from the fourth target feature set; Cluster analysis is performed on the fifth target feature set to obtain the second signal pathway; The dataset and the second signaling pathway are input into a pre-trained heterogeneous graph neural network model to obtain the user's disease prediction results.

6. The artificial intelligence-based drug efficacy prediction method as described in claim 5, characterized in that, Before inputting the dataset and the second signaling pathway into a pre-trained heterogeneous graph neural network model to obtain the user's disease prediction result, the method further includes: Obtain data from multiple users, including each user's historical data prior to the onset of illness; Obtain the heterogeneous graph corresponding to the heterogeneous medical knowledge graph of the historical data; Obtain the signal pathways that interact with the heterogeneous graph; The training set is composed of multiple historical data and signal paths of the multiple users. The pre-trained heterogeneous graph neural network model is trained using the training set to obtain the heterogeneous graph neural network model.

7. A drug efficacy prediction device based on artificial intelligence, characterized in that, The device includes: The acquisition module is used to acquire the user's original dataset, which includes multiple time series datasets at multiple time points from before the user takes medication to the medication process. The first extraction module is used to extract features from the original dataset to obtain a first target feature set, wherein the first target feature set includes a first common component feature set and a first key component feature set for each time node; The second extraction module is used to extract a second target feature set from the first common component feature set of each time node, and to extract a third target feature set from the first key component feature set of each time node; The analysis module is used to analyze the second target feature set and the third target feature set of each time node respectively, and obtain the analysis results of the first signal path of the corresponding time node; An input module is used to input the analysis results of the first signal pathway, the second target feature set, the third target feature set, and the corresponding drug efficacy results of each time point into a pre-trained time series prediction model to obtain the drug efficacy prediction result for each time point. The pre-training process of the time series prediction model includes: acquiring historical sequence datasets of multiple users, wherein each user's historical sequence dataset includes historical sequence data of multiple time points from before medication to the medication process; extracting features from the historical sequence dataset to obtain a first historical feature set, wherein the first historical feature set includes a common component feature set and a key component feature set for each time point; extracting a second historical feature set from the common component feature set for each time point, and extracting a third historical feature set from the key component feature set for each time point; analyzing the second historical feature set and the third historical feature set for each time point to obtain the analysis results of the historical signal pathway; using the analysis results of multiple historical signal pathways, multiple second historical feature sets, multiple third historical feature sets, and the corresponding drug efficacy results of multiple time points of multiple users as a training set; and training a pre-set graph neural network based on the training set to obtain the time series prediction model.

8. An electronic device, characterized in that, The electronic device includes a processor and a memory, wherein the processor is used to execute a computer program stored in the memory to implement the artificial intelligence-based drug efficacy prediction method as described in any one of claims 1 to 6.

9. A computer-readable storage medium storing a computer program thereon, characterized in that, When the computer program is executed by the processor, it implements the artificial intelligence-based drug efficacy prediction method as described in any one of claims 1 to 6.