A pharmaceutical knowledge graph construction method and system based on artificial intelligence
Through the analysis and time series processing of drug molecule and protein feature sets, the flexibility and accuracy issues of pharmaceutical knowledge graphs under dynamic changes are solved, and the optimization of drug development and support for personalized treatment are achieved.
Patent Information
- Application Number
- CN202510390758.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-31
- Publication Date
- 2025-09-12
- Estimated Expiration
- 2045-03-31
AI Technical Summary
Existing pharmaceutical knowledge graph technology lacks flexibility and accuracy when processing real-time data and dynamic changes, making it difficult to update and reflect the latest medical information in real time. This results in the inability to make real-time adjustments and optimizations during drug development, and side effect predictions rely on historical data, lacking consideration of the dynamic impact on actual biological pathways.
Through database query, we collect drug molecular structures and their target protein sequences, perform numerical coding and feature extraction, analyze drug-protein interactions, predict drug resistance pathways, conduct time series analysis, identify drug side effects, draw occurrence maps, and generate drug effect correlation analysis results.
It significantly enhances the understanding of drug-protein interactions, dynamically tracks drug effects, optimizes treatment plans, provides data support for personalized medicine, and improves drug development efficiency and safety.
Smart Images

Figure CN120375910B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of pharmaceutical knowledge graph technology, and in particular to a pharmaceutical knowledge graph construction method and system based on artificial intelligence. Background Art
[0002] The field of pharmaceutical knowledge graph technology focuses on leveraging artificial intelligence (AI), particularly data science and machine learning algorithms, to integrate, analyze, and visualize large amounts of pharmaceutical-related data. This area involves building complex data models capable of representing and processing drug discovery, drug interactions, biomarkers, and the relationship between disease and treatment. These knowledge graphs encompass not only chemical and biological information but may also integrate data from clinical trials, patient health records, and scientific literature. By creating such graphs, researchers can quickly access and leverage interdisciplinary information to support the development of new drugs and the more efficient use of existing medications.
[0003] Among them, the AI-powered pharmaceutical knowledge graph construction method involves developing and applying algorithms and computational technologies to construct a graph encompassing pharmaceutical knowledge. This approach aims to improve the efficiency and accuracy of drug development and accelerate the transition from the laboratory to the clinic through intelligent data analysis. Specific applications include, but are not limited to, discovering new drug candidate molecules, predicting drug interactions, and optimizing treatment plans. Furthermore, this technology can help healthcare practitioners and researchers better understand complex pharmacological mechanisms and disease pathways, thereby promoting the development of personalized medicine and precision medicine.
[0004] Although existing pharmaceutical knowledge graph technologies provide powerful tools for data integration, analysis, and visualization, they often lack sufficient flexibility and accuracy when dealing with real-time data and dynamic changes. Existing technologies are inadequate in tracking the biological activity of drugs over time, resulting in limited understanding of the long-term effects and potential side effects of drugs. In addition, existing methods often have difficulty updating and reflecting the latest medical information in real time when integrating clinical trials and patient health records, which limits their application value in a rapidly changing medical environment. This lack of dynamic analysis capabilities makes it impossible to make real-time adjustments and optimizations during the drug development process, which may delay the time to market for drugs or lead to inappropriate drug use. The prediction of side effects also relies heavily on historical data and static information, lacking consideration of the dynamic impact of actual biological pathways, which to some extent limits the personalization and precision of treatment plans. Summary of the Invention
[0005] The purpose of the present invention is to solve the shortcomings of the existing technology and propose a method and system for constructing a pharmaceutical knowledge graph based on artificial intelligence.
[0006] In order to achieve the above objectives, the present invention adopts the following technical solution: a method for constructing a pharmaceutical knowledge graph based on artificial intelligence, comprising the following steps:
[0007] S1: Through database query, the molecular structure of the drug and its corresponding target protein sequence are collected, the molecular structure data of the drug is numerically encoded, the molecular fingerprint and protein domain characteristics are extracted, and the drug chemical properties and protein sequence characteristics are combined to form a drug and protein feature set;
[0008] S2: Based on the drug and protein feature set, statistically analyze the interaction between the drug molecule and the protein target, analyze the impact of the interaction on the protein function, identify key protein targets, and obtain interaction analysis results;
[0009] S3: Based on the interaction analysis results, predict and evaluate new drug resistance pathways through quantitative analysis, and simultaneously evaluate the variability of drug-protein interactions and their manifestation in clinical samples to generate drug resistance pathway prediction results;
[0010] S4: Using the drug resistance pathway prediction results, perform time series analysis on the interaction between the drug and the biological pathway, analyze the dynamic changes of the drug's impact, analyze the changes in the drug's effect over time, and generate drug dynamic network analysis results;
[0011] S5: Combined with the results of the drug dynamic network analysis, identify and predict the side effects of the drug, draw a side effect occurrence map by comparing and analyzing the biological pathways affected by the drug and the known side effect database, identify the lesion area and change trend in the image, describe the interaction between the drug treatment effect and pathological characteristics, and generate drug effect correlation analysis results.
[0012] As a further solution of the present invention, the drug and protein feature set includes molecular property characteristics, protein structure characteristics, and interaction potential indicators. The interaction analysis results include protein function perturbation analysis, candidate resistance targets, and molecular interaction strength scores. The resistance pathway prediction results include emerging resistance pathways, pathway stability scores, and resistance impact predictions. The drug dynamic network analysis results include time series response graphs, biological response pathway analysis, and dynamic effect evaluation. The drug effect correlation analysis results include side effect frequency, pathological response patterns, and efficacy correlation assessments.
[0013] As a further solution of the present invention, the molecular structure of the drug and its corresponding target protein sequence are collected by database query, the molecular structure data of the drug is numerically encoded, the molecular fingerprint and protein domain characteristics are extracted, and the specific steps of forming a drug and protein feature set by combining the drug chemical properties and protein sequence characteristics are as follows:
[0014] S101: Summarize drug molecular structure information and target protein sequence data through database query, gradually extract the chemical property parameters of the molecule and the biological property parameters of the protein, classify the molecular structure characteristics and protein sequence characteristics, clean up redundant data item by item, screen and integrate data that meets the constraints, and generate the original data set;
[0015] S102: Based on the original data set, analyze the molecular structures one by one, extract the molecular coordinates and perform standardization adjustment on the normative parameters of the coordinates, analyze the key chemical types and ring structures between the molecules, calculate and generate the molecular fingerprint code based on the analysis results and the weight parameters, integrate the coding results with the target protein sequence tag parameters, and construct the coding data set;
[0016] S103: Based on the encoded data set, a step-by-step comparison process is performed on the molecular fingerprint information, and a multi-dimensional characteristic comparison is performed in combination with the chemical characteristics and similarity parameters between molecules. The target protein characteristics are extracted and cross-calculated with the molecular parameters. The key characteristic records between molecules and proteins are gradually classified and organized to generate a drug and protein feature set.
[0017] As a further solution of the present invention, the molecular fingerprint coding calculation formula is specifically:
[0018]
[0019] Among them, F represents the molecular fingerprint code value, x i Represents the actual three-dimensional coordinate value of each atom in the molecular coordinates, y i Represents the molecular standardized reference coordinate value, z i A relative strength factor that represents the distance between atoms, where n is the total number of atoms in the molecule.
[0020] As a further embodiment of the present invention, based on the drug and protein feature set, the interaction between the drug molecule and the protein target is statistically analyzed, the effect of the interaction on the protein function is analyzed, and the key protein targets are identified. The specific steps for obtaining the interaction analysis results are as follows:
[0021] S201: Extracting characteristic parameters based on the structural characteristics of the drug and protein feature sets and performing grouping operations on the drugs and proteins one by one, calculating the similarity between the two, performing item-by-item verification according to the characteristic comparison rules, eliminating paired records that do not meet the grouping conditions, and organizing and classifying paired characteristic records that meet the rules to construct a paired dataset;
[0022] S202: extracting the interaction characteristic parameters between the molecules and the structural characteristic parameters of the target protein for each pair of drugs and proteins in the paired data set, performing multiple rounds of cross-analysis based on the interaction parameter set, screening and classifying the interaction data, and generating an interaction index set;
[0023] S203: Based on the interaction indicator set, target records with potential drug resistance characteristics are gradually recorded, target parameters are extracted and characteristic comparison analysis is performed, target parameter records with no influence are eliminated through the characteristic set, key characteristic data are gradually classified and integrated to form interaction analysis results.
[0024] As a further embodiment of the present invention, based on the interaction analysis results, new drug resistance pathways are predicted and evaluated by quantitative analysis, while the variability of drug-protein interactions and their manifestation in clinical samples are evaluated. The specific steps for generating drug resistance pathway prediction results are as follows:
[0025] S301: Based on the interaction analysis results, extract the interaction parameters between the drug and the protein, perform a step-by-step classification analysis on the characteristic set of interaction variability and consistency records, extract the distribution of interaction data among different drugs and clinical samples, and combine the classification data to generate a variability analysis result.
[0026] S302: Based on the variability analysis results, interactive pathways with high drug resistance risk are extracted, path parameters are parsed layer by layer using a multi-level feature set, parameter characteristics are gradually grouped and compared, a path characteristic correlation threshold is set, and path characteristic data with a value less than the path characteristic correlation threshold is eliminated to generate a drug resistance path set;
[0027] S303: Extract pathway characteristics item by item based on the drug resistance pathway set, conduct verification and comparison of pathway characteristic parameters item by item in combination with experimental and clinical data, eliminate records with characteristic deviations, and integrate them to form an optimized pathway set to generate prediction results for drug resistance pathways.
[0028] As a further embodiment of the present invention, the drug resistance pathway prediction results are used to perform time series analysis on the interaction between drugs and biological pathways, analyze the dynamic changes of drug effects, and analyze the changes in drug effects over time. The specific steps for generating drug dynamic network analysis results are as follows:
[0029] S401: Based on the drug resistance pathway prediction results, extract drug usage records and related biological data, gradually extract characteristic parameters according to time series and associate them with segmented time nodes, construct a mapping relationship between time labels and characteristic parameters, extract the dynamic change pattern of parameters at time nodes, and generate a time series data set;
[0030] S402: Extracting time node characteristic parameters from the time series data set, gradually parsing the drug effect intensity data over time periods, analyzing changes in drug effect intensity at different time nodes, and calculating decay rates. The dynamically changing characteristic parameters are sorted and classified item by item to generate analysis results of the dynamic changes in drug effects.
[0031] S403: Based on the dynamic change analysis results, extract the time-related parameters of the drug, and through segment-by-segment comparative analysis of the time series data and long-term usage records, extract the correspondence between key characteristics and the disease treatment process, and classify and integrate them to generate standardized analysis results of the drug dynamic network.
[0032] As a further embodiment of the present invention, the formula for calculating the attenuation rate of the drug effect intensity is specifically:
[0033]
[0034] Among them, R t Indicates the attenuation rate of drug effect intensity, A j is the drug effect intensity value at time node j, A j-1 is the drug effect intensity value at time node j-1, T j is the time interval length of time node j, D j is the dose-normalized value at time node j, and m is the total number of time nodes.
[0035] As a further embodiment of the present invention, the side effects of drugs are identified and predicted by combining the results of the drug dynamic network analysis. By comparing and analyzing the biological pathways affected by the drug and the known side effect database, a side effect occurrence map is drawn, the lesion area and change trend in the image are identified, and the interaction between the drug treatment effect and the pathological characteristics is depicted. The specific steps for generating the drug effect correlation analysis results are as follows:
[0036] S501: Based on the results of the drug dynamic network analysis, extracting biological pathway characteristic data related to the drug, analyzing the correlation characteristics between molecules and side effects in the biological pathway layer by layer, setting a data sensitivity threshold, gradually comparing and eliminating low-sensitivity data according to segmented characteristics, classifying and integrating the screened high-sensitivity side effect records, and generating a standardized side effect prediction data set;
[0037] S502: Based on the side effect prediction dataset, extract the dosage conditions and time node characteristic parameters of the drug, analyze the distribution trend item by item, calculate the characteristic change, classify and integrate the key characteristic records according to the trend characteristics of the time node, and form a side effect trend map;
[0038] S503: Based on the side effect trend map, the time period characteristic records related to the lesion area are gradually screened, key characteristic parameters of trend changes are extracted, grouped and sorted based on the correlation characteristics of the lesion parameters, and gradually integrated to form dynamic characteristic records of the drug and the lesion area, generating correlation analysis results of the drug effect.
[0039] A pharmaceutical knowledge graph construction system based on artificial intelligence, including:
[0040] The molecular feature extraction module collects the molecular structure of the drug and its corresponding target protein sequence through database query, numerically encodes the molecular structure data of the drug, extracts the molecular fingerprint and protein domain features, and forms a drug and protein feature set;
[0041] The interaction statistics module statistically analyzes the interactions between drug molecules and protein targets based on the drug and protein feature set, identifies key protein targets, and obtains interaction analysis results;
[0042] The drug resistance pathway prediction module predicts and evaluates new drug resistance pathways based on the interaction analysis results through quantitative analysis, while also evaluating the variability of drug-protein interactions to generate drug resistance pathway prediction results;
[0043] The dynamic network analysis module uses the drug resistance pathway prediction results to perform time series analysis on the interaction between the drug and the biological pathway, analyzes the effect change of the drug over time, and generates drug dynamic network analysis results;
[0044] The side effect prediction module combines the results of the drug dynamic network analysis to identify and predict the side effects of the drug, and draws a side effect occurrence map by comparing and analyzing the biological pathways affected by the drug and the known side effect database;
[0045] The effect analysis module identifies the lesion area and change trend in the image based on the occurrence map of the side effects, depicts the interaction between the drug treatment effect and the pathological characteristics, and generates drug effect correlation analysis results.
[0046] Compared with the prior art, the advantages and positive effects of the present invention are:
[0047] In this invention, through precise analysis of the drug's molecular structure and its target protein sequence, the innovative solution significantly enhances the understanding of the interaction between drugs and proteins, allowing researchers to directly extract key features from the data and monitor the dynamic changes in drug effects. This not only accelerates the drug development process and optimizes treatment plans, but also provides strong data support for personalized medicine by dynamically tracking the interaction between drug side effects and pathological characteristics. BRIEF DESCRIPTION OF THE DRAWINGS
[0048] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.
[0049] Figure 1 Schematic diagram of the steps of the present invention;
[0050] Figure 2 is a flow chart of the steps of S1 of the present invention;
[0051] Figure 3 This is a flow chart of the steps of S2 of the present invention;
[0052] Figure 4 This is a flow chart of the steps of S3 of the present invention;
[0053] Figure 5 This is a flow chart of the steps of S4 of the present invention;
[0054] Figure 6 This is a flow chart of the steps of S5 of the present invention;
[0055] Figure 7 This is a flow chart of the system of the present invention. DETAILED DESCRIPTION
[0056] The technical solution of the present invention is described below in conjunction with the accompanying drawings.
[0057] In the embodiments of the present invention, words such as "exemplarily" and "for example" are used to indicate examples, illustrations, or explanations. Any embodiment or design described as an "exemplary" in the present invention should not be interpreted as being preferred or advantageous over other embodiments or designs. Rather, the use of the word "exemplary" is intended to present concepts in a concrete manner. Furthermore, in the embodiments of the present invention, "and / or" can mean both or either of the two.
[0058] In the embodiments of the present invention, the terms "image" and "picture" may be used interchangeably. It should be noted that, when the distinction between them is not emphasized, their intended meanings are the same. The terms "of," "corresponding," and "corresponding" may be used interchangeably. It should be noted that, when the distinction between them is not emphasized, their intended meanings are the same.
[0059] In the embodiments of the present invention, sometimes a subscript such as W1 may be written as a non-subscript such as W1. When the difference is not emphasized, the meanings to be expressed are the same.
[0060] In order to make the technical problems, technical solutions and advantages to be solved by the present invention clearer, a detailed description will be given below with reference to the accompanying drawings and specific embodiments.
[0061] See also Figure 1 , a pharmaceutical knowledge graph construction method based on artificial intelligence, comprising the following steps:
[0062] S1: Through database query, the molecular structure of the drug and its corresponding target protein sequence are collected, the molecular structure data of the drug is numerically encoded, the molecular fingerprint and protein domain characteristics are extracted, and the drug chemical properties and protein sequence characteristics are combined to form a drug and protein feature set;
[0063] S2: Based on the drug and protein feature sets, statistically analyze the interactions between drug molecules and protein targets, analyze the impact of interactions on protein functions, identify key protein targets, and obtain interaction analysis results;
[0064] S3: Based on the results of the interaction analysis, new drug resistance pathways are predicted and evaluated through quantitative analysis. The variability of drug-protein interactions and their manifestation in clinical samples are also evaluated to generate drug resistance pathway prediction results.
[0065] S4: Using the drug resistance pathway prediction results, conduct time series analysis on the interactions between drugs and biological pathways, analyze the dynamic changes of drug effects, analyze the changes in drug effects over time, and generate drug dynamic network analysis results;
[0066] S5: Combined with the results of drug dynamic network analysis, identify and predict the side effects of drugs. By comparing and analyzing the biological pathways affected by drugs and the database of known side effects, draw a map of the occurrence of side effects, identify the lesion areas and change trends in the image, describe the interaction between drug treatment effects and pathological characteristics, and generate drug effect correlation analysis results.
[0067] The drug and protein feature set includes molecular property characteristics, protein structure characteristics, and interaction potential indicators. The interaction analysis results include protein function perturbation analysis, candidate resistance targets, and molecular interaction strength scores. The resistance pathway prediction results include emerging resistance pathways, pathway stability scores, and resistance impact predictions. The drug dynamic network analysis results include time series response diagrams, biological response pathway analysis, and dynamic effect evaluation. The drug effect correlation analysis results include side effect frequency, pathological response patterns, and efficacy correlation assessments.
[0068] See also Figure 2 , the specific steps of S1 are:
[0069] S101: Summarize drug molecular structure information and target protein sequence data through database query, gradually extract the chemical property parameters of the molecule and the biological property parameters of the protein, classify the molecular structure characteristics and protein sequence characteristics, clean up redundant data item by item, screen and integrate data that meets the constraints, and generate the original data set;
[0070] First, the structural information of drug molecules and the sequence data of target proteins are obtained from the database, and then the chemical properties of the molecules and the biological properties of the proteins are identified through analytical tools. This process includes detailed classification of drug molecular structures and protein sequences. These data are classified and processed to facilitate subsequent operations and analysis. Subsequently, duplicate items in the data are cleaned up and data items that do not meet the conditions are removed to ensure the accuracy and availability of the data. After these processes, the generated original data set will be used for preliminary research on drug development, which lays a solid foundation for subsequent drug research.
[0071] S102: Based on the original data set, analyze the molecular structures one by one, extract the molecular coordinates and perform standardization adjustment on the normative parameters of the coordinates, analyze the intermolecular chemical bond types and ring structures, calculate and generate molecular fingerprint codes based on the analysis results and weight parameters, integrate the coding results with the target protein sequence tags, and construct the coding data set;
[0072] The molecular fingerprint coding calculation formula is as follows:
[0073]
[0074] Among them, F represents the molecular fingerprint code value, x i Represents the actual three-dimensional coordinate value of each atom in the molecular coordinates, y i Represents the molecular standardized reference coordinate value, z i A relative strength factor that represents the distance between atoms, where n is the total number of atoms in the molecule.
[0075] Parameter value setting basis:
[0076] x i The values of are obtained by high-resolution molecular modeling technology, such as molecular coordinates x1 = (1.2, 2.3, 1.8),
[0077] x2=(0.8,1.9,2.0), x3=(1.0,2.0,1.5), the unit is angstrom.
[0078] y i Generate reference coordinate values using a molecular alignment tool (such as alignment based on a reference molecule), such as y1 = (1.0, 2.2, 1.9), y2 = (0.7, 2.0, 2.1), and y3 = (1.1, 1.9, 1.6), in angstroms.
[0079] z i The value is the normalization factor of the distance between adjacent atoms, which is measured experimentally or quantified by the formula where d i is the actual distance between atoms, in angstroms, for example, z1=0.9, z2=0.8, z3=0.7.
[0080] n is the total number of atoms in the molecule, which is directly obtained using molecular analysis tools. In this example, n = 3.
[0081] Calculation steps:
[0082] Calculate the absolute value of the deviation of each set of atomic coordinates and multiply by the corresponding distance factor:
[0083] For the first atom:
[0084] |x1-y1|=|(1.2,2.3,1.8)-(1.0,2.2,1.9)|=(0.2,0.1,0.1);
[0085] Multiply by the corresponding z1=0.9:
[0086] |x1-y1|·z1=(0.2,0.1,0.1)·0.9=(0.18,0.09,0.09);
[0087] Quadratic sum:
[0088] ∑(0.18 2 +0.09 2 +0.09 2 )=0.0324+0.0081+0.0081=0.0486;For the second atom:
[0089] |x2-y2|=|(0.8,1.9,2.0)-(0.7,2.0,2.1)|=(0.1,0.1,0.1);
[0090] Multiply by the corresponding z2=0.8:
[0091] |x2-y2|·z2=(0.1,0.1,0.1)·0.8=(0.08,0.08,0.08);
[0092] Quadratic sum:
[0093] ∑(0.08 2 +0.08 2 +0.08 2 )=0.0064+0.0064+0.0064=0.0192;For the third atom:
[0094] |x3-y3|=|(1.0,2.0,1.5)-(1.1,1.9,1.6)|=(0.1,0.1,0.1);
[0095] Multiply by the corresponding z3=0.7:
[0096] |x3-y3|·z3=(0.1,0.1,0.1)·0.7=(0.07,0.07,0.07);
[0097] Quadratic sum:
[0098] ∑(0.07 2 +0.07 2 +0.07 2 ) = 0.0049 + 0.0049 + 0.0049 = 0.0147; add the results calculated for each atom:
[0099] ∑ i 3 =1 (|x i -y i |·z i ) 2 =0.0486+0.0192+0.0147=0.0825;
[0100] Calculate the normalization factor:
[0101] Calculate the denominator:
[0102] ∑ i 3 =1 z i 2 =0.9 2 +0.8 2 +0.7 2 =0.81+0.64+0.49=1.94;
[0103] Application formula:
[0104]
[0105] Result interpretation:
[0106] The results showed that the molecular fingerprint encoding value was 0.206, reflecting the combined information of molecular coordinate standardization and chemical bond characteristics, used to generate a unified encoding data set. The magnitude of the encoding value directly reflects the degree of deviation of the molecule from the reference standardized structure, which facilitates subsequent target analysis and parameter integration.
[0107] S103: Based on the encoded data set, the molecular fingerprint information is compared step by step, and multi-dimensional characteristics are compared by combining the chemical characteristics and similarity parameters between molecules. The target protein characteristics are extracted and cross-calculated with the molecular parameters. The key characteristics between molecules and proteins are gradually classified and organized to generate drug and protein feature sets.
[0108] Perform step-by-step comparison processing on the molecular fingerprint information, according to the formula W=∑(ci ·S mi ) calculates the key characteristic records between molecules and proteins. i Represents the weight coefficient of chemical properties, S mi Represents the characteristic score of the molecule. Taking into account the differences in chemical properties and similarity parameters between molecules, the characteristic score of each molecule S mi It can be calculated by structural similarity score. For example, for two molecules with similar chemical structures, their structural similarity score can be determined by comparing the number of shared elements and bond types in their chemical structures. If a molecule contains a chemical structure that matches the target protein binding site, its similarity score is higher. By setting the weight coefficient c i , you can adjust the influence of different chemical properties on the overall score. Calculation example: Assume that molecule A and the binding site of the protein have a high structural similarity, and their structural similarity score is S mA =0.8, chemical property weight c ia =0.5, then its contribution score is 0.4. By accumulating the scores of all related molecules, the total key feature record score is obtained, which reflects the interaction potential of the molecule with the target protein.
[0109] See also Figure 3 , the specific steps of S2 are:
[0110] S201: Based on the structural characteristics of the drug and protein feature sets, characteristic parameters are extracted and grouped one by one. The similarity between the two is calculated. Item-by-item verification is performed according to the characteristic comparison rules. Paired records that do not meet the grouping conditions are eliminated. Paired characteristic records that meet the rules are sorted and classified to construct a paired dataset.
[0111] The characteristic parameters of each drug and protein are extracted, and the extracted parameters are divided into different groups according to predefined rules. The grouping and pairing of drugs and proteins are verified one by one according to the characteristic comparison rules. The pairing records that do not meet the grouping conditions are eliminated, and the pairing records that meet the grouping conditions are reclassified. The characteristic records are further sorted according to the grouping and pairing relationship. The sorted records are matched according to the structural characteristics to form a standardized format data set, and finally a paired data set that meets the pairing rules is generated.
[0112] S202: Based on each pair of drugs and proteins in the paired data set, extract the interaction characteristic parameters between the molecules and the structural characteristic parameters of the target protein, conduct multiple rounds of cross-analysis based on the interaction parameter set, screen and classify the interaction data, and generate an interaction index set;
[0113] Based on the interaction characteristic parameters of each pair of drugs and proteins in the paired data set, according to the formula
[0114]
[0115] Calculate interaction indices.
[0116] In the formula, I represents the interaction index value, k i Represents the interaction feature weight factor, x i is the chemical interaction parameter between drug molecules, y i is the structural characteristic parameter of the target protein, z i is the interaction score between the two, and n is the total logarithm of the characteristics of the drug and protein.
[0117] First, obtain the parameters of drugs and proteins through experiments, including x i (such as the binding free energy between molecules, the unit is kcal / mol), y i (e.g., surface hydrophobicity score of a protein, unit is dimensionless score) and z i (e.g. similarity score of binding sites, in percentage).
[0118] Determine the weight factor k i This factor is obtained based on the analysis of experimental statistical data. For example, the weight factor can be determined based on the range of the matching rate between the binding sites of the molecule and the protein.
[0119] The calculation process takes specific data as an example: assuming that the parameters of a pair of drugs and proteins are x1=-8.2, y1=0.75, z1=5.1, and the weight factor k1=1.2.
[0120] First calculate the sub-expression of each term: x1·y1=-6.15,
[0121] Substitute it into the formula to calculate the interaction index:
[0122] I1=k1·(-6.15+26.01)=1.2·19.86=23.83;
[0123] For multiple drug-protein pairs (e.g., n=3), the total interaction index value is obtained by summing the interaction index values of each pair. Assuming that the calculated results of the other two pairs are 25.76 and 18.45 respectively, the total interaction index is:
[0124] I 总 =23.83+25.76+18.45=68.04;
[0125] The results show that the calculated interaction index can quantify the interaction characteristics between drugs and target proteins, provide a basis for screening efficient drug-target pairs, and can be used for subsequent data classification analysis and optimization design.
[0126] S203: Based on the interaction indicator set, gradually record targets with potential drug resistance characteristics, extract target parameters and conduct characteristic comparison analysis, eliminate target parameter records with no influence through the characteristic set, gradually classify and integrate key characteristic data to form interaction analysis results;
[0127] Data on targets with potential drug resistance characteristics were extracted, and combined with interaction characteristic parameters, the key characteristic sets of each target were compared and analyzed with other data sets. Redundant or non-influential target parameter records in the analysis were eliminated. The elimination process was achieved by comparing the influence values of different characteristic sets. Finally, the remaining key target data were classified and integrated, and organized into interaction analysis results that can be directly applied to subsequent analysis and research.
[0128] See also Figure 4 , the specific steps of S3 are:
[0129] S301: Based on the interaction analysis results, the interaction parameters between the drug and protein are extracted. By performing a step-by-step classification analysis on the characteristic set of interaction variability and consistency records, the distribution of interaction data among different drugs and clinical samples is extracted. The classification data is then combined and collated to generate the variability analysis results.
[0130] The interaction parameters of each drug and protein are extracted from the interaction analysis results. The variability and consistency of these parameters are carefully analyzed using data analysis tools. The interaction data are classified according to the analysis results to show the distribution characteristics between different drugs and clinical samples. This method can identify variation patterns with potential clinical relevance, and further classify and organize the data according to these patterns, ultimately generating a detailed variability analysis result. These results help researchers better understand the dynamic interaction behavior between drugs and proteins.
[0131] S302: Based on the variability analysis results, interactive pathways with high drug resistance risk are extracted. Pathway parameters are parsed layer by layer using a multi-level feature set. Pathway parameters are grouped and compared step by step based on their characteristics. A path feature correlation threshold is set, and path feature data with a value less than the path feature correlation threshold is eliminated to generate a drug resistance pathway set.
[0132] Extract the interaction pathways with high resistance risk and group them according to the pathway characteristics.
[0133]
[0134] Calculate the path feature relevance threshold. Where R represents the path feature relevance threshold, t i is the characteristic value of the path, p i is the occurrence frequency of the path, and n is the total number of paths.
[0135] For each path, t i (characteristic value) can be the chemical activity of the interaction between the drug and the protein, and p i (Occurrence frequency) is determined based on the number of times the drug appears in the clinical sample. Taking the specific data as an example, suppose there are three paths, whose characteristic values t are 2.1, 3.5, and 4.0, and the occurrence frequencies p are 10, 20, and 15 times respectively. Calculate the product t of each item i ·p i : 21, 70, 60. Calculate the sum of the total products: 151. Calculate the sum of the frequencies: 45. Apply the formula to calculate R:
[0136]
[0137] The results indicate that the calculated association threshold can be used to identify key pathways associated with specific drug resistance characteristics, thereby helping researchers screen out interaction pathways with potential high drug resistance risks.
[0138] S303: Extract pathway characteristics item by item based on the drug resistance pathway set, conduct verification and comparison of pathway characteristic parameters item by item based on experimental and clinical data, eliminate records with characteristic deviations, and integrate them to form an optimized pathway set to generate prediction results for drug resistance pathways;
[0139] Pathway characteristics are compared using laboratory and clinical data. Through this comparison, records whose characteristic parameters deviate from the normal range can be discovered and eliminated. During the elimination process, each pathway characteristic is evaluated according to preset standards to ensure that only pathways that meet the expected characteristics are retained in the optimized pathway set. This step helps researchers accurately predict which pathways may lead to drug resistance. The generated prediction results provide a scientific basis for subsequent drug development and therapy optimization.
[0140] See also Figure 5 , the specific steps of S4 are:
[0141] S401: Based on the drug resistance pathway prediction results, extract drug usage records and related biological data, gradually extract characteristic parameters according to the time series and associate them with segmented time nodes, build a mapping relationship between time labels and characteristic parameters, extract the dynamic change pattern of parameters at time nodes, and generate a time series dataset;
[0142] Through the time series processing method, the characteristic parameters of each time node are gradually extracted, and these parameters are associated with the segmented time nodes, and a mapping relationship between time labels and characteristic parameters is constructed. In this process, by analyzing the changes in the characteristic parameters of different time nodes, their dynamic change laws are extracted, and the dynamic change characteristics are recorded in a time series format, finally forming a complete time series data set, providing basic data support for subsequent drug dynamic characteristic analysis.
[0143] S402: Extracting time node characteristic parameters from the time series data set, gradually parsing the drug effect intensity data over time periods, analyzing changes in drug effect intensity at different time nodes, and calculating decay rates. The dynamically changing characteristic parameters are then sorted and classified item by item to generate analysis results of the dynamic changes in drug effects.
[0144] The formula for calculating the attenuation rate of drug effect intensity is as follows:
[0145]
[0146] Among them, R t Indicates the attenuation rate of drug effect intensity, A j is the drug effect intensity value at time node j, A j-1 is the drug effect intensity value at time node j-1, T j is the time interval length of time node j, D j is the dose-normalized value at time node j, and m is the total number of time nodes.
[0147] Parameter value setting basis:
[0148] A j The value is obtained through experimental monitoring. For example, the effect intensity of a drug at time nodes 1, 2, and 3 are A1=50, A2=40, and A3=30, respectively, and the unit is mg / L.
[0149] A j-1 It is directly related to the action strength of the adjacent time node. For example, the previous node strength corresponding to node 2 is A1=50.
[0150] T j The time interval is obtained through experimental design or monitoring data, for example, the time interval from node 1 to node 2 is T1 = 5 hours, and the time interval from node 2 to node 3 is T2 = 5 hours.
[0151] D j The dose normalization value is calculated by the ratio of the dose value to the maximum dose value. For example, the dose values of nodes 1, 2, and 3 are D1=100, D2=80, and D3=60, respectively. The maximum dose is 100, and the normalized values are D1=1.0, D2=0.8, and D3=0.6.
[0152] Calculation process:
[0153] Calculate the absolute value of the intensity change at each time node:
[0154] Node 1 to Node 2:
[0155] |A2-A1|=|40-50|=10;
[0156] Node 2 to Node 3:
[0157] |A3-A2|=|30-40|=10;
[0158] The product of the intensity change at each time point and the time interval was calculated and adjusted by the square root of the dose-normalized value:
[0159] Node 1 to Node 2:
[0160]
[0161] Node 2 to Node 3:
[0162]
[0163] Find the sum of squares:
[0164]
[0165] Compute the sum of time intervals:
[0166]
[0167] Substituting the above results into the formula:
[0168]
[0169] Result interpretation:
[0170] The results indicate that the decay rate of drug effect intensity is 23.72, representing the average change in drug effect intensity per unit time. A high decay rate may reflect a rapid drug loss over time, providing important insights for optimizing drug dosage and duration of action. The formula incorporates dose normalization and square root operations, ensuring flexibility and accuracy for varying dosages and time intervals.
[0171] S403: Based on the results of the dynamic change analysis, the time-related parameters of the drug are extracted. By comparing and analyzing the time series data with the long-term usage records, the corresponding relationship between the key characteristics and the disease treatment process is extracted, and the standardized analysis results of the drug dynamic network are generated by classification and integration;
[0172] Time-related parameters of drugs are extracted from time series data. By comparing long-term usage records with time series data section by section, key characteristic parameters of different stages are screened and analyzed. These key characteristic parameters are corresponded to the progress of disease treatment, and the parameter change patterns of each time period are classified and integrated. Finally, a standardized drug dynamic network analysis result is formed, which provides important basic data for dynamic characteristic drug research and clinical application.
[0173] See also Figure 6 , the specific steps of S5 are:
[0174] S501: Based on the results of drug dynamic network analysis, extract the biological pathway characteristic data related to the drug, analyze the correlation characteristics between molecules and side effects in the biological pathway layer by layer, set the data sensitivity threshold, gradually compare and eliminate low-sensitivity data according to the segmented characteristics, classify and integrate the screened high-sensitivity side effect records, and generate a standardized side effect prediction data set;
[0175] By comparing the sensitivity data records of each biological pathway, eliminating low-sensitivity data according to the set sensitivity threshold, and analyzing the pathway characteristic records layer by layer, the highly sensitive side effect records are screened, classified and integrated, and the screened side effect records are standardized to generate a side effect prediction data set, providing basic data for further analysis of drug safety and potential side effects.
[0176] S502: Based on the side effect prediction data set, extract the drug dosage conditions and time node characteristic parameters, analyze the distribution trend item by item, calculate the characteristic change, classify and integrate the key characteristic records according to the trend characteristics of the time node, and form a side effect trend map;
[0177] Extract drug dosage conditions and time node characteristic parameters, analyze their distribution trends item by item, and use the formula
[0178]
[0179] Calculate the time node distribution center of the characteristic change trend.
[0180] Where T represents the center of gravity of the time node, d i is the drug dosage intensity at time node i, t i is the time position of time node i, and n is the total number of time nodes.
[0181] d i The value of is obtained through experiments or records, indicating the dose intensity of the drug effect at each time point, in mg / L or similar units; i Indicates the time value of each time node (such as days).
[0182] During the calculation process, the numerator is the sum of the products of the dose intensity and the time position at each time node, and the denominator is the sum of the dose intensities of all nodes.
[0183] Assume that there are three time nodes, whose dose intensities are d1=20, d2=15, and d3=10, and the corresponding time positions are t1=2, t2=4, and t3=6.
[0184] Calculate the molecular part:
[0185]
[0186] Calculate the denominator:
[0187]
[0188] Apply the formula to calculate:
[0189]
[0190] The results show that the distribution center of the time node is located near the time value of 3.56 days. This value reflects the central trend of drug dosage intensity changes over time and can be used as an important basis for analyzing the time effect and characteristic changes of drug effects.
[0191] S503: Based on the side effect trend map, the time period characteristic records related to the lesion area are gradually screened, key characteristic parameters of trend changes are extracted, and grouped and sorted based on the correlation characteristics of the lesion parameters. The dynamic characteristic records of the drug and the lesion area are gradually integrated to generate the correlation analysis results of the drug effect;
[0192] The lesion parameters are grouped and organized according to their correlation characteristics. By analyzing the changing patterns of the parameters in different time periods, dynamic characteristic records are gradually integrated. These records can reflect the specific interaction patterns between the drug and the lesion area, and ultimately generate correlation analysis results of the drug effect, providing guiding data for drug effect evaluation and optimization.
[0193] See also Figure 7 , a pharmaceutical knowledge graph construction system based on artificial intelligence, including:
[0194] The molecular feature extraction module collects the molecular structure of the drug and its corresponding target protein sequence through database query, numerically encodes the molecular structure data of the drug, extracts the molecular fingerprint and protein domain features, and forms a drug and protein feature set;
[0195] The interaction statistics module statistically analyzes the interactions between drug molecules and protein targets based on the drug and protein feature sets, identifies key protein targets, and obtains interaction analysis results;
[0196] The drug resistance pathway prediction module predicts and evaluates new drug resistance pathways through quantitative analysis based on the results of interaction analysis, while also evaluating the variability of drug-protein interactions to generate drug resistance pathway prediction results;
[0197] The dynamic network analysis module uses the drug resistance pathway prediction results to perform time series analysis on the interactions between drugs and biological pathways, analyze the changes in drug effects over time, and generate drug dynamic network analysis results;
[0198] The side effect prediction module combines the results of drug dynamic network analysis to identify and predict the side effects of drugs. By comparing and analyzing the biological pathways affected by drugs and the database of known side effects, it draws a map of the occurrence of side effects.
[0199] The effect analysis module identifies the lesion areas and change trends in the image based on the occurrence map of side effects, depicts the interaction between drug treatment effects and pathological characteristics, and generates drug effect correlation analysis results.
[0200] The above description is merely a specific embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any modifications or substitutions that can be easily conceived by a person skilled in the art within the technical scope disclosed in the present invention should be included in the scope of protection of the present invention. Therefore, the scope of protection of the present invention should be based on the scope of protection of the claims.
Claims
1. A method for constructing a pharmaceutical knowledge graph based on artificial intelligence, characterized in that: The following steps are involved: Through database query, the molecular structure of the drug and its corresponding target protein sequence are collected, the molecular structure data of the drug is numerically encoded, the molecular fingerprint and protein domain characteristics are extracted, and the drug chemical properties and protein sequence characteristics are combined to form a drug and protein feature set; Based on the drug and protein feature set, statistically analyze the interaction between the drug molecule and the protein target, analyze the impact of the interaction on the protein function, identify key protein targets, and obtain interaction analysis results; Based on the interaction analysis results, new drug resistance pathways are predicted and evaluated through quantitative analysis, while the variability of drug-protein interactions and their manifestations in clinical samples are evaluated to generate drug resistance pathway prediction results; Using the drug resistance pathway prediction results, a time series analysis is performed on the interaction between the drug and the biological pathway, the dynamic changes of the drug effect are analyzed, and the changes in the drug effect over time are analyzed to generate drug dynamic network analysis results; Combined with the results of the drug dynamic network analysis, the side effects of the drug are identified and predicted. By comparing and analyzing the biological pathways affected by the drug and the known side effect database, a side effect occurrence map is drawn, the lesion areas and change trends in the image are identified, the interaction between the drug treatment effect and the pathological characteristics is depicted, and the drug effect correlation analysis results are generated.
2. The method for constructing a pharmaceutical knowledge graph based on artificial intelligence according to claim 1, characterized in that: The drug and protein feature set includes molecular property characteristics, protein structure characteristics, and interaction potential indicators. The interaction analysis results include protein function perturbation analysis, candidate resistance targets, and molecular interaction strength scores. The resistance pathway prediction results include emerging resistance pathways, pathway stability scores, and resistance impact predictions. The drug dynamic network analysis results include time series response diagrams, biological response pathway analysis, and dynamic effect evaluation. The drug effect correlation analysis results include side effect frequency, pathological response patterns, and efficacy correlation assessments.
3. The method for constructing a pharmaceutical knowledge graph based on artificial intelligence according to claim 1, characterized in that: The specific steps for collecting the molecular structure of drugs and their corresponding target protein sequences through database query, numerically encoding the molecular structure data of drugs, extracting molecular fingerprints and protein domain features, and combining the chemical properties of drugs with the characteristics of protein sequences to form a drug and protein feature set are as follows: By querying and summarizing drug molecular structure information and target protein sequence data through database query, the chemical property parameters of the molecule and the biological property parameters of the protein are gradually extracted, the molecular structure characteristics and protein sequence characteristics are classified and processed, and redundant data are cleaned item by item. The data that meets the constraints is screened and integrated to generate the original data set; Based on the original data set, the molecular structures are analyzed one by one, the molecular coordinates are extracted and the normative parameters of the coordinates are standardized, the intermolecular chemical bond types and ring structures are analyzed, and the molecular fingerprint codes are calculated and generated based on the analysis results combined with the weight parameters. The coding results are integrated with the target protein sequence tags to construct a coding data set; Based on the encoded data set, a step-by-step comparison process is performed on the molecular fingerprint information, and a multi-dimensional characteristic comparison is performed in combination with the chemical properties and similarity parameters between molecules. The target protein characteristics are extracted and cross-calculated with the molecular parameters. The key characteristic records between molecules and proteins are gradually classified and organized to generate a drug and protein feature set.
4. The method for constructing a pharmaceutical knowledge graph based on artificial intelligence according to claim 3, characterized in that: The molecular fingerprint coding calculation formula is specifically: Among them, F represents the molecular fingerprint code value, x i Represents the actual three-dimensional coordinate value of each atom in the molecular coordinates, y i Represents the molecular standardized reference coordinate value, z i A relative strength factor that represents the distance between atoms, where n is the total number of atoms in the molecule.
5. The method for constructing a pharmaceutical knowledge graph based on artificial intelligence according to claim 1, characterized in that: Based on the drug and protein feature set, the interaction between the drug molecule and the protein target is statistically analyzed, the impact of the interaction on the protein function is analyzed, and the key protein targets are identified. The specific steps for obtaining the interaction analysis results are as follows: Based on the structural characteristics of the drug and protein feature sets, characteristic parameters are extracted and grouped one by one. The similarity between the two is calculated. Item-by-item verification is performed according to the characteristic comparison rules. Paired records that do not meet the grouping conditions are eliminated. Paired characteristic records that meet the rules are sorted and classified to construct a paired dataset. Based on each pair of drugs and proteins in the paired data set, extract the interaction characteristic parameters between the molecules and the structural characteristic parameters of the target protein, conduct multiple rounds of cross-analysis based on the interaction parameter set, screen and classify the interaction data, and generate an interaction index set; According to the interaction indicator set, targets with potential drug resistance characteristics are gradually recorded, target parameters are extracted and characteristic comparison analysis is carried out, and non-influential target parameter records are eliminated through the characteristic set. Key characteristic data are gradually classified and integrated to form interaction analysis results.
6. The method for constructing a pharmaceutical knowledge graph based on artificial intelligence according to claim 1, characterized in that: Based on the interaction analysis results, new resistance pathways are predicted and evaluated through quantitative analysis, while also evaluating the variability of drug-protein interactions and their manifestation in clinical samples. The specific steps for generating resistance pathway prediction results are as follows: Based on the interaction analysis results, the interaction parameters between the drug and protein are extracted. By performing a step-by-step classification analysis on the characteristic set of interaction variability and consistency records, the distribution of interaction data among different drugs and clinical samples is extracted, and the classification data is combined to generate variability analysis results. Based on the variability analysis results, interactive pathways with high drug resistance risk are extracted, path parameters are parsed layer by layer through a multi-level feature set, and parameter characteristics are gradually grouped and compared. A path characteristic correlation threshold is set, and path characteristic data with a value less than the path characteristic correlation threshold is eliminated to generate a drug resistance path set. Based on the drug resistance pathway set, the pathway characteristics are extracted item by item, and the pathway characteristic parameters are verified and compared item by item in combination with experimental and clinical data. Records with characteristic deviations are eliminated, and the optimized pathway set is integrated to generate the prediction results of the drug resistance pathway.
7. The method for constructing a pharmaceutical knowledge graph based on artificial intelligence according to claim 1, characterized in that: The specific steps for generating drug dynamic network analysis results by using the drug resistance pathway prediction results to perform time series analysis on the interaction between drugs and biological pathways, analyze the dynamic changes of drug effects, and analyze the changes in drug effects over time are as follows: Based on the drug resistance pathway prediction results, drug usage records and related biological data are extracted, characteristic parameters are gradually extracted according to the time series and associated with segmented time nodes, a mapping relationship between time labels and characteristic parameters is constructed, and the dynamic change pattern of parameters at time nodes is extracted to generate a time series data set; Extracting time node characteristic parameters based on the time series data set, gradually parsing the drug effect intensity data over time periods, analyzing changes in drug effect intensity at different time nodes, and calculating decay rates, sorting and classifying the dynamic change characteristic parameters item by item, and generating analysis results of dynamic changes in drug effects; Based on the analysis results of the dynamic changes, the time-related parameters of the drug are extracted. Through a segment-by-segment comparative analysis of the time series data and the long-term usage records, the correspondence between the key characteristics and the disease treatment process is extracted, and the standardized analysis results of the drug dynamic network are generated by classification and integration.
8. The method for constructing a pharmaceutical knowledge graph based on artificial intelligence according to claim 7, characterized in that: The formula for calculating the attenuation rate of the drug effect intensity is specifically: Among them, R t Indicates the attenuation rate of drug effect intensity, A j is the drug effect intensity value at time node j, A j-1 is the drug effect intensity value at time node j-1, T j is the time interval length of time node j, D j is the dose-normalized value at time node j, and m is the total number of time nodes.
9. The method for constructing a pharmaceutical knowledge graph based on artificial intelligence according to claim 1, characterized in that: Combined with the results of the drug dynamic network analysis, the side effects of the drug are identified and predicted. By comparing and analyzing the biological pathways affected by the drug and the known side effect database, a side effect occurrence map is drawn, the lesion area and change trend in the image are identified, and the interaction between the drug treatment effect and the pathological characteristics is depicted. The specific steps for generating the drug effect correlation analysis results are as follows: Based on the results of the drug dynamic network analysis, biological pathway characteristic data related to the drug are extracted, the correlation characteristics between molecules and side effects in the biological pathway are analyzed layer by layer, a data sensitivity threshold is set, low-sensitivity data are gradually eliminated according to the segmented characteristics, and the highly sensitive side effect records screened out are classified and integrated to generate a standardized side effect prediction data set; Based on the side effect prediction data set, extract the dosage conditions and time node characteristic parameters of the drug, analyze the distribution trend item by item, calculate the characteristic change, classify and integrate the key characteristic records according to the trend characteristics of the time node, and form a side effect trend map; According to the side effect trend map, the time period characteristic records related to the lesion area are gradually screened, the key characteristic parameters of the trend change are extracted, and grouping and sorting are carried out in combination with the correlation characteristics of the lesion parameters. The dynamic characteristic records of the drug and the lesion area are gradually integrated to generate the correlation analysis results of the drug effect.
10. A pharmaceutical knowledge graph construction system based on artificial intelligence, characterized by: The method for constructing a pharmaceutical knowledge graph based on artificial intelligence according to any one of claims 1 to 9, wherein the system comprises: The molecular feature extraction module collects the molecular structure of the drug and its corresponding target protein sequence through database query, numerically encodes the molecular structure data of the drug, extracts the molecular fingerprint and protein domain features, and forms a drug and protein feature set; The interaction statistics module statistically analyzes the interactions between drug molecules and protein targets based on the drug and protein feature sets, identifies key protein targets, and obtains interaction analysis results; The drug resistance pathway prediction module predicts and evaluates new drug resistance pathways through quantitative analysis based on the results of interaction analysis, while also evaluating the variability of drug-protein interactions to generate drug resistance pathway prediction results; The dynamic network analysis module uses the drug resistance pathway prediction results to perform time series analysis on the interactions between drugs and biological pathways, analyze the changes in drug effects over time, and generate drug dynamic network analysis results; The side effect prediction module combines the results of drug dynamic network analysis to identify and predict the side effects of drugs. By comparing and analyzing the biological pathways affected by drugs and the database of known side effects, it draws a map of the occurrence of side effects. The effect analysis module identifies the lesion areas and change trends in the image based on the occurrence map of side effects, depicts the interaction between drug treatment effects and pathological characteristics, and generates drug effect correlation analysis results.
Citation Information
Patent Citations
Drug-drug interaction prediction method and system based on multi-modal knowledge graph
CN118430639A
Screening method for multi-target drugs and / or drug combinations
US20180080913A1