Machine learning-based drug target prediction and screening method and system
By using machine learning to screen drug targets in a hierarchical manner, and combining physiological feature analysis and graph vector features, the problem of high risk of combined drug use due to drug target prediction in the treatment of multiple diseases in elderly patients was solved, thus achieving the safety and effectiveness of individualized treatment.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-18
- Publication Date
- 2026-03-27
AI Technical Summary
Existing drug target prediction methods pose a high risk of combined drug use in elderly patients with multiple co-existing diseases, making it difficult to meet the needs of individualized treatment and medication safety.
By acquiring the homeostatic and offset characteristics of the patient's physiological system, a machine learning model is used to screen drug targets in a hierarchical manner. Combined with adaptation indicators and graph vector features, the output weight of the target is dynamically adjusted to optimize the medication regimen.
It improves the accuracy of individual suitability assessment, reduces the risks of combination therapy, and optimizes the safety and efficacy of medication regimens.
Smart Images

Figure CN120998357B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of precision medicine, and in particular to a drug target prediction and screening method and system based on machine learning. BACKGROUND
[0002] Drug target prediction refers to identifying the biological molecular targets, such as proteins, enzymes or receptors, to which a certain compound may act through calculation, bioinformatics or experimental means, which is of great significance in the drug development process. This technology can be used to discover the potential mechanism of action of candidate drugs, assist in new drug target screening, drug repositioning and side effect prediction, effectively improve drug development efficiency and success rate. With the development of artificial intelligence, multi-omics data fusion and network pharmacology, target prediction has become one of the core technologies in precision medicine and systems pharmacology research.
[0003] The current mainstream drug target prediction methods mainly include ligand-based methods based on drug structure similarity, structure-based methods based on protein structure, and network reasoning models integrating multi-omics information. These methods are generally based on biological interaction data between targets and drugs and disease association information, and by constructing a target-drug-disease mapping model, high-throughput prediction and screening of disease treatment targets are achieved to some extent. This approach shows good prediction accuracy and universality in single disease or standardized patient population. However, for the elderly population, due to the general multi-system functional decline in their physiological state, and often suffering from multiple chronic diseases, complex situations of multiple disease co-treatment often occur in clinical treatment. Such patients often need long-term combined treatment with multiple drugs, which may cause metabolic interaction, target interference and action conflict between drugs. If only relying on biological interaction data between targets and drugs and disease association information for target prediction, the problem of high risk of clinical drug use although the target screening is effective will occur, which is difficult to meet the actual needs of individualized treatment and drug safety for such patients.
[0004] Therefore, a drug target prediction and screening method and system based on machine learning are proposed. SUMMARY
[0005] In view of the above prior art situation, the present application is proposed. The embodiments of the present application provide a drug target prediction and screening method and system based on machine learning, which can improve the individual adaptability evaluation accuracy, reduce the risk of combined drug use and optimize the safety of drug use scheme.
[0006] According to an aspect of the present application, a method for predicting and screening drug targets based on machine learning is provided, comprising: obtaining a set of candidate targets related to a disease suffered by a target patient; obtaining a first feature vector representing a physiological system steady state feature of the target patient at a first time point before combination drug use, and a second feature vector representing a physiological function deviation feature at a second time point after combination drug use; calculating an adaptation index of each candidate target in the set of candidate targets according to the first feature vector and the second feature vector, the adaptation index representing the individual adaptation of the candidate target to the target patient; dividing the target points in the set of candidate targets into a first candidate set, a second candidate set and a third candidate set from high to low according to the numerical interval of the adaptation index; inputting the first candidate set into a first machine learning model to obtain a first target set with direct intervention value meeting a first prediction index; for each target point in the second candidate set, generating a graph vector feature representing its structural attribute and upstream and downstream dependency in a preset target-protein interaction network, and inputting it into a second machine learning model to obtain a second target set with potential intervention value meeting a second prediction index; extracting output proportion coefficients of the first target set and the second target set through a preset first mapping table according to the average values of the adaptation indexes corresponding to the first target set and the second target set respectively; merging the first target set and the second target set according to the proportion coefficients to obtain the individualized drug target screening result of the target patient.
[0007] According to another aspect of the present application, a machine learning-based drug target prediction and screening system is provided, comprising: a disease-related target acquisition module configured to acquire a candidate target set related to a disease suffered by a target patient; a physiological feature extraction module configured to acquire a first feature vector representing a physiological system steady state feature of the target patient at a first time point before combination medication, and a second feature vector representing a physiological function deviation feature at a second time point after combination medication; an adaptation index calculation module configured to calculate an adaptation index of each candidate target in the candidate target set according to the first feature vector and the second feature vector, the adaptation index representing the individual adaptation of the candidate target to the target patient; a candidate set division module configured to divide the target points in the candidate target set into a first candidate set, a second candidate set and a third candidate set from high to low according to the numerical interval where the adaptation index is located; a direct intervention prediction module configured to input the first candidate set into a first machine learning model to obtain a first target set with a direct intervention value meeting a first prediction index; a potential intervention prediction module configured to generate a graph vector feature representing the structural attribute and upstream and downstream dependency of each target in the second candidate set in a preset target-protein interaction network, and input the graph vector feature into a second machine learning model to obtain a second target set with a potential intervention value meeting a second prediction index; a proportion coefficient extraction module configured to extract an output proportion coefficient of the first target set and the second target set according to the average value of the adaptation index corresponding to the first target set and the second target set respectively through a preset first mapping table; and a result merging module configured to merge the first target set and the second target set according to the proportion coefficient to obtain an individualized drug target screening result of the target patient.
[0008] According to another aspect of the present application, an electronic device is provided, comprising a memory and a processor, the memory being configured to store computer executable instructions, and the processor being configured to execute the computer executable instructions, which when executed by the processor implement the steps of the method described above.
[0009] According to another aspect of the present application, a computer storage medium is provided, having stored thereon computer executable instructions, which when executed by a processor implement the steps of the method described above.
[0010] Compared with the prior art, the machine learning-based drug target prediction and screening method and system according to the embodiments of the present application can hierarchically evaluate the intervention value of the target points by dividing the first candidate set, the second candidate set and the third candidate set, dynamically adjust the output proportion of the target points in combination with the adaptation index, quantify the target conflict relationship and optimize the safety of the medication in the combination medication scenario, and has the advantages of improving the evaluation accuracy of the individual adaptation, reducing the risk of combination medication and optimizing the safety of the medication scheme. BRIEF DESCRIPTION OF DRAWINGS
[0011] The above and other objects, features and advantages of the present application will become more apparent from the following detailed description thereof taken in conjunction with the accompanying drawings, in which:
[0012] Figure 1 A flowchart of the method for predicting and screening drug targets based on machine learning of the present application.
[0013] Figure 2 A flowchart of the method for predicting and screening drug targets based on machine learning of the present application.
[0014] Figure 3 A flowchart of the method for predicting and screening drug targets based on machine learning of the present application.
[0015] Figure 4 A block diagram of the system for predicting and screening drug targets based on machine learning of the present application.
[0016] Figure 5 A block diagram of an electronic device of the present application. DETAILED DESCRIPTION
[0017] Hereinafter, example embodiments according to the present application will be described in detail with reference to the accompanying drawings. It should be apparent to those skilled in the art that the described embodiments are merely a portion of the embodiments of the present application and thus do not limit the present application thereto as the present application can be implemented in many different forms.
[0018] SUMMARY
[0019] The current mainstream drug target prediction method is generally based on the biological interaction data between the target and the drug and the disease association information, and through the construction of a target-drug-disease mapping model, high-throughput prediction and screening of disease treatment targets are achieved to some extent. However, for the elderly patients with multiple diseases, due to the reason of combination drug use, if only relying on the biological interaction data between the target and the drug and the disease association information for target prediction, the problem of high risk of clinical drug use although the target screening is effective is prone to occur.
[0020] For example, when elderly patients receive combined anticoagulant and anti-inflammatory drug therapy, their cardiovascular homeostasis characteristics show a combined shift in platelet aggregation rate and inflammatory factor levels after medication. If the screening is based solely on the correlation between the target and the disease, ignoring the strength of the target's response to the shift in the patient's individual physiological function, targets highly correlated with coagulation function may be preferentially selected, while ignoring the sensitivity of the target to changes in the activity of inflammatory pathways. Although such targets meet the needs of disease treatment, they may exacerbate metabolic interference between drugs, leading to excessive suppression of coagulation function or an imbalance in the immune response.
[0021] If the above problems are not addressed, the target screening results will not be able to match the individualized physiological state of patients, leading to potential synergistic toxicity risks in clinical medication regimens. Target conflicts under the action of multiple drugs may trigger unexpected activation of biological pathways or inhibition of key metabolic nodes, reducing treatment efficacy and increasing the incidence of adverse reactions. In the long run, this will hinder the application of precision medicine in complex medication scenarios and affect the quality of prognosis management for patients with chronic diseases.
[0022] Faced with the aforementioned problems, this application proposes the following approach: First, consider how to integrate dynamic changes in the patient's physiological system under combined medication scenarios into the target prediction process. To this end, this application attempts to capture patient-specific response patterns by quantifying physiological characteristic shifts before and after medication. The key lies in constructing a dynamic adaptation index that can characterize the correlation between homeostatic changes and target function. Further analysis reveals that the response mechanisms of direct intervention targets and potential synergistic targets to patient physiological shifts differ, requiring stratified processing. Therefore, this application introduces a hierarchical screening mechanism based on adaptation indices. By dividing the candidate set and adopting a differentiated modeling strategy, potential synergistic targets are mined while prioritizing the effectiveness of direct intervention. Finally, the risk-benefit balance between the two types of targets is achieved by dynamically adjusting the output weight, thereby balancing efficacy and safety in complex medication scenarios.
[0023] Exemplary methods
[0024] Figure 1 The illustration shows a machine learning-based drug target prediction and screening method according to an embodiment of this application, including steps S1 to S8.
[0025] like Figure 1 As shown, in step S1, a set of candidate targets related to the disease suffered by the target patient is obtained.
[0026] The candidate target set refers to the set of potential drug targets related to the disease of the target patient. Specifically, it can be obtained by integrating disease-target association databases, drug-target interaction databases and patient genomic data for screening, and is used to provide a basic data range for subsequent target screening.
[0027] In step S2, a first feature vector representing the homeostatic characteristics of the target patient's physiological system at a first time point before combined medication, and a second feature vector representing the physiological function shift characteristics at a second time point after combined medication are obtained.
[0028] The first feature vector is a multidimensional feature vector constructed by collecting homeostatic parameters of the target patient's physiological system before combined drug administration. Specifically, it can be quantified using blood biochemical indicators, metabolomics data, and protein expression levels, and is used to characterize the patient's baseline physiological state before medication. The second feature vector is two vectors of the same dimension and format as the first feature vector. Essentially, they are the values of the same type of parameter at two time points. The second feature vector is used to characterize the dynamic response of the patient's physiological state after drug intervention.
[0029] In step S3, the fit index of each candidate target in the candidate target set is calculated based on the first feature vector and the second feature vector. The fit index represents the individual fit between the candidate target and the target patient.
[0030] Among them, the fit index refers to a comprehensive scoring index that quantifies the fit between candidate targets and individual patients. Specifically, it can be achieved by calculating the weighted value of the target's response intensity to physiological functional deviations and changes in pathway activity, and is used to screen targets that dynamically match the patient's pathological characteristics and drug effects.
[0031] Figure 2 The illustration shows a flowchart of the adaptation index calculation according to an embodiment of this application.
[0032] like Figure 2 As shown, in step S31, the first feature vector is calculated. With the second eigenvector Offset vector: .
[0033] In step S32, candidate target points are extracted according to a preset second mapping table. Response weight vectors that influence intensity along each dimension of the first eigenvector / second eigenvector .
[0034] Specifically, the response weight vector is assigned values based on the correlation strength between the target and the physiological dimension. For example, cardiovascular-related targets have higher weight values for the dimensions of blood pressure and heart rate.
[0035] In step S33, the response weight vector is... With offset vector Perform inner product operation to obtain the representation of candidate target points. The expression of the degree of functional shift response in the target patient: (Shift intensity component) .
[0036] In step S34, the candidate target point is obtained The pathway activity indicators of all involved biological pathways at the first time point and the second time point 、 The pathway activity indicator offset value of the candidate target point is calculated:
[0037] ;
[0038] wherein, represents the number of biological pathways.
[0039] Specifically, it is to quantify the dynamic regulation ability of the candidate target point in the biological network, wherein the pathway activity indicator is a numerical indicator used to quantify the functional state or activity level of a biological pathway at a specific time point, which usually reflects the expression level, activity change or signal transduction intensity of the related genes, proteins or metabolites in the pathway. Specifically, the pathway activity indicator can be calculated by integrating multi-omics data (such as transcriptome, proteome, metabolome, etc.), common methods include gene set enrichment analysis, pathway score calculation, gene expression weighted summation or network topology-based activity score. This indicator is used to evaluate the physiological or pathological state changes of the candidate target point related pathway at different time points of the patient, and to assist in judging the effect and treatment potential of the target point.
[0040] In step S35, the pathway activity indicator offset value and the expression offset intensity component are weighted and summed to obtain the adaptation indicator: wherein, 、 is a preset weight coefficient, satisfying .
[0041] wherein, 、 can be set to be dynamically configured according to the patient's age and the number of complications, for example, for elderly patients with multiple system decline, the weight coefficient of the pathway activity indicator offset value can be set to 0.6 to strengthen the consideration of the network effect of the target point. This calculation method can simultaneously capture the direct response ability of the target point to the physiological changes of the patient individual and its regulation potential in the biological system, thereby improving the evaluation accuracy of the adaptation indicator in complex medication scenarios.
[0042] Returning to Figure 1 In step S4, according to the numerical interval where the adaptation indicator is located, the target points in the candidate target point set are divided into a first candidate set, a second candidate set and a third candidate set from high to low.
[0043] The first candidate set, the second candidate set and the third candidate set are used to realize fine classification processing of target screening. The first candidate set represents highly adaptive target points, which can be directly used for intervention value judgment, and is preferentially screened in a high confidence model. The second candidate set is a medium adaptive target point, which is suitable for combining its structural characteristics in a target-protein interaction network to evaluate its upstream and downstream dependencies and potential action mechanisms through a complex model such as a graph neural network. The third candidate set has low adaptability and can be used as a supplement when there is resource redundancy or insufficient individual target carrying capacity. This hierarchical strategy not only helps to improve the stability and individual matching degree of the prediction result, but also effectively controls the risk level of model calculation resources and target screening.
[0044] In step S5, the first candidate set is input into the first machine learning model to obtain a first target set with direct intervention value meeting a first prediction index.
[0045] The first machine learning model refers to a supervised classification model used to predict whether a candidate target point has a direct intervention value. The model can be trained by algorithms such as support vector machines and random forests. The training samples include target points with known direct therapeutic effects as positive samples, and target points with no significant intervention effect or side effects as negative samples. The input features are individual adaptation characteristics of the candidate target point in the target patient, including but not limited to the response degree of the target point to expression bias and the association strength with the steady-state regulation path. The first prediction index is used to measure whether the candidate target point has a direct intervention value for the target patient. Specifically, it can include the ability of the target point to control the key pathological pathways of the patient and the potential contribution of the target point to the recovery of the system steady state after intervention. The index is quickly distinguished by the trained first machine learning model in the high adaptability target point range, and the output result constitutes the first target set, which is used for subsequent individualized drug target intervention strategy formulation.
[0046] In step S6, for each target point in the second candidate set, a graph vector feature representing its structural attributes and upstream and downstream dependency relationships in a preset target-protein interaction network is generated and input into a second machine learning model to obtain a second target set with potential intervention value meeting a second prediction index.
[0047] The graph vector features refer to the embedding vectors that characterize the topological properties and dependencies of a target in a protein-protein interaction network. Specifically, they can be generated by extracting node centrality and adjacency matrix features through graph neural networks, and are used to uncover the system-level regulatory effects of potential targets. The second machine learning model refers to a graph neural network or other deep learning model suitable for graph-structured data used to identify targets with potential interventional value. Its training samples consist of historical data with known indirect therapeutic effects or validated as potential targets. The input is the graph vector features of each candidate target in the target-protein interaction network. The second prediction index refers to a comprehensive evaluation standard used to assess whether candidate targets have potential interventional value at the system level. It reflects the ability of a target to indirectly regulate pathological processes through network propagation mechanisms, including its role in redundant regulation of signaling pathways and the strength of its interaction with known therapeutic targets. By modeling and predicting targets in the second candidate set using the second machine learning model, potential targets that may not directly affect core pathological pathways but could achieve therapeutic gains through network cascade effects can be effectively identified, thereby improving the comprehensiveness and foresight of drug target identification.
[0048] In step S7, the output weight coefficients of the first target set and the second target set are extracted by means of a preset first mapping table based on the average values of the adaptation indicators corresponding to the first target set and the second target set.
[0049] Specifically, the first target set represents high-trust targets that can be directly acted upon by drugs under the current individual state, while the second target set reflects more the potential regulatory capacity and structural dependence of the system. The average fit index of the two reflects their overall fit with the patient's physiological state. By quantifying these two average values and combining them with a pre-set mapping table (which can be established based on clinical experience or system simulation results), the corresponding output weight coefficients can be extracted. The composition ratio of the final target combination can be flexibly adjusted. This can ensure that targets with significant efficacy are given priority in the screening results, and can also appropriately include potential targets with systemic regulatory value but without direct action evidence, thereby enhancing the coverage and robustness of the treatment strategy. This step is essentially a trade-off mechanism for individual target combination strategies, taking into account both the "credibility" and "systemicity" of target intervention.
[0050] However, since patients are often in a state of multi-system imbalance, there may be functional antagonism, metabolic competition or negative regulatory relationship on signaling pathways between some targets. If these conflicting targets are retained in personalized medicine, the intervention effect may cancel each other out or even produce new adverse reactions.
[0051] Therefore, the application further proposes that before extracting the output proportionality coefficient, the following steps are further included: identifying candidate target point pairs with conflict relationships in the first target point set, the second target point set, and the third candidate set according to a preset target point conflict map; and deleting candidate target points with lower corresponding adaptation indexes from the corresponding first target point set, the second target point set, and the third candidate set.
[0052] Through the above technical solutions, the application can reduce the mutual interference and conflict between target points, improve the consistency and reliability of the screening results, and better adapt to the clinical needs of multi-target joint intervention and reduce the potential risk of adverse reactions, because the complex interactions between target points are considered. In addition, by deleting the options with poor adaptability in the conflict target points, the calculation efficiency of subsequent target point screening can be optimized, and the quality of the final screening results can be improved.
[0053] In step S8, the first target point set and the second target point set are merged according to the proportionality coefficient to obtain the individualized drug target point screening result of the target patient.
[0054] Specifically, merging the first target point set and the second target point set includes: extracting the maximum allowed target point number of the target patient according to the preset bearing capacity mapping table and the drug regimen of the combination drug; calculating the first target point amount required to be extracted from the first target point set and the second target point amount required to be extracted from the second target point set according to the maximum allowed target point number and the output proportionality coefficient; extracting a corresponding number of candidate target points from the first target point set and the second target point set according to the first target point amount and the second target point amount; and merging all the extracted candidate target points to obtain the individualized drug target point screening result.
[0055] For example, for a 65-year-old elderly patient with hypertension, diabetes, and coronary heart disease, the maximum allowed target point number is 10, which is obtained by querying the bearing capacity mapping table according to the intervention intensity, toxicity risk, and patient tolerance of the drug combination. The first target point amount required to be extracted from the first target point set and the second target point amount required to be extracted from the second target point set are calculated according to the maximum allowed target point number and the output proportionality coefficient. Assuming that the output proportionality coefficient is 0.6:0.4, the first target point amount is 6, and the second target point amount is 4, then 6 target points are finally selected from the first target point set, and 4 target points are selected from the second target point set. Here, the selection can be random selection or selection according to the adaptation index from high to low.
[0056] In the above scheme, since the same target point can be acted on by multiple drugs, if not further screened, there can be repetition of drug action and functional overlap of target points, resulting in waste of resources. Therefore, the application further proposes: extracting a corresponding number of candidate target points comprises: calculating the redundancy of each candidate target point in the first target point set and the second target point set according to the drug combination scheme and the preset drug-target mapping atlas, the redundancy representing the degree of overlap of the candidate target point being acted on by multiple drugs; according to the order from low to high of the redundancy, a corresponding number of candidate target points are extracted from the first target point set and the second target point set respectively.
[0057] In the above scheme, the calculation formula of the redundancy is: , wherein, represents the set of all drugs in the drug combination scheme, is an indicator function, indicating that when the candidate target point is a known target point of the drug , , otherwise , represents the size of the drug set .
[0058] Specifically, after calculating the redundancy of each candidate target point, the target points in the first target point set and the second target point set are sorted in order from high to low, and for the N target points to be extracted from the first target point set, the first N target points with the highest redundancy are selected, and the same sorting rule is used for the second target point set. For example, when the first target point set contains a target point A with a redundancy of 0.5 and a target point B with a redundancy of 0.8, and one target point needs to be extracted, target point A is preferred. This way can effectively identify target points that are acted on by multiple drugs, and in the merging process, target points with high repeated intervention are preferentially excluded, thereby reducing the metabolic pressure and adverse reaction risk caused by excessive activation of target points when multiple drugs are used.
[0059] In the above scheme, there can be a case where the maximum allowed number of target points of the patient is greater than the sum of the number of candidate target points in the first target point set and the second target point set. In this case, if no supplement is performed, the number of target points in the individualized drug target screening result is insufficient, which can limit the coverage of drug intervention and fail to fully exert the synergistic therapeutic effect of combination therapy, thereby reducing the comprehensiveness and accuracy of treatment, especially in elderly patients with multiple system damage or multiple disease co-treatment, it is difficult to meet the individualized treatment needs and improve the safety of drug use. Therefore, the present application further proposes that before merging all the extracted candidate target points, the following steps are further included: extracting all candidate target points with an adaptation index greater than a predetermined threshold from the third candidate set to form a candidate target set; when the maximum allowed number of target points is greater than the sum of the number of candidate target points in the first target point set and the second target point set, the corresponding candidate target points are extracted from the candidate target set in order of high to low according to the adaptation index for supplement until the maximum allowed number of target points is met or the candidate target set is exhausted.
[0060] It should be noted that the adaptation index of the candidate target points in the third candidate set is lower, indicating that the individual adaptation of the candidate target points to the current physiological state and disease characteristics of the patient is limited, and therefore the priority is lower, and the candidate target points are not suitable for directly entering the fine screening stage of the machine learning model. In addition, due to the large number and lower priority, directly performing complex machine learning prediction and redundancy detection on the candidate target points will increase the computational burden and limited benefits. Therefore, the candidate target points are used as candidate target points, and are only supplemented when the first and second target point sets cannot meet the maximum allowed number of target points, which ensures the efficiency and focus of target point screening, and also takes into account the flexibility and comprehensiveness of the treatment plan, avoiding resource waste and potential risk of excessive intervention.
[0061] Exemplary system
[0062] Figure 4The system for predicting and screening drug targets based on machine learning according to the embodiment of the present application is illustrated, comprising: a disease-related target acquisition module configured to acquire a candidate target set related to a disease suffered by a target patient; a physiological characteristic extraction module configured to acquire a first feature vector representing a physiological system steady state characteristic of the target patient at a first time point before combination medication, and a second feature vector representing a physiological function deviation characteristic at a second time point after combination medication; an adaptation index calculation module configured to calculate an adaptation index of each candidate target in the candidate target set according to the first feature vector and the second feature vector, the adaptation index representing individual adaptation of the candidate target to the target patient; a candidate set division module configured to divide the target in the candidate target set into a first candidate set, a second candidate set and a third candidate set from high to low according to a numerical interval where the adaptation index is located; a direct intervention prediction module configured to input the first candidate set into a first machine learning model to obtain a first target set with direct intervention value meeting a first prediction index; a potential intervention prediction module configured to generate a graph vector feature representing structural attribute and upstream / downstream dependency relationship of each target in the second candidate set in a preset target-protein interaction network, and input into a second machine learning model to obtain a second target set with potential intervention value meeting a second prediction index; a proportion coefficient extraction module configured to extract an output proportion coefficient of the first target set and the second target set according to an average value of the adaptation index corresponding to the first target set and the second target set respectively through a preset first mapping table; and a result merging module configured to merge the first target set and the second target set according to the proportion coefficient to obtain an individualized drug target screening result of the target patient.
[0063] In one example, the adaptation index calculation module calculates the adaptation index, comprising: calculating a deviation vector of the first feature vector and the second feature vector ; extracting a response weight vector representing influence intensity of the candidate target on each dimension of the first feature vector / second feature vector according to a preset second mapping table; performing an inner product operation on the response weight vector and the deviation vector to obtain an expression deviation intensity component representing a response degree of the candidate target to the function deviation of the target patient: acquiring path activity indexes of all biological pathways participated by the candidate target at the first time point and the second time point, and calculating a path activity index deviation value of the candidate target :
[0064] ;
[0065] wherein, represents the number of biological pathways; the pathway activity index offset value and the expression offset intensity component weighted sum, to obtain the fitting index: wherein, , is a preset weight coefficient, satisfying .
[0066] In one example, the proportion coefficient extraction module further includes, before extracting the output proportion coefficient: identifying candidate target pairs in the first target set, the second target set and the third candidate set that have a conflict relationship according to a preset target point conflict map; and deleting the candidate target with a lower corresponding fitting index in the candidate target pair from the corresponding first target set, the second target set and the third candidate set.
[0067] In one example, the result merging module merging the first target set and the second target set includes: extracting the maximum allowed target number of the target patient according to a preset bearing capacity mapping table and the drug regimen of the combination drug; calculating the first target amount required to be extracted from the first target set and the second target amount required to be extracted from the second target set according to the maximum allowed target number and the output proportion coefficient; extracting a corresponding number of candidate targets from the first target set and the second target set according to the first target amount and the second target amount; and merging all extracted candidate targets to obtain the individualized drug target screening result.
[0068] In one example, the result merging module extracting a corresponding number of candidate targets includes: calculating the redundancy of each candidate target in the first target set and the second target set according to the drug regimen of the combination drug and a preset drug-target mapping map, the redundancy indicating the degree of coincidence of the candidate target being acted on by multiple drugs; and extracting a corresponding number of candidate targets from the first target set and the second target set in order from low to high according to the redundancy. In one example, the calculation formula of the redundancy is:
[0069] ;
[0070] wherein, represents the set of all drugs in the drug regimen of the combination drug, is an indicator function, indicating when the candidate target is a known target of the drug , , otherwise , represents the size of the drug set .
[0071] In one example, the result merging module further includes, before merging all the extracted candidate target points, extracting all the candidate target points with the fitting index greater than the pre-set threshold from the third candidate set to form a candidate target set; and when the maximum allowed target point amount is greater than the sum of the candidate target points in the first target set and the second target set, extracting the corresponding candidate target points from the candidate target set in order from high to low according to the fitting index to supplement until the maximum allowed target point amount is met or the candidate target set is exhausted.
[0072] Exemplary electronic device
[0073] Figure 5 An electronic device according to embodiments of the present application is illustrated. The electronic device can be the mobile device itself, or a stand-alone device independent of the mobile device, which can communicate with the mobile device to receive the acquired input signals therefrom and send the selected target driving behavior thereto.
[0074] Figure 5 A block diagram of an electronic device according to embodiments of the present application is illustrated.
[0075] As Figure 5 illustrated, the electronic device includes one or more processors and a memory.
[0076] The processor can be a central processing unit (CPU) or other form of processing unit having data processing and / or instruction execution capabilities, and can control other components in the electronic device to perform desired functions.
[0077] The memory can include one or more computer program products, which can include various forms of computer-readable storage media, such as volatile memory and / or non-volatile memory. The volatile memory, for example, can include random access memory (RAM), cache, and / or the like. The non-volatile memory, for example, can include read-only memory (ROM), hard disk, flash memory, and / or the like. One or more computer program instructions can be stored on the computer-readable storage media, and the processor can execute the program instructions to implement the driving behavior decision method of various embodiments of the present application described above and / or other desired functions.
[0078] In one example, the electronic device can further include input and output devices, which are interconnected through a bus system and / or other forms of connection mechanism (not shown).
[0079] Of course, for simplicity, Figure 3 only some of the components in the electronic device related to the present application are shown in the figure, and components such as buses, input / output interfaces, and the like are omitted. In addition, the electronic device can include any other appropriate components according to specific application cases.
[0080] Exemplary computer-readable medium
[0081] Embodiments of the present application can also be computer readable storage medium having stored thereon computer program instructions which, when executed by a processor, cause the processor to perform steps of the driving behavior decision method according to various embodiments of the present application described in the above "Exemplary Method" section of the present specification.
[0082] The computer readable storage medium can be any combination of one or more computer readable medium. The computer readable medium can be a computer readable signal medium or a computer readable storage medium. The computer readable storage medium can include, for example, but not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any suitable combination of the above. More specific examples (a non-exhaustive list) of the computer readable storage medium include an electrical connection having one or more wires, a portable disc, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above.
[0083] The above describes the basic principles of the present application in conjunction with specific embodiments, but it should be noted that the advantages, benefits, effects and the like mentioned in the present application are only examples and are not limiting, and these advantages, benefits, effects and the like cannot be considered as necessary for each embodiment of the present application. In addition, the above specific details are only for the purpose of example and understanding, and are not limiting, and the above details do not limit the present application to the above specific details.
[0084] The block diagrams of the devices, apparatuses, equipment, systems involved in the present application are only illustrative examples and are not intended to require or imply that the connections, arrangements, configurations must be as shown in the block diagrams. As those skilled in the art will recognize, these devices, apparatuses, equipment, systems can be connected, arranged, configured in any manner. Words such as "include", "contain", "have" and the like are open-ended words that mean "including but not limited to", and can be used interchangeably. The words "or" and "and" used herein mean the word "and / or", and can be used interchangeably unless the context clearly indicates otherwise. The word "such as" used herein means the phrase "such as but not limited to", and can be used interchangeably.
[0085] It should also be noted that in the devices, apparatuses and methods of the present application, each component or each step can be decomposed and / or recombined. These decompositions and / or recombinations should be considered as equivalent solutions of the present application.
[0086] The above description of the disclosed aspects is provided to enable any person skilled in the art to make or use the application. Various modifications to these aspects will be readily apparent to those skilled in the art, and the generic principles defined herein can be applied to other aspects without departing from the scope of the application. Thus, the present application is not intended to be limited to the aspects shown herein but is to be accorded the widest scope consistent with the principles and novel features disclosed herein.
[0087] The above description has been presented to enable any person skilled in the art to make or use the application. Numerous modifications to the aspects described herein will be readily apparent to those skilled in the art, and the principles defined herein can be applied to other aspects without departing from the scope of the application. Thus, the present application is not intended to be limited to the aspects shown herein but is to be accorded the widest scope consistent with the principles and novel features disclosed herein.
Claims
1. A method for drug target prediction and screening based on machine learning, characterized in that, The method comprises the following steps: obtaining a set of candidate target points related to a disease suffered by a target patient; obtaining a first feature vector representing a physiological system steady state feature of the target patient at a first time point before combination medication, and a second feature vector representing a physiological function deviation feature at a second time point after combination medication; calculating an adaptation index of each candidate target point in the set of candidate target points according to the first feature vector and the second feature vector, the adaptation index representing the individual adaptation of the candidate target point to the target patient; dividing the target points in the set of candidate target points into a first candidate set, a second candidate set and a third candidate set from high to low according to the numerical interval of the adaptation index; inputting the first candidate set into a first machine learning model to obtain a first target point set with direct intervention value meeting a first prediction index; for each target point in the second candidate set, generating a graph vector feature representing its structural attribute and upstream and downstream dependence in a preset target point-protein interaction network, and inputting the graph vector feature into a second machine learning model to obtain a second target point set with potential intervention value meeting a second prediction index; extracting an output proportion coefficient of the first target point set and the second target point set through a preset first mapping table according to the average value of the adaptation index corresponding to the first target point set and the second target point set respectively; merging the first target point set and the second target point set according to the proportion coefficient to obtain an individualized drug target point screening result of the target patient; wherein the merging of the first target point set and the second target point set comprises: extracting a maximum allowed target point number of the target patient according to a preset bearing capacity mapping table and a drug regimen of the combination medication; calculating a first target point amount required to be extracted from the first target point set and a second target point amount required to be extracted from the second target point set according to the maximum allowed target point number and the output proportion coefficient; extracting a corresponding number of candidate target points from the first target point set and the second target point set according to the first target point amount and the second target point amount; and merging all extracted candidate target points to obtain the individualized drug target point screening result; the extracting of the corresponding number of candidate target points comprises: calculating a redundancy of each candidate target point in the first target point set and the second target point set according to a drug regimen of the combination medication and a preset drug-target mapping atlas, the redundancy representing the coincidence degree of the candidate target point being acted on by multiple drugs; and extracting a corresponding number of candidate target points from the first target point set and the second target point set in order from low to high according to the redundancy; the merging of all extracted candidate target points further comprises: extracting all candidate target points with an adaptation index greater than a preset threshold from the third candidate set to form a candidate target set; and when the maximum allowed target point amount is greater than the sum of the number of candidate target points in the first target point set and the second target point set, extracting corresponding candidate target points from the candidate target set in order from high to low according to the adaptation index to supplement until the maximum allowed target point amount is met or the candidate target set is exhausted. 2.The method of claim 1, wherein, the calculation of the adaptation index comprises: computing the first feature vector a displacement vector from the second feature vector ; extracting a candidate target point according to a preset second mapping table a response weight vector affecting intensity in each dimension of the first feature vector / second feature vector ; The response weight vector With the offset vector Perform inner product operation to obtain the representation of the candidate target point. The expression of the degree of functional shift response in the target patient: ; acquiring the candidate target point an indicator of pathway activity of all biological pathways involved at the first time point and at the second time point 、 and calculating an indicator of pathway activity offset value of the candidate target point : wherein, denotes the number of biological pathways; shift the channel activity indicator with the expression shift intensity component weighting and summing to obtain the fit indicator: wherein the , is a preset weight coefficient, satisfying . 3.The machine learning based drug target prediction and screening method according to claim 1, characterized in that, the extraction of the output proportion coefficient further comprises: According to a preset target conflict map, the first target set, the second target set and the candidate target set in the third candidate target set are identified to have a conflict relationship; The candidate target corresponding to the lower adaptation index is deleted from the first target set, the second target set and the third candidate target set. 4.The method of claim 1, wherein the method is characterized by, The calculation formula of the redundancy is: ; wherein, denotes the set of all drugs in the drug regimen of the combination therapy, and is an indicator function, denoting when the candidate target is a known target of a drug , , otherwise , denotes the size of the set of drugs .
5. A method and system for predicting and screening drug targets based on machine learning, characterized in that, Comprise: A disease-related target acquisition module is configured to acquire a candidate target set related to a disease suffered by a target patient; A physiological characteristic extraction module is configured to acquire a first feature vector representing a physiological system steady state characteristic of the target patient at a first time point before combination drug use, and a second feature vector representing a physiological function deviation characteristic at a second time point after combination drug use; An adaptation index calculation module is configured to calculate an adaptation index of each candidate target in the candidate target set according to the first feature vector and the second feature vector, the adaptation index representing the individual adaptation of the candidate target to the target patient; A candidate set division module is configured to divide the target points in the candidate target set into a first candidate set, a second candidate set and a third candidate set from high to low according to the numerical interval of the adaptation index; A direct intervention prediction module is configured to input the first candidate set into a first machine learning model to obtain a first target set with a direct intervention value meeting a first prediction index; A potential intervention prediction module is configured to generate a graph vector feature representing the structural attribute and upstream and downstream dependence relationship of each target in the second candidate set in a preset target-protein interaction network, and input the graph vector feature into a second machine learning model to obtain a second target set with a potential intervention value meeting a second prediction index; A proportion coefficient extraction module is configured to extract output proportion coefficients of the first target set and the second target set through a preset first mapping table according to the average values of the adaptation indexes corresponding to the first target set and the second target set; A result merging module is configured to merge the first target set and the second target set according to the proportion coefficients to obtain an individualized drug target screening result of the target patient; The merging of the first target set and the second target set comprises: extracting a maximum allowed target number of the target patient according to a preset bearing capacity mapping table and the drug regimen of the combination drug; calculating a first target amount required to be extracted from the first target set and a second target amount required to be extracted from the second target set according to the maximum allowed target number and the output proportion coefficients; extracting a corresponding number of candidate targets from the first target set and the second target set according to the first target amount and the second target amount; and merging all extracted candidate targets to obtain the individualized drug target screening result. The extracting a corresponding number of candidate targets comprises: calculating redundancy of each candidate target in the first target set and the second target set according to a drug regimen of the combination medication and a preset drug-target mapping, the redundancy representing a coincidence degree of the candidate target being acted on by multiple drugs; and extracting a corresponding number of candidate targets from the first target set and the second target set in order from low to high according to the redundancy. The merging all extracted candidate targets further comprises: extracting all candidate targets with an adaptation index greater than a preset threshold from the third candidate set to form a candidate target set; and when the maximum allowed target amount is greater than a sum of the number of candidate targets in the first target set and the second target set, extracting corresponding candidate targets from the candidate target set in order from high to low according to the adaptation index to supplement, until the maximum allowed target amount is met or the candidate target set is exhausted. 6.An electronic device comprising a memory and a processor, the electronic device characterized by: The memory is configured to store computer executable instructions, and the processor is configured to execute the computer executable instructions, so as to implement the steps of the method in any one of claims 1-4.
7. A computer storage medium having stored thereon computer- executable instructions which, when executed by a computer, cause the computer to carry out the steps of claim 1. The computer executable instructions, when executed by the processor, implement the steps of the method in any one of claims 1-4.
Citation Information
Patent Citations
High-throughput automatic method for drug screening
CN119360949A
Artificial intelligence-driven polycystic ovarian syndrome drug target screening method
CN119517152A