Prediction method of drug activity

Through Gaussian correlation embedding technology, the dependency neglect problem in multi-dimensional feature modeling of drug molecular data is solved, and the accuracy of drug characteristic prediction and candidate drug screening is improved, and the efficiency and success rate of drug development is improved.

CN120183539APending Publication Date: 2025-06-20JINAN UNIVERSITY
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510475693.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-16
Publication Date
2025-06-20

AI Technical Summary

Technical Problem

The prior art is difficult to effectively model the multidimensional characteristics of drug molecular data, especially ignoring the dependence between different molecular features, resulting in incomplete data feature representation and insufficient model prediction accuracy.

Method used

Using Gaussian correlation embedding technology, the multi-dimensional nonlinear dependent structure is mapped to the Gaussian distribution space by introducing Gaussian correlation function, and a unified and consistent feature embedding representation is generated to effectively integrate heterogeneous data such as chemical structure, biological activity and drug interaction of the compound.

Benefits of technology

It significantly improves the accuracy and robustness of drug characteristic prediction and candidate drug screening, solves the shortcomings of existing models in complex data integration and information utilization, and improves the efficiency and success rate of drug research and development.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120183539A_ABST
    Figure CN120183539A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of structural analysis and learning of pharmaceutical activity and pharmaceutical molecules, in particular to a prediction method of pharmaceutical activity. According to the technical scheme, the method comprises the steps of extracting and preprocessing data features, conducting joint modeling on drug molecular features through a Gaussian correlation embedding mechanism to form a prediction model so as to process the complex relation between different modals and heterogeneous data features, and obtaining a drug function activity prediction network through the prediction model. The method can effectively capture and model multi-dimensional dependency in drug molecule features, and can map a multi-dimensional nonlinear dependency structure to a Gaussian distribution space by introducing a Gaussian correlation function, so as to generate uniform and consistent feature embedding representation among the multi-dimensional features. According to the method, heterogeneity data such as chemical structures, biological activity and drug interaction of compounds can be effectively integrated, and the deep dependency relationship of the heterogeneity data is captured, so that the precision and robustness of drug characteristic prediction and candidate drug screening are remarkably improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of the analysis and learning of drug activity and the structure of drug molecules, and particularly relates to a method for predicting drug activity. Background Art

[0002] With the continuous development of modern drug research and development, simple drug research and development technologies can no longer meet its development needs. The rapid rise of information science in recent years has brought great opportunities to the research and development work in the drug field. The process of drug research and development involves various types of data, and its data forms involve various styles, such as its chemical structure, donor biological activity, drug interaction structure, etc. The complexity and diversity presented by these data provide rich semantic information in drug research and development. How to efficiently integrate these data to improve the accuracy of drug property prediction and candidate drug screening is of great significance for enhancing the robustness and effectiveness of drug research and development. This makes the development of an embedding model that can process and fuse various types of data a key direction for improving the efficiency and success rate of drug research and development.

[0003] A key challenge in drug research and development is how to effectively model the multi-dimensional features of molecular data, and there are often complex non-linear correlations between these features. Existing feature embedding methods mostly rely on independent modeling, ignoring the dependence relationships between different molecular features, which may lead to problems such as incomplete data feature representation and insufficient model prediction accuracy. Specific to drug discovery tasks, the dependence relationships of molecular features are crucial for the activity, selectivity, toxicity, etc. of drugs, but existing technologies have not fully revealed these feature dependencies. Summary of the Invention

[0004] The object of the present invention is to address the problems in the background art and propose a method for predicting drug activity, which can effectively integrate heterogeneous data such as the chemical structure, biological activity, and drug interaction of compounds, capture their deep dependence relationships, thereby significantly improving the accuracy and robustness of drug property prediction and candidate drug screening. To solve the deficiencies of existing models in complex data integration and information utilization, through Gaussian correlation embedding technology, it can effectively capture and model the multi-dimensional dependencies in drug molecular features. By introducing the Gaussian correlation function, it is possible to map the multi-dimensional non-linear dependence structure to the Gaussian distribution space, and then generate a unified and consistent feature embedding representation among multi-dimensional features, which can effectively integrate heterogeneous data such as the chemical structure, biological activity, and drug interaction of compounds, capture their deep dependence relationships, thereby significantly improving the accuracy and robustness of drug property prediction and candidate drug screening.

[0005] Technical solution of the present invention: A method for predicting drug activity, data feature extraction and preprocessing, obtaining physicochemical property data and pharmacological experiment data in drug data, using a context - associated feature engineering method to capture complex relationships between data, extracting core information of drug physicochemical property data through feature decomposition, and performing feature modeling on pharmacological experiment data by combining context and time - series characteristics;

[0006] Gaussian correlation embedding mechanism, jointly modeling drug molecular features through the Gaussian correlation embedding mechanism to form a prediction model for handling complex relationships between different modalities and heterogeneous data features;

[0007] Gradient optimization and parameter update, the prediction model uses context information and the interaction between features to gradually optimize each parameter, making the final feature representation more accurate;

[0008] Input the drug molecular features after Gaussian correlation embedding and parameter optimization into the prediction model to obtain a drug functional activity prediction network.

[0009] Preferably, the physicochemical property data in the drug data includes melting point, boiling point, solubility, lipophilic - hydrophilic partition coefficient, dissociation degree, volatility, hygroscopicity, efflorescence, surface activity and crystal form;

[0010] The pharmacological experiment data in the drug data includes biological activity, pharmacodynamic response, metabolic stability, cytotoxicity.

[0011] Preferably, the drug data extracts core information through feature decomposition and forms samples, performs independent feature extraction on each sample, and combines the association relationship of context nodes to form a context - associated feature representation.

[0012] Preferably, an effective context - independent feature matrix is constructed through a context model, which combines the embedding vectors of each independent feature with the weight values of their corresponding nodes to generate an effective context representation for each sample, and its mathematical expression is:

[0013]

[0014] where, R n is the embedding vector of the effective context representation of each sample, W is the number of context nodes, α t,ω is the embedding vector of the corresponding node, val ω is the value of the context node.

[0015] Preferably, the context-independent feature matrix includes the physical and chemical property feature matrix M1 of the drug and the feature matrix M2 of the pharmacodynamic experimental data, and the Gaussian correlation function is used to capture the statistical correlation between different features. The joint distribution of the features can be calculated by the Gaussian correlation expression:

[0016] GC(u R ,u T )=Φ(Φ - (u R ),Φ -1 (u T );ρ)

[0017] Wherein, u R and u T respectively represent the marginal distribution values of different feature variables, Φ is the cumulative distribution function of the standard normal distribution, Φ -1 represents its corresponding inverse function, and ρ is the correlation parameter between the features.

[0018] Preferably, based on the physical and chemical property feature matrix M1 and the feature matrix M2 of the pharmacodynamic experimental data, the feature representation is further optimized and updated through the Gaussian correlation embedding mechanism to form a feature matrix representation based on context information. During the context-dependent modeling process, the embedding vector R n characterizes the relationship between each central node and its context nodes to adjust the feature weights and optimize the feature representation. The formula for optimizing the feature representation is:

[0019] M context =τM1+(1-τ)M2

[0020] Wherein, τ and 1-τ are dynamically adjusted weights controlled by the Gaussian correlation embedding mechanism, used to balance the importance of different features, thereby generating an integrated feature representation optimized by context. M context refers to the feature matrix dynamically updated based on the context relationship.

[0021] Preferably, in each iteration of the prediction model, the feature values of each sample and its context nodes are traversed through the backpropagation mechanism to calculate the gradients of the model parameters.

[0022] Preferably, during each iteration of the prediction model, the model complexity is reduced and the risk of overfitting is reduced through the parameter sharing strategy. The same model parameter θ shared, is shared throughout the iteration process and can be expressed as:

[0023] θ shared =θ1=θ2=…=θ n

[0024] Wherein, θ sharedModel parameters shared among various parts of the model or different samples, θ1, θ2, … θ n respectively represent the same parameters used for each subtask or sample;

[0025] Through multiple iterations of gradient optimization and parameter update, the feature representation of the model is gradually optimized, and its specific expression is:

[0026]

[0027] Among them, F (t) represents the feature representation at the t-th iteration, Λ is the feature optimization function, and usually the feature is updated and optimized based on the gradient backpropagation mechanism, represents the gradient at the t-th iteration, and θ shared represents the common parameters in the model.

[0028] The present invention proposes a Gaussian correlation embedding technique to effectively capture and model the multi-dimensional dependencies in drug molecular features. By introducing a Gaussian correlation function, the present invention can map the multi-dimensional non-linear dependence structure into the Gaussian distribution space, and then generate a unified and consistent feature embedding representation among multi-dimensional features. This method solves the limitations of the prior art in multi-dimensional dependence modeling, enabling the embedded features to more truly reflect the molecular structure and drug properties.

[0029] The Gaussian correlation technique in the present invention is mainly applied to drug activity prediction, toxicity assessment, and drug similarity analysis. Compared with the existing independent modeling methods, the present invention innovatively uses the correlation function to model the complex dependencies of different features, improving the expression ability of data embedding. This method has significant advantages in capturing the potential structure of molecular data, improving model performance, etc., and is applicable to the processing scenarios of high-dimensional and complex-dependent molecular data.

[0030] By applying Gaussian correlation embedding to actual drug data, the present invention has achieved performance improvement in drug prediction tasks, including higher prediction accuracy and more stable model performance. In addition, the present invention has strong scalability and can be used for different types of molecular data analysis, bringing new ideas and technological breakthroughs to data processing in drug research and development.

[0031] Preferably, based on the integration method of multi-level feature representation, multi-scale feature integration processing is performed on the optimized feature representation. The prediction model includes a function classification module for predicting the specific functional categories of drugs, and also includes an active region regression module for locating the action regions of drugs in biological systems, and the two form a drug function activity prediction network.

[0032] Compared with the existing technologies, the beneficial effects of the present invention are as follows: The present invention realizes the fusion of multiple types of data in drug research and development, greatly improves the accuracy of drug property prediction and the robustness of the system. At the same time, it optimizes the computational complexity and improves the overall efficiency. The Gaussian correlation embedding model can effectively utilize the complementary information of compound structures, biological activities, and interaction data, and shows excellent performance in complex data environments. Through shared feature representations, the present invention effectively controls the consumption of computing resources while improving performance, making it suitable for resource-constrained environments. The optimized drug screening framework processes complex embedded features more efficiently and supports more accurate and rapid drug screening. Description of the Drawings

[0033] Figure 1 is an exemplary flowchart of the present invention;

[0034] Figure 2 is an exemplary flowchart of the data feature extraction and preprocessing section of the present invention. Detailed Embodiments

[0035] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present application.

[0036] The embodiments of the present application provide a method for predicting drug activity. The core is to effectively capture and model the multi-dimensional dependencies in drug molecular features through Gaussian correlation embedding technology. By introducing the Gaussian correlation function, the multi-dimensional non-linear dependence structure can be mapped into the Gaussian distribution space, and then a unified and consistent feature embedding representation can be generated among multi-dimensional features. It can effectively integrate heterogeneous data such as the chemical structure, biological activity, and drug interaction of compounds, capture their deep dependence relationships, and thus significantly improve the accuracy and robustness of drug property prediction and candidate drug screening.

[0037] To better understand the above technical solutions, the above technical solutions will be described in detail below in conjunction with the accompanying drawings of the specification and specific embodiments. Refer to Figure 1 As shown, this figure is an exemplary flowchart of a method for predicting drug activity according to the present embodiment of the present application. The prediction method includes the following steps:

[0038] In step S1, data feature extraction and preprocessing are performed to obtain the physicochemical property data and pharmacological experiment data in the drug data. A context - associated feature engineering method is used to capture the complex relationships between the data. The core information of the drug physicochemical property data is extracted through feature decomposition, and the pharmacological experiment data is feature - modeled by combining context and time - series characteristics.

[0039] In this embodiment, the physicochemical property data in the drug data includes melting point, boiling point, solubility, lipophilicity - hydrophilicity partition coefficient, dissociation degree, volatility, hygroscopicity, efflorescence, surface activity, and crystal form.

[0040] The pharmacological experiment data in the drug data includes biological activity, pharmacodynamic response, metabolic stability, and cytotoxicity. The drug data extracts core information through feature decomposition to form samples, performs independent feature extraction on each sample, and combines the association relationships of context nodes to form a context - associated feature representation.

[0041] It also constructs an effective context - independent feature matrix through a context model. This matrix combines the embedding vectors of each independent feature with the weight values of their corresponding nodes to generate an effective context representation for each sample. Its mathematical expression is:

[0042]

[0043] where, R n is the embedding vector of the effective context representation of each sample, W is the number of context nodes, α t,ω is the embedding vector of the corresponding node, and val ω is the value of the context node.

[0044] Reference Figure 2 As shown, based on step S1, it is specifically divided into:

[0045] S11: Obtain data, and obtain the physicochemical property data and pharmacological experiment data in the drug data.

[0046] S12: Data preprocessing and representation. In the data preprocessing stage, the drug data extracts core information through feature decomposition to form samples, performs independent feature extraction on each sample, and combines the association relationships of context nodes to form a context - associated feature representation.

[0047] S13: Calculate effective context information. An effective context - independent feature matrix is constructed through a context model. This matrix combines the embedding vectors of each independent feature with the weight values of their corresponding nodes to generate an effective context representation for each sample. Its mathematical expression is:

[0048]

[0049] where, R n is the embedding vector of the effective context representation of each sample, W is the number of context nodes, and α y,ω is the embedding vector of the corresponding node, and val ω is the value of the context node.

[0050] Moreover, in this embodiment, regarding the above steps S12 and S13, it should be noted that for each input sample, an effective feature representation is calculated through context nodes and a background embedding matrix. The features of the context nodes and the embedding vector of the central node jointly construct the relationship between nodes, and its mathematical expression is obtained by normalizing the embedding values of each context node.

[0051] It should also be noted that for feature extraction, when separately extracting features for each type of attribute, statistical methods can be used to analyze the distribution characteristics of individual attributes, and domain knowledge can be utilized to create new features, such as predicting certain biological activities by combining molecular structure information.

[0052] In the context - associated feature engineering method, by combining different types of features, the interactions between them are discovered, such as the relationship between physicochemical properties and drug efficacy. Then, deep learning techniques (such as neural networks) are applied to automatically learn useful feature representations from a large amount of raw data. Additionally, for datasets with clear interactions, a graph model can be constructed to capture the relationships between entities using the relationships between nodes and edges.

[0053] By capturing the complex relationships between data, important properties such as the activity and toxicity of drugs can be predicted more accurately. When external knowledge is fused, it helps to understand the scientific basis behind model decisions, enhances researchers' trust in the model, and efficient feature engineering methods can quickly identify potential effective candidate drugs, shortening the new drug development cycle.

[0054] Regarding step S1, through this diverse feature extraction method, the feature representation of the drug not only retains the static information of physicochemical properties but also captures the dynamic performance under experimental conditions. These multi - level features provide rich inputs for subsequent Gaussian - associated embedding, ensuring that the model can more accurately characterize drug features.

[0055] In step S2, the Gaussian - associated embedding mechanism jointly models the drug molecular features through the Gaussian - associated embedding mechanism to form a prediction model to handle the complex relationships between different modalities and heterogeneous data features;

[0056] Specifically, the context-independent feature matrix includes the physicochemical property feature matrix M1 of the drug and the feature matrix M2 of the pharmacodynamic experiment data, and the Gaussian correlation function is used to capture the statistical correlation between different features. The joint distribution of the features can be calculated by the Gaussian correlation as follows:

[0057] GC(u R ,u T )=Φ(Φ -1 (u R ),Φ - 1(u T );ρ)

[0058] Where u R and u T represent the marginal distribution values of different feature variables respectively, Φ is the cumulative distribution function of the standard normal distribution, Φ -1 represents its corresponding inverse function, and ρ is the correlation parameter between the features.

[0059] In step S3, gradient optimization and parameter update are performed. The prediction model utilizes the context information and the interaction between features to gradually optimize each parameter, making the final feature representation more accurate;

[0060] Specifically, based on the physicochemical property feature matrix M1 and the feature matrix M2 of the pharmacodynamic experiment data, the feature representation is further optimized and updated through the Gaussian correlation embedding mechanism to form a feature matrix representation based on context information. And in the context-dependent modeling process, the embedding vector R n characterizes the relationship between each central node and its context nodes to adjust the feature weights and optimize the feature representation. The formula for optimizing the feature representation is:

[0061] M conteixt =τM1+(1-τ)M2

[0062] Where τ and 1-τ are dynamically adjusted weights controlled by the Gaussian correlation embedding mechanism, used to balance the importance of different features, thereby generating a context-optimized comprehensive feature representation. M context refers to the feature matrix dynamically updated based on the context relationship.

[0063] The Gaussian correlation embedding technology can effectively capture and model the multi-dimensional dependencies in drug molecular features. By introducing the Gaussian correlation function, the present invention can map the multi-dimensional non-linear dependence structure into the Gaussian distribution space, and then generate a unified and consistent feature embedding representation among multi-dimensional features. This method solves the limitations of the prior art in multi-dimensional dependence modeling, enabling the embedded features to more truly reflect the molecular structure and drug properties.

[0064] It should be noted that the Gaussian correlation technology in the present invention is mainly used for drug activity prediction, toxicity assessment and drug similarity analysis. Compared with the existing independent modeling methods, the present invention innovatively uses correlation functions to model the complex dependencies of different features, which improves the expressive power of data embedding. This method has significant advantages in capturing the potential structure of molecular data and improving model performance. It is suitable for high-dimensional and complex dependent molecular data processing scenarios. By applying Gaussian correlation embedding to actual drug data, the present invention achieves performance improvement in drug prediction tasks, including higher prediction accuracy and more stable model performance. In addition, the present invention has strong scalability and can be used for different types of molecular data analysis, bringing new ideas and technical breakthroughs to data processing in drug research and development.

[0065] It should be noted that, in this embodiment, in each iteration of the prediction model, the feature value of each sample and its context node are traversed through the back-propagation mechanism to calculate the gradient of the model parameters;

[0066] Specifically, in the gradient calculation process, the model models the relationship between feature embedding and context vectors, considers the dependency between the central node and its context nodes, and calculates the corresponding gradient. The present invention adopts a negative sampling strategy, that is, randomly extracts some negative samples from a large number of possible context nodes for gradient calculation, and uses the Adam optimizer to gradually update and adjust the parameters based on the calculated model parameter gradients;

[0067] It should also be noted that the present invention adopts a negative sampling method to randomly extract a few negative samples from the context for gradient calculation. Compared with positive sampling, negative sampling greatly reduces the computational complexity and avoids the computational burden of a complete traversal of all context nodes, while effectively retaining the most important context information, thereby improving training efficiency and model robustness. In addition, negative sampling is more efficient when the number of samples is unbalanced, allowing the model to focus more on extracting effective information from important samples, further reducing dependence on sparse context.

[0068] Furthermore, in each iteration of the prediction model, the model complexity is reduced and the risk of overfitting is reduced through parameter sharing strategy, and the same model parameters θ are shared in all iterations. shared , which can be expressed as:

[0069] θ shared =θ1=θ2=…=θ n

[0070] Among them, θ shared represents the model parameters shared among various parts of the model or different samples, θ1, θ2, …θ nRespectively represent the same parameters used for each subtask or sample;

[0071] Through multiple iterations of gradient optimization and parameter update, the feature representation of the model is gradually optimized, and its specific expression is:

[0072]

[0073] Among them, F (t) represents the feature representation at the t-th iteration, Λ is the feature optimization function, and usually updates and optimizes the features based on the gradient backpropagation mechanism. represents the gradient at the t-th iteration, and θ shared represents the common parameters in the model.

[0074] In step S4, the drug molecule features after Gaussian correlation embedding and parameter optimization are input into the prediction model to obtain a drug functional activity prediction network;

[0075] In this embodiment, based on the integration method of multi-level feature representation, multi-scale feature integration processing is performed on the optimized feature representation. The prediction model includes a function classification module for predicting the specific function categories of drugs, and also includes an active region regression module for locating the action region of drugs in the biological system. The two form a drug functional activity prediction network.

[0076] Specifically, in this embodiment, the function classification module is used to predict the specific function categories of drugs, such as their biological targets or specific biological activities. This part of the network usually consists of several convolutional layers or fully connected layers, and finally outputs the probability of each category through the softmax layer;

[0077] Specifically, for functional activity prediction: The functional activity of drug molecules is predicted by using the logistic regression model Ψ(.). By analyzing the relationship between the embedded features and the context information, the model can more accurately predict the biological functions and action mechanisms of drugs;

[0078] Furthermore, for fine localization of the active region: By further analyzing the drug feature matrix, the active region of the drug and its binding ability to specific targets are identified. The prediction of the active region is finely located by the logistic regression model Ψ(.), which helps to identify the most active part of the drug and enhance the interpretability of its functional performance;

[0079] In addition, the active region regression module is responsible for locating the action region of drugs in the biological system and outputs four regression values, which respectively represent the central coordinates and size of the active site. Assume M context represents the context feature matrix updated dynamically. The task of the active prediction network is to convert these features into specific functional activity prediction results, and this process can be expressed as:

[0080] y class = Ψ claas (M context , κ class ), y reg = Ψ reg (M context , κ reg )

[0081] where y class represents the predicted functional class probabilities, usually a vector where each element represents the probability of a certain class, and y reg represents the predicted active region boundary coordinates, including information such as the center coordinates, width, and height. Ψ class and Ψ reg are the logistic regression models of the functional classification and region regression modules respectively, responsible for converting the fused features into prediction results, and κ class and κ reg are the network parameters of the classification and regression modules.

[0082] The present invention proposes the Gaussian correlation embedding technique to effectively capture and model the multi-dimensional dependencies in drug molecular features. By introducing the Gaussian correlation function, the present invention can map the multi-dimensional non-linear dependence structure into the Gaussian distribution space, and then generate a unified and consistent feature embedding representation among multi-dimensional features. This method solves the limitations of the prior art in multi-dimensional dependence modeling, enabling the embedded features to more truly reflect the molecular structure and drug properties.

[0083] The Gaussian correlation technique in the present invention is mainly applied to drug activity prediction, toxicity assessment, and drug similarity analysis. Compared with the existing independent modeling methods, the present invention innovatively uses the correlation function to model the complex dependencies of different features, improving the expressive ability of data embedding. This method has significant advantages in capturing the potential structure of molecular data, improving model performance, etc., and is applicable to the processing scenarios of high-dimensional and complex-dependent molecular data.

[0084] By applying the Gaussian correlation embedding to actual drug data, the present invention has achieved performance improvement in drug prediction tasks, including higher prediction accuracy and more stable model performance. In addition, the present invention has strong scalability and can be used for different types of molecular data analysis, bringing new ideas and technological breakthroughs to data processing in drug research and development.

[0085] This application is described with reference to the flowcharts and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the application. It should be understood that each flow and / or block in the flowcharts and / or block diagrams can be implemented by computer program instructions, and the combination of flows and / or blocks in the flowcharts and / or block diagrams can also be implemented. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing devices generate means for implementing the functions specified in one flow Figure 1 one flow or multiple flows and / or blocks Figure 1 or in multiple blocks.

[0086] Those of ordinary skill in the art can understand that all or part of the steps in the various methods of the above embodiments can be completed by instructing relevant hardware through a program. This program can be stored in a computer-readable storage medium, which includes read-only memory (ROM), random access memory (RAM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), one-time programmable read-only memory (OTPROM), electrically-erasable programmable read-only memory (EEPROM), compact disc read-only memory (CD-ROM), or other optical disc memories, magnetic disk memories, tape memories, or any other medium that can be used to carry or store data and is computer-readable.

[0087] It should also be noted that the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, commodity or device including a series of elements not only includes those elements, but also includes other elements not explicitly listed, or also includes elements inherent to such process, method, commodity or device. Without further limitation, an element defined by the statement "including one..." does not exclude the existence of another identical element in the process, method, commodity or device including the element.

Claims

1. A method for predicting drug activity, characterized in that: The prediction method has the following steps: Data feature extraction and preprocessing: obtaining the physicochemical property data and pharmacological experimental data from drug data, capturing the complex relationship between data based on context-related feature engineering methods, extracting the core information of drug physicochemical property data through feature decomposition, and performing feature modeling on pharmacological experimental data in combination with context and time series characteristics; Gaussian correlation embedding mechanism, which jointly models the characteristics of drug molecules to form a prediction model to handle the complex relationship between different modalities and heterogeneous data features; Gradient optimization and parameter updating: the prediction model uses the interaction between context information and features to gradually optimize each parameter, making the final feature representation more accurate. The drug molecular features after Gaussian correlation embedding and parameter optimization are input into the prediction model to obtain the drug functional activity prediction network.

2. A method for predicting drug activity according to claim 1, characterized in that: The physicochemical property data in the drug data include melting point, boiling point, solubility, lipid-water partition coefficient, degree of dissociation, volatility, hygroscopicity, weathering, surface activity and crystal form; The pharmacological experimental data in the drug data include biological activity, pharmacodynamic response, metabolic stability, and cytotoxicity.

3. A method for predicting drug activity according to claim 2, characterized in that: The drug data extracts core information and forms samples through feature decomposition, performs independent feature extraction on each sample, and combines the association relationship of context nodes to form a feature representation with context association.

4. A method for predicting drug activity according to claim 3, characterized in that: The context model is used to construct an effective context-independent feature matrix, which combines the embedding vectors of each independent feature with the weight values ​​of its corresponding nodes to generate an effective context representation for each sample. Its mathematical expression is: Among them, R n is the embedding vector of the effective context representation of each sample, W is the number of context nodes, α t,ω is the embedding vector of the corresponding node, val ω is the value of the context node.

5. A method for predicting drug activity according to claim 4, characterized in that: The context-independent feature matrix includes the feature matrix M1 of the physicochemical properties of the drug and the feature matrix M2 of the efficacy experimental data, and uses the Gaussian correlation function to capture the statistical correlation between different features. The joint distribution of the features can be calculated by the Gaussian correlation expression: GC(u R ,u T )=Φ(Φ -1 (u R ),F -1 (u T );p) Among them, u R and u T They represent the marginal distribution values ​​of different characteristic variables, Φ is the cumulative distribution function of the standard normal distribution, Φ -1 represents its corresponding inverse function, and ρ is the correlation parameter between features.

6. A method for predicting drug activity according to claim 5, characterized in that: Based on the physicochemical property feature matrix M1 and the feature matrix M2 of the efficacy experimental data, the feature representation is further optimized and updated through the Gaussian correlation embedding mechanism to form a feature matrix representation based on context information, and the embedding vector R of the effective context representation is used in the context dependency modeling process. n Characterize the relationship between each central node and its context nodes to adjust the feature weight and optimize the feature representation. The formula for feature representation optimization is: M context =τM1+(1-τ)M2 Among them, τ and 1-τ are dynamically adjusted weights controlled by the Gaussian correlation embedding mechanism, which are used to balance the importance of different features, thereby generating a comprehensive feature representation after context optimization. context Refers to the feature matrix that is dynamically updated based on the contextual relationship.

7. A method for predicting drug activity according to claim 6, characterized in that: In each iteration of the prediction model, the feature value of each sample and its context node are traversed through the back-propagation mechanism to calculate the gradient of the model parameters.

8. A method for predicting drug activity according to claim 7, characterized in that: In each iteration of the prediction model, the parameter sharing strategy is used to reduce the model complexity and reduce the risk of overfitting. The same model parameters θ are shared in all iterations. shared , which can be expressed as: i shared =θ1=θ2…=θ n Among them, θ shared represents the model parameters shared among various parts of the model or different samples, θ1, θ2, …θ n Represent the same parameters used by each subtask or sample; Through multiple iterations of gradient optimization and parameter update, the feature representation of the model is gradually optimized. The specific expression is: Among them, F (t) represents the feature representation of the tth iteration, Λ is the feature optimization function, which usually updates and optimizes the features based on the gradient back propagation mechanism. represents the gradient of the tth iteration, θ shared Represents a common parameter in the model.

9. A method for predicting drug activity according to claim 8, characterized in that: The optimized feature representation is subjected to multi-scale feature integration processing based on the integration method of multi-level feature representation. The prediction model includes a functional classification module for predicting the specific functional category of the drug, and an active region regression module for locating the action area of ​​the drug in the biological system. The two modules form a drug functional activity prediction network.

Citation Information

Cited By

  • Intelligent analysis method for bulk drug crystal form transformation risk

    CN120878010A