Chinese herbal medicine-gene association relationship prediction method and system and storage medium
By obtaining and processing the structural information of Chinese herbal ingredients and gene sequences, and using sub-structure perception networks and graph neural networks for feature extraction and fusion, the problems of insufficient accuracy and poor universality in the existing methods are solved, and more accurate and efficient prediction of Chinese herbal-gene association relationships are achieved.
Patent Information
- Application Number
- CN202510662375.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-22
- Publication Date
- 2025-06-24
- Estimated Expiration
- 2045-05-22
AI Technical Summary
The existing Chinese herbal medicine-gene association prediction method ignores the structural information of Chinese herbal medicine and genes, resulting in insufficient prediction accuracy and poor universality, making it difficult to systematically reveal the overall relationship between Chinese herbal medicine and genes.
By obtaining the ingredients and gene sequences of Chinese herbal medicines, the components are converted into molecular maps, and the initial molecular structure characteristics are generated, and the target molecular structure characteristics are obtained through the optimization of the substructure perception network. At the same time, the gene sequence was divided to construct a k-mers map, the gene sequence characteristics were extracted, and the component embedding representation and gene embedding representation were obtained through multi-level feature fusion, and their joint probability distribution was calculated to predict the association relationship between Chinese herbal medicine and genes.
It improves the accuracy and efficiency of predicting the relationship between Chinese herbal medicine-gene, has high universality, and can more accurately reveal the overall relationship between Chinese herbal medicine and genes.
Smart Images

Figure CN120196962A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of bioinformatics technology, and more specifically, to a method, system and storage medium for predicting the association relationship between Chinese herbal medicines and genes. Background Art
[0002] As an important part of traditional Chinese medicine in China, Chinese herbal medicines not only inherit the profound cultural heritage of China but also are an important supplement to the modern medical system. Modern pharmacological research shows that the active ingredients of Chinese herbal medicines often exert their therapeutic effects by regulating the expression or function of specific genes.
[0003] Currently, the research on predicting the association relationship between Chinese herbal medicines and genes mainly focuses on network topology analysis methods. However, only through network topology analysis methods, the structural information of Chinese herbal medicines and genes is often ignored, resulting in insufficient accuracy of the predicted association relationship between Chinese herbal medicines and genes. Moreover, the existing methods mainly focus on specific Chinese herbal medicines. Although this method can uncover the action mechanisms of specific Chinese herbal medicines, Chinese herbal medicines usually contain multiple active ingredients, and these ingredients may exert their effects through multiple targets and multiple pathways. Studying a single Chinese herbal medicine or ingredient in isolation is difficult to systematically reveal the overall association between Chinese herbal medicines and genes, and the universality is poor. Summary of the Invention
[0004] In view of this, the present invention provides a method, system and storage medium for predicting the association relationship between Chinese herbal medicines and genes, which fully considers the structural information of Chinese herbal medicine components and genes, the prediction results of the association relationship are more accurate, the prediction efficiency is improved, and it has high universality.
[0005] To achieve the above object, the following solutions are proposed: A method for predicting the association relationship between Chinese herbal medicines and genes, comprising: Obtaining Chinese herbal medicine components and gene sequences; Converting the Chinese herbal medicine components into a molecular graph to generate initial molecular structural features; Optimizing the initial molecular structural features through a substructure-aware network to obtain target molecular structural features; Dividing the gene sequence into several subsequences and constructing a k-mers graph based on the gene subsequences to obtain gene subsequence features; Through multi-level feature fusion, aggregating the target molecular structural features and the gene subsequence features to obtain component embedding representations and gene embedding representations; Calculating the joint probability distribution between the component embedding representations and the gene embedding representations to obtain the predicted relationship between the Chinese herbal medicine and the gene.
[0006] Preferably, the process of generating the initial molecular structural features includes: Generate an undirected graph according to the molecular graph through the component substructure adaptive extraction network to obtain the atomic nodes of the components. Aggregate the features of all adjacent nodes within H-hop of the atoms in the molecular graph of Chinese herbal medicine components through an aggregation function to obtain the initial component substructure features.
[0007] Preferably, the process of aggregating the features of all adjacent nodes within H-hop of the atoms in the molecular graph of Chinese herbal medicine components through an aggregation function includes: For the atomic nodes of the component molecules, the substructure embedded by its neighbor nodes is represented as: ; Among them, is the initial component substructure feature of the atomic node , represents the set composed of all adjacent nodes within the jump of the atomic node and H of the given node, is the atomic node of the component molecular graph G, is the number of hidden layers of the component substructure adaptive extraction network MPNN, represents the coefficient of the nodes within the jump, is the feature of the central atom based on different jumps; , ; Among them, is the feature representation of the atom in the previous layer, , and respectively represent the message aggregation function, the aggregation function and the update operation of MPNN.
[0008] Preferably, the process of optimizing the initial component substructure features through the substructure perception network includes: Integrate the initial component substructure features through the self-attention mechanism in the substructure perception network: ; Among them, is the component substructure feature information, is the topology perception function centered on the atomic node , is the topology perception function centered on the atomic node , is the graph kernel function composed of the adjacency linear representation functions of the atomic node , is the linear transformation function of the absolute position encoding of the atomic node ; Aggregate the molecular substructure feature information and atomic nodes through the multi-head self-attention mechanism in the substructure-aware network to obtain aggregated feature information: ; Among them, is the aggregated feature information of the atomic node , is the output of the -th attention head in the self-attention mechanism; Fuse the atomic node features and the aggregated feature information to obtain the graph attention features of the atomic nodes: ; Among them, is the graph attention feature of the atomic node at the -th layer, , is the number of network layers, is the graph attention feature of the atomic node feature at the -th layer, represents the normalization in the substructure-aware network, represents the aggregated feature of the atomic node at the -th layer, is the training parameter, represents the feature dimension of the node ; Generate the target molecular substructure feature through the feed-forward neural network in the substructure-aware network according to the atomic node features, the initial molecular substructure features, and the graph attention features: ; Among them, is the weight matrix in the feed-forward neural network, is the residual term of the feed-forward neural network.
[0009] Preferably, the process of extracting the gene subsequence features includes: Segment the gene sequence into several k-mer subsequences with a length of k-mer that form a k-mers group through gap k-mer encoding; Calculate the frequency of occurrence of each k-mer respectively to obtain a k-mers graph; Characterize the nodes in the k-mers graph through a graph neural network and perform feature aggregation on the nodes to obtain gene subsequence features.
[0010] Preferably, the process of characterizing the nodes in the k-mers graph through a graph neural network and performing feature aggregation on the nodes includes: The nodes in the k-mers graph are represented by a graph neural network; Aggregate the features of the neighbor nodes of each node to obtain the aggregated feature of the node: ; Among them, represents the aggregated feature of node at the th layer, is the set of adjacent nodes of node , and are respectively and the degrees of the nodes, represents the feature of node at the th layer, is the weight matrix for linearly transforming the features; Enhance the node representation through linear transformation and non-linear activation function to obtain the gene subsequence feature: ; Among them, represents the representation matrix of node at the th layer, represents the non-linear activation function, specifically here. By weighted aggregation and non-linear transformation of adjacent nodes, a new feature representation of the current layer node can be obtained, and the gene subsequence feature can be extracted.
[0011] Preferably, the aggregation of the target component substructure feature and the gene subsequence feature through multi-level feature fusion includes: Calculate the interaction probability between each target component substructure feature and the gene subsequence feature to obtain a component substructure descending matrix, a substructure importance coefficient matrix, a gene subsequence descending matrix, and a gene subsequence importance coefficient matrix; Calculate the component embedding representation based on the component substructure descending matrix and the substructure importance coefficient matrix; Calculate the gene embedding representation based on the gene subsequence descending matrix and the gene subsequence importance coefficient matrix.
[0012] Preferably, it includes: the process of calculating the interaction probability between each target component substructure feature and the gene subsequence feature to obtain each matrix includes: Construct a triple between the target component substructure feature, the gene subsequence feature, and the association relationship: Among them, is the target component substructure feature, It is a gene subsequence feature, It is the association relationship between Chinese herbal medicine and genes; Calculate the interaction probability between each component substructure and the gene subsequence in different Chinese herbal medicine-gene combinations through the crossover probability function: {\beta}^{p}_{{N}_{i}}={∑}^{Q}_{q=1}{∑}^{{S}_{D}}_{i=1}{∑}^{{S}_{T}}_{j=1}\gamma \left [ {{R}_{pq}\cdot \sigma \left ( {\ast \cdot \ast} \right )-\left ( {1-{R}_{pq}} \right )\cdot \left ( {1-\sigma \left ( {\ast \cdot \ast} \right )} \right )} \right ] ; ; ; Among them, is the number of combinations between Chinese herbal medicine and genes, is the number of target component substructure features, is the number of gene subsequence features, is the number of association relationships, is from represents the probability function, and respectively represent the target component substructure feature and the gene subsequence feature, and are weight matrices; Calculate the importance coefficients of each target component substructure feature and gene subsequence feature in different Chinese herbal medicine-gene combinations, and sort the substructure matrix and the subsequence matrix in descending order according to the obtained importance coefficients: ; ; Among them, is the component substructure descending matrix, is the substructure importance coefficient matrix, is the gene subsequence descending matrix, is the gene subsequence importance coefficient matrix.
[0013] A prediction system for the association relationship between Chinese herbal medicine and genes, including: A data acquisition module for obtaining Chinese herbal medicine components and gene sequences; The substructure feature extraction module is used to convert Chinese herbal medicine components into molecular graphs and generate initial molecular substructure features; The substructure feature optimization module is used to optimize the initial molecular substructure features through a substructure-aware network to obtain target molecular substructure features; The gene sequence feature extraction module is used to divide a gene sequence into several subsequences and construct a k-mers graph based on the gene subsequences to obtain gene subsequence features; The feature fusion module is used to aggregate the target molecular substructure features and the gene subsequence features through multi-level feature fusion to obtain component embedding representations and gene embedding representations; The association relationship prediction module is used to calculate the joint probability distribution between the component embedding representations and the gene embedding representations to obtain the predicted relationship between the Chinese herbal medicine and the gene.
[0014] A storage medium stores a computer program thereon. When the computer program is executed by a processor, each step of the aforementioned Chinese herbal medicine-gene association relationship prediction method is implemented.
[0015] According to the specific embodiments provided by the present invention, the present invention discloses the following technical effects: The Chinese herbal medicine-gene association relationship prediction method provided by the present invention, first, obtains Chinese herbal medicine components and gene sequences; converts the Chinese herbal medicine components into molecular graphs and generates initial molecular substructure features; optimizes the initial molecular substructure features through a substructure-aware network to obtain target molecular substructure features; divides the gene sequence into several subsequences and constructs a k-mers graph based on the gene subsequences to obtain gene subsequence features; aggregates the target molecular substructure features and the gene subsequence features through multi-level feature fusion to obtain component embedding representations and gene embedding representations; calculates the joint probability distribution between the component embedding representations and the gene embedding representations to obtain the predicted relationship between the Chinese herbal medicine and the gene. The Chinese herbal medicine-gene association relationship prediction method based on core substructure perception of the present invention not only extracts the features of the gene sequence, but also extracts the structural features of the Chinese herbal medicine components, fully considers the structural information of the Chinese herbal medicine components and the genes, the prediction result of the association relationship is more accurate, improves the prediction efficiency, and has high universality. Description of the Drawings
[0016] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are only the embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained according to the provided drawings without creative efforts.
[0017] Figure 1 Flow chart of a traditional Chinese medicine-gene association relationship prediction method provided by an embodiment of the present invention; Figure 2 Schematic structural diagram of a traditional Chinese medicine-gene association relationship prediction system provided by an embodiment of the present invention. Detailed implementation manners
[0018] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0019] First, in combination with Figure 1 a traditional Chinese medicine-gene association relationship prediction method provided by an embodiment of the present invention is introduced. As Figure 1 shown, the prediction method includes: Step S01, obtaining traditional Chinese medicine components and gene sequences.
[0020] Specifically, the data of the traditional Chinese medicine components in the traditional Chinese medicine formula can be obtained from databases such as TCMSP and TCMBank, and the gene expression data can be obtained from the NCBI database. The obtained traditional Chinese medicine component data and gene expression data are de-duplicated and cleaned to reduce data redundancy and eliminate homology bias. Among them, the traditional Chinese medicine components are represented by SMILES molecular strings, and the genes are represented by base sequences. Provide a training data set and initial data support for the prediction of the traditional Chinese medicine-gene association relationship; In addition, a traditional Chinese medicine-gene interaction association relationship can also be constructed based on the TCMBank database. The herb-gene interaction information is represented by the numerical values of the corresponding traditional Chinese medicine-gene interaction relationships.
[0021] Step S02, converting the traditional Chinese medicine components into molecular graphs to generate initial molecular structure features.
[0022] Specifically, the SMILES of the traditional Chinese medicine components are converted into molecular graphs through the RDKit tool.
[0023] Then, an undirected graph is generated according to the molecular graph through the molecular structure adaptive extraction network to obtain the atomic nodes of the components. Among them, the molecular structure adaptive extraction network is constructed based on MPNN (Message Passing Neural Network). The present invention uses MPNN to generate the embedding representation of the nodes, and the molecular structure adaptive extraction network represents the topological structure of the traditional Chinese medicine components with an undirected graph where, is the set of atomic nodes in the component molecular graph, It represents the adjacency matrix between atomic nodes.
[0024] Finally, the features of all adjacent nodes within H-hop of the atoms in the molecular graph of the component are aggregated through an aggregation function. An aggregation function is used to aggregate the features of atomic nodes of all adjacent nodes within the hop are used for entity representation to obtain the initial substructure features of the component. For the atomic nodes of a given component molecule , from its hop neighbor nodes, the initial substructure features of the component are embedded and represented as: ; where is the initial substructure feature of the atomic node , represents the set composed of the atomic nodes of the given node and H all adjacent nodes within the hop, is the atomic node of the component molecular graph G, is the number of hidden layers of the substructure adaptive extraction network MPNN of the component, represents the coefficient of the nodes within the hop number, is the feature of the central atom based on different hop numbers; .
[0025] ; where is the feature representation of the atom in the previous layer, , and represent the message aggregation function, the aggregation function, and the update operation of MPNN respectively.
[0026] Step S03, optimize the initial substructure features of the component through a substructure perception network to obtain the target substructure features of the component.
[0027] Specifically, using the initial substructure features obtained in step S02 as the input, through the constructed substructure perception network. The substructure perception network can be a Transformer network. Further extract the substructure of the Chinese herbal medicine component through the graph self-attention mechanism and the feed-forward neural network and generate the optimized target substructure feature representation to improve the prediction performance of the Chinese herbal medicine-gene association relationship prediction.
[0028] The Transformer network introduces a graph Transformer with a graph self-attention mechanism as the backbone network to further integrate the feature information of nodes. The self-attention mechanism in the sub-structure perception network integrates the initial component sub-structure features: ; Among them, is the component sub-structure feature information, is a parameterized asymmetric exponential function, is a topology-aware function centered on the atomic node , is a topology-aware function centered on the atomic node , is a graph kernel function composed of the adjacency linear representation functions of the atomic node , is a linear transformation function of the absolute position encoding of the atomic node .
[0029] Based on the multi-head self-attention mechanism of the Transformer network encoder, the feature information in the self-attention mechanism is calculated with the aggregation of atoms. Through the multi-head self-attention mechanism in the sub-structure perception network, the component sub-structure feature information is aggregated with the atomic nodes to obtain the aggregated feature information: ; Among them, is the aggregated feature information of the atomic node , is the output of the th attention head in the self-attention mechanism.
[0030] The aggregated feature information that fuses the feature representation of atoms and the graph self-attention example features generates graph attention features. Calculate the graph self-attention feature of the th layer of the atomic node . The atomic node features are fused with the aggregated feature information to obtain the graph attention features of the atomic node: ; Among them, is the graph attention feature of the atomic node of the th layer, , is the number of network layers, is the graph attention feature of the atomic node of the th layer, represents the normalization in the sub-structure perception network, represents the atomic node The aggregation feature of the layer, is a training parameter, indicating the feature dimension of the node.
[0031] Finally, according to the atomic node features, the initial molecular substructure features, and the graph attention features, the target molecular substructure features are generated through the feed - forward neural network in the sub - structure perception network. In the sub - structure perception network (Transformer network), , and are passed into the feed - forward neural network (FFN), and through the residual connection and the normalization layer, the target molecular substructure feature embedding representation is generated: ; where is the trainable weight matrix in the feed - forward neural network, is the residual term of the feed - forward neural network.
[0032] Step S04: The gene sequence is segmented into several subsequences, and a k - mers graph is constructed based on the gene subsequences to obtain the gene subsequence features.
[0033] Specifically, for the gene sequence, it can be segmented into multiple k - mer subsequences of length that form a k - mers group using gap k - mer encoding. Then, a k - mers graph is constructed according to the base sequence in the gene. Calculate the frequency of each k - mer occurrence to form the k - mers graph , represents the node set, and the feature information is determined by the occurrence frequency of its respective k - mer, is the connected edge set, and the feature information is determined by the frequency of two k - mers appearing together. A graph neural network is used to represent the nodes in the k - mers graph and the node representation is used as the embedding representation of the gene subsequence features. Each node aggregates the neighbor node features to obtain the aggregation feature on the layer, and the calculation method is: ; where represents the aggregation feature of node on the layer, is the set of adjacent nodes of node , and are respectively and Degree of a node represents the node at the layer with features and is the weight matrix for linearly transforming the features
[0034] After feature aggregation of the nodes, the node representation is enhanced through linear transformation and non - linear activation function to obtain the embedded representation of the gene subsequence features : ; Among them represents the node at the layer's representation matrix represents the non - linear activation function, specifically here. By performing weighted aggregation and non - linear transformation on adjacent nodes, a new feature representation of the nodes in the current layer can be obtained , thereby extracting the gene subsequence features .
[0035] Step S05: Through multi - level feature fusion, aggregate the target component sub - structure features and gene subsequence features to obtain the component embedded representation and gene embedded representation
[0036] Specifically, based on the original data, calculate the interaction probability between each sub - structure and sub - gene sequence in different Chinese herbal medicine - gene combinations to quantify the importance of the Chinese herbal medicine component sub - structure and gene subsequence. Calculate the interaction probability between each target component sub - structure feature and gene subsequence feature to obtain the component sub - structure descending matrix, sub - structure importance coefficient matrix, gene subsequence descending matrix, and gene subsequence importance coefficient matrix. And extract the target component sub - structure features of the Chinese herbal medicine as the graph - level embedding for information transmission and enhancement of node representation, and segment and extract the gene subsequence features as the graph - level embedding for feature aggregation and transformation. The process is as follows (1) Construct a triple among the target component sub - structure features, gene subsequence features, and association relationships , where is the target component sub - structure feature is the gene subsequence feature is the association relationship between Chinese herbal medicine and gene
[0037] (2) Calculate the interaction probability between each component sub - structure and sub - gene sequence in different Chinese herbal medicine - gene combinations through the cross - probability function : \({\beta}^{p}_{{N}_{i}}=\sum_{q = 1}^{Q}\sum_{i = 1}^{{S}_{D}}\sum_{j = 1}^{{S}_{T}}\gamma \left [ {{R}_{pq}\cdot \sigma \left ( {\ast \cdot \ast} \right )-\left ( {1-{R}_{pq}} \right )\cdot \left ( {1-\sigma \left ( {\ast \cdot \ast} \right )} \right )} \right ]\) ; ; ; Among them, is the number of combinations between Chinese herbal medicines and genes, is the number of target component sub - structure features, is the number of gene sub - sequence features, is the number of association relationships, is the probability function represented by , and represent the target component sub - structure feature and the gene sub - sequence feature respectively, and are weight matrices.
[0038] (3) Calculate the importance coefficients of each target component sub - structure feature and gene sub - sequence feature in different Chinese herbal medicine - gene combinations, and sort the sub - structure matrix and sub - sequence matrix in descending order according to the obtained importance coefficients: ; ; Among them, is the descending matrix of component sub - structures, is the sub - structure importance coefficient matrix, is the descending matrix of gene sub - sequences, is the gene sub - sequence importance coefficient matrix.
[0039] Perform graph - level embedding based on the descending matrix of component sub - structures and the sub - structure importance coefficient matrix to obtain the component embedding representation : .
[0040] Perform graph - level embedding based on the descending matrix of gene sub - sequences and the gene sub - sequence importance coefficient matrix to obtain the gene embedding representation : 。
[0041] Step S06: Calculate the joint probability distribution between the component embedding representation and the gene embedding representation to obtain the predicted relationship between the Chinese herbal medicine and the gene.
[0042] Specifically, fuse the features of the Chinese herbal medicine component molecules and the gene sequences to enhance the feature expression of the core sub-structures and subsequences, and provide triple data for the prediction of the Chinese herbal medicine-gene association relationship. Construct triples based on the component embedding representation, the gene subsequence embedding representation, and the association relationship. 。
[0043] Judge whether the Chinese herbal medicine and the gene interact with each other, and reconstruct the prediction of the Chinese herbal medicine-gene association relationship into a joint probability distribution represented by a function: ; where is the predicted probability score, and is the training parameter matrix in the predictor.
[0044] In addition, the binary cross-entropy loss minimization method can be used for model optimization and parameter learning. Select the optimal model to predict the association relationship between multiple Chinese herbal medicines and genes. The loss function is: ; where is the number of data contained in the dataset.
[0045] The Chinese herbal medicine-gene association relationship prediction method provided by the embodiments of the present invention is based on core sub-structure perception, making the association relationship prediction method highly efficient, accurate, easy to implement, and highly practical, significantly reducing the costs and time consumption in related fields.
[0046] Under the premise of fully considering the structural information of Chinese herbal medicines and genes, the embodiments of the present invention construct a Chinese herbal medicine-gene association relationship prediction method based on core sub-structure perception, providing intuitive results and clear influence mechanisms for the prediction of Chinese herbal medicine-gene association relationships, and providing auxiliary support for the optimization of Chinese herbal medicine innovative drug research and development and target finding.
[0047] The embodiments of the present invention effectively solve the problems of missing, lack of systematicness, and low prediction accuracy of the structural information of Chinese herbal medicines and genes while meeting the prediction requirements of Chinese herbal medicine-gene association relationships, and reduce the research costs in related fields. This has very important practical significance and application value for promoting the modernization of traditional Chinese medicine and the development of precision medicine in China.
[0048] The Chinese herbal medicine-gene association relationship prediction system provided by the embodiments of the present invention will be described below. The Chinese herbal medicine-gene association relationship prediction system described below can be correspondingly referred to the Chinese herbal medicine-gene association relationship prediction method described above.
[0049] First, in combination with Figure 2 , the Chinese herbal medicine-gene association relationship prediction system will be introduced. As Figure 2 shown, the Chinese herbal medicine-gene association relationship prediction system may include: A data acquisition module 100, configured to obtain Chinese herbal medicine components and gene sequences; A sub-structure feature extraction module 200, configured to convert Chinese herbal medicine components into molecular graphs and generate initial component molecular structure features; A sub-structure feature optimization module 300, configured to optimize the initial component molecular structure features through a sub-structure perception network to obtain target component molecular structure features; A gene sequence feature extraction module 400, configured to segment gene sequences into several sub-sequences and construct k-mers graphs according to the gene sub-sequences to obtain gene sub-sequence features; A feature fusion module 500, configured to aggregate the target component molecular structure features and the gene sub-sequence features through multi-level feature fusion to obtain component embedding representations and gene embedding representations; An association relationship prediction module 600, configured to calculate the joint probability distribution between the component embedding representations and the gene embedding representations to obtain the predicted relationship between Chinese herbal medicines and genes.
[0050] The embodiments of the present invention further provide a storage medium, which can store a program suitable for execution by a processor, and the program is used to implement each processing flow in the foregoing Chinese herbal medicine-gene association relationship prediction solution.
[0051] Finally, it should also be noted that in this article, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "comprising", "including" or any other variant thereof is intended to cover a non-exclusive inclusion, so that a process, method, article or device including a series of elements not only includes those elements, but also includes other elements not expressly listed, or also includes elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "including a..." does not exclude the existence of additional identical elements in the process, method, article or device including the element.
[0052] The various embodiments in this specification are described in a progressive manner. Each embodiment focuses on the differences from other embodiments. For the same or similar parts among the various embodiments, reference can be made to each other.
[0053] The above description of the disclosed embodiments enables those skilled in the art to implement or use the present invention. Various modifications to these embodiments will be obvious to those skilled in the art. The general principles defined herein can be implemented in other embodiments without departing from the spirit or scope of the present invention. Therefore, the present invention will not be limited to these embodiments shown herein, but rather to the widest scope consistent with the principles and novel features disclosed herein.
Claims
1. A method for predicting the relationship between Chinese herbal medicines and genes, characterized in that Including: Obtain Chinese herbal medicine components and gene sequences; Convert the Chinese herbal medicine components into molecular graphs to generate initial molecular structure features; Optimize the initial molecular structure features through a substructure-aware network to obtain target molecular structure features; Divide the gene sequence into several subsequences, and construct a k-mers graph based on the gene subsequences to obtain gene subsequence features; Through multi-level feature fusion, aggregate the target molecular structure features and gene subsequence features to obtain component embedding representations and gene embedding representations; Calculate the joint probability distribution between the component embedding representation and the gene embedding representation to obtain the prediction relationship between the Chinese herbal medicine and the gene.
2. The traditional Chinese medicine-gene association relationship prediction method according to claim 1, wherein The process of generating the initial molecular structure features includes: Generate an undirected graph according to the molecular graph through a molecular structure adaptive extraction network to obtain the atomic nodes of the components; Aggregate the features of all adjacent nodes within H-hop of the atoms in the Chinese herbal medicine component molecular graph through an aggregation function to obtain the initial molecular structure features.
3. The Chinese herbal medicine-gene association relationship prediction method according to claim 2, characterized in that The process of aggregating the features of all adjacent nodes within H-hop of the atoms in the Chinese herbal medicine component molecular graph through the aggregation function includes: For the atomic nodes of the component molecule, the substructure embedding representation composed of its neighbor nodes is: ; Among them, is the initial component substructure feature of the atomic node , represents the set composed of the atomic nodes of the given node and H all adjacent nodes within the jump, is the atomic node of the component molecular graph G, is the number of hidden layers of the component substructure adaptive extraction network MPNN, represents the coefficient of the nodes within the jump number, is the feature of the central atom based on different jump numbers; , ; Among them, is an atom is the feature representation of the previous layer, , and represent the message aggregation function, the aggregation function, and the update operation of the MPNN, respectively.
4. The Chinese herbal medicine-gene association relationship prediction method according to claim 3, wherein The process of optimizing the initial molecular structure features through the substructure-aware network includes: Integrate the initial molecular structure features through the self-attention mechanism in the substructure-aware network: ; Among them, is the molecular substructure feature information, is the topology-aware function centered on the atomic node ; is the topology-aware function centered on the atomic node ; is the graph kernel function composed of the adjacency linear representation functions of the atomic node ; is the linear transformation function of the absolute position encoding of the atomic node . Aggregate the molecular structure feature information and atomic nodes through the multi-head self-attention mechanism in the substructure-aware network to obtain aggregated feature information: ; Among them, is the aggregated feature information of the atomic node , is the output of the th attention head in the self-attention mechanism; Fuse the atomic node features and the aggregated feature information to obtain the graph attention features of the atomic nodes: ; Among them, is an atomic node The graph attention feature of the -th layer, where is the network layer number, The graph attention feature of the -th layer, represents the normalization in the sub-structure perception network, The aggregated feature of the -th layer of atomic nodes, where is the training parameter, representing the feature dimension of node Generate target molecular structure features through the feed-forward neural network in the substructure-aware network according to the atomic node features, the initial molecular structure features, and the graph attention features: ; Among them, is the weight matrix in the feedforward neural network, is the residual term of the feedforward neural network.
5. The traditional Chinese medicine-gene association relationship prediction method according to claim 1, wherein The extraction process of the gene subsequence features includes: Divide the gene sequence into several k-mer subsequences with a length of k-mer that form a k-mers group through gap k-mer encoding; Calculate the occurrence frequency of each k-mer respectively to obtain a k-mers graph; Characterize the nodes in the k-mers graph through a graph neural network and perform feature aggregation on the nodes to obtain gene subsequence features.
6. The traditional Chinese medicine-gene association relationship prediction method according to claim 5, wherein The process of characterizing the nodes in the k-mers graph through the graph neural network and performing feature aggregation on the nodes includes: Characterize the nodes in the k-mers graph through the graph neural network; Aggregate the neighbor node features of each node to obtain the aggregated features of the nodes: ; Among them, represents the aggregated feature of the node at the layer, is the set of adjacent nodes of the node , and are respectively and the degrees of the nodes, represents the feature of the node at the layer, is the weight matrix for linearly transforming the feature; Enhance the node representation through linear transformation and non-linear activation functions to obtain gene subsequence features: ; Among them, represents the node in the representation matrix of the layer, represents the non-linear activation function, specifically here. By performing weighted aggregation and non-linear transformation on adjacent nodes, a new feature representation of the nodes in the current layer can be obtained, and the gene subsequence features can be extracted.
7. The method for predicting the correlation between Chinese herbal medicines and genes according to claim 4, wherein The aggregation of the target molecular structure features and the gene subsequence features through multi-level feature fusion includes: Calculate the interaction probability between each target molecular structure feature and the gene subsequence feature to obtain a molecular structure descending matrix, a substructure importance coefficient matrix, a gene subsequence descending matrix, and a gene subsequence importance coefficient matrix; Calculate the component embedding representation based on the molecular structure descending matrix and the substructure importance coefficient matrix; The gene embedding representation is calculated based on the gene subsequence descending matrix and the gene subsequence importance coefficient matrix.
8. The Chinese herbal medicine-gene association relationship prediction method according to claim 7, wherein Including: The process of calculating the interaction probability between each target component substructure feature and the gene subsequence feature to obtain each matrix, including: Constructing a triple among the target component substructure feature, the gene subsequence feature, and the association relationship: , wherein, is the target molecular substructure feature, is the gene subsequence feature, is the association relationship between Chinese herbal medicine and genes; Calculating the interaction probability between each component substructure and the sub-gene sequence in different Chinese herbal medicine-gene combinations through the cross probability function: \beta_{N_i}^p = \sum_{q = 1}^{Q}\sum_{i = 1}^{S_D}\sum_{j = 1}^{S_T}\gamma \left[R_{pq} \cdot \sigma(\ast \cdot \ast) - (1 - R_{pq}) \cdot (1 - \sigma(\ast \cdot \ast))\right] ; ; ; Among them, is the number of combinations between Chinese herbal medicines and genes, is the number of target component substructure features, is the number of gene subsequence features, is the number of association relationships, is from denotes the probability function, and respectively denote the target component substructure feature and the gene subsequence feature, and are weight matrices; Calculating the importance coefficients of each target component substructure feature and the gene subsequence feature in different Chinese herbal medicine-gene combinations, and sorting the substructure matrix and the subsequence matrix in descending order according to the obtained importance coefficients: ; ; Among them, is the descending matrix of component substructures, is the matrix of substructure importance coefficients, is the descending matrix of gene subsequences, is the matrix of gene subsequence importance coefficients.
9. A traditional Chinese medicine-gene association relationship prediction system, characterized in that, Including: A data acquisition module for obtaining Chinese herbal medicine components and gene sequences; A substructure feature extraction module for converting Chinese herbal medicine components into molecular graphs and generating initial component substructure features; A substructure feature optimization module for optimizing the initial component substructure features through a substructure perception network to obtain target component substructure features; A gene sequence feature extraction module for splitting gene sequences into several subsequences and constructing a k-mers graph based on the gene subsequences to obtain gene subsequence features; A feature fusion module for aggregating the target component substructure features and the gene subsequence features through multi-level feature fusion to obtain component embedding representations and gene embedding representations; An association relationship prediction module for calculating the joint probability distribution between the component embedding representation and the gene embedding representation to obtain the predicted relationship between Chinese herbal medicine and genes.
10. A storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements each step of the Chinese herbal medicine-gene association relationship prediction method according to any one of claims 1-8.
Citation Information
Patent Citations
Potential policy distribution for hypothesis in network
CN116324810A
Gene sequence for lotus root chalcone synthase and its use
CN1603412A
Characterization of interactions between compounds and polymers using negative pose data and model conditioning
US20240395364A1
Personalized accurate recommendation method driven by knowledge graph
WO2021103789A1