A method, system and storage medium for predicting the correlation between Chinese herbal medicines and genes
By converting Chinese herbal ingredients into molecular maps and optimizing them, combining the k-mers maps of gene sequences and multi-level features, the problem of missing structural information in the prediction of the related relationship of traditional Chinese herbal medicines and genes is solved, and a more accurate and universal correlation prediction is achieved.
Patent Information
- Application Number
- CN202510662375.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-22
- Publication Date
- 2025-07-25
- Estimated Expiration
- 2045-05-22
AI Technical Summary
The prior art Chinese herbal medicine-gene association prediction methods ignore the structural information of Chinese herbal medicine and genes, resulting in insufficient accuracy of association relationships and poor universality.
By converting Chinese herbal ingredients into molecular maps, initially into molecular structural features are generated, and substructure perception network is used for optimization, combining k-mers plots and multi-level feature fusion of gene sequences, the joint probability distribution between component embedding representations and gene embedding representations is calculated, and the association relationship between Chinese herbal medicines and genes is predicted.
It improves the accuracy and universality of prediction of Chinese herbal medicine-gene association relationships, reduces research costs and time consumption, and provides systematic association relationship prediction results.
Smart Images

Figure CN120196962B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of bioinformatics technology, and more specifically, to a method, system and storage medium for predicting the association relationship between Chinese herbal medicines and genes. Background Art
[0002] As an important part of traditional Chinese medicine in China, Chinese herbal medicines not only inherit the profound cultural heritage of China, but also are an important supplement to the modern medical system. Modern pharmacological research shows that the active ingredients of Chinese herbal medicines often play their therapeutic roles by regulating the expression or function of specific genes.
[0003] At present, the research on predicting the association relationship between Chinese herbal medicines and genes mainly focuses on network topology analysis methods. However, only through the method of network topology analysis, the structural information of Chinese herbal medicines and genes is often ignored, resulting in insufficient accuracy of the predicted association relationship between Chinese herbal medicines and genes. Moreover, the existing methods mainly focus on specific Chinese herbal medicines. Although this method can uncover the action mechanisms of specific Chinese herbal medicines, Chinese herbal medicines usually contain multiple active ingredients, and these ingredients may act through multiple targets and multiple pathways. Studying a single Chinese herbal medicine or ingredient in isolation is difficult to systematically reveal the overall association between Chinese herbal medicines and genes, and the universality is poor. Summary of the Invention
[0004] In view of this, the present invention provides a method, system and storage medium for predicting the association relationship between Chinese herbal medicines and genes, which fully considers the structural information of Chinese herbal medicine components and genes, and the prediction results of the association relationship are more accurate, improving the prediction efficiency and having high universality.
[0005] To achieve the above object, the following solutions are proposed:
[0006] A method for predicting the association relationship between Chinese herbal medicines and genes, comprising:
[0007] Obtaining Chinese herbal medicine components and gene sequences;
[0008] Converting Chinese herbal medicine components into molecular graphs to generate initial molecular structural features;
[0009] Optimizing the initial molecular structural features through a substructure-aware network to obtain target molecular structural features;
[0010] Dividing the gene sequence into several subsequences, and constructing a k-mers graph according to the gene subsequences to obtain gene subsequence features;
[0011] Through multi-level feature fusion, aggregating the target molecular structural features and the gene subsequence features to obtain component embedding representations and gene embedding representations;
[0012] Calculate the joint probability distribution between the computational component embedding representation and the gene embedding representation to obtain the predicted relationship between Chinese herbal medicines and genes.
[0013] Preferably, the process of generating the initial substructure features of the component includes:
[0014] Generate an undirected graph according to the molecular graph through the component substructure adaptive extraction network to obtain the atomic nodes of the component;
[0015] Aggregate the features of all adjacent nodes within H-hop of the atoms in the molecular graph of the Chinese herbal medicine component through an aggregation function to obtain the initial substructure features of the component.
[0016] Preferably, the process of aggregating the features of all adjacent nodes within H-hop of the atoms in the molecular graph of the Chinese herbal medicine component through an aggregation function includes:
[0017] For the atomic nodes of the component molecule, the substructure embedding representation composed of its neighbor nodes is:
[0018] ;
[0019] Where is the initial substructure feature of the atomic node , represents the set composed of the atomic nodes of the given node and H all adjacent nodes within the jump, is the atomic node of the component molecule graph G, is the number of hidden layers of the component substructure adaptive extraction network MPNN, represents the coefficient of the nodes within the jump number, is the feature of the central atom based on different jump numbers; , ;
[0020] Where is the feature representation of the atom in the previous layer, , and respectively represent the message aggregation function, the aggregation function and the update operation of MPNN.
[0021] Preferably, the process of optimizing the initial substructure features of the component through the substructure perception network includes:
[0022] Integrate the initial substructure features of the component through the self-attention mechanism in the substructure perception network:
[0023] ;
[0024] Among them, is the component sub - structure feature information, is the topology - aware function centered on the atomic node , is the atomic node centered topology - aware function, is the atomic node graph kernel function composed of the adjacency linear representation functions of, is the atomic node linear transformation function of the absolute position encoding;
[0025] The component sub - structure feature information and the atomic nodes are aggregated through the multi - head self - attention mechanism in the sub - structure perception network to obtain the aggregated feature information:
[0026] ;
[0027] Among them, is the aggregated feature information of the atomic node , is the output of the th attention head in the self - attention mechanism;
[0028] The atomic node features and the aggregated feature information are fused to obtain the graph attention features of the atomic nodes:
[0029] ;
[0030] Among them, is the graph attention feature of the atomic node at the th layer, , is the number of network layers, is the atomic node feature at the th layer of the graph attention feature, represents the normalization in the sub - structure perception network, represents the atomic node at the th layer of the aggregated feature, is the training parameter, represents the feature dimension of the node ;
[0031] The target component sub - structure features are generated through the feed - forward neural network in the sub - structure perception network based on the atomic node features, the initial component sub - structure features, and the graph attention features:
[0032] ;
[0033] Among them, is the weight matrix in the feedforward neural network, is the residual term of the feedforward neural network.
[0034] Preferably, the process of extracting the gene subsequence features includes:
[0035] Segment the gene sequence into a number of length-k-mer subsequences that make up the k-mers group through gap k-mer encoding;
[0036] Calculate the frequency of occurrence of each k-mer respectively to obtain the k-mers graph;
[0037] Characterize the nodes in the k-mers graph through a graph neural network and perform feature aggregation on the nodes to obtain the gene subsequence features.
[0038] Preferably, the process of characterizing the nodes in the k-mers graph through a graph neural network and performing feature aggregation on the nodes includes:
[0039] Characterize the nodes in the k-mers graph through a graph neural network;
[0040] Aggregate the features of the neighbor nodes of each node to obtain the aggregated feature of the node:
[0041] ;
[0042] where, represents the aggregated feature of node at the -th layer, is the set of adjacent nodes of node , and are respectively and the degrees of the nodes, represents the feature of node at the -th layer, is the weight matrix for linearly transforming the features;
[0043] Enhance the node representation through linear transformation and non-linear activation function to obtain the gene subsequence features:
[0044] ;
[0045] where, represents the representation matrix of node at the -th layer, represents the non-linear activation function, specifically , by performing weighted aggregation and non - linear transformation on adjacent nodes, a new feature representation of the current layer nodes can be obtained , and the gene subsequence features are extracted.
[0046] Preferably, the aggregation of the target component sub - structure features and the gene subsequence features through multi - level feature fusion includes:
[0047] Calculating the interaction probability between each target component sub - structure feature and the gene subsequence feature to obtain a component sub - structure descending matrix, a sub - structure importance coefficient matrix, a gene subsequence descending matrix, and a gene subsequence importance coefficient matrix;
[0048] Calculating the component embedding representation based on the component sub - structure descending matrix and the sub - structure importance coefficient matrix;
[0049] Calculating the gene embedding representation based on the gene subsequence descending matrix and the gene subsequence importance coefficient matrix.
[0050] Preferably, it includes: the process of calculating the interaction probability between each target component sub - structure feature and the gene subsequence feature to obtain each matrix, including:
[0051] Constructing a triple between the target component sub - structure feature, the gene subsequence feature, and the association relationship: , where is the target component sub - structure feature, is the gene subsequence feature, is the association relationship between Chinese herbal medicine and gene;
[0052] Calculating the interaction probability between each component sub - structure and the sub - gene sequence in different Chinese herbal medicine - gene combinations through a cross - probability function:
[0053] {\beta}^{p}_{{N}_{i}}={∑}^{Q}_{q=1}{∑}^{{S}_{D}}_{i=1}{∑}^{{S}_{T}}_{j=1}\gamma \left [ {{R}_{pq}\cdot \sigma \left ( {\ast \cdot \ast} \right )-\left ( {1-{R}_{pq}} \right )\cdot \left ( {1-\sigma \left ( {\ast \cdot \ast} \right )} \right )} \right ] ;
[0054] ;
[0055] ;
[0056] wherein, is the number of combinations between Chinese herbal medicines and genes, is the number of molecular structure features of the target component, is the number of gene subsequence features, is the number of association relationships, is from denotes the probability function, and respectively represent the molecular structure features of the target component and the gene subsequence features, and are weight matrices;
[0057] Calculate the importance coefficients of each molecular structure feature of the target component and the gene subsequence feature in different Chinese herbal medicine-gene combinations, and sort the sub-structure matrix and the sub-sequence matrix in descending order according to the obtained importance coefficients:
[0058] ;
[0059] ;
[0060] wherein, is the molecular structure descending matrix, is the sub-structure importance coefficient matrix, is the gene subsequence descending matrix, is the gene subsequence importance coefficient matrix.
[0061] A Chinese herbal medicine-gene association relationship prediction system, comprising:
[0062] A data acquisition module for obtaining Chinese herbal medicine components and gene sequences;
[0063] A sub-structure feature extraction module for converting Chinese herbal medicine components into molecular graphs and generating initial molecular structure features;
[0064] A sub-structure feature optimization module for optimizing the initial molecular structure features through a sub-structure perception network to obtain the target molecular structure features;
[0065] A gene sequence feature extraction module for splitting gene sequences into several sub-sequences and constructing k-mers graphs according to the gene subsequences to obtain gene subsequence features;
[0066] A feature fusion module for aggregating the target molecular structure features and the gene subsequence features through multi-level feature fusion to obtain component embedding representations and gene embedding representations;
[0067] The association relationship prediction module is used to calculate the joint probability distribution between the component embedding representation and the gene embedding representation, and obtain the predicted relationship between the Chinese herbal medicine and the gene.
[0068] A storage medium stores a computer program thereon. When the computer program is executed by a processor, each step of the aforementioned Chinese herbal medicine-gene association relationship prediction method is implemented.
[0069] According to the specific embodiments provided by the present invention, the present invention discloses the following technical effects:
[0070] The Chinese herbal medicine-gene association relationship prediction method provided by the present invention first obtains Chinese herbal medicine components and gene sequences; converts the Chinese herbal medicine components into molecular graphs to generate initial molecular sub-structure features; optimizes the initial molecular sub-structure features through a sub-structure perception network to obtain target molecular sub-structure features; divides the gene sequences into several sub-sequences, and constructs a k-mers graph according to the gene sub-sequences to obtain gene sub-sequence features; aggregates the target molecular sub-structure features and the gene sub-sequence features through multi-level feature fusion to obtain component embedding representations and gene embedding representations; calculates the joint probability distribution between the component embedding representations and the gene embedding representations to obtain the predicted relationship between the Chinese herbal medicine and the gene. The Chinese herbal medicine-gene association relationship prediction method based on core sub-structure perception of the present invention not only extracts the features of gene sequences, but also extracts the structural features of Chinese herbal medicine components, fully considers the structural information of Chinese herbal medicine components and genes, the prediction result of the association relationship is more accurate, improves the prediction efficiency, and has high universality. Description of the Drawings
[0071] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are only the embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained according to the provided drawings without creative efforts.
[0072] Figure 1 It is a flowchart of a Chinese herbal medicine-gene association relationship prediction method provided by an embodiment of the present invention;
[0073] Figure 2 It is a schematic structural diagram of a Chinese herbal medicine-gene association relationship prediction system provided by an embodiment of the present invention. Detailed Embodiments
[0074] Next, in combination with the accompanying drawings in the embodiments of the present invention, the technical solutions in the embodiments of the present invention will be clearly and completely described. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.
[0075] First, in combination with Figure 1 a method for predicting the association relationship between Chinese herbal medicines and genes provided by the embodiments of the present invention will be introduced. As Figure 1 shown, the prediction method includes:
[0076] Step S01, obtain Chinese herbal medicine components and gene sequences.
[0077] Specifically, the data of the Chinese herbal medicine components in the Chinese herbal medicine formula can be obtained from databases such as TCMSP and TCMBank, and the gene expression data can be obtained from the NCBI database. The obtained Chinese herbal medicine component data and gene expression data are de-duplicated and cleaned to reduce data redundancy and eliminate homology bias. Among them, the Chinese herbal medicine components are represented by SMILES molecular strings, and the genes are represented by base sequences. Provide a training data set and initial data support for predicting the association relationship between Chinese herbal medicines and genes;
[0078] In addition, the association relationship between Chinese herbal medicines and genes can also be constructed based on the TCMBank database. The information on the interaction between Chinese herbal medicines and genes is represented by the numerical values of the corresponding interaction relationships between Chinese herbal medicines and genes.
[0079] Step S02, convert the Chinese herbal medicine components into molecular graphs to generate initial molecular structure features.
[0080] Specifically, the SMILES of the Chinese herbal medicine components are converted into molecular graphs through the RDKit tool.
[0081] Then, an undirected graph is generated according to the molecular graph through the molecular structure adaptive extraction network to obtain the atomic nodes of the components. Among them, the molecular structure adaptive extraction network is constructed based on MPNN (Message Passing Neural Network). The present invention uses MPNN to generate the embedding representation of the nodes, and the molecular structure adaptive extraction network represents the topological structure of the Chinese herbal medicine components with an undirected graph where, is the set of atomic nodes in the component molecular graph, is the adjacency matrix representing the adjacency between atomic nodes.
[0082] Finally, the features of all adjacent nodes within H-hop of the atoms in the molecular graph of the components are aggregated through an aggregation function. The aggregation function is used to aggregate the atomic nodes of The features of all adjacent nodes within the jump are used for entity representation to obtain the initial component substructure features For the atomic nodes of a given component molecule , from its The initial component substructure features composed of the jump neighbor nodes The embedding representation is:
[0083] ;
[0084] Among them, is the initial component substructure feature of the atomic node , represents the set composed of the atomic nodes of the given node and H all adjacent nodes within the jump, is the atomic node of the component molecule graph G, is the number of hidden layers of the component substructure adaptive extraction network MPNN, represents the coefficient of the nodes within the jump number, is the feature of the central atom based on different jump numbers; .
[0085] ;
[0086] Among them, is the feature representation of the atom in the previous layer, , and respectively represent the message aggregation function, the aggregation function, and the update operation of the MPNN.
[0087] Step S03, optimize the initial component substructure features through the substructure perception network to obtain the target component substructure features.
[0088] Specifically, using the initial component substructure features obtained in step S02 as the input, through the constructed substructure perception network. The substructure perception network can be a Transformer network. Further extract the Chinese herbal medicine component substructure through the graph self-attention mechanism and the feed-forward neural network and generate the optimized target component substructure feature representation to improve the prediction performance of the Chinese herbal medicine-gene association relationship prediction.
[0089] The Transformer network introduces a graph Transformer with a graph self-attention mechanism as the backbone network to further integrate the feature information of the nodes. Integrate the initial component substructure features through the self-attention mechanism in the substructure perception network:
[0090] ;
[0091] Among them, is the component sub - structure feature information, is a parameterized asymmetric exponential function, is a topology - aware function centered on the atomic node ; is the atomic node is a topology - aware function centered on; is the atomic node is the graph kernel function composed of the adjacency linear representation functions of the atomic node is the atomic node is the linear transformation function of the absolute position encoding of the atomic node.
[0092] Based on the multi - head self - attention mechanism of the Transformer network encoder, the feature information in the self - attention mechanism is calculated with the aggregation of atoms. Through the multi - head self - attention mechanism in the sub - structure perception network, the component sub - structure feature information is aggregated with the atomic nodes to obtain the aggregated feature information:
[0093] ;
[0094] Among them, is the aggregated feature information of the atomic node , is the output of the th attention head in the self - attention mechanism.
[0095] The aggregated feature information that fuses the feature representation of atoms and the graph self - attention example features generates graph attention features. Calculate the graph self - attention feature of the th layer of the atomic node . The atomic node features are fused with the aggregated feature information to obtain the graph attention features of the atomic node:
[0096] ;
[0097] Among them, is the graph attention feature of the atomic node at the th layer, , is the number of network layers, is the atomic node at the th layer of the graph attention feature, represents the normalization in the sub - structure perception network, represents the aggregated feature of the atomic node at the th layer, is a training parameter, represents the node feature dimension.
[0098] Finally, according to the atomic node features, the initial component substructure features, and the graph attention features, the target component substructure features are generated through the feed-forward neural network in the substructure-aware network. In the substructure-aware network (Transformer network), , and are fed into the feed-forward neural network (FFN), and through the residual connection and the normalization layer, the target component substructure feature embedding representation is generated:
[0099] ;
[0100] where, is the trainable weight matrix in the feed-forward neural network, is the residual term of the feed-forward neural network.
[0101] Step S04: The gene sequence is segmented into several subsequences, and a k-mers graph is constructed according to the gene subsequences to obtain gene subsequence features.
[0102] Specifically, for the gene sequence, gap k-mer encoding can be used to segment it into multiple k-mer subsequences of length . Then, according to the base sequence in the gene, the k-mers group is used to construct a k-mers graph. Calculate the frequency of occurrence of each k-mer to form the k-mers graph , represents the node set, and the feature information is determined by the frequency of occurrence of its respective k-mer, is the connected edge set, and the feature information is determined by the frequency of co-occurrence of two k-mers. A graph neural network is used to represent the nodes in the k-mers graph and the node representation is used as the embedding representation of the gene subsequence features. Each node aggregates the neighbor node features to obtain the aggregated feature at the layer. The calculation method is:
[0103] ;
[0104] where, represents the aggregated feature of node at the layer, is the set of adjacent nodes of node , and are respectively and the degree of the node, represents the feature of the node at the layer, is the weight matrix for linearly transforming the feature.
[0105] After aggregating the features of the nodes, the node representation is enhanced through linear transformation and non - linear activation function to obtain the embedded representation of the gene subsequence feature :
[0106] ;
[0107] wherein, represents the representation matrix of the node at the layer, represents the non - linear activation function, specifically here. By performing weighted aggregation and non - linear transformation on adjacent nodes, the new feature representation of the nodes in the current layer can be obtained , thereby extracting the gene subsequence feature .
[0108] Step S05, through multi - level feature fusion, aggregate the target component sub - structure feature and the gene subsequence feature to obtain the component embedded representation and the gene embedded representation.
[0109] Specifically, calculate the interaction probability between each sub - structure and sub - gene sequence in different Chinese herbal medicine - gene combinations based on the original data to quantify the importance of the Chinese herbal medicine component sub - structure and the gene subsequence. Calculate the interaction probability between each target component sub - structure feature and the gene subsequence feature to obtain the component sub - structure descending matrix, the sub - structure importance coefficient matrix, the gene subsequence descending matrix, and the gene subsequence importance coefficient matrix. And extract the target component sub - structure feature of the Chinese herbal medicine as the graph - level embedding for information transmission and enhancement of the node representation, and segment and extract the gene subsequence feature as the graph - level embedding for feature aggregation and transformation. The process is as follows:
[0110] (1) Construct a triple between the target component sub - structure feature, the gene subsequence feature, and the association relationship: , wherein, is the target component sub - structure feature, is the gene subsequence feature, is the association relationship between the Chinese herbal medicine and the gene.
[0111] (2) Calculate the interaction probability between each component sub - structure and sub - gene sequence in different Chinese herbal medicine - gene combinations through the cross - probability function :
[0112] \(\beta^{p}_{N_{i}}=\sum_{q = 1}^{Q}\sum_{i = 1}^{S_{D}}\sum_{j = 1}^{S_{T}}\gamma\left[R_{pq}\cdot\sigma(\ast\cdot\ast)-(1 - R_{pq})\cdot(1-\sigma(\ast\cdot\ast))\right]\) ;
[0113] ;
[0114] ;
[0115] where is the number of combinations between Chinese herbal medicines and genes, is the number of structural features of the target component, is the number of gene subsequence features, is the number of association relationships, is the probability function represented by , and represent the structural features of the target component and the gene subsequence features respectively, and are weight matrices.
[0116] (3)Calculate the importance coefficients of each target component structural feature and gene subsequence feature in different Chinese herbal medicine - gene combinations, and sort the sub - structure matrix and sub - sequence matrix in descending order according to the obtained importance coefficients:
[0117] ;
[0118] ;
[0119] where is the descending matrix of component structures, is the sub - structure importance coefficient matrix, is the descending matrix of gene subsequences, is the gene subsequence importance coefficient matrix.
[0120] Perform graph - level embedding based on the descending matrix of component structures and the sub - structure importance coefficient matrix to obtain the component embedding representation :
[0121] 。
[0122] Perform graph-level embedding based on the gene subsequence descending matrix and the gene subsequence importance coefficient matrix to obtain the gene embedding representation :
[0123] 。
[0124] Step S06, calculate the joint probability distribution between the component embedding representation and the gene embedding representation to obtain the predicted relationship between the Chinese herbal medicine and the gene.
[0125] Specifically, fuse the features of the Chinese herbal medicine component molecules and the gene sequences to enhance the feature expression of the core substructures and subsequences, and provide triple data for the prediction of the Chinese herbal medicine-gene association relationship. Construct triples based on the component embedding representation, the gene subsequence embedding representation, and the association relationship 。
[0126] Judge whether the Chinese herbal medicine and the gene interact, and reconstruct the prediction of the Chinese herbal medicine-gene association relationship into a joint probability distribution represented by a function:[[]]
[0127] ;
[0128] wherein, is the predicted probability score, is the training parameter matrix in the predictor.
[0129] In addition, the binary cross-entropy loss minimization method can be used for model optimization and parameter learning, and the optimal model is selected to predict the association relationship between various Chinese herbal medicines and genes. The loss function is:[[]]
[0130] ;
[0131] wherein, is the number of data contained in the dataset.
[0132] The method for predicting the Chinese herbal medicine-gene association relationship provided by the embodiment of the present invention, based on core substructure perception, enables the association relationship prediction method to have high working efficiency, high accuracy, be easy to implement, have strong practicability, and significantly reduce the cost and time consumption in related fields.
[0133] On the premise of fully considering the structural information of Chinese herbal medicines and genes, the embodiment of the present invention constructs a method for predicting the Chinese herbal medicine-gene association relationship based on core substructure perception, provides intuitive results and clear influence mechanisms for the prediction of the Chinese herbal medicine-gene association relationship, and provides auxiliary support for the optimization of the research and development of innovative Chinese herbal medicines and the search for targets.
[0134] In the embodiments of the present invention, in the case of meeting the requirements for predicting the relationship between Chinese herbal medicines and genes, the problems of missing structural information of Chinese herbal medicines and genes, lack of systematicness, and low prediction accuracy are effectively solved, and the research costs in related fields are reduced. This has very important practical significance and application value for promoting the modernization of traditional Chinese medicine and the development of precision medicine in China.
[0135] The following describes the system for predicting the relationship between Chinese herbal medicines and genes provided by the embodiments of the present invention. The system for predicting the relationship between Chinese herbal medicines and genes described below can be correspondingly referred to the method for predicting the relationship between Chinese herbal medicines and genes described above.
[0136] First, in combination with Figure 2 , the system for predicting the relationship between Chinese herbal medicines and genes is introduced. As Figure 2 shown, the system for predicting the relationship between Chinese herbal medicines and genes may include:
[0137] A data collection module 100, configured to obtain Chinese herbal medicine components and gene sequences;
[0138] A sub-structure feature extraction module 200, configured to convert Chinese herbal medicine components into molecular graphs and generate initial component molecular structure features;
[0139] A sub-structure feature optimization module 300, configured to optimize the initial component molecular structure features through a sub-structure perception network to obtain target component molecular structure features;
[0140] A gene sequence feature extraction module 400, configured to divide gene sequences into several sub-sequences and construct k-mers graphs based on the gene sub-sequences to obtain gene sub-sequence features;
[0141] A feature fusion module 500, configured to aggregate the target component molecular structure features and the gene sub-sequence features through multi-level feature fusion to obtain component embedding representations and gene embedding representations;
[0142] A relationship prediction module 600, configured to calculate the joint probability distribution between the component embedding representations and the gene embedding representations to obtain the predicted relationship between Chinese herbal medicines and genes.
[0143] The embodiments of the present invention also provide a storage medium, which can store a program suitable for being executed by a processor, and the program is used to implement each processing flow in the foregoing solution for predicting the relationship between Chinese herbal medicines and genes.
[0144] Finally, it should also be noted that in this text, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the terms "comprising", "including" or any other variant thereof are intended to cover non-exclusive inclusion, such that a process, method, article or device comprising a series of elements not only includes those elements but also includes other elements not expressly listed, or elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "comprising an..." does not exclude the presence of additional identical elements in the process, method, article or device comprising said element.
[0145] The various embodiments in this specification are described in a progressive manner, with each embodiment focusing on the differences from other embodiments. The same or similar parts among the embodiments can be referred to each other.
[0146] The above description of the disclosed embodiments enables those skilled in the art to implement or use the present invention. Various modifications to these embodiments will be obvious to those skilled in the art, and the general principles defined herein can be implemented in other embodiments without departing from the spirit or scope of the present invention. Therefore, the present invention will not be limited to the embodiments shown herein, but will be accorded the widest scope consistent with the principles and novel features disclosed herein.
Claims
1. A method for predicting the correlation between Chinese herbal medicines and genes, characterized in that, Including: Obtain Chinese herbal medicine components and gene sequences; Convert the Chinese herbal medicine components into molecular graphs to generate initial molecular structure features; Optimize the initial molecular structure features through a substructure-aware network to obtain target molecular structure features; Divide the gene sequence into several subsequences, and construct a k-mers graph based on the gene subsequences to obtain gene subsequence features; Through multi-level feature fusion, aggregate the target molecular structure features and gene subsequence features to obtain component embedding representations and gene embedding representations; Calculate the joint probability distribution between the component embedding representation and the gene embedding representation to obtain the prediction relationship between the Chinese herbal medicine and the gene; The process of generating the initial molecular structure features includes: Generate an undirected graph according to the molecular graph through a molecular structure adaptive extraction network to obtain the atomic nodes of the components; Aggregate the features of all adjacent nodes within H-hop of the atoms in the Chinese herbal medicine component molecular graph through an aggregation function to obtain the initial molecular structure features; The process of aggregating the features of all adjacent nodes within H-hop of the atoms in the Chinese herbal medicine component molecular graph through the aggregation function includes: For the atomic nodes of the component molecule, the substructure embedding representation composed of its neighbor nodes is: Among them, χ subg (N c , G) is the initial component substructure feature of atomic node N c , represents the set composed of atomic node N of a given node c and all adjacent nodes within H hops, N c is the atomic node of the component molecular graph G, L M is the number of hidden layers of the component substructure adaptive extraction network MPNN, ε h represents the coefficient of nodes within the number of hops, is the feature of the central atom based on different numbers of hops; Among them, is the atomic N c is the feature representation of the previous layer, and M(), Agg(), and U() respectively represent the message aggregation function, the aggregation function, and the update operation of the MPNN; The process of optimizing the initial molecular structure features through the substructure-aware network includes: Integrate the initial molecular structure features through the self-attention mechanism in the substructure-aware network: Among them, ATT G is the component substructure feature information, δ(N c , N m ) is the topological perception function centered on the atomic node N c , δ(N w , N m ) is the topological perception function centered on the atomic node N w , γ(ε c ) is the graph kernel function composed of the adjacency linear representation functions of the atomic node N c ; is the linear transformation function of the absolute position encoding of the atomic node N c ; Aggregate the molecular structure feature information and atomic nodes through the multi-head self-attention mechanism in the substructure-aware network to obtain aggregated feature information: Among them, is the aggregation feature information of the atomic node N c , is the output of the L G -th attention head in the self-attention mechanism; Fuse the atomic node features and the aggregated feature information to obtain the graph attention features of the atomic nodes: Among them, is the atomic node N c is the l t -th layer of graph attention feature, where l t ∈{1,2,…L T}, and L T is the number of network layers, is the atomic node feature N c is the l t -1-th layer of graph attention feature, and Norm{} represents the normalization in the sub-structure perception network, represents the aggregated feature of the atomic node N c at the l t -th layer, W V is the training parameter, represents the feature dimension of the node N c ; Generate target molecular structure features through the feed-forward neural network in the substructure-aware network according to the atomic node features, the initial molecular structure features, and the graph attention features: Among them, W F is the weight matrix in the feedforward neural network, and b is the residual term of the feedforward neural network.
2. The traditional Chinese medicine-gene association relationship prediction method according to claim 1, wherein The extraction process of the gene subsequence features includes: Divide the gene sequence into several k-mer subsequences with a length of k-mer that form a k-mers group through gap k-mer encoding; Calculate the occurrence frequency of each k-mer respectively to obtain a k-mers graph; Characterize the nodes in the k-mers graph through a graph neural network and perform feature aggregation on the nodes to obtain gene subsequence features.
3. The traditional Chinese medicine-gene association relationship prediction method according to claim 2, wherein The process of characterizing the nodes in the k-mers graph through a graph neural network and performing feature aggregation on the nodes includes: Characterize the nodes in the k-mers graph through a graph neural network; Aggregate the neighbor node features of each node to obtain the aggregated features of the nodes: Among them, represents the aggregated feature of node v i on the s-th layer, N (i) is the set of adjacent nodes of node v i , d i and d j are the degrees of nodes v i and v j respectively; represents the feature of node v i on the (s - 1)-th layer, W (s-1) is the weight matrix for linearly transforming the feature; Enhance the node representation through linear transformation and non-linear activation functions to obtain gene subsequence features: Among them, represents the node v i representation matrix at the s-th layer, σ represents the non-linear activation function, specifically ReLU here. By performing weighted aggregation and non-linear transformation on adjacent nodes, a new feature representation H of the current layer nodes can be obtained i , and the gene subsequence features are extracted.
4. The traditional Chinese medicine-gene association relationship prediction method according to claim 1, wherein The aggregation of the target molecular structure features and the gene subsequence features through multi-level feature fusion includes: Calculate the interaction probability between each target molecular structure feature and the gene subsequence feature to obtain a molecular structure descending matrix, a substructure importance coefficient matrix, a gene subsequence descending matrix, and a gene subsequence importance coefficient matrix; Calculate the component embedding representation based on the molecular structure descending matrix and the substructure importance coefficient matrix; The gene embedding representation is calculated based on the gene subsequence descending matrix and the gene subsequence importance coefficient matrix.
5. The Chinese herbal medicine-gene correlation prediction method according to claim 4, wherein Including: The process of calculating the interaction probability between each target component substructure feature and the gene subsequence feature to obtain each matrix includes: Construct a triple among the substructure features of the target component, the subsequence features of the gene, and the association relationship: R(D, T, R DT ), where D is the substructure feature of the target component, T is the subsequence feature of the gene, and R DT is the association relationship between Chinese herbal medicine and genes; Calculating the interaction probability between each component substructure and the sub-gene sequence in different Chinese herbal medicine-gene combinations through a cross probability function: Among them, Q is the number of combinations between Chinese herbal medicines and genes, S D is the number of target component substructure features, S T is the number of gene subsequence features, γ is the number of association relationships, σ(*·*) is the probability function represented by softmax(), and respectively represent the target component substructure feature and the gene subsequence feature, W D and W T are weight matrices; Calculating the importance coefficients of each target component substructure feature and the gene subsequence feature in different Chinese herbal medicine-gene combinations, and sorting the substructure matrix and the subsequence matrix in descending order according to the obtained importance coefficients: Among them, D Sub is the descending matrix of component sub-structures, SI D is the sub-structure importance coefficient matrix, T Sub is the descending matrix of gene subsequences, SI T is the gene subsequence importance coefficient matrix.
6. A traditional Chinese medicine-gene association relationship prediction system, characterized in that, Including: A data acquisition module for obtaining Chinese herbal medicine components and gene sequences; A substructure feature extraction module for converting Chinese herbal medicine components into molecular graphs and generating initial component substructure features; A substructure feature optimization module for optimizing the initial component substructure features through a substructure perception network to obtain target component substructure features; A gene sequence feature extraction module for splitting gene sequences into several subsequences and constructing k-mers graphs based on the gene subsequences to obtain gene subsequence features; A feature fusion module for aggregating the target component substructure features and the gene subsequence features through multi-level feature fusion to obtain component embedding representations and gene embedding representations; An association relationship prediction module for calculating the joint probability distribution between the component embedding representation and the gene embedding representation to obtain the predicted relationship between Chinese herbal medicine and genes; The process by which the substructure feature extraction module generates initial component substructure features includes: Generating an undirected graph according to the molecular graph through a component substructure adaptive extraction network to obtain the atomic nodes of the component; Aggregating the features of all adjacent nodes within H-hop of the atoms in the Chinese herbal medicine component molecular graph through an aggregation function to obtain the initial component substructure features; The process of aggregating the features of all adjacent nodes within H-hop of the atoms in the Chinese herbal medicine component molecular graph through the aggregation function includes: For the atomic nodes of the component molecule, the substructure embedding representation composed of its neighbor nodes is: Among them, χ subg (N c , G) is the initial component sub - structure feature of atomic node N c , represents the set composed of atomic node N c of the given node and all adjacent nodes within H hops, N c is the atomic node of component molecular graph G, L M is the number of hidden layers of the component sub - structure adaptive extraction network MPNN, ε h represents the coefficient of nodes within the number of hops, is the feature of the central atom based on different numbers of hops; Among them, is the atomic N c is the feature representation of the previous layer, and M(), Agg(), and U() respectively represent the message aggregation function, the aggregation function, and the update operation of the MPNN; The process by which the substructure feature optimization module optimizes the initial component substructure features through the substructure perception network includes: Integrating the initial component substructure features through the self-attention mechanism in the substructure perception network: Among them, ATT G is the molecular structure feature information, δ(N c ,N m ) is the topological perception function centered on the atomic node N c , δ(N w ,N m ) is the topological perception function centered on the atomic node N w , γ(ε c ) is the graph kernel function composed of the adjacency linear representation functions of the atomic node N c ; is the linear transformation function of the absolute position encoding of the atomic node N c ; Aggregating the component substructure feature information and the atomic nodes through the multi-head self-attention mechanism in the substructure perception network to obtain aggregated feature information: Among them, is the aggregation feature information of the atomic node N c , is the output of the L-th attention head in the self-attention mechanism G ; Fusing the atomic node features and the aggregated feature information to obtain the graph attention features of the atomic nodes: Among them, is the atomic node N c the l t -th layer of graph attention feature, where l t ∈{1, 2, …, L T}, and L T is the number of network layers, is the atomic node feature N c the l t -1-th layer of graph attention feature, and Norm{} represents the normalization in the sub-structure perception network, represents the aggregated feature of the atomic node N c at the l t -th layer, where W V is a training parameter, represents the feature dimension of the node N c . Generating target component substructure features through a feed-forward neural network in the substructure perception network according to the atomic node features, the initial component substructure features, and the graph attention features: Among them, W F is the weight matrix in the feedforward neural network, and b is the residual term of the feedforward neural network.
7. A storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements each step of the Chinese herbal medicine-gene association relationship prediction method according to any one of claims 1-5.
Citation Information
Patent Citations
Gene sequence for lotus root chalcone synthase and its use
CN1603412A
Characterization of interactions between compounds and polymers using negative pose data and model conditioning
US20240395364A1