A brain region feature extraction system based on motif-driven evolution

By using a brain region feature extraction system driven by motifs, a brain region-gene directed graph is constructed, and key brain region-gene directed graphs are detected and generated. This solves the problem that the complexity of the relationship between brain regions and genes is difficult to reveal in traditional methods, and achieves high-accuracy diagnosis of brain lesions and prediction of AD risk.

CN118866317BActive Publication Date: 2025-11-21HUNAN NORMAL UNIVERSITY
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202410879688.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-07-02
Publication Date
2025-11-21
Estimated Expiration
2044-07-02

AI Technical Summary

Technical Problem

Existing brain lesion diagnostic systems are based on traditional Euclidean methods, which make it difficult to reveal the complex relationships between brain regions and genes at the network level, resulting in low diagnostic accuracy.

Method used

A brain region feature extraction system based on motif-driven evolution is adopted. An initial brain region-gene directed graph is constructed through a generator, motifs of preset brain lesions are detected, and key brain region-gene directed graphs are generated. The similarity is calculated by a discriminator to optimize the generator model parameters. Combined with the motif discovery and evolution module, subgraph partitioning, directed subgraph evolution and fusion module, key brain region features are extracted.

Benefits of technology

It improves the accuracy of brain lesion diagnosis, can accurately extract key features of pre-defined brain lesions, provides auxiliary decision support for AD treatment strategies, and reveals the evolutionary pattern in the transformation process from LMCI to AD.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118866317B_ABST
    Figure CN118866317B_ABST
Patent Text Reader

Abstract

The application provides a brain region feature extraction system based on motif-driven evolution, which comprises the following steps: acquiring brain image data and gene data of a target object, constructing an initial brain region-gene directed graph; constructing a generator based on a motif-driven evolution model, detecting a motif of a preset brain lesion from the initial brain region-gene directed graph, and generating a key brain region-gene directed graph for detecting the preset brain lesion according to the evolution result of the motif; constructing a discriminator to calculate the similarity between the generated key brain region-gene directed graph and a real key brain region-gene directed graph, and optimizing the model parameters of the generator according to the similarity. Compared with the prior art, the application can accurately extract the key features of the preset brain lesion, and provides support for accurately predicting brain lesions.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of biological information, and particularly relates to a brain region feature extraction system based on motif-driven evolution. BACKGROUND

[0002] The course of AD usually goes through the early stage of MCI, including the EMCI stage of subjective cognitive decline and the late LMCI stage. During this period, the core pathological structure of the physiological functional network is first affected, triggering a series of chain reactions, including changes in gene expression patterns and dynamic changes in brain function states, which are closely related to the gradual degradation of cognitive function, and promote the development and deterioration of AD. The existing brain lesion diagnosis system is generally constructed based on the traditional Euclidean method. However, the traditional Euclidean method is often difficult to reveal the complex relationship between brain regions and genes at the network level, resulting in low accuracy of the existing brain lesion diagnosis system. SUMMARY

[0003] The present application provides a brain region feature extraction system based on motif-driven evolution, which is used to solve the technical problem of low accuracy of the existing brain lesion diagnosis system.

[0004] To solve the above technical problems, the technical scheme provided by the present application is as follows:

[0005] A brain region feature extraction system based on motif-driven evolution, comprising:

[0006] A preprocessing module is configured to acquire brain image data and gene data of a target object, and construct an initial brain region-gene directed graph.

[0007] A generator is configured to detect motifs of a preset brain lesion from the initial brain region-gene directed graph based on a motif-driven evolution model, and generate a key brain region-gene directed graph for detecting the preset brain lesion according to the evolution result of the motifs.

[0008] A discriminator is configured to calculate the similarity between the generated key brain region-gene directed graph and the real key brain region-gene directed graph, and optimize the model parameters of the generator according to the similarity.

[0009] Preferably, the generator comprises:

[0010] A motif discovery and evolution module is configured to detect motif structures in the initial gene-brain region directed graph and perform evolution.

[0011] A subgraph division module is configured to divide subgraphs according to the influence range of the motifs.

[0012] A directed subgraph evolution module is configured to extract motif trees in each subgraph through motif convolution, and form corresponding evolution subgraphs.

[0013] The evolutionary subgraph fusion module is used to integrate these evolutionary subgraphs to form an evolved gene-brain region directed graph.

[0014] Preferably, the phantom discovery and evolution module includes:

[0015] The motif construction unit is used to detect the motifs of the initial brain region-gene directed graph and to construct the motif adjacency matrix and the motif recording matrix.

[0016] The phantom selection unit is used to select the target phantom with the largest directed topological potential from the phantom adjacency matrix and update the phantom record matrix with the target phantom.

[0017] The first update unit updates the feature information matrix of nodes within the module based on the information transmission between nodes within the module.

[0018] The second update unit updates the module adjacency matrix based on the feature information of the updated nodes;

[0019] The third update unit calculates the weight information of newly generated edges within the module and updates the weight matrix.

[0020] Preferably, the phantom building unit includes:

[0021] The first search subunit is used to find all subgraphs of a given size in the initial brain region-gene directed graph by enumeration, record their node information using the initial motif adjacency matrix, and continue enumerating until the search is completed.

[0022] Randomly generated sub-units: used to classify the subgraph into the corresponding isomorphic class, randomly generating units that are related to the initial brain region-gene. Figure One A randomized plot of sample size;

[0023] The adjacency matrix generates sub-units. The occurrence frequency of multivariate subgraphs in the random graph and the initial brain region-gene directed graph is compared. Multivariate subgraphs with a higher occurrence frequency than the random network are marked and recorded as motifs. At the same time, multivariate subgraphs that do not belong to the motifs are deleted from the adjacency matrix to obtain the final motif adjacency matrix.

[0024] The record matrix generates sub-units. Based on the found motifs, the nodes of all motifs are counted, and a motif record matrix is ​​constructed.

[0025] and / or

[0026] The phantom screening unit includes:

[0027] Computational subunit: Obtain the node features of each module from the module record matrix, calculate the out-degree module potential and in-degree module potential of each module, and calculate the directed topological potential of each module based on the out-degree module potential and in-degree module potential of each module;

[0028] comparing subunit: comparing the directed topological potentials of each motif, selecting the target motif with the largest directed topological potential, and updating the motif record matrix;

[0029] and / or

[0030] The first updating unit comprises:

[0031] Update subunit A: for calculating the correlation coefficient of the feature information of each node and other nodes based on the feature information of the updated node, and determining that there is an edge between the two nodes when the correlation coefficient of the feature information of any two nodes is greater than a preset value;

[0032] Update subunit B: for judging the node edge direction based on the feature information of the updated node and the potential theory, for any two nodes V i and V j , if the potential energy of V i is higher than that of V j , it is determined that there is an edge from V i to V j .

[0033] Preferably, the subgraph division module is used to divide the subgraph according to the influence range of each motif after evolution, taking the directed topological potential field of the calculation motif as the influence range of the motif.

[0034] Preferably, the directed subgraph evolution module comprises:

[0035] The correlation degree calculation unit is used to calculate the correlation degree of the gene nodes and the motifs in the directed subgraph and sort them to obtain a first correlation degree sequence, and store the first correlation degree sequence in a motif-node correlation matrix S (k,0) , calculate the correlation degree of the brain region nodes and the motifs in the directed subgraph and sort them to obtain a first correlation degree sequence, and store the first correlation degree sequence in a motif-node correlation matrix S k , 1 );

[0036] The motif tree construction unit is used to construct a motif tree according to the motif-node correlation matrix S k ,1), connect the gene nodes in the directed subgraph, and update a motif tree weight matrix connect the brain region nodes in the directed subgraph, and update a motif tree adjacency matrix

[0037] The motif tree evolution unit is used to transmit evolution information based on the connection structure of the motif tree, update a motif tree node feature matrix H k ), re-calculate the edges based on the updated feature information of the nodes, and update an evolution subgraph weight matrix

[0038] Preferably, the correlation degree calculation unit comprises:

[0039] an average feature calculation subunit configured to calculate the feature information MH(k) of the k-order motif based on the k-order screening of the feature information of the nodes contained in the motif, k the feature of the i-th motif is obtained by averaging the features of all nodes in the motif; the nodes are gene nodes or brain region nodes;

[0040] a motif-node correlation degree matrix construction subunit configured to calculate the correlation coefficient of each node feature and the feature information MH(k), sort the correlation coefficient corresponding to each node feature to obtain a correlation degree sequence, and store the result in a motif-node correlation degree matrix.

[0041] and / or

[0042] The motif tree construction unit comprises:

[0043] a first index construction subunit configured to take the gene node with the highest correlation degree in the first correlation degree sequence as a leaf to be included in the motif tree, and take the correlation coefficient of the included node as the weight thereof; iteration: calculate the correlation coefficient between each node not included and the nodes in the motif tree already included to obtain an auxiliary correlation degree sequence, aggregate the auxiliary correlation degree sequence and the correlation degree sequence, and select a gene node with the highest correlation degree from the aggregation as a leaf to be included in the motif tree, and take the correlation coefficient of the included node as the weight thereof; repeat the iteration step until all gene nodes are included in the motif tree, and then update the weight matrix with the weights of all gene nodes;

[0044] a second index construction subunit configured to take the brain region node with the highest correlation degree in the second correlation degree sequence as a leaf to be included in the motif tree, and take the correlation coefficient of the included node as the weight thereof; iteration: calculate the correlation coefficient between each node not included and the nodes in the motif tree already included to obtain an auxiliary correlation degree sequence, aggregate the auxiliary correlation degree sequence and the correlation degree sequence, and select a gene node with the highest correlation degree from the aggregation as a leaf to be included in the motif tree, and take the correlation coefficient of the included node as the weight thereof; repeat the iteration step until all gene nodes are included in the motif tree, and then update the weight matrix with the weights of all gene nodes;

[0045] and / or

[0046] The motif tree evolution unit comprises:

[0047] ​​​The motif tree evolution update subunit: make the nodes on the motif tree update the node feature information from top to bottom, and then update the motif tree node feature matrix H (k) , the update formula is as follows:

[0048]

[0049] , wherein H (k) represents the feature information of the nodes in the k-order motif tree, H (k-1) represents the feature information of the nodes in the k-1-order motif tree, WW T (k-1) [i] represents the weight information of the i th node in the k-1-order motif tree, H (k-1) [i] represents the feature information of the i th node in the k-1-order motif tree, represents the update information ΔH transmitted downward along the out-degree edge of the i th node, and θ represents the learnable parameter matrix; N represents the total number of nodes;

[0050] The weight evolution update subunit: based on the updated motif tree node feature matrix H (k) , the edges and weights between nodes are recalculated to form an evolution subgraph, and the evolution subgraph weight matrix is updated based on the evolution subgraph weight matrix, wherein the calculation formula is as follows:

[0051]

[0052] , wherein, represents the feature information of the i th node after the k th information propagation, represents the feature information of the j th node after the t th information propagation, represents the weight of the edge between the i th node and the j th node after the k th information propagation.

[0053] Preferably, the evolution subgraph fusion module fuses the evolution subgraphs corresponding to each motif in the k-1-order gene-brain region directed graph to form a k-order gene-brain region directed graph, and the calculation formula is as follows:

[0054]

[0055] , wherein, represents all evolution subgraphs formed by the i th subject in the k th evolution, Fuse() represents a graph fusion function; k is the number of evolutions in the generator; N is the total number of evolutions in the generator; M is the total number of evolution subgraphs; D (i) is the directed graph after evolution of the i th subject; D is the set of directed graphs after evolution of all subjects.

[0056] Preferably, the discriminator comprises:

[0057] a directed graph convolutional layer, configured to extract a feature matrix of the generated key brain region-gene directed graph and a real key brain region-gene directed graph;

[0058] a full connection layer, configured to calculate a similarity of the feature matrix of the generated key brain region-gene directed graph and the real key brain region-gene directed graph, and determine the authenticity of the generated key brain region-gene directed graph according to the similarity;

[0059] and / or

[0060] a loss function of the brain region feature extraction system is as follows:

[0061]

[0062] wherein the generator and the discriminator are represented by G and D respectively; E(·) represents an expected value of a variable; x represents a real gene-brain region directed graph; P data (x) represents a distribution of data x, and D(x) represents a probability that x is a real gene-brain region directed graph; G(z) represents a generated gene-brain region directed graph of a normal person or an AD patient; P z (z) represents a distribution of data z, and D(G(z)) represents a probability that G(z) is a real gene-brain region directed graph.

[0063] Preferably, the brain region feature extraction system further comprises:

[0064] a risk prediction module, configured to compare the key brain region-gene directed graph output by the trained generator with a real key brain region-gene directed graph of a diseased brain region, calculate a similarity of the two, and determine a risk of disease according to the similarity, wherein a calculation formula is as follows:

[0065]

[0066]

[0067] wherein S rc represents a value of an rth row and a cth column of a similarity matrix, represents a value of an rth row and a cth column of a weight matrix of the generated gene-brain region directed graph; represents a value of an rth row and a cth column of a weight matrix of the real key brain region-gene directed graph;

[0068] and / or

[0069] a diseased gene / brain region extraction module, configured to calculate a topological potential energy change of each node before and after each stage of evolution, and obtain a potential energy increment matrix corresponding to n nodes Summing the n node potential energy increment matrices can obtain a total node potential energy increment matrix ΔT k; according to the total increment matrix of node potential energy of all subjects ΔT1, ΔT2,..., ΔT m , calculate the average change of node potential energy; based on the average change of node potential energy, the final pathogenic gene / brain region is screened out by the incremental search method.

[0070] The present application has the following beneficial effects:

[0071] 1、The brain region feature extraction system based on model driven evolution in the application, through obtaining brain image data and gene data of the target object, an initial brain region-gene directed graph is constructed; a generator based on the model driven evolution model is constructed, the model of the generator is optimized according to the similarity between the generated key brain region-gene directed graph and the real key brain region-gene directed graph, and the model parameters of the generator are optimized according to the similarity, compared with the prior art, the present application can accurately extract the key features of the preset brain lesion, and provide support for accurately predicting brain lesions.

[0072] 2、In the preferred scheme, the present application can predict the risk of future development of AD in LMCI patients, thereby providing important auxiliary decision support for AD treatment strategies.

[0073] 3、In the preferred scheme, the present application can reveal the evolution mode in the conversion process of LMCI to AD, and successfully extract the key lesion brain regions and risk genes related to disease progression.

[0074] In addition to the purposes, features and advantages described above, the present application has other purposes, features and advantages. The present application will be further described in detail below with reference to the accompanying drawings. BRIEF DESCRIPTION OF DRAWINGS

[0075] The accompanying drawings, which form a part of the present application, are intended to provide further understanding of the present application, and the illustrative embodiments of the present application and their description serve the purpose of explaining the present application. The present application is not limited by the improper limitation of the accompanying drawings. In the drawings:

[0076] Figure 1 The model driven directed graph evolution process provided by the present application;

[0077] Figure 2 The workflow diagram of the brain region feature extraction system based on model driven evolution provided by the present application;

[0078] Figure 3 The overall framework diagram of the generator provided by the present application;

[0079] Figure 4A discovery and evolution process diagram of the motif provided by the present application is shown in FIG. 1.

[0080] Figure 5 An operation diagram of motif evolution provided by the present application is shown in FIG. 2.

[0081] Figure 6 A schematic diagram of motif evolution provided by the present application is shown in FIG. 3.

[0082] Figure 7 A subgraph partitioning schematic diagram provided by the present application is shown in FIG. 4.

[0083] Figure 8 A directed subgraph evolution process diagram based on motif convolution provided by the present application is shown in FIG. 5.

[0084] Figure 9 An operation diagram of motif convolution generating an evolved subgraph is shown in FIG. 6.

[0085] Figure 10 A directed subgraph evolution schematic diagram based on motif convolution provided by the present application is shown in FIG. 7.

[0086] Figure 11 An internal structure diagram of the discriminator provided by the present application is shown in FIG. 8.

[0087] Figure 12 An internal structure diagram of the full connection layer provided by the present application is shown in FIG. 9.

[0088] Figure 13 A feature extraction process example diagram provided by the present application is shown in FIG. 10.

[0089] Figure 14 A disease risk prediction schematic diagram provided by the present application is shown in FIG. 11. DETAILED DESCRIPTION

[0090] The embodiments of the present application are described in detail below with reference to the accompanying drawings, but the present application can be implemented in various different ways as limited and covered by the claims.

[0091] In order to accurately capture the evolution pattern of LMCI to AD and find the key pathogenic factors, the risk prediction of the disease can be realized. This section constructs a motif-driven directed graph evolution model. The design idea of the model is that in the evolution process of neurodegenerative diseases such as AD, some genes and corresponding brain regions that are highly affected by pathology play a core role in transmitting pathological signals. They show pathological activity first and preferentially transmit pathological information to adjacent genes or brain region nodes that are highly coupled with them. As these initial pathological modules continue to worsen, their pathological effects spread in the network, causing abnormal gene expression and brain region function. Ultimately, it leads to clinical manifestations such as cognitive decline and memory impairment. In order to simulate the cascade reaction of disease deterioration from partial areas, this chapter uses motif structure to describe the more important node combination in the gene-brain region directed graph, and designs an evolution mode suitable for the motif in the directed graph. Based on the influence range of the evolved motif, the directed graph is divided into multiple subgraphs. According to the correlation degree of the nodes in the subgraph and the motif, a motif tree is constructed, and the motif evolution information is propagated based on the motif tree structure. Based on the characteristic information of the change of the nodes in the motif tree, the edges between the nodes are reconstructed to form an evolution subgraph. Finally, all evolution subgraphs are integrated to obtain the gene-brain region directed graph of the next evolution stage, and the final generated gene-brain region directed graph is obtained.

[0092] As shown in Figure 1 , a motif-driven directed graph evolution can be mainly divided into motif evolution part (a) and motif-driven directed subgraph evolution part (b). First, all motifs in the gene-brain region directed graph are found, and motifs with low importance are screened based on the directed topological potential of the motifs to avoid the same node appearing in multiple motifs. Second, the node information in the motif is changed and the node connection is updated by performing node information transmission for the nodes in the motif, and the motif evolution is completed. After the motif evolution is completed, the gene-brain region directed graph is divided into multiple directed subgraphs based on the influence range of each motif. Third, a motif tree is constructed based on the correlation degree of the motif and each node, and the motif evolution information is diffused to the remaining leaf nodes in the tree from the root node. Finally, the directed subgraph topology structure is updated to generate an evolution subgraph according to the node evolution information.

[0093] The brain region feature extraction system based on motif-driven evolution provided in this embodiment establishes a motif-driven evolution model of the directed graph based on the gene-brain region directed graph. Second, a directed graph generative adversarial network is designed based on the model. In the generator, first, the motif structure in the initial gene-brain region directed graph is detected and evolved. Second, the subgraphs are divided according to the influence range of the motifs. Then, the motif tree is extracted in each subgraph through motif convolution, and the corresponding evolution subgraph is formed. Finally, these evolution subgraphs are integrated to form the evolved gene-brain region directed graph. The discriminator is responsible for evaluating the similarity between the generated graph and the real graph. The generator and the discriminator continuously optimize the parameters to reduce the difference between them.

[0094] As Figure 2 shown, the brain region feature extraction system based on motif-driven evolution provided by the embodiment specifically comprises:

[0095] a preprocessing module, configured to acquire brain image data and gene data of a target object, and construct an initial brain region-gene directed graph;

[0096] a generator, constructed based on a motif-driven evolution model, configured to detect a motif of a preset brain lesion from the initial brain region-gene directed graph, and generate a key brain region-gene directed graph for detecting the preset brain lesion according to a result of evolution of the motif;

[0097] a discriminator, configured to calculate a similarity between the generated key brain region-gene directed graph and a real key brain region-gene directed graph, and optimize model parameters of the generator according to the similarity;

[0098] a risk prediction module, configured to compare the key brain region-gene directed graph output by the trained generator with a real key brain region-gene directed graph of a lesion, calculate a similarity between the two, and determine a risk of disease according to the similarity;

[0099] a lesion gene / brain region extraction module, configured to calculate a topological potential energy change of each node before and after each stage of evolution, and obtain a potential energy increment matrix corresponding to n nodes Summing the n node potential energy increment matrices can obtain a total node potential energy increment matrix ΔT k ; according to the total node potential energy increment matrices ΔT1, ΔT2,..., ΔT m of all subjects, an average change amount of node potential energy is calculated; based on the average change amount of node potential energy, an incremental search method is used to screen out the final lesion gene / brain region.

[0100] 1. Generator

[0101] As Figure 3 shown, the generator comprises:

[0102] an input layer, which carries and expresses the multi-dimensional and internal relations of the data to be processed;

[0103] a motif discovery and evolution module, configured to detect motif structures in the initial gene-brain region directed graph and perform evolution;

[0104] a subgraph division module, configured to divide subgraphs according to the influence range of the motifs;

[0105] a directed subgraph evolution module, configured to extract motif trees in each subgraph through motif convolution, and form corresponding evolution subgraphs;

[0106] An evolutionary subgraph fusion module is configured to integrate the evolutionary subgraphs to form an evolved gene-brain region directed graph.

[0107] 1.1 Input layer

[0108] Specifically, the input layer is composed of a series of matrices, which serve as the basic input units of the network and carry and express the multiple dimensions and internal relations of the data to be processed. The input matrix is as follows:

[0109] (1) Node feature information matrix H. N is the total number of nodes, and l is the dimension of the feature information of each node. H ij represents the j-dimensional feature information of the i-th node, and the feature matrix H can be represented as follows:

[0110]

[0111] (2) Inter-node edge weight matrix W. W ∈ R N×N , W ij represents the weight of the edge between node i and node j, and the weight matrix W can be represented as follows:

[0112]

[0113] (3) Inter-node adjacency matrix A. A ∈ {0, 1}, A ij represents the connection between node i and node j, and its value can be 0 or 1, 0 representing no edge and 1 representing an edge, and the inter-node adjacency matrix A can be represented as follows:

[0114]

[0115] (4) Motif adjacency matrix M A . M A ∈ {0, 1}. represents the connection between node i and node j, MA ij = 1 represents the existence of an edge from node i to node j, represents that node i does not point to node j, and the motif adjacency matrix M A can be represented as follows:

[0116]

[0117] (5) Motif internal node feature matrix MH. MH ∈ R N×l , N is the total number of nodes in the motif, and l is the dimension of the feature information of each node. MH ij represents the j-dimensional feature information of the i-th node, and the motif internal node feature matrix MH can be represented as follows:

[0118]

[0119] (6) The edge weight matrix W between nodes in the motif M . W M ∈R N×N , denotes the weight of the edge between node i and node j in the motif, the inter-node weight matrix W M can be represented as follows:

[0120]

[0121] (7) Motif-node association matrix S. S ∈ {0, 1}, S ij denotes the association between node i and node (motif) j, the i-th row and the j-th column are 0, indicating that the association degree is 0, and are not 0, indicating that the association degree is S ij , then the inter-node motif-node association matrix S can be represented as follows:

[0122]

[0123] 1.2 Motif discovery and evolution module:

[0124] Specifically, the motif discovery and evolution module comprises:

[0125] a motif construction unit, configured to detect motifs of the initial brain region-gene directed graph, and construct a motif adjacency matrix and a motif record matrix;

[0126] a motif screening unit, configured to screen a target motif with the largest directed topological potential from the motif adjacency matrix, and update the motif record matrix with the target motif;

[0127] a first updating unit, configured to update the inter-node feature information matrix based on information transmission of nodes in a motif;

[0128] a second updating unit, configured to update the motif adjacency matrix based on the feature information of the updated nodes;

[0129] a third updating unit, configured to calculate weight information of newly generated edges in the motif, and update the weight matrix.

[0130] The motif construction unit comprises:

[0131] a first search subunit, configured to find all subgraphs of a given size in the initial brain region-gene directed graph by enumeration, record node information of the initial motif adjacency matrix, and continuously enumerate until the search is completed;

[0132] a random generation subunit, configured to classify the subgraphs into corresponding isomorphic classes, and randomly generate an initial motif adjacency matrix corresponding to the initial brain region-gene directed graphFigure One A randomized plot of sample size;

[0133] The adjacency matrix generates sub-units. The occurrence frequency of multivariate subgraphs in the random graph and the initial brain region-gene directed graph is compared. Multivariate subgraphs with a higher occurrence frequency than the random network are marked and recorded as motifs. At the same time, multivariate subgraphs that do not belong to the motifs are deleted from the adjacency matrix to obtain the final motif adjacency matrix.

[0134] The record matrix generates sub-units. Based on the found motifs, the nodes of all motifs are counted, and a motif record matrix is ​​constructed.

[0135] The phantom screening unit includes:

[0136] Computational subunit: Obtain the node features of each module from the module record matrix, calculate the out-degree module potential and in-degree module potential of each module, and calculate the directed topological potential of each module based on the out-degree module potential and in-degree module potential of each module;

[0137] Comparison sub-units: Compare the directed topological potentials of each motif, select the target motif with the largest directed topological potential, and update the motif record matrix;

[0138] The first update unit includes:

[0139] Update subunit A: is used to calculate the correlation coefficient between each node and the feature information of other nodes based on the updated node feature information. When the correlation coefficient between any two nodes is greater than a preset value, it is determined that the two nodes are connected to each other.

[0140] Update subunit B: Used to determine the direction of node connections based on the updated node's feature information and potential theory, for any two nodes V i and V j If V i The potential energy is greater than V j If it is high, then determine if V exists. i Point to V j The connecting edges.

[0141] Based on the above structure, such as Figure 4 As shown, the phantom discovery and evolution module mainly performs the following steps:

[0142] The first step is to detect and label all motifs in the gene-brain region directed network to obtain the motif adjacency matrix M. A (k-1) With phantom record matrix

[0143] First, enumerate all subgraphs of a given size in the target network and use the motif adjacency matrix M. ARecord its node information and continuously enumerate until the search is complete. Then, classify these subgraphs into their corresponding isomorphic classes. Next, randomly generate genes-brain region directed graphs. Figure One The random graph is then used to calculate the frequency of triplets. Finally, the occurrence frequencies of triplets in the random graph and the gene-brain region directed graph are compared. Triplets with higher frequencies than those in the random network are marked and recorded as motifs. Triplets that do not belong to motifs are removed from the adjacency matrix, resulting in the final motif adjacency matrix M. A After identifying all motifs in the network, the nodes constituting the motifs are counted, and the motifs detected by the algorithm are recorded in an M-row, 4-column motif record matrix M. F In this diagram, each row represents a module, the first three columns are the nodes that constitute that module, and the last column is the directed topological potential of the module. For example, if nodes i, j, and k constitute module a, then its M... F [a] = (i, j, k.TP) a The calculation formula is as follows:

[0144]

[0145] in, A is the adjacency matrix of the motif detected in the directed gene-brain region graph of the i-th subject, A is the adjacency matrix of the directed gene-brain region graph D, and K is the number of nodes of the detected motif. In this paper, we mainly study ternary motifs, so K is taken as 3.

[0146] The second step is motif selection. For motif structures with shared nodes, only the motif with the highest directed topological potential is retained, and the motif adjacency matrix M is updated. A (k,0) With phantom record matrix

[0147] First, in the phantom recording matrix M obtained in the previous step F In this process, if a node participates in the construction of multiple motifs, the motif with the highest directed topological potential is selected for evolution. This reduces redundant evolution and improves the accuracy and stability of the evolution. The directed topological potential formula for a node is divided into in-degree and out-degree topological potentials. The in-degree topological potential is calculated from the information of all nodes that can affect the node, while the out-degree topological potential is calculated from the information of all nodes that can affect the node. The formula for calculating the in-degree topological potential is shown below. The out-degree topological potential TP of a node can be calculated in the same way. out (i).

[0148]

[0149] Among them, MH j For node v j The characteristic information, σ, is used to control the influence range of a node and can be calculated based on the node's topological potential entropy E. wherein is a standardization factor. The directed topological potential of the node is brought in, and then it is formulated as a function of the directed topological potential and the topological potential entropy. The minimum point of the function is solved, and the horizontal coordinate is the optimal influence factor.

[0150] The out-degree motif potential and the in-degree motif potential are added to obtain the directed topological potential of the node. The calculation formula is as shown below

[0151] TP(i) = TP out (i) + TP in (i)

[0152] Secondly, since the constituent nodes of each motif have been obtained in the motif record matrix M F , the directed topological potential of each motif can be obtained through the directed topological potential of the node, and recorded in the last column of the motif record matrix M F . The calculation formula is as shown below:

[0153] M TP = TP(V i ) + TP(V j ) + TP(V k )

[0154] wherein M TP is the directed topological potential of the motif, and TP(V i ), TP(V j ), and TP(V k ) represent the directed topological potentials of the nodes constituting the motif, respectively.

[0155] Finally, in the motif record matrix M F , if the same node number appears repeatedly in multiple rows, it means that the node participates in the construction of multiple motifs. The directed topological potentials of the motifs participated by the node are compared, and only the motif with the largest directed topological potential is retained to complete the motif screening. The calculation formula is as shown below:

[0156] M F = Renew(Max(M F [i][4], M F [j][4],..., M F [k][4]));

[0157] if n = M F [i][], M F [j][],..., M F [k]]

[0158] wherein M F represents the updated motif record matrix, M F [][4] is the directed topological potential of the multiple motifs participated by the node in the construction, and MF [i][], [j][], ..., [k][] are the same node numbers in different rows of the matrix.

[0159] The third step involves the nodes within the schema transmitting information and updating the feature information matrix MH of the nodes within the schema. (k) .

[0160] Within the schema, nodes transmit their own information to connected nodes via out-degree edges, thereby updating the information of nodes within the schema. After the t-th inter-layer feature information transmission, the formula for calculating node feature information is as follows:

[0161]

[0162] Among them MH (k) and These represent the feature information of the i-th node within the phantom after the k-th and (k-1)-th feature information transmissions, respectively. This represents the feature information of node j pointing to node i within the modal after the (k-1)th information transmission. This represents the weight of the directed edge from node j to node i within the module.

[0163] The fourth step is to calculate the node connection probability based on the updated node feature information and update the modal adjacency matrix M. A (k ,1) .

[0164] Changes in node feature information affect the probability of edges existing between nodes. The probability of directed edges obtained through the Pearson correlation coefficient can be used to quantify the likelihood of a connection between two nodes. If the Pearson correlation coefficient exceeds a threshold α, then an edge exists. This is reflected in the modulus adjacency matrix M. A The adjacency relationship between two nodes is set to 1. The calculation formula is as follows:

[0165]

[0166] Among them, M A [i][j], [j][i] represent two nodes V in the modality adjacency matrix. i and V j The connection relationship between them, Pearson(V) in the formula i V j ) represents the node V calculated using the Pearson correlation coefficient. i V j The probability of connecting edges between them.

[0167] The fifth step involves calculating the directed topological potential based on the updated node feature information, determining the node connection directions according to potential theory, and updating the module adjacency matrix M. A(k,2) .

[0168] Based on the updated node feature information MH (k) The node directed topology potential is calculated. Secondly, the direction of the newly generated edge between nodes in the motif is determined according to the potential theory. The potential theory means that for any two nodes V i and V j in the network, if the potential energy of V i is higher than that of V j , there may be an edge from V i to V j . The calculation formula is as follows:

[0169]

[0170] Where M A [i][j] is the edge relationship between two nodes V i and V j in the motif adjacency matrix.

[0171] Sixth, the weight information of the newly generated edge in the motif is calculated, and the weight matrix W

[0172] The edge weight between nodes in the motif is recalculated based on the updated connection, and the weight matrix in the motif is updated The calculation formula is as follows:

[0173]

[0174] Where, represents the feature information of the i-th node in the motif after the k-th feature information transmission, represents the feature information of the j-th node pointing to the i-th node in the motif after the k-th information transmission, represents the directed edge weight from the j-th node to the i-th node in the motif, represents the directed edge weight from the j-th node to the i-th node in the initial gene-brain region directed graph.

[0175] In summary, the steps of motif detection and evolution are represented by the following formula:

[0176]

[0177] Next, the operation process of motif evolution is explained by taking a subgraph of a gene-brain region directed graph with 6 nodes as an example, as shown in Figure 5 In the gene-brain region directed graph shown in the figure, nodes 1, 2, 3 and nodes 3, 4, 5 form two motifs in the subgraph respectively.

[0178] Firstly, the motif record matrix is updated based on the motif adjacency matrix. There are two motifs in the graph, so the motif record matrix M F There are two rows. Since node 3 participates in the construction of two motifs, it is necessary to compare the sizes of the directed topological potentials of motif 1 and motif 2, and only keep the motif structure with a higher directed topological potential, which also represents a higher influence of the motif in the network. The motif with a lower directed topological potential may carry more redundant information and interfere with the evolution accuracy of the network. Then the motif adjacency matrix and the motif record matrix are updated, and the information about motif 2 is deleted. Secondly, the node information is transmitted within motif 1, and node 3 and node 1 propagate information along the directed edge, update the node feature matrix, and finally determine the node edge probability within the motif according to the updated information of the nodes within the motif. The adjacency relationship between the nodes with a probability of generating a connection is set to 1, such as nodes 1 and 2, and 3 in the figure have a probability of generating a connection, so the 1st row and 2nd and 3rd columns in the motif adjacency matrix are set to 1. The directed topological potential of the node is calculated based on the updated feature information of the node within the motif, and the direction of the node edge is determined according to the potential theory. At this time, the motif adjacency matrix and the motif record matrix are updated. The motif evolution is completed.

[0179] Significance of motif evolution: The evolution of motifs reveals the dynamic changes of structural units within the motif over time or pathological process. By observing the evolution of the connection type, frequency and distribution of gene-brain region nodes, not only can the specific relationship between gene nodes and brain region nodes be revealed, but also the evolution pattern of gene-brain region directed network and how these structural changes interact with the functional activities of biological systems can be understood. Furthermore, by analyzing these changes, functional abnormalities and regulatory mechanisms related to the development of specific diseases can be revealed. By analyzing the role changes, interaction relationships and regularities of gene nodes and brain region nodes in motif connection over time, the dynamic evolution pattern of gene-brain region directed network can be systematically grasped, thereby providing deep theoretical basis and practical guidance for analyzing the complex network biology mechanism of disease occurrence.

[0180] 1.2 Subgraph division module

[0181] As shown in Figures 6-7 , the subgraph division module divides the directed subgraph according to the influence range of each motif after evolution. First, the influence range of the motif is calculated. The influence factor σ obtained when calculating the topological potential is the influence range of the node. According to the principle of Gaussian potential function "3σ", the influence range l of each node is a field with the node as the center and the radius of . When , the influence range of the node is l-hop neighbor nodes. When calculating the influence range of the motif, the smallest influence range of the three-node group within the motif is taken as the influence range σ of the motif MWhen multiple motifs contain duplicate nodes within their influence range, the duplicate nodes are assigned to the motif with the largest directed topological potential, as motifs with larger directed topological potentials have greater influence. Finally, the subgraph is partitioned, and the corresponding feature matrix H is obtained. SG and weight matrix W SG The calculation formula is as follows:

[0182]

[0183] in, H represents the feature matrix and weight matrix of the i-th directed subgraph in the k-th evolution stage, respectively. (k,m) W (k,m) Let σ represent the node feature matrix and weight matrix after the motif evolution in the k-th evolutionary stage, respectively. M This refers to the influence range of the model.

[0184] 1.3 Directed Subgraph Evolution Module

[0185] like Figure 8 As shown, the directed subgraph evolution module includes:

[0186] The correlation calculation unit is used to calculate and sort the correlation between gene nodes and motifs within the directed subgraph, obtain a first correlation sequence, and store the first correlation sequence into the motif-node correlation matrix S. (k,0) The association degree between brain region nodes and phantoms within the directed subgraph is calculated and sorted to obtain the first association degree sequence, which is then stored in the phantom-node association matrix S. (k,1) ;

[0187] The phantom tree construction unit is used to construct the phantom-node association matrix S. (k,1) Construct a motif tree, connect gene nodes within the directed subgraph, and update the motif tree weight matrix. Connect the nodes of the brain regions within the directed subgraph and update the adjacency matrix of the motif tree.

[0188] Motif tree evolution unit: used to transmit evolutionary information based on the motif tree connection structure and update the feature matrix H of the motif tree nodes. (k) Based on the node update feature information, the edges are recalculated, and the weight matrix of the evolutionary subgraph is updated.

[0189] The correlation calculation unit includes:

[0190] The average feature calculation subunit is used to select the feature information MH of the nodes contained in the phantom based on the k-th order. (k) Calculate the feature information of the k-th order modulus. The i-th module Features averaged from the features of all nodes in the motif; the nodes are gene nodes or brain region nodes;

[0191] The motif-node correlation degree matrix construction subunit is configured to calculate the correlation coefficient of each node feature and the feature information , sort the correlation coefficient corresponding to each node feature, obtain a correlation degree sequence, and store the result in a motif-node correlation degree matrix.

[0192] The motif tree construction unit comprises:

[0193] The first index construction subunit is configured to take the gene node with the highest correlation degree in the first correlation degree sequence as a leaf to be included in the motif tree, and take the correlation coefficient of the included node as its weight; iteration: calculate the correlation coefficient between each node not included and the nodes in the motif tree already included, obtain an auxiliary correlation degree sequence, aggregate the auxiliary correlation degree sequence and the correlation degree sequence, and select a gene node with the highest correlation degree as a leaf to be included in the motif tree, and take the correlation coefficient of the included node as its weight; repeat the iteration step until all gene nodes are included in the motif tree, and then update the weight matrix with the weights of all gene nodes;

[0194] The second index construction subunit is configured to take the brain region node with the highest correlation degree in the second correlation degree sequence as a leaf to be included in the motif tree, and take the correlation coefficient of the included node as its weight; iteration: calculate the correlation coefficient between each node not included and the nodes in the motif tree already included, obtain an auxiliary correlation degree sequence, aggregate the auxiliary correlation degree sequence and the correlation degree sequence, and select a gene node with the highest correlation degree as a leaf to be included in the motif tree, and take the correlation coefficient of the included node as its weight; repeat the iteration step until all gene nodes are included in the motif tree, and then update the weight matrix with the weights of all gene nodes;

[0195] The motif tree evolution unit comprises:

[0196] The motif tree evolution update subunit is configured to update the node feature information of the nodes on the motif tree from top to bottom, and further update the motif tree node feature matrix H (k) The weight evolution update subunit is configured to re-calculate the edges and weights between nodes based on the updated motif tree node feature matrix H (k) , form an evolution subgraph, and update the subgraph weight matrix based on the weights in the evolution subgraph Based on the above structure, as shown in Figure 8 the directed subgraph evolution module performs the following steps:

[0197] The first step is to calculate and sort the association degree between gene nodes and motifs within the directed subgraph, and then update the motif-node association matrix S. (k,0) The second step is to calculate and sort the association degrees between brain region nodes and phantoms within the directed subgraph, and then update the phantom-node association matrix S. (k ,1) The third step is to determine the relationship between the phantom and the node based on the phantom-node association matrix S. (k,1) Construct a motif tree, connect gene nodes, and update the motif tree weight matrix. The fourth step is to connect the brain region nodes and update the motif tree weight matrix. Fifth, based on the phantom tree connection structure, transmit phantom evolution information and update the phantom tree node feature matrix H. (k) Step 6: Recalculate the edges between nodes based on the node update feature information, and update the weight matrix of the evolutionary subgraph. The specific algorithm steps are as follows.

[0198] The first step is to calculate and sort the association degree between gene nodes and motifs within the directed subgraph, and then update the motif-node association matrix S. (k,0) ;

[0199] First, based on the k-order selection, the feature information MH of the nodes contained in the phantom is obtained. (k) Calculate the feature information of the k-th order modulus. The i-th module Features The formula is as follows: It is obtained by averaging the features of all nodes in the phantom.

[0200]

[0201] in This represents the motif feature information in the k-th evolutionary stage; motif The number of nodes contained, i.e., the size of the phantom; motif The feature information of the j-th node in the dataset.

[0202] Secondly, a motif-node association matrix is ​​constructed. The i-th row and j-th column of the matrix represents the association degree between the i-th node and the j-th node (motif) (if i or j is 0, it represents a motif in the subgraph, i.e., the similarity sequence between the motif and the remaining nodes is the first row of the matrix). First, the association degree between gene nodes and motifs M is calculated. i The correlation. The results are then stored in the phantom-node correlation matrix S. (k) The calculation formula is as follows:

[0203]

[0204] in, represents the sequence of the correlation degree between the gene node and the motif in the tthsubgraph, represents the feature information of the ithgene node in the subgraph except the motif in the kthevolution process, represents the feature information of the motif after the kthmotif evolution in the subgraph.

[0205] Secondly, the correlation degree between the brain region node and the motif in the directed subgraph is calculated and sorted, and the motif-node correlation matrix S is updated (k,1) ;

[0206] The correlation degree between the brain region node and the motif in the subgraph is calculated in the same way and the result is stored in the motif-node correlation matrix S (k) . The calculation formula is as follows:

[0207]

[0208] wherein, represents the sequence of the correlation degree between the brain region node and the motif in the tthsubgraph, represents the feature information of the ithbrain region node in the subgraph except the motif in the kthevolution process, represents the feature information of the motif after the kthmotif evolution in the subgraph.

[0209] Thirdly, the motif tree is constructed according to the motif-node correlation matrix S (k,1) , the gene nodes in the directed subgraph are connected, and the weight matrix of the motif tree is updated

[0210] Firstly, since the brain region node is often regulated by the gene node, and a single gene node will interact with multiple brain regions or genes, resulting in the importance of the gene node being often higher than that of the brain region node, therefore, when constructing the motif tree, the connection between the gene node and the driving motif in the subgraph is considered first, and then the brain region node is considered. The similarity of the gene node and the motif is sorted, the node with the highest similarity will be directly connected as the leaf node of the motif and connected to the tree structure, and the similarity value calculated by the Pearson coefficient will be used as the weight of the edge. The calculation formula is as follows:

[0211]

[0212] Secondly, the similarity between the node and the motif and the remaining nodes connected to the tree is calculated according to the correlation degree between the gene node and the motif, when the remaining nodes are connected to the tree, the similarity between the node and all the nodes in the tree including the motif needs to be calculated, and the most similar node is connected as the leaf node of the tree, the similarity is calculated by the Pearson coefficient, and the result is stored in the motif-node correlation matrix S as the weight of the edge. The calculation formula is as follows:

[0213] S[i][j] = Pearson((M k , V j ), V i )

[0214] where V i represents the gene node being judged this time, V j represents the node that has been connected to the motif tree, M k represents the driving motif in the motif tree, (M k , V j ) represents the possible motif M k or node V j participating in the judgment.

[0215] Finally, the weight matrix is updated based on the updated information. The calculation formula is as follows:

[0216] W T [j][i] = S[i][j] if S[i][j] = max(S[i])

[0217] where max(S[i]) represents the highest value of the correlation degree in the ith column, S[i][j] represents the highest value of the correlation degree in the ith column, W T [j][i] = S[i][j] represents filling the correlation degree value as the weight in the weight matrix, which is abstracted as connecting the ith node as a leaf node to the jth node (motif) in the motif tree.

[0218] Step 4: Connect the nodes of the directed subgraph of the brain region, and update the motif tree adjacency matrix

[0219] After all the gene nodes are added to the motif tree, the same steps are used to calculate the brain region nodes. The calculation formula is as follows:

[0220] S[i][j] = Pearson((M k , V j ), V i )

[0221] where V i represents the brain region node being judged this time, V j represents the node that has been connected to the motif tree, M k represents the driving motif in the motif tree, (M k , V j ) represents the possible motif M k or node V j participating in the judgment. The calculation formula is as follows:

[0222] W T [j][i] = S[i][j] if S[i][j] = max(S[i])

[0223] where max(S[i]) means the highest correlation value in the ith column, S[i][j] means the highest correlation value in the ith column is located in the jth column, W T [j][i] = S[i][j] means filling the correlation value as weight in the weight matrix, which is abstracted as connecting the ith node as a leaf node to the jth node in the motif tree.

[0224] The fifth step is to transfer the evolution information based on the motif tree connection structure and update the motif tree node feature matrix H (k)

[0225] According to the constructed motif tree, the node information is transferred. The changed motif transfers its change information to the lower layer nodes, and the lower layer nodes complete the information transfer in turn until all the node information in the tree is changed. The calculation formula is as follows:

[0226]

[0227] where H (k-1) represents the feature information of the nodes in the k-1 order motif tree, W T (k-1) [i] represents the weight information of the ith node in the k-1 order, H (k-1) [i] represents the feature information of the ith node in the k-1 order, represents the update information AH of the ith node transmitted along the out-degree edge, and θ represents the learnable parameter matrix.

[0228] The sixth step is to recompute the edges based on the updated node feature information and update the evolution subgraph weight matrix

[0229] According to the updated node feature information, the edges and weights between nodes are recalculated to form the evolution subgraph. The calculation formula is as follows:

[0230]

[0231] where, represents the feature information of the ith node after the kth information propagation, the feature information of the jth node after the tth information propagation, represents the weight of the edge between the ith node and the jth node after the kth information propagation.

[0232] As Figure 9 shown, the following explains the generation process of the evolution subgraph with a 7-node subgraph as an example:

[0233] Firstly, the correlation degree of gene nodes and motifs is calculated based on the motif-node association matrix and stored in S[0]. The gene node with the highest correlation degree in S[0] is found, which is node 3. The correlation degree of node 3 with other nodes is calculated and stored in S[3][0] together with the correlation degree of the motif. Node 3 is added as a leaf node in the tree with the Pearson value as the edge weight. The correlation degree of node 6 is second only to node 3, and node 6 is processed in the same way. Secondly, the brain region nodes are connected in the same way. Then, according to the motif tree structure, the motifs propagate their change information to update the information of the remaining nodes in the tree. The upper nodes in the tree transmit their change information to update the lower nodes in the tree. The node connection is rejudged according to the updated node information, the evolution subgraph is constructed, and the generation of the evolution subgraph is completed.

[0234] The significance of motif convolution: First, it reveals the correlation between motifs and nodes through the generated motif tree, clearly showing the correlation between the two, which helps to simplify the organization of the gene-brain region directed network. Second, the connection mode and position of the key nodes in the motif tree reflect their importance in the biological network. It helps to determine the gene-brain region nodes that have a significant impact on the system function and stability. Then, the construction of the motif tree layers the nodes according to the correlation degree, forming a hierarchical structure. It helps to understand the organization of the gene-brain region directed network and the connection and regulation between different levels. Finally, the motif tree can also be used to study the evolution of the gene-brain region directed network, the functional changes of brain regions, and the pathogenesis of AD.

[0235] 1.4 Evolution subgraph fusion module

[0236] The evolution subgraph fusion module fuses the evolution subgraphs corresponding to each motif in the k-1 order gene-brain region directed graph to form a k order gene-brain region directed graph, and the calculation formula is as follows:

[0237]

[0238] wherein, represents all evolution subgraphs formed by the ith subject in the kth evolution, Fuse() represents the graph fusion function; k is the number of evolutions in the generator; N is the total number of evolutions in the generator; M is the total number of evolution subgraphs; D (i) is the directed graph after evolution of the ith subject; D is the set of directed graphs after evolution of all subjects.

[0239] 2. Discriminator

[0240] As shown in Figure 11 , the discriminator is a network structure composed of a directed graph convolution layer and a fully connected layer, and its purpose is to distinguish whether the input brain region-gene directed network is real or generated by the generation model.

[0241] The following is a specific introduction to each component of the discriminator.

[0242] The input of the directed graph convolution layer is the real or generator reconstructed gene-brain region directed network, and the calculation formula of the n-order directed graph convolution is as follows:

[0243]

[0244] The internal result of the fully connected layer is as shown in the following formula (3) : Figure 12 The gene-brain region directed graph input to the discriminator is obtained after multiple convolutions, and an updated feature matrix H (2) is obtained, and then an updated matrix F (2) is calculated according to the updated feature matrix H (2) , and is flattened into one-dimensional data to obtain Y (0) as the input of the fully connected layer, and outputs a vector of size 2*1, indicating the true or false discrimination score of the gene-brain region directed graph, and the calculation formula of the fully connected layer is as follows:

[0245]

[0246] where W (l) represents the weight matrix between the lth and l+1th fully connected layers; Y (l) represents one-dimensional data of the lth fully connected layer; b (l) represents the bias between the lth and l+1th fully connected layers; ReLU(·) and Softmax(·) represent activation functions; Y (l+1) represents one-dimensional data of the l+1th fully connected layer.

[0247] In this embodiment, the loss function calculation formula of the brain region feature extraction system is as follows:

[0248]

[0249] where G and D represent the generator and the discriminator, respectively; E(·) represents the expected value of a variable; x represents the gene-brain region directed graph of a normal person or an AD patient (real gene-brain region directed graph); P data (x) represents the distribution of data x, and D(x) represents the probability that x is a real gene-brain region directed graph; G(z) represents a generated gene-brain region directed graph of a normal person or an AD patient, and P z (z) represents the distribution of data z, and D(G(z)) represents the probability that G(z) is a real gene-brain region directed graph.

[0250] The formulaic expression of the training steps of the discriminator and the generator is as follows:

[0251]

[0252] where max D V(D, G) is to maximize the discriminant accuracy of the discriminator, min G V(D, G) is to minimize the probability of the pseudo sample being discriminated as true.

[0253] The formula represents the maximum and minimum optimization problem of the generative adversarial network, that is, to maximize the discriminant accuracy of the discriminator and to minimize the difference between the directed graph generated by the generator and the real directed graph. The first formula is to train the discriminator, which means that the probability of discriminating the real sample as true should be as large as possible, and the probability of discriminating the pseudo sample as false should be as large as possible. The second formula is to train the generator, which means that the probability of discriminating the pseudo sample as true should be as small as possible, and the difference between the generated data and the real data should be as small as possible.

[0254] The two steps are repeatedly performed, so that the generator and the discriminator are continuously optimized in the mutual game and converge to the highest accuracy.

[0255] 3. Lesion gene / brain region extraction module

[0256] The lesion gene / brain region extraction module comprises:

[0257] The LMCI patient gene-brain region directed graph is input into the feature extraction system for n times of iterative evolution, and the generated directed graph is output. Then, the generated directed graph and the real AD gene-brain region directed graph are jointly input into the discriminator. In the iterative process of the model, the dynamic competition between the discriminator and the generator is continuously performed until the discriminator cannot determine true or false. At this time, the network converges, and the model successfully learns and captures the evolution pattern of the disease.

[0258] After the model training is completed, by exploring the change of the node directed topological potential in the disease deterioration process, the nodes with significant changes in potential energy information and practical prevention and treatment significance in the disease deterioration process can be found, so as to extract abnormal genes and lesion brain regions. The process example diagram is shown in FIG. 1. Figure 13 The specific steps are as follows:

[0259] In the first step, by calculating the topological potential change of each node before and after each stage evolution, the potential increment matrix corresponding to n nodes can be obtained. The calculation formula of the node potential increment matrix corresponding to the tth evolution stage is as follows:

[0260]

[0261] wherein and respectively represent the node potential matrix of the t+1 order gene-brain region directed graph and the t order gene-brain region directed graph of the kth subject. represents the node potential increment matrix of the kth subject in the tth evolution stage.

[0262] Secondly, summing up the n node potential increment matrices, the total node potential increment matrix AT can be obtained. k The calculation formula is as follows:

[0263]

[0264] wherein, represents the node potential increment matrix corresponding to the tth evolution stage; AT k represents the total node potential increment matrix of the kth subject.

[0265] Thirdly, according to the total node potential increment matrices AT1, AT2,..., AT m of all m LMCI subjects, the average change of the node potential is calculated, and the calculation formula is as follows:

[0266]

[0267] wherein, m represents the number of LMCI subjects; represents the weighted total node potential increment matrix of the ith subject; represents the average change of the node potential in the process of evolution from LMCI to AD.

[0268] Fourthly, the final features are screened out by the incremental search method, and the function expression formula for finding the optimal brain region feature subset and gene feature subset is as follows:

[0269]

[0270] s.t.1≤i≤N1,1≤j≤N2

[0271] wherein, s.t. represents a constraint condition, and i and j represent the numbers of the brain region feature subset and the gene feature subset, respectively. The brain regions and genes contained in the optimal feature subset combination are the diseased brain regions and pathogenic genes captured by the model.

[0272] 4. Risk prediction module

[0273] The converged feature extraction system can output the corresponding generated gene-brain region directed graph according to the current data of the LMCI patient. By comparing the generated gene-brain region directed graph with the real gene-brain region directed graph of AD, the risk probability of the patient converting to AD in the future can be predicted.

[0274] Firstly, the similarity matrix is calculated based on two weight matrices, and the calculation formula is as follows:

[0275]

[0276] wherein, represents the value of the rth row and cth column of the weight matrix of the generated gene-brain region directed graph; represents the value of the rth row and cth column of the weight matrix of the real AD gene-brain region directed graph;

[0277] Secondly, the similarity P is calculated and the risk prediction is completed according to P, and the calculation formula of the similarity P is as follows:

[0278]

[0279] wherein, S rc represents the value (1 or 0) of the rth row and cth column of the similarity matrix.

[0280] As Figure 14 shown, taking a gene-brain region directed graph containing 10 nodes as an example, the steps of calculating the similarity P between two brain networks are illustrated.

[0281] The risk prediction module comprises the following procedures:

[0282] The generated gene-brain region directed graph and the AD actual gene-brain region directed graph are represented by figures (a) and (b) respectively, and the corresponding weight matrices are shown by figures (c) and (d). First, the similarity between the edges of the generated gene-brain region directed graph of the patient and the AD real gene-brain region directed graph is calculated. For example, the similarity of the 2nd row and 3rd column of the matrix is calculated, that is, The value of the 2nd row and 3rd column of the similarity matrix S is 0. Secondly, the proportion of the number of elements with a value of 1 in the similarity matrix S is counted, and the similarity P value of the two networks is obtained as 0.30. It is indicated that the similarity of the predicted patient network and the real AD network is 0.30. The probability of the patient suffering from AD in the future is 30%.

[0283] The above only describes the preferred embodiments of the present application and is not used to limit the present application. Various modifications and changes can be made by those skilled in the art based on the principles and technical solutions of the present application. Any modification, equivalent replacement, improvement, etc. made within the spirit and principles of the present application shall be included in the protection scope of the present application.

Claims

1. A brain region feature extraction system based on motif-driven evolution, characterized in that, include: The preprocessing module is used to acquire brain imaging data and gene data of the target object and construct an initial brain region-gene directed graph; The generator, constructed based on the motif-driven evolution model, is used to detect motifs of preset brain lesions from the initial brain region-gene directed graph, and to generate key brain region-gene directed graphs for detecting preset brain lesions based on the results of motif evolution. A discriminator is used to calculate the similarity between the generated key brain region-gene directed graph and the real key brain region-gene directed graph, and to optimize the model parameters of the generator based on the similarity. The generator includes: The motif discovery and evolution module is used to detect and evolve motif structures in the initial gene-brain region directed graph; The subgraph partitioning module is used to divide the subgraph based on the directed topological potential field of the computational module as the influence range of the module. The directed subgraph evolution module is used to extract the motif tree from each subgraph through motif convolution and form the corresponding evolution subgraph; The evolutionary subgraph fusion module is used to integrate these evolutionary subgraphs to form an evolved gene-brain region directed graph; The motif discovery and evolution module includes: The motif construction unit is used to detect the motifs of the initial brain region-gene directed graph and to construct the motif adjacency matrix and the motif recording matrix. The phantom selection unit is used to select the target phantom with the largest directed topological potential from the phantom adjacency matrix and update the phantom record matrix with the target phantom. The first update unit updates the feature information matrix of nodes within the module based on the information transmission between nodes within the module. The second update unit updates the module adjacency matrix based on the feature information of the updated nodes; The third update unit calculates the weight information of newly generated edges within the module and updates the weight matrix. The directed subgraph evolution module includes: The correlation calculation unit is used to calculate and sort the correlation between gene nodes and motifs within the directed subgraph, obtain a first correlation sequence, and store the first correlation sequence into the motif-node correlation matrix S. (k,0) The association degree between brain region nodes and phantoms within the directed subgraph is calculated and sorted to obtain the second association degree sequence, which is then stored in the phantom-node association matrix S. (k,1) ; The phantom tree construction unit is used to construct the phantom-node association matrix S. (k,1) Construct a motif tree, connect gene nodes within the directed subgraph, and update the motif tree weight matrix. Connect the nodes of the brain regions within the directed subgraph and update the adjacency matrix of the motif tree. Motif tree evolution unit: used to transmit evolutionary information based on the motif tree connection structure and update the feature matrix H of the motif tree nodes. (k) Based on the node update feature information, the edges are recalculated, and the weight matrix of the evolutionary subgraph is updated.

2. The brain region feature extraction system based on motif-driven evolution according to claim 1, characterized in that, The phantom construction unit includes: The first search subunit is used to find all subgraphs of a given size in the initial brain region-gene directed graph by enumeration, record their node information using the initial motif adjacency matrix, and continue enumerating until the search is completed. Randomly generated sub-units: used to classify the sub-graph into the corresponding isomorphic class, and randomly generate a random graph with the same degree as the initial brain region-gene directed graph; The adjacency matrix generates sub-units. The occurrence frequency of multivariate subgraphs in the random graph and the initial brain region-gene directed graph is compared. Multivariate subgraphs with a higher occurrence frequency than the random network are marked and recorded as motifs. At the same time, multivariate subgraphs that do not belong to the motifs are deleted from the adjacency matrix to obtain the final motif adjacency matrix. The record matrix generates sub-units. Based on the found phantoms, the nodes of all phantoms are counted, and a phantom record matrix is ​​constructed. and / or The phantom screening unit includes: Computational subunit: Obtain the node features of each module from the module record matrix, calculate the out-degree module potential and in-degree module potential of each module, and calculate the directed topological potential of each module based on the out-degree module potential and in-degree module potential of each module; Comparison sub-units: Compare the directed topological potentials of each motif, select the target motif with the largest directed topological potential, and update the motif record matrix; and / or The first update unit includes: Update subunit A: is used to calculate the correlation coefficient between each node and the feature information of other nodes based on the updated node feature information. When the correlation coefficient between any two nodes is greater than a preset value, it is determined that the two nodes are connected to each other. Update subunit B: Used to determine the direction of node connections based on the updated node's feature information and potential theory, for any two nodes V i and V j If V i The potential energy is greater than V j If it is high, then determine if V exists. i Point to V j The connecting edges.

3. The brain region feature extraction system based on motif-driven evolution according to claim 2, characterized in that, The correlation calculation unit includes: The average feature calculation subunit is used to select the feature information NH of the nodes contained in the phantom based on the k-th order. (k) Calculate the feature information of the k-th order modulus. The i-th module Features It is obtained by averaging the features of all nodes in the phantom; the nodes are gene nodes or brain region nodes. The motif-node correlation matrix construction sub-unit is used to calculate the feature of each node and the feature information. The correlation coefficients are calculated, and the correlation coefficients corresponding to each node feature are sorted to obtain a correlation degree sequence. The results are then stored in the modulus-node correlation degree matrix. and / or The phantom tree construction unit includes: The first index construction subunit is used to include the gene node with the highest correlation in the first correlation sequence as a leaf node in the motif tree, and to use the correlation coefficient of the included node as its weight. Iteration: Calculate the correlation coefficient between each unincluded node and the nodes already included in the motif tree to obtain an auxiliary correlation sequence. Summarize the auxiliary correlation sequence and the correlation sequence, and select the gene node with the highest correlation as a leaf node in the motif tree, and use the correlation coefficient of the included node as its weight. Repeat the iteration steps until all gene nodes are included in the motif tree, and then update the weight matrix with the weights of all gene nodes. The second index construction subunit is used to include the brain region node with the highest correlation in the second correlation sequence as a leaf node in the motif tree, and to use the correlation coefficient of the included node as its weight. Iteration: Calculate the correlation coefficient between each unincluded node and the nodes already included in the motif tree to obtain an auxiliary correlation sequence. Summarize the auxiliary correlation sequence and the correlation sequence, and select the gene node with the highest correlation as a leaf node in the motif tree, and use the correlation coefficient of the included node as its weight. Repeat the iteration steps until all gene nodes are included in the motif tree, and then update the weight matrix with the weights of all gene nodes. and / or The motif tree evolution unit includes: The motif tree evolution update subunit updates the node feature information of each node in the motif tree sequentially from top to bottom, thereby updating the motif tree node feature matrix H. (k) The update formula is as follows: Among them, H (k) H represents the feature information of nodes within a k-order motif tree. (k-1) Represents the feature information of nodes within a k-1 order motif tree, WW T (k-1) [i] represents the weight information of the i-th node in the (k-1)-th order, H (k-1) [i] represents the feature information of the i-th node in the (k-1)-th order. This represents the calculation of the update information ΔH propagated downwards along the out-degree edge from the i-th node, where θ represents the learnable parameter matrix and N represents the total number of nodes. Weighted evolution update subunit: based on the updated feature matrix H of the phantom tree nodes (k) The edges and weights between nodes are recalculated to form an evolutionary subgraph, and the weight matrix of the evolutionary subgraph is updated based on the weights in the evolutionary subgraph. The calculation formula is as follows: in, This represents the feature information of the i-th node after the k-th information propagation. The feature information of the j-th node after the t-th information propagation. This represents the weight of the edge between node i and node j after the k-th information propagation.

4. The brain region feature extraction system based on phantom-driven evolution according to claim 2, characterized in that, The evolutionary subgraph fusion module fuses the evolutionary subgraphs corresponding to each motif in the k-1 order gene-brain region directed graph to form a k order gene-brain region directed graph. The calculation formula is as follows: in, Let represent all evolutionary subgraphs formed by the i-th subject in the k-th evolution; Fuse() represents the graph fusion function; k is the number of evolutions performed in the generator; N is the total number of evolutions in the generator; M is the total number of evolutionary subgraphs; D (i) Let be the directed graph evolved by the i-th subject; D is the set of all directed graphs evolved by all subjects.

5. The brain region feature extraction system based on motif-driven evolution according to any one of claims 1-4, characterized in that, The discriminator includes: Directed graph convolutional layers are used to extract the feature matrices of the generated key brain region-gene directed graph and the real key brain region-gene directed graph; Fully connected layer: used to calculate the similarity between the generated key brain region-gene directed graph and the feature matrix of the real key brain region-gene directed graph, and to determine the authenticity of the generated key brain region-gene directed graph based on the similarity. and / or The loss function of the brain region feature extraction system is: In this context, the generator and discriminator are represented by G and D, respectively; E(·) represents the expected value of the variable; x represents the true gene-brain region directed graph; and P... data (x) represents the distribution of data x, D(x) represents the probability that x is a true gene-brain region directed graph; G(z) represents the gene-brain region directed graph generated by normal individuals or AD patients, P z (z) represents the distribution of data z, and D(G(z)) represents the probability of judging G(z) as a real gene-brain region directed graph.

6. The brain region feature extraction system based on phantom-driven evolution according to any one of claims 1-4, characterized in that, Also includes: The risk prediction module compares the key brain region-gene directed graph output by the trained generator with the directed graphs of diseased and real key brain regions, calculates the similarity between the two, and determines the risk of disease based on the similarity. The calculation formula is as follows: Among them, S rc This represents the value in the r-th row and c-th column of the similarity matrix. This represents the value in the r-th row and c-th column of the weight matrix of the generated gene-brain region directed graph; The value in the r-th row and c-th column of the weight matrix representing the actual key brain region-gene directed graph; and / or It also includes a disease gene / brain region extraction module, which is used to calculate the topological potential energy change of each node before and after each stage of evolution, and obtain the potential energy increment matrix corresponding to n nodes. Summing the potential energy increment matrices of the n nodes yields the total potential energy increment matrix ΔT. k Based on the total incremental matrix of nodal potential energy for all subjects, ΔT1, ΔT2, ..., ΔT m The average change in node potential energy is calculated; based on the average change in node potential energy, the final pathological genes / brain regions are selected by incremental search.

Citation Information

Patent Citations

  • Electronic device for predicting Alzheimer's disease based on fuzzy structure entropy propagation model

    CN115101208A