Global interpretation method for discrete dynamic graph neural network

Through the improved static graph interpretation method and point-of-time occlusion technology, a global interpretation sequence of discrete dynamic graph neural network is generated, which solves the problem of lack of global interpretation in the existing technology, and realizes the category global interpretation of discrete dynamic graph neural networks.

CN120259839APending Publication Date: 2025-07-04SOUTH CHINA UNIV OF TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510325496.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-19
Publication Date
2025-07-04

AI Technical Summary

Technical Problem

Existing discrete dynamic graph neural network models are usually used as black box models, lack transparency and credibility, and existing interpretation methods are difficult to provide global explanations and cannot generate global explanations for corresponding dynamic graphs for specific categories.

Method used

By designing a global interpretation method for discrete dynamic graph neural networks, including graph embedding module, sequence embedding module and classification module, using time points to mask non-important time points, improving static graph interpretation methods, generating conceptual existing sequences, and training sequence generation models to extract global interpretation sequences of each category.

Benefits of technology

The global perspective interpretation of discrete dynamic graph neural network is realized, and the problem of migration of static graph local interpretation method to discrete dynamic graph local interpretation is solved, reducing the difficulty of convergence during the training process of concept classification model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120259839A_ABST
    Figure CN120259839A_ABST
Patent Text Reader

Abstract

The invention discloses a global explanation method for a discrete dynamic graph neural network. The method comprises the following steps: expanding a local explanation method of a static graph and extracting local explanation of a dynamic graph; training a concept classification model at time points one by one, and extracting a concept existence sequence of each local interpretation sequence by using the model; training a sequence generation model by using the concept existence sequence, and generating a global interpretation sequence corresponding to each category; and converting each global interpretation sequence into a corresponding dynamic graph sequence by utilizing the generated global interpretation sequences and mapping of local interpretations to concepts by a concept classification model so as to obtain global level interpretations about each category. By providing global level interpretation, an interpretation result with global knowledge is provided for the user, and the user can conveniently understand information learned by the model in a simpler mode.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of brain network analysis, and in particular to a global interpretation method for discrete dynamic graph neural networks. Background Art

[0002] With the development of graph neural networks, people often construct functional magnetic resonance imaging (fMRI) data into dynamic brain functional network data and use discrete dynamic graph neural networks to analyze these fMRI data.

[0003] Due to the complexity of discrete dynamic graph neural networks, existing discrete dynamic graph neural network models are usually used as black-box models, which challenges the transparency and credibility of the models. The interpretation methods of discrete dynamic graph neural networks aim to reveal the decision-making mechanism of the models on dynamic graph data and help users understand how the models make predictions based on the temporal evolution of the graph structure and the changes in node features. Due to the spatio-temporal complexity of dynamic graph data, traditional model interpretation methods are difficult to be directly applied to the interpretation of dynamic graph neural network models. Therefore, specialized dynamic graph model interpretation methods are needed to address their unique challenges. Existing methods usually perform local interpretation for specific samples by means of perturbation, gradient, and attention, and give important edges, nodes, or time points of the dynamic graph.

[0004] However, the existing interpretation methods of discrete dynamic graph neural networks often focus on local and specific-sample interpretations and cannot generate global interpretations of the dynamic graphs corresponding to specific categories, making it still difficult to comprehensively understand the decision-making of dynamic graph neural network models. Summary of the Invention

[0005] The object of the present invention is to overcome the deficiencies of the prior art and propose a global interpretation method for discrete dynamic graph neural networks, which can generate corresponding global interpretation graph sequences for each category to be interpreted.

[0006] To achieve the above object, the technical solution provided by the present invention is: a global interpretation method for discrete dynamic graph neural networks, including the following steps:

[0007] 1) Obtain the fMRI data of the subjects, and calculate the dynamic brain functional network data of the subjects. The dynamic brain functional network data corresponding to all subjects jointly constitute the dynamic brain functional network dataset corresponding to the fMRI data. Among them, threshold the matrix at each time point in the dynamic functional connection sequence of the fMRI data of the subjects to obtain the adjacency matrix corresponding to each time point, which is used as the graph corresponding to the time point. Each node in the graph represents a brain region, and each node feature represents the corresponding brain region feature. Arrange each graph in the original time order to obtain the discrete dynamic graph sequence corresponding to the subject. The discrete dynamic graph sequence of the subject and the label corresponding to the subject category jointly constitute the dynamic brain functional network data;

[0008] 2) Design and train a discrete dynamic graph neural network based on the characteristics of the dynamic brain functional network dataset as the discrete dynamic graph neural network to be explained. The discrete dynamic graph neural network consists of a graph embedding module, a sequence embedding module, and a classification module. For the discrete dynamic graph sequence corresponding to each subject, the graph embedding module calculates the embedding vector of the graph through the adjacency matrix and node features corresponding to each time point. The sequence embedding module organizes the embedding vectors of the graphs at each time point into a sequence to obtain a time series embedding vector. The classification module uses this time series embedding vector to obtain the probability of each label corresponding to the discrete dynamic graph sequence as the output of the discrete dynamic graph neural network;

[0009] 3) For the discrete dynamic graph neural network to be explained, extract the importance of each time point on each discrete dynamic graph sequence in the dynamic brain functional network dataset corresponding to the given fMRI data one by one;

[0010] 4) Use the importance of the time points extracted in step 3) and utilize a discrete dynamic graph interpretation method improved based on a static graph interpretation method to extract the dynamic graph local interpretation sequence for each discrete dynamic graph sequence. The specific improvement is as follows: Use the importance of the time points to mask the unimportant time points in the dynamic graph local interpretation sequence, so that the sequence embedding module calculates the sequence embedding only based on the information of the important time points, and use this sequence embedding to continue to input the subsequent classification module, so that the process of learning the interpretation is only affected by the important time points;

[0011] 5) Use the local interpretation sequence of the dynamic graph obtained in step 4) to train an improved concept classification model, enabling it to use the input local interpretation sequence of the dynamic graph to best fit the output of the discrete dynamic graph neural network based on the discrete dynamic graph sequence; after training, the trained concept classification model can convert each sequence into a concept existence sequence and use this concept existence sequence for classification; the specific improvement of the concept classification model is: replacing the gumbel softmax of the hard trick originally assigned to different concepts with a softmax with temperature, and the temperature gradually decreases during the training process; where gumbel softmax is to use the reparameterization trick to learn a parameterized distribution, which is used to sample a probability value between [0,1], and the hard trick is used to convert the sampled probability value into a value of 0 or 1 with a specified threshold while retaining the gradient.

[0012] 6) Use the concept classification model trained in step 5) to extract the concept existence sequence of each discrete dynamic graph sequence.

[0013] 7) Use the importance of the time points calculated in step 3) to select the starting time point of the data; concatenate the category, the starting time point, and the concept existence sequence obtained in step 6) to train a sequence generation model.

[0014] 8) Use the sequence generation model trained in step 7), set the category to the category to be explained, and set the starting time point to 0 to generate an explanation sequence for the corresponding category, obtaining a global explanation for the category.

[0015] Furthermore, in step 1), for each subject, divide the fMRI data into brain regions to obtain an fMRI data sequence of the activation of each brain region changing over time, and at the same time obtain a label corresponding to the category of the corresponding subject; calculate the corresponding dynamic functional connection sequence for the fMRI data sequence of each subject using the given step size and window length, where each time point is the correlation coefficient calculated for the activation between brain regions within a time window; threshold the matrix at each time point in the dynamic functional connection sequence to obtain the adjacency matrix corresponding to each time point, which serves as the graph corresponding to the time point. Each node in the graph represents a brain region, and each node feature represents the corresponding brain region feature. Arrange each graph in the original time order to obtain the discrete dynamic graph sequence corresponding to the subject. Among them, the thresholding is based on a given threshold, setting the positions in the matrix greater than the threshold to 1 and the remaining positions to 0.

[0016] Furthermore, in step 3), use a perturbation-based method to calculate the time importance; during perturbation, replace the embedding vectors of randomly selected time point graphs in the original sequence with noise vectors The changes in the final classification results are evenly distributed to the time points replaced this time. After sampling a preset number of times, the importance of the time points is obtained. The perturbation-based method determines the importance of each edge in the discrete dynamic graph sequence by masking a part of the edges on the input discrete dynamic graph sequence and examining the impact of the masked edges on the output of the discrete dynamic graph neural network to be explained.

[0017] Further, in step 4), the importance threshold is calculated using the importance of the time points obtained in step 3) to filter the important time points. Specifically, a threshold a is selected i-thresh , and the points with importance lower than this threshold a i-thresh are set to 0, and the points higher than this threshold a i-thresh are set to 1 to obtain the time mask of the sample, which is used to identify whether a time point is an important time point; the process of extracting the local interpretation sequence of the dynamic graph is as follows:

[0018] For a discrete dynamic graph sequence G i , and the trained discrete dynamic graph neural network φ(·), train an improved dynamic graph local interpreter φ(·) so that it can obtain the local interpretation sequence of the dynamic graph based on the discrete dynamic graph sequence G i ; The improvement of the dynamic graph local interpreter lies in: during the forward propagation of the dynamic graph local interpreter, for the local interpretation embedding sequence of the dynamic graph obtained using the local interpretation sequence of the dynamic graph, the corresponding time points are masked using the time mask of the sample. For unimportant time points, they are replaced with noise vectors , and the remaining points retain the graph embeddings obtained based on the interpretation graph; the training process of the dynamic graph local interpreter is to minimize the objective function:

[0019]

[0020] In the formula, J Local (·) represents the loss function during the training of the dynamic graph local interpreter, G i represents the i-th discrete dynamic graph sequence, n represents the number of all discrete dynamic graph sequences in the dataset, represents the output of the discrete dynamic graph neural network to be explained, CrossEntropy represents the cross-entropy loss function, and φ θ (·) represents the dynamic graph local interpreter φ(·) using the parameter θ; finally, use the trained dynamic graph local interpreter to calculate and obtain the local interpretation of the dataset

[0021] Further, in step 5), the concept classification model consists of a graph embedding module, a concept module, a sequence module, and a classification module. Its forward propagation process is as follows: For a dynamic graph local interpretation sequence at each time point, it is disassembled into corresponding connected components, where represents the graph corresponding to the j-th time point on the i-th dynamic graph local interpretation sequence; during the forward propagation of the concept classification model, each connected component first passes through the graph embedding module to obtain the corresponding graph embedding vector, and then the concept module maps these vectors to a set of randomly initialized learnable concept vectors On, the concept existence vector of each connected component is obtained by calculating the assignment function and assigning each connected component to one of the concept vectors:

[0022]

[0023] The result of the obtained assign function is the concept existence vector corresponding to the connected component , represents the corresponding graph of the k-th connected component; above, p m represents the m-th concept vector, n c represents the total number of concept vectors, represents the embedding vector obtained by the corresponding connected component passing through the graph embedding module of the discrete dynamic graph neural network. τ is the temperature parameter, which is used to control the closeness of the result of the assign function to 0 or 1 and gradually decreases during the training process. sim is the similarity function for calculating the similarity between two vectors;

[0024] Taking the element-wise maximum of the concept existence vectors corresponding to all connected components in a graph, a concept existence vector corresponding to the graph is obtained; concatenating the concept existence vectors of all graphs in a dynamic graph local interpretation sequence in chronological order to obtain a concept existence sequence; after obtaining the concept existence sequence through the concept module, the corresponding sequence is provided to the subsequent sequence module, and the sequence module outputs the concept existence sequence as a sequence embedding vector. The classification module uses this sequence embedding vector for classification and outputs the probability corresponding to each category;

[0025] The concept classification model learns its behavior by fitting the output of the discrete dynamic graph neural network Φ(·). The optimal parameters are found using the following formula during the training process:

[0026]

[0027] In the formula, ω * represents the optimal parameters, CrossEntropy is the cross-entropy loss, and MLP ω(·) represents the classification module, TempAGG ω (·) represents the time embedding module, GNN ω (·) represents the graph embedding module, G i represents the i-th discrete dynamic graph sequence in the original dataset. The subscript ω represents the model using ω as a parameter, argmin ω means to optimize ω to the minimum function value.

[0028] Furthermore, in step 6), based on the optimal parameter ω of the concept classification model obtained in step 5) * , the concept existence sequence set of the dataset can be extracted where represents the concept existence sequence corresponding to the i-th dynamic graph local interpretation sequence, represents the concept existence vector at the j-th time point of the i-th concept existence sequence, and there is where represents the application of the optimal parameter ω * to the graph embedding module in the concept classification model.

[0029] Furthermore, in step 7), the sequence generation model extracts the concept global information corresponding to each category by fitting the concept existence sequence corresponding to each category; the sequence generation model consists of a sequence embedding module and a classification module. The input data is the vocabulary sequence organized by the concept existence sequence obtained in step 6) in terms of time and feature dimensions, and autoregressive learning is performed on this data; the sequence embedding module follows the formula:

[0030]

[0031] In the formula, represents the hidden vector at the t-th position on the i-th vocabulary sequence, represents the hidden vector at the (t - 1)-th position on the i-th vocabulary sequence, represents the actual value at the t-th position of the i-th vocabulary sequence, f Seq (·) represents the sequence embedding module; the classification module reads the hidden state of the previous moment of the sequence embedding module and outputs the probability of each vocabulary that may be filled in at the current moment, which follows the following formula:

[0032]

[0033] In the formula, is the probability that the value at the t-th position of the i-th vocabulary sequence corresponds to each word in the vocabulary, f pred (·) is the corresponding classification module; the initial state It is calculated from the category to which the corresponding vocabulary sequence belongs and a marker of the starting position of a concept, where the starting position of the concept is represented by the subscript of the first non-0 position in the time mask obtained in step 4); the optimization process of the sequence generation model is obtained by minimizing the following loss function:

[0034]

[0035] In the formula, L(·) is the loss function, and log(·) is the logarithmic function. denotes the probability that the sequence generation model outputs a vocabulary sequence under the condition of the initial state and the current parameter β; here represents that the vocabulary sequence output before time point t is is the condition for the conditional probability; T is the total length of the vocabulary sequence.

[0036] Furthermore, in step 8), based on the sequence generation model trained in step 7), the category to generate the explanatory graph and the starting position 0 are jointly calculated to obtain the initial state, and the initial state is used to generate a sequence of length T; after obtaining the initial state, the complete vocabulary sequence from 1 to T is gradually calculated to obtain the global explanation on the category, and this explanation is composed of the concept existence sequences of each category; using the concept existence vectors of the connected components of each local explanatory sequence of the dynamic graph in step 5), the connected component most similar to each prototype can be selected as the representative graph of the corresponding concept vector; by selecting and putting the concept existence vectors from the representative graphs into the sequence, the discrete dynamic graph sequence corresponding to the global explanation can be obtained.

[0037] Compared with the prior art, the present invention has the following advantages and beneficial effects:

[0038] 1. The present invention provides an explanation for the discrete dynamic graph neural network from a global perspective based on categories.

[0039] 2. Aiming at the problem existing in directly migrating the static graph local explanation method to the discrete dynamic graph local explanation method, the present invention uses time points to mask unimportant time points, enabling the static graph local explanation method to be applied to the discrete dynamic graph local explanation.

[0040] 3. The present invention improves the softmax process in the concept assignment process by removing the hard trick and introducing the temperature coefficient, reducing the difficulty of convergence in the training process of the concept classification model. BRIEF DESCRIPTION OF THE DRAWINGS

[0041] Figure 1 is the overall process of obtaining the global explanation in the embodiment.

[0042] Figure 2Schematic diagram of the structure of the discrete dynamic graph neural network in the embodiment; in the figure, GNN represents the graph embedding module, cls_token represents the token for extracting the time series aggregation, and y i ' represents the output of the model.

[0043] Figure 3 Schematic diagram of the local interpretation extraction process in the embodiment; in the figure, TGNN represents the corresponding discrete dynamic graph neural network, mask represents the masking operation, represents the output obtained by the interpretation, and y i ' represents the original output of the model.

[0044] Figure 4 Schematic diagram of the concept classification model in the embodiment; in the figure, TGNN represents the corresponding discrete dynamic graph neural network, and GNN represents the graph embedding module.

[0045] Figure 5 Schematic diagram of the sequence generation model in the embodiment; in the figure, TGNN represents the corresponding discrete dynamic graph neural network, represents the concept existence vector at the j-th time point of the i-th concept existence sequence, and respectively represent the probabilities that the values at the t-th position of the i-th sequence in the sequence are 0 and 1 in the vocabulary. Detailed implementation manners

[0046] The present invention will be further described in detail below in conjunction with the embodiments and the accompanying drawings, but the implementation manners of the present invention are not limited thereto.

[0047] As Figures 1 to 5 shown, this embodiment discloses a global interpretation method for a discrete dynamic graph neural network. This method extracts the local interpretation sequence of the dynamic graph through an improved static graph interpretation method, uses the local interpretation sequence of the dynamic graph to perform concept assignment to obtain the concept existence sequence, and finally uses the sequence generation model to extract the class-specific global interpretation.

[0048] The specific implementation of this method includes the following steps:

[0049] 1) Obtain the fMRI data of the subject, preprocess the fMRI data, and obtain the corresponding labels at the same time; the preprocessing includes time slice correction, head motion correction, global normalization, and spatial standardization operations.

[0050] 2) Perform brain region partitioning on the preprocessed fMRI data. For the fMRI data sequence of each subject, use a sliding window with a length of 50 and a step size of 10 to divide the sequence into multiple windows. For each window, calculate the functional connectivity matrix at the corresponding time step by computing the Pearson correlation coefficient; organize the functional connectivity matrices by time to obtain the dynamic functional connectivity sequence of the subject.

[0051] 3) Threshold the matrix at each time point in the dynamic functional connectivity matrix of each subject to obtain the adjacency matrix corresponding to each time point, which serves as the graph corresponding to the time point. Each node in the graph represents a brain region, and each node feature represents the corresponding brain region feature. Arrange each graph in the original time order to obtain the discrete dynamic graph sequence corresponding to the subject. Among them, the thresholding is based on a given threshold, setting the positions in the matrix greater than the threshold to 1 and the remaining positions to 0; the discrete dynamic graph sequence of the subject and the label of the corresponding category of the subject jointly form the dynamic brain functional network data; the dynamic brain functional network data corresponding to all subjects jointly form the dynamic brain functional network dataset corresponding to the fMRI data.

[0052] 4) As Figure 2 shown, design a discrete dynamic graph neural network composed of a graph embedding module, a sequence embedding module, and a classification module. The input of this network is a discrete dynamic graph sequence, and the category corresponding to the discrete dynamic graph sequence is given.

[0053] 5) Calculate the temporal importance item by item on the dynamic brain functional network dataset obtained from the fMRI data using a perturbation-based method. First, calculate the average of all graph embeddings as the noise vector during the perturbation process, where n represents the number of all discrete dynamic graph sequences in the dataset, and l represents the length of each discrete dynamic graph sequence, represents the graph embedding of the graph g corresponding to the j-th time point of the i-th discrete dynamic graph sequence, obtained from i (l) , where f is obtained, and f GNN (·) represents the graph embedding module of the discrete dynamic graph neural network. During perturbation, randomly select a random number of time points from all time points, replace the output of the graph embedding module at this time point with the average value, and evenly distribute the change in the final classification result to the time points replaced this time. After a certain number of samplings, obtain the importance of the time points.

[0054] 6) Extract the local interpretation sequence of the dynamic graph. First, calculate the importance threshold using the importance of the time points to filter the important time points. Specifically, use the finite difference method to calculate the second derivative of the data and use the elbow method to select a threshold a i-thresh, set the points with importance lower than this threshold to 0 and the points higher than this threshold to 1 to obtain the time mask of the sample, which is used to identify whether a time point is an important time point. The process of extracting the local interpretation sequence of the dynamic graph is as follows:

[0055] As Figure 3 shown, for a discrete dynamic graph sequence G i , and the discrete dynamic graph neural network Φ(·) trained above, train an improved dynamic graph local interpreter φ(·) so that it can obtain the local interpretation sequence of the dynamic graph based on the discrete dynamic graph sequence G i where the main improvement of the dynamic graph local interpreter lies in: during the forward propagation of the dynamic graph local interpreter, for the corresponding graph embedding sequence obtained by using the local interpretation sequence of the dynamic graph, mask the corresponding time points with the time mask of the sample. For unimportant time points, in the embodiment, the already calculated graph average value is used for replacement, and the remaining points retain the graph embeddings obtained based on the interpreted graph. The training process of the dynamic graph local interpreter is to minimize the objective function In the formula, J

[0056]

[0057] where J(·) represents the loss function during the training of the dynamic graph local interpreter, Local represents the output of the discrete dynamic graph neural network to be interpreted, CrossEntropy represents the cross-entropy loss function, and φ (·) represents the dynamic graph local interpreter φ(·) using the parameter θ; finally, use the trained dynamic graph local interpreter to calculate and obtain the local interpretation of the data set θ (·); finally, use the trained dynamic graph local interpreter to calculate and obtain the local interpretation of the data set

[0058] 7) Train an improved concept classification model to make it use the input local interpretation sequence of the dynamic graph and fit the output of the discrete dynamic graph neural network to be interpreted based on the discrete dynamic graph sequence as much as possible. As Figure 4 shown, the concept classification model consists of a graph embedding module, a concept module, a sequence module, and a classification module. Its forward propagation process is as follows: for each time point on an input local interpretation sequence of the dynamic graph , where represents the graph corresponding to the j-th time point on the i-th local interpretation sequence of the dynamic graph; during the forward propagation of the model, first, each connected component passes through the graph embedding module to obtain the corresponding graph embedding vector, and then the concept module maps these vectors to a set of randomly initialized learnable concept vectors , and through calculating the assignment function:

[0059]

[0060] The result of the obtained assign function is then the connected component The corresponding concept existence vector, indicating the corresponding graph of the k-th connected component; in the above formula, p m represents the m-th concept vector, and n c represents the total number of concept vectors, indicating the corresponding connected component The embedding vector obtained by the graph embedding module of the discrete dynamic graph neural network, τ is the temperature parameter, used to control the closeness of the result of the assign function to 0 or 1. In the embodiment, the temperature linearly decreases from 1 to 0.01 as the number of training rounds increases. The sim in the embodiment is calculated by the following formula

[0061]

[0062] In the formula, log represents the logarithmic function, represents the square of the Euclidean distance, and ∈ represents the parameter p for smoothing m represents the m-th concept vector, and is selected as 0.001 in the embodiment.

[0063] Taking the element-wise maximum of the concept existence vectors corresponding to all the connected components in a graph to obtain a concept existence vector corresponding to the graph. Concatenating the concept existence vectors of all the graphs in a dynamic graph local interpretation sequence in chronological order to obtain a concept existence sequence. After obtaining the concept existence sequence through the concept module, providing the corresponding sequence to the subsequent sequence module, the sequence module outputs the concept existence sequence as a sequence embedding vector, and the classification module uses this sequence embedding vector for classification and outputs the probability corresponding to each category.

[0064] The concept classification model learns its behavior by fitting the output of the discrete dynamic graph neural network Φ(·), and the optimal parameters are found using the following formula during the training process

[0065]

[0066] In the formula, ω * represents the optimal parameter, CrossEntropy is the cross-entropy loss, and MLP ω (·) represents the classification module, and TempAGG ω (·) represents the time embedding module, and GNN ω (·) represents the graph embedding module, and G i represents the i-th discrete dynamic graph sequence in the original dataset, and the subscript of ω represents the model using ω as the parameter, and argmin ω represents optimizing ω to the minimum function value.

[0067] 8) Extract the concept existence sequence. Using the optimal parameter ω of the concept classification model * , the concept existence sequence set of the dataset can be extracted where represents the concept existence sequence corresponding to the i-th local interpretation sequence of the dynamic graph, represents the concept existence vector at the j-th time point of the i-th concept existence sequence, and there is where represents the graph embedding module in the concept classification model that applies the optimal parameter ω * ;

[0068] 9) As Figure 5 shown, fit a generative model using the concept existence sequence. The function of the sequence generation model is to calculate the probability that the next point in the sequence is 0 or 1 based on the concept existence sequence corresponding to each category. Specifically, the input data is the concept existence sequence obtained in step 8) flattened after discretization with a threshold of 0.5. The original concept existence sequence with a length of l and a dimension of n c at each time point is flattened into a one-dimensional vocabulary sequence with a length of l·n c , and each position is 0 or 1. Therefore, in the embodiment, the vocabulary of the vocabulary sequence has only two values, 0 and 1. The described sequence generation model consists of a sequence embedding module and a classification module. The sequence embedding module follows the formula:

[0069]

[0070] In the formula, represents the hidden vector at the t-th position on the i-th vocabulary sequence, represents the hidden vector at the (t - 1)-th position on the i-th vocabulary sequence, represents the actual value at the t-th position of the i-th vocabulary sequence, and f Seq (·) is the sequence embedding module; the classification module reads the hidden state of the previous moment of the sequence module and outputs the probability of each vocabulary that may be filled in at the current moment, which follows the following formula:

[0071]

[0072] In the formula, is the probability that the value at the t-th position of the i-th vocabulary sequence corresponds to each word in the vocabulary, and f pred (·) is the corresponding classification module; the initial state Calculated from the category to which the corresponding vocabulary sequence belongs and a marker of the starting position of a concept, where the starting position of the sequence is represented by the subscript of the first non-zero position in the time mask obtained. In the embodiment, an embedding layer is used to number 0 and 1 of the category to which it belongs as corresponding numerical values respectively, and the starting position is numbered as the original numerical value + 2 to form an additional token, and a learnable embedding layer is used to embed it into a numerical value, which is spliced in front of the flattened concept existence sequence as a starting symbol, and the output of the sequence model for these two inputs is directly discarded without supervision, and the model is trained to learn a suitable The optimization process of the sequence generation model is obtained by minimizing the following loss function:

[0073]

[0074] In the formula, log(·) is the logarithmic function, Indicates that the probability that the sequence generation model outputs the sequence under the condition of the initial state and the current parameter β, Here, it represents that the vocabulary sequence output before time point t is is the condition for the conditional probability. T is the total length of the vocabulary sequence.

[0075] Furthermore, based on the trained sequence generation model, the category for which the explanatory graph needs to be generated and the starting position 0 are jointly calculated to obtain the initial state, and the initial state is used to generate a vocabulary sequence of length T. After obtaining the initial state, the complete vocabulary sequence from 1 to T is gradually calculated to obtain the global explanation in terms of categories, which is composed of the concept existence sequences of each category. Using the previous concept existence row vectors to collect the prototype graphs that each concept vector can represent, the connected component closest to each prototype can be selected as the representative graph of the corresponding concept vector. Using the concept existence vectors to select and put into the sequence from the representative graphs, the discrete dynamic graph sequence corresponding to the global explanation can be obtained.

[0076] The above embodiments are the preferred embodiments of the present invention, but the embodiments of the present invention are not limited by the above embodiments. Any other changes, modifications, substitutions, combinations, and simplifications made without departing from the spirit and principle of the present invention shall be equivalent replacement methods and are all included in the protection scope of the present invention.

Claims

1. A global interpretation method for discrete dynamic graph neural networks, characterized in that The steps include the following: 1) Obtain the fMRI data of the subjects, and calculate the dynamic brain functional network data of the subjects. The dynamic brain functional network data corresponding to all subjects together constitute the dynamic brain functional network dataset corresponding to the fMRI data. Among them, threshold the matrix at each time point in the dynamic functional connection sequence of the fMRI data of the subjects to obtain the adjacency matrix corresponding to each time point, which is used as the graph corresponding to the time point. Each node in the graph represents a brain region, and each node feature represents the corresponding brain region feature. Arrange each graph in the original time order to obtain the discrete dynamic graph sequence corresponding to the subject. The discrete dynamic graph sequence of the subject and the label corresponding to the category of the subject together constitute the dynamic brain functional network data; 2) Design and train a discrete dynamic graph neural network according to the characteristics of the dynamic brain functional network dataset as the discrete dynamic graph neural network to be explained. The discrete dynamic graph neural network consists of a graph embedding module, a sequence embedding module, and a classification module. For the discrete dynamic graph sequence corresponding to each subject, the graph embedding module calculates the embedding vector of the graph through the adjacency matrix and node features corresponding to each time point. The sequence embedding module organizes the embedding vectors of the graphs at each time point into a sequence to obtain the time series embedding vector. The classification module uses this time series embedding vector to obtain the probability of each label corresponding to the discrete dynamic graph sequence as the output of the discrete dynamic graph neural network; 3) For the discrete dynamic graph neural network to be explained, extract the importance of each time point on each discrete dynamic graph sequence in the dynamic brain functional network dataset corresponding to the given fMRI data one by one; 4) Use the importance of the time points extracted in step 3), and use a discrete dynamic graph interpretation method improved based on the static graph interpretation method to extract the dynamic graph local interpretation sequence for each discrete dynamic graph sequence. The specific improvement is as follows: Use the importance of the time points to mask the unimportant time points in the dynamic graph local interpretation sequence, so that the sequence embedding module calculates the sequence embedding only based on the information of the important time points, and use this sequence embedding to continue to input the subsequent classification module, so that the process of learning the interpretation is only affected by the important time points; 5) Use the local interpretation sequence of the dynamic graph obtained in step 4) to train an improved concept classification model, enabling it to use the input local interpretation sequence of the dynamic graph to best fit the output of the discrete dynamic graph neural network based on the discrete dynamic graph sequence; after training, the trained concept classification model can convert each sequence into a concept existence sequence and use this concept existence sequence for classification; the specific improvement of the concept classification model is: replace the gumbel softmax of the hard trick originally assigned to different concepts with a softmax with temperature, and the temperature gradually decreases during the training process; where gumbel softmax is to use the reparameterization trick to learn a parameterized distribution, which is used to sample a probability value between [0,1], and the hard trick is used to convert the sampled probability value into a value of 0 or 1 at a specified threshold while retaining the gradient; 6) Use the concept classification model trained in step 5) to extract the concept existence sequence of each discrete dynamic graph sequence; 7) Use the importance of the time points calculated in step 3) to select the starting time point of the data; concatenate the category, the starting time point, and the concept existence sequence obtained in step 6) to train a sequence generation model; 8) Use the sequence generation model trained in step 7), set the category to the category to be explained, and generate an explanation sequence for the corresponding category with the starting time point being 0 to obtain the global explanation for the category.

2. The global interpretation method for a discrete dynamic graph neural network according to claim 1, characterized in that: In step 1), for each subject, divide the fMRI data into brain regions to obtain an fMRI data sequence of the activation situation of each brain region changing over time, and at the same time obtain a label corresponding to the category of the corresponding subject; for the fMRI data sequence of each subject, calculate the corresponding dynamic functional connectivity sequence using the given step size and window length, where each time point is the correlation coefficient calculated for the activation situation between brain regions within a time window; Threshold the matrix at each time point in the dynamic functional connectivity sequence to obtain the adjacency matrix corresponding to each time point as the graph corresponding to the time point. Each node in the graph represents a brain region, and each node feature represents the corresponding brain region feature. Arrange each graph in the original time order to obtain the discrete dynamic graph sequence corresponding to the subject. Among them, the thresholding is based on a given threshold, setting the positions in the matrix greater than the threshold to 1 and the remaining positions to 0.

3. The global interpretation method for a discrete dynamic graph neural network according to claim 2, characterized in that: In step 3), a perturbation-based method is used to calculate the time importance; during perturbation, the embedding vectors of randomly selected time-point graphs in the original sequence are replaced with noise vectors. The changes in the final classification results are evenly distributed to the time points replaced this time. After sampling a preset number of times, the importance of the time points is obtained. The perturbation-based method determines the importance of each edge in the discrete dynamic graph sequence by masking a part of the edges on the input discrete dynamic graph sequence and examining the impact of the masked edges on the output of the discrete dynamic graph neural network to be explained.

4. A global interpretation method for discrete dynamic graph neural networks according to claim 3, characterized in that: In step 4), the importance threshold is calculated using the importance of the time points obtained in step 3) to filter important time points. Specifically, a threshold a is selected i-thresh , and the points with importance lower than this threshold a i-thresh are set to 0, and the points higher than this threshold a i-thresh are set to 1 to obtain the time mask of the sample, which is used to identify whether a time point is an important time point. The process of extracting the local interpretation sequence of the dynamic graph is as follows: For a discrete dynamic graph sequence G i , and the trained discrete dynamic graph neural network Φ(·), train an improved dynamic graph local interpreter φ(·) to enable it to be based on the discrete dynamic graph sequence G i Get the local explanation sequence of the dynamic graph The improvement of the local interpreter of dynamic graphs is as follows: in the process of forward propagation of the local interpreter of dynamic graphs, for the local interpretation embedding sequence of dynamic graphs obtained by using the local interpretation sequence of dynamic graphs, the corresponding time points are masked by the time mask of the sample, and the unimportant time points are replaced by the noise vector e, and the remaining points retain the graph embedding obtained based on the interpretation graph; the training process of the local interpreter of dynamic graphs is to minimize the objective function: Where J Local (·) represents the loss function when training the dynamic graph local interpreter, G i represents the i-th discrete dynamic graph sequence, n represents the number of all discrete dynamic graph sequences in the dataset, represents the output of the discrete dynamic graph neural network to be interpreted, CrossEntropy represents the cross-entropy loss function, φ θ (·) represents the dynamic graph local interpreter φ(·) using the parameter θ; finally, the local interpretation of the dataset is calculated and obtained using the trained dynamic graph local interpreter 5. The global interpretation method for a discrete dynamic graph neural network according to claim 4, characterized in that: In step 5), the concept classification model consists of a graph embedding module, a concept module, a sequence module, and a classification module. Its forward propagation process is as follows: For each time point on an input dynamic graph local interpretation sequence , it is disassembled into corresponding connected components, where represents the graph corresponding to the j-th time point on the i-th dynamic graph local interpretation sequence; during the forward propagation of the concept classification model, each connected component is first passed through the graph embedding module to obtain the corresponding graph embedding vector, and then the concept module maps these vectors to a set of randomly initialized learnable concept vectors P = {p1, p2, …, p m , …, p nc}, and the concept existence vector of the connected component is obtained by calculating the assignment function to assign each connected component to one of the concept vectors: The result of the obtained assign function is the connected component The corresponding concept existence vector, indicating the corresponding graph of the k-th connected component; above, p m represents the m-th concept vector, n c represents the total number of concept vectors, indicating the corresponding connected component The embedding vector obtained by the graph embedding module of the discrete dynamic graph neural network. τ is the temperature parameter, which is used to control the closeness of the result of the assign function to 0 or 1 and gradually decreases during the training process. sim is the similarity function for calculating the similarity between two vectors; Take the element-wise maximum value of the concept existence vectors corresponding to all connected components in a graph to obtain a concept existence vector corresponding to the graph; concatenate the concept existence vectors of all graphs in a local interpretation sequence of the dynamic graph in time order to obtain a concept existence sequence; after obtaining the concept existence sequence through the concept module, provide the corresponding sequence to the subsequent sequence module. The sequence module outputs the concept existence sequence as a sequence embedding vector, and the classification module uses this sequence embedding vector for classification and outputs the probability corresponding to each category; The described concept classification model learns its behavior by fitting the output of the discrete dynamic graph neural network Φ(·), and the optimal parameters are found using the following formula during the training process: where ω * represents the optimal parameter, CrossEntropy is the cross-entropy loss, and MLP ω (·) represents the classification module, and TempAGG ω (·) represents the temporal embedding module, and GNN ω (·) represents the graph embedding module, and G i represents the i-th discrete dynamic graph sequence in the original dataset, the subscript of ω represents the model using ω as a parameter, and argmin ω means optimizing ω to the minimum function value.

6. The global interpretation method for a discrete dynamic graph neural network according to claim 5, characterized in that: In step 6), based on the optimal parameter ω of the concept classification model obtained in step 5) * , the concept existence sequence set of the data set can be extracted where represents the concept existence sequence corresponding to the i-th local interpretation sequence of the dynamic graph, represents the concept existence vector at the j-th time point of the i-th concept existence sequence, and there is where represents the graph embedding module in the concept classification model that applies the optimal parameter ω * of the concept classification model.

7. A global interpretation method for discrete dynamic graph neural networks according to claim 6, characterized in that: In step 7), the sequence generation model extracts the concept global information corresponding to each category by fitting the concept existence sequence corresponding to each category; the sequence generation model consists of a sequence embedding module and a classification module, and the input data is the vocabulary sequence organized by the concept existence sequence obtained in step 6) in the dimensions of time and features, and autoregressive learning is performed on this data; the sequence embedding module follows the formula: Wherein, represents the hidden vector at the t-th position in the i-th word sequence, represents the hidden vector at the (t-1)-th position in the i-th word sequence, represents the actual value at the t-th position in the i-th word sequence, f Seq (·) represents the sequence embedding module; the classification module reads in the hidden state of the previous moment of the sequence embedding module and outputs the probability of each word that may be filled in at the current moment, which follows the following formula: In the formula, is the probability that the value corresponding to the t-th position of the i-th vocabulary sequence is each word in the vocabulary, and f pred (·) is the corresponding classification module; the initial state is calculated from the category to which the corresponding vocabulary sequence belongs and a marker at the starting position of a concept, and the starting position of the concept is represented by the subscript of the first non-0 position in the time mask obtained in step 4); the optimization process of the sequence generation model is obtained by minimizing the following loss function: where \(L(\cdot)\) is the loss function and \(\log(\cdot)\) is the logarithmic function, denotes the probability that the sequence generation model outputs a vocabulary sequence under the condition of the initial state and the current parameter \(\beta\), where \(\left.\right.\) represents the vocabulary sequence output before time point \(t\) is the condition for the conditional probability; \(T\) is the total length of the vocabulary sequence.

8. A global interpretation method for discrete dynamic graph neural networks according to claim 7, characterized in that: In step 8), based on the sequence generation model trained in step 7), the category that needs to generate the explanatory graph and the starting position 0 are jointly used to calculate the initial state, and a sequence of length T is generated using the initial state; After obtaining the initial state, the complete vocabulary sequence from 1 to T is gradually calculated to obtain the global explanation on the category, and this explanation is composed of the concept existence sequences of each category; using the concept existence vectors of the connected components of each dynamic graph local explanation sequence in step 5), the connected component most similar to each prototype can be selected as the representative graph of the corresponding concept vector; by selecting and putting the concept existence vectors from the representative graph into the sequence, the discrete dynamic graph sequence corresponding to the global explanation can be obtained.