Method and device for generating trusted computation representing time-space diagram neural network decision factors
By constructing a spatiotemporal subset alliance through the Monte Carlo algorithm and game theory interaction theory, the node interactions of the human skeleton spatiotemporal graph are quantified and node feature masks are generated. This solves the problem that the spatiotemporal interaction relationship is not fully considered in the existing methods, and improves the explanatory power and reliability of the spatiotemporal graph neural network.
Patent Information
- Application Number
- CN202510723216.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-30
- Publication Date
- 2025-09-26
AI Technical Summary
Existing Shapley-based trusted computing methods have difficulty revealing the interactive effects between skeleton nodes in the spatiotemporal graph of human skeleton behavior recognition. Traditional methods fail to fully consider the spatiotemporal interaction relationship, resulting in limited interpretation capabilities of graph neural networks.
Using the Monte Carlo algorithm and game theory interaction theory, the spatiotemporal subset alliance of the human skeleton spatiotemporal graph is constructed to quantify the spatiotemporal interaction value of the nodes, generate the node feature mask, and optimize the node feature mask to reflect the key spatiotemporal interaction relationship.
The explanatory power of spatiotemporal graph neural networks has been improved, which can more clearly reveal the interactions between skeleton nodes and enhance the reliability of behavior recognition.
Smart Images

Figure CN120708273A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of artificial intelligence technology, and in particular relates to a method and device for generating trusted calculations that represent decision factors of spatiotemporal graph neural networks. Background Art
[0002] Human skeleton behavior recognition involves identifying and classifying human behavior by analyzing the human skeleton. By integrating graph neural networks with various temporal learning methods, spatiotemporal graph neural networks can extract complex spatiotemporal dependencies. This allows for human skeleton behavior recognition and predictions. Understanding and interpreting the predictions of spatiotemporal graph networks for human skeleton behavior recognition is crucial for improving the security and reliability of behavior recognition algorithms based on spatiotemporal graph networks.
[0003] Existing Shapley-based trustworthy computational methods for interpreting graph neural networks assume that the interactions between skeleton nodes are evaluated independently. However, the spatiotemporal dependencies between skeleton nodes in the spatiotemporal human skeleton graph for action recognition are very complex, and skeleton nodes associated with target categories often involve dynamic interactions between multiple skeleton nodes. Trustworthy computational methods based on Shapley values can only provide explanations based on the contributions of individual nodes, making it difficult to reveal the impact of interactions between skeleton nodes, which are critical for the recognition and classification of spatiotemporal graph neural networks. Other traditional trustworthy computational methods focus only on individual skeleton nodes or edges and fail to fully consider the spatiotemporal interactions between skeleton nodes, resulting in limited overall interpretation capabilities of graph neural networks. Summary of the Invention
[0004] To solve the above technical problems, the present invention proposes a method and device for generating trusted computing that characterizes decision factors of spatiotemporal graph neural networks.
[0005] According to a first aspect of the present invention, a method for generating a trusted computation representing a decision factor of a spatiotemporal graph neural network is provided, the method comprising the following steps:
[0006] Step S1: Obtain the recognition result of the to-be-explained spatiotemporal graph neural network for skeleton behavior recognition and the input of the to-be-explained spatiotemporal graph neural network corresponding to the recognition result, where the input is a video;
[0007] Step S2: Monte Carlo sampling is performed on the spatiotemporal graph of the human skeleton corresponding to the video to obtain several skeleton node subsets, and a spatiotemporal subset union of each skeleton node is generated based on the skeleton node subsets; a skeleton node is the entire skeleton of a pedestrian in the video frame; wherein, based on the spatiotemporal physical connectivity of the skeleton nodes in the spatiotemporal graph of the human skeleton, the skeleton nodes in the skeleton node subset are subjected to second-order sampling to generate a spatiotemporal subset union of each skeleton node;
[0008] Step S3: determining the spatiotemporal interaction value and marginal effect value of each skeleton node based on the spatiotemporal graph of the human skeleton and the spatiotemporal subset union of each skeleton node; constructing the spatiotemporal interaction matrix of the skeleton nodes based on the spatiotemporal interaction value and marginal effect value of each skeleton node;
[0009] Step S4: Randomly generate node feature masks based on the human skeleton spatiotemporal graph, optimize the node feature masks based on the human skeleton spatiotemporal graph and the skeleton node spatiotemporal interaction matrix, and obtain the final node feature mask as the mask to characterize the decision factors of the spatiotemporal graph neural network.
[0010] Preferably, step S2 performs Monte Carlo sampling on the human skeleton spatiotemporal graph corresponding to the video to obtain a plurality of skeleton node subsets, and generates a spatiotemporal subset union of each skeleton node based on the skeleton node subsets; the skeleton node is the entire skeleton of a pedestrian in the video frame; wherein, based on the spatiotemporal physical connectivity of the skeleton nodes in the human skeleton spatiotemporal graph, the skeleton nodes in the skeleton node subset are subjected to second-order sampling to generate the spatiotemporal subset union of each skeleton node, including:
[0011] Step S21: Generate a human skeleton spatiotemporal graph G = {V, E} corresponding to the video, where the skeleton node set The edge set E=E of the human skeleton space-time graph S ∩E F , T is the number of video frames, is the i-th skeleton node in the t-th frame, t is the frame number of the video, N is the total number of skeleton nodes, i, j are the skeleton node numbers; E S is a set of connected edges, where the connected edges are the edges that connect the skeleton nodes with natural connection relationships in the same frame of the video. is the jth skeleton node in the tth frame, H is the index set of skeleton nodes in several human skeletons in one frame; E F is the time edge set, Temporal edges are edges connecting the same skeleton nodes in consecutive frames. is the i-th skeleton node in the t+1th frame;
[0012] Step S22: Monte Carlo sampling is performed on the human skeleton spatiotemporal graph to obtain several skeleton node subsets, and a spatiotemporal subset union of each skeleton node is generated based on the skeleton node subsets; the skeleton node is the entire skeleton of a pedestrian in the video frame; wherein, second-order sampling is performed on the skeleton nodes in the skeleton node subset based on the spatiotemporal physical connectivity of the skeleton nodes in the human skeleton spatiotemporal graph to generate a spatiotemporal subset union of each skeleton node.
[0013] Preferably, step S3 determines the spatiotemporal interaction value and marginal effect value of each skeleton node based on the spatiotemporal graph of the human skeleton and the spatiotemporal subset union of each skeleton node; and constructs a spatiotemporal interaction matrix of skeleton nodes based on the spatiotemporal interaction value and marginal effect value of each skeleton node, including:
[0014] Step S31: Determine the spatiotemporal interaction value of each skeleton node based on the spatiotemporal graph of the human skeleton and the spatiotemporal subset alliance of each skeleton node. The spatiotemporal interaction value is calculated as follows:
[0015]
[0016] in, is the spatiotemporal interaction value of the skeleton node, For skeleton nodes The second-order spacetime subset union of For skeleton nodes The kth second-order spacetime subset of For skeleton nodes Add to The subsequent profit function, for The profit function, for size;
[0017] Step S32: Determine the marginal effect value of each skeleton node based on the spatiotemporal interaction value of each skeleton node. The calculation formula of the marginal effect value of the skeleton node is:
[0018] and
[0019] in, express A collection of mid-skeleton node features, express The skeleton node feature is added to The collection of skeleton node features after express The prediction results input into the spatiotemporal graph network model, express The prediction results input into the spatiotemporal graph network model;
[0020] Step S33: Based on the spatiotemporal interaction value and marginal effect value of each skeleton node, construct the spatiotemporal interaction matrix M of the skeleton nodes H :
[0021]
[0022] Among them, F(V) represents the set of skeleton node features in the skeleton node set V, and Φ(F(V)) represents the prediction result of F(V) input into the spatiotemporal graph network model.
[0023] Preferably, step S4: randomly generating a node feature mask based on the human skeleton spatiotemporal graph, optimizing the node feature mask based on the human skeleton spatiotemporal graph and the skeleton node spatiotemporal interaction matrix, and obtaining a final node feature mask as a mask for characterizing the decision factors of the spatiotemporal graph neural network, includes:
[0024] Step S41: Construct a mask inference module, which randomly generates node feature masks M based on the human skeleton spatiotemporal graph. N ;
[0025] Step S42: Based on the initial node feature mask M N and the skeleton node spatiotemporal interaction matrix M H , get the node feature mask M based on spatiotemporal interaction HN ,in,
[0026] Step S43: Construct the loss function, initialize the current number of iterations k, set k = 1; set the number of iterations num;
[0027] Step S44: Optimize the node feature mask M based on spatiotemporal interaction based on Adam optimizer HN Optimize and obtain the current node feature mask M based on spatiotemporal interaction HN ;
[0028] Based on the current spatiotemporal interaction-based node feature mask M HN Calculate the loss function value; assign k to k+1;
[0029] Step S45: If the number of iterations is greater than num or the accuracy of the spatiotemporal graph neural network to be interpreted has not been improved in five consecutive iterations, the current spatiotemporal interaction-based node feature mask M is HN As the final node feature mask, the final node feature mask is used as the mask to characterize the decision factors of the spatiotemporal graph neural network; otherwise, go to step S44.
[0030] Preferably, the loss function is:
[0031]
[0032] Among them, Loss is the loss function, is the structural loss function, is the sparse loss function, M H The element in row p and column q in MN The element in the pth row and qth column, λ is the weight, set to 0.01, and n is M HN The number of elements in M HN The element at row p and column q in .
[0033] According to a second aspect of the present invention, there is provided an apparatus for generating a trusted computation representing a decision factor of a spatiotemporal graph neural network, the apparatus comprising:
[0034] Initialization module: configured to obtain the recognition results of the to-be-explained spatiotemporal graph neural network for skeletal behavior recognition and the input of the to-be-explained spatiotemporal graph neural network corresponding to the recognition results, where the input is a video;
[0035] The subset module is configured to perform Monte Carlo sampling on the spatiotemporal graph of the human skeleton corresponding to the video to obtain several skeleton node subsets, and generate a spatiotemporal subset union of each skeleton node based on the skeleton node subsets; a skeleton node is the entire skeleton of a pedestrian in the video frame; wherein, the skeleton nodes in the skeleton node subset are subjected to second-order sampling based on the spatiotemporal physical connectivity of the skeleton nodes in the human skeleton spatiotemporal graph to generate a spatiotemporal subset union of each skeleton node;
[0036] Spatiotemporal interaction module: configured to determine the spatiotemporal interaction value and marginal effect value of each skeleton node based on the spatiotemporal graph of the human skeleton and the spatiotemporal subset alliance of each skeleton node; based on the spatiotemporal interaction value and marginal effect value of each skeleton node, construct the spatiotemporal interaction matrix of the skeleton nodes;
[0037] Mask generation module: configured to randomly generate node feature masks based on the human skeleton spatiotemporal graph, optimize the node feature masks based on the human skeleton spatiotemporal graph and the skeleton node spatiotemporal interaction matrix, and obtain the final node feature mask as the mask to characterize the decision factors of the spatiotemporal graph neural network.
[0038] According to a third aspect of the present invention, there is provided an electronic device, comprising:
[0039] A processor, which is used to execute multiple instructions;
[0040] A memory for storing a plurality of instructions;
[0041] The plurality of instructions are used to be stored by the memory and loaded and executed by the processor to implement the method as described above.
[0042] According to a fourth aspect of the present invention, a computer-readable storage medium is provided, wherein a plurality of instructions are stored in the storage medium; the plurality of instructions are used for a processor to load and execute the aforementioned method.
[0043] This paper utilizes the Monte Carlo algorithm and the second-order spatiotemporal physical relationships in the human skeleton space-time graph to construct a space-time subset alliance of the human skeleton space-time graph. Based on the interaction theory in game theory, the spatiotemporal interaction value of skeleton nodes and their neighbors in the human skeleton space-time graph is quantified, and the impact of different space-time subset alliances on the whole is quantified. A masked inference module is constructed. Based on the principle of interactive guidance, the masked inference module iteratively optimizes the mask representing the decision factors of the space-time graph neural network, and ultimately generates a mask representing the decision factors of the space-time graph neural network.
[0044] The above description is only an overview of the technical solution of the present invention. In order to more clearly understand the technical means of the present invention and implement it according to the contents of the specification, the following is a detailed description of the preferred embodiments of the present invention with reference to the accompanying drawings. BRIEF DESCRIPTION OF THE DRAWINGS
[0045] The accompanying drawings, which constitute part of the present invention, are used to provide a further understanding of the present invention. The present invention is described with the following accompanying drawings. In the accompanying drawings:
[0046] Figure 1 A flowchart of a method for generating a trusted computation mask representing decision factors of a spatiotemporal graph neural network according to one embodiment of the present invention. DETAILED DESCRIPTION
[0047] First combine Figure 1 The following describes a method for generating a trustworthy computation representing a decision factor of a spatiotemporal graph neural network according to an embodiment of the present invention. Figure 1 As shown, the method includes the following steps:
[0048] Step S1: Obtain the recognition result of the to-be-explained spatiotemporal graph neural network for skeleton behavior recognition and the input of the to-be-explained spatiotemporal graph neural network corresponding to the recognition result, where the input is a video;
[0049] Step S2: Monte Carlo sampling is performed on the spatiotemporal graph of the human skeleton corresponding to the video to obtain several skeleton node subsets, and a spatiotemporal subset union of each skeleton node is generated based on the skeleton node subsets; a skeleton node is the entire skeleton of a pedestrian in the video frame; wherein, based on the spatiotemporal physical connectivity of the skeleton nodes in the spatiotemporal graph of the human skeleton, the skeleton nodes in the skeleton node subset are subjected to second-order sampling to generate a spatiotemporal subset union of each skeleton node;
[0050] Step S3: determining the spatiotemporal interaction value and marginal effect value of each skeleton node based on the spatiotemporal graph of the human skeleton and the spatiotemporal subset union of each skeleton node; constructing the spatiotemporal interaction matrix of the skeleton nodes based on the spatiotemporal interaction value and marginal effect value of each skeleton node;
[0051] Step S4: Randomly generate node feature masks based on the human skeleton spatiotemporal graph, optimize the node feature masks based on the human skeleton spatiotemporal graph and the skeleton node spatiotemporal interaction matrix, and obtain the final node feature mask as the mask to characterize the decision factors of the spatiotemporal graph neural network.
[0052] Furthermore, in step S2, Monte Carlo sampling is performed on the spatiotemporal graph of the human skeleton corresponding to the video to obtain several skeleton node subsets, and a spatiotemporal subset union of each skeleton node is generated based on the skeleton node subsets; the skeleton node is the entire skeleton of a pedestrian in the video frame; wherein, based on the spatiotemporal physical connectivity of the skeleton nodes in the spatiotemporal graph of the human skeleton, the skeleton nodes in the skeleton node subset are subjected to second-order sampling to generate the spatiotemporal subset union of each skeleton node, including:
[0053] Step S21: Generate a human skeleton spatiotemporal graph G = {V, E} corresponding to the video, where the skeleton node set The edge set E=E of the human skeleton space-time graph S ∩E F , T is the number of video frames, is the i-th skeleton node in the t-th frame, t is the frame number of the video, N is the total number of skeleton nodes, i, j are the skeleton node numbers; E S is a set of connected edges, where the connected edges are the edges that connect the skeleton nodes with natural connection relationships in the same frame of the video. is the jth skeleton node in the tth frame, H is the index set of skeleton nodes in several human skeletons in one frame; E F is the time edge set, Temporal edges are edges connecting the same skeleton nodes in consecutive frames. is the i-th skeleton node in the t+1th frame;
[0054] Step S22: Monte Carlo sampling is performed on the human skeleton spatiotemporal graph to obtain several skeleton node subsets, and a spatiotemporal subset union of each skeleton node is generated based on the skeleton node subsets; the skeleton node is the entire skeleton of a pedestrian in the video frame; wherein, second-order sampling is performed on the skeleton nodes in the skeleton node subset based on the spatiotemporal physical connectivity of the skeleton nodes in the human skeleton spatiotemporal graph to generate a spatiotemporal subset union of each skeleton node.
[0055] In the present invention, second-order spatiotemporal relationship sampling is performed on each skeleton node in the skeleton node subset, and each node generates K second-order spatiotemporal subsets. Therefore, there will be multiple spatiotemporal subset alliances.
[0056] The step S3 determines the spatiotemporal interaction value and marginal effect value of each skeleton node based on the spatiotemporal graph of the human skeleton and the spatiotemporal subset alliance of each skeleton node; constructs a spatiotemporal interaction matrix of skeleton nodes based on the spatiotemporal interaction value and marginal effect value of each skeleton node, including:
[0057] Step S31: Determine the spatiotemporal interaction value of each skeleton node based on the spatiotemporal graph of the human skeleton and the spatiotemporal subset alliance of each skeleton node. The spatiotemporal interaction value is calculated as follows:
[0058]
[0059] in, is the spatiotemporal interaction value of the skeleton node, For skeleton nodes The second-order spacetime subset union of For skeleton nodes The kth second-order spacetime subset of For skeleton nodes Add to The subsequent profit function, for The profit function, for size;
[0060] Step S32: Determine the marginal effect value of each skeleton node based on the spatiotemporal interaction value of each skeleton node. The calculation formula of the marginal effect value of the skeleton node is:
[0061] and
[0062] in, express A collection of mid-skeleton node features, express The skeleton node feature is added to The collection of skeleton node features after express The prediction results input into the spatiotemporal graph network model, express The prediction results input into the spatiotemporal graph network model;
[0063] Step S33: Based on the spatiotemporal interaction value and marginal effect value of each skeleton node, construct the spatiotemporal interaction matrix M of the skeleton nodes H :
[0064]
[0065] Among them, F(V) represents the set of skeleton node features in the skeleton node set V, and Φ(F(V)) represents the prediction result of F(V) input into the spatiotemporal graph network model.
[0066] Step S4: randomly generate node feature masks based on the human skeleton spatiotemporal graph, optimize the node feature masks based on the human skeleton spatiotemporal graph and the skeleton node spatiotemporal interaction matrix, and obtain the final node feature mask as the mask representing the decision factors of the spatiotemporal graph neural network, including:
[0067] Step S41: Construct a mask inference module, which randomly generates node feature masks M based on the human skeleton spatiotemporal graph. N ;
[0068] Step S42: Based on the initial node feature mask M N and the skeleton node spatiotemporal interaction matrix M H , get the node feature mask M based on spatiotemporal interaction HN ,in,
[0069] Step S43: construct the loss function, initialize the current number of iterations k, set k = 1; set the number of iterations num;
[0070] Step S44: Optimize the node feature mask M based on spatiotemporal interaction based on Adam optimizer HN Optimize and obtain the current node feature mask M based on spatiotemporal interaction HN ;
[0071] Based on the current spatiotemporal interaction-based node feature mask M HN Calculate the loss function value; assign k to k+1;
[0072] Step S45: If the number of iterations is greater than num or the accuracy of the spatiotemporal graph neural network to be interpreted has not improved in five consecutive iterations, the current spatiotemporal interaction-based node feature mask M is HN As the final node feature mask, the final node feature mask is used as the mask to characterize the decision factors of the spatiotemporal graph neural network; otherwise, go to step S44.
[0073] Furthermore, the loss function is:
[0074]
[0075] Among them, Loss is the loss function, is the structural loss function, is the sparse loss function, M H The element in row p and column q in MN The element in the pth row and qth column, λ is the weight, set to 0.01, and n is M HN The number of elements in M HN The element at row p and column q in .
[0076] In the present invention, the mask M is updated by back propagation HN , in order to gradually obtain the key spatiotemporal interaction relationships corresponding to the target category. Finally, an early stopping mechanism is introduced to avoid overfitting. When the spatiotemporal graph neural network model Φ does not show significant improvement in multiple consecutive iterations, the iteration process is terminated. Finally, the mask M HN Used to characterize decision factors of spatiotemporal graph neural networks.
[0077] The present invention adopts the Monte Carlo method to randomly sample the human skeleton spatiotemporal graph, and randomly sorts the combinations of all nodes to obtain Q node subsets, which are expressed as
[0078] For each node in the node subset, according to its spatiotemporal physical connection relationship E in the human skeleton spatiotemporal graph S and E F Construct a spatiotemporal subset alliance for each node. The spatiotemporal subset alliance captures the spatiotemporal interaction relationship between nodes by considering the second-order adjacency structure of the nodes. Perform second-order sampling on each node in the node subset according to the spatiotemporal physical connection relationship to obtain the spatiotemporal subset of each node. node The second-order neighbor set N2(v i ) is defined as follows.
[0079]
[0080] in, Represents the edge set E at time S and the spatial edge set E F midpoint The first-order neighbor set of Therefore, the nodes Get the second-order space-time subset alliance according to its neighbor relationship It includes K second-order space-time subsets, which are defined as follows.
[0081]
[0082] in, and satisfy This not only takes into account the natural connection between temporal edge sets and human joints in the human skeleton spatiotemporal graph, but also accurately simulates the natural patterns of human movements, effectively capturing the continuity of movements between different time frames and the spatiotemporal interactions between nodes.
[0083] There is not only a spatial topological relationship between the nodes in the human skeleton spatiotemporal graph, but also a dynamic interactive relationship in the time dimension. This complex spatiotemporal interactivity makes it difficult for traditional graph neural network trustworthy computing methods to fully capture the degree of dependence of the model on the interaction structure of the nodes and their neighborhoods. In order to solve this problem, the present invention can characterize the spatiotemporal interaction value of the node in its neighborhood, thereby revealing the spatiotemporal dependency of the node and more clearly explaining its key role in the model prediction results. The mask generated by the present invention has a good explanatory effect on the ST-GCN network model and the HD-GCN network model.
[0084] set up is the set of all nodes in the human skeleton spatiotemporal graph of the human skeleton, where each node There are corresponding node features at different time frames t in the human skeleton spatiotemporal graph. For nodes All possible unions of spacetime subsets of Representing a union of spacetime subsets The profit function of the node The spatiotemporal interaction value in all possible spatiotemporal subset alliances, that is, the spatiotemporal interaction value of the skeleton nodes is shown in the following formula.
[0085]
[0086] in, Spacetime Subset Alliance The size of Representation node Add to Spacetime Subset Alliance The subsequent profit function, Representation node Alliance of spacetime subsets S k The marginal effect value of the node Join the Spacetime Subset Alliance The change in the time-benefit function is the gain of the node to the overall prediction result. It is the node It is determined by the three-dimensional characteristics and topological structure of the image.
[0087] Let Φ(·) be the prediction result of the spatiotemporal graph network model based on behavior recognition, and the node The marginal effect value in the space-time graph network is defined as follows.
[0088]
[0089] in, Representing a union of spacetime subsets The collection of midpoint features, Representation node The node feature is added to The set of node features after Representing a union of spacetime subsets The prediction results input into the spatiotemporal graph network model, and the benefit function depends on the three-dimensional feature information of the node and its spatiotemporal interaction relationship, can fully reflect the role of the node in the human skeleton spatiotemporal graph and its effect on the model prediction results.
[0090] Then, construct the skeleton node spatiotemporal interaction matrix M H .set up is the set of all nodes in the human skeleton spatiotemporal graph, and the spatiotemporal interaction value M of all nodes in all spatiotemporal subsets H , M H Expressed in matrix form, it is defined as follows.
[0091]
[0092] For all nodes in all spatiotemporal subsets, the spatiotemporal interaction value M H , its profit function Dependence on spacetime subset alliances However, the core of spatiotemporal graph networks lies in the spatiotemporal interaction between nodes, and the spatiotemporal subset alliance In fact, the local node features in the human skeleton spatiotemporal graph can only reflect the information of some nodes, thus ignoring the influence of the global structure of the human skeleton spatiotemporal graph. Therefore, the profit function Converting to Φ(F(V)) more comprehensively captures the relationship between nodes and the entire human skeleton spatiotemporal graph, thereby more accurately quantifying the marginal effect of nodes in the spatiotemporal graph network. By uniformly using the features F(V) of all nodes in the human skeleton spatiotemporal graph, bias introduced by local subsets is avoided, ensuring global consistency in the calculation of node marginal effects. In summary, the above equation can be rewritten as the following.
[0093]
[0094] The mask reasoning module uses spatiotemporal interaction values to guide the importance evaluation process of node features, and infers the interaction-guided masks, helping the model focus on key nodes related to the target category and their spatiotemporal interaction features, thereby effectively improving the model's ability to identify important interaction relationships in complex spatiotemporal sequences.
[0095] Randomly generate node feature mask M according to the input human skeleton spatiotemporal graph N ∈(0,1), using the normalized spatiotemporal interaction value M H and node feature mask M N Combined with the optimization reasoning process of guiding the node feature mask, the initial node feature mask M based on interactive guidance HN The definition is shown below.
[0096]
[0097] Among them, M HN The value of reflects the importance of each node feature in the model's prediction results. To avoid imbalanced node feature weights during subsequent inference, the interactively guided mask is normalized. That is, the contribution values of all interactively guided node feature masks are scaled to within (0, 1). This ensures stability during mask inference.
[0098] Then, the optimizer is used to optimize the node feature mask M based on the interaction guidance. HN Perform iterative optimization and update the mask through back propagation. In each iteration, the mask M HN Applied to the spatiotemporal graph network model, perturbs the model and generates prediction results. During the inference process, the weight of the loss term is dynamically adjusted according to the node interaction relationship to ensure that the mask M HN It can accurately capture important spatiotemporal interactions that play a key role in model decision-making.
[0099] In addition, in order to prevent the model from overfitting during the optimization process, an early stopping mechanism is introduced. When the model prediction result Φ(F(V); M HN ) When the performance does not improve in 5 consecutive iterations, the iteration process automatically terminates. Finally, the optimal mask M is obtained. HN , in order to more accurately capture and reflect the spatiotemporal interaction relationship related to the target category, and thus achieve an effective interpretation of the spatiotemporal graph network model.
[0100] The loss function includes a structural loss term Sparse loss term This loss function constrains the mask optimization process so that the mask can capture important interactive features that conform to the original human skeleton spatiotemporal graph structure. The spatiotemporal structure consistency loss function is defined as follows.
[0101]
[0102] in, is the node feature mask M based on interaction guidance HN The behavior recognition network model after perturbation Φ(F(V);M HN) is the weight value of the edge.
[0103] Structural loss term Used to quantize the current mask M HN The global spatiotemporal structure ensures the balance of node interaction feature selection during the optimization process, and the structural loss term The expression is shown below.
[0104]
[0105] Sparse loss term It can make the generated mask as sparse as possible, encouraging the recognition of fewer and more important node features rather than all node features. The sparse loss term The expression is shown below.
[0106]
[0107] The final total loss function is obtained by weighted combination of the above two loss terms to guide the model to strike a balance between sparsity, structural information preservation and mask size when generating masks.
[0108] The present invention provides a device for generating a trusted computation representing decision factors of a spatiotemporal graph neural network, the device comprising:
[0109] Initialization module: configured to obtain the recognition results of the to-be-explained spatiotemporal graph neural network for skeletal behavior recognition and the input of the to-be-explained spatiotemporal graph neural network corresponding to the recognition results, where the input is a video;
[0110] The subset module is configured to perform Monte Carlo sampling on the spatiotemporal graph of the human skeleton corresponding to the video to obtain several skeleton node subsets, and generate a spatiotemporal subset union of each skeleton node based on the skeleton node subsets; a skeleton node is the entire skeleton of a pedestrian in the video frame; wherein, the skeleton nodes in the skeleton node subset are subjected to second-order sampling based on the spatiotemporal physical connectivity of the skeleton nodes in the human skeleton spatiotemporal graph to generate a spatiotemporal subset union of each skeleton node;
[0111] Spatiotemporal interaction module: configured to determine the spatiotemporal interaction value and marginal effect value of each skeleton node based on the spatiotemporal graph of the human skeleton and the spatiotemporal subset alliance of each skeleton node; based on the spatiotemporal interaction value and marginal effect value of each skeleton node, construct the spatiotemporal interaction matrix of the skeleton nodes;
[0112] Mask generation module: configured to randomly generate node feature masks based on the human skeleton spatiotemporal graph, optimize the node feature masks based on the human skeleton spatiotemporal graph and the skeleton node spatiotemporal interaction matrix, and obtain the final node feature mask as the mask to characterize the decision factors of the spatiotemporal graph neural network.
[0113] An embodiment of the present invention further provides an electronic device, including:
[0114] A processor, which is used to execute multiple instructions;
[0115] A memory for storing a plurality of instructions;
[0116] The plurality of instructions are used to be stored by the memory and loaded and executed by the processor to implement the method as described above.
[0117] An embodiment of the present invention further provides a computer-readable storage medium, wherein a plurality of instructions are stored in the storage medium; the plurality of instructions are used for a processor to load and execute the method described above.
[0118] It should be noted that, in the absence of conflict, the embodiments of the present invention and the features in the embodiments may be combined with each other.
[0119] In the several embodiments provided by the present invention, it should be understood that the disclosed systems, devices and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of the units is merely a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. In addition, the mutual coupling or direct coupling or communication connection shown or discussed may be an indirect coupling or communication connection through some interface, device or unit, which may be electrical, mechanical or other forms.
[0120] The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of these units may be selected to achieve the purpose of this embodiment according to actual needs.
[0121] In addition, the functional units in various embodiments of the present invention may be integrated into a single processing unit, each unit may exist physically separately, or two or more units may be integrated into a single unit. The aforementioned integrated units may be implemented in the form of hardware or hardware plus software functional units.
[0122] The above-mentioned integrated unit implemented in the form of a software functional unit can be stored in a computer-readable storage medium. The above-mentioned software functional unit is stored in a storage medium and includes a number of instructions for causing a computer device (which can be a personal computer, a physical server, or a network cloud server, etc., and requires the Ubuntu operating system to be installed) to perform some of the steps of the method described in various embodiments of the present invention. The aforementioned storage medium includes: a USB flash drive, a mobile hard drive, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disk, etc., various media that can store program code.
[0123] The above description is merely a preferred embodiment of the present invention and does not constitute any form of limitation to the present invention. Any simple modifications, equivalent changes and modifications made to the above embodiment based on the technical essence of the present invention still fall within the scope of the technical solution of the present invention.
Claims
1. A method for generating trusted computations representing decision factors of spatiotemporal graph neural networks, characterized in that: Methods include: Step S1: Obtain the recognition result of the to-be-explained spatiotemporal graph neural network for skeleton behavior recognition and the input of the to-be-explained spatiotemporal graph neural network corresponding to the recognition result, where the input is a video; Step S2: Monte Carlo sampling is performed on the spatiotemporal graph of the human skeleton corresponding to the video to obtain several skeleton node subsets, and a spatiotemporal subset union of each skeleton node is generated based on the skeleton node subsets; a skeleton node is the entire skeleton of a pedestrian in the video frame; wherein, based on the spatiotemporal physical connectivity of the skeleton nodes in the spatiotemporal graph of the human skeleton, the skeleton nodes in the skeleton node subset are subjected to second-order sampling to generate a spatiotemporal subset union of each skeleton node; Step S3: determining the spatiotemporal interaction value and marginal effect value of each skeleton node based on the spatiotemporal graph of the human skeleton and the spatiotemporal subset union of each skeleton node; constructing the spatiotemporal interaction matrix of the skeleton nodes based on the spatiotemporal interaction value and marginal effect value of each skeleton node; Step S4: Randomly generate node feature masks based on the human skeleton spatiotemporal graph, optimize the node feature masks based on the human skeleton spatiotemporal graph and the skeleton node spatiotemporal interaction matrix, and obtain the final node feature mask as the mask to characterize the decision factors of the spatiotemporal graph neural network.
2. The method according to claim 1, wherein Step S2, performing Monte Carlo sampling on the spatiotemporal graph of the human skeleton corresponding to the video to obtain a plurality of skeleton node subsets, and generating a spatiotemporal subset union of each skeleton node based on the skeleton node subsets; a skeleton node is the entire skeleton of a pedestrian in the video frame; wherein, based on the spatiotemporal physical connectivity of the skeleton nodes in the spatiotemporal graph of the human skeleton, second-order sampling is performed on the skeleton nodes in the skeleton node subset to generate a spatiotemporal subset union of each skeleton node, including: Step S21: Generate a human skeleton spatiotemporal graph G = {V, E} corresponding to the video, where the skeleton node set The edge set E=E of the human skeleton space-time graph S ∩E F , T is the number of video frames, is the i-th skeleton node of the t-th frame, t is the frame number of the video, N is the total number of skeleton nodes, i, j are the skeleton node numbers; E S is a set of connected edges, where the connected edges are the edges that connect the skeleton nodes with natural connection relationships in the same frame of the video. is the jth skeleton node in the tth frame, H is the index set of skeleton nodes in several human skeletons in one frame; E F is the time edge set, Temporal edges are edges connecting the same skeleton nodes in consecutive frames. is the i-th skeleton node of the t+1th frame; Step S22: Monte Carlo sampling is performed on the human skeleton spatiotemporal graph to obtain several skeleton node subsets, and a spatiotemporal subset union of each skeleton node is generated based on the skeleton node subsets; the skeleton node is the entire skeleton of a pedestrian in the video frame; wherein, second-order sampling is performed on the skeleton nodes in the skeleton node subset based on the spatiotemporal physical connectivity of the skeleton nodes in the human skeleton spatiotemporal graph to generate a spatiotemporal subset union of each skeleton node.
3. The method according to claim 2, wherein The step S3 is to determine the spatiotemporal interaction value and marginal effect value of each skeleton node based on the spatiotemporal graph of the human skeleton and the spatiotemporal subset alliance of each skeleton node; Based on the spatiotemporal interaction value and marginal effect value of each skeleton node, the spatiotemporal interaction matrix of the skeleton nodes is constructed, including: Step S31: Determine the spatiotemporal interaction value of each skeleton node based on the spatiotemporal graph of the human skeleton and the spatiotemporal subset alliance of each skeleton node. The spatiotemporal interaction value is calculated as follows: in, is the spatiotemporal interaction value of the skeleton node, For skeleton nodes The second-order spacetime subset union of For skeleton nodes The kth second-order spacetime subset of For skeleton nodes Add to The subsequent profit function, for The profit function, for size; Step S32: Determine the marginal effect value of each skeleton node based on the spatiotemporal interaction value of each skeleton node. The calculation formula of the marginal effect value of the skeleton node is: and in, express A collection of mid-skeleton node features, express The skeleton node feature is added to The collection of skeleton node features after express The prediction results input into the spatiotemporal graph network model, express The prediction results input into the spatiotemporal graph network model; Step S33: Based on the spatiotemporal interaction value and marginal effect value of each skeleton node, construct the spatiotemporal interaction matrix M of the skeleton nodes H : Among them, F(V) represents the set of skeleton node features in the skeleton node set V, and Φ(F(V)) represents the prediction result of F(V) input into the spatiotemporal graph network model.
4. The method according to claim 3, wherein Step S4: randomly generate node feature masks based on the human skeleton spatiotemporal graph, optimize the node feature masks based on the human skeleton spatiotemporal graph and the skeleton node spatiotemporal interaction matrix, and obtain the final node feature mask as the mask representing the decision factors of the spatiotemporal graph neural network, including: Step S41: Construct a mask inference module, which randomly generates node feature masks M based on the human skeleton spatiotemporal graph. N ; Step S42: Based on the initial node feature mask M N and the skeleton node spatiotemporal interaction matrix M H , get the node feature mask M based on spatiotemporal interaction HN ,in, Step S43: construct the loss function, initialize the current number of iterations k, set k = 1; set the number of iterations num; Step S44: Optimize the node feature mask M based on spatiotemporal interaction based on Adam optimizer HN Optimize and obtain the current node feature mask M based on spatiotemporal interaction HN ; Based on the current spatiotemporal interaction-based node feature mask M HN Calculate the loss function value; assign k to k+1; Step S45: If the number of iterations is greater than num or the accuracy of the spatiotemporal graph neural network to be interpreted has not been improved in five consecutive iterations, the current spatiotemporal interaction-based node feature mask M is HN As the final node feature mask, the final node feature mask is used as the mask to characterize the decision factors of the spatiotemporal graph neural network; otherwise, go to step S44.
5. The method according to claim 4, wherein The loss function is: Among them, Loss is the loss function, is the structural loss function, is the sparse loss function, M H The element in row p and column q in M N The element in the pth row and qth column, λ is the weight, set to 0.01, and n is M HN The number of elements in M HN The element at row p and column q in .
6. A device for generating a trusted computation representing decision factors of a spatiotemporal graph neural network, characterized in that: The device includes: Initialization module: configured to obtain the recognition results of the to-be-explained spatiotemporal graph neural network for skeletal behavior recognition and the input of the to-be-explained spatiotemporal graph neural network corresponding to the recognition results, where the input is a video; The subset module is configured to perform Monte Carlo sampling on the spatiotemporal graph of the human skeleton corresponding to the video to obtain several skeleton node subsets, and generate a spatiotemporal subset union of each skeleton node based on the skeleton node subsets; a skeleton node is the entire skeleton of a pedestrian in the video frame; wherein, the skeleton nodes in the skeleton node subset are subjected to second-order sampling based on the spatiotemporal physical connectivity of the skeleton nodes in the human skeleton spatiotemporal graph to generate a spatiotemporal subset union of each skeleton node; Spatiotemporal interaction module: configured to determine the spatiotemporal interaction value and marginal effect value of each skeleton node based on the spatiotemporal graph of the human skeleton and the spatiotemporal subset alliance of each skeleton node; based on the spatiotemporal interaction value and marginal effect value of each skeleton node, construct the spatiotemporal interaction matrix of the skeleton nodes; Mask generation module: configured to randomly generate node feature masks based on the human skeleton spatiotemporal graph, optimize the node feature masks based on the human skeleton spatiotemporal graph and the skeleton node spatiotemporal interaction matrix, and obtain the final node feature mask as the mask to characterize the decision factors of the spatiotemporal graph neural network.
7. An electronic device comprising: A processor, which is used to execute multiple instructions; A memory for storing a plurality of instructions; The plurality of instructions are used to be stored in the memory and loaded and executed by the processor according to any one of claims 1 to 5.
8. A computer-readable storage medium, wherein a plurality of instructions are stored in the storage medium; the plurality of instructions are used for a processor to load and execute the method according to any one of claims 1 to 5.