Interpretability Method, Visualization Method and Related Devices for Deep Neural Networks
By obtaining and processing the motion vectors and mask vectors of key points of the human body, calculating the motion sequence and iteratively updating the reference value of the optimization function, the problem of the difficulty of explaining deep neural networks in human behavior recognition is solved, and the recognition accuracy and efficiency are improved.
Patent Information
- Application Number
- CN202111454454.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-12-01
- Publication Date
- 2025-05-30
- Estimated Expiration
- 2041-12-01
AI Technical Summary
Existing deep neural networks are difficult to explain the failure and effectiveness of model in human behavior recognition, and cannot accurately explain the importance of model input features, resulting in poor interpretation effects.
By obtaining the motion vector and mask vector of the human body's key points, compute the motion sequence, and input it into the pre-trained behavior recognition model, iteratively updates the reference value of the optimization function to obtain the target mask vector, and then explains the behavior recognition model.
The recognition accuracy and recognition efficiency of the behavior recognition model are improved, and the accuracy of the recognition results is targeted by explaining the impact of the difference information on the model judgment results.
Smart Images

Figure CN114419726B_ABST
Abstract
Description
Technical Field
[0001] Embodiments of the present application relate to the field of artificial intelligence technology, and in particular, to a method for interpreting a deep neural network, a visualization method, and related devices. Background Art
[0002] With the development of computer vision, the use of computer vision to recognize human behaviors has been widely applied in many fields, such as intelligent video surveillance, public security, behavior analysis, human-computer interaction, and intelligent robots. In related technologies, highly accurate depth sensors are mostly used to obtain human pose information, and human pose estimation algorithms are used to extract the positions of human skeleton joint points. The joint point coordinate positions of the entire skeleton sequence are structured into coordinate vectors, and then a behavior recognition model is constructed based on deep learning, such as Recurrent Neural Networks (RNNs), Long-Short Term Memory (LSTM), Convolutional Neural Networks (CNNs), etc., so as to learn the spatial and temporal information of the human skeleton sequence and perform behavior recognition.
[0003] However, the models used in these neural network recognition methods are all black-box models, and it is difficult to accurately explain when the model will fail and be effective. In related technologies, methods such as using Laplacian matrices are used to estimate the importance of model input features. However, this method can only roughly explain whether a certain model input feature is important to the model, and cannot more accurately explain the model, resulting in poor model interpretation effects. In the behavior recognition method, the changes in the motion states of different skeleton nodes at different times have differences, that is, different skeleton nodes have different weights in different motion states. This difference information affects the judgment result of the behavior recognition model. If this difference information is not considered, the accuracy of the behavior recognition result cannot be improved targeted. Summary of the Invention
[0004] The following is an overview of the subject matter described in detail in this article. This overview is not intended to limit the scope of protection of the claims.
[0005] Embodiments of the present application provide a method for interpreting a deep neural network, a visualization method, and related devices, which can use a target mask vector to explain the influence of difference information on the judgment result of a behavior recognition model, thereby improving the recognition accuracy of the behavior recognition model targeted.
[0006] In a first aspect, embodiments of the present application provide a method for interpreting a deep neural network, including:
[0007] Obtain the motion vector and mask vector of human key points;
[0008] Calculate a motion sequence based on the mask vector and the motion vector;
[0009] Input the motion sequence into a pre-trained behavior recognition model to obtain a predicted behavior recognition result, where the behavior recognition model is a deep neural network model;
[0010] Use the predicted behavior recognition result to iteratively update the reference value of the optimization function to obtain a target mask vector for characterizing the weight of the motion vector.
[0011] In an optional implementation, the calculating a motion sequence based on the mask vector and the motion vector includes:
[0012] Calculate the motion data information of the motion vectors and the corresponding mask vectors at all times before the acquisition time;
[0013] Accumulate the motion data information to obtain the motion sequence.
[0014] In an optional implementation, the calculating the motion data information of the motion vectors and the corresponding mask vectors at all times before the acquisition time includes:
[0015] Calculate the Hadamard product between the motion vector and the corresponding mask vector to obtain the motion data information;
[0016] It is expressed as:
[0017]
[0018] Where, represents the Hadamard product, Z C,t,V,M represents the motion vector at time t, G C,t,V,M represents the mask vector at time t, X′ C,t,V,M represents the motion data information at time t, C represents the motion three-dimensional coordinate information, V represents the number of joint points of each person, and M represents the number of people in the motion sequence.
[0019] In an optional implementation, the using the predicted behavior recognition result to iteratively update the reference value of the optimization function to obtain a target mask vector for characterizing the weight of the motion vector includes:
[0020] Calculate the reference value of the optimization function according to the predicted behavior recognition result and the mask vector;
[0021] Adjust the mask vector according to the reference value until the optimization function reaches the convergence condition to obtain the target mask vector.
[0022] In an optional implementation manner, adjusting the mask vector according to the reference value until the optimization function reaches the convergence condition to obtain the target mask vector includes:
[0023] Adjust the mask vector to obtain an optimized mask vector;
[0024] Calculate an optimized motion sequence according to the optimized mask vector and the motion vector;
[0025] Input the optimized motion sequence into a pre-trained behavior recognition model to obtain an optimized predicted behavior recognition result;
[0026] Calculate the reference value of the optimization function corresponding to the optimized predicted behavior recognition result;
[0027] Judge whether the optimization function reaches the convergence condition according to the reference value. If the convergence condition is reached, use the optimized mask vector as the target mask vector; otherwise, repeat the above steps until the optimization function reaches the convergence condition.
[0028] In an optional implementation manner, the optimization function is a mutual information optimization function, expressed as:
[0029] MI(Y,(G,Z))=H(Y|Z)-H(Y|G,Z)
[0030] Where Y represents the predicted behavior recognition result, Z represents the motion vector, G represents the mask vector, H(·) represents the entropy function, and MI(Y,(G,Z)) represents the reference value.
[0031] In an optional implementation manner, the convergence condition of the optimization function is: maximizing the reference value of the optimization function.
[0032] In a second aspect, an embodiment of the present application provides a visualization method, including:
[0033] Obtain a motion sequence and a target mask vector, where the motion sequence and the target mask vector are calculated by using the deep neural network interpretable method according to any item in the first aspect;
[0034] Sample the motion sequence frame by frame to obtain a skeleton motion vector;
[0035] Visualize and display the weight change of the skeleton motion vector between different frames according to the target mask vector.
[0036] In a third aspect, an embodiment of the present application provides a deep neural network interpretable device, including:
[0037] A first acquisition module, configured to acquire a motion vector and a mask vector of human body key points;
[0038] A motion sequence calculation module, configured to calculate a motion sequence according to the mask vector and the motion vector;
[0039] A preliminary prediction module, configured to input the motion sequence into a pre-trained behavior recognition model to obtain a predicted behavior recognition result;
[0040] An iterative optimization module, configured to iteratively update a reference value of an optimization function by using the predicted behavior recognition result to obtain a target mask vector for interpreting the behavior recognition model, where the target mask vector represents the weight of the motion vector.
[0041] In a fourth aspect, an embodiment of the present application provides a visualization device, including:
[0042] A second acquisition module, configured to acquire a motion sequence and a target mask vector, where the motion sequence and the target mask vector are calculated by using the deep neural network interpretable method according to any item in the first aspect;
[0043] A motion vector sampling module, configured to sample a skeleton motion vector frame by frame according to the motion sequence;
[0044] A visualization module, configured to visually display the weight change of the skeleton motion vector between different frames according to the target mask vector.
[0045] In a fifth aspect, a computer device includes a processor and a memory;
[0046] The memory is used to store a program;
[0047] The processor is configured to execute the deep neural network interpretable method according to any item in the first aspect or the visualization method according to the second aspect according to the program.
[0048] In a sixth aspect, an embodiment of the present application provides a computer-readable storage medium, storing computer-executable instructions, where the computer-executable instructions are used to execute the deep neural network interpretable method according to any item in the first aspect or the visualization method according to the second aspect.
[0049] The first aspect of the embodiment of the present application provides a deep neural network interpretable method. Compared with the related art, the motion vector and mask vector of the key points of the human body are obtained, and then the motion sequence is calculated based on the mask vector and the motion vector, and then the motion sequence is input into the pre-trained behavior recognition model to obtain the predicted behavior recognition result, and finally the reference value of the optimization function is iteratively updated using the predicted behavior recognition result to obtain the target mask vector representing the weight of the motion vector, and then the behavior recognition model is interpreted using the target mask vector. In this embodiment, the target mask vector is obtained by iterative updating based on the mask vector, and the weight of the key points of the human body of the motion vector is represented by the target mask vector, taking into account the influence of the weights of different features in the motion vector on the judgment result of the behavior recognition model, thereby improving the accuracy of the behavior recognition result and the recognition efficiency.
[0050] It can be understood that the beneficial effects of the second to sixth aspects compared with the related art are the same as the beneficial effects of the first aspect compared with the related art. Please refer to the relevant description in the first aspect, and no further details will be given here. BRIEF DESCRIPTION OF THE DRAWINGS
[0051] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the drawings required for use in the embodiments or related technical descriptions will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative labor.
[0052] Figure 1 is a schematic diagram of an exemplary system architecture provided by an embodiment of the present application;
[0053] Figure 2 is a flowchart of a deep neural network interpretable method provided by an embodiment of the present application;
[0054] Figure 3 is a schematic diagram of joints of a human skeleton provided by an embodiment of the present application;
[0055] Figure 4 is another flow chart of a deep neural network interpretable method provided by an embodiment of the present application;
[0056] Figure 5 is another flow chart of a deep neural network interpretable method provided by an embodiment of the present application;
[0057] Figure 6 is another flow chart of a deep neural network interpretable method provided by an embodiment of the present application;
[0058] Figure 7It is a flowchart of a visualization method provided by an embodiment of the present application;
[0059] Figure 8 It is a visualization schematic diagram provided by an embodiment of the present application;
[0060] Figure 9 It is a structural block diagram of an interpretable device for a deep neural network provided by an embodiment of the present application;
[0061] Figure 10 It is a structural block diagram of a visualization device provided by an embodiment of the present application. Detailed implementation manners
[0062] In the following description, specific details such as specific system structures and technologies are presented for the purpose of illustration rather than limitation, so as to thoroughly understand the embodiments of the present application. However, those skilled in the art should clearly understand that the embodiments of the present application can also be implemented in other embodiments without these specific details. In other cases, detailed descriptions of well-known systems, devices, circuits, and methods are omitted to avoid unnecessary details from interfering with the description of the embodiments of the present application.
[0063] It should be noted that although the logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in a different order from that in the flowchart. Terms such as "first" and "second" in the specification, claims, and the above-mentioned drawings are used to distinguish similar objects and do not necessarily need to describe a specific order or sequence.
[0064] It should also be understood that references to "an embodiment" or "some embodiments" etc. described in the specification of the embodiments of the present application mean that specific features, structures, or characteristics described in connection with the embodiment are included in one or more embodiments of the present application. Thus, statements such as "in an embodiment", "in some embodiments", "in other some embodiments", "in still other embodiments" etc. appearing in different places in this specification do not necessarily all refer to the same embodiment, but mean "one or more but not all embodiments", unless otherwise specifically emphasized in other ways. Terms such as "comprising", "including", "having" and their variants all mean "including but not limited to", unless otherwise specifically emphasized in other ways.
[0065] With the development of computer vision, the use of computer vision to recognize human behaviors has been widely applied in many fields, such as intelligent video surveillance, public security, behavior analysis, human-computer interaction, and intelligent robots. In related technologies, highly accurate depth sensors are mostly used to obtain human pose information, and human pose estimation algorithms are used to extract the positions of human skeleton joint points. The joint point coordinate positions of the entire skeleton sequence are structured into coordinate vectors, and then a behavior recognition model is constructed based on deep learning, such as Recurrent Neural Networks (RNNs), Long-Short Term Memory (LSTM), Convolutional Neural Networks (CNNs), etc., so as to learn the spatial and temporal information of the human skeleton sequence for behavior recognition.
[0066] However, the models used in these neural network recognition methods are all black-box models, and it is difficult to accurately explain when the model will fail and be effective. In related technologies, methods such as using Laplacian matrices are used to estimate the importance of model input features. However, this method can only roughly explain whether a certain model input feature is important to the model and cannot provide a more accurate explanation of the model, resulting in a poor model explanation effect. In behavior recognition methods, the changes in the motion states of different skeleton nodes at different times have differences, that is, different skeleton nodes have different weights in different motion states. This difference information affects the judgment results of the behavior recognition model. If this difference information is not considered, the accuracy of the behavior recognition results cannot be improved targeted.
[0067] Therefore, the embodiments of the present application provide an interpretable method for deep neural networks, which obtains the motion vector and mask vector of human key points, then calculates the motion sequence according to the mask vector and the motion vector, and then inputs the motion sequence into a pre-trained behavior recognition model to obtain the predicted behavior recognition result. Finally, the predicted behavior recognition result is used to iteratively update the reference value of the optimization function to obtain the target mask vector, and then the target mask vector is used to interpret the behavior recognition model. In this embodiment, the target mask vector is obtained through iterative update according to the mask vector, and the target mask vector is used to characterize the weights of the human key points of the motion vector. Considering the influence of the weights of different features in the motion vector on the judgment results of the behavior recognition model, the accuracy and efficiency of behavior recognition are improved targeted.
[0068] The following further elaborates on the embodiments of the present application in conjunction with the accompanying drawings.
[0069] Figure 1 The figure shows a schematic diagram of an exemplary system architecture to which the technical solutions of the embodiments of the present application can be applied.
[0070] As shown Figure 1 in FIG. 1, the system architecture 100 may include a terminal device (such as Figure 1 one or more of the desktop computer 101, tablet computer 102, and portable computer 103 shown in FIG. 1, and of course it may also be other terminal devices with a display screen, etc.), a network 104, and a server 105. The network 104 is used to provide a medium for a communication link between the terminal device and the server 105. The network 104 may include various connection types, such as a wired communication link, a wireless communication link, etc.).
[0071] It should be understood that Figure 1 the numbers of the terminal devices, networks, and servers in FIG. 1 are only illustrative. According to the implementation requirements, there may be any number of terminal devices, networks, and servers. For example, the server 105 may be a server cluster composed of multiple servers, etc.).
[0072] In an embodiment of the present invention, a user may use the terminal device 101 (or it may also be the terminal device 102 or 103) to upload the motion vectors collected by the collection device to the server 105. It may be to obtain human body posture information using a highly accurate depth sensor, extract the positions of human skeleton joint points using a human body pose estimation algorithm, and structure the joint point coordinate positions of the entire skeleton sequence into coordinate vectors. After the server 105 obtains the motion vectors, it randomly generates a mask vector, then calculates a motion sequence based on the mask vector and the motion vectors, inputs the motion sequence into a pre-trained behavior recognition model to obtain a predicted behavior recognition result, uses the predicted behavior recognition result to iteratively update the reference value of the optimization function to obtain a target mask vector, and uses the target mask vector to explain the influence of the difference information on the judgment result of the behavior recognition model, so as to specifically improve the recognition accuracy of the behavior recognition model. Solve the problem in the related art that only a rough explanation can be made on whether a certain model input feature is important to the model, and the model cannot be more accurately explained, resulting in a poor model explanation effect. Considering the influence of the weights of different features in the motion vectors on the judgment result of the behavior recognition model, the accuracy and efficiency of the behavior recognition result are specifically improved.
[0073] It should be noted that the deep neural network interpretable method or behavior recognition method provided by the embodiments of the present invention is generally executed by the server 105. Correspondingly, the deep neural network interpretable device or behavior recognition device is generally set in the server 105. However, in other embodiments of the present invention, the terminal device may also have a similar function to the server, so as to execute the deep neural network interpretable method or visualization method provided by the embodiments of the present invention.
[0074] The system architecture and application scenarios described in the embodiments of this application are to more clearly illustrate the technical solutions of the embodiments of this application, and do not constitute a limitation on the technical solutions provided by the embodiments of this application. Those skilled in the art will know that with the evolution of the system architecture and the emergence of new application scenarios, the technical solutions provided by the embodiments of this application are equally applicable to similar technical problems. Those skilled in the art can understand that Figure 1 The system architecture shown in does not constitute a limitation on the embodiments of this application, and may include more or fewer components than shown in the figure, or combine certain components, or have different component arrangements.
[0075] Based on the above system architecture, various embodiments of the deep neural network interpretability method or visualization method of the embodiments of this application are proposed.
[0076] As Figure 2 shown, Figure 2 is a flowchart of a deep neural network interpretability method provided by an embodiment of this application, including but not limited to steps S110 to S140.
[0077] Step S110, obtain the motion vector and mask vector of human key points.
[0078] In one embodiment, highly accurate depth sensors can be used to extract human pose information from video frames, or the Kinect depth camera can be directly used to capture human pose information, or 3D pose estimation algorithms can be used to obtain human pose information from ordinary videos. After obtaining the human pose information, the coordinate data of human key points (for example, the human key points can be human skeleton joint points) are extracted, that is, the position data of human key points are obtained from the information of the coordinates changing with time, and the motion vector is calculated according to the position data of the entire skeleton sequence.
[0079] In one embodiment, 15 joint points of the human skeleton are selected as human key points, and the relative distances between the joint points and the three-dimensional coordinates of each joint point are calculated. Refer to Figure 3 , which is a schematic diagram of the human skeleton joint points in this embodiment. The 15 human key points 300 in the figure include: head, neck, left shoulder, right shoulder, left elbow, right elbow, left wrist, right wrist, waist, left hip, right hip, left knee, right knee, left ankle, and right ankle. It can be understood that the joint points in this embodiment are only for illustration, and do not mean that only the above 15 human key points 300 can be selected in this embodiment. More joint points can also be selected according to the calculation accuracy requirements and the computing power of the computing device, such as selecting 25 human key points, etc., which are not specifically limited here.
[0080] In one embodiment, the motion vector contains the motion information of more than one human body. If there are multiple human bodies, local feature regions of all human body images are segmented in sequence. When necessary, the local feature regions of all human bodies also need to be corrected for inclination to obtain the skeleton sequence corresponding to each human body, and then the motion vector is generated by using the skeleton sequence corresponding to each human body.
[0081] In one embodiment, since the importance of each joint point in the skeleton sequence corresponding to different motion behaviors of the human body is different. For example, in the "applauding" behavior, the motion of the hand joint points is more important (i.e., has a higher weight) than that of other joint points. Therefore, the joint points in each skeleton sequence cannot be processed with the same weight. In this embodiment, a mask vector is used to block some joint points, so as to achieve the effect that different joint points correspond to different importance levels. In this embodiment, a vector with the same dimension as the motion vector is randomly initialized and generated as the mask vector. For example, the mask vector can be a matrix of all 1s.
[0082] Step S120, calculate the motion sequence according to the mask vector and the motion vector.
[0083] In one embodiment, in the human body skeleton sequence, if the position of the current joint point changes, it will cause the cumulative position change of the current joint point at subsequent times. Therefore, in this embodiment, considering the cumulative position change factor, the motion sequence is obtained by performing an accumulation operation according to the mask vector and the motion vector.
[0084] In one embodiment, referring to Figure 4 , step S120 includes but is not limited to steps S121 to S122:
[0085] Step S121, calculate the motion data information of the motion vectors and the corresponding mask vectors at all times before the acquisition moment.
[0086] In one embodiment, the motion data information is obtained by calculating the Hadamard product between the motion vector and the corresponding mask vector.
[0087] Step S122, accumulate the motion data information to obtain the motion sequence.
[0088] In one embodiment, the motion vector is denoted as Z C,T,V,M ;
[0089] where C represents the number of features. For example, the three-dimensional motion coordinate information of each joint point included in the skeleton sequence. A three-dimensional coordinate system can be set, the coordinate origin is selected, and the coordinate information of each joint point included in the skeleton sequence is obtained, denoted as (x, y, z);
[0090] T represents the length of the time series, that is, the data information of this time series needs to be collected to form the motion vector;
[0091] V represents the number of joint points of each person. For example, V = 15 or V = 25;
[0092] M represents the number of people in the motion sequence. For example, M = 1 or M = 2.
[0093] Since in the human body skeleton sequence, if the position of the current joint point changes, it will cause the cumulative position change of the current joint point at subsequent times. Therefore, the motion sequence at the acquisition moment is related to the motion vectors at all times before the acquisition moment. The motion data information at a certain moment is expressed as:
[0094]
[0095] Among them, represents the Hadamard product, Z C,t,V,M represents the motion vector at time t, G C,t,V,M represents the mask vector at time t, X' C,t,V,M represents the motion data information at time t.
[0096] In one embodiment, the Hadamard product is a type of operation on matrices. If two matrices of the same order: The dimension is m×n, and A = {a ij}, B = {b ij}, then the Hadamard product of matrix A and matrix B is expressed as:
[0097]
[0098] From the above formula, it can be seen that the Hadamard product of matrix A and matrix B is obtained by multiplying the corresponding elements of matrix A and matrix B to get the values at the corresponding positions.
[0099] In one embodiment, if G C,k,V,M is a matrix of all 1s, then the motion data information is the same as the motion vector.
[0100] Therefore, in this embodiment, after accumulation, the obtained motion sequence X C,T,V,M , this motion sequence represents the cumulative position change of all motion vectors before time T. According to the principle of difference, the motion data information is expressed as:
[0101] X' C,t,V,M = X C,t,V,M - X C,t-1,V,M
[0102] Then the motion sequence X C,T,V,M is expressed as:
[0103] X C,t,V,M = X' C,t,V,M + X C,t-1,V,M
[0104] where \(t = 1, 2, 3, 4, \ldots, T - 1\), and further derivation is expressed as:
[0105]
[0106] In one embodiment, in order to improve the operation efficiency of matrix operations, the above-mentioned motion sequence \(X\) C,t,V,M is subjected to matrix transformation, and the specific transformation steps are as follows:
[0107]
[0108] is simplified to:
[0109]
[0110] If there is:
[0111]
[0112] Further reasoning gives:
[0113]
[0114]
[0115] Further expressed as:
[0116]
[0117] where \(U\) is an upper triangular matrix with non-zero elements being 1, that is:
[0118] Step S130, input the motion sequence into the pre-trained behavior recognition model to obtain the predicted behavior recognition result.
[0119] In one embodiment, after obtaining the motion sequence \(X\) in the above steps, the motion sequence \(X\) is input into the pre-trained behavior recognition model to obtain the predicted behavior recognition result \(Y\).
[0120] In one embodiment, the behavior recognition model can be a deep neural network model, such as a graph neural network model (denoted as \(P\)), which is pre-trained using a training data set composed of motion vectors generated from a large number of human skeleton sequences. The training objective is to perform behavior recognition, input the motion vectors and output the behavior recognition result. In this embodiment, the pre-trained model parameters are denoted as \(\varphi\). It can be understood that in this embodiment, the behavior recognition model can also be other types of models, as long as it can achieve behavior recognition based on the motion sequence, and the model structure is not limited here.
[0121] Step S140: Iteratively update the reference value of the optimization function using the predicted behavior recognition result to obtain the target mask vector.
[0122] In one embodiment, the target mask vector represents the weights of the human key points of the motion vector and is used to interpret the behavior recognition model. Refer to Figure 5 , Step S140 includes but is not limited to Steps S141 to S142:
[0123] Step S141: Calculate the reference value of the optimization function based on the predicted behavior recognition result and the mask vector.
[0124] Step S142: Adjust the mask vector according to the reference value until the optimization function reaches the convergence condition to obtain the target mask vector.
[0125] In one embodiment, refer to Figure 6 , Step S142 includes but is not limited to Steps S1421 to S1425:
[0126] Step S1421: Adjust the mask vector to obtain the optimized mask vector.
[0127] In one embodiment, a sub - graph of the mask vector G is generated through a random graph vector as the optimized mask vector.
[0128] Step S1422: Calculate the optimized motion sequence based on the optimized mask vector and the motion vector.
[0129] Step S1423: Input the optimized motion sequence into the pre - trained behavior recognition model to obtain the optimized predicted behavior recognition result.
[0130] Step S1424: Calculate the reference value of the optimization function corresponding to the optimized predicted behavior recognition result.
[0131] The method for calculating the reference value in the above Steps S1422 to S1424 is the same as that in Steps S120 to S140, except that the optimized mask vector is used instead of the mask vector.
[0132] Step S1425: Determine whether the optimization function reaches the convergence condition according to the reference value. If it reaches the convergence condition, take the optimized mask vector as the target mask vector; otherwise, repeat the above steps until the optimization function reaches the convergence condition.
[0133] In one embodiment, if the convergence condition is not reached, continuously adjust the optimized mask vector according to Step S1421 until the convergence condition is reached.
[0134] In one embodiment, an optimization function is constructed based on the mutual information quantity. The mutual information quantity is used to represent whether there is a relationship between two variables A and variable B (here, variables A and B are only for illustration and have no substantial meaning), and the strength of the relationship, which is expressed as:
[0135]
[0136] Among them, MI(A,B) represents the value of the mutual information quantity, P(A,B) represents the joint probability density between variables A and B, P(A) represents the probability density of variable A, and P(B) represents the probability density of variable B. It can be seen from the above formula that if variables A and B are independent, then P(A,B) = P(A)P(B), and MI(A,B) is obtained as 0, indicating that variables A and B are not related.
[0137] By derivation and transformation, we get:
[0138] MI(A,B) = H(B) - H(B|A)
[0139] H(B) = -∫ B P(B)logP(B)
[0140] Among them, H(B) represents the entropy of variable B, which is a parameter used to measure the uncertainty of variable B. The more discrete the distribution of variable B is, the higher this value is. H(B|A) represents the uncertainty of variable B under the condition that variable A is known. In this embodiment, MI(A,B) can be interpreted as the amount by which the uncertainty of variable B is reduced due to the introduction of variable A, and this reduced amount is expressed as H(B|A). Therefore, if variables A and B are more closely related, MI(A,B) is larger. At this time, H(B|A) is 0, indicating that variables A and B are completely related. Under the condition that variable A is determined, variable B is a fixed value, and there is no probability of other uncertain situations. Therefore, H(B|A) is 0. When the value of MI(A,B) is 0, it indicates that variables A and B are independent. At this time, H(B) = H(B|A), indicating that the occurrence of variable A does not affect variable B.
[0141] In one embodiment, an optimization function is constructed according to the meaning of the above mutual information quantity. Since the purpose of this embodiment is to obtain the correlation between the predicted behavior recognition result, the motion vector, and the mask vector corresponding to the motion vector, the optimization function is expressed as:
[0142] MI(Y,(G,Z)) = H(Y|Z) - H(Y|G,Z)
[0143] Among them, Y represents the predicted behavior recognition result, Z represents the motion vector, G represents the mask vector, H(·) represents the entropy function, and MI(Y,(G,Z)) represents the reference value.
[0144] According to the purpose of the optimization function, this convergence condition means finding a suitable mask vector G such that there is the strongest correlation between the predicted behavior recognition result Y, the motion vector Z, and the mask vector G corresponding to the motion vector Z. The strongest correlation is manifested as the maximum value of the mutual information of (G, Z) with respect to Y, that is, the maximum reference value.
[0145] Therefore, in the above embodiment, the convergence condition of the optimization function is to maximize the reference value of the optimization function, which is expressed as:
[0146]
[0147] In one embodiment, H(Y|Z) represents the uncertainty of the predicted behavior recognition result Y when the motion vector Z is known. Since the relationship between the motion vector Z and the predicted behavior recognition result Y is known, H(Y|Z) is a constant.
[0148] And expanding the conditional information quantity H(Y|G,Z) gives:
[0149]
[0150] Where P represents the graph neural network model, φ represents the model parameters, and E represents the expectation.
[0151] In one embodiment, the above formula is interpreted as a variational approximation of the G subgraph distribution using continuous relaxation, and iterative optimization is performed by gradient descent to obtain the target mask vector to explain the importance of each action in the skeleton.
[0152] The convergence condition of the above formula can be rewritten as:
[0153]
[0154] According to Jensen's inequality, the convergence condition is transformed to get:
[0155]
[0156] The above embodiment uses the gradient descent method to adjust the reference value until the convergence condition is reached, and then the target mask vector representing the weights of the human key points of the motion vector can be obtained.
[0157] An embodiment of the present application provides a method for interpreting a deep neural network. The motion vectors and mask vectors of human key points are obtained, and then a motion sequence is calculated based on the mask vectors and motion vectors. The motion sequence is then input into a pre-trained behavior recognition model to obtain a predicted behavior recognition result. Finally, the reference value of the optimization function is iteratively updated using the predicted behavior recognition result to obtain a target mask vector, and then the behavior recognition model is interpreted using the target mask vector. In this embodiment, the target mask vector is obtained by iterative update according to the mask vector, and the weights of the human key points of the motion vectors are characterized by the target mask vector. Considering the influence of the weights of different features in the motion vectors on the judgment result of the behavior recognition model, the accuracy and recognition efficiency of the behavior recognition result are improved accordingly.
[0158] In addition, based on obtaining the target mask vector for interpreting the behavior recognition model, an embodiment of the present application proposes a visualization method, referring to Figure 7 , which is a flowchart of the visualization method in this embodiment, including but not limited to steps S210 to S230.
[0159] Step S210, obtain the motion sequence and the target mask vector.
[0160] In one embodiment, the motion sequence and the target mask vector are calculated using the method for interpreting a deep neural network according to any one of the above embodiments.
[0161] Step S220, sample the skeleton motion vectors frame by frame according to the motion sequence.
[0162] Step S230, visually display the weight changes of the skeleton motion vectors between different frames according to the target mask vector.
[0163] In one embodiment, referring to Figure 8 , it is a schematic diagram of the visualization in this embodiment.
[0164] First, select the minimum coordinate in the motion sequence (i.e., the sequence containing human skeleton motion information), and set it as the origin of the drawing coordinates, that is, the coordinates of each point in the motion sequence are subtracted by this minimum coordinate and then drawn in a three-dimensional coordinate system; then sample the motion sequence frame by frame, for example, sample once every 5 frames to obtain the skeleton motion vectors; use a drawing tool (such as matplotlib) to draw the skeleton motion vectors to obtain a schematic diagram of the skeleton motion vectors as shown in Figure 8 shown, Figure 8As shown, the human key points 300 (such as skeleton nodes) in the skeleton motion vector and the connections of the skeleton are plotted on a three-dimensional coordinate according to the coordinates and time sequence. The action interval is set to 1.5 times the width of the previous frame's action to improve the visualization display effect. Finally, according to the generated target mask vector (i.e., the weight G obtained in the above embodiment), the importance degree of the skeleton motion vector between different frames is visually displayed. The importance degree can be presented in the form of a heat map. In this embodiment, referring to Figure 8 , the importance of each human key point in the motion vector is represented by the depth of color, so as to interpret the behavior recognition model and visually illustrate the influence degree of the different weights of different skeleton nodes in different motion states on the judgment result of the behavior recognition model. Furthermore, according to this weight, the behavior recognition model can be adjusted to obtain a more accurate predicted behavior recognition result.
[0165] It can be understood that the above Figure 8 is only a visualization schematic diagram of this embodiment and does not limit this embodiment.
[0166] In addition, an embodiment of the present application also provides an interpretable device for a deep neural network. Referring to Figure 9 , the device includes:
[0167] A first acquisition module 910, configured to acquire the motion vector and the mask vector of the human key points;
[0168] A motion sequence calculation module 920, configured to calculate a motion sequence according to the mask vector and the motion vector;
[0169] A preliminary prediction module 930, configured to input the motion sequence into a pre-trained behavior recognition model to obtain a predicted behavior recognition result;
[0170] An iterative optimization module 940, configured to iteratively update and optimize the reference value of the optimization function by using the predicted behavior recognition result to obtain a target mask vector, where the target mask vector is used to characterize the weight of the human key points of the motion vector.
[0171] In one embodiment, the motion sequence calculation module 920 is further configured to calculate the motion data information of the motion vectors and the corresponding mask vectors at all previous moments before the acquisition moment, and accumulate the motion data information to obtain a motion sequence.
[0172] In one embodiment, the iterative optimization module 940 is further configured to calculate a reference value of the optimization function according to the predicted behavior recognition result and the mask vector, and adjust the mask vector according to the reference value until the optimization function reaches the convergence condition to obtain a target mask vector.
[0173] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, that is, they may be located in one place or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution in this embodiment.
[0174] It should be noted that the deep neural network interpretable device in this embodiment can execute the deep neural network interpretable method in the embodiment as Figure 2 shown. That is, the deep neural network interpretable device in this embodiment and the deep neural network interpretable method in the embodiment as Figure 2 shown belong to the same inventive concept. Therefore, these embodiments have the same implementation principle and technical effects, and will not be elaborated here.
[0175] In addition, an embodiment of the present application also provides a visualization device. Referring to Figure 10 , the device includes:
[0176] A second acquisition module 1010, configured to acquire a motion sequence and a target mask vector, where the motion sequence and the target mask vector are calculated by using the deep neural network interpretable method in any one of the above embodiments;
[0177] A motion vector sampling module 1020, configured to sample the skeleton motion vector frame by frame according to the motion sequence;
[0178] A visualization module 1030, configured to visually display the weight change of the skeleton motion vector between different frames according to the target mask vector.
[0179] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, that is, they may be located in one place or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution in this embodiment.
[0180] It should be noted that the visualization device in this embodiment can execute the visualization method in the embodiment as Figure 7 shown. That is, the visualization device in this embodiment and the visualization method in the embodiment as Figure 7 shown belong to the same inventive concept. Therefore, these embodiments have the same implementation principle and technical effects, and will not be elaborated here.
[0181] In addition, an embodiment of the present application also provides a computer device, which includes: a memory, a processor, and a computer program stored on the memory and executable on the processor.
[0182] The processor and the memory can be connected via a bus or other means.
[0183] As a non-transitory computer-readable storage medium, the memory can be used to store non-transitory software programs and non-transitory computer-executable programs. In addition, the memory can include high-speed random access memory, and can also include non-transitory memory, such as at least one magnetic disk storage device, a flash memory device, or other non-transitory solid-state storage devices. In some embodiments, the memory optionally includes a memory remotely located relative to the processor, and these remote memories can be connected to the processor through a network. Examples of the above-mentioned network include but are not limited to the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.
[0184] The non-transitory software programs and instructions required to implement the deep neural network interpretability method or visualization method of the above embodiments are stored in the memory. When executed by the processor, the deep neural network interpretability method or visualization method in the above embodiments is executed. For example, the method steps S110 to S140 described above are executed. Figure 2 in the method steps S110 to S140 in Figure 4 the method steps S121 to step S122 in Figure 7 the method steps S210 to S230 in etc.
[0185] In addition, an embodiment of the present application further provides a computer-readable storage medium. The computer-readable storage medium stores computer-executable instructions. When the computer-executable instructions are executed by a processor or a controller, for example, executed by a processor in the above computer device embodiment, the above processor can execute the deep neural network interpretability method or visualization method in the above embodiments. For example, the method steps S110 to S140 described above are executed. Figure 2 in the method steps S110 to S140 in Figure 4 the method steps S121 to step S122 in Figure 7 the method steps S210 to S230 in etc.
[0186] For another example, when executed by a processor in the above computer device embodiment, the above processor can execute the deep neural network interpretability method or visualization method in the above embodiments. For example, the method steps S110 to S140 described above are executed. Figure 2 in the method steps S110 to S140 in Figure 4 the method steps S121 to step S122 in Figure 7 the method steps S210 to S230 in etc.
[0187] Those of ordinary skill in the art will understand that all or some of the steps and systems disclosed above in the methods can be implemented as software, firmware, hardware, and appropriate combinations thereof. Some physical components or all physical components can be implemented as software executed by a processor, such as a central processing unit, a digital signal processor, or a microprocessor, or can be implemented as hardware, or can be implemented as an integrated circuit, such as an application specific integrated circuit. Such software can be distributed on a computer-readable medium, which can include a computer storage medium (or non-transitory medium) and a communication medium (or transitory medium). As is well known to those of ordinary skill in the art, the term computer storage medium includes volatile and non-volatile, removable and non-removable media implemented in any method or technology for storing information, such as computer-readable instructions, data structures, program modules, or other data. Computer storage media includes, but is not limited to, RAM, ROM, EEPROM, flash memory or other memory technologies, CD-ROM, digital versatile disks (DVD) or other optical disk storage, magnetic cassettes, magnetic tape, magnetic disk storage or other magnetic storage devices, or any other medium that can be used to store the desired information and can be accessed by a computer. In addition, it is well known to those of ordinary skill in the art that a communication medium typically contains computer-readable instructions, data structures, program modules, or other data in a modulated data signal such as a carrier wave or other transmission mechanism, and can include any information delivery medium.
[0188] The above is a specific description of the preferred implementation of the embodiments of the present application. However, the embodiments of the present application are not limited to the above implementation manners. Those skilled in the art can also make various equivalent deformations or substitutions without departing from the spirit of the embodiments of the present application. These equivalent deformations or substitutions are all included within the scope defined by the claims of the embodiments of the present application.
Claims
1. A method for interpreting a deep neural network, characterized in that, it includes: Obtain the motion vector and mask vector of human key points; Calculate a motion sequence based on the mask vector and the motion vector; Input the motion sequence into a pre-trained behavior recognition model to obtain a predicted behavior recognition result, and the behavior recognition model is a deep neural network model; Use the predicted behavior recognition result to iteratively update the reference value of the optimization function to obtain a target mask vector for characterizing the weight of the motion vector; The calculating the motion sequence based on the mask vector and the motion vector includes: calculating the motion data information of the motion vectors and the corresponding mask vectors at all moments before the acquisition moment; accumulating the motion data information to obtain the motion sequence; The calculating the motion data information of the motion vectors and the corresponding mask vectors at all moments before the acquisition moment includes: calculating the Hadamard product between the motion vector and the corresponding mask vector to obtain the motion data information; expressed as: Among them, represents the Hadamard product, represents the motion vector at time t, represents the mask vector at time t, represents the motion data information at time t, C represents the motion three-dimensional coordinate information, V represents the number of joint points of each person, and M represents the number of people in the motion sequence; The using the predicted behavior recognition result to iteratively update the reference value of the optimization function to obtain a target mask vector for characterizing the weight of the motion vector includes: calculating the reference value of the optimization function according to the predicted behavior recognition result and the mask vector; adjusting the mask vector according to the reference value until the optimization function reaches the convergence condition to obtain the target mask vector.
2. The method for interpreting a deep neural network according to claim 1, characterized in that, the adjusting the mask vector according to the reference value until the optimization function reaches the convergence condition to obtain the target mask vector includes: Adjust the mask vector to obtain an optimized mask vector; Calculate an optimized motion sequence based on the optimized mask vector and the motion vector; Input the optimized motion sequence into a pre-trained behavior recognition model to obtain an optimized predicted behavior recognition result; Calculate the reference value of the optimization function corresponding to the optimized predicted behavior recognition result; Judge whether the optimization function reaches the convergence condition according to the reference value. If the convergence condition is reached, use the optimized mask vector as the target mask vector; otherwise, repeat the above steps until the optimization function reaches the convergence condition.
3. The method for interpreting a deep neural network according to any one of claims 1 or 2, characterized in that, the optimization function is a mutual information optimization function, expressed as: wherein, represents the predicted behavior recognition result, represents the motion vector, represents the mask vector, represents the entropy function, represents the reference value.
4. The method for interpreting a deep neural network according to claim 3, characterized in that, the convergence condition of the optimization function is: maximizing the reference value of the optimization function.
5. A visualization method, characterized in that, it includes: Obtain a motion sequence and a target mask vector, and the motion sequence and the target mask vector are calculated by using the method for interpreting a deep neural network according to any one of claims 1 to 4; Sample the skeleton motion vector frame by frame according to the motion sequence; Visualize and display the weight change of the skeleton motion vector between different frames according to the target mask vector.
6. A device for interpreting a deep neural network, characterized in that, it includes: The first acquisition module is configured to acquire the motion vectors and mask vectors of human key points; The motion sequence calculation module is configured to calculate a motion sequence based on the mask vector and the motion vector; The preliminary prediction module is configured to input the motion sequence into a pre-trained behavior recognition model to obtain a predicted behavior recognition result; The iterative optimization module is configured to iteratively update the reference value of the optimization function by using the predicted behavior recognition result to obtain a target mask vector for characterizing the weight of the motion vector; The calculating the motion sequence based on the mask vector and the motion vector includes: calculating the motion data information of the motion vectors and the corresponding mask vectors at all times before the acquisition time; accumulating the motion data information to obtain the motion sequence; The calculating the motion data information of the motion vectors and the corresponding mask vectors at all times before the acquisition time includes: calculating the Hadamard product between the motion vector and the corresponding mask vector to obtain the motion data information; expressed as: Among them, represents the Hadamard product, represents the motion vector at time t, represents the mask vector at time t, represents the motion data information at time t, C represents the motion three-dimensional coordinate information, V represents the number of joint points of each person, and M represents the number of people in the motion sequence; The iterative updating the reference value of the optimization function by using the predicted behavior recognition result to obtain a target mask vector for characterizing the weight of the motion vector includes: calculating the reference value of the optimization function according to the predicted behavior recognition result and the mask vector; adjusting the mask vector according to the reference value until the optimization function reaches the convergence condition to obtain the target mask vector.
7. A visualization device Characterized in that Comprising: A second acquisition module configured to acquire a motion sequence and a target mask vector, where the motion sequence and the target mask vector are calculated by using the deep neural network interpretable method according to any one of claims 1 to 4; The motion vector sampling module is configured to sample the skeleton motion vectors frame by frame according to the motion sequence; The visualization module is configured to visually display the weight change of the skeleton motion vectors between different frames according to the target mask vector.
8. A computer device Characterized in that Comprising a processor and a memory; The memory is used to store programs; The processor is configured to execute the deep neural network interpretable method according to any one of claims 1 to 4, or the visualization method according to claim 5 according to the program.
9. A computer-readable storage medium storing computer-executable instructions for executing the deep neural network interpretable method according to any one of claims 1 to 4, or the visualization method according to claim 5.
Citation Information
Patent Citations
Graph convolutional neural network action recognition method based on attention mechanism
CN113128424A
Harmonized local illumination compensation and modified inter coding tools
CN113287317A