Interpretability Behavior Recognition Method Guided by Natural Language Knowledge Description
By constructing an interpretable behavior recognition method based on natural language description, RGB video encoding is used as a timing diagram structure, combined with object detectors, spatiotemporal feature extraction and description prefix tree, transparent interpretation of video behavior is achieved, solving the problem of lack of evidence support and vulnerability in the existing technology, and improving the interpretability and security of identification.
Patent Information
- Application Number
- CN202211448907.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-11-18
- Publication Date
- 2025-07-22
- Estimated Expiration
- 2042-11-18
AI Technical Summary
The existing deep learning-based video behavior recognition methods lack clear evidence support and interpretability, and are vulnerable to attacks, limiting their application in real-world scenarios with strict security requirements.
By constructing an interpretable behavior recognition method based on natural language description, RGB video encoding is used as a timing diagram structure, combining object detectors, spatiotemporal feature extraction, description prefix tree and feature space discriminator, interpretive recognition of video behavior and transparency of decision-making process.
It improves the interpretability and reliability of behavior recognition, reduces the risk of confrontational attacks, and makes the method more useful in scenarios with strict security requirements.
Smart Images

Figure CN115761886B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of artificial intelligence and computer vision, and particularly relates to an interpretable behavior recognition method guided by natural language knowledge description. Background Art
[0002] In recent years, with the wide application of devices such as cameras, it has become increasingly important to recognize human behaviors in videos, and behavior recognition based on deep learning technology has attracted much attention. By automatically and correctly recognizing human behaviors in videos, the service level of intelligent devices can be improved or serious safety problems caused by human behaviors can be avoided. With the development of various well-designed neural architectures and end-to-end learning algorithms, the recognition of human behaviors in videos has made great progress in recent years. Currently, the behavior recognition methods based on deep technology mainly have the following problems:
[0003] 1. The current methods input a video clip and show the confidence of each behavior category through multi-layer calculations. Such a black-box prediction mechanism does not clearly provide convincing evidence about the behavior, such as the time / location / cause of the behavior occurrence.
[0004] 2. The current methods are unreliable and unexplainable, and may lead to serious safety problems when attacked.
[0005] Currently, the interpretable behavior recognition methods based on videos usually use the state or relationship changes of objects in the video to explain the behavior recognition process, without using the free-form language description, a way that humans can more easily understand, to prove the decisions of the behavior recognition method, and at the same time, without explaining the decision-making process of each segment of the video, which limits its application in many real scenarios with strict safety requirements. Summary of the Invention
[0006] In order to overcome the above-mentioned deficiencies of the prior art, the purpose of the present invention is to provide an interpretable behavior recognition method guided by natural language knowledge description. Its first purpose is to explain the process of the neural network recognizing video behaviors by constructing a method for aligning natural language description-video segment features, and at the same time, to explain the decision-making process of each segment of the video; its second purpose is to avoid the application limitations caused by the vulnerability of the recognition network to attacks, so as to be applied to real scenarios with strict safety requirements.
[0007] In order to achieve the above purpose, the technical solution adopted by the present invention is:
[0008] An interpretable behavior recognition method guided by natural language knowledge description, comprising the following steps:
[0009] (1) Encode the RGB video into a temporal graph structure;
[0010] (1a) The input is an RGB video frame {I t}, where t = 1, …, T, and the action label is a. An offline object detector is used to detect N entities in each frame, where the entities are people or objects. Using these entities as nodes, all pairs of nodes within a frame are connected, and nodes of the same entity across frames are connected in chronological order to form a temporal graph structure;
[0011] (1b) According to the detection results of the object detector, the appearance features, spatial features, and semantic features of each entity are extracted, and these features are combined into node features The corresponding nodes in the graph are instantiated;
[0012] (2) Spatiotemporal feature extraction;
[0013] (2a) In each frame, centered on any node, through convolution centered on a person (object), the features of itself and neighboring object (person) nodes are aggregated to obtain the updated node features Specifically, if there are only objects in the scene, a virtual node is set as a person node for feature aggregation, and the virtual node is initialized with a zero vector of the same scale as other nodes.
[0014] (2b) Repeat step (2a) several times to obtain refined node features, and use average pooling to extract the spatial feature vector m for each frame t ; Use a bidirectional GRU to extract the spatiotemporal feature vector h for the spatial feature vector m of each frame t to encode the temporal dependencies between the features of each frame; t
[0015] (2c) Use a differentiable discrete value estimator GSM to estimate the spatial feature vector m t and the spatiotemporal feature vector h t to determine whether the current time t is the end of a segment. If so, the spatiotemporal feature h at the current time is used as the spatiotemporal feature v of a segment t The spatiotemporal features of video segments are represented as V = {v k}, where k = 1, …, K; k
[0016] (3) Action classifier;
[0017] (3a) The spatiotemporal features of video segments respectively pass through a multi-layer perceptron to obtain the action classification scores of each segment, and the average classification score of all segments is taken as the final score s p ;
[0018] (4) Construct a description prefix tree;
[0019] (4a) The input is a description document database corresponding to video behavior tags Match the most relevant set of documents for behavior tag a and entity category in step (1a) Specifically, the entity category is used for matching in the inference stage;
[0020] (4b) Parse the sentences in each document into a linked list of phrases, merge the repeated sentences in the linked list, and use a prefix tree to store and organize the elements in the linked list;
[0021] (4c) Use the Bert model to tokenize and encode the sentences in the prefix tree to obtain word-level features w i , and input them into a bidirectional GRU to obtain the context features s of the sentences i ;
[0022] (4d) Organize the context features si of each sentence according to the structure of the prefix tree in (4b) to obtain a description prefix tree;
[0023] (5) Feature space discrimination;
[0024] (5a) The input is a GRU (m t ) that encodes the temporal dependence between spatial features of each frame and a GRU (w i ) that encodes the context features of sentences;
[0025] (5b) Use the domain classifier D of the GAN network to discriminate the output of the GRU network to align the cross-domain feature distributions;
[0026] (5c) Use a Gradient Reversal Layer (GRL) to connect the GRU network and the domain classifier D. The GRL inputs the output of the GRU network into the domain classifier D, makes no changes to the elements of the output of the GRU network during forward propagation, and reverses the gradient during backpropagation;
[0027] (6) Feature similarity matching;
[0028] (6a) The input is the segmented spatio-temporal feature v corresponding to the current time step k and the subtree node features {s i} corresponding to the selected node in the description prefix tree at the previous time step, i = 1,..., M;
[0029] (6b) Calculate the similarity scores between each subtree node feature and the segmented spatio-temporal feature using Cosine Similarity Obtain the text description corresponding to the subtree node with the highest score as the description of the video segment;
[0030] (6c) Repeat steps (6a) and (6b) to traverse all segmented spatio-temporal features, and obtain the segmented description of the entire video. The description of each segment explains the behavior recognition process.
[0031] Further, the person (object)-centered convolution calculation formula described in step (2a) is as follows:
[0032]
[0033] where c e and c r are the types of the node itself and the neighboring nodes respectively, is the set of neighboring nodes with different types from the current node type; specifically, the implementation of HOC is as follows:
[0034]
[0035]
[0036] where u is the feature at the query position, {v k} k=1,…,n enumerates the features at all possible positions, and W, φ, and φ′ are linear projection functions for node feature transformation.
[0037] Further, the formula for segmenting spatio-temporal features using GSM described in step (2c) is as follows:
[0038] u t = GSM{γ([m t , h t )}
[0039] where the binary output u t = 1 indicates that the current time t is the last frame of the current segment, and the spatio-temporal feature h t instantiates the spatio-temporal feature v k .
[0040] Further, the calculation of taking the average classification score of all segments as the final score described in step (3a) is as follows:
[0041]
[0042] where f is a two-layer fully connected network.
[0043] Further, the loss function of feature similarity matching described in step (6) is as follows:
[0044]
[0045] where α is the similarity margin hyperparameter. When vk when the video segment where it is located matches the s i in the text description where it is located, t sim = 1, otherwise t sim = 0.
[0046] Furthermore, in step (5a), the bidirectional GRU shares parameters with the GRU in (2c).
[0047] Furthermore, the calculation formula of the gradient reversal layer in step (5c) is as follows:
[0048] R λ (x) = x
[0049]
[0050] where R λ represents the gradient reversal layer (GRL) operation, and x ∈ {GRU(mt), GRU(wi)} represents the video segment feature and the sentence context feature.
[0051] Furthermore, during the forward propagation in the training process of step (5c), the GRL does not perform any operation on the passing elements, and during the backward propagation, the gradient propagated to the GRL is multiplied by -λ.
[0052] Furthermore, the loss function for feature space discrimination in step (5) is as follows:
[0053]
[0054] where Θ GRU and Θ D are the parameters of the GRU and the domain classifier D respectively, represents the expectation of the distribution.
[0055] Advantages of the present invention:
[0056] First, the present invention retrieves descriptions related to videos from the description document database, constructs a description prefix tree after parsing, and through text-visual feature space discrimination and similarity matching strategies, effectively interprets the decision-making process of video recognition using free-form language descriptions, overcoming the problem that the black-box prediction mechanism of the prior art cannot clearly provide convincing evidence about behaviors, making the present invention have better interpretability and reliability, and thus the recognition model is not easily vulnerable to adversarial attacks and other advantages.
[0057] Second, the present invention interprets the decision-making process of video recognition through free-form language descriptions, matches the video segment features by traversing the nodes in the description prefix tree hierarchically, improves the interpretability of the intermediate features, and makes the interpretation of the behavior recognition process easier to understand, making the present invention have the advantage of interpreting behavior recognition in a higher dimension.
[0058] Thirdly, the present invention extracts the features of videos and texts based on the GRU with shared parameters, and uses a domain classifier to discriminate the output features, effectively narrowing the domain gap existing between the video domain and the text domain, enabling knowledge to be better transferred between the two domains, and enabling the present invention to achieve better classification results in the action recognition task. Brief Description of the Drawings
[0059] Figure 1 is the overall framework of the interpretable behavior recognition method guided by natural language knowledge description.
[0060] Figure 2 is a schematic diagram of the description prefix tree corresponding to the "drink water" behavior.
[0061] Figure 3 is a schematic diagram for calculating the similarity between the features of the sub-tree nodes of the description prefix tree and the segmented spatio-temporal features. Detailed Description of the Preferred Embodiments
[0062] The present invention will be further described in detail below with reference to the accompanying drawings.
[0063] Refer to the attached Figure 1 , and further describe the specific steps of the present invention.
[0064] Step 1. Encode the RGB video into a temporal graph structure;
[0065] The input is the RGB video frames {I t}, t = 1, …, T, and the behavior label a. Use an offline object detector to detect N entities in each frame, where the entities are people or objects; use these entities as nodes, connect each pair of nodes within the frame, and connect the same entity nodes between frames in chronological order to form a temporal graph structure;
[0066] Extract the appearance feature, spatial feature, and semantic feature of each entity according to the detection result, and combine these features into node features Instantiate the corresponding nodes in the graph;
[0067] Step 2. Spatio-temporal feature extraction;
[0068] In each frame, centered on any node, aggregate the features of itself and neighboring object (person) nodes through person (object)-centered convolution to obtain the updated node features If there are only objects in the scene, set a virtual node as the person node for feature aggregation, and the virtual node is initialized with a all-zero vector of the same scale as other nodes;
[0069] Repeat the convolution centered on the person (object) twice to obtain refined node features, and use the average pooling method to extract the spatial feature vector m for each frame t ; Use a bidirectional GRU for the spatial feature vector m of each frame t to extract the feature vector h t , encoding the temporal dependencies between the features of each frame;
[0070] Use the differentiable discrete value estimator GSM to estimate the spatial feature vector m t and the spatio-temporal feature vector h t to determine whether the current time t is the end of a segment. If so, use the spatio-temporal feature h at the current time t as the spatio-temporal feature v of a segment k , and the spatio-temporal features of the video segments are represented as V = {v k}, k = 1, …, K;
[0071] Step 3. Action classifier;
[0072] The spatio-temporal features of the video segments are respectively passed through a multi-layer perceptron to obtain the behavior classification scores of each segment, and the average classification score of all segments is taken as the final score s p ;
[0073] Step 4. Construct a descriptive prefix tree;
[0074] The input is the descriptive document database corresponding to the video behavior labels to match the most relevant set of documents for the behavior label a and entity category in step (1a) In particular, the entity category is used for matching in the inference stage;
[0075] Parse the sentences in each document into a linked list of phrases and merge the repeated sentences in the linked list, and use a prefix tree to store and organize the elements in the linked list, as Figure 2 shown;
[0076] Use the Bert model to tokenize and encode the sentences in the prefix tree to obtain word-level features w i , and input them into a bidirectional GRU to obtain the context features s of the sentences i ; In particular, this bidirectional GRU shares parameters with the GRU in video segment feature extraction;
[0077] Organize the context features si of each sentence according to the structure of the prefix tree in (4b) to obtain a descriptive prefix tree;
[0078] Step 5. Feature space discrimination;
[0079] The input is the GRU that encodes the temporal dependencies between the spatial features of each frame (m t) and a GRU (w that encodes sentence context features i );
[0080] The domain classifier D of the GAN network is used to discriminate the output of the GRU network to align the feature distributions across domains;
[0081] The Gradient Reversal Layer (GRL) is used to connect the GRU network and the domain classifier D. The GRL inputs the output of the GRU network into the domain classifier D, making no changes to the elements of the GRU network output during forward propagation, while reversing the gradient during backpropagation;
[0082] Step 6. Feature similarity matching;
[0083] The input is the segmented spatio-temporal feature v corresponding to the current time step k and the sub-tree node features {s corresponding to the node selected in the description prefix tree at the previous time step i} , i = 1, …, M;
[0084] Cosine Similarity is used to calculate the similarity scores between each sub-tree node feature and the segmented spatio-temporal feature The text description corresponding to the sub-tree node with the highest score is taken as the description of the video segment, as Figure 3 shown;
[0085] The matching is repeated to traverse all the segmented spatio-temporal features to obtain the segmented descriptions of the entire video. The description of each segment explains the behavior recognition process. In particular, the segmented description of the entire video output during the inference phase can explain the recognition process of the video.
Claims
1. An interpretable behavior recognition method guided by natural language knowledge description, characterized in that, Including the following steps: (1) Encode the RGB video into a temporal graph structure; (1a) The input is an RGB video frame {I t}, where \(t = 1,\ldots,T\), and the action label \(a\). An offline object detector is used to detect \(N\) entities in each frame, where the entities are people or objects. These entities are used as nodes, and each pair of nodes within a frame is connected. Nodes of the same entity across frames are connected in chronological order to form a temporal graph structure. (1b)Extract the appearance features, spatial features, and semantic features of each entity based on the detection results of the target detector, and combine these features into node features Instantiate the corresponding nodes in the graph; (2) Spatiotemporal feature extraction; (2a) In each frame, centered around any node, aggregate the features of itself and neighboring object nodes through human-centered convolution Obtain the updated node features If there are only objects in the scene, set a virtual node as the human node for feature aggregation, and the virtual node is initialized with a all-zero vector of the same scale as other nodes; (2b) Repeat step (2a) several times to obtain refined node features, and use average pooling to extract the spatial feature vector m for each frame. t ; Use bidirectional GRU for the spatial feature vector m of each frame t to extract the feature vector h t , and encode the temporal dependencies between the features of each frame; (2c) Estimate the spatial feature vector m using the differentiable discrete value estimator GSM t and the spatio-temporal feature vector h t to determine whether the current time t is the end of a segment. If so, use the spatio-temporal feature h at the current time t as a spatio-temporal feature v of a segment k , and the spatio-temporal features of video segments are represented as V = {v k}, k = 1, …, K; (3) Action classifier; (3a) The spatio-temporal features of each video segment are respectively passed through a multi-layer perceptron to obtain the behavior classification score of each segment, and the average classification score of all segments is taken as the final score s p ; (4) Construct a description prefix tree; (4a) The input is a description document database corresponding to video behavior tags Match the most relevant set of documents for the behavior tag a and entity category in step (1a) Specifically, the entity category is used for matching in the inference phase; (4b) Parse the sentences in each document into a linked list of phrases and merge the repeated sentences in the linked list, and use a prefix tree to store and organize the elements in the linked list; (4c) Tokenize and encode the sentences in the prefix tree using the Bert model to obtain word-level features w i and input them into a bidirectional GRU to obtain the context features s of the sentence i ; (4d) Organize the context features s of each sentence according to the structure of the prefix tree in step (4b). i Obtain the described prefix tree. (5) Feature space discrimination; (5a) The inputs are a GRU(m t ) that encodes the temporal dependencies between the spatial features of each frame and a GRU(w i ) that encodes the sentence context features; (5b) Use the domain classifier D of the GAN network to discriminate the output of the GRU network to align the cross-domain feature distributions; (5c) Use the Gradient Reversal Layer (GRL) to connect the GRU network and the domain classifier D. The GRL inputs the output of the GRU network into the domain classifier D, and does not make any changes to the elements output by the GRU network during forward propagation, while reversing the gradient during backpropagation; (6) Feature similarity matching; (6a) The input is the segmented spatio-temporal feature v corresponding to the current time step k and the sub-tree node features {s i} i = 1, …, M corresponding to the node selected in the description prefix tree at the previous time step; (6b) Calculate the similarity score between each subtree node feature and the segmented spatio-temporal feature using Cosine Similarity Obtain the text description corresponding to the subtree node with the highest score as the description of the video segment; (6c) Repeat steps (4a) and (4b) to traverse all segmented spatiotemporal features, and obtain the segmented description of the entire video. The description of each segment explains the behavior recognition process.
2. The interpretable behavior recognition method based on natural language knowledge description guidance according to claim 1, wherein The human-centered convolution calculation formula described in step (2a) is as follows: where c e and c r are the type of the node itself and the type of the neighboring node, respectively, is the set of neighboring nodes with a type different from the current node type; specifically, the HOC is implemented as follows: where u is the feature of the query position, {v k} k=1,…,n enumerate the features of all possible positions, W, φ, and φ ′ are linear projection functions for node feature transformation.
3. The interpretable behavior recognition method based on natural language knowledge description guidance according to claim 1, characterized in that The calculation formula for segmenting the spatiotemporal features using GSM described in step (2c) is as follows: u t = GSM{γ([m t ,h t )} where the binary output u t = 1 indicates that the current time t is the last frame of the current segment, and the spatio-temporal feature h at this time is taken t Instantiate the spatio-temporal feature v of the current segment k .
4. The interpretable behavior recognition method based on natural language knowledge description guidance according to claim 1, wherein The calculation formula for taking the average classification score of all segments as the final score described in step (3a) is as follows: where f is a two-layer fully connected network and K is the number of segmented spatiotemporal features.
5. The interpretable behavior recognition method based on natural language knowledge description guidance according to claim 1, characterized in that The loss function of the feature similarity matching described in step (6) is as follows: where α is the similarity margin hyperparameter, when the video segment where v k is located matches the text description where s i is located, t sim = 1, otherwise t sim = 0.
6. The interpretable behavior recognition method based on natural language knowledge description guidance according to claim 1, characterized in that The bidirectional GRU in step (5a) shares parameters with the GRU in step (2b).
7. An interpretable behavior recognition method based on natural language knowledge description guidance according to claim 1, characterized in that The segmented description of the entire video output in the inference stage of step (6c) can explain the video recognition process.
8. An interpretable behavior recognition method based on natural language knowledge description guidance according to claim 1, characterized in that The calculation formula of the Gradient Reversal Layer described in step (5c) is as follows: R λ f(x) = x where R λ represents a gradient reversal layer (GRL) operation, and x ∈ {GRU(m t ), GRU(w i )} represents video segment features and sentence context features.
9. The interpretable behavior recognition method based on natural language knowledge description guidance according to claim 1, characterized in that During the forward propagation in the training process of step (5c), the GRL does not perform any operations on the passed elements, and during the backpropagation, the gradient propagated to the GRL will be multiplied by -λ.
10. An interpretable behavior recognition method based on natural language knowledge description guidance according to claim 1, characterized in that, The loss function of the feature space discrimination described in step (5) is as follows: where Θ GRU and Θ D are the parameters of the GRU and the domain classifier D respectively, denotes the expectation of the distribution.
Citation Information
Patent Citations
Video identification method and related device
CN112203115A
System and iterative method for lexicon, segmentation and language model joint optimization
CN1387651A