A Multi-Task Method for Behavior Recognition and Object Detection by Fusing Human-Object Interaction Analysis
By simultaneously extracting the features of behavior recognition and object detection tasks in a single neural network model, and using long and short-term spatial variation convolution and graph network model for feature correction, the problem of unused connection between the two tasks in the existing technology is solved, and more efficient video analysis performance is achieved.
Patent Information
- Application Number
- CN202111297058.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-11-03
- Publication Date
- 2025-06-13
- Estimated Expiration
- 2041-11-03
AI Technical Summary
The prior art regards behavior recognition tasks and object detection tasks as independent tasks, failing to make full use of the relationship between the two, resulting in insufficient performance.
A single neural network model is used to simultaneously extract the time and spatial features required for behavior recognition and object detection tasks, and reduce the noise impact of spatial features through long and short-term spatial variation convolution, and use the graph network model to analyze the interaction process between people and objects to correct the features.
Through multi-task learning and feature correction, the performance of behavior recognition and object detection tasks is improved, higher quality spatial and temporal characteristics are provided, and the effect of video analysis is improved.
Smart Images

Figure CN114120440B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical fields of video analysis and machine learning, and in particular, to a behavior recognition method and an object detection method. Background Art
[0002] Video analysis tasks are responsible for helping a computer achieve information perception and logical acquisition of an input video. As the cornerstone tasks of video analysis, behavior recognition tasks and object detection tasks have always been in the hot research fields. Among them, the behavior recognition task is committed to understanding human behaviors by analyzing temporal information and spatial information in a video sequence; the object detection task is responsible for identifying objects that appear in a video and obtaining their position information and categories. However, existing research and inventions regard behavior recognition tasks and object detection tasks as two independent tasks for separate analysis, ignoring the connection between the two tasks and how to use this connection to improve the performance of behavior recognition tasks and object detection tasks. Summary of the Invention
[0003] In order to overcome the deficiencies of the prior art and make full use of the connection between behavior recognition tasks and object detection tasks during the interaction between humans and objects to improve the performance of the two tasks, the present invention provides a multi-task method for behavior recognition and object detection that integrates the analysis of human-object interaction.
[0004] The object of the present invention can be achieved by the following technical solutions:
[0005] A multi-task method for behavior recognition and object detection that integrates the analysis of human-object interaction, the method includes the following steps:
[0006] Step S1: Simultaneously extract the temporal features and spatial features required for behavior recognition tasks and object detection tasks based on a single neural network model;
[0007] Step S2: Use long-short-term spatio-temporal variant convolution to process the spatial features of adjacent frames as an effective supplement to the spatial features of the current frame to reduce the influence of the spatial features of the current frame by situations such as motion blur, defocus, rare perspectives, and partial occlusion;
[0008] Step S3: Use a graph network model to analyze the interaction process between humans and objects to achieve the correction of features related to behavior recognition and object detection;
[0009] Step S4: Analyze the corrected features through sub-neural networks corresponding to behavior recognition tasks and object detection tasks respectively to obtain behavior recognition results and object detection results that integrate the analysis of human-object interaction.
[0010] Further, the process of step S1 is as follows:
[0011] A single neural network model is constructed to be responsible for extracting the features required for the action recognition task and the object detection task. The model contains two different paths. Among them, the path responsible for extracting spatial features has sparse input frames. In this path, the size of the feature map is large and it does not contain time downsampling operations. The path responsible for extracting temporal features has a high refresh rate of input frames and contains time downsampling operations. The residual structures of the two feature extraction paths both adopt time-aligned residual operations, which are:
[0012]
[0013] where f 0 represents the feature output by the shortcut path of the residual structure, f n represents the source feature before connection, T n is the time range corresponding to the source feature, T is the time range corresponding to the feature output by the shortcut path, and n is the number of all source features within the time range.
[0014] Furthermore, the process of step S2 is as follows:
[0015] Construct a long-short-term spatially variant convolution kernel generation model. Subsequently, use this model to generate convolution kernels with different values at different spatial positions. Then, perform spatially variant convolution operations on the spatial features of adjacent frames to achieve effective correction of the spatial features of adjacent frames, so that the spatial features of adjacent frames can be used as an effective supplement to the spatial features of the current frame. The convolution kernel generation module contains two branches. The inputs of the branches are the temporal features of different levels output in step S1, corresponding to short-term temporal information and long-term temporal information respectively. Concatenate the outputs of the two branches and continue to perform information abstraction using a convolutional neural network. Finally, complete the normalization operation through the softmax layer. The detailed calculation process of the spatially variant convolution is:
[0016]
[0017] where, F s (x, y) is the value of the adjacent frame feature map at the (x, y) position, F p (x, y) is the value of the feature map generated after spatially variant convolution at the (x, y) position, W xy is the value of the convolution kernel at the (x, y) position, and K represents the size of the convolution kernel W xy .
[0018] Even further, the process of step S3 is as follows:
[0019] The design graph neural network analyzes the interaction between humans and objects through unsupervised learning, and uses this to correct the relevant features of action recognition and object detection tasks, achieving the use of the clues generated during the interaction between humans and objects to guide the action recognition task and the object detection task; the graph neural network consists of three parts: graph initialization, message function, and update function. Regarding graph initialization, first, the nodes in the graph need to be defined, and then each node is initialized by the corresponding features extracted respectively; in each iteration process, the calculation formula for summarizing the input information is:
[0020]
[0021] where is the input information summarized by node i at the s-th iteration, is the hidden state of node i, and the node connection metric W i,j represents the weight of the information flow between nodes, f() represents the message function, and the calculation process of f() is:
[0022]
[0023] This process is to input the hidden state of node i into the fully connected layer characterized by parameters θ and φ for processing, where [,] represents the concatenation operation; using a recurrent neural network as the update function, the hidden state of the node is updated according to the input message, and the update function is:
[0024]
[0025] where is the hidden state of the node, represents the input message.
[0026] The process of step S4 is as follows:
[0027] Sub-neural networks are constructed respectively according to the task characteristics of action recognition and object detection. Considering that action recognition is essentially a multi-classification problem, the sub-neural network for the action recognition task is built based on a dual-path structure similar to that in step S1. The two paths are respectively responsible for processing the temporal features and spatial features output by step S3, and the temporal feature map can be converted to the size of the spatial feature map using the temporal stride convolution method after a certain layer of convolution operation, so as to fuse the temporal features into the spatial features. Finally, global average pooling operations are performed on the outputs of each path, and the feature vectors of the two paths are concatenated and then input into the fully connected layer. The output of the fully connected layer is the result of action recognition; the object detection task is to process the spatial features output by step S3 through a static object detector to obtain the position information and category information of the object.
[0028] The beneficial effects of the present invention are manifested in:
[0029] Based on the idea of multi-task learning, the present invention uses a single neural network model to simultaneously extract the features required for action recognition and object detection tasks, which helps to strengthen the inductive bias of the extracted features, thereby improving the performance of the action recognition task and the object detection task; the dual-path structure (one path has a high input frame refresh rate and a low number of features, and the other path has a low input frame refresh rate and a high number of features) helps to decouple the spatial information and temporal information of the video signal, providing high-quality spatial features and temporal features for subsequent video analysis tasks; the graph convolutional neural network can achieve unsupervised analysis of the relationships between input features and perform feature correction based on these relationships. Brief Description of the Drawings
[0030] Figure 1 is a flowchart of a multi-task system for action recognition and object detection that integrates human-object interaction analysis according to the present invention;
[0031] Figure 2 is an illustration of the spatio-temporal variant convolution of long and short time in the present invention;
[0032] Figure 3 is an illustration of the graph convolutional neural network in the present invention. Detailed Embodiments
[0033] Next, the technical solutions in the method of the present invention will be clearly and completely described in conjunction with the accompanying drawings. Obviously, the described embodiments are part of the embodiments of the present invention, rather than all embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the scope of protection of the present invention.
[0034] Refer to Figures 1 to 3 , a multi-task method for action recognition and object detection that integrates human-object interaction analysis, the method includes the following steps:
[0035] Step S1: Simultaneously extract the temporal features and spatial features required for the action recognition task and the object detection task according to a single neural network model;
[0036] The process of step S1 is as follows:
[0037] A single neural network model is constructed to be responsible for extracting the features required for the action recognition task and the object detection task. This model contains two different paths. Among them, the path responsible for extracting spatial features has sparse input frames and a large feature map, and does not include temporal downsampling operations; the path responsible for extracting temporal features has a high refresh rate of input frames and includes temporal downsampling operations. The residual structures of the two feature extraction paths both use time-aligned residual operations, which are:
[0038]
[0039] where f 0 represents the feature formed after connection in the residual structure, and f n represents the source feature before connection, and T n is the time range corresponding to the source feature, T is the time range corresponding to the feature after connection, and n is the number of all features within the time range.
[0040] Step S2: Use long - short - term spatio - variant convolution to process the spatial features of adjacent frames as an effective supplement to the spatial features of the current frame, so as to reduce the influence of the spatial features of the current frame caused by motion blur, defocus, rare viewpoints, local occlusion, etc.;
[0041] The process of the above - mentioned step S2 is as follows:
[0042] Construct a long - short - term spatio - variant convolution kernel generation model, then use this model to generate convolution kernels with different values at different spatial positions, and then perform spatio - variant convolution operations on the spatial features of adjacent frames to achieve effective correction of the spatial features of adjacent frames, thereby enhancing the information supplement effect of the spatial features of adjacent frames on the spatial features of the current frame; the convolution kernel generation module includes two branches, and the inputs of the branches are the time features of different levels output in step S1, corresponding to short - term time information and long - term time information respectively. Concatenate the outputs of the two branches, and continue to perform information abstraction using a convolutional neural network, and finally complete the normalization operation through a softmax layer. The detailed calculation process of the spatio - variant convolution is as follows:
[0043]
[0044] where, F s (x, y) is the value of the adjacent - frame feature map at the position (x, y), and F p (x, y) is the value of the feature map generated after spatio - variant convolution at the position (x, y), W xy is the value of the convolution kernel at the position (x, y), and K represents the size of the convolution kernel W xy
[0045] Step S3: Use a graph network model to analyze the interaction process between people and objects, so as to achieve the correction of the features related to behavior recognition and target detection;
[0046] The process of the above - mentioned step S3 is as follows:
[0047] The design graph neural network analyzes the interaction between humans and objects through unsupervised learning, and uses this to correct the relevant features of the behavior recognition and object detection tasks, achieving the use of the clues generated during the interaction between humans and objects to guide the behavior recognition task and the object detection task; the graph neural network consists of three parts: graph initialization, message function, and update function. Regarding graph initialization, first, the nodes in the graph need to be defined, and then each node is initialized by the corresponding features extracted respectively; in each iteration process, the calculation formula for summarizing the input information is:
[0048]
[0049] Among them is the input information summarized by node i at the s-th iteration, is the hidden state of node i, and the node connection metric W i,j represents the weight of the information flow between nodes, f() represents the message function, and the calculation process of f() is:
[0050]
[0051] This process is to input the hidden state of node i into the fully connected layer characterized by parameters θ and φ for processing, where [,] represents the concatenation operation; using a recurrent neural network as the update function, the hidden state of the node is updated according to the input message, and the update function is:
[0052]
[0053] Among them is the hidden state of the node, represents the input message.
[0054] Step S4: Analyze the corrected features through the sub-neural networks corresponding to the behavior recognition task and the object detection task respectively, and obtain the behavior recognition result and object detection result that integrate the analysis of the interaction between humans and objects;
[0055] The process of step S4 is as follows:
[0056] Sub-neural networks are constructed separately according to the respective task characteristics of behavior recognition and object detection. Considering that behavior recognition is essentially a multi-classification problem, the sub-neural network for the behavior recognition task is built based on a dual-path structure similar to that in step S1. The two paths are respectively responsible for processing temporal features and spatial features, and after a certain layer of convolution operation, the method of temporal stride convolution can be used to convert the temporal feature map to the size of the spatial feature map, so as to fuse the temporal features into the spatial features. Finally, global average pooling operations are performed on the outputs of each path, and the feature vectors of the two paths are concatenated and then input into the fully connected layer, and the output of the fully connected layer is the result of behavior recognition; the object detection task is to process the corrected spatial features through a static object detector.
[0057] In this embodiment, in step S1, the constructed single neural network model includes two different feature extraction paths, and each path includes 4 convolutional neural network modules (Res 1-Res 4). Among them, the path responsible for extracting spatial features has only 4 input frames, and the distribution is relatively sparse, with a large number of channels, which is 8 times that of the path for extracting temporal features; the path responsible for extracting temporal features has a higher frame refresh rate, with 33 frames, and includes a temporal downsampling operation. The residual structures of the two feature extraction paths both adopt time-aligned residual operations, specifically:
[0058]
[0059] where f 0 represents the feature formed after connection in the residual structure, f n represents the source feature before connection, T n is the time range corresponding to the source feature, T is the time range corresponding to the feature after connection, and n is the number of all features within the time range.
[0060] In step 2, as Figure 2 shown, the long- and short-term spatially variant convolution kernel generation model includes two branches, corresponding to short-term temporal information and long-term temporal information respectively. Each branch is composed of 4 convolutional layers and ReLU layers intertwined, and the input is the features at different levels output by the path responsible for extracting temporal features in step S1. We concatenate the outputs of the two branches, then perform 2-layer convolution operations, and finally normalize them using the softmax layer to achieve the generation of long- and short-term spatially variant convolution kernels.
[0061] In step S3, first, the spatial feature maps of adjacent processed frames and the spatial feature map of the current frame are added together as a new spatial feature map, which is used as the input of the graph neural network together with the temporal features. As Figure 3 shown, the graph neural network includes three parts: graph initialization, message function, and update function.
[0062] Regarding graph initialization, three graph nodes (behavior node, interaction object node, and environmental object node) are first defined, and then each node is initialized by its corresponding input features;
[0063] In each iteration step, the calculation formula for summarizing the input information is:
[0064]
[0065] Where is the input information summarized by node i at the s-th iteration, is the hidden state of node i, and the node connection index W i,j represents the weight of the information flow between nodes, and f() represents the message function. The calculation process of f() used in the present invention is:
[0066]
[0067] This process is to input the hidden state of node i into the fully connected layer characterized by parameters θ and φ for processing, where [,] represents the concatenation operation;
[0068] Furthermore, a Gated Recurrent Unit is used as the update function to update the hidden state of the node according to the input message, and the update function is:
[0069]
[0070] Where is the hidden state of the node, represents the input message.
[0071] In step S4, two task-specific sub-neural networks are used to obtain the behavior recognition result and the target detection result respectively. The sub-neural network specific to the behavior recognition task is built based on a dual-path structure, including two convolutional neural network modules Re5 and Re6, and the time convolution method is used to connect the time features to the spatial features (this operation exists after the operation of the Re4 and Re5 convolutional neural network modules), and then the global average pooling operation is performed on the output of each path. After concatenating the feature vectors of the two paths, they are input into the fully connected layer; the sub-neural network specific to the target detection task consists of a part of the ResNet-101 network and R-FCN. The part of the ResNet-101 network further abstracts and reduces the dimensions of the features, and the static target detector R-FCN is responsible for outputting the position information and category information of the object.
[0072] Based on the above method, the present invention is verified on the Kinetics (action recognition), Charades (action recognition), and ImageNetVID (video object detection) datasets. The results are shown in Table 1, which is a performance comparison table between the present invention and other methods on the Kinetics dataset, Table 2, which is a performance comparison table between the present invention and other methods on the Charades dataset, and Table 3, which is a performance comparison table between the present invention and other methods on the ImageNet VID dataset. It can be seen from this that compared with other methods, our invention has obvious performance advantages. Regarding the improvement effects of human-object interaction analysis and long-short-term spatio-variant convolution on the action recognition task and the object detection task, sufficient ablation experiments have also been carried out on the three datasets, and the results are shown in Table 4, which is the ablation experiment on the Kinetics, Charades, and ImageNet VID datasets.
[0073]
[0074] Table 1
[0075] Name Pre-trained dataset mAP (%) CoViAR, R-50 ImageNet 21.9 Asyn-TF, VGG16 ImageNet 22.4 MultiScale TRN ImageNet 25.2 Nonlocal, R101 ImageNet + Kinetics-400 37.5 STRG, R101+NL ImageNet + Kinetics-400 39.7 SlowFast Kinetics-400 42.1 The present invention Kinetics-400 42.3
[0076] Table 2
[0077]
[0078] Table 3
[0079]
[0080] Table 4
[0081] The multi-task method for action recognition and object detection that fuses human-object interaction analysis in this embodiment constructs a single neural network model to simultaneously extract the temporal features and spatial features required for the action recognition task and the object detection task, designs long-short-term spatio-variant convolution to process the spatial features of adjacent frames as an effective supplement to the spatial features of the current frame, and constructs a graph network model to analyze the interaction clues between humans and objects, so as to realize the correction of the features related to action recognition and the features related to object detection. Finally, sub-neural network models are built for the action recognition task and the object detection task respectively to obtain the action recognition and object detection results that fuse human-object interaction analysis.
[0082] As described above, the above are only the specific embodiments of the present invention, but the protection scope of the present invention is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present invention can easily think of various equivalent modifications or substitutions, and these modifications or substitutions should all be covered within the protection scope of the present invention. Therefore, the protection scope of the present invention should be subject to the protection scope of the claims.
Claims
1. A multi - task method for behavior recognition and object detection that integrates human - object interaction analysis, characterized in that, the method comprises the following steps: Step S1: Simultaneously extract the temporal features and spatial features required for video behavior recognition tasks and object detection tasks based on a single neural network model; Step S2: Use long - short - term spatio - variant convolutions to process the spatial features of adjacent frames as an effective supplement to the spatial features of the current frame, so as to reduce the influence of motion blur, defocus, rare viewpoints, and local occlusion on the spatial features of the current frame; Step S3: Use a graph network model to analyze the interaction process between humans and objects, thereby realizing the correction of features related to behavior recognition and object detection; Step S4: Analyze the corrected features through sub - neural networks corresponding to the behavior recognition task and the object detection task respectively, and obtain the behavior recognition result and object detection result that integrate human - object interaction analysis; The process of the above - mentioned Step S1 is: Construct a single neural network model responsible for extracting the features required for behavior recognition tasks and object detection tasks. This model contains two different paths. Among them, the path responsible for extracting spatial features has a sparse input frame, the size of the feature map in this path is larger, and it does not contain temporal down - sampling operations; the path responsible for extracting temporal features has a higher input frame refresh rate and contains temporal down - sampling operations. The residual structures of the two feature extraction paths both adopt temporally aligned residual operations, which are: Among them, f 0 represents the feature formed after connection in the residual structure, and f n represents the source feature before connection, T n is the time range corresponding to the source feature, T is the time range corresponding to the feature after connection, and n is the number of all features within the time range.
2. The multi - task method for behavior recognition and object detection that integrates human - object interaction analysis according to claim 1, characterized in that, the process of the above - mentioned Step S2 is: Construct a long - short - term spatio - variant convolution kernel generation model, and then use this model to generate convolution kernels with different values at different spatial positions. Then perform spatio - variant convolution operations on the spatial features of adjacent frames to achieve effective correction of the spatial features of adjacent frames, so that the spatial features of adjacent frames can be used as an effective supplement to the spatial features of the current frame; the convolution kernel generation module contains two branches. The inputs of the branches are the temporal features of different levels output in Step S1, corresponding to short - term temporal information and long - term temporal information respectively. Concatenate the outputs of the two branches, and continue to perform information abstraction using a convolutional neural network. Finally, complete the normalization operation through a softmax layer. The detailed calculation process of the spatio - variant convolution is: Among them, F s (x, y) is the value of the adjacent frame feature map at the (x, y) position, and F p (x, y) is the value of the feature map generated after spatial variant convolution at the (x, y) position, and W xy is the value of the convolution kernel at the (x, y) position, and K represents the size of the convolution kernel W xy of.
3. The multi - task method for behavior recognition and object detection that integrates human - object interaction analysis according to claim 1, characterized in that, the process of the above - mentioned Step S3 is: Design a graph neural network to analyze the interaction between humans and objects through unsupervised learning, and use this to correct the features related to behavior recognition and object detection tasks, so as to realize using the clues generated by the human - object interaction process to guide the behavior recognition task and the object detection task; the graph neural network contains three parts: graph initialization, message function, and update function. Regarding graph initialization, first, it is necessary to define the nodes in the graph, and then each node is initialized by the corresponding extracted features; In each iteration process, the calculation formula for summarizing the input information is: Among which F i s is the input information aggregated by node i at the s-th iteration, is the hidden state of node i, and the node connection metric W i,j represents the weight of the information flow between nodes, and f() represents the message function. The calculation process of f() is as follows: The process is to input the hidden state of node i into a fully connected layer characterized by parameters θ and φ for processing, where [,] represents the concatenation operation; use a recurrent neural network as the update function to update the hidden state of the node according to the input message, and the update function is: Among them is the hidden state of the node, F i s represents the input message.
4. The multi-task method for behavior recognition and object detection integrating human-object interaction analysis according to claim 1, characterized in that, the process of step S4 is as follows: Sub-neural networks are respectively constructed according to the task characteristics of behavior recognition and object detection. Considering that behavior recognition is essentially a multi-classification problem, the sub-neural network for the behavior recognition task is built based on the same dual-path structure as in step S1. The two paths are respectively responsible for processing the temporal features and spatial features output by step S3, and after a certain layer of convolution operation, the method of temporal stride convolution is used to convert the temporal feature map to the size of the spatial feature map, so as to fuse the temporal features into the spatial features. Finally, global average pooling operations are performed on the outputs of each path, and the feature vectors of the two paths are concatenated and then input into the fully connected layer, and the output of the fully connected layer is the result of behavior recognition; For the object detection task, the spatial features output by step S3 are processed by a static object detector to obtain the position information and category information of the object.
Citation Information
Patent Citations
Sparse low-rank structure based multi-task learning behavior identification method
CN105809119A
Image classification system based on channel importance pruning and binary quantization
CN113177580A