A method and device for behavior recognition with enhanced dynamic information

By introducing dynamic feature network branches into the behavior recognition model, and using differential computing and feature encoding technology to generate dynamic feature representations for behavior recognition, the problem of poor action recognition effect in the existing technology is solved, and higher accuracy and real-timeness are achieved.

CN114120445BActive Publication Date: 2025-06-10BEIJING YIDA TURING TECH CO LTD +2
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202111371379.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-11-18
Publication Date
2025-06-10
Estimated Expiration
2041-11-18

AI Technical Summary

Technical Problem

In the prior art, 3D convolutional neural networks are difficult to effectively learn dynamic information of behavior in videos, resulting in poorer recognition of continuous changes in action.

Method used

By introducing dynamic feature network branches into the behavior recognition model, differential operations are performed on the apparent feature map sequence in the image sequence to obtain the dynamic feature map sequence, and feature encoding is performed in combination with the apparent feature map sequence to generate dynamic feature representations for behavior recognition.

Benefits of technology

It improves the accuracy of behavior recognition, reduces the calculation amount of dynamic information extraction, enhances the real-time nature of behavior recognition, and has higher application value.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114120445B_ABST
    Figure CN114120445B_ABST
Patent Text Reader

Abstract

The present invention provides a method and apparatus for behavior recognition with enhanced dynamic information. The method includes: determining an image sequence of a video to be recognized; inputting the image sequence into a behavior recognition model to obtain a behavior recognition result output by the behavior recognition model, where the behavior recognition model is trained based on a sample image sequence and a sample behavior recognition result of a sample video; wherein, the behavior recognition model is used to extract features from each frame image in the image sequence to obtain an appearance feature map sequence, perform a difference operation on every two adjacent appearance feature maps in the appearance feature map sequence to obtain a dynamic feature map sequence, perform feature encoding on the dynamic feature map sequence and the appearance feature map sequence to obtain a dynamic feature representation, and perform behavior recognition based on the dynamic feature representation. The method, apparatus, electronic device and storage medium provided by the present invention improve the accuracy and real-time performance of behavior recognition while having higher application value.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of computer vision technology, and in particular, to a behavior recognition method and device with enhanced dynamic information. Background Art

[0002] Human behavior recognition is a hot topic in computer vision research, which requires automatically recognizing the ongoing behavior from a video segment or an image sequence. The application scope of human behavior recognition is very wide, mainly concentrated in the field of intelligent video surveillance, including intelligent behavior analysis and management, intelligent transportation, human-computer interaction, intelligent security, etc. In the face of massive data, computer and machine learning are needed to automatically analyze the data.

[0003] Currently, most of the behavior recognition methods using video as input adopt 3D (3 dimensions, three-dimensional) convolutional neural networks. The advantage of 3D convolutional neural networks is that the convolutional kernel can perform filtering in the time dimension, so it can learn simultaneously in the spatial dimension (width, height) and the time dimension. However, although 3D convolutional neural networks can capture the relationship between adjacent moments within a certain range, they cannot learn the dynamic information of the behavior in the video, resulting in poor recognition effects for continuously changing actions. Summary of the Invention

[0004] The present invention provides a behavior recognition method and device with enhanced dynamic information to solve the defect of poor recognition effects for continuously changing actions in the prior art and achieve an improvement in the accuracy of behavior recognition.

[0005] The present invention provides a behavior recognition method with enhanced dynamic information, including:

[0006] Determine the image sequence of the video to be recognized;

[0007] Input the image sequence into a behavior recognition model to obtain a behavior recognition result output by the behavior recognition model. The behavior recognition model is trained based on the sample image sequence and sample behavior recognition result of a sample video;

[0008] Wherein, the behavior recognition model is used to extract features from each frame image in the image sequence to obtain a sequence of apparent feature maps, perform a difference operation on every two adjacent apparent feature maps in the sequence of apparent feature maps to obtain a sequence of dynamic feature maps, perform feature encoding on the sequence of dynamic feature maps and the sequence of apparent feature maps to obtain a dynamic feature representation, and perform behavior recognition based on the dynamic feature representation.

[0009] According to a behavior recognition method with enhanced dynamic information provided by the present invention, the behavior recognition model includes a first 2D convolutional layer, a dynamic feature network branch, an apparent feature network branch, and a behavior recognition layer;

[0010] Inputting the image sequence into the behavior recognition model to obtain the behavior recognition result output by the behavior recognition model includes:

[0011] Inputting the image sequence into the first 2D convolutional layer to obtain the apparent feature map sequence output by the first 2D convolutional layer;

[0012] Inputting the apparent feature map sequence into the dynamic feature network branch to obtain the dynamic feature representation output by the dynamic feature network branch;

[0013] Inputting the apparent feature map sequence into the apparent feature network branch to obtain the apparent feature representation output by the apparent feature network branch;

[0014] Inputting the dynamic feature representation and the apparent feature representation into the behavior recognition layer to obtain the behavior recognition result output by the behavior recognition layer.

[0015] According to a behavior recognition method with enhanced dynamic information provided by the present invention, the dynamic feature network branch includes a difference operation layer, a sequence fusion layer, and a 3D convolutional layer;

[0016] Inputting the apparent feature map sequence into the dynamic feature network branch to obtain the dynamic feature representation output by the dynamic feature network branch includes:

[0017] Inputting the apparent feature map sequence into the difference operation layer to obtain the dynamic feature map sequence output by the difference operation layer;

[0018] Inputting the apparent feature map sequence and the dynamic feature map sequence into the sequence fusion layer to obtain the sequence fusion result output by the sequence fusion layer;

[0019] Inputting the sequence fusion result into the 3D convolutional layer to obtain the dynamic feature representation output by the 3D convolutional layer.

[0020] According to a behavior recognition method with enhanced dynamic information provided by the present invention, the difference operation layer is used to perform a difference operation on each apparent feature map except the last one in the apparent feature map sequence with its adjacent next apparent feature map, and perform a difference operation on the last apparent feature map and the first apparent feature map to obtain the dynamic feature map sequence; the sequence fusion layer is used to splice the apparent feature map sequence and the dynamic feature map sequence along the channel direction to obtain the sequence fusion result.

[0021] A behavior recognition method with enhanced dynamic information according to the present invention, wherein the appearance feature network branch includes a second 2D convolutional layer and an average pooling layer;

[0022] Inputting the sequence of appearance feature maps into the appearance feature network branch to obtain the appearance feature representation output by the appearance feature network branch includes:

[0023] Inputting the sequence of appearance feature maps into the second 2D convolutional layer to obtain a sequence of appearance-enhanced feature maps output by the second 2D convolutional layer;

[0024] Inputting the sequence of appearance-enhanced feature maps into the average pooling layer to obtain the appearance feature representation output by the average pooling layer.

[0025] A behavior recognition method with enhanced dynamic information according to the present invention, wherein inputting the dynamic feature representation and the appearance feature representation into the behavior recognition layer to obtain the behavior recognition result output by the behavior recognition layer includes:

[0026] Inputting the dynamic feature representation and the appearance feature representation into the behavior recognition layer, and the behavior recognition layer fuses the dynamic feature representation and the appearance feature representation to obtain a feature fusion result, and performs a linear transformation on the feature fusion result to obtain the behavior recognition result output after the linear transformation by the behavior recognition layer.

[0027] The present invention also provides a behavior recognition device with enhanced dynamic information, including:

[0028] A determination module for determining an image sequence of a video to be recognized;

[0029] A recognition module for inputting the image sequence into a behavior recognition model to obtain a behavior recognition result output by the behavior recognition model, and the behavior recognition model is trained based on a sample image sequence and a sample behavior recognition result of a sample video;

[0030] Wherein, the behavior recognition model is used to extract features from each frame image in the image sequence to obtain a sequence of appearance feature maps, perform a difference operation on every two adjacent appearance feature maps in the sequence of appearance feature maps to obtain a sequence of dynamic feature maps, perform feature encoding on the sequence of dynamic feature maps and the sequence of appearance feature maps to obtain a dynamic feature representation, and perform behavior recognition based on the dynamic feature representation.

[0031] The present invention also provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor, and when the processor executes the program, the steps of the behavior recognition method with enhanced dynamic information as described in any one of the above are implemented.

[0032] The present invention also provides a non-transitory computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps of the behavior recognition method with dynamic information enhancement as described in any one of the above are implemented.

[0033] The present invention also provides a computer program product, including a computer program. When the computer program is executed by a processor, the steps of the behavior recognition method with dynamic information enhancement as described in any one of the above are implemented.

[0034] A behavior recognition method and device with dynamic information enhancement provided by the present invention extract features from each frame image in an image sequence through a behavior recognition model, and perform a difference operation on every two adjacent apparent feature maps in the obtained sequence of apparent feature maps to obtain a sequence of dynamic feature maps, greatly reducing the computational amount of dynamic information extraction. On this basis, behavior recognition is performed based on the dynamic feature representation determined by the sequence of apparent feature maps and the sequence of dynamic feature maps, improving both the accuracy and real-time performance of behavior recognition, and making the behavior recognition method have higher application value. BRIEF DESCRIPTION OF THE DRAWINGS

[0035] In order to more clearly illustrate the technical solutions in the present invention or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0036] Figure 1 is a schematic flowchart of the behavior recognition method with dynamic information enhancement provided by the present invention;

[0037] Figure 2 is a schematic structural diagram of the behavior recognition model provided by the present invention;

[0038] Figure 3 is a schematic structural diagram of the behavior recognition device with dynamic information enhancement provided by the present invention;

[0039] Figure 4 is a schematic structural diagram of the electronic device provided by the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0040] To make the objectives, technical solutions, and advantages of the present invention clearer, the following will clearly and completely describe the technical solutions in the present invention with reference to the drawings in the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art without creative efforts based on the embodiments in the present invention belong to the scope of protection of the present invention.

[0041] An embodiment of the present invention provides a behavior recognition method with enhanced dynamic information. Figure 1 It is a schematic flowchart of the behavior recognition method with enhanced dynamic information provided by the present invention. As Figure 1 shown, the method includes:

[0042] Step 110, determine the image sequence of the video to be recognized.

[0043] Specifically, the video to be recognized is the video for which behavior recognition is required. Here, the video to be recognized can be a pre-shot and stored video or a real-time captured video stream. The embodiments of the present invention do not make specific limitations on this. The image sequence is obtained by sampling the video to be recognized. The image sequence contains multiple frames of images. Each frame of image is derived from the video to be recognized, and the multiple frames of images are arranged in the time order in the video to be recognized, thus forming an image sequence. Here, the specific way of sampling the video to be recognized can be to divide the video to be recognized into several video segments, and then extract one frame of image from each video segment to form the image sequence.

[0044] Step 120, input the image sequence into the behavior recognition model to obtain the behavior recognition result output by the behavior recognition model. The behavior recognition model is trained based on the sample image sequence and sample behavior recognition result of the sample video;

[0045] Among them, the behavior recognition model is used to extract features from each frame of image in the image sequence to obtain a sequence of appearance feature maps, perform differential operations on every two adjacent appearance feature maps in the sequence of appearance feature maps to obtain a sequence of dynamic feature maps, perform feature encoding on the sequence of dynamic feature maps and the sequence of appearance feature maps to obtain a dynamic feature representation, and perform behavior recognition based on the dynamic feature representation.

[0046] Specifically, after the image sequence is input into the behavior recognition model, the behavior recognition model can first extract features from each frame of image in the image sequence to obtain the appearance feature maps corresponding to each frame of image. The appearance feature maps corresponding to each frame of image are arranged in the time order of each frame of image, thus forming a sequence of appearance feature maps. Here, the appearance feature map is a feature map used to represent the appearance information of the corresponding image.

[0047] Next, considering that the existing behavior recognition methods using 3D convolutional neural networks cannot learn the dynamic information of behaviors in videos, resulting in poor recognition effects for continuously changing actions. In this regard, in the embodiments of the present invention, the behavior recognition model can perform differential operations on every two adjacent apparent feature maps in the sequence of apparent feature maps, thereby obtaining dynamic feature maps composed of the differences between every two adjacent apparent feature maps. Then, these dynamic feature maps are combined into a sequence of dynamic feature maps, and this sequence of dynamic feature maps is applied to behavior recognition, so as to improve the recognition effect of continuously changing actions.

[0048] It should be noted that although there are currently some behavior recognition methods that use additional modalities such as optical flow to supplement the extraction of dynamic information, this method requires a huge amount of computational power and storage space, and it is difficult to meet the real-time requirements in the actual recognition process. In the embodiments of the present invention, by directly performing differential operations on every two adjacent apparent feature maps in the sequence of apparent feature maps obtained by feature extraction, a sequence of dynamic feature maps is obtained, thereby utilizing the difference reflected by two consecutive apparent feature maps in the time dimension to obtain dynamic feature maps that can represent the dynamic information of behaviors in the video. The whole process is convenient and fast, greatly reducing the computational amount of dynamic information extraction and helping to improve the real-time performance of behavior recognition.

[0049] Subsequently, the behavior recognition model can perform feature encoding on the sequence of dynamic feature maps and the sequence of apparent feature maps, thereby obtaining a dynamic feature representation that simultaneously covers apparent information and dynamic information, and applying this dynamic feature representation to behavior classification, so as to obtain a behavior recognition result with high accuracy and low computational amount. Here, the specific behavior classification method can be to perform behavior classification only based on the dynamic feature representation, or to combine the dynamic feature representation with other features for behavior classification. The embodiments of the present invention do not make specific limitations in this regard. The behavior recognition result is used to indicate the behaviors existing in the video to be recognized and the specific behavior categories.

[0050] Before performing step 120, a behavior recognition model can also be pre-trained. Specifically, the behavior recognition model can be trained in the following way: First, a large number of sample videos are collected, the sample image sequences of the sample videos are extracted, and the sample behavior recognition results of the sample videos are obtained through manual annotation. Subsequently, the initial model is trained using the sample image sequences of the sample videos and the sample behavior recognition results, thereby obtaining the behavior recognition model.

[0051] The method provided by the embodiment of the present invention extracts features from each frame image in the image sequence through a behavior recognition model, and performs a difference operation on every two adjacent apparent feature maps in the obtained sequence of apparent feature maps to obtain a sequence of dynamic feature maps, greatly reducing the computational amount of dynamic information extraction. On this basis, behavior recognition is performed based on the dynamic feature representation determined by the sequence of apparent feature maps and the sequence of dynamic feature maps, improving the accuracy of behavior recognition while also enhancing the real-time performance of behavior recognition, making the behavior recognition method have higher application value.

[0052] Based on any of the above embodiments, the behavior recognition model includes a first 2D convolutional layer, a dynamic feature network branch, an apparent feature network branch, and a behavior recognition layer;

[0053] Step 120 includes:

[0054] Input the image sequence into the first 2D convolutional layer to obtain a sequence of apparent feature maps output by the first 2D convolutional layer;

[0055] Input the sequence of apparent feature maps into the dynamic feature network branch to obtain a dynamic feature representation output by the dynamic feature network branch;

[0056] Input the sequence of apparent feature maps into the apparent feature network branch to obtain an apparent feature representation output by the apparent feature network branch;

[0057] Input the dynamic feature representation and the apparent feature representation into the behavior recognition layer to obtain a behavior recognition result output by the behavior recognition layer.

[0058] Specifically, the behavior recognition model may include a first 2D convolutional layer, a dynamic feature network branch, an apparent feature network branch, and a behavior recognition layer. Among them, the first 2D convolutional layer may be a convolutional layer in a 2D convolutional neural network, which is used to extract features from each frame image in the input image sequence, form a sequence of apparent feature maps corresponding to each frame image extracted, and output the sequence of apparent feature maps.

[0059] The dynamic feature network branch is used to perform a difference operation on every two adjacent apparent feature maps in the input sequence of apparent feature maps to obtain a sequence of dynamic feature maps, and perform feature encoding on the sequence of dynamic feature maps and the sequence of apparent feature maps, so as to obtain a dynamic feature representation and output it to the behavior recognition layer. Here, the specific way of feature encoding can be to first fuse the sequence of dynamic feature maps and the sequence of apparent feature maps, and then perform feature encoding on the obtained fusion result to obtain a dynamic feature representation, or to first perform feature encoding on the sequence of dynamic feature maps and the sequence of apparent feature maps respectively, and then fuse the respectively obtained feature vectors to obtain a dynamic feature representation, or to first perform feature encoding on one of the two sequences, fuse the obtained intermediate variable with the other sequence, and then perform feature encoding on the fusion result to obtain a dynamic feature representation. The embodiments of the present invention do not make specific limitations on this.

[0060] The apparent feature network branch is used to perform feature encoding on the input sequence of apparent feature maps, so as to obtain an apparent feature representation and output it to the behavior recognition layer. The behavior recognition layer is used to respectively obtain the outputs of the two branches, namely the dynamic feature network branch and the apparent feature network branch, and thus combine the dynamic feature representation and the apparent feature representation to perform behavior classification on the video to be recognized, so as to obtain a behavior recognition result. It can be understood that this behavior recognition result is both the output of the behavior recognition layer and the final output of the entire behavior recognition model.

[0061] It should be noted that this behavior recognition method performs behavior recognition by combining a dynamic feature representation with strong dynamic feature expression ability and an apparent feature representation with strong apparent feature expression ability, thus accurately taking into account the apparent information of each frame image in the image sequence and the dynamic information during the change process of each frame image, and further greatly improving the accuracy of behavior recognition.

[0062] Based on any of the above embodiments, the dynamic feature network branch includes a difference operation layer, a sequence fusion layer, and a 3D convolutional layer;

[0063] Inputting the sequence of apparent feature maps into the dynamic feature network branch to obtain the dynamic feature representation output by the dynamic feature network branch includes:

[0064] Inputting the sequence of apparent feature maps into the difference operation layer to obtain a sequence of dynamic feature maps output by the difference operation layer;

[0065] Inputting the sequence of apparent feature maps and the sequence of dynamic feature maps into the sequence fusion layer to obtain a sequence fusion result output by the sequence fusion layer;

[0066] Inputting the sequence fusion result into the 3D convolutional layer to obtain the dynamic feature representation output by the 3D convolutional layer.

[0067] Specifically, the dynamic feature network branch includes a difference operation layer, a sequence fusion layer, and a 3D convolution layer. Among them, the difference operation layer is used to perform difference operations on every two adjacent apparent feature maps in the input sequence of apparent feature maps, so as to obtain a sequence of dynamic feature maps and output them to the sequence fusion layer. The sequence fusion layer is used to fuse the two input sequences, namely the sequence of apparent feature maps and the sequence of dynamic feature maps, so as to obtain a sequence fusion result that covers both apparent information and dynamic information. Here, the specific fusion method can be operations such as splicing, convolution, or addition of the two sequences, and the embodiments of the present invention do not make specific limitations on this.

[0068] In order to further learn temporal information and spatial information simultaneously, the sequence fusion result can be input into the 3D convolution layer, and the 3D convolution layer performs feature encoding on the sequence fusion result, so as to obtain a dynamic feature representation and output the dynamic feature representation. It can be understood that, based on the fact that the sequence fusion result covers both apparent information and dynamic information, inputting the sequence fusion result into the 3D convolution layer for further learning of temporal information and spatial information can learn a dynamic feature representation with stronger dynamic feature expression ability, realizing the enhancement of dynamic information. Here, the 3D convolution layer can be the convolution layer in a 3D convolutional neural network.

[0069] Based on any of the above embodiments, the difference operation layer is used to perform difference operations on each apparent feature map in the sequence of apparent feature maps except the last one with its adjacent next apparent feature map, and perform a difference operation on the last apparent feature map and the first apparent feature map, so as to obtain a sequence of dynamic feature maps; the sequence fusion layer is used to splice the sequence of apparent feature maps and the sequence of dynamic feature maps along the channel direction to obtain a sequence fusion result.

[0070] Specifically, the difference operation layer included in the dynamic feature network branch can be used for: for each apparent feature map in the sequence of apparent feature maps except the last one, a pixel-level difference operation can be performed between this apparent feature map and its adjacent next apparent feature map to extract the dynamic features at adjacent moments and obtain a dynamic feature map; for the last apparent feature map in the sequence of apparent feature maps, a pixel-level difference operation can be performed between this apparent feature map and the first apparent feature map to obtain a dynamic feature map; all the obtained dynamic feature maps are arranged in the order of the corresponding apparent feature maps to form a sequence of dynamic feature maps. The sequence fusion layer included in the dynamic feature network branch can be used to splice the sequence of apparent feature maps and the sequence of dynamic feature maps output by the first 2D convolution layer and the difference operation layer along the channel direction, so as to obtain a sequence fusion result that covers both apparent information and dynamic information, and output the sequence fusion result to the 3D convolution layer.

[0071] Based on any of the above embodiments, the appearance feature network branch includes a second 2D convolutional layer and an average pooling layer;

[0072] Inputting the sequence of appearance feature maps into the appearance feature network branch to obtain the appearance feature representation output by the appearance feature network branch, including:

[0073] Inputting the sequence of appearance feature maps into the second 2D convolutional layer to obtain a sequence of appearance-enhanced feature maps output by the second 2D convolutional layer;

[0074] Inputting the sequence of appearance-enhanced feature maps into the average pooling layer to obtain the appearance feature representation output by the average pooling layer.

[0075] Specifically, in order to further learn the appearance features of behaviors, the appearance feature network branch may include a second 2D convolutional layer, which is used to perform convolutional transformation and downsampling on the sequence of appearance feature maps output by the first 2D convolutional layer, so as to obtain a sequence of appearance-enhanced feature maps with stronger appearance feature expression ability. Here, the second 2D convolutional layer and the first 2D convolutional layer may be different convolutional layers in the same 2D convolutional neural network, or may belong to different 2D convolutional neural networks respectively. The embodiments of the present invention do not make specific limitations on this.

[0076] Then, the average pooling layer included in the appearance feature network branch performs average pooling operations on the sequence of appearance-enhanced feature maps along the spatial dimension and the temporal dimension, so as to obtain the appearance feature representation and output the appearance feature representation.

[0077] It should be noted that different from the existing behavior recognition method that simply uses a 3D convolutional neural network, since the increase in the dimension of the convolutional kernel increases the number of network parameters, this kind of network is more likely to have the problem of overfitting. Therefore, it is necessary to use a large-scale action dataset for pre-training to improve its generalization performance. In the embodiments of the present invention, a combination of 2D and 3D networks is used for behavior recognition. The appearance feature representation can be extracted only by a 2D convolutional neural network. Although the dynamic feature representation applies a 3D convolutional neural network, only the part of the structure in the 3D convolutional neural network that can realize feature encoding is required. Therefore, compared with the behavior recognition method that simply uses a 3D convolutional neural network, it can greatly reduce the computational amount, improve the computational efficiency, and reduce the occupied computational resources.

[0078] Based on any of the above embodiments, inputting the dynamic feature representation and the appearance feature representation into the behavior recognition layer to obtain the behavior recognition result output by the behavior recognition layer, including:

[0079] Input the dynamic feature representation and the appearance feature representation into the behavior recognition layer. The behavior recognition layer fuses the dynamic feature representation and the appearance feature representation to obtain a feature fusion result, and performs a linear transformation on the feature fusion result to obtain the behavior recognition result output after the linear transformation of the behavior recognition layer.

[0080] Specifically, input the dynamic feature representation output by the 3D convolutional layer and the appearance feature representation output by the average pooling layer into the behavior recognition layer. The behavior recognition layer fuses the dynamic feature representation and the appearance feature representation to obtain a feature fusion result that simultaneously covers appearance information and dynamic information. Here, the specific fusion method can be to perform a splicing process on the dynamic feature representation and the appearance feature representation.

[0081] On this basis, the behavior recognition layer performs a linear transformation on the feature fusion result, so as to obtain a high-precision behavior recognition result and output it as the result of the entire behavior recognition model. Here, the specific method of the linear transformation can be implemented through the fully connected layer included in the behavior recognition layer.

[0082] Based on any of the above embodiments, the image sequence of the video to be recognized can be obtained in the following specific way: First, divide the video to be recognized into several video segments, extract one frame of image from each video segment to form an initial image sequence. Then, perform normalization processing on each color channel in the image sequence to control the pixel values between (-1, 1). Next, perform cropping and stretching processing on the image sequence to form a fixed size to obtain the final image sequence as the input of the behavior recognition model.

[0083] Figure 2 is a schematic structural diagram of the behavior recognition model provided by the present invention. As Figure 2 shown, the behavior recognition model includes a first 2D convolutional layer, a differential operation layer (i.e., Figure 2 the dynamic enhancement module in), a sequence fusion layer, a 3D convolutional layer, a second 2D convolutional layer, an average pooling layer, and a behavior recognition layer. Based on this behavior recognition model, the specific process of the behavior recognition method provided by the embodiments of the present invention is as follows:

[0084] 1. Input the image sequence into the first 2D convolutional layer for feature extraction to form a sequence of appearance feature maps F with a size of T×H×W×C. Wherein, T represents the number of frames of images in the image sequence; H, W, and C respectively represent the height, width, and number of channels of the corresponding appearance feature maps. For example, T×H×W×C = 16×14×14×256.

[0085] 2. On the one hand, input the sequence of appearance feature maps extracted in step 1 into the dynamic enhancement module for dynamic information enhancement to obtain a sequence of dynamic feature maps F diffHere, the specific way of enhancing dynamic information can be to perform pixel-level difference operations between each apparent feature map and the adjacent next apparent feature map to extract the dynamic features at adjacent moments. For the last apparent feature map, a pixel-level difference operation is performed with the first apparent feature map.

[0086] 3. Input the sequence of dynamic feature maps obtained in step 2 and the sequence of apparent feature maps obtained in step 3 into the sequence fusion layer. The sequence fusion layer concatenates these two sequences along the channel dimension to obtain the sequence fusion result that combines apparent features and dynamic features, namely the enhanced feature map sequence F. con At this time, the size of the enhanced feature map sequence is T×H×W×(C + C).

[0087] 4. Input the enhanced feature map sequence obtained in step 3 into the 3D convolutional layer for further learning of temporal and spatial information, and perform average pooling operations on the learned sequence features of size T / 4×H / 2×W / 2×2C along the spatial and temporal dimensions to obtain a 2C-dimensional feature vector as the dynamic feature representation after further dynamic information enhancement.

[0088] 5. On the other hand, input the sequence of apparent feature maps obtained in step 1 into the second 2D convolutional layer for convolutional transformation and downsampling, and input the obtained sequence of apparent enhanced feature maps of size T×H / 2×W / 2×4C into the average pooling layer. The average pooling layer performs average pooling operations on the sequence of apparent enhanced feature maps along the spatial and temporal dimensions to obtain a 4C-dimensional feature vector as the apparent feature representation.

[0089] 6. Input the dynamic feature representation obtained in step 4 and the apparent feature representation obtained in step 5 into the behavior recognition layer. The behavior recognition layer performs concatenation processing on the dynamic feature representation and the apparent feature representation, and performs linear transformation on the concatenated feature representation based on the fully connected layer included in the behavior recognition layer, and then normalizes the output of the fully connected layer to obtain the probability distribution of the behavior in the video to be recognized over each behavior category. The behavior category with the highest probability is used as the behavior recognition result of the video to be recognized.

[0090] Before performing the above steps, a behavior recognition model can be pre-trained: Use the sample image sequences of a large number of collected sample videos and the sample behavior recognition results as the training sample set to train the initial model. In the training stage, use the cross-entropy loss function as the objective function and combine the stochastic gradient descent algorithm to train and optimize the network parameters of the initial model, thereby obtaining the behavior recognition model. Optionally, the momentum coefficient in the stochastic gradient descent algorithm is set to 0.9.

[0091] The training process of the specific behavior recognition model can refer to the above steps 1-6. It should be noted that in order to avoid overfitting problems, during the training stage, the behavior recognition layer will discard features of the dynamic feature representation and the appearance feature representation with a probability of 0.5, and splice the dynamic feature representation and the appearance feature representation after feature discarding. After linear transformation through the fully connected layer, the probability distribution of the sample behavior in each behavior category in the sample video is obtained; during the testing and application stages, the behavior recognition layer does not discard features of the dynamic feature representation and the appearance feature representation, but directly performs splicing processing and inputs it into the fully connected layer.

[0092] The embodiment of the present invention uses a two-stream neural network as the main structure and discloses a dynamic enhanced two-stream behavior recognition method. In the embodiment of the present invention, the network structures of steps 1-4 are used as the dynamic feature stream of the behavior recognition model, and the network structures of steps 1 and 5 are used as the appearance feature stream of the behavior recognition model, thus forming a two-stream neural network. By combining 2D convolutional neural network and 3D convolutional neural network for feature extraction at different levels, while reducing the number of parameters, the recognition accuracy of behaviors in videos is improved, which can play a key role in the application of real-time behavior recognition.

[0093] Based on the above embodiments, the behavior recognition method provided by the present invention only uses the image sequence of the video as the input, and can extract dynamic information through the dynamic enhancement module, that is, the differential operation layer, without the need for additional modalities such as optical flow, and combines the appearance features to accurately judge the behavior categories in the video. At the same time, the 2D+3D fusion network reduces the number of parameters and improves the calculation speed compared with the pure 3D convolutional neural network. Therefore, this technical solution has wide application value in the fields of intelligent transportation, smart city, intelligent monitoring, etc.

[0094] To illustrate the effectiveness of the method provided by the present invention, the Penn-Action dataset and the UCF11 dataset are used to verify the performance of the method. Penn-Action contains 2326 video segments and 15 behavior categories, such as bowling, clean and jerk, golf swing, etc. The annotation information of the Penn-Action dataset includes behavior category annotations. The UCF11 dataset contains more than 1600 videos of 11 behavior categories. During the video shooting process, there is a large degree of randomness, accompanied by changes in perspective, focal length, and camera movement. The action background is complex, and there are many interference factors such as blurring and occlusion, which is a very challenging behavior recognition task. The comparison results of the behavior recognition model with and without the dynamic enhancement module on the Penn-Action dataset and the UCF11 dataset are shown in Table 1:

[0095] Table 1

[0096] Method UCF11 Penn-Action Does not contain a dynamic enhancement module 92.0 94.3 Contains a dynamic enhancement module 94.3 96.4

[0097] The dynamic enhancement module proposed by the present invention effectively improves the performance of the behavior recognition model on two public action datasets, demonstrating the effectiveness of this module for video-based behavior recognition tasks.

[0098] The following describes the behavior recognition device with dynamic information enhancement provided by the present invention. The behavior recognition device with dynamic information enhancement described below can be correspondingly referred to the behavior recognition method with dynamic information enhancement described above.

[0099] Based on any of the above embodiments, Figure 3 is a schematic structural diagram of the behavior recognition device with dynamic information enhancement provided by the present invention. As Figure 3 shown, the device includes:

[0100] A determination module 310, configured to determine the image sequence of the video to be recognized;

[0101] A recognition module 320, configured to input the image sequence into the behavior recognition model to obtain the behavior recognition result output by the behavior recognition model. The behavior recognition model is trained based on the sample image sequence and the sample behavior recognition result of the sample video;

[0102] Wherein, the behavior recognition model is used to extract features from each frame image in the image sequence to obtain an apparent feature map sequence, perform a difference operation on every two adjacent apparent feature maps in the apparent feature map sequence to obtain a dynamic feature map sequence, perform feature encoding on the dynamic feature map sequence and the apparent feature map sequence to obtain a dynamic feature representation, and perform behavior recognition based on the dynamic feature representation.

[0103] The device provided by the embodiment of the present invention extracts features from each frame image in the image sequence through the behavior recognition model, and performs a difference operation on every two adjacent apparent feature maps in the obtained apparent feature map sequence to obtain a dynamic feature map sequence, greatly reducing the computational amount of dynamic information extraction. On this basis, behavior recognition is performed based on the dynamic feature representation determined by the apparent feature map sequence and the dynamic feature map sequence, improving the accuracy of behavior recognition while also improving the real-time performance of behavior recognition, making the behavior recognition method have higher application value.

[0104] Based on any of the above embodiments, the behavior recognition model includes a first 2D convolutional layer, a dynamic feature network branch, an apparent feature network branch, and a behavior recognition layer;

[0105] The recognition module 320 is specifically configured to:

[0106] Input the image sequence into the first 2D convolutional layer to obtain the apparent feature map sequence output by the first 2D convolutional layer;

[0107] Input the sequence of appearance feature maps into the dynamic feature network branch to obtain the dynamic feature representation output by the dynamic feature network branch;

[0108] Input the sequence of appearance feature maps into the appearance feature network branch to obtain the appearance feature representation output by the appearance feature network branch;

[0109] Input the dynamic feature representation and the appearance feature representation into the behavior recognition layer to obtain the behavior recognition result output by the behavior recognition layer.

[0110] Based on any of the above embodiments, the dynamic feature network branch includes a differential operation layer, a sequence fusion layer, and a 3D convolutional layer;

[0111] Inputting the sequence of appearance feature maps into the dynamic feature network branch to obtain the dynamic feature representation output by the dynamic feature network branch includes:

[0112] Input the sequence of appearance feature maps into the differential operation layer to obtain the sequence of dynamic feature maps output by the differential operation layer;

[0113] Input the sequence of appearance feature maps and the sequence of dynamic feature maps into the sequence fusion layer to obtain the sequence fusion result output by the sequence fusion layer;

[0114] Input the sequence fusion result into the 3D convolutional layer to obtain the dynamic feature representation output by the 3D convolutional layer.

[0115] Based on any of the above embodiments, the differential operation layer is used to perform a differential operation on each appearance feature map in the sequence of appearance feature maps except the last one with its adjacent next appearance feature map, and perform a differential operation on the last appearance feature map and the first appearance feature map to obtain the sequence of dynamic feature maps; the sequence fusion layer is used to splice the sequence of appearance feature maps and the sequence of dynamic feature maps along the channel direction to obtain the sequence fusion result.

[0116] Based on any of the above embodiments, the appearance feature network branch includes a second 2D convolutional layer and an average pooling layer;

[0117] Inputting the sequence of appearance feature maps into the appearance feature network branch to obtain the appearance feature representation output by the appearance feature network branch includes:

[0118] Input the sequence of appearance feature maps into the second 2D convolutional layer to obtain the sequence of appearance enhanced feature maps output by the second 2D convolutional layer;

[0119] Input the sequence of appearance enhanced feature maps into the average pooling layer to obtain the appearance feature representation output by the average pooling layer.

[0120] Based on any of the above embodiments, input the dynamic feature representation and the appearance feature representation into the behavior recognition layer to obtain the behavior recognition result output by the behavior recognition layer, including:

[0121] Input the dynamic feature representation and the appearance feature representation into the behavior recognition layer. The behavior recognition layer fuses the dynamic feature representation and the appearance feature representation to obtain a feature fusion result, and performs a linear transformation on the feature fusion result to obtain the behavior recognition result output after the linear transformation by the behavior recognition layer.

[0122] Figure 4 An example of the physical structure diagram of an electronic device is shown as Figure 4 shown. The electronic device may include: a processor 410, a communication interface 420, a memory 430, and a communication bus 440. Among them, the processor 410, the communication interface 420, and the memory 430 complete communication with each other through the communication bus 440. The processor 410 can call the logical instructions in the memory 430 to execute the behavior recognition method with enhanced dynamic information. The method includes: determining an image sequence of a video to be recognized; inputting the image sequence into a behavior recognition model to obtain the behavior recognition result output by the behavior recognition model. The behavior recognition model is trained based on the sample image sequence and the sample behavior recognition result of the sample video. Among them, the behavior recognition model is used to extract features from each frame image in the image sequence to obtain a sequence of appearance feature maps, perform a difference operation on every two adjacent appearance feature maps in the sequence of appearance feature maps to obtain a sequence of dynamic feature maps, perform feature encoding on the sequence of dynamic feature maps and the sequence of appearance feature maps to obtain a dynamic feature representation, and perform behavior recognition based on the dynamic feature representation.

[0123] In addition, when the logical instructions in the above-mentioned memory 430 are implemented in the form of software functional units and sold or used as an independent product, they can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or a part of this technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present invention. The foregoing storage medium includes: various media such as a USB flash drive, a mobile hard disk, a read-only memory (ROM, Read-Only Memory), a random access memory (RAM, Random Access Memory), a magnetic disk, or an optical disc that can store program codes.

[0124] On the other hand, the present invention also provides a computer program product, which includes a computer program. The computer program can be stored on a non-transitory computer-readable storage medium. When the computer program is executed by a processor, the computer can execute the dynamic information enhanced behavior recognition method provided by the above-mentioned various methods. The method includes: determining an image sequence of a video to be recognized; inputting the image sequence into a behavior recognition model to obtain a behavior recognition result output by the behavior recognition model, where the behavior recognition model is trained based on a sample image sequence of a sample video and a sample behavior recognition result; wherein, the behavior recognition model is used to extract features from each frame image in the image sequence to obtain an appearance feature map sequence, perform a difference operation on every two adjacent appearance feature maps in the appearance feature map sequence to obtain a dynamic feature map sequence, perform feature encoding on the dynamic feature map sequence and the appearance feature map sequence to obtain a dynamic feature representation, and perform behavior recognition based on the dynamic feature representation.

[0125] In another aspect, the present invention also provides a non-transitory computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it realizes the dynamic information enhanced behavior recognition method provided by the above-mentioned various methods. The method includes: determining an image sequence of a video to be recognized; inputting the image sequence into a behavior recognition model to obtain a behavior recognition result output by the behavior recognition model, where the behavior recognition model is trained based on a sample image sequence of a sample video and a sample behavior recognition result; wherein, the behavior recognition model is used to extract features from each frame image in the image sequence to obtain an appearance feature map sequence, perform a difference operation on every two adjacent appearance feature maps in the appearance feature map sequence to obtain a dynamic feature map sequence, perform feature encoding on the dynamic feature map sequence and the appearance feature map sequence to obtain a dynamic feature representation, and perform behavior recognition based on the dynamic feature representation.

[0126] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment. Those of ordinary skill in the art can understand and implement it without creative labor.

[0127] Through the description of the above embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus a necessary general hardware platform, and of course, it can also be implemented by hardware. Based on such an understanding, the essence of the above technical solution, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions to enable a computer device (which can be a personal computer, server, or network device, etc.) to execute the methods described in each embodiment or some parts of the embodiments.

[0128] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements for some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. A method for behavior recognition with enhanced dynamic information, characterized in that, it includes: Determine the image sequence of the video to be recognized; Input the image sequence into a behavior recognition model to obtain the behavior recognition result output by the behavior recognition model, where the behavior recognition model is trained based on the sample image sequence and sample behavior recognition result of the sample video; Among them, the behavior recognition model is used to extract features from each frame image in the image sequence to obtain an apparent feature map sequence, perform differential operations on every two adjacent apparent feature maps in the apparent feature map sequence to obtain a dynamic feature map sequence, perform feature encoding on the dynamic feature map sequence and the apparent feature map sequence to obtain a dynamic feature representation, and perform behavior recognition based on the dynamic feature representation; The behavior recognition model includes a first 2D convolutional layer, a dynamic feature network branch, an apparent feature network branch, and a behavior recognition layer; The step of inputting the image sequence into the behavior recognition model to obtain the behavior recognition result output by the behavior recognition model includes: Input the image sequence into the first 2D convolutional layer to obtain the apparent feature map sequence output by the first 2D convolutional layer; Input the apparent feature map sequence into the dynamic feature network branch to obtain the dynamic feature representation output by the dynamic feature network branch; Input the apparent feature map sequence into the apparent feature network branch to obtain the apparent feature representation output by the apparent feature network branch; Input the dynamic feature representation and the apparent feature representation into the behavior recognition layer to obtain the behavior recognition result output by the behavior recognition layer; The dynamic feature network branch includes a differential operation layer, a sequence fusion layer, and a 3D convolutional layer; The step of inputting the apparent feature map sequence into the dynamic feature network branch to obtain the dynamic feature representation output by the dynamic feature network branch includes: Input the apparent feature map sequence into the differential operation layer to obtain the dynamic feature map sequence output by the differential operation layer; Input the apparent feature map sequence and the dynamic feature map sequence into the sequence fusion layer to obtain the sequence fusion result output by the sequence fusion layer; Input the sequence fusion result into the 3D convolutional layer to obtain the dynamic feature representation output by the 3D convolutional layer.

2. The method for behavior recognition with enhanced dynamic information according to claim 1, characterized in that, The differential operation layer is used to perform differential operations on each apparent feature map in the apparent feature map sequence except the last one with its adjacent next apparent feature map, and perform a differential operation on the last apparent feature map and the first apparent feature map to obtain the dynamic feature map sequence; the sequence fusion layer is used to splice the apparent feature map sequence and the dynamic feature map sequence along the channel direction to obtain the sequence fusion result.

3. The method for behavior recognition with enhanced dynamic information according to claim 1, characterized in that, The apparent feature network branch includes a second 2D convolutional layer and an average pooling layer; Inputting the apparent feature map sequence into the apparent feature network branch to obtain the apparent feature representation output by the apparent feature network branch includes: Inputting the apparent feature map sequence into the second 2D convolutional layer to obtain the sequence of apparent enhancement feature maps output by the second 2D convolutional layer; Inputting the sequence of apparent enhancement feature maps into the average pooling layer to obtain the apparent feature representation output by the average pooling layer.

4. The dynamic information enhanced behavior recognition method according to claim 1, wherein, Inputting the dynamic feature representation and the apparent feature representation into the behavior recognition layer to obtain the behavior recognition result output by the behavior recognition layer includes: Inputting the dynamic feature representation and the apparent feature representation into the behavior recognition layer, and the behavior recognition layer fuses the dynamic feature representation and the apparent feature representation to obtain a feature fusion result, and performs a linear transformation on the feature fusion result to obtain the behavior recognition result output after the linear transformation of the behavior recognition layer.

5. A dynamic information enhanced behavior recognition device, wherein, comprising: a determination module for determining the image sequence of the video to be recognized; a recognition module for inputting the image sequence into a behavior recognition model to obtain the behavior recognition result output by the behavior recognition model, and the behavior recognition model is trained based on the sample image sequence and the sample behavior recognition result of the sample video; wherein, the behavior recognition model is used to extract features from each frame image in the image sequence to obtain an apparent feature map sequence, perform a difference operation on every two adjacent apparent feature maps in the apparent feature map sequence to obtain a dynamic feature map sequence, perform feature encoding on the dynamic feature map sequence and the apparent feature map sequence to obtain a dynamic feature representation, and perform behavior recognition based on the dynamic feature representation; the behavior recognition model includes a first 2D convolutional layer, a dynamic feature network branch, an apparent feature network branch, and a behavior recognition layer; Inputting the image sequence into the behavior recognition model to obtain the behavior recognition result output by the behavior recognition model includes: Inputting the image sequence into the first 2D convolutional layer to obtain the sequence of apparent feature maps output by the first 2D convolutional layer; Inputting the sequence of apparent feature maps into the dynamic feature network branch to obtain the dynamic feature representation output by the dynamic feature network branch; Inputting the sequence of apparent feature maps into the apparent feature network branch to obtain the apparent feature representation output by the apparent feature network branch; Inputting the dynamic feature representation and the apparent feature representation into the behavior recognition layer to obtain the behavior recognition result output by the behavior recognition layer; the dynamic feature network branch includes a difference operation layer, a sequence fusion layer, and a 3D convolutional layer; Inputting the sequence of apparent feature maps into the dynamic feature network branch to obtain the dynamic feature representation output by the dynamic feature network branch includes: Inputting the sequence of apparent feature maps into the difference operation layer to obtain the sequence of dynamic feature maps output by the difference operation layer; Input the apparent feature map sequence and the dynamic feature map sequence into the sequence fusion layer to obtain the sequence fusion result output by the sequence fusion layer; Input the sequence fusion result into the 3D convolutional layer to obtain the dynamic feature representation output by the 3D convolutional layer.

6. An electronic device, comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, wherein, when the processor executes the program, the steps of the dynamic information enhanced behavior recognition method according to any one of claims 1 to 4 are implemented.

7. A non-transitory computer-readable storage medium, on which a computer program is stored, wherein, when the computer program is executed by a processor, the steps of the dynamic information enhanced behavior recognition method according to any one of claims 1 to 4 are implemented.

8. A computer program product, comprising a computer program, wherein, when the computer program is executed by a processor, the steps of the dynamic information enhanced behavior recognition method according to any one of claims 1 to 4 are implemented.

Citation Information

Patent Citations

  • Video-based behavior recognition method and device, electronic equipment and storage medium

    CN111242068A