Behavior recognition method and device, equipment and medium

By extracting frames and keyframes on the video stream, combined with pedestrian detection and behavior recognition models, the problem of difficult to take into account the speed and accuracy of character behavior recognition in the video is solved, and efficient behavior recognition effect is achieved.

CN119942631APending Publication Date: 2025-05-06CHINA TELECOM CLOUD TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202411730450.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-11-28
Publication Date
2025-05-06

AI Technical Summary

Technical Problem

The prior art is difficult to take into account both the recognition accuracy and recognition speed when recognizing character behavior in videos, resulting in poor application results.

Method used

By extracting frames on the target video stream, obtaining image frame sequences and determining keyframes, calling the pedestrian detection model to determine whether pedestrians exist in the keyframe. If there is, the detection box information and spatiotemporal feature tensors are input to the behavior recognition model for processing to determine the behavior category.

Benefits of technology

The speed and accuracy of character behavior recognition in video streams are achieved, and the application effect of behavior recognition is improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119942631A_ABST
    Figure CN119942631A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of computer vision, and discloses a behavior recognition method and device, equipment and a medium, and the method comprises the steps: obtaining a target video stream, carrying out the frame extraction of the target video stream, obtaining an image frame sequence of the target video stream, and determining a key frame in the image frame sequence; preprocessing the image frame sequence to obtain a spatial-temporal feature tensor corresponding to the image frame sequence, and calling a pedestrian detection model to determine whether a pedestrian exists in the key frame; if yes, detecting frame information corresponding to the pedestrians in the key frames and the spatial-temporal feature tensor are input into a behavior recognition model; the detection frame information and the spatial-temporal feature tensor are processed through the behavior recognition model, and the behavior category of the pedestrian in the target video stream is determined. Whether the pedestrian exists in the key frame is judged, so that the behavior of the pedestrian in the video stream is recognized according to the detection frame information and the spatial-temporal feature tensor corresponding to the pedestrian in the key frame; and the application effect of identifying the character behavior in the video stream is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of computer vision technology, and in particular to a behavior recognition method, device, equipment and medium. Background Art

[0002] With the continuous development of technology, video surveillance systems have become an indispensable part of modern society. However, traditional video surveillance systems mainly rely on manual monitoring and cannot automatically analyze the behavior of pedestrians in the surveillance video. Therefore, behavior recognition technology based on video surveillance has emerged. Behavior recognition technology based on video surveillance can realize automatic recognition and analysis of human behavior by processing video images with technologies such as deep learning and machine learning. This technology can be widely used in public security, intelligent transportation, smart home and other fields, bringing more convenience and safety to people's lives.

[0003] However, current behavior detection and recognition methods for people in videos often fail to balance recognition accuracy and recognition speed during behavior recognition, resulting in poor application results. Summary of the invention

[0004] In view of this, the present invention provides a behavior recognition method, device, equipment and medium to solve the problem of poor application effect when recognizing human behavior in a video.

[0005] In a first aspect, the present invention provides a behavior recognition method, the method comprising:

[0006] Acquire a target video stream, perform frame extraction processing on the target video stream, obtain an image frame sequence of the target video stream and determine a key frame in the image frame sequence;

[0007] Preprocessing the image frame sequence to obtain a spatiotemporal feature tensor corresponding to the image frame sequence, and calling a pedestrian detection model to determine whether there is a pedestrian in the key frame;

[0008] If so, the detection frame information corresponding to the pedestrian in the key frame and the spatiotemporal feature tensor are input into the behavior recognition model;

[0009] The detection frame information and the spatiotemporal feature tensor are processed by the behavior recognition model to determine the behavior category of the pedestrian in the target video stream.

[0010] The method provided in this aspect extracts frames from the target video stream to obtain a corresponding image frame sequence, thereby determining the key frames and spatiotemporal feature tensors corresponding to the image frame sequence. When a pedestrian is identified in the key frame, the detection box information and the spatiotemporal feature tensor in the key frame are processed by a preset behavior recognition model to determine the behavior category of the pedestrian in the target video stream. This can ensure the speed and accuracy of identifying the behavior of people in the video stream and improve the application effect of identifying the behavior of people in the video stream through the model.

[0011] In an optional embodiment, the method further includes:

[0012] If there is no pedestrian in the key frame, re-performing frame extraction processing on the target video stream, and updating the image frame sequence and the key frame in the image frame sequence;

[0013] Based on the updated image frame sequence and the key frames in the image frame sequence, performing the steps of preprocessing the image frame sequence and calling a pedestrian detection model to determine whether there is a pedestrian in the key frames;

[0014] Until the number of repeated frame extractions reaches a preset number, if there is still no pedestrian in the key frame, the recognition of the target video stream is terminated.

[0015] In this embodiment, when no pedestrian is recognized in a key frame, frame extraction is performed again and subsequent processing is performed until no pedestrian is recognized in the key frame after repeated frame extraction. Then, recognition of the target video stream is terminated, thereby avoiding omissions in recognition of pedestrian behavior in the video stream.

[0016] In an optional embodiment, the target recognition model includes a first convolution block and at least one other convolution block connected in sequence;

[0017] The step of processing the detection frame information and the spatiotemporal feature tensor by the behavior recognition model to determine the behavior category of the pedestrian in the target video stream includes:

[0018] Determine a target mask corresponding to the pedestrian detection frame according to pixel information and position information corresponding to the pedestrian detection frame in the key frame;

[0019] Using each of the multiple convolution blocks in the target recognition model, convolution processing is performed on the spatiotemporal enhanced feature tensor to obtain an updated spatiotemporal enhanced feature tensor; the spatiotemporal enhanced feature tensor is obtained by performing feature enhancement fusion on the original input tensor using the target mask; the original input tensor of the first convolution block is the spatiotemporal feature tensor; the original input tensor of the other convolution blocks is the output tensor of the previous convolution block;

[0020] Based on the spatiotemporal enhanced feature tensor updated by the last other convolutional block, the behavior category of the pedestrian in the target video stream is determined.

[0021] In this implementation, when the behavior recognition model processes the spatiotemporal feature tensor, the target mask corresponding to the pedestrian detection box is used to perform feature fusion enhancement on the original spatiotemporal feature tensor corresponding to the current convolution block before each convolution process, and then convolution processing is performed on it, and the output tensor after the convolution processing is used as the original spatiotemporal feature tensor of the next convolution block. After multiple layers of feature fusion enhancement, the recognition effect of the behavior recognition model can be effectively improved.

[0022] In an optional implementation, preprocessing the image frame sequence to obtain a spatiotemporal feature tensor corresponding to the image frame sequence includes:

[0023] The image frame sequence is size-formatted, regularized, and dimensionally converted to obtain a spatiotemporal feature tensor corresponding to the image frame sequence.

[0024] In this embodiment, the image frame sequence is resized, regularized, and dimensionally converted to obtain a feature tensor corresponding to each frame of the image, thereby ensuring the validity of the spatiotemporal feature tensor corresponding to the entire image frame sequence.

[0025] In an optional implementation manner, the key frames in the image frame sequence are determined as follows:

[0026] An image frame located in the middle of the image frame sequence is determined as a key frame.

[0027] In this embodiment, the image frame in the middle position is determined as the key frame so that the pedestrian detection frame information determined subsequently is the position of the person in the video stream at a relatively middle moment, thereby ensuring the recognition effect when subsequently performing behavior recognition based on the pedestrian detection frame and the spatiotemporal feature tensor.

[0028] In an optional implementation, determining the target mask corresponding to the pedestrian detection frame according to the pixel information and position information corresponding to the pedestrian detection frame in the key frame includes:

[0029] Determining the average value of pixels inside the pedestrian detection box in the key frame;

[0030] The pixel values ​​inside the pedestrian detection frame in the key frame are replaced with the pixel average value, and the pixel values ​​outside the pedestrian detection frame in the key frame are replaced with 0.

[0031] In this implementation, the pixel values ​​inside the pedestrian detection frame are replaced with the pixel average value to ensure the fusion effect when the feature enhancement fusion is performed based on the target mask.

[0032] In an optional implementation, the step of performing feature enhancement fusion on the original input tensor using the target mask includes:

[0033] The feature tensor corresponding to each frame of the original input tensor and the target mask are superimposed to obtain a spatiotemporal enhanced feature tensor.

[0034] In this implementation, the feature tensor and the target mask are fused by superposition, which can quickly and efficiently achieve enhanced fusion of the feature tensor.

[0035] In a second aspect, the present invention provides a behavior recognition device, the device comprising:

[0036] A video stream data acquisition module is used to acquire a target video stream, perform frame extraction processing on the target video stream, obtain an image frame sequence of the target video stream and determine a key frame in the image frame sequence;

[0037] An image frame preprocessing module is used to preprocess the image frame sequence to obtain a spatiotemporal feature tensor corresponding to the image frame sequence, and at the same time call a pedestrian detection model to determine whether there is a pedestrian in the key frame;

[0038] A feature tensor input module, used for inputting the detection frame information corresponding to the pedestrian in the key frame and the spatiotemporal feature tensor into the behavior recognition model when there is a pedestrian in the key frame;

[0039] The behavior category determination module is used to process the detection frame information and the spatiotemporal feature tensor through the behavior recognition model to determine the behavior category of the pedestrian in the target video stream.

[0040] In a third aspect, the present invention provides a computer device, comprising: a memory and a processor, the memory and the processor being communicatively connected to each other, the memory storing computer instructions, and the processor executing the behavior recognition method of the first aspect or any corresponding embodiment thereof by executing the computer instructions.

[0041] In a fourth aspect, the present invention provides a computer-readable storage medium having computer instructions stored thereon, the computer instructions being used to enable a computer to execute the behavior recognition method of the first aspect or any corresponding embodiment thereof. BRIEF DESCRIPTION OF THE DRAWINGS

[0042] In order to more clearly illustrate the specific implementation methods of the present invention or the technical solutions in the prior art, the drawings required for use in the specific implementation methods or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are some implementation methods of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.

[0043] Figure 1 is a flow chart of a behavior recognition method according to an embodiment of the present invention;

[0044] Figure 2 is a flow chart of another behavior recognition method according to an embodiment of the present invention;

[0045] Figure 3 is an example flow chart of a behavior detection and recognition method according to an embodiment of the present invention;

[0046] Figure 4 is a structural example diagram of a behavior detection and recognition system according to an embodiment of the present invention;

[0047] Figure 5 is a structural block diagram of a behavior recognition device according to an embodiment of the present invention;

[0048] Figure 6 It is a schematic diagram of the hardware structure of a computer device according to an embodiment of the present invention. DETAILED DESCRIPTION

[0049] In order to make the purpose, technical solution and advantages of the embodiments of the present invention clearer, the technical solution in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative work are within the scope of protection of the present invention.

[0050] With the continuous development of technology, video surveillance systems have become an indispensable part of modern society. However, traditional video surveillance systems mainly rely on manual monitoring and cannot automatically analyze the behavior of pedestrians in the surveillance video. Therefore, behavior recognition technology based on video surveillance has emerged. Behavior recognition technology based on video surveillance can realize automatic recognition and analysis of human behavior by processing video images with technologies such as deep learning and machine learning. This technology can be widely used in public security, intelligent transportation, smart home and other fields, bringing more convenience and safety to people's lives.

[0051] However, current behavior detection and recognition methods for people in videos often fail to balance recognition accuracy and recognition speed during behavior recognition, resulting in poor application results.

[0052] To this end, an embodiment of the present invention provides a behavior recognition method, which extracts frames from a target video stream to obtain a corresponding image frame sequence, thereby determining key frames and spatiotemporal feature tensors corresponding to the image frame sequence. When a pedestrian is identified in a key frame, the detection box information and spatiotemporal feature tensor in the key frame are processed by a preset behavior recognition model to determine the behavior category of the pedestrian in the target video stream. This can ensure the speed and accuracy of identifying the behavior of people in the video stream and improve the application effect of behavior recognition through the model.

[0053] According to an embodiment of the present invention, an embodiment of a behavior recognition method is provided. It should be noted that the steps shown in the flowchart of the accompanying drawings can be executed in a computer system such as a set of computer executable instructions, and although a logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in an order different from that shown here.

[0054] In this embodiment, a behavior recognition method is provided, which can be used in the above scenario of recognizing the behavior of people in a video stream through a model. Figure 1 is a flow chart of a behavior recognition method according to an embodiment of the present invention. Figure 1 As shown, the process includes the following steps:

[0055] Step S101, obtaining a target video stream, performing frame extraction processing on the target video stream, obtaining an image frame sequence of the target video stream and determining key frames in the image frame sequence.

[0056] First, obtain the video stream that needs to be used for character behavior recognition, that is, the target video stream. For the video stream, there are multiple continuous frames of images. According to the pre-set frame extraction method, such as random frame extraction, frame extraction with a fixed frame interval, etc., the target video stream is subjected to frame extraction. For the multiple frames of images obtained by frame extraction, they are sorted according to time to obtain an image frame sequence. A frame is selected as a key frame in the image frame sequence. For example, an image frame closest to the middle moment of the video stream can be selected as a key frame according to the time of each image frame in the video. Regarding the specific key frame selection method, it can be selected according to actual needs. For example, the key frame is selected from the perspective of time. For example, the image frame closest to the middle moment can be selected as the key frame; or the key frame is selected from the perspective of image information. For example, the image frame with the richest image information can be selected as the key frame. There is no limitation here.

[0057] Step S102 , preprocessing the image frame sequence to obtain the spatiotemporal feature tensor corresponding to the image frame sequence, and at the same time calling the pedestrian detection model to determine whether there is a pedestrian in the key frame.

[0058] It can be understood that the image frame sequence is a plurality of frames of images sorted by time. Each frame of the image frame sequence is preprocessed to obtain the feature tensor corresponding to the frame of the image. The feature tensors corresponding to each frame of the image are combined to obtain the spatiotemporal feature tensor corresponding to the entire image frame sequence, including the feature tensors of each frame of the image sorted by time. For specific preprocessing methods, reference can be made to feature extraction methods related to image processing, such as alignment processing, filtering, grayscale extraction, etc., which are not limited here.

[0059] While preprocessing, a pre-set target detection model is called to detect the key frames. The target detection model can be a commonly used person recognition model, which is used to identify the presence of pedestrians in the key frames. If present, the identified pedestrians will be marked through the corresponding person detection frame.

[0060] Step S103: If it exists, the detection box information and spatiotemporal feature tensor corresponding to the pedestrian in the key frame are input into the behavior recognition model.

[0061] When the pedestrian detection model detects the presence of pedestrians in the key frame, it indicates that there are people whose behavior needs to be recognized in the target video stream, namely, the detected pedestrians. Therefore, the detection frame information corresponding to the pedestrians detected in the key frame and the spatiotemporal feature tensor corresponding to the image frame sequence can be input into the behavior recognition model. The detection frame information may include: detection frame position information, pixel information inside the detection frame. The behavior recognition model can be understood as a behavior recognition model pre-trained by deep learning related machine learning methods.

[0062] Step S104: Process the detection frame information and the spatiotemporal feature tensor through the behavior recognition model to determine the behavior category of the pedestrian in the target video stream.

[0063] The behavior recognition model can estimate the relevant positions of pedestrians in the entire image frame sequence based on the detection frame information. For example, if the positions corresponding to the pedestrian detection frames in the key frames are the lower right corner and the upper left corner, then the pedestrians are most likely located around these two positions in the entire video stream. The position of the behavior detection frame or the area within the preset range near it can be directly used as the main activity area of ​​pedestrians in the entire video stream. Therefore, in the process of the behavior recognition model processing the spatiotemporal feature tensor, the image features corresponding to the detection frame in each image frame in the spatiotemporal feature tensor are assigned higher weights or the corresponding feature parameters are enhanced to ensure the final effect of pedestrian behavior recognition in the video stream.

[0064] For the behavior recognition model, it can be a behavior recognition model built and trained based on the principles of convolutional neural networks, residual neural networks, etc. When it recognizes human behavior based on the spatiotemporal feature tensor corresponding to the image frame sequence, it only needs to focus on processing the feature information of the relevant position of the detection frame based on its own recognition mechanism. For example, more processing weights can be allocated, or the features of the relevant position of the detection frame can be enhanced, etc., which can be set in combination with the recognition mechanism of the behavior recognition model itself.

[0065] The behavior recognition method provided in this embodiment extracts frames from the target video stream to obtain a corresponding image frame sequence, thereby determining the key frames and spatiotemporal feature tensors corresponding to the image frame sequence. When a pedestrian is identified in the key frame, the detection box information and the spatiotemporal feature tensor in the key frame are processed by a preset behavior recognition model to determine the behavior category of the pedestrian in the target video stream. This can ensure the speed and accuracy of identifying the behavior of people in the video stream and improve the application effect of identifying the behavior of people in the video stream through the model.

[0066] According to an embodiment of the present invention, another behavior recognition method embodiment is provided, which can be used in the above scenario of recognizing the behavior of people in a video stream through a model. Figure 2 is a flow chart of another behavior recognition method according to an embodiment of the present invention. Figure 2 As shown, the process includes the following steps:

[0067] Step S201, obtaining a target video stream, performing frame extraction processing on the target video stream, obtaining an image frame sequence of the target video stream and determining key frames in the image frame sequence.

[0068] Specifically, in step S201, the key frame in the image frame sequence is determined as follows:

[0069] An image frame located in the middle of the image frame sequence is determined as a key frame.

[0070] It can be understood that in the image frame sequence obtained by extracting frames from the target video stream, the image frames are arranged according to time. Under such an arrangement, the image frames closer to the middle position in the image frame sequence can relatively better represent the center position of the pedestrian activities corresponding to each image frame in the image frame sequence. For example, when the frame extraction is performed at the same frame interval, the position of the pedestrian in the image frame at the middle position belongs to the center position of the pedestrian activities corresponding to each image frame in the image frame sequence.

[0071] To help understand the simplest case, assume that the image width in the video stream is 10 units in length, represented by the horizontal axis, that is, 0-10, and the corresponding image frame sequence is 5 frames in total. The video shows a pedestrian walking on the road, and the horizontal positions of the pedestrian in these five frames are 4, 5, 7, 8, and 9. The position of the pedestrian corresponding to the image frame in the middle position can better reflect the main position of the pedestrian in the video, that is, it can be used as the center position of the pedestrian's activities in the video. After the frame is determined as a key frame, the detection box information corresponding to the pedestrian can help in subsequent image recognition, and more feature enhancement or more weight allocation to the main area where the pedestrian exists, thereby improving recognition accuracy.

[0072] The above example is based on the case where there are pedestrians in the key frame to explain why the image frame in the middle position is selected as the key frame. In actual situations, there may be no pedestrians in the target video stream, or no pedestrians are detected in the image frame number in the middle position of the extracted image frame sequence. In this case, using the image frame in the middle position as the key frame can also more quickly determine whether there are pedestrians in the video stream, thereby avoiding the detection time consumed by implementing pedestrian detection for each frame of the image.

[0073] In one example, for the image frame sequence, the image frame closest to the middle moment of the video can also be selected as the key frame according to the corresponding time of each image frame in the video. Specifically, the key frame can be determined in combination with the frame extraction method, as long as it can be ensured that the pedestrian position in the determined image frame can reflect the main activity area of ​​the pedestrian in each image in the image frame sequence.

[0074] Step S202 , preprocessing the image frame sequence to obtain the spatiotemporal feature tensor corresponding to the image frame sequence, and at the same time calling the pedestrian detection model to determine whether there is a pedestrian in the key frame.

[0075] Specifically, in step S202, the image frame sequence is preprocessed to obtain a spatiotemporal feature tensor corresponding to the image frame sequence, including:

[0076] The image frame sequence is resized, regularized and dimensionally transformed to obtain the spatiotemporal feature tensor corresponding to the image frame sequence.

[0077] It can be understood that for different video streams, due to their different video specifications, the size of the image frames in the image frame sequence may not meet the requirements of the recognition model. Therefore, it is necessary to format the size of each frame in the image frame sequence to ensure that the image size of each frame is consistent and meets the requirements of the recognition model.

[0078] Next, it is necessary to regularize the pixels in each frame of the image, that is, compress the pixel values ​​within a preset range. Specifically, in order to avoid gradient disappearance in the subsequent image recognition process, the preset range can be set to [-1,1]. At the same time, each frame of the image in the image frame sequence is an RGB image, that is, it has three channels. The single-channel grayscale images corresponding to different channels of the image can be regularized to obtain the feature vectors corresponding to each image frame. After summarizing these feature vectors, they are converted into dimensions. For example, the channel dimension and the time dimension can be exchanged to obtain the spatiotemporal feature tensor corresponding to the image frame sequence.

[0079] The above steps can be exemplified by the following example to assist understanding. Assuming that the image frame sequence has T frame images, size formatting can be understood as formatting the sizes of all T frame images into H*W, where H and W are the height and width of the image respectively.

[0080] Regularization can be understood as compressing the pixel values ​​of the image in [0, 1]. In order to prevent the gradient from disappearing during the training process of the model, the present invention compresses the pixel values ​​in [-1, 1]. The calculation formula is:

[0081]

[0082] Among them, T, C, H, and W are the number of frames in the sequence, the number of channels of the image (the number of channels of the RGB image is 3), the height, and the width, respectively.

[0083] Dimension conversion is to exchange the time dimension (number of frames) and channel dimension in F obtained above to obtain the spatiotemporal feature tensor I∈R C×T×H×W .

[0084] Step S203: If the pedestrian exists, the detection frame information and spatiotemporal feature tensor corresponding to the pedestrian in the key frame are input into the behavior recognition model. Figure 1 Step S103 of the illustrated embodiment will not be described in detail here.

[0085] Furthermore, if there is no pedestrian in the key frame, the target video stream is re-framed to update the image frame sequence and the key frame in the image frame sequence;

[0086] Based on the updated image frame sequence and the key frames in the image frame sequence, performing the steps of preprocessing the image frame sequence and calling a pedestrian detection model to determine whether there is a pedestrian in the key frame;

[0087] Until the number of repeated frame extractions reaches the preset number, if there is still no pedestrian in the key frame, the recognition of the target video stream is terminated.

[0088] It can be understood that when no pedestrian is detected in the key frame, it may be that there are no pedestrians in the target video stream, or there may be no pedestrians at the corresponding moment of the key frame during the frame extraction process. In this case, it is necessary to return to step S201 and re-extract the frame. When re-extracting the frame, it is necessary to determine whether to adjust the frame extraction rules based on the actual frame extraction method. If a random frame extraction method is used, no adjustment is required. If the frame extraction is performed using an interval frame number, the interval frame number or the starting frame extraction position needs to be changed to ensure that the image frame sequence obtained by re-extracting the frame changes.

[0089] After re-extracting the frame, the subsequent step S202 is re-executed. If there is still no pedestrian in the key frame, it can be decided whether to extract the frame again according to the preset number of cycles actually set. If there is no pedestrian in the key frame during multiple frame extractions, it can be said that there is no pedestrian to be detected in the target video stream, and the recognition of the target video stream can be terminated, and the recognized video stream can be replaced to perform a new round of behavior recognition.

[0090] Step S204: Process the detection frame information and the spatiotemporal feature tensor through the behavior recognition model to determine the behavior category of the pedestrian in the target video stream.

[0091] Specifically, the target recognition model includes a first convolution block and at least one other convolution block connected in sequence;

[0092] In step S204, it includes:

[0093] Step S204 - 1 , determining a target mask corresponding to the pedestrian detection frame according to the pixel information and position information corresponding to the pedestrian detection frame in the key frame.

[0094] For the pedestrian detection frame in the key frame detected in the above step, the mask corresponding to the pedestrian detection frame can be determined according to the corresponding specific pixels and the pixel positions inside it. For example, the pixel values ​​inside the detection frame remain unchanged and the pixel values ​​outside are set to 0; or the pixel values ​​inside are set to 1 and the pixel values ​​outside are set to 0, etc., and the setting is made in combination with the actual feature enhancement and fusion requirements.

[0095] Specifically, in step S204-1, according to the pixel information and position information corresponding to the pedestrian detection frame in the key frame, determining the target mask corresponding to the pedestrian detection frame includes:

[0096] Determine the average value of pixels inside the pedestrian detection box in the key frame;

[0097] The pixel values ​​inside the pedestrian detection box in the key frame are replaced with the pixel average value, and the pixel values ​​outside the pedestrian detection box in the key frame are replaced with 0.

[0098] It can be understood that when determining the target mask, the pixel average value corresponding to each pixel inside the detection box can be first determined, the pixel value of each pixel is replaced by the average value, and the pixel value outside the detection box is set to 0.

[0099] In another example, the inner pixel value can be directly set to 1 and the outer pixel value can be set to 0.5. For such a mask, when performing feature enhancement fusion later, the corresponding pixel values ​​are multiplied to perform feature enhancement fusion. The specific mask setting method can be set in combination with the specific method of subsequent feature enhancement fusion, which is not limited here.

[0100] Step S204-2, using each of the multiple convolution blocks in the target recognition model, convolution processing is performed on the spatiotemporal enhanced feature tensor to obtain an updated spatiotemporal enhanced feature tensor; the spatiotemporal enhanced feature tensor is obtained by using the target mask to perform feature enhancement fusion on the original input tensor; the original input tensor of the first convolution block is the spatiotemporal feature tensor; the original input tensors of other convolution blocks are the output tensors of the previous convolution blocks.

[0101] It can be understood that for a person recognition model with multiple convolution blocks, it is necessary to perform multiple rounds of convolution processing on the features when performing person recognition, that is, multiple convolution blocks are set up, and after each convolution block performs convolution processing on the input feature tensor, it outputs the processing result to the next convolution block for a new round of convolution processing.

[0102] The feature tensor input to each convolution block is a spatiotemporal enhanced feature tensor obtained by fusion of the original input tensor with enhanced features through the target mask.

[0103] Different convolution blocks correspond to different original input tensors. For the first convolution block in the behavior recognition model, the original input tensor corresponding to it is the spatiotemporal feature tensor initially input to the behavior recognition model. For each convolution block after the first convolution block, the original input tensor corresponding to it is the feature tensor output by the previous convolution block.

[0104] Specifically, in step S204-2, the target mask is used to perform feature enhancement fusion on the original input tensor, including:

[0105] The feature tensor corresponding to each frame of the original input tensor and the target mask are superimposed to obtain a spatiotemporal enhanced feature tensor.

[0106] It can be understood that when performing feature enhancement fusion through the target mask, the feature tensor corresponding to each frame image in the original input tensor and the target mask can be directly superimposed, thereby achieving enhanced fusion of the spatiotemporal feature tensor.

[0107] For example, the following formula can be used to assist understanding:

[0108]

[0109] Where I is the original feature tensor corresponding to a convolution block, that is, the spatiotemporal feature tensor that has not yet been fused, t represents the number of image frames, M is the target mask, and the Sigmoid function is used to normalize the spatiotemporal feature tensor after feature enhancement, so that the spatiotemporal feature tensor after feature enhancement is more coordinated in scale during subsequent model calculations.

[0110] Taking a channel in the feature tensor corresponding to a certain frame of the spatiotemporal feature tensor as an example, assume that the feature tensor I(t) under this channel is:

[0111]

[0112] The target mask M is:

[0113]

[0114] After feature fusion enhancement, the matrix corresponding to I(t)+M is:

[0115]

[0116] Taking this as an example, the feature tensors corresponding to different frames and channels in the spatiotemporal feature tensor are enhanced and fused. For example, in addition to the above-mentioned superposition method, the pixel values ​​of the area corresponding to the target mask in the feature tensor can also be multiplied by a specific multiple to achieve feature enhancement and fusion. The specific setting can be combined with the actual situation. The above is only an exemplary implementation method for enhancing and fusing the feature tensor.

[0117] Step S204-3, based on the spatiotemporal enhanced feature tensor updated by the last other convolutional block, determine the behavior category of the pedestrian in the target video stream.

[0118] The spatiotemporal enhanced feature tensor after the convolution processing of the last convolution block in the behavior recognition model can be identified through its features in the classification layer of the behavior recognition model to determine the behavior category of the pedestrian. Specifically, the spatiotemporal enhanced feature tensor can be input into the maximum pooling layer, the fully connected layer and the SoftMax layer for classification and recognition. Specifically, the specific classification operation can be performed in combination with the network architecture corresponding to the behavior recognition model, which will not be repeated here.

[0119] The behavior recognition method provided by the embodiment of the present invention extracts frames from the target video stream to obtain a corresponding image frame sequence, thereby determining the key frames and spatiotemporal feature tensors corresponding to the image frame sequence. When a pedestrian is identified in the key frame, the detection box information and spatiotemporal feature tensor in the key frame are processed by a preset behavior recognition model to determine the behavior category of the pedestrian in the target video stream. This can ensure the speed and accuracy of identifying the behavior of people in the video stream and improve the application effect of behavior recognition through the model.

[0120] In order to facilitate understanding of the above method embodiment, according to an embodiment of the present invention, a behavior detection and identification method flow chart is also provided. Figure 3 shown.

[0121] First, obtain the surveillance video stream for pedestrian behavior recognition, extract T frames from the surveillance video stream, and combine them into an image frame sequence in chronological order, where T is a positive even number. For example, one frame can be extracted every five consecutive frames, and T is 16.

[0122] Next, extract the T / 2th frame of the above T-frame image sequence as the key frame, use the target detection model to detect pedestrians on the key frame, obtain the pedestrian target frame, and obtain the target frame coordinates [x, y, w, h] of each pedestrian, where (x, y) is the coordinate of the upper left corner of the target frame in the frame, and w, h are the width and height of the frame respectively. If no pedestrian is detected in the key frame, return to the above frame extraction link to re-extract the frame. Exemplarily, in this embodiment, the target detection model adopts the lightweight model yolov5s.

[0123] At the same time, the extracted T-frame image sequence is preprocessed and converted into a spatiotemporal feature tensor, wherein the preprocessing process may include size formatting, regularization, and dimension conversion.

[0124] The size formatting is to format the size of the T frame image into H×W, where H and W are the height and width of the image respectively. For example, H×W is 224×224, and the formatting method adopts a direct resize function.

[0125] Regularization is to normalize the image value to the range of [0, 1]. In order to prevent the gradient from vanishing during the training process and causing non-convergence, the present invention compresses the pixel value to [-1, 1]. The calculation formula is:

[0126]

[0127] Wherein, T, C, H, and W are the number of frames in the sequence, the number of channels of the image (the number of channels of the RGB image is 3), the height, and the width. In this embodiment, T is 16 obtained in step 2, C is 3, which is the number of channels of the original image, and H and W are 224 respectively in the above steps.

[0128] Dimension conversion is to exchange the time dimension (number of frames) of F with the channel dimension to obtain the spatiotemporal feature tensor I∈R C×T×H×W .

[0129] When the target detection model recognizes that there is a pedestrian detection box in the key frame, the key frame is scaled according to the size format of the above-mentioned size normalization, and the pedestrian detection box in the key frame and the spatiotemporal feature tensor obtained after the above-mentioned data preprocessing are input into the behavior recognition model, and finally the behavior category of each pedestrian is output.

[0130] When the behavior recognition model identifies the behavior category of each pedestrian based on the input spatiotemporal feature tensor and pedestrian detection frame, it first generates the corresponding mask M based on the target frame coordinates of the pedestrian, where M∈R C×H×W In this embodiment, the mask M replaces the pixel values ​​in the target box by calculating the average value of the pixels in the target box, and all pixels outside the target box are set to 0.

[0131] The behavior recognition model uses a three-dimensional convolutional network, which is composed of multiple three-dimensional convolutional modules, and finally connects the maximum pooling layer, the fully connected layer and the SoftMax layer to output the behavior category of each pedestrian; in this embodiment, before the spatiotemporal feature tensor is input into each three-dimensional convolutional block and the maximum pooling layer, the spatiotemporal feature enhancement fusion processing is performed using a mask, and the calculation process is:

[0132]

[0133] For each convolution block, the spatiotemporal feature tensor received after feature enhancement and fusion is convolved and activated as follows:

[0134] I″=Conv(I′)

[0135] I″′=ReLU(I″)

[0136] For example, in the above formula, ReLU(x)=max(0,x), Conv is a convolution operation.

[0137] Further, for the above method embodiment, according to an embodiment of the present invention, a structural example diagram of a behavior detection and recognition system is also provided, such as Figure 4 shown.

[0138] As shown in the figure, the image frame in the middle position of the image frame sequence corresponding to the T frame image is extracted as the key frame, and the pedestrian detection frame is identified through the target detection model. At the same time, the preprocessed image frame sequence is input into the behavior recognition model, and convolution processing is performed through multiple layers of convolution blocks. Before each convolution processing, the feature tensor to be input into the convolution is enhanced and fused through the mask corresponding to the behavior detection frame. Finally, when the feature tensor output by the last convolution block is input into the maximum pooling layer, the feature enhancement processing is also performed in this way, and then the behavior category of each pedestrian finally recognized is obtained through the fully connected layer and SoftMax.

[0139] The embodiment of the present invention proposes an end-to-end behavior detection and recognition method, which adopts a parallel structure for the target detection model and the behavior recognition model, and simultaneously realizes pedestrian detection and behavior recognition. Compared with the traditional serial structure, pedestrians are first detected, then tracked and image frames are collected, and then behavior recognition is performed on each pedestrian, which greatly shortens the reasoning time and ensures the high real-time performance of real-time recognition.

[0140] In addition, the present invention proposes a spatiotemporal feature fusion method, which obtains the target frame of the pedestrian through the target detection model, and performs spatiotemporal feature enhancement processing on the spatiotemporal feature tensor by generating a mask, which helps to enhance the representation of the spatiotemporal action characteristics of the pedestrian, suppress the interference of background information, and improve the accuracy of behavior recognition. At the same time, by extracting the key frames of the image sequence, detecting the pedestrians in the key frames through the target detection model, and mapping them to the image sequence using the spatiotemporal feature enhancement fusion method, and finally sending them to the behavior recognition model, it is possible to realize behavior recognition of multiple people at the same time, and improve the overall throughput of the model.

[0141] In this embodiment, a behavior recognition device is also provided, which is used to implement the above-mentioned embodiments and preferred implementation modes, and the descriptions that have been made will not be repeated. As used below, the term "module" can implement a combination of software and / or hardware of a predetermined function. Although the devices described in the following embodiments are preferably implemented in software, the implementation of hardware, or a combination of software and hardware, is also possible and conceivable.

[0142] This embodiment provides a behavior recognition device, such as Figure 5 As shown, including:

[0143] The video stream data acquisition module 401 is used to acquire a target video stream, perform frame extraction processing on the target video stream, obtain an image frame sequence of the target video stream, and determine key frames in the image frame sequence.

[0144] The image frame preprocessing module 402 is used to preprocess the image frame sequence to obtain the spatiotemporal feature tensor corresponding to the image frame sequence, and at the same time call the pedestrian detection model to determine whether there is a pedestrian in the key frame.

[0145] The feature tensor input module 403 is used to input the detection frame information and the spatiotemporal feature tensor corresponding to the pedestrian in the key frame into the behavior recognition model when there is a pedestrian in the key frame;

[0146] The behavior category determination module 404 is used to process the detection frame information and the spatiotemporal feature tensor through the behavior recognition model to determine the behavior category of the pedestrian in the target video stream.

[0147] In some optional implementations, the feature tensor input module 403 is further used to, when there is no pedestrian in the key frame, re-perform frame extraction processing on the target video stream to update the image frame sequence and the key frame in the image frame sequence;

[0148] Based on the updated image frame sequence and the key frames in the image frame sequence, the steps of preprocessing the image frame sequence and calling the pedestrian detection model to determine whether there is a pedestrian in the key frame are performed.

[0149] Until the number of repeated frame extractions reaches the preset number, if there is still no pedestrian in the key frame, the recognition of the target video stream is terminated.

[0150] In some optional embodiments, the object recognition model includes a first convolution block and at least one other convolution block connected in sequence;

[0151] The behavior category determination module 404, when processing the detection frame information and the spatiotemporal feature tensor through the behavior recognition model to determine the behavior category of the pedestrian in the target video stream, includes:

[0152] Determine the target mask corresponding to the pedestrian detection frame according to the pixel information and position information corresponding to the pedestrian detection frame in the key frame;

[0153] Using each of the multiple convolution blocks in the target recognition model, the spatiotemporal enhanced feature tensor is convolved to obtain an updated spatiotemporal enhanced feature tensor; the spatiotemporal enhanced feature tensor is obtained by performing feature enhancement fusion on the original input tensor using the target mask; the original input tensor of the first convolution block is the spatiotemporal feature tensor; the original input tensors of other convolution blocks are the output tensors of the previous convolution blocks;

[0154] Based on the spatiotemporal enhanced feature tensor updated by the last other convolutional block, the behavior category of the pedestrian in the target video stream is determined.

[0155] In some optional implementations, the image frame preprocessing module 402, when preprocessing the image frame sequence to obtain the spatiotemporal feature tensor corresponding to the image frame sequence, includes:

[0156] The image frame sequence is resized, regularized and dimensionally transformed to obtain the spatiotemporal feature tensor corresponding to the image frame sequence.

[0157] In an optional implementation manner, a key frame in an image frame sequence is determined as follows:

[0158] The image frame located in the middle of the image frame sequence is determined as a key frame

[0159] In an optional implementation, the behavior category determination module 404, when determining the target mask corresponding to the pedestrian detection frame according to the pixel information and position information corresponding to the pedestrian detection frame in the key frame, includes:

[0160] Determine the average value of pixels inside the pedestrian detection box in the key frame;

[0161] The pixel values ​​inside the pedestrian detection box in the key frame are replaced with the pixel average value, and the pixel values ​​outside the pedestrian detection box in the key frame are replaced with 0.

[0162] In an optional implementation, the behavior category determination module 404, when performing feature enhancement fusion on the original input tensor using the target mask, includes:

[0163] The feature tensor corresponding to each frame of the original input tensor and the target mask are superimposed to obtain a spatiotemporal enhanced feature tensor.

[0164] The further functional description of each of the above modules and units is the same as that of the above corresponding embodiments and will not be repeated here.

[0165] The behavior recognition device in this embodiment is presented in the form of a functional unit, where the unit refers to an ASIC (Application Specific Integrated Circuit) circuit, a processor and memory that executes one or more software or fixed programs, and / or other devices that can provide the above functions.

[0166] The embodiment of the present invention also provides a computer device having the above Figure 5 The behavior recognition device shown.

[0167] See also Figure 6 , Figure 6 is a schematic diagram of the structure of a computer device provided by an optional embodiment of the present invention, such as Figure 6As shown, the computer device includes: one or more processors 10, a memory 20, and interfaces for connecting various components, including high-speed interfaces and low-speed interfaces. Various components are connected to each other using different buses for communication, and can be installed on a common mainboard or installed in other ways as needed. The processor can process the instructions executed in the computer device, including instructions stored in or on the memory to display the graphical information of the GUI on an external input / output device (such as, a display device coupled to the interface). In some optional embodiments, if necessary, multiple processors and / or multiple buses can be used together with multiple memories and multiple memories. Similarly, multiple computer devices can be connected, and each device provides some necessary operations (for example, as a server array, a group of blade servers, or a multi-processor system). Figure 6 A processor 10 is taken as an example.

[0168] The processor 10 may be a central processing unit, a network processor or a combination thereof. The processor 10 may further include a hardware chip. The hardware chip may be a dedicated integrated circuit, a programmable logic device or a combination thereof. The programmable logic device may be a complex programmable logic device, a field programmable gate array, a general purpose array logic or any combination thereof.

[0169] The memory 20 stores instructions executable by at least one processor 10, so that at least one processor 10 executes the method shown in the above embodiment.

[0170] The memory 20 may include a program storage area and a data storage area, wherein the program storage area may store an operating system, an application required for at least one function; the data storage area may store data created according to the use of the computer device, etc. In addition, the memory 20 may include a high-speed random access memory, and may also include a non-transient memory, such as at least one disk storage device, a flash memory device, or other non-transient solid-state storage device. In some optional embodiments, the memory 20 may optionally include a memory remotely arranged relative to the processor 10, and these remote memories may be connected to the computer device via a network. Examples of the above-mentioned network include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.

[0171] The memory 20 may include a volatile memory, such as a random access memory; the memory may also include a non-volatile memory, such as a flash memory, a hard disk or a solid state drive; the memory 20 may also include a combination of the above types of memory.

[0172] The computer device also includes an input device 30 and an output device 40. The processor 10, the memory 20, the input device 30 and the output device 40 may be connected via a bus or other means. Figure 6 The example of connecting through bus is taken in the following.

[0173] The input device 30 can receive input digital or character information, and generate key signal input related to the user settings and function control of the computer device, such as a touch screen, a keypad, a mouse, a track pad, a touch pad, an indicator bar, one or more mouse buttons, a trackball, a joystick, etc. The output device 40 may include a display device, an auxiliary lighting device (e.g., an LED) and a tactile feedback device (e.g., a vibration motor), etc. The above-mentioned display device includes but is not limited to a liquid crystal display, a light emitting diode, a display and a plasma display. In some optional embodiments, the display device can be a touch screen.

[0174] The embodiment of the present invention also provides a computer-readable storage medium. The method according to the embodiment of the present invention can be implemented in hardware, firmware, or can be implemented as a computer code that can be recorded in a storage medium, or can be implemented as a computer code that is originally stored in a remote storage medium or a non-temporary machine-readable storage medium and will be stored in a local storage medium through a network download, so that the method described herein can be stored in such software processing on a storage medium using a general-purpose computer, a dedicated processor, or programmable or dedicated hardware. Among them, the storage medium can be a magnetic disk, an optical disk, a read-only storage memory, a random access memory, a flash memory, a hard disk or a solid-state hard disk, etc.; further, the storage medium can also include a combination of the above types of memories. It can be understood that a computer, a processor, a microprocessor controller, or programmable hardware includes a storage component that can store or receive software or computer code. When the software or computer code is accessed and executed by a computer, a processor, or hardware, the method shown in the above embodiment is implemented.

[0175] Although the embodiments of the present invention have been described in conjunction with the accompanying drawings, those skilled in the art may make various modifications and variations without departing from the spirit and scope of the present invention, and such modifications and variations are all within the scope defined by the appended claims.

Claims

1. A behavior recognition method, characterized in that: The method comprises: Acquire a target video stream, perform frame extraction processing on the target video stream, obtain an image frame sequence of the target video stream and determine a key frame in the image frame sequence; Preprocessing the image frame sequence to obtain a spatiotemporal feature tensor corresponding to the image frame sequence, and calling a pedestrian detection model to determine whether there is a pedestrian in the key frame; If so, the detection frame information corresponding to the pedestrian in the key frame and the spatiotemporal feature tensor are input into the behavior recognition model; The detection frame information and the spatiotemporal feature tensor are processed by the behavior recognition model to determine the behavior category of the pedestrian in the target video stream.

2. The method according to claim 1, characterized in that The method further comprises: If there is no pedestrian in the key frame, re-performing frame extraction processing on the target video stream, and updating the image frame sequence and the key frame in the image frame sequence; Based on the updated image frame sequence and the key frames in the image frame sequence, performing the steps of preprocessing the image frame sequence and calling a pedestrian detection model to determine whether there is a pedestrian in the key frames; Until the number of repeated frame extractions reaches a preset number, if there is still no pedestrian in the key frame, the recognition of the target video stream is terminated.

3. The method according to claim 1, characterized in that The target recognition model includes a first convolution block and at least one other convolution block connected in sequence; The step of processing the detection frame information and the spatiotemporal feature tensor by the behavior recognition model to determine the behavior category of the pedestrian in the target video stream includes: Determine a target mask corresponding to the pedestrian detection frame according to pixel information and position information corresponding to the pedestrian detection frame in the key frame; Using each of the multiple convolution blocks in the target recognition model, convolution processing is performed on the spatiotemporal enhanced feature tensor to obtain an updated spatiotemporal enhanced feature tensor; the spatiotemporal enhanced feature tensor is obtained by performing feature enhancement fusion on the original input tensor using the target mask; the original input tensor of the first convolution block is the spatiotemporal feature tensor; the original input tensor of the other convolution blocks is the output tensor of the previous convolution block; Based on the spatiotemporal enhanced feature tensor updated by the last other convolutional block, the behavior category of the pedestrian in the target video stream is determined.

4. The method according to claim 1, characterized in that: The preprocessing of the image frame sequence to obtain a spatiotemporal feature tensor corresponding to the image frame sequence includes: The image frame sequence is size-formatted, regularized, and dimensionally converted to obtain a spatiotemporal feature tensor corresponding to the image frame sequence.

5. The method according to claim 1, characterized in that The key frames in the image frame sequence are determined as follows: An image frame located in the middle of the image frame sequence is determined as a key frame.

6. The method according to claim 3, characterized in that The determining, according to the pixel information and position information corresponding to the pedestrian detection frame in the key frame, a target mask corresponding to the pedestrian detection frame comprises: Determining the average value of pixels inside the pedestrian detection box in the key frame; The pixel values ​​inside the pedestrian detection frame in the key frame are replaced with the pixel average value, and the pixel values ​​outside the pedestrian detection frame in the key frame are replaced with 0.

7. The method according to any one of claims 3 or 6, characterized in that: The step of using the target mask to perform feature enhancement fusion on the original input tensor includes: The feature tensor corresponding to each frame of the original input tensor and the target mask are superimposed to obtain a spatiotemporal enhanced feature tensor.

8. A behavior recognition device, characterized in that: The device comprises: A video stream data acquisition module is used to acquire a target video stream, perform frame extraction processing on the target video stream, obtain an image frame sequence of the target video stream and determine a key frame in the image frame sequence; An image frame preprocessing module is used to preprocess the image frame sequence to obtain a spatiotemporal feature tensor corresponding to the image frame sequence, and at the same time call a pedestrian detection model to determine whether there is a pedestrian in the key frame; A feature tensor input module, used for inputting the detection frame information corresponding to the pedestrian in the key frame and the spatiotemporal feature tensor into the behavior recognition model when there is a pedestrian in the key frame; The behavior category determination module is used to process the detection frame information and the spatiotemporal feature tensor through the behavior recognition model to determine the behavior category of the pedestrian in the target video stream.

9. A computer device, characterized in that: include: A memory and a processor, wherein the memory and the processor are communicatively connected to each other, the memory stores computer instructions, and the processor executes the behavior recognition method according to any one of claims 1 to 7 by executing the computer instructions.

10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores computer instructions, and the computer instructions are used to enable a computer to execute the behavior recognition method according to any one of claims 1 to 7.