Behavior Recognition Method, Apparatus, Device, and Storage Medium

By selecting high-quality image areas in the video frame sequence for behavior recognition, the problem that irrelevant information in camera video affects the recognition accuracy is solved, and more efficient and accurate behavior recognition is achieved.

CN114550049BActive Publication Date: 2025-07-18SHANGHAI SENSETIME INTELLIGENT TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210166617.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-02-23
Publication Date
2025-07-18
Estimated Expiration
2042-02-23

AI Technical Summary

Technical Problem

In the prior art, in human-centered video behavior recognition, since the video captured by the camera contains a large amount of irrelevant information, the accuracy of behavior recognition is reduced, and it is difficult to effectively identify preset behaviors.

Method used

By determining the image area of the object to be identified in the video frame sequence, selecting an image area that meets the preset conditions for behavior classification and recognition, reducing the calculation amount and improving the recognition accuracy.

Benefits of technology

By selecting high-quality image areas for behavior recognition, the calculation amount is reduced and the effective receptive field of the recognition network is improved, thereby improving the accuracy and recall of behavior recognition.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114550049B_ABST
    Figure CN114550049B_ABST
Patent Text Reader

Abstract

An embodiment of the present application provides a behavior recognition method, device, equipment and storage medium. Among them, the method includes: determining a video frame sequence in a video stream including an object to be recognized; determining at least one first image area where the object to be recognized is located in the video frame sequence; classifying the behavior of the object to be recognized based on the at least one first image area to obtain a classification result; selecting a second image area in the at least one first image area where the classification result meets a preset condition; recognizing the behavior of the object to be recognized based on the second image area to obtain a recognition result.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The embodiments of the present application relate to the field of computer vision, including but not limited to a method, device, equipment and storage medium for behavior recognition. Background Art

[0002] For human-centered video behavior recognition, the input video sequence is subjected to full-image data augmentation and then sent to a classification model for prediction. Since the video captured by the camera often contains more information and has a larger field of view. In this way, the occurrence location and human body scale of the target event of the pedestrian are also random, which affects the accuracy of behavior recognition. Summary of the Invention

[0003] The embodiments of the present application provide a technical solution for behavior recognition.

[0004] The technical solution of the embodiments of the present application is implemented as follows:

[0005] The embodiments of the present application provide a method for behavior recognition, the method includes:

[0006] In a video stream including an object to be recognized, determine a video frame sequence;

[0007] In the video frame sequence, determine at least one first image region where the object to be recognized is located;

[0008] Based on the at least one first image region, classify the behavior of the object to be recognized to obtain a classification result;

[0009] In the at least one first image region, select a second image region whose classification result meets a preset condition;

[0010] Based on the second image region, recognize the behavior of the object to be recognized to obtain a recognition result.

[0011] In some embodiments, the step of determining at least one first image region corresponding to the object to be recognized in the video frame sequence includes: detecting the object to be recognized in each video frame to obtain a plurality of detection frames of the object to be recognized; adjusting the areas of the plurality of detection frames in each video frame to obtain a plurality of adjusted regions; in the plurality of adjusted regions of each video frame, determine the at least one first image region. In this way, by adjusting the areas of the plurality of detection frames in each video frame and then selecting a part of the plurality of adjusted regions as the first image region of this video frame, duplicate recognition can be reduced.

[0012] In some embodiments, determining the at least one first image region among the multiple adjusted regions in each video frame includes: determining, among the multiple adjusted regions in each video frame, the first adjusted region with the highest first confidence level of the detection box; determining a second adjusted region whose overlap with the first adjusted region is greater than a preset overlap threshold; and removing, among the multiple adjusted regions in each video frame, the second adjusted regions with an area smaller than a preset area threshold, to obtain the at least one first image region in each video frame. In this way, the computational complexity of behavior recognition in the adjusted regions can be reduced, and the behavior recognition can be performed using the first image regions with relatively high quality, thereby improving the accuracy of recognition.

[0013] In some embodiments, classifying the behavior of the object to be recognized based on the at least one first image region to obtain a classification result includes: selecting, from the video frame sequence, video frames with a number less than a preset number of frames as target video frames; and classifying the behavior of the object to be recognized based on the first image regions in each target video frame to obtain the classification result. In this way, by selecting a small number of target video frames in the video frame sequence for object behavior classification, the computational complexity of behavior classification can be reduced.

[0014] In some embodiments, selecting, from the video frame sequence, video frames with a number less than a preset number of frames as target video frames includes: selecting the first video frame, the middle video frame, and the last video frame from the video frame sequence as the target video frames. In this way, by selecting the first video frame, the middle video frame, and the last video frame from the video frame sequence as the target video frames for subsequent processing, the complexity of subsequent calculations can be reduced.

[0015] In some embodiments, when the target video frames in the video sequence include at least one first image region, selecting, among the at least one first image region, a second image region whose classification result meets a preset condition includes: determining the second confidence level that the classification result of each first image region in the target video frame is a preset category; and determining, in the target video frame, the first image regions with a second confidence level greater than a preset confidence threshold as the second image regions. In this way, by selecting the first image regions with relatively high second confidence levels as the second image regions in the target video frames, the performance of subsequent behavior recognition based on the second image regions can be improved.

[0016] In some embodiments, when the target video frame is at least one frame, the method of identifying the behavior of the object to be identified based on the second image region and obtaining an identification result includes: determining at least one target region sequence in the video frame sequence based on the second image region in the at least one target video frame; and identifying the behavior of the object to be identified in the at least one target region sequence to obtain the identification result. In this way, by inputting multiple target region sequences into the behavior recognition network, the behavior recognition network can be more focused on identifying the behavior of the object to be identified and pay more attention to how to distinguish different motion details of the object to be identified.

[0017] In some embodiments, the method of determining at least one target region sequence in the video frame sequence based on the second image region in the at least one target video frame includes: selecting any second image region in the second image region of each target video frame in the at least one target video frame to obtain at least one set of second image regions; merging the second image regions in each set of second image regions to obtain at least one merged region; and determining a target region sequence matching each merged region in the video frame sequence to obtain the at least one target region sequence. In this way, the content of the picture in the target region sequence can be more focused on the object to be identified itself, improving the effective receptive field in practice.

[0018] In some embodiments, when the target video frame includes the first video frame, the intermediate video frame, and the last video frame, the method of selecting any second image region in the second image region of each target video frame in the at least one target video frame to obtain at least one set of second image regions includes: selecting one second image region from at least one second image region of the first video frame, the intermediate video frame, and the last video frame to obtain the at least one set of second image regions. In this way, by obtaining multiple sets of second image regions, it is convenient to perform subsequent merging according to the multiple sets of second image regions, enriching the merged regions.

[0019] In some embodiments, the method of identifying the behavior of the object to be identified in the at least one target region sequence and obtaining the identification result includes: adjusting the side length of the target region in each target region sequence to a preset side length to obtain an adjusted target region sequence; and identifying the behavior of the object to be identified in each adjusted target region sequence to obtain the identification result. In this way, by adjusting the side length of the target region in the target region sequence to a unified length, it is convenient to perform subsequent behavior recognition and can improve the efficiency of behavior recognition.

[0020] An embodiment of the present application provides a behavior recognition device, and the device includes:

[0021] A first determination module, configured to determine a video frame sequence in a video stream including an object to be recognized;

[0022] A second determination module, configured to determine at least one first image region where the object to be recognized is located in the video frame sequence;

[0023] A first classification module, configured to classify the behavior of the object to be recognized based on the at least one first image region to obtain a classification result;

[0024] A first selection module, configured to select a second image region in the at least one first image region, where the classification result of the second image region meets a preset condition;

[0025] A first recognition module, configured to recognize the behavior of the object to be recognized based on the second image region to obtain a recognition result.

[0026] Correspondingly, an embodiment of the present application provides a computer storage medium, on which computer-executable instructions are stored. After being executed, the computer-executable instructions can implement the above-mentioned behavior recognition method.

[0027] An embodiment of the present application provides an electronic device, which includes a memory and a processor. Computer-executable instructions are stored on the memory, and when the processor runs the computer-executable instructions on the memory, the above-mentioned behavior recognition method can be implemented.

[0028] An embodiment of the present application provides a behavior recognition method, device, equipment and storage medium. In the video frame sequence of a video stream, first determine the first image region of the object to be recognized in each video frame; then, classify the behavior of the object to be recognized in the first image region, and select a second image region that meets the preset condition from the first image regions of each frame according to the classification result; in this way, the subsequent recognition times can be reduced and the calculation amount can be reduced. Finally, recognize the behavior of the object to be recognized through the second image region in the video frame to obtain a recognition result; in this way, based on the second image region whose classification effect meets the preset condition, the behavior of the object to be recognized is recognized, which can improve the effective receptive field of the recognition network, thereby improving the accuracy of behavior recognition. Description of the Drawings

[0029] Figure 1 It is a schematic implementation flowchart of the behavior recognition method provided by an embodiment of the present application;

[0030] Figure 2 It is another schematic implementation flowchart of the behavior recognition method provided by an embodiment of the present application;

[0031] Figure 3Another schematic diagram of the implementation process of the behavior recognition method provided by the embodiments of the present application;

[0032] Figure 4 Yet another schematic diagram of the implementation process of the behavior recognition method provided by the embodiments of the present application;

[0033] Figure 5 Schematic diagram of the application scenario of the behavior recognition method provided by the embodiments of the present application;

[0034] Figure 6 Another schematic diagram of the application scenario of the behavior recognition method provided by the embodiments of the present application;

[0035] Figure 7 Schematic diagram of the structural composition of the behavior recognition device provided by the embodiments of the present application;

[0036] Figure 8 Schematic diagram of the composition structure of the electronic device provided by the embodiments of the present application. Detailed implementation manners

[0037] To make the objectives, technical solutions, and advantages of the embodiments of the present application clearer, the following will further describe the specific technical solutions of the invention in detail with reference to the accompanying drawings in the embodiments of the present application. The following embodiments are used to illustrate the present application but are not intended to limit the scope of the present application.

[0038] In the following description, "some embodiments" are involved, which describe a subset of all possible embodiments. However, it can be understood that "some embodiments" can be the same subset or different subsets of all possible embodiments, and can be combined with each other without conflict.

[0039] In the following description, the terms "first / second / third" involved are only used to distinguish similar objects and do not represent a specific order for the objects. It can be understood that "first / second / third" can be interchanged with a specific order or sequence when allowed, so that the embodiments of the present application described here can be implemented in an order other than that illustrated or described here.

[0040] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the technical field to which this application belongs. The terms used herein are only for the purpose of describing the embodiments of the present application and are not intended to limit the present application.

[0041] Before further elaborating on the embodiments of the present application, the nouns and terms involved in the embodiments of the present application are explained. The nouns and terms involved in the embodiments of the present application are applicable to the following explanations.

[0042] 1) Computer vision refers to machine vision that uses cameras and computers to replace human eyes for tasks such as object recognition, tracking, and measurement, and further performs graphics processing to make the computer-processed images more suitable for human eye observation or transmission to instrument detection.

[0043] 2) Non-Maximum Suppression (NMS) is to search for local maxima and suppress the maxima. Taking object detection as an example. During the process of object detection, a large number of candidate bounding boxes will be generated at the position of the same object. These candidate bounding boxes may overlap with each other. At this time, non-maximum suppression needs to be used to find the best object bounding box and eliminate redundant bounding boxes.

[0044] The following describes the exemplary applications of the behavior recognition device provided in the embodiments of the present application. The device provided in the embodiments of the present application can be implemented as various types of user terminals such as laptops, tablets, desktop computers, and mobile devices (for example, personal digital assistants, dedicated messaging devices, portable game devices) with data processing functions, or can also be implemented as a server. Below, the exemplary applications when the device is implemented as a terminal or a server will be described.

[0045] This method can be applied to an electronic device. The functions implemented by this method can be realized by a processor in the electronic device calling program code. Of course, the program code can be stored in a computer storage medium. It can be seen that the electronic device at least includes a processor and a storage medium.

[0046] The embodiments of the present application provide a behavior recognition method, as Figure 1 shown, and will be described in conjunction with the steps as Figure 1 shown:

[0047] Step S101, in a video stream including an object to be recognized, determine a video frame sequence.

[0048] In some embodiments, the video stream can be video data collected in any scene. For example, video data collected by a camera in the scene in any scene, or video data received from other devices. The object to be recognized can be a movable object in the scene where the video stream is located; the object to be recognized can be one or more. For example, if the video stream is collected for pedestrians, then the object to be recognized is pedestrians. If the video stream is collected in a traffic scene, then the object to be recognized can be vehicles in the traffic scene; if the video stream is an image collected on the grassland, then the object to be recognized can be animals such as cattle and sheep on the grassland.

[0049] In some possible implementations, the video stream may be video data collected by a camera in a short period of time in the scene, for example, a 3-second video stream collected for a pedestrian. The video frame sequence of the video stream is a plurality of video frames obtained by sampling the video stream; for example, by equally spaced sampling of a 3-second video stream, 8 video frames are obtained, that is, the video frame sequence; it may also be a plurality of video frames obtained by randomly sampling a 3-second video stream.

[0050] Step S102, in the video frame sequence, determine at least one first image region where the object to be recognized is located.

[0051] In some embodiments, in each video frame of the video frame sequence, determine the first image region where the object to be recognized is located in each video frame. The first image region is obtained by preprocessing the detection box of the object to be recognized in the video frame, and the preprocessing process includes expanding the size of the detection box and screening the expanded detection box.

[0052] In some possible implementations, first, in each video frame of the video frame sequence, detect the object to be recognized to obtain the detection box of each object to be recognized in the video frame; if there are multiple objects to be recognized in the video frame, then detect the multiple objects to be recognized in the video frame to obtain the detection box of each object to be recognized, that is, multiple detection boxes. Then, adaptively expand the length and width of each detection box by a certain proportion. Finally, in each video frame, perform non-maximum suppression in space on the multiple expanded detection boxes to eliminate redundant expanded detection boxes. In this way, by expanding the detection box of the object to be recognized in the video frame, the first image region with higher picture quality is selected, redundant detection boxes can be eliminated, and thus the calculation amount can be reduced.

[0053] Step S103, based on the at least one first image region, classify the behavior of the object to be recognized to obtain a classification result.

[0054] In some embodiments, the first image region in each video frame can be obtained through the above step S102. The first image region in each video frame can be input into a classification network, and based on the first image region in each video frame, classify the behavior of the object to be recognized; it can also be to select a target video frame at a preset position in the video frame sequence, and input the first image region in the target video frame into the classification network to classify the behavior of the object to be recognized.

[0055] In some possible implementation manners, first, a target video frame arranged at a preset position is selected from a video frame sequence; then, a first image region in each target video frame is input into a classification network to obtain the category to which the behavior of the object to be recognized in each target video frame belongs. The classification result includes the confidence of the behavior of the object to be recognized belonging to each category, where the classification categories include: preset behaviors and non-preset behaviors. In a specific example, if the object to be recognized is a vehicle, then the preset behavior may be the behavior of the vehicle having a traffic accident, for example, collision, rolling, scratching, etc.; the non-preset behavior is the behavior other than the vehicle having a traffic accident, for example, normal driving, normal parking, etc. If the object to be recognized is multiple pedestrians, then the preset behavior may be whether there is a fighting behavior among the multiple pedestrians; the non-preset behavior is the behavior other than fighting, for example, normal walking or multiple people walking together normally, etc. In this way, in the classification result, the confidence of the behavior of the object to be recognized belonging to the preset behavior can be obtained.

[0056] Step S104, in the at least one first image region, select a second image region whose classification result meets a preset condition.

[0057] In some embodiments, in the first image region included in each video frame for behavior classification, a second image region whose classification result meets a preset condition is selected. For example, if there are three video frames for behavior classification, then in the first image region included in each of these three video frames, a second image region whose classification result meets a preset condition is selected, so that a second image region whose classification result meets the preset condition in each video frame can be obtained. The preset condition may be that the confidence of the preset behavior in the classification result of the behavior of the object to be recognized is greater than a certain threshold, that is, the confidence of the behavior of the object to be recognized belonging to the preset behavior is relatively high. For at least one first image region of each video frame, select a first image region whose confidence of the category corresponding to the preset behavior in the classification result is greater than a certain threshold to obtain the second image region of this video frame.

[0058] In some possible implementation manners, if the first image regions of multiple target video frames are input into a classification network to classify whether the behavior of the object to be recognized belongs to a preset behavior; then, in at least one first image region of each target video frame, select a first image region whose confidence of belonging to the preset behavior is greater than a certain threshold as the second image region, so that the second image regions corresponding to each target video frame can be obtained. In a specific example, if the object to be recognized is multiple pedestrians and the preset behavior is the behavior of multiple people fighting, after classifying the behavior of the object to be recognized in each first image region of the target video frame, select a first image region whose confidence of the behavior belonging to the fighting behavior is greater than a certain threshold as the second image region of this target video frame, so that the second image regions of each target video frame can be obtained.

[0059] Step S105: Based on the second image region, identify the behavior of the object to be recognized to obtain an identification result.

[0060] In some embodiments, there is at least one second image region. The second image regions in each video frame are obtained through the above step S104. By merging the obtained multiple second image regions according to different video frames, the behavior of the object to be recognized can be identified based on the merged region, and the identification result can be obtained. The identification result includes the category to which the behavior of the object to be recognized belongs.

[0061] In some possible implementation manners, for a video frame including a second image region, select any one second image region in each video frame, and spatially merge the second image regions selected from multiple video frames to obtain a merged region. Identifying the behavior of the object to be recognized based on this merged region can obtain an accurate identification result.

[0062] In the embodiments of the present application, in the video frame sequence of the video stream, first determine the first image region of the object to be recognized in each video frame; then, classify the behavior of the object to be recognized in the first image region, and select a second image region that meets the preset conditions from the first image region of each frame according to the classification result; in this way, the subsequent number of identifications can be reduced and the calculation amount can be reduced. Finally, identify the behavior of the object to be recognized through the second image region in the video frame to obtain an identification result; in this way, identifying the behavior of the object to be recognized based on the second image region whose classification effect meets the preset conditions can improve the effective receptive field of the recognition network, thereby improving the accuracy of behavior recognition.

[0063] In some embodiments, by preprocessing the detection frame of the object to be recognized in the video frame, the first image region in each video frame is obtained, that is, the above step S102 can be implemented through Figure 2 the steps shown as follows:

[0064] Step S201: Detect the object to be recognized in each video frame to obtain multiple detection frames of the object to be recognized.

[0065] In some embodiments, there are at least two objects to be recognized. In each video frame of the video frame sequence, detect the object to be recognized to obtain the detection frames of each object to be recognized, that is, multiple detection frames.

[0066] In some possible implementation manners, the video frame sequence is input into a detection network to detect the object to be recognized, and the detection result marked with a detection box is output. For example, if the object to be recognized is multiple pedestrians, then by detecting the multiple pedestrians in each video frame, the detection box of each pedestrian in the video frame is obtained, so that each video frame includes at least one detection box.

[0067] Step S202, adjust the areas of the multiple detection boxes in each of the video frames to obtain multiple adjusted regions.

[0068] In some embodiments, the length and width sizes of each detection box are adaptively expanded outwards according to the preset ratio to increase the coverage area of the detection box. For example, the length and width of each detection box are adaptively expanded by 1.5 times to obtain the adjusted region corresponding to the detection box. In the video frame where the detection box is located, the side length of the detection box is expanded outwards by the preset ratio to obtain the adjusted region.

[0069] Step S203, determine the at least one first image region from the multiple adjusted regions in each of the video frames.

[0070] In some embodiments, select some of the adjusted regions from the multiple adjusted regions in each video frame as the at least one first image region. For example, select one adjusted region from the multiple adjusted regions in each video frame as the first image region within the video frame. Or, screen the multiple adjusted regions in each video frame to obtain at least one first image region within the video frame. In this way, after adjusting the areas of the multiple detection boxes in each video frame, select a part of the multiple adjusted regions as the first image region of the video frame, which can reduce duplicate recognition.

[0071] In some embodiments, perform non-maximum suppression in the spatial domain on the multiple adjusted regions of each video frame to eliminate redundant regions in at least one adjusted region, and obtain a first image region with higher image quality. That is, the above step S203 can be implemented through the following steps S231 to S233 (not shown in the figure):

[0072] Step S231, determine the first adjusted region with the highest confidence from the multiple adjusted regions in each of the video frames.

[0073] In some embodiments, in the multiple adjusted regions of the video frame, determine the adjusted region with the highest confidence of the detection box, that is, the first adjusted region; the first adjusted region obtained in this way is the region with the highest confidence in detecting the object to be recognized, indicating that the clarity and integrity of the object to be recognized detected in the first adjusted region are the best.

[0074] Step S232: determining a second adjusted region whose degree of overlap with the first adjusted region is greater than a preset overlap threshold.

[0075] In some embodiments, among the multiple adjusted regions in each video frame, the overlap between each adjusted region and the first adjusted region is determined, so as to determine the adjusted region whose overlap is greater than a preset overlap threshold, i.e., the second adjusted region. The second adjusted region may be one or more. In a video frame, if there are 5 adjusted regions, by determining the overlap between four of the adjusted regions and the first adjusted region, if the overlap between two of the adjusted regions and the first adjusted region is greater than the preset overlap threshold, then the two adjusted regions are used as the second adjusted region in the video frame.

[0076] Step S233 , among the multiple adjusted regions of each video frame, remove the second adjusted region whose area is smaller than a preset area threshold, to obtain the at least one first image region of each video frame.

[0077] In some embodiments, among the second adjusted regions in the video frame, a second adjusted region whose area is smaller than a preset area threshold is selected; among the multiple adjusted regions in the video frame, such second adjusted regions are deleted, and the remaining adjusted regions are used as higher quality first image regions.

[0078] In an embodiment of the present application, multiple detection frames of the object to be identified in each video frame are enlarged and then the adjusted area with higher quality is selected as the first image area; thereby, the amount of calculation for behavior recognition in the adjusted area can be reduced, and the behavior recognition can be performed using the first image area with higher quality, which can improve the accuracy of recognition.

[0079] In some embodiments, by selecting a small number of target video frames in the video frame sequence to classify the behavior of the object to be identified, the computational overhead can be reduced, that is, the above step S103 can be implemented by the following steps S131 and S132 (not shown):

[0080] Step S131 : Selecting video frames with a number less than a preset number of frames from the video frame sequence as target video frames.

[0081] In some embodiments, the preset number of frames is less than the total number of frames in the video frame sequence, for example, the preset number of frames is set to be much smaller than the total number of frames. In the video frame sequence, a small number of video frames are selected as target video frames. In some possible implementations, the video frames arranged at preset positions in the video frame sequence are determined as target video frames, so that the same number of target video frames are determined at several preset positions. For example, the preset positions are the first position, the middle position, and the last position, then the target video frames include the first frame, the middle frame, and the last frame.

[0082] Step S132: Classify the behavior of the object to be recognized based on the first image region in each target video frame to obtain the classification result.

[0083] In some embodiments, input the first image region in each target video frame into a behavior classification network to classify the behavior of the object to be recognized, and obtain the classification results of multiple first image regions in the target video frame. Taking the target video frame including the first frame, intermediate frames, and last frame as an example, input the multiple first image regions in the first frame into the behavior classification network to obtain the classification result corresponding to each first image region in the first frame; at the same time, input the multiple first image regions in the intermediate frames and the last frame into the behavior classification network respectively to obtain the classification result corresponding to each first image region in the intermediate frames and the classification result corresponding to each first image region in the last frame.

[0084] In some possible implementation manners, it may be to classify whether the behavior of the object to be recognized is an abnormal behavior based on the first image region in each target video frame. Then, the classification result is the confidence that the behavior of the object to be recognized in the first image region belongs to an abnormal behavior, and the confidence that the behavior of the object to be recognized in the first image region belongs to a non-abnormal behavior.

[0085] In the embodiments of the present application, by selecting a small number of target video frames in the video frame sequence for object behavior classification, the computational amount for behavior classification can be reduced.

[0086] In some embodiments, in at least one first image region of each target video frame, select a second image region from the at least one first image region according to the classification result. That is, the above step S104 can be implemented by the following steps S141 and S142 (not shown in the figure):

[0087] Step S141: Determine the second confidence that the classification result of each first image region in the target video frame is a preset category.

[0088] In some embodiments, for each target video frame in the video frame sequence, determine the second confidence that the classification result of each first image region in the target video frame is a preset category. For example, if there are three first image regions in the target video frame, determine the second confidence that the classification result of the third first image region is a preset category. The preset category is determined based on the categories included in the classification process. For example, if the classification categories include abnormal behavior and non-abnormal behavior, then the preset category can be abnormal behavior. In a specific example, if the object to be recognized is a vehicle and a traffic accident is set as an abnormal behavior, then determine the confidence that a traffic accident occurs to the vehicle in each first image region.

[0089] Step S142: In the target video frame, determine the first image region where the second confidence level is greater than the preset confidence threshold as the second image region.

[0090] In some embodiments, the preset confidence threshold is greater than or equal to the minimum second confidence level corresponding to the target video frame, or the preset confidence threshold is custom - set to a relatively large value. For example, the preset confidence threshold is set to 0.8. Among at least one first image region of the target video frame, the first image region with a classification result of a preset category and a confidence level greater than the preset confidence threshold is used as the second image region; thus, at least one second image region with a relatively high confidence level can be selected in each target video frame.

[0091] In some possible implementation manners, the user custom - sets the preset confidence threshold; or different confidence thresholds can be set for different target video frames, and the confidence threshold can be set based on the minimum second confidence level corresponding to at least one first image region of the target video frame. For example, the confidence threshold is set to be greater than or equal to the minimum second confidence level. In this way, by analyzing the second confidence level, in at least one first image region of each target video frame, the first image region with a relatively high second confidence level within the target video frame can be filtered out, and this first image region with a relatively high second confidence level is used as the second image region. In this way, by selecting the first image region with a relatively high second confidence level in the target video frame as the second image region, it is convenient to improve the performance of subsequent behavior recognition based on the second image region.

[0092] In some embodiments, when the target video frame includes at least one frame, by merging any second image regions in different target video frames and performing behavior recognition on the original video frame sequence based on the merged image region, the perception of the effective region during the recognition process is improved. That is, the above - mentioned step S105 can be implemented through Figure 3 the steps shown as follows:

[0093] Step S301: Based on the second image regions in the at least one target video frame, determine at least one target region sequence in the video frame sequence.

[0094] In some embodiments, each target video frame includes at least one second image region. Any second image regions in different target video frames are merged spatially, and according to the merged region, region extraction is performed on the original video frame sequence to obtain the target region sequence. In this way, based on multiple merged regions, multiple target region sequences can be extracted from the video frame sequence.

[0095] In some possible implementation manners, by merging a selected second image region in each target video frame and cropping a region in the original video frame sequence according to the merged region, a target region sequence is obtained. That is, the above step S301 can be implemented by the following steps S311 to S313 (not shown in the figure):

[0096] Step S311, select any second image region in the second image regions of each target video frame of the at least one target video frame to obtain at least one set of second image regions.

[0097] In some embodiments, a second image region is randomly selected in each frame of the target video frame. Thus, for several frames of the target video frame, several second image regions are obtained. For example, if the target video frame includes a first frame, an intermediate frame, and a last frame, a second image region is respectively selected in the first frame, the intermediate frame, and the last frame to obtain a set of second image regions, and the set of second image regions includes three second image regions; in at least one second image region of the first frame video frame, the intermediate frame video frame, and the last frame video frame, one second image region is selected from each to obtain the at least one set of second image regions. For example, if each of the first frame, the intermediate frame, and the last frame includes two second image regions, and one second image region is randomly selected from the first frame, the intermediate frame, and the last frame respectively, then eight sets of second image regions can be obtained, and each set of second image regions includes three second image regions. In this way, by obtaining multiple sets of second image regions, it is convenient to perform merging according to the multiple sets of second image regions subsequently, and the merged regions are enriched.

[0098] Step S312, merge the second image regions in each set of second image regions to obtain at least one merged region.

[0099] In some embodiments, by selecting a second image region in each target video frame and merging the selected second image regions, multiple merged regions can be obtained based on the second image regions in multiple target video frames. The multiple second image regions in each set of second image regions are merged spatially to obtain a merged region corresponding to the set of second image regions. In this way, the number of merged regions is the same as the number of sets.

[0100] In some possible implementation manners, multiple second image regions in a set of second image regions may be enclosed in a frame, and the region covered by this frame is the merged region.

[0101] Step S313, in the video frame sequence, determine a target region sequence matching each merged region to obtain the at least one target region sequence.

[0102] In some embodiments, for any merging region, matte extraction is performed on the video frame sequence according to the merging region to extract the image region corresponding to the merging region, thereby obtaining a sequence of target regions. In this way, for several merging regions, the same number of sequences of target regions can be determined. The sequence of target regions can represent the behavior trajectory of the object to be recognized.

[0103] In the embodiments of the present application, by selecting a second image region in each target video frame, merging the selected second image regions, and performing region extraction on the video frame sequence according to the merging region, a sequence of target regions matching each merging region can be obtained, enabling the picture content in the sequence of target regions to focus more on the object to be recognized itself and improving the effective receptive field in practice.

[0104] Step S302: Identify the behavior of the object to be recognized in the at least one sequence of target regions to obtain the recognition result.

[0105] In some embodiments, since the picture of the target region in the sequence of target regions focuses on the object to be recognized and reduces the interference of most irrelevant information in the picture, by inputting multiple sequences of target regions into the behavior recognition network, the behavior recognition network can be more focused on recognizing the behavior of the object to be recognized and more concerned about how to distinguish different motion details of the object to be recognized, rather than the differences between the object to be recognized and irrelevant objects.

[0106] In some possible implementation manners, by adjusting the side length of the target region, the side lengths of the target regions in the sequence of target regions are unified, facilitating the behavior recognition of the object to be recognized in the sequence of target regions. That is, the above step S302 can be implemented through the following steps S321 and S322 (not shown in the figure):

[0107] Step S321: Adjust the side length of the target region in each sequence of target regions to a preset side length to obtain an adjusted sequence of target regions.

[0108] In some embodiments, the side length of each target region is adjusted using a preset side length to obtain an adjusted target region. For the target region in any sequence of target regions, the length and width of the target region are adjusted to the preset side length. For example, the length and width of the target region are both adjusted to 224. In this way, the size of the adjusted target region in the adjusted sequence of target regions is 224×224.

[0109] Step S322: Identify the behavior of the object to be recognized in each adjusted sequence of target regions to obtain the recognition result.

[0110] In some embodiments, by inputting each adjusted target region sequence into a behavior recognition network for behavior recognition, the confidence level of whether the behavior of the object to be recognized is a preset behavior is obtained. In this way, by adjusting the side length of the target region in the target region sequence to a unified length, it is convenient for subsequent behavior recognition and can improve the efficiency of behavior recognition.

[0111] Next, an exemplary application of the embodiments of the present application in an actual application scenario will be described. Taking the recognition of the preset behavior of pedestrians in a video collected in a complex scenario as an example, the description will be carried out.

[0112] Anomaly detection in videos is an important issue in the field of computer vision and has a wide range of applications in the field of video management, such as detecting traffic accidents and some uncommon events, etc. Tens of thousands of video capture cameras are deployed worldwide. However, most cameras only record the dynamics at each moment and do not have the ability of automatic management (usually requiring personnel to view manually). Due to the huge number of videos, it is obviously not very realistic to filter the content in the videos only by manpower. Therefore, it is necessary to use computer vision and deep learning technologies to automatically detect abnormal events occurring in the videos.

[0113] In the related art, it is extremely difficult to recognize preset behaviors in videos. For example, due to small-probability events, the scarcity of labeled data, large inter-class / intra-class variances, subjective differences in the definition of abnormal events, low resolution of managed videos, etc.

[0114] For the detection of preset behaviors in video scenarios, how to accurately locate the preset behavior occurrence area in the entire picture (from different perspectives) of the video frame sequence, and then use the local area to replace the entire picture and input it into the recognition network for behavior classification, which helps to improve the effective perception range of the machine for target events, reduce the interference of most irrelevant information in the picture, and make the model more focused on how to distinguish different motion details of the main person rather than the differences between the target population and irrelevant passers-by. At the same time, it supports indoor and outdoor general scenarios such as urban streets and rail transit, enabling the automatic analysis of preset behaviors in video content to provide convenient services for users.

[0115] In the related art, in the scenario of behavior recognition, after performing full-image data augmentation or other preprocessing on the input video sequence, it is sent to the classification model for prediction. However, this method is only applicable to human-centered video behavior recognition. For videos captured by cameras, they often contain more information and cover a larger field of view. At the same time, the occurrence position of the preset behavior and the human body scale are also random. Therefore, simply using the full image as the model input is obviously unreasonable. In this way, by introducing prior information related to the category for region extraction, the results of each frame are unstable, which easily leads to an overly large extraction range.

[0116] Based on this, an embodiment of the present application provides a behavior recognition method, as Figure 4 shown Figure 4 is another schematic diagram of the implementation process of the behavior recognition method provided by the embodiment of the present application. The following description will be made in combination with Figure 4 the steps shown below:

[0117] Step S401: In the video frame sequence, extract the detection frame of the pedestrian in each video frame.

[0118] In some embodiments, obtain the video data captured by the camera, sample the video data to obtain the full-image video frame sequence; call the upstream structured detection model to extract the detection frame of the pedestrian in the video frame sequence, as Figure 5 shown by pedestrian detections 501 and 502 in

[0119] Step S402: Enlarge each detection frame by m times to obtain the adjusted area corresponding to each detection frame.

[0120] In some embodiments, both the length and width of each detection frame are enlarged by m times to increase the area of the detection frame. For example, m can be set to 1.5. In this way, the image area corresponding to each detection frame is the area obtained by enlarging both the length and width of the detection frame by m times; as Figure 5 shown, the adjusted area 511 obtained by enlarging the pedestrian detection frame 501 by m times, and the adjusted area 521 obtained by enlarging the pedestrian detection frame 502 by m times.

[0121] Step S403: Sort each adjusted area based on the area of the adjusted area to obtain a sorting result.

[0122] In some embodiments, the sorting result is obtained by sorting the enlarged detection frames in descending order according to the area of the adjusted area.

[0123] Step S404: Based on the sorting result, perform non-maximum suppression in the space on multiple adjusted areas within the same frame to obtain the first image area within the same frame.

[0124] In some embodiments, since the collected video is collected for pedestrians, there may be people close to each other in the same frame image. At this time, the corresponding adjusted areas will overlap in space. To reduce duplicate recognition, perform non-maximum suppression on the adjusted areas, sort the multiple adjusted areas in descending order according to the area of the adjusted area, and discard the adjusted areas with high overlap and small area within the same frame to obtain the first image area within the frame.

[0125] Step S405: Perform binary classification on multiple first image regions corresponding to the first frame, middle frames, and last frame of the video frame sequence to obtain a classification result.

[0126] In some embodiments, multiple first image regions corresponding to the first frame, middle frames, and last frame in the video frame sequence are input into a binary classification model for binary classification to determine whether the picture content in the first image region belongs to the abnormal class or the non-abnormal class. The classification result includes the confidence that each first image region belongs to the abnormal class.

[0127] Step S406: Based on the classification result, determine second image regions in the multiple first image regions where the confidence of the abnormal class is greater than a preset confidence threshold to obtain a set of second image regions.

[0128] In some embodiments, after the multiple first image regions are discriminated by the classification network, a score indicating that each first image region belongs to the abnormal class will be output. The first image regions in the first half of the scores in each of the first frame, middle frames, and last frame are used as the second image regions. As Figure 6 shown, the set of second image regions includes those shown as second image regions 61 to 68; among them, the second image region 64 is an image region where the detection fails, that is, no valid pedestrians are detected in this image region.

[0129] Step S407: Take one second image region from each of the first frame, middle frames, and last frame respectively and merge them to obtain multiple merged regions.

[0130] In some embodiments, in the set of second image regions, one second image region is taken from each of the first frame, middle frames, and last frame respectively and merged to obtain one merged region, and so on, to obtain multiple merged regions. As Figure 6 shown, by merging one second image region in different frames among the second image regions 61 to 68, merged regions 601, merged regions 602, and merged regions 603 are obtained.

[0131] Step S408: Based on the merged regions, extract corresponding image region sequences in the video frame sequence to obtain multiple target region sequences.

[0132] In some embodiments, corresponding regions are extracted from the merged regions on the original video frame sequence respectively to obtain multiple target region sequences.

[0133] In some possible implementation manners, taking the determination of the Kth target region as an example, the long side of the Kth target region is scaled to 224, the short side of the Kth target region is scaled proportionally, and black edges are filled up and down for the regions in the target region that are less than 224. Finally, the size of the target region is 224×224.

[0134] Step S409: Based on multiple target region sequences, identify the behaviors of pedestrians in the picture.

[0135] In some embodiments, input multiple target region sequences into a video classification network to identify whether the behaviors of pedestrians in the multiple target regions are abnormal. As Figure 5 shown, first, after performing spatial non-maximum suppression on the adjusted region 511 and the adjusted region 521, input the obtained first image regions into a classification network model to obtain the scores of abnormal behaviors of pedestrians included in each first image region; then, use the first image regions with scores greater than a preset threshold in the classification results as the second image regions to obtain a set of second image regions; by merging the second image regions within the same frame and cropping the image regions based on the merged regions in the original frame; finally, input the target region sequences into the network model 503 to identify whether the regions include abnormal behaviors of pedestrians, obtain the scores 512 of abnormal behaviors of pedestrians included in each target region sequence, and output the target regions with scores greater than the preset threshold in the set 504; output the target regions with scores less than the preset threshold in the set 505. In this way, the perception ability of the network model for target regions in the video is effectively improved, the retrieval range and the amount of calculation are greatly reduced, the preprocessing estimation results of the network model for different abnormal behavior labels are more stable, the effective receptive field of the event is improved, and the recall rate of the preprocessing is also improved.

[0136] In the embodiments of the present application, the target regions are determined based on pedestrian detection frames and learnable preprocessed second image regions, which increases the effective perception region of the model and reduces the retrieval range of irrelevant backgrounds; moreover, the behaviors of the first image regions after non-maximum suppression in the first, last, and middle frames are identified in a learnable manner, which can improve the recognition accuracy; the abnormal second image regions with higher scores predicted in each frame are merged into corresponding target regions to identify the behaviors respectively, which can improve the recall rate of the second image regions and the effective perception region of the network model for events.

[0137] The embodiments of the present application provide a behavior recognition device. Figure 7 It is a schematic structural composition diagram of the behavior recognition device in the embodiments of the present application. As Figure 7 shown, the behavior recognition device 700 includes:

[0138] A first determination module 701, configured to determine a video frame sequence in a video stream including an object to be recognized;

[0139] A second determination module 702, configured to determine at least one first image region where the object to be recognized is located in the video frame sequence;

[0140] The first classification module 703 is configured to classify the behavior of the object to be recognized based on the at least one first image region, and obtain a classification result;

[0141] The first selection module 704 is configured to select a second image region whose classification result meets a preset condition from the at least one first image region;

[0142] The first recognition module 705 is configured to recognize the behavior of the object to be recognized based on the second image region, and obtain a recognition result.

[0143] In some embodiments, the second determination module 702 includes:

[0144] The first detection sub-module is configured to detect the object to be recognized in each video frame, and obtain a plurality of detection frames of the object to be recognized;

[0145] The second detection sub-module is configured to adjust the areas of the plurality of detection frames in each video frame, and obtain a plurality of adjusted regions;

[0146] The first adjustment sub-module is configured to determine the at least one first image region from the plurality of adjusted regions in each video frame.

[0147] In some embodiments, the first adjustment sub-module includes:

[0148] The first determination unit is configured to determine a first adjusted region with the highest first confidence level of the detection frame from the plurality of adjusted regions in each video frame;

[0149] The second determination unit is configured to determine a second adjusted region whose overlap with the first adjusted region is greater than a preset overlap threshold;

[0150] The first adjustment unit is configured to remove, from the plurality of adjusted regions in each video frame, the second adjusted regions whose areas are smaller than a preset area threshold, and obtain the at least one first image region of each video frame.

[0151] In some embodiments, the first classification module 703 includes:

[0152] The first selection sub-module is configured to select, from the video frame sequence, video frames with a number less than a preset number of frames as target video frames;

[0153] The first classification sub-module is configured to classify the behavior of the object to be recognized based on the first image region in each target video frame, and obtain the classification result.

[0154] In some embodiments, the first selection sub-module is further configured to: select the first video frame, the middle video frame, and the last video frame from the video frame sequence as the target video frames.

[0155] In some embodiments, when the target video frames in the video sequence include at least one first image region, the second determination module 702 includes:

[0156] A first determination sub-module, configured to determine a second confidence level that the classification result of each first image region in the target video frame is a preset category;

[0157] A second determination sub-module, configured to determine, in the target video frame, a first image region whose second confidence level is greater than a preset confidence threshold as the second image region.

[0158] In some embodiments, when the target video frames are at least one frame, the first recognition module 705 includes:

[0159] A third determination sub-module, configured to determine at least one target region sequence in the video frame sequence based on the second image regions in the at least one frame of target video frames;

[0160] A first recognition sub-module, configured to recognize the behavior of the object to be recognized in the at least one target region sequence to obtain the recognition result.

[0161] In some embodiments, the third determination sub-module includes:

[0162] A first selection unit, configured to select any second image region from the second image regions of each target video frame in the at least one target video frame to obtain at least one set of second image regions;

[0163] A first merging unit, configured to merge the second image regions in each set of second image regions to obtain at least one merged region;

[0164] A third determination unit, configured to determine, in the video frame sequence, a target region sequence matching each merged region to obtain the at least one target region sequence.

[0165] In some embodiments, when the target video frames include the first video frame, the middle video frame, and the last video frame, the first selection unit is further configured to: select one second image region from at least one second image region of the first video frame, the middle video frame, and the last video frame to obtain the at least one set of second image regions.

[0166] In some embodiments, the first recognition sub-module includes:

[0167] A second adjustment unit, configured to adjust the side length of the target region in each target region sequence to a preset side length, so as to obtain an adjusted target region sequence;

[0168] A first recognition unit, configured to recognize the behavior of the object to be recognized in each of the adjusted target region sequences, so as to obtain the recognition result.

[0169] It should be noted that the description of the above device embodiments is similar to the description of the above method embodiments, and has beneficial effects similar to those of the method embodiments. For the technical details not disclosed in the device embodiments of the present application, please refer to the description of the method embodiments of the present application for understanding.

[0170] It should be noted that in the embodiments of the present application, if the above-mentioned behavior recognition method is implemented in the form of a software function module and sold or used as an independent product, it can also be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the embodiments of the present application, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for causing an electronic device (which may be a terminal, a server, etc.) to execute all or part of the methods described in the embodiments of the present application. The foregoing storage medium includes: various media such as a USB flash drive, a removable hard disk, a read-only memory (ROM), a magnetic disk, or an optical disc that can store program codes. In this way, the embodiments of the present application are not limited to any specific combination of hardware and software.

[0171] Correspondingly, the embodiments of the present application further provide a computer program product, where the computer program product includes computer-executable instructions, and after the computer-executable instructions are executed, the steps in the behavior recognition method provided by the embodiments of the present application can be implemented.

[0172] Correspondingly, the embodiments of the present application further provide a computer storage medium, where computer-executable instructions are stored on the computer storage medium, and when the computer-executable instructions are executed by a processor, the steps of the behavior recognition method provided by the above embodiments are implemented.

[0173] Correspondingly, the embodiments of the present application provide an electronic device, Figure 8 which is a schematic structural diagram of the electronic device in the embodiments of the present application, as Figure 8As shown, the electronic device 800 includes: a processor 801, at least one communication bus, a communication interface 802, at least one external communication interface, and a memory 803. Among them, the communication interface 802 is configured to implement connection communication between these components. Among them, the communication interface 802 may include a display screen, and the external communication interface may include a standard wired interface and a wireless interface. Among them, the processor 801 is configured to execute an image processing program in the memory to implement the steps of the behavior recognition method provided in the above embodiments.

[0174] The descriptions of the above embodiments of the behavior recognition device, electronic device, and storage medium are similar to the descriptions of the above method embodiments, and have similar technical descriptions and beneficial effects to the corresponding method embodiments. Due to space limitations, reference may be made to the records of the above method embodiments, so they will not be repeated here. For the technical details not disclosed in the embodiments of the behavior recognition device, electronic device, and storage medium of this application, please refer to the descriptions of the method embodiments of this application for understanding.

[0175] It should be understood that the term "one embodiment" or "an embodiment" mentioned throughout the specification means that a specific feature, structure, or characteristic related to the embodiment is included in at least one embodiment of the present application. Therefore, the appearances of "in one embodiment" or "in an embodiment" throughout the specification do not necessarily refer to the same embodiment. In addition, these specific features, structures, or characteristics may be combined in any suitable manner in one or more embodiments. It should be understood that in various embodiments of the present application, the magnitudes of the serial numbers of the above processes do not mean the order of execution, and the execution order of each process should be determined by its function and internal logic, and should not constitute any limitation to the implementation process of the embodiments of the present application. The serial numbers of the embodiments of the present application are only for description and do not represent the advantages and disadvantages of the embodiments.

[0176] It should be noted that in this article, the term "comprising", "including" or any other variation thereof is intended to cover a non-exclusive inclusion, so that a process, method, article or device including a series of elements not only includes those elements, but also includes other elements not explicitly listed, or further includes elements inherent to such process, method, article or device. Without further limitations, the element defined by the statement "including one..." does not exclude the existence of additional identical elements in the process, method, article or device including the element.

[0177] In several embodiments provided by the present application, it should be understood that the disclosed devices and methods can be implemented in other ways. The device embodiments described above are only illustrative. For example, the division of the units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined, or can be integrated into another system, or some features can be ignored, or not executed. In addition, the coupling, direct coupling, or communication connection between the components shown or discussed with each other can be through some interfaces. The indirect coupling or communication connection of devices or units can be electrical, mechanical, or other forms.

[0178] The units described above as separate components may or may not be physically separated. The components shown as units may or may not be physical units; they can be located in one place or distributed to multiple network units; some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0179] In addition, each functional unit in the embodiments of the present application can be all integrated in a processing unit, or each unit can be separately used as a unit, or two or more units can be integrated in one unit; the above-mentioned integrated units can be implemented in the form of hardware, or in the form of a combination of hardware and software functional units. Those of ordinary skill in the art can understand that all or part of the steps of implementing the above method embodiments can be completed by hardware related to program instructions. The foregoing program can be stored in a computer-readable storage medium. When the program is executed, it executes the steps including the above method embodiments; and the foregoing storage media include: removable storage devices, read-only memory (ROM), magnetic disks, or optical disks and other various media that can store program codes.

[0180] Alternatively, if the above-integrated units of the present application are implemented in the form of software functional modules and sold or used as independent products, they can also be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the embodiments of the present application, in essence or the part that contributes to the prior art, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for causing an electronic device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the methods described in the various embodiments of the present application. The foregoing storage medium includes: various media that can store program codes such as removable storage devices, ROMs, magnetic disks, or optical discs. The above is only the specific implementation manner of the present application, but the protection scope of the present application is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present application can easily think of changes or substitutions, which should all be covered by the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims.

Claims

1. A behavior recognition method, characterized in that, The method includes: In a video stream including an object to be recognized, determining a sequence of video frames; In the sequence of video frames, determining at least one first image region where the object to be recognized is located; Based on the at least one first image region, classifying the behavior of the object to be recognized to obtain a classification result; When a target video frame in the sequence of video frames includes at least one first image region, determining a second confidence level that the classification result of each first image region in the target video frame is a preset class; in the target video frame, determining a first image region with the second confidence level greater than a preset confidence threshold as a second image region; the target video frame is a video frame in the sequence of video frames with a number less than a preset number of frames; When there is at least one target video frame, selecting any second image region from the second image regions in each target video frame of the at least one target video frame to obtain at least one set of second image regions; merging the second image regions in each set of second image regions to obtain at least one merged region; in the sequence of video frames, determining a target region sequence matching each merged region to obtain at least one target region sequence; Recognizing the behavior of the object to be recognized in the at least one target region sequence to obtain a recognition result.

2. The method according to claim 1, wherein The determining, in the sequence of video frames, at least one first image region corresponding to the object to be recognized includes: Detecting the object to be recognized in each video frame to obtain a plurality of detection frames of the object to be recognized; Adjusting the areas of the plurality of detection frames in each video frame to obtain a plurality of adjusted regions; In the plurality of adjusted regions in each video frame, determining the at least one first image region.

3. The method according to claim 2, wherein The determining, in the plurality of adjusted regions in each video frame, the at least one first image region includes: In the plurality of adjusted regions in each video frame, determining a first adjusted region with the highest first confidence level of the detection frame; Determining a second adjusted region with an overlap degree greater than a preset overlap degree threshold with the first adjusted region; Excluding, in the plurality of adjusted regions in each video frame, a second adjusted region with an area less than a preset area threshold to obtain the at least one first image region in each video frame.

4. The method according to any one of claims 1 to 3, characterized in that, The classifying, based on the at least one first image region, the behavior of the object to be recognized to obtain a classification result includes: Selecting, from the sequence of video frames, video frames with a number less than a preset number of frames as target video frames; Based on the first image regions in each target video frame, classifying the behavior of the object to be recognized to obtain the classification result.

5. The method according to claim 4, characterized in that The selecting, from the sequence of video frames, video frames with a number less than a preset number of frames as target video frames includes: Selecting the first video frame, the middle video frame, and the last video frame from the sequence of video frames as the target video frames.

6. The method according to claim 5, characterized in that, When the target video frames include the first video frame, the middle video frames, and the last video frame, selecting any second image region in the second image regions of each target video frame among the at least one target video frame to obtain at least one set of second image regions, including: Selecting one second image region from at least one second image region of the first video frame, the middle video frames, and the last video frame to obtain the at least one set of second image regions.

7. The method according to claim 1, wherein Identifying the behavior of the object to be identified in the at least one target region sequence to obtain the identification result, including: Adjusting the side length of the target regions in each target region sequence to a preset side length to obtain an adjusted target region sequence; Identifying the behavior of the object to be identified in each of the adjusted target region sequences to obtain the identification result.

8. An action recognition device, characterized in that The apparatus includes: A first determination module, configured to determine a video frame sequence in a video stream including an object to be identified; A second determination module, configured to determine at least one first image region where the object to be identified is located in the video frame sequence; A first classification module, configured to classify the behavior of the object to be identified based on the at least one first image region to obtain a classification result; A first selection module, configured to, when the target video frames in the video frame sequence include at least one first image region, determine a second confidence level that the classification result of each first image region in the target video frame is a preset category; in the target video frame, determine a first image region with the second confidence level greater than a preset confidence threshold as a second image region; the target video frame is a video frame in the video frame sequence that is less than a preset number of frames; A first identification module, configured to, when there is at least one target video frame, select any second image region in the second image regions of each target video frame among the at least one target video frame to obtain at least one set of second image regions; merge the second image regions in each set of second image regions to obtain at least one merged region; in the video frame sequence, determine a target region sequence matching each merged region to obtain at least one target region sequence; identify the behavior of the object to be identified in the at least one target region sequence to obtain the identification result.

9. A computer storage medium, characterized in that, The computer storage medium stores computer-executable instructions, which, when executed, can implement the method steps of any one of claims 1 to 7.

10. An electronic device, characterized in that, The electronic device includes a memory and a processor. When the processor runs the computer-executable instructions stored on the memory, it can implement the method steps of any one of claims 1 to 7.

Citation Information

Patent Citations

  • Video processing method, device and equipment and storage medium

    CN111741329A

  • Behavior recognition method and device, equipment and storage medium

    CN113920585A