Target behavior recognition method, electronic device, and storage medium

By performing object detection, morphology detection, attribute detection and behavior detection on video frames, and identifying target behaviors with the detection results, the problem of inability to accurately identify target behaviors in the prior art is solved, and higher recognition accuracy is achieved.

CN115578668BActive Publication Date: 2025-08-15ZHEJIANG DAHUA TECH CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202211124689.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-09-15
Publication Date
2025-08-15
Estimated Expiration
2042-09-15

AI Technical Summary

Technical Problem

The prior art cannot accurately identify the behavior of targets.

Method used

By performing object detection on the video to be identified, the target detection results of each video frame are obtained, and target morphology detection, target attribute detection and target behavior detection are performed separately. Combining the morphology detection results, attribute detection results and behavior detection results, the target behavior recognition results are obtained.

Benefits of technology

It improves the accuracy of target behavior recognition, adds detection dimensions, and can more accurately identify target behaviors.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115578668B_ABST
    Figure CN115578668B_ABST
Patent Text Reader

Abstract

This application discloses a target behavior recognition method, electronic device, and storage medium. The target behavior recognition method includes: obtaining a video to be recognized; performing target detection on the video to be recognized to obtain a target detection result corresponding to each video frame; performing target morphology detection, target attribute detection, and target behavior detection on each video frame based on the target detection result to obtain a morphology detection result, attribute detection result, and behavior detection result corresponding to each video frame; and obtaining a target behavior recognition result corresponding to the video to be recognized based on the morphology detection result, attribute detection result, and behavior detection result. Through the above methods, the application can accurately identify target behaviors.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of target recognition technology, and in particular to a target behavior recognition method, electronic device, and storage medium. Background Art

[0002] In the context of artificial intelligence, target behavior recognition has undergone significant changes in data collection scale, sample data form, and behavior analysis methods. Target behavior recognition has gradually become automated, information-based, and intelligent.

[0003] However, the shortcoming is that the related target behavior recognition cannot accurately identify the target behavior. Summary of the Invention

[0004] The main technical problem solved by this application is to provide a target behavior recognition method, electronic device and storage medium, which can accurately identify the target behavior.

[0005] In order to solve the above technical problems, a technical solution adopted in this application is: to provide a target behavior recognition method, the method comprising: obtaining a video to be recognized; wherein the video to be recognized includes continuous video frames; performing target detection on the video to be recognized, and obtaining a target detection result corresponding to each video frame; based on the target detection result, performing target morphology detection, target attribute detection and target behavior detection on each video frame respectively, and obtaining a morphology detection result, attribute detection result and behavior detection result corresponding to each video frame; obtaining a target behavior recognition result corresponding to the video to be recognized based on the morphology detection result, attribute detection result and behavior detection result.

[0006] Among them, based on the target detection results, target morphology detection, target attribute detection and target behavior detection are performed on each video frame respectively to obtain the morphology detection results, attribute detection results and behavior detection results corresponding to each video frame, including: performing target tracking on each video frame based on the target detection results to obtain the target tracking results corresponding to each video frame; performing key point detection on the video frame to obtain the key point detection results corresponding to each video frame; performing target morphology detection on the video frame based on the key point detection results and the target detection results to obtain the morphology detection results corresponding to each video frame; and performing target attribute detection on the video frame based on the target tracking results and the key point detection results to obtain the attribute detection results corresponding to each video frame; and performing target behavior detection on all video frames based on the target tracking results and the target detection results to obtain the behavior detection results corresponding to each video frame.

[0007] The target tracking result includes trajectory information. Target tracking is performed on each video frame based on the target detection result to obtain the target tracking result corresponding to each video frame, including: determining a target video frame from all video frames based on the target detection result; wherein the target video frame includes at least one target object; and forming trajectory information of the target object based on the target object in the target video frame and the target objects in the remaining video frames.

[0008] Among them, target morphology detection is performed on the video frames based on the key point detection results and the target detection results to obtain the morphology detection results corresponding to each video frame, including: determining the target video frame where the target object exists based on the target detection results; connecting the key points in the key point detection results corresponding to each target video frame according to a preset method to form a key point image of the target object; performing target morphology detection on the key point image to obtain the morphology detection results corresponding to each target video frame.

[0009] Among them, the target tracking result includes trajectory information, and target attribute detection is performed on the video frame based on the target tracking result and the key point detection result to obtain the attribute detection result corresponding to each video frame, including: determining the target video frame where the target object exists based on the target tracking result; performing target attribute detection on the target video frame based on the trajectory information corresponding to each target video frame and the key points in the key point detection result to determine the attribute detection result of the target object, wherein the attribute detection result includes at least one of a backpack, a hat and a water bottle.

[0010] Among them, target behavior detection is performed on all video frames based on the target tracking results and the target detection results to obtain the behavior detection results corresponding to each video frame, including: performing event analysis on all video frames based on the target tracking results and the target detection results to obtain the target event corresponding to each video frame; classifying the target events to obtain the behavior detection results.

[0011] Among them, the target tracking results include trajectory information. Based on the target tracking results and target detection results, event analysis is performed on all video frames to obtain the target event corresponding to each video frame, including: if the trajectory information of the target object in the video to be identified is abnormal in two adjacent video frames, the target event corresponding to the target object will be regarded as the key event.

[0012] Among them, before performing key point detection on the video frames and obtaining the key point detection results corresponding to each video frame, it includes: screening all video frames based on the target detection results and the target tracking results, and screening out video frames that meet the preset conditions; performing key point detection on the video frames to obtain the key point detection results corresponding to each video frame, including: performing key point detection on the video frames that meet the preset conditions, and obtaining the key point detection results corresponding to each video frame.

[0013] Among them, the target detection results include at least one of the head information, shoulder information, upper body information, front information, side information and back information of the target object in each video frame, and the preset condition is that the score of the head information, shoulder information, upper body information, front information, side information or back information is greater than the preset score.

[0014] Among them, after screening out the video frames that meet the preset conditions, it includes: selecting from the video frames that meet the preset conditions according to a preset ratio to obtain selected video frames; performing key point detection on the video frames to obtain the key point detection results corresponding to each video frame, including: performing key point detection on the selected video frames to obtain the key point detection results corresponding to each video frame.

[0015] In order to solve the above technical problems, another technical solution adopted in this application is: to provide an electronic device, which includes a memory and a processor, the memory is used to store program data, and the processor is used to execute the program data to implement the target behavior recognition method as described above.

[0016] In order to solve the above technical problems, another technical solution adopted in this application is: providing a computer-readable storage medium, which stores program data. When the program data is executed by a processor, it is used to implement the target behavior recognition method as described above.

[0017] The beneficial effect of the present application is that, different from the prior art, the present application performs relatively comprehensive target morphology detection, target attribute detection and target behavior detection on the video frames after target detection, and then obtains the target behavior recognition result corresponding to the video to be identified based on the morphology detection results, attribute detection results and behavior detection results, thereby adding detection dimensions such as target morphology detection and target attribute detection, and can improve the accuracy of target behavior recognition. BRIEF DESCRIPTION OF THE DRAWINGS

[0018] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following briefly introduces the drawings required for describing the embodiments. Obviously, the drawings described below are only some embodiments of the present application. For those skilled in the art, other drawings can be obtained based on these drawings without inventive efforts. Among them:

[0019] Figure 1 This is a flowchart of the first embodiment of the target behavior recognition method provided by this application;

[0020] Figure 2 It is a structural diagram of the key point image provided by this application;

[0021] Figure 3This is a flowchart of a complete embodiment of the target behavior recognition method provided by this application;

[0022] Figure 4 This is a structural diagram of an embodiment of an electronic device provided by the present application;

[0023] Figure 5 It is a structural diagram of an embodiment of a computer-readable storage medium provided by this application. DETAILED DESCRIPTION

[0024] The following will be combined with the drawings in the embodiments of this application to clearly and completely describe the technical solutions in the embodiments of this application. Obviously, the embodiments described are only part of the embodiments of this application, not all of the embodiments. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.

[0025] like Figure 1 As shown, the target behavior recognition method described in this application may include: Step 100: Obtain the video to be recognized. Step 200: Perform target detection on the video to be recognized to obtain the target detection result corresponding to each video frame. Step 300: Based on the target detection result, perform target morphology detection, target attribute detection, and target behavior detection on each video frame to obtain the morphology detection result, attribute detection result, and behavior detection result corresponding to each video frame. Step 400: Obtain the target behavior recognition result corresponding to the video to be recognized based on the morphology detection result, attribute detection result, and behavior detection result.

[0026] That is to say, this application targets the video to be identified, obtains the target detection results corresponding to each video frame, and then performs more comprehensive target morphology detection, target attribute detection and target behavior detection on the video frames after target detection. According to the morphology detection results, attribute detection results and behavior detection results, the target behavior recognition results corresponding to the video to be identified are obtained, which adds detection dimensions such as target morphology detection and target attribute detection, and can improve the accuracy of target behavior recognition.

[0027] The following is a detailed description of the first embodiment of the target behavior recognition method of the present application.

[0028] Step 100: Obtain the video to be identified.

[0029] The video to be identified includes continuous video frames.

[0030] In some embodiments, the video to be identified may be shot by a monocular camera or a binocular camera.

[0031] Step 200: Perform target detection on the video to be identified, and obtain target detection results corresponding to each video frame.

[0032] Optionally, an object detection algorithm may be used to detect the video to be identified. The object detection algorithm may be based on a deep learning model, and the application does not limit the specific type of the algorithm.

[0033] Among them, object detection is to detect all objects of interest in the image, such as people, animals, or other creatures, and then determine the category of the object and the position of the object in the image or in world coordinates.

[0034] In some embodiments, in addition to determining the category and location of the target, target detection may also detect the size of the target or various shapes of the target.

[0035] For example, assuming that the target object to be detected is a person, the target detection result obtained may include information such as the head, shoulders, upper body, front, side, and back of the person.

[0036] Optionally, in order to facilitate subsequent operations, the target detection results can be used to obtain the target detection results using the detection data set Ω odj {odj1, odj2, ..., odj n} for record keeping.

[0037] Among them, each data element in the detection data set represents the target detection result corresponding to each video frame, for example, odj n Represents the target detection result corresponding to the nth video frame.

[0038] In some embodiments, there may be multiple targets in the video to be identified. Therefore, in step 200, target detection is performed on the video to be identified, and the positions of multiple targets in the video to be identified can be determined. For example, the target detection results obtained after target detection can include the positional relationships of various targets, such as human bodies, objects, etc.

[0039] In some embodiments, an edge detection algorithm may be used to determine the position of an object in a video to be identified.

[0040] Step 300: Based on the target detection results, target morphology detection, target attribute detection and target behavior detection are performed on each video frame to obtain the morphology detection result, attribute detection result and behavior detection result corresponding to each video frame.

[0041] Optionally, a multi-task model can be used to perform target morphology detection, target attribute detection, and target behavior detection on each video frame.

[0042] For example, assuming that the target object to be detected is a person, the target morphology detection can be the detection of the human body's morphology. The human body's morphology can be the human body's limb movement information, which can be understood as a static movement representation. For example, whether the person in the video to be identified is in a state of holding his head with both hands, raising his hands, waving, pointing, falling to the ground, sitting on the ground, lying down, etc.

[0043] Similarly, assuming that the target object to be detected is a person, the target attribute detection can be to detect whether the person in the video to be identified is carrying a bag or wearing a hat, etc., or to detect the posture of the person in the video to be identified, such as whether the hands are crossed on the chest from the front or hunched from the back, or to detect the clothes of the person in the video to be identified, such as the color of the top, the style of the pants, the hairstyle, etc., and the gender of the person can be detected by the person's clothes, hairstyle and posture.

[0044] Generally speaking, human behavior can be roughly determined based on the object's morphology. However, the object's morphology represents the state presented independently in each video frame and cannot accurately determine the object's behavior. Therefore, it is necessary to detect the target's essential behavior based on the object's trajectory information or the relationship between different objects to obtain the behavior detection result.

[0045] Taking human form as an example, while it may represent a person's current external state, it doesn't necessarily represent their underlying behavior. For example, while object form detection may reveal a person lying prone, object behavior detection, based on object trajectory information or the relationships between different objects, can also detect information about the person's previous and subsequent states. For example, if the person was standing before lying prone, and then standing again afterward, based on the previous and subsequent states, it can be inferred that the person may have fallen and then stood up.

[0046] Step 400: Obtain a target behavior recognition result corresponding to the video to be recognized based on the morphology detection result, the attribute detection result, and the behavior detection result.

[0047] It's important to note that the target behavior recognition result here can differ from the behavior detection result mentioned in the previous step. This target behavior recognition result can include both body shape and attributes. For example, if the target behavior is someone falling and then standing up, and the target attribute result is someone carrying a bag, the target behavior recognition result for the video to be recognized might be someone carrying a bag and then falling and then standing up.

[0048] We can also say that the target behavior recognition result is a conscious activity, including the behavior subject, behavior object, behavior environment, and behavior means. For example, in the scenario of students in class, the target behavior recognition results may include the student lying on the table, the student listening attentively, or the student playing with the phone.

[0049] Among them, when students are playing with mobile phones, the students are the subject of the behavior, the mobile phone is the object of the behavior, and playing is the method used by the subject students to act on the object mobile phone. Class can be the objective environment for students to play with mobile phones.

[0050] Since target detection tends to locate and identify objects in a single frame, it cannot accurately identify the target's location due to factors such as the surrounding environment's spatiotemporal information. Therefore, in order to more efficiently determine the target's location, some embodiments will further track the target based on target detection and predict the target's trajectory information.

[0051] Specific steps may include:

[0052] Step 1: Track the target for each video frame based on the target detection result to obtain the target tracking result corresponding to each video frame;

[0053] For example, the target detection result of a certain video frame may be a person standing on the playground. After target tracking, the target tracking result may be two people playing ball on the playground.

[0054] Optionally, the target tracking result includes trajectory information, amplitude information, information on the relationship between multiple targets, and the like.

[0055] Similarly, in order to facilitate subsequent operations, the target tracking results can be obtained by first using the tracking dataset Ω otj {otj1, otj2, ..., otj n} for record keeping.

[0056] Among them, each data element in the tracking dataset represents the target tracking result corresponding to each video frame, for example, otj n Represents the target tracking result corresponding to the nth video frame.

[0057] In some embodiments, when the target tracking result includes trajectory information, performing target tracking on each video frame based on the target detection result to obtain the target tracking result corresponding to each video frame may include the following sub-steps:

[0058] Step 11: Based on the target detection results, determine the target video frame from all video frames.

[0059] The target video frame at least includes a target object.

[0060] Step 12: Generating trajectory information of the target object based on the target object in the target video frame and the target objects in the remaining video frames.

[0061] For example, the different positions of a target object in different video frames, or the relationship between multiple target objects in different video frames.

[0062] Specifically, assuming that target detection can detect that there are two cars in a certain video frame, such as car A and car B, target tracking can determine the corresponding trajectory information of the two cars based on the positional relationship between the two cars in the remaining video frames, and accurately distinguish which car is car A and which car is car B.

[0063] By tracking the target, we can make full use of the inter-frame information between the target video frames, the environmental information around the target, etc. to obtain the target's trajectory information, thereby identifying the target more efficiently and accurately.

[0064] Since the video to be identified includes many continuous video frames, after the video to be identified is subjected to target detection and target tracking, the number of video frames containing the target object obtained is still large, and the target morphology detection, target attribute detection and target behavior detection performed subsequently are all performed separately on each video frame. Therefore, in order to improve the speed of detection, some embodiments can perform preliminary screening on the video frames containing the target object obtained after the video to be identified is subjected to target detection and target tracking. The specific screening process can be based on the target detection results and target tracking results to screen all video frames and filter out video frames that meet the preset conditions. The preset conditions can be determined based on the target detection results and target tracking results. For example, all video frames can be screened based on whether there are target detection results and target tracking results. If the target detection result corresponding to the video frame is no target, it is considered that the video frame does not meet the preset conditions. If the target detection result corresponding to the video frame is that there is a target, it is considered that the video frame meets the preset conditions.

[0065] In some embodiments, the target detection result may include at least one of head information, shoulder information, upper body information, front information, side information, and back information of the target object in each video frame.

[0066] The preset condition may be that the score of the head information, shoulder information, upper body information, front information, side information or back information is greater than a preset score.

[0067] For example, a human body optimization algorithm can be used to score human body parts in target detection results based on clarity, occlusion range, posture, angle, and other related conditions. Alternatively, the trajectory information of target tracking results can be scored based on trajectory completeness and trajectory clarity. Based on the obtained scores, the optimal selection is performed to obtain the target optimization result. The optimally selected video frames are those that meet the preset conditions.

[0068] Optionally, in order to facilitate subsequent operations, the target optimization result can be used to obtain the target optimization result using the optimization data set Ω qej {qej1, qej2, ..., qej n} for record keeping.

[0069] Among them, each data element in the preferred data set represents the target preferred result corresponding to each video frame, for example, qej n Represents the target optimization result corresponding to the nth video frame.

[0070] In some embodiments, the scoring mechanism may train the network in advance based on a standard comparison library or the weight of each part information, and use the trained network for scoring.

[0071] In addition, when the number of video frames contained in the video to be identified is very large, although after performing target detection and target tracking on the video to be identified, the obtained video frames containing the target object are initially screened out to a certain extent, some video frames may still be screened out, but there may still be a large number of video frames obtained after the initial screening. Therefore, in some embodiments, after screening out the video frames that meet the preset conditions, further selection may be performed, specifically:

[0072] The video frames meeting the preset conditions are selected according to a preset ratio to obtain selected video frames.

[0073] For example, the target tracking results and the target optimization results can be combined, and then the target selection algorithm can be used to analyze and select the target to obtain the target selection result. The target optimization result includes the video frames that meet the preset conditions, and the target selection result can be the result obtained by further selecting the target optimization results according to the preset ratio.

[0074] Optionally, in order to facilitate subsequent operations, the target selection results can be used to select the target data set Ω. spi {spi1, spi2, ..., spi n} for record keeping.

[0075] Among them, each data element in the selected data set represents the target selection result corresponding to each video frame, such as spin Represents the target selection result corresponding to the nth video frame.

[0076] Among them, after each video frame completes the corresponding target detection, target tracking, and target optimization, the optimized video frame actually also has the corresponding target detection result and target tracking result.

[0077] The preset ratio can be 1:100 or 1:1000, or can be selected based on the frame rate. For example, the ratio n / m is used, where n and m are natural numbers greater than 1, n is less than m, and m represents the frame rate. For example, m is 30, 60, 90, or 120. This application does not limit the specific ratio relationship.

[0078] In addition, in order to better identify the shape and attributes of the target, key point analysis can be performed on the target after detection and tracking, that is, step 1: target tracking is performed on each video frame based on the target detection result. After obtaining the target tracking result corresponding to each video frame, step 2 can be performed.

[0079] Among them, step 2 can be to perform key point detection on the video frame to obtain the key point detection result corresponding to each video frame.

[0080] Among them, key points can be extracted through a top-down or bottom-up method.

[0081] Assuming that the detection target is a person, the detected key points can be various parts and joints of the human body, such as nose, right eye, left eye, right ear, left ear, right shoulder, left shoulder, right elbow, left elbow, right wrist, left wrist, right knee, left knee, right ankle, left ankle, neck, etc.

[0082] In some embodiments, key point detection may also be performed on video frames that meet preset conditions to obtain key point detection results corresponding to each video frame.

[0083] Alternatively, in some embodiments, key point detection may be performed on the selected video frames to obtain a key point detection result corresponding to each video frame.

[0084] For example, a key point algorithm analysis can be performed based on the target detection results, target tracking results, and target image picking results to obtain a key point detection result. In order to facilitate subsequent operations, the key point detection result can be used to obtain the key point detection result using the key point dataset Ω. kpi (kpi1, kpi2, ..., kpi n} for record keeping.

[0085] Among them, each data element in the key point data set represents the key point detection result corresponding to each video frame, such as kpin Represents the key point detection result corresponding to the nth video frame.

[0086] After the key point detection of the target is performed as described above, the target can be subjected to morphological detection, attribute detection, and behavior detection. For specific detection methods, please refer to the following steps 3, 4, and 5.

[0087] For example, step 3 may be to perform target morphology detection on the video frames based on the key point detection results and the target detection results to obtain a morphology detection result corresponding to each video frame.

[0088] For example, a preliminary morphological image can be obtained by connecting all key points, and then the preliminary morphological image is compared with a preset human morphology, and combined with the target detection result to determine the final target morphology.

[0089] Optionally, in order to facilitate subsequent operations, the morphological detection results can be first obtained by using the morphological data set Ω bai {bai1, bai2, ..., bai n} for record keeping.

[0090] Among them, each data element in the morphological data set represents the morphological detection result corresponding to each video frame, for example, n Represents the morphological detection result corresponding to the nth video frame.

[0091] Since key point detection can be implemented using a trained network model, such as a deep learning model, after detecting key points, the corresponding information of each key point can be obtained. For example, based on key point 0, it can be concluded that the part corresponding to the key point is the nose. Therefore, in some embodiments, all key points are connected according to a preset method to obtain the target shape. Specifically, it can be as follows:

[0092] 1) Determine the target video frame where the target object exists based on the target detection result.

[0093] 2) The key points in the key point detection results corresponding to each target video frame are connected according to a preset method to form a key point image of the target object.

[0094] The preset method can be combined according to the characteristics of the human body structure. For example, if the key points detected are: "0"-"13" respectively correspond to the nose, right eye, left eye, right ear, left ear, right shoulder, left shoulder, right elbow, left elbow, right wrist, left wrist, right knee, left knee, and neck of the human body. Then, according to the structure of the human body, the numbers corresponding to each part or joint are connected to obtain the following: Figure 2 The keypoint image shown.

[0095] 3) Perform target morphology detection on the key point image to obtain the morphology detection result corresponding to each target video frame.

[0096] For example, Figure 2 The key point image shown is used to detect the target shape. By comparing it with the preset shape, it can be obtained Figure 2 The key point image in the middle corresponds to the shape of raising the hand.

[0097] Step 4: Based on the target tracking results and key point detection results, target attribute detection is performed on the video frames to obtain the attribute detection results corresponding to each video frame.

[0098] Target attributes such as backpacks, hats, and glasses can be determined based on target tracking results and key point locations. For example, if tracking a video frame reveals a backpack with no track and remains stationary, and the backpack is located at a key point on the human body, such as an arm, the target person can be identified as carrying a backpack.

[0099] Optionally, in order to facilitate subsequent operations, the attribute detection results can be used to obtain the attribute detection results using the attribute data set Ω pedi {ped1, ped2, ..., pedi n} for record keeping.

[0100] Among them, each data element in the attribute data set represents the attribute detection result corresponding to each video frame, such as pedi n Represents the attribute detection result corresponding to the nth video frame.

[0101] Optionally, in some embodiments, target attribute detection may be performed on the target video frame based on the target's trajectory information and key point information, as follows:

[0102] (1) Determine the target video frame where the target object exists based on the target tracking result.

[0103] (2) Target attribute detection is performed on the target video frame based on the trajectory information corresponding to each target video frame and the key points in the key point detection result, and an attribute detection result of the target object is determined, wherein the attribute detection result includes at least one of a backpack, a hat, and a water bottle.

[0104] Step 5: Based on the target tracking results and target detection results, target behavior detection is performed on all video frames to obtain the behavior detection results corresponding to each video frame.

[0105] Optionally, in some embodiments, event analysis can be performed on all video frames based on the target tracking results and target detection results. After obtaining the target event corresponding to each video frame, the target event is classified and the behavior detection result is obtained.

[0106] Among them, target events can be classified according to contingency, necessity, etc., or events can be classified according to the environment. For example, they can be emergencies, such as natural disasters, or public health events, social security events, etc.

[0107] This application does not limit how target events are specifically classified.

[0108] Due to limited resources, different types of events receive different levels of attention.

[0109] In some embodiments, some target events may be marked as events of varying degrees for attention.

[0110] For example, we can analyze key events based on the target detection results and target tracking results to obtain key event results. Similarly, in order to facilitate subsequent operations, we can first use the target detection results to obtain key event results using the key event dataset Ω. iej {iej1, iej2, ..., iej n} for record keeping.

[0111] Among them, each data element in the key event dataset represents the key event result corresponding to each video frame, for example, iej n Represents the target detection result corresponding to the nth video frame.

[0112] Among them, key events can be accidental injury, illness, fainting, etc. For example, iej1 can represent accidental injury, iej2 can represent illness, etc.

[0113] Similarly, in order to improve the detection speed, some embodiments can screen and select the obtained video frames containing the target object after performing target detection and target tracking on the video to be identified, or before performing event analysis on all video frames, and then perform key event analysis and behavior classification algorithm analysis based on the screening and selection results to obtain the target behavior detection results.

[0114] In order to facilitate subsequent operations, the behavior detection results can be used to obtain the behavior data set Ω scj {scj1, scj2, ..., scj n} for record keeping.

[0115] Among them, each data element in the behavior dataset represents the behavior detection result corresponding to each video frame, for example, scj n Represents the behavior detection result corresponding to the nth video frame.

[0116] Regarding the target screening and selection process involved in the behavior detection process, please refer to the relevant description of the above steps, and this application will not go into details here.

[0117] Based on the morphological detection results obtained in step 3, the attribute detection results obtained in step 4, and the behavior detection results obtained in step 5, the target behavior recognition result corresponding to the video to be recognized is finally obtained. Similarly, the target behavior recognition result can be obtained by using the recognition dataset Ω sci {Ω sci1 ,Ω sci2 ,…,Ω scin}express.

[0118] For example, taking the test object as a student, where Ω sci1 It can indicate the behavior of students lying on the table, Ω sci2 It can indicate that students listen carefully to the class. sci3 It can indicate students' behavior of playing with mobile phones, sci4 Can indicate the student falling down, Ω sci5 It can indicate students sitting on the ground, etc.

[0119] In other words, by adding detection dimensions such as target morphology detection and target attribute detection, the accuracy of target behavior recognition can be improved.

[0120] In combination with the above embodiments, the following describes a relatively complete embodiment of the present application. For details, please refer to Figure 3 , Figure 3 This is a flowchart of a complete embodiment of the present application, which may include the following steps:

[0121] (1) First obtain the video to be identified;

[0122] (2) performing target detection on the acquired video to be identified, and obtaining the target detection result corresponding to each video frame;

[0123] (3) Tracking the video frame after target detection through the target tracking algorithm to obtain the target tracking result;

[0124] (4) Analyze the target detection results and target tracking results through the target optimization algorithm to obtain the target optimization result;

[0125] (5) Merge the target tracking results and the target optimization results, and then use the image selection algorithm to select the target to obtain the target selection result;

[0126] (6) Based on the target detection results and target selection results, perform key point analysis on the video frame to obtain the key point detection results;

[0127] (7) Based on the target detection results and key point detection results, the human body morphology analysis is performed on the video frame to obtain the morphology detection results;

[0128] (8) Based on the target tracking results and key point detection results, the human attributes of the video frame are analyzed to obtain the attribute detection results;

[0129] (9) The optimized target detection results and target tracking results are first subjected to key event identification to obtain key event results, and then the key event results are analyzed by a behavior classification algorithm to obtain behavior detection results;

[0130] (10) Finally, a comprehensive analysis is performed based on the morphological detection results, attribute detection results, and behavior detection results to obtain the target behavior recognition results.

[0131] See Figure 4 , Figure 4 1 is a schematic diagram of the structure of an electronic device according to an embodiment of the present application. The electronic device 130 includes a memory 131 and a processor 132. The memory 131 is used to store program data, and the processor 132 is used to execute the program data to implement the following method:

[0132] Obtain a video to be identified; wherein the video to be identified includes continuous video frames; perform target detection on the video to be identified to obtain a target detection result corresponding to each video frame; based on the target detection result, perform target morphology detection, target attribute detection, and target behavior detection on each video frame to obtain a morphology detection result, an attribute detection result, and a behavior detection result corresponding to each video frame; obtain a target behavior recognition result corresponding to the video to be identified based on the morphology detection result, the attribute detection result, and the behavior detection result.

[0133] It can be understood that the processor 132 is further configured to execute program data to implement the method of any of the above embodiments.

[0134] Optionally, in one embodiment, the electronic device 130 can be a chip, a field programmable gate array (FPGA), a single chip microcomputer, etc., wherein the chip can be a processing chip such as a CPU, GPU, MCU, etc., or a storage chip such as a DRAM, SRAM, etc.

[0135] See Figure 5 , Figure 51 is a schematic diagram of the structure of an embodiment of a computer-readable storage medium provided by the present application. The computer-readable storage medium 140 stores program data 141. When the program data 141 is executed by a processor, it is used to implement the following method:

[0136] Obtain a video to be identified; wherein the video to be identified includes continuous video frames; perform target detection on the video to be identified to obtain a target detection result corresponding to each video frame; based on the target detection result, perform target morphology detection, target attribute detection, and target behavior detection on each video frame to obtain a morphology detection result, an attribute detection result, and a behavior detection result corresponding to each video frame; obtain a target behavior recognition result corresponding to the video to be identified based on the morphology detection result, the attribute detection result, and the behavior detection result.

[0137] It can be understood that when the program data 141 is executed by the processor, it is also used to implement the method of any of the above embodiments.

[0138] When the embodiments of the present application are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application is essentially or the part that contributes to the prior art or all or part of the technical solution can be embodied in the form of a software product, and the computer software product is stored in a storage medium, including a number of instructions to enable a computer device (which can be a personal computer, a server, or a network device, etc.) or a processor to execute all or part of the steps of the various implementation methods of the present application. The aforementioned storage medium includes: various media that can store program codes, such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk.

[0139] The above is only an implementation method of the present application and does not limit the patent scope of the present application. Any equivalent structure or equivalent process transformation made using the contents of the description and drawings of this application, or directly or indirectly applied in other related technical fields, are also included in the patent protection scope of the present application.

Claims

1. A target behavior recognition method, characterized in that: The method comprises: Acquire a video to be identified; wherein the video to be identified includes continuous video frames; Performing target detection on the video to be identified to obtain a target detection result corresponding to each video frame; Based on the target detection result, target morphology detection, target attribute detection and target behavior detection are performed on each of the video frames to obtain a morphology detection result, an attribute detection result and a behavior detection result corresponding to each of the video frames, including: performing target tracking on each of the video frames based on the target detection result to obtain a target tracking result corresponding to each of the video frames; performing key point detection on the video frames to obtain a key point detection result corresponding to each of the video frames; performing target morphology detection on the video frames based on the key point detection result and the target detection result to obtain the morphology detection result corresponding to each of the video frames; and performing target attribute detection on the video frames based on the target tracking result and the key point detection result to obtain the attribute detection result corresponding to each of the video frames; and performing target behavior detection on all of the video frames based on the target tracking result and the target detection result to obtain the behavior detection result corresponding to each of the video frames; A target behavior recognition result corresponding to the video to be recognized is obtained based on the morphology detection result, the attribute detection result, and the behavior detection result.

2. The method according to claim 1, characterized in that The target tracking result includes trajectory information. The target tracking is performed on each of the video frames based on the target detection result to obtain a target tracking result corresponding to each video frame, including: Based on the target detection result, determining a target video frame from all the video frames; wherein the target video frame includes at least one target object; The trajectory information of the target object is formed based on the target object in the target video frame and the target objects in the remaining video frames.

3. The method according to claim 1, characterized in that The performing target morphology detection on the video frames based on the key point detection results and the target detection results to obtain the morphology detection results corresponding to each of the video frames includes: Determining a target video frame in which a target object exists based on the target detection result; Connecting the key points in the key point detection results corresponding to each target video frame according to a preset method to form a key point image of the target object; Perform target morphology detection on the key point image to obtain the morphology detection result corresponding to each target video frame.

4. The method according to claim 1, wherein The target tracking result includes trajectory information, and the target attribute detection is performed on the video frame based on the target tracking result and the key point detection result to obtain the attribute detection result corresponding to each video frame, including: Determining a target video frame in which a target object exists based on the target tracking result; Based on the trajectory information corresponding to each target video frame and the key points in the key point detection result, target attribute detection is performed on the target video frame to determine the attribute detection result of the target object, wherein the attribute detection result includes at least one of a backpack, a hat and a water bottle.

5. The method according to claim 1, wherein The performing target behavior detection on all the video frames based on the target tracking result and the target detection result to obtain the behavior detection result corresponding to each video frame includes: Performing event analysis on all the video frames based on the target tracking result and the target detection result to obtain a target event corresponding to each video frame; The target event is classified to obtain the behavior detection result.

6. The method according to claim 5, characterized in that The target tracking result includes trajectory information, and the event analysis is performed on all the video frames based on the target tracking result and the target detection result to obtain the target event corresponding to each video frame, including: If the trajectory information of the target object in the to-be-identified video is abnormal in two adjacent video frames, the target event corresponding to the target object is taken as a key event.

7. The method according to claim 1, characterized in that Before performing key point detection on the video frames to obtain key point detection results corresponding to each video frame, the method includes: Filtering all the video frames based on the target detection result and the target tracking result to select video frames that meet preset conditions; The performing key point detection on the video frames to obtain a key point detection result corresponding to each of the video frames includes: Key point detection is performed on the video frames that meet the preset conditions to obtain a key point detection result corresponding to each video frame.

8. The method according to claim 7, characterized in that The target detection result includes at least one of the head information, shoulder information, upper body information, front information, side information and back information of the target object in each video frame, and the preset condition is that the score of the head information, the shoulder information, the upper body information, the front information, the side information or the back information is greater than the preset score.

9. The method according to claim 7, characterized in that After the video frames meeting the preset conditions are screened out, the method includes: Selecting video frames that meet preset conditions according to a preset ratio to obtain the selected video frames; The performing key point detection on the video frames to obtain a key point detection result corresponding to each of the video frames includes: Key point detection is performed on the selected video frames to obtain a key point detection result corresponding to each video frame.

10. An electronic device, characterized in that: The electronic device includes a memory and a processor, the memory is used to store program data, and the processor is used to execute the program data to implement the target behavior recognition method according to any one of claims 1 to 9.

11. A computer-readable storage medium, characterized in that The computer-readable storage medium stores program data, and when the program data is executed by a processor, it is used to implement the target behavior recognition method according to any one of claims 1 to 9.

Citation Information

Patent Citations

  • Behavior detection method, equipment and storage medium

    CN113673351A