Action Recognition Method, Device, Storage Medium, and Computer Equipment

By extracting the characteristics of the character and object detection frame of the central video frame from the space-time features of the video, the problem of low action recognition efficiency in the prior art is solved, and more efficient and accurate action recognition is achieved.

CN114495272BActive Publication Date: 2025-07-08GUANGDONG OPPO MOBILE TELECOMMUNICATIONS CORP LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202210069070.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-01-19
Publication Date
2025-07-08
Estimated Expiration
2042-01-19

AI Technical Summary

Technical Problem

The prior art is less efficient in the process of human body movement recognition in videos, and complicated image processing is required for multiple video frames.

Method used

By obtaining the central video frame of the target video and determining the detection frames of the characters and objects, based on these detection frames, the space-time characteristics of the characters and objects are extracted from the space-time characteristics of the video, and the action recognition is directly performed to avoid frame-by-frame processing.

Benefits of technology

The efficiency of action recognition is improved, the complicated process of image processing is reduced, and the accuracy of action recognition is improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114495272B_ABST
    Figure CN114495272B_ABST
Patent Text Reader

Abstract

The present application discloses an action recognition method, apparatus, storage medium, and computer device. The method includes: obtaining a central video frame in a target video; determining a person detection box corresponding to a target person in the central video frame and object detection boxes corresponding to at least one object in the central video frame, where the target person is any one of multiple persons in the central video frame; obtaining a person spatio-temporal feature corresponding to the target person in the video spatio-temporal feature corresponding to the target video based on the person detection box; obtaining an object spatio-temporal feature corresponding to the at least one object in the video spatio-temporal feature corresponding to the target video based on the object detection box; and recognizing the action of the target person based on the person spatio-temporal feature and the object spatio-temporal feature. By adopting the present application, the action recognition efficiency is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of computer vision technology, and more particularly, to a method and apparatus for action recognition, a storage medium, and a computer device. Background Art

[0002] Action recognition refers to recognizing the actions performed by a person in image data such as videos and pictures based on the received image data, and is mainly applied to the process of human-computer interaction, that is, based on the recognized actions, relevant control of the computer is performed. Summary of the Invention

[0003] This application provides a method and apparatus for action recognition, a storage medium, and a computer device, which can solve the technical problem of how to improve the efficiency of action recognition.

[0004] In a first aspect, an embodiment of this application provides a method for action recognition, the method comprising:

[0005] Obtain a central video frame in a target video, where the central video frame is a video frame located in the middle position in a video frame sequence corresponding to the target video, and the video frame sequence includes a plurality of the video frames arranged in the acquisition order;

[0006] Determine a person detection box corresponding to a target person in the central video frame and an object detection box corresponding to at least one object in the central video frame, where the target person is any one of multiple persons in the central video frame;

[0007] Based on the person detection box, obtain a person spatio-temporal feature corresponding to the target person in a video spatio-temporal feature corresponding to the target video;

[0008] Based on the object detection box, obtain an object spatio-temporal feature corresponding to the at least one object in a video spatio-temporal feature corresponding to the target video;

[0009] Based on the person spatio-temporal feature and the object spatio-temporal feature, recognize the action of the target person.

[0010] In a second aspect, an embodiment of this application provides an apparatus for action recognition, comprising:

[0011] A video frame acquisition module, configured to obtain a central video frame in a target video, where the central video frame is a video frame located in the middle position in a video frame sequence corresponding to the target video, and the video frame sequence includes a plurality of the video frames arranged in the acquisition order;

[0012] The detection box determination module is used to determine the person detection box corresponding to the target person in the central video frame and the object detection boxes corresponding to at least one object in the central video frame, where the target person is any one of multiple persons in the central video frame;

[0013] The feature acquisition module is used to acquire the person spatio-temporal feature corresponding to the target person in the video spatio-temporal feature corresponding to the target video based on the person detection box;

[0014] The feature acquisition module is further used to acquire the object spatio-temporal feature corresponding to the at least one object in the video spatio-temporal feature corresponding to the target video based on the object detection box;

[0015] The action recognition module is used to recognize the action of the target person based on the person spatio-temporal feature and the object spatio-temporal feature.

[0016] In a third aspect, an embodiment of the present application provides a storage medium storing a computer program, and the computer program is adapted to be loaded and executed by a processor to perform the steps of the above method.

[0017] In a fourth aspect, an embodiment of the present application provides a computer device, including a memory, a processor, and a computer program stored on the memory and executable on the processor, and when the processor executes the program, it implements the steps of the above method.

[0018] In the embodiment of the present application, by determining the central video frame in the target video, and then obtaining the person detection box and the item detection box in the central video frame, the person detection box and the object detection box can represent the approximate positions of the person and the item in the entire target video, so as to directly extract the relevant features for action recognition from the video spatio-temporal feature corresponding to the target video based on the person detection box and the item detection box, thereby eliminating the need to process each frame of the target video to obtain the effective features (i.e., person features and item features) in each video frame, and further reducing the complicated image processing process in the action recognition process and improving the action recognition efficiency. BRIEF DESCRIPTION OF THE DRAWINGS

[0019] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the following drawings are only some embodiments of the present application, and those skilled in the art can obtain other drawings without creative efforts based on these drawings.

[0020] Figure 1 It is a schematic flowchart of an action recognition method provided by an embodiment of the present application;

[0021] Figure 2 It is an example schematic diagram of each detection box in the central video frame provided by the embodiment of the present application;

[0022] Figure 3 It is a schematic flow chart of an action recognition method provided by the embodiment of the present application;

[0023] Figure 4 It is an example schematic diagram of each human body key point corresponding to the target person in the central video frame provided by the embodiment of the present application;

[0024] Figure 5 It is an example schematic diagram of each detection box in the central video frame provided by the embodiment of the present application;

[0025] Figure 6 It is a schematic flow chart of an action recognition method provided by the embodiment of the present application;

[0026] Figure 7 It is a schematic structural diagram of an action recognition device provided by the embodiment of the present application;

[0027] Figure 8 It is a schematic structural diagram of an action recognition device provided by the embodiment of the present application;

[0028] Figure 9 It is a schematic structural diagram of a computer device provided by the embodiment of the present application. Detailed implementation manners

[0029] To make the features and advantages of the present application more obvious and understandable, the technical solutions in the embodiments of the present application will be clearly and completely described below in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative efforts belong to the scope of protection of the present application.

[0030] In the existing action recognition method, when recognizing the human body actions in a video, first perform image processing on each video frame in multiple video frames in the video, then extract the human feature data in each video frame, and then perform action recognition based on multiple consecutive human feature data, resulting in low action recognition efficiency.

[0031] The following will be combined with Figures 1-6 to introduce in detail the action recognition method provided by the embodiment of the present application.

[0032] Please refer to Figure 1 , which provides a schematic flow chart of an action recognition method for the embodiment of the present application. As shown in Figure 1As shown, the method may include the following steps S101 - S104.

[0033] S101, obtain a central video frame in the target video, where the central video frame is the video frame located in the middle position in the video frame sequence corresponding to the target video, and the video frame sequence includes a plurality of the video frames arranged in the acquisition order.

[0034] In one embodiment, the action recognition device may include an image acquisition device such as a camera to collect video data in the environment where the action recognition device is located through the camera. The action recognition device may also directly receive video data sent by other devices.

[0035] The action recognition device obtains video data for identifying human actions as the target video, and then performs decoding processing on the target video to obtain a continuous video frame sequence. It should be noted that the video frame sequence includes a plurality of video frames arranged in the image acquisition order. Then, in the continuous video frame sequence, the video frame located in the middle position is determined, and this video frame is used as the central video frame in the target video.

[0036] Exemplarily, if the video frame sequence is A, B, C, D, E, then the central video frame is C. If the video frame sequence is A, B, C, D, then the central video frame can be B or C, which is not limited herein.

[0037] Further, in order to improve the accuracy of action recognition, when the action recognition device obtains video data, it splits the video data according to the video upper limit to obtain a plurality of target videos. Exemplarily, if the video upper limit is 10k and the obtained video data is 25k, then a 10k first target video, a 10k second target video, and a 5k third target video can be obtained. It should be noted that in this embodiment, feature extraction is performed on the entire target video based on the positions of the person and objects in the central video frame. Since the human body and objects cannot move over a large range in a short time, therefore, this embodiment can improve the accuracy of feature extraction by restricting the size of the target video, and further improve the accuracy of action recognition.

[0038] S102, determine a person detection box corresponding to the target person in the central video frame and an object detection box corresponding to at least one object in the central video frame, where the target person is any one of the multiple persons in the central video frame.

[0039] In one embodiment, the person detection box is used to represent the specific position of the person in the central video frame, that is, the area where the person is located. The object detection box is used to represent the specific position of the object in the central video frame, that is, the area where the object is located. Exemplarily, as Figure 2 shown Figure 2is the central video frame 100 corresponding to the target video, where Figure 2 also shown is a person detection box 110 corresponding to the target person in the central video frame 100, and an item detection box 120 corresponding to at least one object.

[0040] The action recognition device performs human body detection and object detection based on the central video frame, and then generates a person detection box and an item detection box based on the detection results. It should be noted that one person will have a corresponding person detection box, and one item will have a corresponding item detection box.

[0041] When there are multiple persons in the central video frame, the action recognition device selects any one of the multiple persons as the target person, and then obtains the person detection box corresponding to the target person. Among them, the action recognition device can first obtain the person detection boxes corresponding to each person in the central video frame, and then obtain the person detection box corresponding to the target person, or it can first determine the target person in the central video frame, and then obtain the person detection box corresponding to the target person based on the central video frame, which is not limited here. Further, the action recognition device will traverse each person in the central video frame one by one to obtain the human actions of each person based on the person detection boxes corresponding to each person.

[0042] S103. Based on the person detection box, in the video spatio-temporal feature corresponding to the target video, obtain the person spatio-temporal feature corresponding to the target person.

[0043] In one embodiment, the video spatio-temporal feature corresponding to the target video refers to the time-space feature of the target video. It can be understood that the video spatio-temporal feature at least includes the human feature, object feature in the target video, and the temporal change feature of each feature in the entire video. The person spatio-temporal feature can represent the posture change of the target person in the video, such as turning the head, waving the hand, etc.

[0044] The action recognition device, based on the person detection box, obtains the person spatio-temporal feature corresponding to the target person in the video spatio-temporal feature corresponding to the target video. Specifically, since the person detection box represents the specific position of the person in the central video frame, the spatio-temporal feature corresponding to this position can be extracted in the video spatio-temporal feature corresponding to the target video as the person spatio-temporal feature.

[0045] S104. Based on the object detection box, in the video spatio-temporal feature corresponding to the target video, obtain the object spatio-temporal feature corresponding to the at least one object.

[0046] In one embodiment, the action recognition device obtains the object spatio-temporal features corresponding to at least one object from the video spatio-temporal features corresponding to the target video based on the object detection frames of the at least one object. Specifically, since the object detection frame represents the specific position of the object in the central video frame, the spatio-temporal features corresponding to this position can be extracted from the video spatio-temporal features corresponding to the target video based on the specific position of the object as the object spatio-temporal features.

[0047] S105. Based on the human spatio-temporal features and the object spatio-temporal features, recognize the action of the target person.

[0048] In one embodiment, the action recognition device recognizes the human posture of the target person based on the human spatio-temporal features, and then, based on the human posture and the object features, assists in judging the action performed by the target person.

[0049] This embodiment can also recognize the actions of multiple persons in the target video simultaneously.

[0050] In the embodiments of the present application, by determining the central video frame in the target video, and then obtaining the human detection frame and the object detection frame in the central video frame, the human detection frame and the object detection frame can represent the approximate positions of the person and the object in the entire target video, so as to directly extract the relevant features for action recognition from the video spatio-temporal features corresponding to the target video, thereby eliminating the need to process each frame of the target video to obtain the effective features (i.e., human features and object features) in each video frame, and further reducing the complicated image processing process in the action recognition process and improving the action recognition efficiency.

[0051] It can be understood that, in order to improve the accuracy of action recognition, the human detection frame may include a human body detection frame and a torso detection frame, and then the relevant features of the target person are extracted based on the human body detection frame and the torso detection frame, so as to perform action recognition based on the features related to the target person. The following will be combined with Figures 3-4 to introduce in detail the action recognition method based on the human body detection frame and the torso detection frame.

[0052] S201. Obtain the central video frame in the target video, where the central video frame is the video frame located in the middle position in the video frame sequence corresponding to the target video, and the video frame sequence includes a plurality of the video frames arranged in the acquisition order.

[0053] For details, reference can be made to S101, which will not be elaborated here.

[0054] S202. Determine the human body detection frame corresponding to the target person in the central video frame.

[0055] In one embodiment, a detection box recognition model may be included in the action recognition device, and the recognition model may recognize the detection box of a person and the detection box of an object.

[0056] The action recognition device uses the detection box recognition model to obtain the human detection box of the target person in the central video frame based on the central video frame.

[0057] S203. Based on the human detection box corresponding to the target person, obtain the human key points corresponding to the target person.

[0058] In one embodiment, the human key points may be the top of the head key point, the neck key point, the left shoulder key point, the right shoulder key point, the left elbow key point, the right elbow key point, the left wrist key point, the right wrist key point, the abdomen key point, the left hip key point, the right hip key point, the left knee key point, the right knee key point, the left ankle key point, and the right ankle key point, etc. The specific number and types of key points depend on the specific situation, and may be more than the human key points exemplified above or less than the human key points exemplified above, and are not limited herein. A human pose estimation model may be included in the action recognition device, and the estimation model may recognize multiple human key points of the target person in the central video frame.

[0059] The action recognition device uses the human pose estimation model to obtain multiple human key points corresponding to the target person in the area corresponding to the human detection box in the central video frame based on the central video frame and the human detection box corresponding to the target person. Exemplarily, as Figure 4 shown, Figure 4 shows multiple human key points corresponding to the target person in the central video frame, where the central video frame is 200, the human detection box is 210, and the human key points are 220.

[0060] S204. Based on the human detection box and the human key points corresponding to the target person, obtain the torso detection box corresponding to the target person.

[0061] In one embodiment, the action recognition device obtains multiple torso detection boxes corresponding to the target person in the central video frame based on the human detection box and multiple human key points corresponding to the target person.

[0062] Exemplarily, as Figure 5 shown, Figure 5 shows multiple torso detection boxes corresponding to the target person in the central video frame, where the central video frame is 300, the human detection box is 310, the human key points are 320, and the torso detection box is 330. It should be noted that for ease of description, Figure 5Only two torso detection frames 330 and the human key points 320 corresponding to the torso detection frames 330 are shown. Specifically, the action recognition device can obtain a head detection frame based on the head key point and the neck key point, an upper left arm detection frame based on the left shoulder key point and the left elbow key point, a lower left arm detection frame based on the left elbow key point and the left wrist key point, an upper right arm detection frame based on the right shoulder key point and the right elbow key point, a lower right arm detection frame based on the right elbow key point and the right wrist key point, and so on. By analogy, no further examples will be given here. It should be noted that the head detection frame, the upper left arm detection frame, the lower left arm detection frame, the upper right arm detection frame, and the lower right arm detection frame in the example are collectively referred to as the torso detection frame.

[0063] S205. Determine the object detection frame corresponding to at least one object in the central video frame.

[0064] In one embodiment, the action recognition device uses a detection frame recognition model to obtain the object detection frame of at least one object of the target person in the central video frame based on the central video frame.

[0065] In the embodiments of the present application, by obtaining multiple torso detection frames, when obtaining the spatio-temporal features of a person, the global spatio-temporal features and local spatio-temporal features of the person can be obtained, so that the possible actions of the target person can be further screened out through the local spatio-temporal features. Exemplarily, for example, playing basketball and playing football are both cases where the human body is close to the ball, but the torso positions in contact are different, so the movements are different, thereby improving the action recognition accuracy.

[0066] Optionally, since there may be unrecognizable objects in the central video frame, in order to improve the action recognition accuracy, determining the object detection frame corresponding to at least one object in the central video frame may include the following steps:

[0067] If there is at least one object in the central video frame, determine the object detection frame corresponding to each object in the at least one object;

[0068] If there is no object in the central video frame, determine the object detection frame corresponding to the background area of the central video frame.

[0069] In one embodiment, when the object recognition device obtains the object detection boxes of the objects in the central video frame through the detection box recognition model, the detection box recognition model first determines at least one object existing in the central video frame and the corresponding regional position of the object, and then generates the object detection boxes corresponding to the respective objects based on the regional position of the object. If the detection and recognition model fails to recognize an object, such as water, in the central video frame, the background area in the central video frame is determined as the corresponding object detection box in the central video frame. Optionally, the entire background area in the central video frame can be used as an object detection box, or the background area in the central video frame can be equally divided into four regions to obtain four object detection boxes.

[0070] Exemplarily, when there is no object in the central video frame, the division of the object detection box can be as Figure 6 shown, where the central video frame is 400, the human detection box is 410, and the object detection box is 420.

[0071] In the embodiment of the present application, the object detection box is obtained based on the background area, thereby avoiding the situation where the object detection box is discarded due to the inability to effectively recognize the object when the object is too large (the object occupies the entire area of the entire central video frame), resulting in the inability to extract the spatio-temporal features of the object. Furthermore, when performing action recognition based on the spatio-temporal features of the person and the spatio-temporal features of the object, the possible actions of the target person are further screened through the spatio-temporal features of the object, improving the accuracy of action recognition.

[0072] S206. Use a three-dimensional convolutional neural network to extract features from the target video to obtain the video spatio-temporal features corresponding to the target video.

[0073] In one embodiment, the three-dimensional convolutional neural network is a video feature extraction model, and through this model, the video spatio-temporal features in the target video can be effectively extracted. It should be noted that the three-dimensional convolutional neural network is different from the common video feature extraction methods and does not require frame-by-frame feature extraction.

[0074] When the action recognition device obtains the video frame sequence corresponding to the target video, it uses a three-dimensional convolutional neural network to perform feature extraction based on the video frame sequence corresponding to the target video, thereby obtaining the video spatio-temporal features corresponding to the target video.

[0075] In the embodiment of the present application, by using a three-dimensional convolutional neural network to obtain the video spatio-temporal features corresponding to the target video, the action recognition device can perform feature extraction in the video spatio-temporal features based on the person detection box and the object detection box, so that it is not necessary to perform frame-by-frame processing on the target video to obtain the effective features (i.e., person features and item features) in each video frame, thereby reducing the complicated image processing process in the action recognition process and improving the action recognition efficiency.

[0076] S207. Extract the human body spatio-temporal feature in the spatio-temporal feature of the target person corresponding to the human body detection box from the spatio-temporal feature of the video.

[0077] In one embodiment, the action recognition device extracts the spatio-temporal feature corresponding to the human body detection box from the spatio-temporal feature of the video based on the regional position of the human body detection box corresponding to the target person, and then uses the spatio-temporal feature corresponding to the human body detection box as the human body spatio-temporal feature corresponding to the target person.

[0078] S208. Extract the torso spatio-temporal feature in the spatio-temporal feature of the target person corresponding to the torso detection box from the spatio-temporal feature of the video.

[0079] In one embodiment, the action recognition device extracts the spatio-temporal features corresponding to multiple torso detection boxes from the spatio-temporal feature of the video based on the regional positions of the multiple torso detection boxes corresponding to the target person, and then uses the spatio-temporal features corresponding to the multiple torso detection boxes as the multiple torso spatio-temporal features corresponding to the target person.

[0080] In the embodiment of the present application, through the human body detection box and the torso detection box, relevant features for action recognition are directly extracted from the spatio-temporal feature of the target video, so that it is not necessary to process each frame of the target video to obtain the effective features (i.e., human features and item features) in each video frame, thereby reducing the complicated image processing process in the action recognition process and improving the action recognition efficiency.

[0081] S209. Obtain the object spatio-temporal feature corresponding to the at least one object in the spatio-temporal feature of the target video corresponding to the object detection box.

[0082] In one embodiment, the action recognition device obtains the object spatio-temporal feature corresponding to the at least one object from the spatio-temporal feature of the target video corresponding to the object detection box of the at least one object. Specifically, since the object detection box represents the specific position of the object in the central video frame, the spatio-temporal feature corresponding to this position can be extracted from the spatio-temporal feature of the target video as the object spatio-temporal feature based on the specific position of the object.

[0083] S210. Obtain the interaction relationship between the target person and each object based on the human body spatio-temporal feature and the object spatio-temporal feature corresponding to each object in the at least one object.

[0084] In one embodiment, the interaction relationship refers to the interaction relationship between the human body and an object. Exemplarily, such as a human body picking up an object, a human body patting an object, the human body having no contact with the object, etc. It should be noted that the foregoing examples are only for auxiliary understanding and do not represent that the interaction relationship between the human body and the object is only the foregoing several kinds.

[0085] The action recognition device performs relationship reasoning based on the spatio-temporal characteristics of the person and the spatio-temporal characteristics of each object in at least one object, so as to obtain the interaction relationship between the target person and each object.

[0086] Optionally, since the spatio-temporal characteristics of the person include the spatio-temporal characteristics of the human body and the spatio-temporal characteristics of the torso, based on the spatio-temporal characteristics of the person and the spatio-temporal characteristics of each object in at least one object, obtaining the interaction relationship between the target person and each object may include the following steps:

[0087] Based on the spatio-temporal characteristics of the human body in the spatio-temporal characteristics of the person and the spatio-temporal characteristics of each object in at least one object, obtain the global interaction relationship between the target person and each object;

[0088] Based on the spatio-temporal characteristics of the human body and the spatio-temporal characteristics of the torso in the spatio-temporal characteristics of the person and the spatio-temporal characteristics of each object, obtain the local interaction relationship between the target person and each object.

[0089] In one embodiment, the action recognition device includes a human body relationship reasoning model and a posture relationship reasoning model. Among them, the human body relationship reasoning model is used to obtain the global interaction relationship between the human body and the object, and the posture relationship reasoning model is used to obtain the local interaction relationship between multiple torsos in the human body and the object.

[0090] The action recognition device uses the human body relationship reasoning model to obtain the global interaction relationship between the target person and the object based on the spatio-temporal characteristics of the human body and the spatio-temporal characteristics of each object; at the same time, it also uses the posture relationship reasoning model to obtain the local interaction relationship between the target person and each object based on the spatio-temporal characteristics of the human body, the spatio-temporal characteristics of the torso, and the spatio-temporal characteristics of each object.

[0091] Exemplarily, if the target person in the target video is playing basketball, the global interaction relationship may be that the human body pats the object, and the multiple local interaction relationships are respectively: the head has no contact with the object, the left upper arm has no contact with the object, the left lower arm has no contact, the right upper arm has no contact with the object, the right lower arm pats the object, etc. It should be noted that the foregoing examples are only for auxiliary understanding and do not represent that the interaction relationship between the human body and the object is only the foregoing several kinds.

[0092] Optionally, the pose relationship inference model can also obtain the interaction relationship between the torso of the target person and other persons in the target video. In this case, the action recognition device inputs all the person spatio-temporal features and object spatio-temporal features in the target video into the pose relationship inference model, and then obtains the global interaction relationships between each target person and each object among multiple persons one by one, the local interaction relationships between each torso and each object among the target persons, and the local interaction relationships between each torso and other persons among the target persons.

[0093] In the embodiment of the present application, the action recognition device obtains the global interaction relationship and multiple local interaction relationships between the target person and each object, and thus filters out the possible actions of the target person through the obtained global interaction relationship and multiple local interaction relationships. Exemplarily, for example, playing basketball and playing football are both cases where the human body is close to the ball, but the torso positions in contact are different, so the performed movements are different, thereby improving the action recognition accuracy.

[0094] S211. Based on the interaction relationship between the target person and each object, obtain multiple action probabilities corresponding to the target person, where the action probability is the similarity between the action performed by the target person and each action type among multiple action types.

[0095] In one embodiment, the action recognition device may include an action probability estimation model, and this action probability estimation model can analyze the probability that the action performed by the target person is an action type based on each interaction relationship. It should be noted that the action probability estimation model is a trained action probability estimation model obtained by training based on a large number of interaction relationships and their corresponding action types. Exemplarily, the action type can be turning around, waving, playing basketball, playing football, swimming, etc., which is not limited herein.

[0096] The action recognition device uses the action probability estimation model to obtain multiple action probabilities corresponding to the target person based on the interaction relationship between the target person and each object.

[0097] Optionally, since the interaction relationship includes a global interaction relationship and a local interaction relationship, obtaining multiple action probabilities corresponding to the target person based on the interaction relationship between the target person and each object may include the following steps:

[0098] Use the action probability estimation model to obtain multiple action probabilities corresponding to the target person based on the person spatio-temporal features, the global interaction relationship and the local interaction relationship between the target person and each object.

[0099] In one embodiment, the action recognition device uses an action probability estimation model to obtain multiple action probabilities corresponding to the target person based on the spatio-temporal characteristics of the person, the global interaction relationship between the target person and each object, and multiple local interaction relationships.

[0100] In the embodiments of the present application, the possible actions of the target person are further screened through local interaction relationships, avoiding recognition errors caused by separately identifying the person's characteristics and global interaction information. For example, when a person is playing football, there is a ball near the person's foot, but the hand is waving up and down. Without the local interaction relationship with the object, it is easy to recognize it as a basketball-playing action, thereby improving the action recognition accuracy.

[0101] S212. Use the action types with action probabilities greater than the probability threshold among the multiple action probabilities as the actions performed by the target person.

[0102] In one embodiment, the action recognition device obtains multiple action types output by the action probability estimation model and their corresponding action probabilities, then compares the action probabilities with the probability threshold one by one, determines the target action probabilities greater than the probability threshold among the multiple action probabilities, and then uses the action types corresponding to the target action probabilities as the actions performed by the target person. It should be noted that if there is no action probability greater than the probability threshold, it is determined that the target person has not performed an action, or a prompt message indicating that the action is not recognized is output. This prompt message can be used to prompt the user to manually identify the action of the target person, or to prompt the user to improve the action probability estimation model.

[0103] In the embodiments of the present application, the possible actions of the target person are screened through each interaction relationship, avoiding recognition errors caused by separately identifying the person's characteristics. For example, if a person's elbow is waving up and down, without the interaction relationship with the object, it is easy to recognize it as a waving action, thereby improving the action recognition accuracy.

[0104] The following will Figure 7 - attached Figure 8 The action recognition device provided in the embodiments of the present application will be introduced in detail. It should be noted that the attached Figure 7 - attached Figure 8 The action recognition device is used to execute the method of the embodiments of the present application Figures 1-6 shown in the embodiments. For the sake of convenience of description, only the parts related to the embodiments of the present application are shown. For the specific technical details not disclosed, please refer to the embodiments Figures 1-6 shown in the present application.

[0105] Please refer to Figure 7 , which is a schematic structural diagram of an action recognition device provided in the embodiments of the present application. As Figure 7As shown in the figure, the action recognition device 1 according to the embodiment of the present application may include: a video frame acquisition module 11, a detection frame determination module 12, a feature acquisition module 13, and an action recognition module 14.

[0106] The video frame acquisition module 11 is configured to acquire a central video frame in a target video, where the central video frame is a video frame located in the middle position in a video frame sequence corresponding to the target video, and the video frame sequence includes a plurality of the video frames arranged in the acquisition order.

[0107] The detection frame determination module 12 is configured to determine a person detection frame corresponding to a target person in the central video frame and an object detection frame corresponding to at least one object in the central video frame, where the target person is any one of multiple persons in the central video frame.

[0108] The feature acquisition module 13 is configured to acquire a person spatio-temporal feature corresponding to the target person in the spatio-temporal feature of the video corresponding to the target video based on the person detection frame.

[0109] The feature acquisition module 13 is further configured to acquire an object spatio-temporal feature corresponding to the at least one object in the spatio-temporal feature of the video corresponding to the target video based on the object detection frame.

[0110] The action recognition module 14 is configured to recognize the action of the target person based on the person spatio-temporal feature and the object spatio-temporal feature.

[0111] In the embodiment of the present application, by determining the central video frame in the target video, and then acquiring the person detection frame and the object detection frame in the central video frame, the person detection frame and the object detection frame can represent the approximate positions of the person and the object in the entire target video, so as to directly extract the relevant features for action recognition from the spatio-temporal feature of the video corresponding to the target video based on the person detection frame and the object detection frame, thereby eliminating the need to process each frame of the target video to obtain the effective features (i.e., person features and object features) in each video frame, and further reducing the complicated image processing process in the action recognition process and improving the action recognition efficiency.

[0112] Optionally, the detection frame determination module 12 is specifically configured to:

[0113] Determine a human body detection frame corresponding to the target person in the central video frame;

[0114] Acquire human body key points corresponding to the target person based on the human body detection frame corresponding to the target person;

[0115] Acquire a torso detection frame corresponding to the target person based on the human body detection frame and the human body key points corresponding to the target person;

[0116] Determine the object detection frames corresponding to at least one object in the central video frame.

[0117] In the embodiments of the present application, by obtaining multiple torso detection frames, when obtaining the spatio-temporal features of a person, the global spatio-temporal features and local spatio-temporal features of the person can be obtained, so that the possible actions of the target person can be further screened out through the local spatio-temporal features. Exemplarily, for example, playing basketball and playing football are both cases where the human body is close to the ball, but the torso positions in contact are different, so the movements are different, thereby improving the accuracy of action recognition.

[0118] Optionally, the detection frame determination module 12 is specifically configured to:

[0119] If there is at least one object in the central video frame, determine the object detection frames corresponding to each object in the at least one object;

[0120] If there is no object in the central video frame, determine the object detection frame corresponding to the background area of the central video frame.

[0121] In one embodiment, when the object recognition device obtains the object detection frames of the objects in the central video frame through the detection frame recognition model, the detection frame recognition model first determines at least one object existing in the central video frame and the regional position corresponding to the object, and then generates the object detection frames corresponding to each object based on the regional position of the object. If the detection and recognition model does not recognize an object in the central video frame, such as water, the background area in the central video frame is determined as the object detection frame corresponding to the central video frame. Optionally, the entire background area in the central video frame can be used as an object detection frame, or the background area in the central video frame can be equally divided into four regions to obtain four object detection frames.

[0122] Optionally, the feature acquisition module 13 is specifically configured to:

[0123] Extract the human body spatio-temporal features in the spatio-temporal features of the target person corresponding to the target person based on the human body detection frame;

[0124] Extract the torso spatio-temporal features in the spatio-temporal features of the target person corresponding to the target person based on the torso detection frame.

[0125] In the embodiments of the present application, through the human body detection frame and the torso detection frame, the relevant features for action recognition are directly extracted from the spatio-temporal features corresponding to the target video, so that it is not necessary to process each frame of the target video to obtain the effective features (i.e., human features and item features) in each video frame, thereby reducing the complicated image processing process in the action recognition process and improving the action recognition efficiency.

[0126] Optionally, the action recognition module 14 is specifically configured to:

[0127] Based on the person spatio-temporal feature and the object spatio-temporal features corresponding to each object in the at least one object, obtain the interaction relationship between the target person and each object;

[0128] Based on the interaction relationship between the target person and each object, obtain multiple action probabilities corresponding to the target person, where the action probability is the similarity between the action executed by the target person and each action type among multiple action types;

[0129] Take the action type with an action probability greater than the probability threshold among the multiple action probabilities as the action executed by the target person.

[0130] In the embodiment of the present application, possible actions of the target person are screened out through each interaction relationship, avoiding recognition errors caused by separately recognizing person features. For example, if a person's elbows are waving up and down, without the interaction relationship with an object, it is easy to be recognized as a waving action, thereby improving the action recognition accuracy.

[0131] Optionally, the action recognition module 14 is specifically configured to: based on the human body spatio-temporal feature in the person spatio-temporal feature and the object spatio-temporal features corresponding to each object in the at least one object, obtain the global interaction relationship between the target person and each object;

[0132] Based on the human body spatio-temporal feature and the torso spatio-temporal feature in the person spatio-temporal feature and the object spatio-temporal features corresponding to each object, obtain the local interaction relationship between the target person and each object.

[0133] In the embodiment of the present application, the action recognition device obtains the global interaction relationship and multiple local interaction relationships between the target person and each object, and thus screens out possible actions of the target person through the obtained global interaction relationship and multiple local interaction relationships. Exemplarily, for example, both playing basketball and playing football are similar in that the human body is close to the ball, but the torso positions in contact are different, so the movements are different, thereby improving the action recognition accuracy.

[0134] Optionally, the action recognition module 14 is specifically configured to: adopt an action probability estimation model, and based on the person spatio-temporal feature, the global interaction relationship and the local interaction relationship between the target person and each object, obtain multiple action probabilities corresponding to the target person.

[0135] In the embodiments of the present application, the possible actions of the target person are further screened through local interaction relationships, avoiding recognition errors caused by separately identifying the person's features and global interaction information. For example, when a person is playing football, there is a ball near the person's feet, but the hands are waving up and down. Without the local interaction relationship with the object, it is easy to recognize it as a basketball-playing action, thereby improving the action recognition accuracy.

[0136] Optionally, please refer to Figure 8 , the action recognition device 1 further includes: a feature extraction module 15.

[0137] Using a three-dimensional convolutional neural network, the target video is subjected to feature extraction to obtain the video spatio-temporal features corresponding to the target video.

[0138] In the embodiments of the present application, by using a three-dimensional convolutional neural network to obtain the video spatio-temporal features corresponding to the target video, the action recognition device can perform feature extraction in the video spatio-temporal features based on the person detection frame and the object detection frame, so that it is not necessary to process each frame of the target video to obtain the effective features (i.e., person features and item features) in each video frame, thereby reducing the complicated image processing process in the action recognition process and improving the action recognition efficiency.

[0139] The embodiments of the present application further provide a storage medium, which can store multiple program instructions, and the program instructions are suitable for being loaded and executed by a processor to perform the method steps of the embodiments as described above Figures 1-6 The specific execution process can refer to Figures 1-6 the specific description of the embodiments shown, and will not be elaborated here.

[0140] Please refer to Figure 9 , which is a schematic structural diagram of a computer device provided by the embodiments of the present application. As Figure 9As shown in the figure, the computer device 1000 may include: at least one processor 1001, at least one communication bus 1002, at least one input / output interface 1003, at least one network interface 1004, and at least one memory 1005. Among them, the processor 1001 may include one or more processing cores. The processor 1001 connects various parts within the entire computer device 1000 through various interfaces and lines, and executes various functions of the terminal 1000 and processes data by running or executing instructions, programs, code sets, or instruction sets stored in the memory 1005, and by calling data stored in the memory 1005. The memory 1005 may be a high-speed RAM memory or a non-volatile memory, such as at least one disk memory. Optionally, the memory 1005 may also be at least one storage device located far from the aforementioned processor 1001. Among them, the network interface 1004 may optionally include a standard wired interface, a wireless interface (such as a WI-FI interface). The communication bus 1002 is used to realize the connection and communication between these components. As Figure 9 shown, the memory 1005, as a storage medium of a terminal device, may include an operating system, a network communication module, an input / output interface module, and an action recognition program.

[0141] In Figure 9 the computer device 1000 shown in the figure, the input / output interface 1003 is mainly used to provide an input interface for users and access devices, and to obtain data input by users and access devices.

[0142] In one embodiment.

[0143] The processor 1001 may be used to call the action recognition program stored in the memory 1005 and specifically perform the following operations:

[0144] Obtain the central video frame in the target video, where the central video frame is the video frame located in the middle position in the video frame sequence corresponding to the target video, and the video frame sequence includes a plurality of the video frames arranged in the acquisition order;

[0145] Determine the person detection box corresponding to the target person in the central video frame and the object detection box corresponding to at least one object in the central video frame, where the target person is any one of the multiple persons in the central video frame;

[0146] Based on the person detection box, obtain the person spatio-temporal feature corresponding to the target person in the spatio-temporal feature of the video corresponding to the target video;

[0147] Based on the object detection box, obtain the object spatio-temporal features corresponding to the at least one object in the video spatio-temporal features corresponding to the target video;

[0148] Based on the human spatio-temporal features and the object spatio-temporal features, identify the actions of the target person.

[0149] Optionally, when the processor 1001 executes to determine the person detection box corresponding to the target person in the central video frame and the object detection box corresponding to at least one object in the central video frame, the following operations are specifically performed:

[0150] Determine the human detection box corresponding to the target person in the central video frame;

[0151] Based on the human detection box corresponding to the target person, obtain the human key points corresponding to the target person;

[0152] Based on the human detection box and the human key points corresponding to the target person, obtain the torso detection box corresponding to the target person;

[0153] Determine the object detection box corresponding to at least one object in the central video frame.

[0154] Optionally, when the processor 1001 executes to determine the object detection box corresponding to at least one object in the central video frame, the following operations are specifically performed:

[0155] If there is at least one object in the central video frame, determine the object detection box corresponding to each object in the at least one object;

[0156] If there is no object in the central video frame, determine the object detection box corresponding to the background area of the central video frame.

[0157] Optionally, when the processor 1001 executes to obtain the human spatio-temporal features corresponding to the target person in the video spatio-temporal features corresponding to the target video based on the person detection box, the following operations are specifically performed:

[0158] Based on the human detection box, extract the human spatio-temporal features in the human spatio-temporal features corresponding to the target person from the video spatio-temporal features;

[0159] Based on the torso detection box, extract the torso spatio-temporal features in the human spatio-temporal features corresponding to the target person from the video spatio-temporal features.

[0160] Optionally, when the processor 1001 executes to identify the actions of the target person based on the human spatio-temporal features and the object spatio-temporal features, the following operations are specifically performed:

[0161] Obtain the interaction relationship between the target person and each object based on the spatio-temporal characteristics of the person and the spatio-temporal characteristics of each object corresponding to each object among the at least one object;

[0162] Based on the interaction relationship between the target person and each object, obtain multiple action probabilities corresponding to the target person, where the action probability is the similarity between the action performed by the target person and each action type among multiple action types;

[0163] Take the action type with an action probability greater than the probability threshold among the multiple action probabilities as the action performed by the target person.

[0164] Optionally, when the processor 1001 executes obtaining the interaction relationship between the target person and each object based on the spatio-temporal characteristics of the person and the spatio-temporal characteristics of each object corresponding to each object among the at least one object, the following operations are specifically performed:

[0165] Obtain the global interaction relationship between the target person and each object based on the human spatio-temporal characteristics in the spatio-temporal characteristics of the person and the spatio-temporal characteristics of each object corresponding to each object among the at least one object;

[0166] Obtain the local interaction relationship between the target person and each object based on the human spatio-temporal characteristics and torso spatio-temporal characteristics in the spatio-temporal characteristics of the person and the spatio-temporal characteristics of each object corresponding to each object.

[0167] Optionally, when the processor 1001 executes obtaining multiple action probabilities corresponding to the target person based on the interaction relationship between the target person and each object, the following operations are specifically performed:

[0168] Adopt an action probability estimation model to obtain multiple action probabilities corresponding to the target person based on the spatio-temporal characteristics of the person, the global interaction relationship and the local interaction relationship between the target person and each object.

[0169] Optionally, before the processor 1001 executes obtaining the spatio-temporal characteristics of the target person in the video spatio-temporal characteristics corresponding to the target video based on the person detection frame, the following operations are also performed:

[0170] Adopt a three-dimensional convolutional neural network to perform feature extraction on the target video to obtain the video spatio-temporal characteristics corresponding to the target video.

[0171] In the embodiments of the present application, by determining the central video frame in the target video, and then obtaining the person detection frame and the object detection frame in the central video frame, the person detection frame and the object detection frame can represent the approximate positions of the person and the object in the entire target video, so as to directly extract the relevant features for action recognition from the video spatio-temporal features corresponding to the target video based on the person detection frame and the object detection frame, thereby eliminating the need to obtain the effective features (i.e., person features and object features) in each video frame by processing the target video frame by frame, and further reducing the complicated image processing process in the action recognition process and improving the action recognition efficiency.

[0172] It should be noted that for the foregoing method embodiments, for the sake of simplicity of description, they are all expressed as a series of action combinations. However, those skilled in the art should know that the present application is not limited by the described action sequence, because according to the present application, certain steps can be performed in other sequences or simultaneously. Secondly, those skilled in the art should also know that the embodiments described in the specification are all preferred embodiments, and the actions and modules involved are not necessarily essential to the present application.

[0173] In the above embodiments, the descriptions of the respective embodiments have their own emphases. For the parts not detailed in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.

[0174] The above is the description of an action recognition method, device, storage medium and equipment provided by the present application. For those skilled in the art, according to the idea of the embodiments of the present application, there will be changes in the specific implementation manners and application scopes. In summary, the content of this specification should not be construed as a limitation to the present application.

Claims

1. A method for action recognition, characterized in that, The method includes: Obtaining a central video frame in the target video, where the central video frame is the video frame located in the middle position in the video frame sequence corresponding to the target video, and the video frame sequence includes a plurality of the video frames arranged in the acquisition order; Determining a person detection box corresponding to the target person in the central video frame and an object detection box corresponding to at least one object in the central video frame, where the target person is any one of multiple persons in the central video frame; Based on the person detection box, obtaining the person spatio-temporal feature corresponding to the target person in the video spatio-temporal feature corresponding to the target video; Based on the object detection box, obtaining the object spatio-temporal feature corresponding to the at least one object in the video spatio-temporal feature corresponding to the target video; Based on the human body spatio-temporal feature in the person spatio-temporal feature and the object spatio-temporal features corresponding to each object in the at least one object, obtaining the global interaction relationship between the target person and each object; Based on the human body spatio-temporal feature and torso spatio-temporal feature in the person spatio-temporal feature and the object spatio-temporal features corresponding to each object, obtaining the local interaction relationship between the target person and each object; Based on the global interaction relationship and the local interaction relationship, obtaining a plurality of action probabilities corresponding to the target person, where the action probability is the similarity between the action executed by the target person and each action type among multiple action types; Taking the action type with an action probability greater than the probability threshold among the plurality of action probabilities as the action executed by the target person.

2. The method according to claim 1, wherein The person detection box includes a human body detection box and a plurality of torso detection boxes; The determining the person detection box corresponding to the target person in the central video frame and the object detection box corresponding to at least one object in the central video frame includes: Determining the human body detection box corresponding to the target person in the central video frame; Based on the human body detection box corresponding to the target person, obtaining the human body key points corresponding to the target person; Based on the human body detection box and human body key points corresponding to the target person, obtaining the torso detection boxes corresponding to the target person; Determining the object detection box corresponding to at least one object in the central video frame.

3. The method according to claim 2, wherein The determining the object detection box corresponding to at least one object in the central video frame includes: If there is at least one object in the central video frame, determining the object detection boxes corresponding to each object in the at least one object; If there is no object in the central video frame, determining the object detection box corresponding to the background area of the central video frame.

4. The method according to claim 2, wherein The obtaining the person spatio-temporal feature corresponding to the target person in the video spatio-temporal feature corresponding to the target video based on the person detection box includes: Based on the human body detection box, extracting the human body spatio-temporal feature in the person spatio-temporal feature corresponding to the target person in the video spatio-temporal feature; Based on the torso detection box, extracting the torso spatio-temporal feature in the person spatio-temporal feature corresponding to the target person in the video spatio-temporal feature.

5. The method according to claim 1, characterized in that, Obtaining multiple action probabilities corresponding to the target person based on the global interaction relationship and the local interaction relationship includes: Using an action probability estimation model, based on the person spatio-temporal features, the global interaction relationship and the local interaction relationship between the target person and each object, to obtain multiple action probabilities corresponding to the target person.

6. The method according to claim 1, wherein Before obtaining the person spatio-temporal features corresponding to the target person in the video spatio-temporal features corresponding to the target video based on the person detection frame, it further includes: Using a three-dimensional convolutional neural network to extract features from the target video to obtain the video spatio-temporal features corresponding to the target video.

7. An action recognition device, characterized in that, It includes: A video frame acquisition module, configured to acquire a central video frame in the target video, where the central video frame is the video frame located in the middle position in the video frame sequence corresponding to the target video, and the video frame sequence includes multiple video frames arranged in the acquisition order; A detection frame determination module, configured to determine a person detection frame corresponding to the target person in the central video frame and an object detection frame corresponding to at least one object in the central video frame, where the target person is any one of multiple persons in the central video frame; A feature acquisition module, configured to obtain the person spatio-temporal features corresponding to the target person in the video spatio-temporal features corresponding to the target video based on the person detection frame; The feature acquisition module is further configured to obtain the object spatio-temporal features corresponding to the at least one object in the video spatio-temporal features corresponding to the target video based on the object detection frame; An action recognition module, configured to obtain the global interaction relationship between the target person and each object based on the human body spatio-temporal features in the person spatio-temporal features and the object spatio-temporal features corresponding to each object in the at least one object; The action recognition module is further configured to obtain the local interaction relationship between the target person and each object based on the human body spatio-temporal features and the torso spatio-temporal features in the person spatio-temporal features and the object spatio-temporal features corresponding to each object; The action recognition module is further configured to obtain multiple action probabilities corresponding to the target person based on the global interaction relationship and the local interaction relationship, where the action probability is the similarity between the action performed by the target person and each action type among multiple action types; The action recognition module is further configured to use the action type with an action probability greater than the probability threshold among the multiple action probabilities as the action performed by the target person.

8. A storage medium, on which a computer program is stored, characterized in that, When the computer program is executed by a processor, it implements the action recognition method according to any one of claims 1-6.

9. A computer device, characterized in that, It includes: A processor and a memory; wherein, the memory stores a computer program, and the computer program is adapted to be loaded and executed by the processor to perform the steps of the action recognition method according to any one of claims 1-6.

Citation Information

Patent Citations

  • Interactive relationship identification method and device, equipment and storage medium

    CN111325141A