Training Method of Interaction Detection Model, Interaction Detection Method and Related Devices
Through the training method of the interaction detection model, using feature extraction and action classification networks, combined with three-dimensional position and interaction score prediction, the problem of insufficient detection accuracy of character interaction relationships in the prior art is solved, and higher detection accuracy and less false detection are achieved.
Patent Information
- Application Number
- CN202210596450.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-05-18
- Publication Date
- 2025-07-01
- Estimated Expiration
- 2042-05-18
AI Technical Summary
The existing spatial and temporal character relationship detection methods mainly focus on the human movements, limiting the improvement of the accuracy of character interaction relationship detection.
Through a training method of interaction detection model, the feature extraction network and action classification network are used to process the sample video data, extract the sample human characteristics and interactive actions, and adjust the model's network parameters through three-dimensional positioning and interaction score prediction to improve the interaction detection accuracy.
The interaction detection model is optimized from the positioning level and classification level, so that the model can pay attention to the position information of human body movements and interactive objects at the same time, improve the accuracy of character interaction relationship detection, and reduce false detection under the long-tail relationship distribution.
Smart Images

Figure CN114898272B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of computer vision technology, and in particular, to a training method for an interaction detection model, an interaction detection method, and related devices. Background Art
[0002] Spatio-temporal human relationship detection aims to detect the human interaction relationships occurring in a video, and spatio-temporal human relationship detection is particularly important for video behavior understanding. In the daily human interaction process, a person may interact with various objects existing in the surrounding environment. For example, when doing housework, a person may pick up or touch dozens of different furniture items.
[0003] Currently, the methods for spatio-temporal human relationship detection usually only focus on the actions of the human body itself, which limits the improvement of the accuracy of human interaction relationship detection. Summary of the Invention
[0004] The present application provides at least a training method for an interaction detection model, an interaction detection method, and related devices.
[0005] In a first aspect of the present application, a training method for an interaction detection model is provided. The method includes: processing a sample image in sample video data based on a feature extraction network of the interaction detection model to obtain a sample human body feature of a sample human body in the sample image; wherein, the sample video data is labeled with a sample score indicating whether a sample object interacts with the sample human body, and a sample interaction action of the sample human body that interacts with the sample object; classifying the sample human body feature based on an action classification network of the interaction detection model to obtain a first predicted interaction action of the sample human body; positioning the three-dimensional position of the sample object based on the two-dimensional position of the sample object and the morphological parameters of the sample human body; wherein, the two-dimensional position is obtained based on the sample human body feature; predicting a predicted score of the sample object based on the morphological parameters and the three-dimensional position; wherein, the predicted score represents the possibility that the sample object has a human interaction; adjusting the network parameters of the interaction detection model based on the difference between the sample interaction action and the first predicted interaction action, and the difference between the sample score and the predicted score.
[0006] Therefore, based on the differences between the sample interaction actions and the first predicted interaction actions, as well as the differences between the sample scores and the predicted scores, the network parameters of the interaction detection model are adjusted. Thus, on the one hand, the predicted score representing the possibility of human interaction with the sample object approaches the sample score indicating whether the sample object interacts with the sample human body. Since the predicted score is obtained based on the three-dimensional position of the sample object, the three-dimensional position of the located sample object approaches the true three-dimensional position of the sample object, that is, it drives the positioning of the sample object to be as accurate as possible. And the positioning of the sample object is achieved based on the sample human body features, so that the interaction detection model can be forced to extract human body features closely related to the interaction actions with the human body as accurately as possible from the positioning level. On the other hand, the first predicted interaction action approaches the sample interaction action, and the first predicted interaction action is obtained based on the human body features, so that the interaction detection model can be forced to extract human body features closely related to the interaction actions with the human body as accurately as possible from the classification level, that is, the human body features extracted by the interaction detection model can be forced to be as accurate as possible from the classification level. Therefore, the interaction detection model is optimized from two dimensions: the positioning level and the classification level, so that when the subsequent interaction detection model classifies and detects the interaction actions of the human body, it can simultaneously focus on the actions of the human body itself and the position information of the interaction objects that have human interaction with the human body, so that the interaction detection model can extract human body features closely related to the human interaction actions, that is, the interaction detection model can accurately extract human body features, thereby improving the detection accuracy of the interaction detection model for human interaction relationships and reducing false detections under the long-tail relationship distribution.
[0007] Among them, the two-dimensional position located based on the sample human body features includes: jointly locating the two-dimensional position of the sample object based on the sample human body features and the sample object features of the sample object.
[0008] Therefore, the two-dimensional position of the sample object can be located through the sample human body features and the sample object features of the sample object, so that the sample object features can assist the sample human body features in the two-dimensional positioning of the sample object, which is beneficial to improving the accuracy of the two-dimensional positioning.
[0009] Among them, jointly locating the two-dimensional position of the sample object based on the sample human body features and the sample object features of the sample object includes: extracting the sample object features of the sample object in the object region based on the sample image features of the sample image and the object region detected in the sample image; predicting based on the sample object features and the sample human body features to obtain an interaction score; where the interaction score represents the closeness of the interaction between the sample human body to which the sample human body features belong and the sample object in the object region; obtaining the two-dimensional position of the sample object based on the object region corresponding to the interaction score that meets the preset conditions.
[0010] Therefore, an interaction score is predicted based on the sample object features and the sample human body features, and the interaction score represents the degree of closeness between the sample human body to which the sample human body features belong and the sample object in the object region. Thus, in response to the interaction score satisfying a preset condition, based on the object region, the two-dimensional position of the sample object is obtained. Therefore, the two-dimensional position of the sample object closely related to the sample human body can be determined as accurately as possible by predicting the interaction score.
[0011] Among them, based on the sample image features of the sample image and the object region detected in the sample image, the sample object features of the sample object in the object region are extracted, including: performing object detection on the sample image to obtain a number of candidate regions; based on the first predicted interaction action, selecting a candidate region as the object region; based on the object region, extracting the sample object features from the sample image features.
[0012] Therefore, a number of candidate regions are screened through the first predicted interaction action, and the remaining candidate regions after screening are used as the object region, thereby reducing the computational amount.
[0013] Among them, based on the feature extraction network of the interaction detection model, the sample human body features of the sample human body in the sample image are obtained, including: extracting the sample image features of the sample image based on the feature extraction network, and performing human body detection on the sample image to obtain the human body region of the sample human body; based on the human body region, extracting the sample human body features from the sample image features.
[0014] Therefore, based on the human body region, the region corresponding to the human body region in the sample image features is determined, and the sample human body features are extracted from the region corresponding to the human body region in the sample image features, so as to accurately extract the sample human body features.
[0015] Among them, positioning is performed based on the two-dimensional position of the sample object and the morphological parameters of the sample human body to obtain the three-dimensional position of the sample object, including: predicting the initial position of the sample object in the three-dimensional space based on the two-dimensional position; based on the initial position and the morphological parameters, constructing an action classification loss with the correction parameter of the initial position as the optimization target; based on the action classification loss, optimizing to obtain the correction parameter, and correcting the position of the initial position based on the correction parameter to obtain the three-dimensional position.
[0016] Therefore, by introducing the action classification loss to optimize the correction parameter of the initial position, the accuracy of the three-dimensional position positioning of the sample object is improved.
[0017] Among them, before adjusting the network parameters of the interaction detection model based on the differences between the sample interaction actions and the first predicted interaction actions, as well as the differences between the sample scores and the predicted scores, the training method of the interaction detection model further includes: predicting based on the morphological parameters and the initial positions to obtain the second predicted interaction action of the sample human body; adjusting the network parameters of the interaction detection model based on the differences between the sample interaction actions and the first predicted interaction actions, as well as the differences between the sample scores and the predicted scores, including: adjusting the network parameters of the interaction detection model based on the differences between the sample interaction actions and the first predicted interaction actions, the differences between the sample scores and the predicted scores, and the differences between the first predicted interaction actions and the second predicted interaction actions.
[0018] Therefore, the interaction detection model is optimized from three dimensions: the positioning level, the classification level, and the action consistency level, so that when the subsequent interaction detection model classifies and detects the interaction actions of the human body, it can extract human body features closely related to the human body interaction actions, that is, the interaction detection model can accurately extract human body features, thereby improving the detection accuracy of the interaction detection model for the human interaction relationship and reducing the false detection under the long-tail relationship distribution.
[0019] Among them, the morphological parameters include human body pose parameters and human body shape parameters. Predicting based on the morphological parameters and the initial positions to obtain the second predicted interaction action of the sample human body includes: encoding based on the human body pose parameters to obtain a pose encoding representation; splicing the pose encoding representation, the human body shape parameters, and the initial positions to obtain a spliced feature representation; predicting based on the spliced feature representation to obtain the second predicted interaction action.
[0020] Therefore, the second predicted interaction action of the sample human body can be determined by the morphological parameters of the sample human body and the initial positions of the sample objects in the three-dimensional space.
[0021] Among them, the sample video data includes a plurality of frames of sample images; classifying the sample human body features based on the action classification network of the interaction detection model to obtain the first predicted interaction action of the sample human body includes: obtaining the sample action trajectory features of the sample human body based on the sample human body features of the same sample human body in each frame of the sample images; classifying based on the sample action trajectory features to obtain the first predicted interaction action of the sample human body.
[0022] Therefore, the interaction actions of the sample human body are classified by the sample action trajectory features to classify the interaction actions of the sample human body from the action level of the sample human body itself.
[0023] The second aspect of the present application provides an interaction detection method, which includes: extracting features of a to-be-detected image in to-be-detected video data based on the feature extraction network of the interaction detection model to obtain human features of a human body in the to-be-detected image; classifying the human features based on the action classification network of the interaction detection model to obtain the interaction action category of the human body; wherein, the interaction detection model is obtained based on the above-mentioned training method of the interaction detection model.
[0024] The third aspect of the present application provides a training device for an interaction detection model, which includes a sample feature extraction module, an interaction action prediction module, a three-dimensional position positioning module, an interaction score prediction module, and a network parameter adjustment module; the sample feature extraction module is used to process a sample image in sample video data based on the feature extraction network of the interaction detection model to obtain sample human features of a sample human body in the sample image; wherein, the sample video data is labeled with a sample score indicating whether a sample object interacts with the sample human body, and a sample interaction action of the sample human body interacting with the sample object; the interaction action prediction module is used to classify the sample human features based on the action classification network of the interaction detection model to obtain a first predicted interaction action of the sample human body; the three-dimensional position positioning module is used to perform positioning based on the two-dimensional position of the sample object and the morphological parameters of the sample human body to obtain the three-dimensional position of the sample object; wherein, the two-dimensional position is obtained based on the sample human features; the interaction score prediction module is used to perform prediction based on the morphological parameters and the three-dimensional position to obtain a predicted score of the sample object; wherein, the predicted score represents the possibility that the sample object has human interaction; the network parameter adjustment module is used to adjust the network parameters of the interaction detection model based on the difference between the sample interaction action and the first predicted interaction action, and the difference between the sample score and the predicted score.
[0025] The fourth aspect of the present application provides an interaction detection device, which includes a feature extraction module and an action classification module; the feature extraction module is used to extract features of a to-be-detected image in to-be-detected video data based on the feature extraction network of the interaction detection model to obtain human features of a human body in the to-be-detected image; the action classification module is used to classify the human features based on the action classification network of the interaction detection model to obtain the interaction action category of the human body; wherein, the interaction detection model is obtained based on the above-mentioned training device of the interaction detection model.
[0026] The fifth aspect of the present application provides an electronic device, which includes a memory and a processor coupled to each other, and the processor is used to execute program instructions stored in the memory to implement the training method of the interaction detection model in the first aspect and the interaction detection method in the second aspect above.
[0027] The sixth aspect of the present application provides a computer-readable storage medium, on which program instructions are stored. When the program instructions are executed by a processor, the training method of the interaction detection model in the above first aspect and the interaction detection method in the above second aspect are implemented.
[0028] In the above solution, based on the differences between the sample interaction actions and the first predicted interaction actions, as well as the differences between the sample scores and the predicted scores, the network parameters of the interaction detection model are adjusted. Therefore, on the one hand, the predicted score representing the possibility of the existence of human interaction with the sample object approaches the sample score representing whether the sample object interacts with the sample human body. Since the predicted score is predicted based on the three-dimensional position of the sample object, the three-dimensional position of the located sample object approaches the true three-dimensional position of the sample object, that is, it drives the positioning of the sample object to be as accurate as possible. And the positioning of the sample object is realized based on the sample human body features. Thus, from the positioning level, the interaction detection model can be forced to extract human body features that are closely related to the interaction actions with the human body as accurately as possible. On the other hand, the first predicted interaction action approaches the sample interaction action, and the first predicted interaction action is predicted based on the human body features. Thus, from the classification level, the interaction detection model can be forced to extract human body features that are closely related to the interaction actions with the human body as accurately as possible. Therefore, the interaction detection model is optimized from two dimensions: the positioning level and the classification level. When the subsequent interaction detection model classifies and detects the interaction actions of the human body, it can simultaneously pay attention to the actions of the human body itself and the position information of the interaction objects that have human interactions with the human body. Thus, the interaction detection model can extract human body features that are closely related to the human interaction actions, that is, the interaction detection model can accurately extract human body features, thereby improving the detection accuracy of the interaction detection model for human interaction relationships and reducing the false detection under the long-tail relationship distribution. BRIEF DESCRIPTION OF THE DRAWINGS
[0029] Figure 1 is a schematic flowchart of an embodiment of the training method of the interaction detection model provided by the present application;
[0030] Figure 2 is a schematic structural diagram of the interaction detection model provided by the present application;
[0031] Figure 3 is a schematic flowchart of an embodiment of predicting and generating the second predicted interaction action provided by the present application;
[0032] Figure 4 is Figure 1 a schematic flowchart of an embodiment of step S11 shown;
[0033] Figure 5 is Figure 1Flow schematic diagram of an embodiment of step S12 shown;
[0034] Figure 6 is Figure 1 Flow schematic diagram of an embodiment of step S13 shown;
[0035] Figure 7 is Figure 6 Flow schematic diagram of an embodiment of step S132 shown;
[0036] Figure 8 Flow schematic diagram of an embodiment of positioning the two - dimensional position of a sample object provided by the present application;
[0037] Figure 9 is Figure 8 Flow schematic diagram of an embodiment of step S81 shown;
[0038] Figure 10 Flow schematic diagram of an embodiment of the interaction detection method provided by the present application;
[0039] Figure 11 Structural schematic diagram of an embodiment of the training device for the interaction detection model provided by the present application;
[0040] Figure 12 Structural schematic diagram of an embodiment of the interaction detection device provided by the present application;
[0041] Figure 13 Structural schematic diagram of an embodiment of the electronic device provided by the present application;
[0042] Figure 14 Structural schematic diagram of an embodiment of the computer - readable storage medium provided by the present application. Detailed implementation manners
[0043] The solutions of the embodiments of the present application will be described in detail below with reference to the accompanying drawings of the specification.
[0044] In the following description, specific details such as specific system structures, interfaces, and technologies are presented for the purpose of illustration rather than limitation, so as to thoroughly understand the present application.
[0045] As used herein, the term "and / or" is merely a description of the relationship between associated objects, indicating that three relationships may exist. For example, A and / or B may represent: A exists alone, A and B exist simultaneously, and B exists alone. In addition, the character " / " in this text generally indicates that the associated objects before and after are in an "or" relationship. Furthermore, "plurality" in this text means two or more than two. In addition, the term "at least one" in this text means any one of a plurality or any combination of at least two of a plurality. For example, including at least one of A, B, and C may mean including any one or more elements selected from the set composed of A, B, and C.
[0046] Please refer to Figure 1 , Figure 1 is a schematic flowchart of an embodiment of the method for training an interaction detection model provided by this application. It should be noted that if there are substantially the same results, this embodiment is not limited to Figure 1 the process sequence shown. As Figure 1 shown, this embodiment includes:
[0047] Step S11: Process the sample images in the sample video data based on the feature extraction network of the interaction detection model to obtain the sample human body features of the sample human body in the sample images.
[0048] The method of this embodiment is used to improve the detection accuracy of the interaction action categories of the human body in video data. The video data described in this text may be a single video data or a combined video data synthesized from multiple video data later. In one implementation manner, the sample video data can be specifically obtained from local storage or cloud storage. It can be understood that in other implementation manners, the sample video data can also be obtained by collecting the current screen through a video capture device.
[0049] In this embodiment, the feature extraction network of the interaction detection model processes the sample images in the sample video data to obtain the sample human body features of the sample human body in the sample images. That is to say, the interaction detection model includes a feature extraction network, and the feature extraction network processes the sample video frame images in the sample video data, that is, the above-mentioned sample images, to obtain the sample human body features of the sample human body in the sample video frame images. In one embodiment, the sample video data includes several frames of sample images, and the feature extraction network of the interaction detection model processes the several frames of sample images in the sample video data frame by frame, that is, the feature extraction network of the interaction detection model processes each frame of corresponding sample images in the sample video data. To reduce the computational complexity and improve the processing efficiency of the feature extraction network of the interaction detection model for the sample images in the sample video data, in other embodiments, the feature extraction network of the interaction detection model may also only process the sample images corresponding to some frames in the sample video data. For example, it processes the sample images corresponding to the even frames or odd frames in the sample video data, etc.
[0050] Among them, the sample video data is labeled with a sample score indicating whether the sample object interacts with the sample human body; and, the sample video data is labeled with the sample interaction actions of the sample human body that interacts with the sample object. By labeling the sample score indicating whether the sample object interacts with the sample human body and the sample interaction actions of the sample human body that interacts with the sample object on the sample video data, it enables subsequent adjustment of the network parameters of the interaction detection model based on the sample score indicating whether the sample object interacts with the sample human body and the sample interaction actions of the sample human body that interacts with the sample object, that is, adjusting the network parameters of the interaction detection model in two aspects to make the interaction detection model converge, thereby improving the detection accuracy of the interaction detection model for human interaction actions.
[0051] In one embodiment, the feature extraction network of the interaction detection model can be directly used to extract features from the sample images in the sample video data to obtain the sample human body features of the sample human body in the sample images. To accurately extract the sample human body features, that is, to improve the feature extraction accuracy, in other embodiments, the human body region of the sample human body in the sample image and the sample image features of the sample image can be obtained first; then, based on the human body region, the sample human body features corresponding to the human body region are extracted from the sample image features.
[0052] Step S12: Classify the sample human body features based on the action classification network of the interaction detection model to obtain the first predicted interaction action of the sample human body.
[0053] In this embodiment, the action classification network of the interaction detection model classifies the sample human body features to obtain the first predicted interaction action of the sample human body. That is to say, as Figure 2As shown Figure 2 is a schematic structural diagram of an interaction detection model provided by the present application. The interaction detection model includes an action classification network, which classifies the interaction actions of a sample human body according to the sample human body features, so as to obtain the first predicted interaction action of the sample human body. Specifically, the sample human body features are sent into the action classification network, such as a Multilayer Perceptron (MLP) network. The MLP network acts as a classifier and classifies the interaction actions of the sample human body based on the sample human body features, so as to obtain the first predicted interaction action of the sample human body.
[0054] In one embodiment, the action classification network of the interaction detection model can directly classify based on the sample human body features of the same sample human body in several frame sample images in the sample video data, so as to obtain the first predicted interaction action of the sample human body. Of course, in other embodiments, the sample human body features of the same sample human body in each frame sample image in the sample video data can also be combined to obtain the sample action trajectory features of the sample human body, and then the trajectory features are classified to obtain the first predicted interaction action of the sample human body.
[0055] Step S13: Locate based on the two-dimensional position of the sample object and the morphological parameters of the sample human body to obtain the three-dimensional position of the sample object.
[0056] In this embodiment, the three-dimensional position of the sample object is obtained by locating based on the two-dimensional position of the sample object and the morphological parameters of the sample human body. Since the sample objects participating in the interaction usually have different degrees of occlusion and different shapes under different perspectives, the predicted score representing the possibility of human-object interaction of the sample object obtained by directly predicting through the two-dimensional position of the sample object and the morphological parameters of the sample human body may be inaccurate in the subsequent process. Therefore, it is necessary to locate the three-dimensional position of the sample object by combining the two-dimensional position of the sample object and the morphological parameters of the sample human body, so that the possibility of human-object interaction of the sample object can be determined based on the three-dimensional position of the sample object and the morphological parameters of the sample human body in the subsequent process.
[0057] In one embodiment, an MLP network model is used to perform 3D position localization of a sample object based on the 2D position of the sample object and the morphological parameters of the sample human body. Specifically, based on the 2D position of the sample object, the MLP network model is used to predict the 3D position of the sample object, such as the central position of the sample object and the radius of the sample object, etc.; combining a loss function (e.g., smooth L1 loss function, L1 loss function, or L2 loss function, etc.) and the difference between the predicted 3D position of the sample object and the labeled 3D position of the sample object, the loss of the MLP network model is obtained; using the loss of the MLP network model, the network parameters of the MLP network model are adjusted to optimize the MLP network model, so that the optimized MLP network model can more accurately localize the 3D position of the sample object. Of course, in other embodiments, a 2D vision neural network (Detailed Joint Representation Network, DJ-RN) can also be used to perform 3D position localization of the sample object based on the 2D position of the sample object and the morphological parameters of the sample human body, which is not specifically limited herein.
[0058] In one embodiment, the sample image in the sample video data can be processed by a method of single-stage regression of all 3D meshes for multiple humans (Regression of Multiple 3D People, ROMP) to obtain the morphological parameters of the sample human body in the sample image. Of course, the morphological parameters of the sample human body in the sample image can also be obtained by other means, which is not specifically limited herein. In a specific embodiment, the morphological parameters include human pose parameters and human shape parameters, that is, subsequently, based on the human pose parameters, human shape parameters, and 3D position, a prediction score of the sample object is obtained. Of course, in other specific embodiments, the morphological parameters may also only include human pose parameters or human shape parameters, etc., which is not specifically limited herein.
[0059] Among them, the 2D position is obtained by localizing based on the sample human body features of the sample human body. In one embodiment, the 2D position of the sample object can be directly obtained by localizing based on the sample human body features of the sample human body. It can be understood that in other embodiments, the 2D position of the sample object can also be obtained by jointly localizing based on the sample human body features of the sample human body and the sample object features of the sample object, so that the sample object features can assist the sample human body features in the 2D localization of the sample object, which is beneficial to improving the accuracy of 2D localization.
[0060] In a specific embodiment, the bounding box regression technique can be used to directly locate the two-dimensional position of the sample object based on the sample human body features of the sample human body. This method is more direct and simple for locating the two-dimensional position of the sample object. Specifically, the sample image features of the sample image and the sample human body features of the sample human body are concatenated in the feature channel dimension; then, an MLP network model with a Sigmoid (Sigmoid function) activation function and two fully connected layers is used to predict the two-dimensional position of the normalized sample object; combining the loss function (e.g., smooth L1 loss function, L1 loss function, or L2 loss function, etc.) and the difference between the predicted two-dimensional position of the sample object and the labeled two-dimensional position of the sample object, the loss of the MLP network model is obtained; using the loss of the MLP network model, the network parameters of the MLP network model are adjusted to optimize the MLP network model, so that the optimized MLP network model can more accurately locate the two-dimensional position of the sample object.
[0061] In other specific embodiments, the offset between the sample object and the sample human body can also be used to locate the two-dimensional position of the sample object, so as to directly locate the two-dimensional position of the sample object based on the sample human body features of the sample human body. Specifically, first calculate the offset between the sample human body and the sample object, and the specific formula is as follows:
[0062]
[0063] where x h represents the abscissa of the sample human body; y h represents the ordinate of the sample human body; x o represents the abscissa of the sample object; y o represents the ordinate of the sample object; w h represents the width of the human body area of the sample human body; h h represents the height of the human body area of the sample human body. It should be noted that the coordinates of the sample human body can be the center point coordinates, the upper left vertex coordinates, or the lower right vertex coordinates of the human body detection frame of the sample human body, etc., which are not specifically limited here; the coordinates of the sample object can be the center point coordinates, the upper left vertex coordinates, or the lower right vertex coordinates of the object detection frame of the sample object, etc., which are not specifically limited here.
[0064] Then, the sample image features of the sample image and the sample human body features of the sample human body are concatenated in the feature channel dimension; then, an MLP network model is used to predict the offset between the normalized sample object and the sample human body; combining a loss function (e.g., smooth L1 loss function, L1 loss function, or L2 loss function, etc.) and the difference between the predicted offset of the sample object and the sample human body and the calculated offset of the sample object and the sample human body, the loss of the MLP network model is obtained; using the loss of the MLP network model, the network parameters of the MLP network model are adjusted to optimize the MLP network model, so that the optimized MLP network model can more accurately predict the offset between the sample object and the sample human body.
[0065] Furthermore, the two-dimensional position of the sample object is calculated using the offset between the sample object and the sample human body.
[0066] Step S14: Based on the morphological parameters and the three-dimensional position, a prediction is made to obtain the prediction score of the sample object.
[0067] In this embodiment, based on the morphological parameters and the three-dimensional position, a prediction is made to obtain the prediction score of the sample object, where the prediction score represents the possibility of human-object interaction for the sample object. The higher the prediction score of the sample object, the closer the interaction relationship between the sample object and the sample human body, that is, the greater the possibility of interaction between the sample object and the sample human body. That is to say, by combining the morphological parameters and the three-dimensional position of the sample object, the possibility of interaction between the sample object and the sample human body can be determined.
[0068] Specifically, the morphological parameters of the sample human body and the three-dimensional position of the sample object are concatenated in the feature channel dimension; then, an MLP network model with two fully connected layers is used to predict the possibility of human-object interaction for the sample object, that is, the prediction score of the sample object is obtained using the MLP network model for prediction.
[0069] Step S15: Based on the difference between the sample interaction action and the first predicted interaction action, and the difference between the sample score and the prediction score, the network parameters of the interaction detection model are adjusted.
[0070] In this embodiment, based on the difference between the sample interaction action and the first predicted interaction action and the difference between the sample score and the prediction score, the network parameters of the interaction detection model are adjusted. Specifically, combining a loss function and the difference between the first predicted interaction action and the sample interaction action and the difference between the sample score and the prediction score, the loss of the interaction detection model is obtained; using the obtained loss of the interaction detection model, the network parameters of the interaction detection model are adjusted; using the above steps to perform iterative training on the interaction detection model, and finally an interaction detection model with network convergence is obtained, and at this time, the training of the interaction detection model is completed.
[0071] Since the network parameters of the interaction detection model are adjusted based on the differences between the sample interaction actions and the first predicted interaction actions, as well as the differences between the sample scores and the predicted scores, by adjusting the network parameters of the interaction detection model, it is possible to minimize the differences between the sample interaction actions and the first predicted interaction actions, and between the sample scores and the predicted scores. On the one hand, the predicted score representing the possibility of human interaction with the sample object approaches the sample score indicating whether the sample object interacts with the sample human body. Since the predicted score is predicted based on the three-dimensional position of the sample object, the three-dimensional position of the located sample object approaches the true three-dimensional position of the sample object, that is, it drives the positioning of the sample object to be as accurate as possible. And the positioning of the sample object is based on the sample human body features, so that the interaction detection model can be forced to extract human body features closely related to the interaction actions with the human body as accurately as possible from the positioning level. On the other hand, the first predicted interaction action approaches the sample interaction action, and the first predicted interaction action is predicted based on the human body features, so that the interaction detection model can be forced to extract human body features closely related to the interaction actions with the human body as accurately as possible from the classification level. Therefore, based on the differences between the sample interaction actions and the first predicted interaction actions, and between the sample scores and the predicted scores, the network parameters of the interaction detection model are adjusted to optimize the interaction detection model from two dimensions: the positioning level and the classification level. Thus, when the subsequent interaction detection model classifies and detects the interaction actions of the human body, it can simultaneously pay attention to the actions of the human body itself and the position information of the interaction objects that have human interactions with the human body, so that the interaction detection model can extract human body features closely related to the human interaction actions, that is, the interaction detection model can accurately extract human body features, and further improve the accuracy of the subsequent interaction detection model in detecting the human interaction relationship based on the human body features, reducing the false detection under the long-tail relationship distribution.
[0072] In order to further improve the detection accuracy of the interaction detection model for the human interaction relationship, in one embodiment, before adjusting the network parameters of the interaction detection model based on the differences between the sample interaction actions and the first predicted actions, as well as the differences between the sample scores and the predicted scores, it is also necessary to predict based on the morphological parameters of the sample human body and the initial position of the sample object in the three-dimensional space to obtain the second predicted interaction action of the sample human body. Among them, the morphological parameters of the sample human body can be human body pose parameters, or human body shape parameters, or can also include both human body pose parameters and human body shape parameters.
[0073] In one embodiment, as Figure 3 shown Figure 3It is a schematic flowchart of an embodiment for predicting and generating a second predicted interaction action provided by this application. The morphological parameters of the sample human body include human posture parameters and human shape parameters. Predicting the second predicted interaction action of the sample human body based on the morphological parameters of the sample human body and the initial position of the sample object in the three-dimensional space specifically includes the following sub-steps:
[0074] Step S31: Encode based on the human posture parameters to obtain a posture encoding representation.
[0075] In this embodiment, encode based on the human posture parameters to obtain a posture encoding representation. Specifically, assume that the human posture parameter is θ h ; Use a posture encoder (e.g., Vposer) to encode the human posture parameter θ h to obtain an implicit representation, that is, a posture encoding representation
[0076] Step S32: Concatenate the posture encoding representation, the human shape parameters, and the initial position to obtain a concatenated feature representation.
[0077] In this embodiment, concatenate the posture encoding representation, the human shape parameters, and the initial position to obtain a concatenated feature representation. Specifically, assume that the human shape parameter is β h and the initial position of the sample object in the three-dimensional space is Concatenate the posture encoding representation the human shape parameter β h and the initial position of the sample object in the three-dimensional space in the feature channel dimension to obtain a concatenated feature representation.
[0078] Step S33: Predict based on the concatenated feature representation to obtain the second predicted interaction action.
[0079] In this embodiment, predict based on the concatenated feature representation to obtain the second predicted interaction action. Since the concatenated feature representation is obtained by concatenating the posture encoding representation the human shape parameter β h and the initial position so it is based on the posture encoding representation the human shape parameter β h and the initial position to predict and obtain the second predicted interaction action of the sample human body. Specifically, use the DJ-RN method to predict based on the two-dimensional position of the sample object to obtain the initial position of the sample object in the three-dimensional space Then, the posture encoding representation corresponding to the human posture parameter θ h the human shape parameter β h and the initial positions of the sample object in three-dimensional space Input into the trained MLP network model with two fully-connected layers to obtain the second predicted interaction action of the sample human body.
[0080] In one embodiment, after predicting the second predicted interaction action of the sample human body, the specific method for adjusting the network parameters of the interaction detection model at this time is: based on the differences between the sample interaction action and the first predicted action, the sample score and the predicted score, and the first predicted interaction action and the second predicted interaction action, adjust the network parameters of the interaction detection model. Since the first predicted interaction action and the second predicted interaction action both correspond to the same sample human body, the second predicted interaction action should be consistent with the first predicted interaction action. Therefore, by introducing the difference between the first predicted interaction action and the second predicted interaction action to adjust the network parameters of the interaction detection model, the second predicted interaction action and the first predicted interaction action are kept consistent. Since the second predicted interaction action is predicted based on the three-dimensional position of the sample object, it can make the located three-dimensional position of the sample object approach the true three-dimensional position of the sample object, that is, drive the positioning of the sample object to be as accurate as possible. And the positioning of the sample object is realized based on the sample human body features. Thus, from the perspective of interaction action consistency, force the interaction detection model to extract human body features that are closely related to the interaction action of the human body as much as possible, that is, force the human body features extracted by the interaction detection model to be as accurate as possible from the level of interaction action consistency. That is to say, optimize the interaction detection model from three dimensions: the positioning level, the classification level, and the action consistency level. Thus, when the subsequent interaction detection model classifies and detects the interaction actions of the human body, it can extract human body features that are closely related to the interaction actions of the human body, that is, the interaction detection model can accurately extract human body features, and then improve the detection accuracy of the interaction detection model for the human interaction relationship and reduce the false detection under the long-tail relationship distribution.
[0081] In the above-described embodiments, based on the differences between the sample interaction actions and the first predicted interaction actions, as well as the differences between the sample scores and the predicted scores, the network parameters of the interaction detection model are adjusted. Therefore, on the one hand, the predicted score representing the possibility of human interaction with the sample object approaches the sample score indicating whether the sample object interacts with the sample human body. Since the predicted score is obtained by predicting based on the three-dimensional position of the sample object, the three-dimensional position of the located sample object approaches the true three-dimensional position of the sample object, that is, it drives the positioning of the sample object to be as accurate as possible. And the positioning of the sample object is realized based on the sample human body features, so that the interaction detection model can be forced to extract human body features closely related to the interaction actions with the human body as accurately as possible from the positioning level. On the other hand, the first predicted interaction action approaches the sample interaction action, and the first predicted interaction action is obtained by predicting based on the human body features, so that the interaction detection model can be forced to extract human body features closely related to the interaction actions with the human body as accurately as possible from the classification level. Therefore, the interaction detection model is optimized from two dimensions: the positioning level and the classification level. When the subsequent interaction detection model classifies and detects the interaction actions of the human body, it can simultaneously pay attention to the actions of the human body itself and the position information of the interaction objects that have human interactions with the human body, so that the interaction detection model can extract human body features closely related to the interaction actions of the human body, that is, the interaction detection model can accurately extract human body features, thereby improving the accuracy of the subsequent interaction detection model for detecting human interaction relationships based on human body features and reducing the false detection under the long-tail relationship distribution.
[0082] Please refer to Figure 4 , Figure 4 is Figure 1 a schematic flowchart of an embodiment of step S11 shown in the figure. It should be noted that if there are substantially the same results, this embodiment is not limited to Figure 4 the process sequence shown in the figure. As Figure 4 shown in the figure, in this embodiment, based on the human body region, the sample human body features corresponding to the human body region are extracted from the sample image features, which specifically include:
[0083] Step S111: Extract the sample image features of the sample image based on the feature extraction network, and perform human body detection on the sample image to obtain the human body region of the sample human body.
[0084] In this embodiment, the sample image features of the sample image are extracted by using the feature extraction network of the interaction detection model. Among them, the feature extraction network is not limited and can be specifically set according to actual usage needs. For example, the feature extraction network can be a SlowFast network, a Multi-Fiber Networks (MF-Net), a TwoStream Network, or a Temporal Segment Networks (TSN), etc.
[0085] In addition, in this embodiment, human detection is also performed on the sample images in the sample video data to obtain the human regions of the sample humans, so as to facilitate the subsequent extraction of sample human features from the sample image features. Among them, the network model used for human detection is not limited and can be specifically set according to actual usage needs. For example, the ResNeXt-101-FPN network, the Region-CNN (R-CNN), the Fast R-CNN network, the Faster R-CNN network, or the Region Proposal Network (RPN), etc. are used to perform human detection on the sample images in the sample video data to obtain the human regions of the sample humans.
[0086] Step S112: Extract sample human features from the sample image features based on the human region.
[0087] In this embodiment, sample human features are extracted from the sample image features based on the human region. That is to say, the region corresponding to the human region in the sample image features is determined according to the human region of the sample human, and then features are extracted from the region corresponding to the sample human region in the sample image features to obtain the sample human features. Among them, the network model for extracting sample human features is not limited and can be specifically set according to actual usage needs. For example, the network model for extracting sample human features from the sample image features can be an ROI-Align network, an ROI Pooling network, etc.
[0088] Specifically, according to the ratio between the resolution of the sample image features and the resolution of the sample image from which the sample image features are extracted, the position of the human region in the sample image features and the size of the human region in the sample image features are determined; then, features are extracted from the region corresponding to the human region in the sample image features, and the extracted features are the sample human features.
[0089] Please refer to Figure 5 , Figure 5 which Figure 1 is a schematic flowchart of an embodiment of step S12 shown. It should be noted that if there are substantially the same results, this embodiment does not Figure 5It is limited to the shown process sequence. For example, Figure 5 As shown, in this embodiment, the sample video data includes several frames of sample images, and the feature extraction network of the interaction detection model processes each frame of the corresponding sample image in the sample video data, specifically including:
[0090] Step S121: Based on the sample human body features of the same sample human body in each frame of the sample images, obtain the sample action trajectory features of the sample human body.
[0091] In this implementation manner, based on the sample human body features of the same sample human body in each frame of the sample images, obtain the sample action trajectory features of the sample human body. Specifically, the sample video data includes several frames of sample images, and the feature extraction network of the interaction detection model processes the several frames of sample images in the sample video data frame by frame, that is, the feature extraction network of the interaction detection model extracts features from each frame of the corresponding sample image in the sample video data to obtain the sample human body features corresponding to each frame of the sample image, and then combines the sample human body features of the same sample human body in each frame of the sample images to obtain the sample action trajectory features of the sample human body.
[0092] Step S122: Classify based on the sample action trajectory features to obtain the first predicted interaction action of the sample human body.
[0093] In this implementation manner, classify based on the sample action trajectory features to obtain the first predicted interaction action of the sample human body. Specifically, the action classification network of the interaction detection model classifies the interaction actions of the sample human body according to the sample action trajectory, so as to obtain the first predicted interaction action of the sample human body.
[0094] Please refer to Figure 6 , Figure 6 which Figure 1 is a schematic flowchart of an embodiment of step S13 shown. It should be noted that if there are substantially the same results, this embodiment is not limited to Figure 6 the shown process sequence. For example, Figure 6 As shown, in this embodiment, the DJ-RN method is used to predict the three-dimensional position of the sample object based on the two-dimensional position of the sample object and the morphological parameters of the sample human body, specifically including:
[0095] Step S131: Based on the two-dimensional position, predict the initial position of the sample object in the three-dimensional space.
[0096] In this embodiment, based on the two-dimensional position, the initial position of the sample object in the three-dimensional space is predicted. That is to say, when using the DJ-RN method to determine the three-dimensional position of the sample object, the two-dimensional position of the sample object is input to determine the three-dimensional position on the basis of the two-dimensional position. Specifically, based on the two-dimensional position of the sample object, the DJ-RN method is used to predict the initial position of the sample object in the three-dimensional space
[0097] Step S132: Based on the initial position and the morphological parameters, construct an action classification loss with the correction parameter of the initial position as the optimization target
[0098] In this embodiment, based on the initial position of the sample object in the three-dimensional space and the morphological parameters of the sample human body, construct an action classification loss with the correction parameter of the initial position as the optimization target. That is to say, the action classification loss with the correction parameter of the initial position of the sample object in the three-dimensional space as the optimization target is used as an implicit guidance centered on the sample human body to correct the initial position of the sample object predicted by the DJ-RN method in the subsequent step
[0099] Please refer to Figure 7 , Figure 7 is Figure 6 the schematic flowchart of an embodiment of step S132 shown. It should be noted that if there are substantially the same results, this embodiment is not limited to Figure 7 the flowchart order shown. As Figure 7 shown, the morphological parameters include human body pose parameters and human body shape parameters. This embodiment includes
[0100] Step S71: Encode based on the human body pose parameters to obtain a pose encoding representation
[0101] In this embodiment, encode based on the human body pose parameters to obtain a pose encoding representation. Specifically, assume that the human body pose parameter is θ h ; use a pose encoder (for example, Vposer) to encode the human body pose parameter θ h to obtain an implicit representation, that is, a pose encoding representation
[0102] Step S72: Based on the pose encoding representation, the human body shape parameters, and the initial position, make a prediction to obtain the third predicted interaction action of the sample human body
[0103] In this embodiment, based on the pose encoding representation, the human body shape parameters, and the initial position, make a prediction to obtain the third predicted interaction action of the sample human body. Specifically, assume that the human body shape parameter is β h and the initial position of the sample object in the three-dimensional space is Predict based on the two-dimensional position of the sample object using the DJ-RN method to obtain the initial position of the sample object in three-dimensional space Then, the human body pose parameter θ h The corresponding pose encoding representation The human body shape parameter β h And the initial position of the sample object in three-dimensional space Are input into the trained MLP network model with two fully connected layers to obtain the third predicted interaction action of the sample human body.
[0104] Step S73: Based on the difference between the sample interaction action and the third predicted interaction action, obtain the action classification loss.
[0105] In this embodiment, based on the difference between the sample interaction action and the third predicted interaction action, obtain the action classification loss. Among them, the third predicted interaction action is predicted based on the human body pose parameter, the human body shape parameter, and the initial position. Therefore, the specific formula of the action classification loss can be shown as follows:
[0106]
[0107] Among them, L represents the action classification loss; θ h Represents the human body pose parameter; β h Represents the human body shape parameter; Represents the initial position of the sample object in three-dimensional space; Represents the correction parameter.
[0108] Step S133: Based on the action classification loss, optimize to obtain the correction parameter, and based on the correction parameter, perform position correction on the initial position to obtain the three-dimensional position.
[0109] In this embodiment, based on the action classification loss, optimize to obtain the correction parameter, and based on the correction parameter, perform position correction on the initial position of the sample object in three-dimensional space, so as to obtain the three-dimensional position of the sample object. That is to say, introduce the action classification loss to guide the optimization of the correction parameter to minimize the action classification loss; then, use the optimized correction parameter to correct the initial position of the sample object in three-dimensional space, so as to obtain the corrected three-dimensional position of the sample object. The corrected three-dimensional position of the sample object approximates the true three-dimensional position of the sample object, that is, introducing the action classification loss improves the accuracy of the three-dimensional position positioning of the sample object. Among them, the formula for the corrected three-dimensional position of the sample object is as follows:
[0110]
[0111] Among them, Is the corrected three-dimensional position of the sample object; Is the action classification loss.
[0112] Please refer to Figure 8 , Figure 8 which is a schematic flowchart of an embodiment for positioning the two-dimensional position of a sample object provided by this application. It should be noted that if there are substantially the same results, this embodiment is not limited to Figure 8 the process sequence shown. As Figure 8 shown, in this embodiment, the two-dimensional position of the sample object is jointly determined based on the sample human body characteristics of the sample human body and the sample object characteristics of the sample object, specifically including:
[0113] Step S81: Based on the sample image features of the sample image and the object region detected in the sample image, extract the sample object characteristics of the sample object in the object region.
[0114] In this embodiment, based on the sample image features of the sample image and the object region detected in the sample image, the sample object characteristics of the sample object in the object region are extracted. Among them, the sample image features are obtained by the feature extraction network extracting the sample image. Specifically, use the feature extraction network of the interactive detection model to extract the sample image features of the sample image; then, perform object detection on the sample image in the sample video data to obtain the object region of the sample object; then, according to the ratio between the resolution of the sample image features and the resolution of the sample image from which the sample image features are extracted, determine the position of the object region in the sample image features and the size of the object region in the sample image features; then, extract the features from the region corresponding to the object region in the sample image features, and the extracted features are the sample object characteristics. It should be noted that after obtaining the above ratio, only the position and region size detected in the sample image need to be scaled according to this ratio to obtain the position of the object region in the sample image features and the size of the object region in the sample image features.
[0115] Among them, the network model for object detection is not limited and can be specifically set according to actual usage needs. For example, use ResNeXt-101-FPN network, Region-CNN (R-CNN), Fast R-CNN network, Faster R-CNN network or Region Proposal Network (RPN), etc. to perform object detection on the sample image in the sample video data to obtain the object region of the sample object.
[0116] Since the sample image may include several objects, when performing object detection on the sample image, several candidate regions may be detected. At this time, in one embodiment, the several candidate regions may be respectively used as object regions to extract the sample object features of the sample objects in the corresponding object regions. For example, the sample image includes three objects a, b, and c. Therefore, when performing object detection on the sample image, candidate regions A corresponding to object a, candidate region B corresponding to object b, and candidate region C corresponding to object c will be detected. The candidate regions A, B, and C are respectively used as object regions to execute step S71.
[0117] To reduce the computational amount, as Figure 9 shown, Figure 9 is Figure 8 a schematic flowchart of an embodiment of step S81 shown. In other embodiments, a preliminary screening will be performed on the several detected candidate regions, and the remaining candidate regions after the preliminary screening will be used as object regions, which specifically includes the following sub-steps:
[0118] Step S811: Perform object detection based on the sample image to obtain several candidate regions.
[0119] In this embodiment, object detection is performed based on the sample image to obtain several candidate regions. Since the sample image may include several objects, when performing object detection on the sample image, candidate regions corresponding to each object will be obtained, that is, several candidate regions will be obtained.
[0120] Among them, the network model used for object detection is not limited and can be specifically set according to actual usage needs. For example, use ResNeXt-101-FPN network, Region-CNN (R-CNN), Fast R-CNN network, Faster R-CNN network, or Region Proposal Network (RPN), etc. to perform object detection on the sample image in the sample video data to obtain several candidate regions.
[0121] Step S812: Select a candidate region as an object region based on the first predicted interaction action.
[0122] In this embodiment, a candidate region is selected as an object region based on the first predicted interaction action. That is to say, through the first predicted interaction action, the several obtained candidate regions can be screened out, and the remaining candidate regions after the screening are used as object regions.
[0123] For example, the first predicted interaction action is stepping; the sample image includes three objects a, b, and c. Therefore, when performing object detection on the sample image, candidate regions A corresponding to object a, candidate region B corresponding to object b, and candidate region C corresponding to object c will be detected. Since candidate region B and candidate region C are located below the sample human body, and candidate region A is located above the sample human body, and the first predicted interaction action is stepping, candidate regions B and C are respectively used as object regions.
[0124] Step S813: Based on the object regions, extract sample object features from the sample image features.
[0125] In this embodiment, based on the object regions, sample object features are extracted from the sample image features. Specifically, according to the ratio between the resolution of the sample image features and the resolution of the sample image from which the sample image features are extracted, determine the position of the object regions in the sample image features and the size of the object regions in the sample image features; then, extract features from the regions corresponding to the object regions in the sample image features, and the extracted features are the sample object features.
[0126] Step S82: Perform prediction based on the sample object features and the sample human body features to obtain an interaction score.
[0127] In this embodiment, prediction is performed based on the sample object features and the sample human body features to obtain an interaction score. Among them, the interaction score represents the degree of interaction tightness between the sample human body to which the sample human body features belong and the sample object in the object region. That is to say, by performing prediction according to the sample object features and the sample human body features, it is possible to determine the possibility of interaction between the sample human body to which the sample human body features belong and the sample object in the object region, so as to determine the object region corresponding to the sample object that interacts with the sample human body, thereby facilitating the subsequent positioning of the two-dimensional position of the sample object.
[0128] Specifically, splice the sample human body features of the sample human body and the sample object features of the sample object in the feature channel dimension; then, use the MLP network model to perform prediction based on the spliced features, and predict whether there is an interaction between the sample human body to which the sample human body features belong and the sample object in the object region and the corresponding interaction score; if it is predicted that there is an interaction between the sample human body to which the sample human body features belong and the sample object in the object region and the interaction score is relatively high, it indicates that the possibility of interaction between the sample human body to which the sample human body features belong and the sample object in this object region is relatively high, and if it is predicted that there is an interaction between the sample human body to which the sample human body features belong and the sample object in the object region and the interaction score is relatively low or there is no interaction and the corresponding non-interaction score is relatively high, it indicates that the possibility of interaction between the sample human body to which the sample human body features belong and the sample object in this object region is relatively low.
[0129] Step S83: Based on the object region corresponding to the interaction score that meets the preset condition, obtain the two-dimensional position of the sample object.
[0130] In this embodiment, based on the object region corresponding to the interaction score that meets the preset condition, obtain the two-dimensional position of the sample object. That is to say, when the interaction score meets the preset requirement, it indicates that the sample object in this object region has the greatest possibility of interaction with the sample human body to which the sample human body feature belongs. Then, this object region is used as the region corresponding to the sample object that interacts with the sample human body. Therefore, according to the determined object region corresponding to the sample object that interacts with the sample human body, the two-dimensional position of the sample object can be determined. For example, the center position of the object region is used as the two-dimensional position of the sample object in the object region, etc.
[0131] Among them, the preset requirement is not limited and can be specifically set according to actual usage needs. For example, the interaction score is greater than the preset value and is the largest, etc. For example, taking the preset requirement as the interaction score being greater than 5 and the interaction score being the largest as an example; perform object detection on the sample image to obtain the object region A corresponding to object a, and the interaction score corresponding to object region A is 7, the object region B corresponding to object b, and the interaction score corresponding to object region B is 9, and the object region C corresponding to object c, and the interaction score corresponding to object region C is 5; since the interaction score of object region A meets the preset requirement, the two-dimensional position of the sample object in object region A is obtained based on object region A.
[0132] Please refer to Figure 10 , Figure 10 which is a schematic flowchart of an embodiment of the interaction detection method provided by this application. It should be noted that if there are substantially the same results, this embodiment is not limited to Figure 10 the shown process sequence. As Figure 10 shown, this embodiment includes:
[0133] Step S101: Based on the feature extraction network of the interaction detection model, extract features from the to-be-detected image in the to-be-detected video data to obtain the human body features of the human body in the to-be-detected image.
[0134] The method of this embodiment is used to improve the detection accuracy of the interaction action category of the human body in the to-be-detected video data. The to-be-detected video data described in this article can be a single video data or a combined video data synthesized from multiple video data later. In one embodiment, the to-be-detected video data can be specifically obtained from local storage or cloud storage. It can be understood that in other embodiments, the to-be-detected video data can also be obtained by collecting the current picture through a video capture device.
[0135] In this embodiment, the feature extraction network of the interaction detection model extracts features from the to-be-tested images in the to-be-tested video data to obtain the human body features of the human body in the to-be-tested images. That is to say, the interaction detection model includes a feature extraction network, and the feature extraction network extracts features from the video frame images in the to-be-tested video data, that is, the above-mentioned to-be-tested images, to obtain the human body features of the human body in the to-be-tested video frame images. In one embodiment, the to-be-tested video data includes several to-be-tested images, and the feature extraction network of the interaction detection model processes the several to-be-tested images in the to-be-tested video data frame by frame, that is, the feature extraction network of the interaction detection model extracts features from each corresponding to-be-tested image in the to-be-tested video data. In order to reduce the computational amount and improve the efficiency of the feature extraction network of the interaction detection model in extracting features from the to-be-tested images in the to-be-tested video data, in other embodiments, the feature extraction network of the interaction detection model may also extract features only from the to-be-tested images corresponding to some frames in the to-be-tested video data. For example, extract features from the to-be-tested images corresponding to the even frames or odd frames in the to-be-tested video data, etc.
[0136] Among them, the interaction detection model is trained by using the training method of the above-mentioned interaction detection model. Therefore, when the interaction detection model extracts features from the to-be-tested images in the to-be-tested video data, it can simultaneously focus on the actions of the human body itself and the position information of the interaction objects that have human interactions with the human body, so as to be able to extract human body features closely related to the human interaction actions, that is, accurately extract human body features, thereby improving the detection accuracy of the subsequent interaction detection model in detecting the human interaction relationship based on the human body features, and reducing the false detection under the long-tail relationship distribution.
[0137] Step S102: Classify the human body features based on the action classification network of the interaction detection model to obtain the interaction action categories of the human body.
[0138] In this embodiment, the action classification network of the interaction detection model classifies the human body features to obtain the interaction action categories of the human body. That is to say, the interaction detection model includes an action classification network, and the action classification network classifies the interaction actions of the human body according to the human body features, so as to obtain the interaction action categories of the human body.
[0139] Please refer to Figure 11 , Figure 11FIG. 0 is a schematic structural diagram of an embodiment of a training device for an interaction detection model. The training device 110 of the interaction detection model includes a sample feature extraction module 111, an interaction action prediction module 112, a three-dimensional position localization module 113, an interaction score prediction module 114, and a network parameter adjustment module 115. The sample feature extraction module 111 is configured to process the sample images in the sample video data based on the feature extraction network of the interaction detection model to obtain the sample human body features of the sample human body in the sample images; wherein, the sample video data is labeled with a sample score indicating whether the sample object interacts with the sample human body, and the sample interaction action of the sample human body interacting with the sample object; the interaction action prediction module 112 is configured to classify the sample human body features based on the action classification network of the interaction detection model to obtain the first predicted interaction action of the sample human body; the three-dimensional position localization module 113 is configured to perform localization based on the two-dimensional position of the sample object and the morphological parameters of the sample human body to obtain the three-dimensional position of the sample object; wherein, the two-dimensional position is obtained by localization based on the sample human body features; the interaction score prediction module 114 is configured to perform prediction based on the morphological parameters and the three-dimensional position to obtain the predicted score of the sample object; wherein, the predicted score represents the possibility of the sample object having a human interaction; the network parameter adjustment module 115 is configured to adjust the network parameters of the interaction detection model based on the difference between the sample interaction action and the first predicted interaction action, and the difference between the sample score and the predicted score.
[0140] Wherein, the training device 110 of the interaction detection model further includes a two-dimensional position localization module 116, and the two-dimensional position localization module 116 is configured to obtain the two-dimensional position by localization based on the sample human body features, specifically including: jointly localizing the two-dimensional position of the sample object based on the sample human body features and the sample object features of the sample object.
[0141] Wherein, the two-dimensional position localization module 116 is configured to jointly localize the two-dimensional position of the sample object based on the sample human body features and the sample object features of the sample object, specifically including: extracting the sample object features of the sample object in the object region based on the sample image features of the sample image and the object region detected in the sample image; predicting an interaction score based on the sample object features and the sample human body features; wherein, the interaction score represents the degree of interaction tightness between the sample human body to which the sample human body features belong and the sample object in the object region; obtaining the two-dimensional position of the sample object based on the object region corresponding to the interaction score that meets the preset conditions.
[0142] Among them, the two-dimensional position positioning module 116 is used to extract the sample object features of the sample object in the object area based on the sample image features of the sample image and the detected object area in the sample image, specifically including: performing object detection on the sample image to obtain a number of candidate areas; selecting a candidate area as the object area based on the first predicted interaction action; and extracting the sample object features from the sample image features based on the object area.
[0143] Among them, the sample feature extraction module 111 is used to process the sample image in the sample video data based on the feature extraction network of the interaction detection model to obtain the sample human body features of the sample human body in the sample image, specifically including: extracting the sample image features of the sample image based on the feature extraction network and performing human body detection on the sample image to obtain the human body area of the sample human body; and extracting the sample human body features from the sample image features based on the human body area.
[0144] Among them, the interaction action prediction module 112 is used to perform positioning based on the two-dimensional position of the sample object and the morphological parameters of the sample human body to obtain the three-dimensional position of the sample object, specifically including: predicting the initial position of the sample object in the three-dimensional space based on the two-dimensional position; constructing an action classification loss with the correction parameter of the initial position as the optimization target based on the initial position and the morphological parameters; optimizing to obtain the correction parameter based on the action classification loss, and performing position correction on the initial position based on the correction parameter to obtain the three-dimensional position.
[0145] Among them, the interaction action prediction module 112 is also used to, before adjusting the network parameters of the interaction detection model based on the difference between the sample interaction action and the first predicted interaction action, and the difference between the sample score and the predicted score, specifically including: predicting the second predicted interaction action of the sample human body based on the morphological parameters and the initial position; the network parameter adjustment module 115 is used to adjust the network parameters of the interaction detection model based on the difference between the sample interaction action and the first predicted interaction action, and the difference between the sample score and the predicted score, specifically including: adjusting the network parameters of the interaction detection model based on the difference between the sample interaction action and the first predicted interaction action, the difference between the sample score and the predicted score, and the difference between the first predicted interaction action and the second predicted interaction action.
[0146] Among them, the above-mentioned morphological parameters include human body pose parameters and human body shape parameters, and the interaction action prediction module 112 is used to predict the second predicted interaction action of the sample human body based on the morphological parameters and the initial position, specifically including: encoding based on the human body pose parameters to obtain a pose encoding representation; splicing the pose encoding representation, the human body shape parameters, and the initial position to obtain a spliced feature representation; and predicting the second predicted interaction action based on the spliced feature representation.
[0147] Among them, the above sample video data includes several frames of sample images; the interaction action prediction module 112 is used to classify the sample human body features based on the action classification network of the interaction detection model to obtain the first predicted interaction action of the sample human body, specifically including: obtaining the sample action trajectory features of the sample human body based on the sample human body features of the same sample human body in each frame of the sample images; classifying based on the sample action trajectory features to obtain the first predicted interaction action of the sample human body.
[0148] Please refer to Figure 12 , Figure 12 FIG. is a schematic structural diagram of an embodiment of the interaction detection device provided by the present application. The interaction detection device 120 includes a feature extraction module 121 and an action classification module 122. The feature extraction module 121 is used to extract features of the human body in the to-be-detected image from the to-be-detected video data based on the feature extraction network of the interaction detection model to obtain the human body features of the human body in the to-be-detected image; the action classification module 122 is used to classify the human body features based on the action classification network of the interaction detection model to obtain the interaction action category of the human body; among them, the interaction detection model is obtained based on the above-mentioned training device 110 of the interaction detection model.
[0149] Please refer to Figure 13 , Figure 13 FIG. is a schematic structural diagram of an embodiment of the electronic device provided by the present application. The electronic device 130 includes a memory 131 and a processor 132 which are coupled to each other. The processor 132 is used to execute the program instructions stored in the memory 131 to implement the steps of any one of the above-mentioned interaction detection model training methods or interaction detection method embodiments. In a specific implementation scenario, the electronic device 130 may include, but is not limited to: a microcomputer, a server. In addition, the electronic device 130 may also include mobile devices such as a laptop computer, a tablet computer, etc., which are not limited herein.
[0150] Specifically, the processor 132 is used to control itself and the memory 131 to implement the steps of the training method of any of the above interaction detection models or the interaction detection method embodiments. The processor 132 may also be referred to as a CPU (Central Processing Unit). The processor 132 may be an integrated circuit chip with signal processing capabilities. The processor 132 may also be a general-purpose processor, a digital signal processor (DSP), an application specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc. Additionally, the processor 132 may be implemented jointly by integrated circuit chips.
[0151] Please refer to Figure 14 , Figure 14 which is a schematic structural diagram of an embodiment of the computer-readable storage medium provided by the present application. The computer-readable storage medium 140 of the embodiments of the present application stores program instructions 141, and when the program instructions 141 are executed, the methods provided by any embodiment of the training method of the interaction detection model or the interaction detection method of the present application and any non-conflicting combination are implemented. Among them, the program instructions 141 may form a program file and be stored in the above computer-readable storage medium 140 in the form of a software product, so that a computer device (which may be a personal computer, a server, or a network device, etc.) can execute all or part of the steps of the methods of various embodiments of the present application. The aforementioned computer-readable storage medium 140 includes: various media that can store program codes, such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disc, or a terminal device such as a computer, a server, a mobile phone, or a tablet.
[0152] If the technical solution of this application involves personal information, before the product applying the technical solution of this application processes personal information, it has clearly informed the personal information processing rules and obtained the individual's independent consent. If the technical solution of this application involves sensitive personal information, before the product applying the technical solution of this application processes sensitive personal information, it has obtained the individual's separate consent and at the same time meets the requirements of "express consent". For example, at personal information collection devices such as cameras, clear and prominent signs are set to inform that the personal information collection scope has been entered and personal information will be collected. If an individual voluntarily enters the collection scope, it is regarded as consenting to the collection of their personal information; or on the personal information processing device, when the personal information processing rules are informed by obvious signs / information, personal authorization is obtained through pop-up messages or by asking the individual to upload their personal information by themselves, etc.; among them, the personal information processing rules may include information such as the personal information processor, the purpose of personal information processing, the processing method, and the types of personal information processed.
[0153] The above are only the implementation manners of this application, and do not limit the patent scope of this application accordingly. Any equivalent structural or equivalent process transformation made by using the content of the specification and drawings of this application, or directly or indirectly applied in other related technical fields, shall be included in the patent protection scope of this application by the same token.
Claims
1. A training method for an interaction detection model, characterized in that, Including: Processing a sample image in sample video data by a feature extraction network based on an interaction detection model to obtain a sample human body feature of a sample human body in the sample image; wherein, the sample video data is labeled with a sample score indicating whether a sample object interacts with the sample human body, and a sample interaction action of the sample human body interacting with the sample object; Classifying the sample human body feature by an action classification network based on the interaction detection model to obtain a first predicted interaction action of the sample human body; Positioning based on the two-dimensional position of the sample object and the morphological parameters of the sample human body to obtain the three-dimensional position of the sample object; wherein, the two-dimensional position is obtained based on the sample human body feature; Predicting based on the morphological parameters and the three-dimensional position to obtain a predicted score of the sample object; wherein, the predicted score represents the possibility that the sample object has a human interaction; Adjusting network parameters of the interaction detection model based on the difference between the sample interaction action and the first predicted interaction action, and the difference between the sample score and the predicted score; 2. The method according to claim 1, wherein The obtaining of the two-dimensional position based on the sample human body feature includes: Jointly positioning the two-dimensional position of the sample object based on the sample human body feature and a sample object feature of the sample object; 3. The method according to claim 2, wherein The jointly positioning the two-dimensional position of the sample object based on the sample human body feature and the sample object feature of the sample object includes: Extracting a sample object feature of the sample object in the object region based on a sample image feature of the sample image and an object region detected in the sample image; Predicting based on the sample object feature and the sample human body feature to obtain an interaction score; wherein, the interaction score represents the interaction tightness between the sample human body to which the sample human body feature belongs and the sample object in the object region; Obtaining the two-dimensional position of the sample object based on the object region corresponding to the interaction score that meets a preset condition; 4. The method according to claim 3, characterized in that, The extracting a sample object feature of the sample object in the object region based on the sample image feature of the sample image and the object region detected in the sample image includes: Performing object detection on the sample image to obtain a plurality of candidate regions; Selecting the candidate region as the object region based on the first predicted interaction action; Extracting the sample object feature from the sample image feature based on the object region; 5. The method according to any one of claims 1 to 4, characterized in that The processing a sample image in sample video data by a feature extraction network based on an interaction detection model to obtain a sample human body feature of a sample human body in the sample image includes: Extracting a sample image feature of the sample image by the feature extraction network and performing human detection on the sample image to obtain a human region of the sample human body; Extracting the sample human body feature from the sample image feature based on the human region; 6. The method according to any one of claims 1 to 4, characterized in that, The positioning based on the two-dimensional position of the sample object and the morphological parameters of the sample human body to obtain the three-dimensional position of the sample object includes: Predict an initial position of the sample object in a three-dimensional space based on the two-dimensional position; Construct an action classification loss with a correction parameter of the initial position as an optimization target based on the initial position and the morphological parameters; Optimize to obtain the correction parameter based on the action classification loss, and perform position correction on the initial position based on the correction parameter to obtain the three-dimensional position.
7. The method according to claim 6, characterized in that, Before adjusting network parameters of the interaction detection model based on a difference between the sample interaction action and the first predicted interaction action, and a difference between the sample score and the predicted score, the method further includes: Perform prediction based on the morphological parameters and the initial position to obtain a second predicted interaction action of the sample human body; Adjusting the network parameters of the interaction detection model based on the difference between the sample interaction action and the first predicted interaction action, and the difference between the sample score and the predicted score, includes: Adjust the network parameters of the interaction detection model based on the difference between the sample interaction action and the first predicted interaction action, the difference between the sample score and the predicted score, and the difference between the first predicted interaction action and the second predicted interaction action.
8. The method according to claim 7, characterized in that The morphological parameters include human body pose parameters and human body shape parameters. Performing prediction based on the morphological parameters and the initial position to obtain a second predicted interaction action of the sample human body includes: Encode based on the human body pose parameters to obtain a pose encoding representation; Concatenate the pose encoding representation, the human body shape parameters, and the initial position to obtain a concatenated feature representation; Perform prediction based on the concatenated feature representation to obtain the second predicted interaction action.
9. The method according to any one of claims 1 to 4 or 7 to 8, characterized in that The sample video data includes a plurality of frames of the sample images. Classifying the sample human body features by an action classification network of the interaction detection model to obtain a first predicted interaction action of the sample human body includes: Obtain a sample action trajectory feature of the sample human body based on sample human body features of the same sample human body in each frame of the sample images; Classify based on the sample action trajectory feature to obtain a first predicted interaction action of the sample human body.
10. An interaction detection method, characterized in that, Includes: Extract features of a to-be-detected image in to-be-detected video data by a feature extraction network of an interaction detection model to obtain human body features of a human body in the to-be-detected image; Classify the human body features by an action classification network of the interaction detection model to obtain an interaction action category of the human body; Wherein, the interaction detection model is obtained based on the training method of the interaction detection model according to any one of claims 1 to 9.
11. A training device for an interaction detection model, characterized in that, Includes: A sample feature extraction module, configured to process a sample image in sample video data by a feature extraction network of an interaction detection model to obtain sample human body features of a sample human body in the sample image; wherein, the sample video data is labeled with a sample score indicating whether a sample object interacts with the sample human body, and a sample interaction action of the sample human body that interacts with the sample object; An interaction action prediction module, configured to classify the sample human body features based on the action classification network of the interaction detection model, so as to obtain a first predicted interaction action of the sample human body; A three-dimensional position positioning module, configured to perform positioning based on the two-dimensional position of the sample object and the morphological parameters of the sample human body, so as to obtain the three-dimensional position of the sample object; wherein, the two-dimensional position is obtained by positioning based on the sample human body features; An interaction score prediction module, configured to perform prediction based on the morphological parameters and the three-dimensional position, so as to obtain a predicted score of the sample object; wherein, the predicted score represents the possibility that there is human interaction with the sample object; A network parameter adjustment module, configured to adjust the network parameters of the interaction detection model based on the difference between the sample interaction action and the first predicted interaction action, and the difference between the sample score and the predicted score.
12. An interaction detection device, characterized in that, Comprising: A feature extraction module, configured to extract features of a to-be-detected image in to-be-detected video data based on a feature extraction network of an interaction detection model, so as to obtain human body features of a human body in the to-be-detected image; An action classification module, configured to classify the human body features based on the action classification network of the interaction detection model, so as to obtain an interaction action category of the human body; Wherein, the interaction detection model is obtained based on the training device of the interaction detection model according to claim 11.
13. An electronic device, characterized in that, Comprising a mutually coupled memory and a processor, the processor is configured to execute program instructions stored in the memory to implement the training method of the interaction detection model according to any one of claims 1 to 9, or to implement the interaction detection method according to claim 10.
14. A computer-readable storage medium having program instructions stored thereon, characterized in that, When the program instructions are executed by the processor, the training method of the interaction detection model according to any one of claims 1 to 9 is implemented, or the interaction detection method according to claim 10 is implemented.
Citation Information
Patent Citations
Figure interaction detection method based on deep learning
CN111914622A
Human-object interaction behavior detection method based on fine-grained multi-mode common representation
CN113468923A