A method of action recognition and related apparatus

CN116824686BActive Publication Date: 2026-10-09TENCENT TECHNOLOGY (SHENZHEN) CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202210278157.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-03-21
Publication Date
2026-10-09
Estimated Expiration
2042-03-21

AI Technical Summary

Technical Problem

[0004]然而,这种方法主要采用大模型如ResNet-50等编码图像特征来实现3D人体动作识别,计算量大,计算时间长,动作识别效率低,难以实现实时动作识别

Benefits of technology

[0022]As can be seen from the above technical solution, when capturing an image frame of the target object to be recognized, the two-dimensional joint position information of the target object in the image frame is first obtained. This two-dimensional joint position information is then used as input to the action recognition model, allowing the model to predict the three-dimensional joint position information based on the two-dimensional joint position information, thus achieving action recognition. Since the input to the action recognition model is two-dimensional joint position information, rather than an image or video, there is no need for the action recognition model to extract the two-dimensional joint position information from the large amount of information contained in the image or video through complex processing. This significantly reduces the computational load and time of the action recognition model, as well as the complexity of its network structure. To predict the three-dimensional joint position information using the action recognition model, the feature generation module of the action recognition model can generate a target feature vector based on the two-dimensional joint position information. Then, based on the target feature vector, the prediction module of the action recognition model can predict the rotation and displacement parameters of each joint of the target object. Finally, based on the rotation and displacement parameters, the kinematic analysis module of the action recognition model can perform kinematic analysis (e.g., forward kinematic analysis) to obtain the corresponding three-dimensional joint position information. As can be seen, this scheme significantly reduces the computational load and time of the action recognition model by directly using the two-dimensional joint position information required for action recognition as input, thereby improving the efficiency of action recognition and facilitating real-time action recognition. Simultaneously, the reduced computational load also greatly reduces the network structure complexity of the action recognition model, making it easier to implement action recognition based on lightweight networks and more suitable for real-time action recognition on mobile terminals.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116824686B_ABST
    Figure CN116824686B_ABST
Patent Text Reader

Abstract

The application discloses a motion recognition method and related device, which can be applied to cloud technology, artificial intelligence, intelligent transportation, auxiliary driving, vehicle-mounted scene and various scenes. When a target object is photographed to obtain a to-be-recognized image frame, two-dimensional joint position information of the target object in the to-be-recognized image frame is acquired, the two-dimensional joint position information is taken as input of a motion recognition model, according to the two-dimensional joint position information, a feature generation module is used for feature generation to obtain a target feature vector, then a prediction module is used for prediction according to the target feature vector, motion rotation parameters and motion displacement parameters of each joint of the target object are obtained, and then kinematics analysis is performed on the motion rotation parameters and the motion displacement parameters by using a kinematics analysis module to obtain three-dimensional joint position information of corresponding joints. Thus, the calculation amount and calculation time of the motion recognition model are greatly reduced, the motion recognition efficiency is improved, and real-time motion recognition is realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of computer vision, and in particular to an action recognition method and related apparatus. Background Technology

[0002] Artificial intelligence (AI) is a new science and technology that studies and develops theories, methods, technologies, and application systems to simulate, extend, and expand human intelligence. In recent years, with the development of AI, computer vision technology based on AI has also developed rapidly. Human motion recognition, as an important method within AI, has significant application prospects in many fields such as security, human-computer interaction, object understanding, object special effects, gaming, film and television production, and 3D modeling.

[0003] Currently, when performing action recognition, the body shape and action parameters of the human body parameterized model can be directly estimated based on the input image or video. For example, a convolutional network can be used to extract the image features of each frame in a video, and then a temporal network module can be used to capture the temporal information of the action sequence to obtain a more accurate action estimate.

[0004] However, this method mainly uses large models such as ResNet-50 to encode image features to achieve 3D human action recognition, which involves a large amount of computation, long computation time, low action recognition efficiency, and difficulty in achieving real-time action recognition. Summary of the Invention

[0005] To address the aforementioned technical issues, this application provides an action recognition method and related apparatus, which significantly reduces the computational load and time of the action recognition model, improves action recognition efficiency, and facilitates real-time action recognition. Simultaneously, the reduced computational load also greatly reduces the network structure complexity of the action recognition model, making it easier to implement action recognition based on lightweight networks and more suitable for real-time action recognition on mobile terminals.

[0006] The embodiments of this application disclose the following technical solutions:

[0007] On one hand, embodiments of this application provide an action recognition method, the method comprising:

[0008] When capturing an image frame to be identified by photographing the target object, the two-dimensional joint point position information of the target object in the image frame to be identified is obtained;

[0009] Based on the two-dimensional joint position information, feature generation is performed using the feature generation module of the action recognition model to obtain the target feature vector.

[0010] Based on the target feature vector, the prediction module of the action recognition model is used to predict and obtain the action rotation parameters and action displacement parameters of each joint of the target object.

[0011] Based on the motion rotation parameters and the motion displacement parameters, kinematic analysis is performed using the kinematic analysis module of the motion recognition model to obtain the three-dimensional joint position information of the corresponding joints.

[0012] On one hand, embodiments of this application provide an action recognition device, the device comprising an acquisition unit, a generation unit, a prediction unit, and an analysis unit:

[0013] The acquisition unit is used to acquire the two-dimensional joint point position information of the target object in the image frame to be identified when the target object is photographed to obtain the image frame to be identified;

[0014] The generation unit is used to generate features based on the two-dimensional joint position information using the feature generation module of the action recognition model to obtain the target feature vector.

[0015] The prediction unit is used to predict the motion rotation parameters and motion displacement parameters of each joint of the target object based on the target feature vector using the prediction module of the motion recognition model.

[0016] The analysis unit is used to perform kinematic analysis using the kinematic analysis module of the motion recognition model based on the motion rotation parameters and the motion displacement parameters, so as to obtain the three-dimensional joint position information of the corresponding joints.

[0017] On one hand, embodiments of this application provide an electronic device for action recognition, the electronic device including a processor and a memory:

[0018] The memory is used to store program code and transmit the program code to the processor;

[0019] The processor is used to execute the action recognition method described above according to the instructions in the program code.

[0020] In one aspect, embodiments of this application provide a computer-readable storage medium for storing program code for executing the action recognition method described above.

[0021] On one hand, embodiments of this application provide a computer program product, including a computer program that, when executed by a processor, implements the action recognition method described above.

[0022] As can be seen from the above technical solution, when capturing an image frame of the target object to be recognized, the two-dimensional joint position information of the target object in the image frame is first obtained. This two-dimensional joint position information is then used as input to the action recognition model, allowing the model to predict the three-dimensional joint position information based on the two-dimensional joint position information, thus achieving action recognition. Since the input to the action recognition model is two-dimensional joint position information, rather than an image or video, there is no need for the action recognition model to extract the two-dimensional joint position information from the large amount of information contained in the image or video through complex processing. This significantly reduces the computational load and time of the action recognition model, as well as the complexity of its network structure. To predict the three-dimensional joint position information using the action recognition model, the feature generation module of the action recognition model can generate a target feature vector based on the two-dimensional joint position information. Then, based on the target feature vector, the prediction module of the action recognition model can predict the rotation and displacement parameters of each joint of the target object. Finally, based on the rotation and displacement parameters, the kinematic analysis module of the action recognition model can perform kinematic analysis (e.g., forward kinematic analysis) to obtain the corresponding three-dimensional joint position information. As can be seen, this scheme significantly reduces the computational load and time of the action recognition model by directly using the two-dimensional joint position information required for action recognition as input, thereby improving the efficiency of action recognition and facilitating real-time action recognition. Simultaneously, the reduced computational load also greatly reduces the network structure complexity of the action recognition model, making it easier to implement action recognition based on lightweight networks and more suitable for real-time action recognition on mobile terminals. Attached Figure Description

[0023] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0024] Figure 1 This application provides an architecture diagram of an action recognition method for various scenarios.

[0025] Figure 2 A flowchart illustrating an action recognition method provided in this application embodiment;

[0026] Figure 3 A structural diagram of an action recognition model provided in an embodiment of this application;

[0027] Figure 4A schematic diagram illustrating an animation of a target object generated based on three-dimensional joint position information, provided in an embodiment of this application;

[0028] Figure 5 A schematic diagram illustrating the results of human motion estimation provided in an embodiment of this application;

[0029] Figure 6 This is a schematic diagram illustrating the effect of 3D character driving as provided in an embodiment of this application;

[0030] Figure 7 A flowchart illustrating a training method for an action recognition model provided in an embodiment of this application;

[0031] Figure 8 The network structures of Fusion Block and FC Block provided in the embodiments of this application;

[0032] Figure 9 A structural diagram of an action recognition device provided in an embodiment of this application;

[0033] Figure 10 A structural diagram of a terminal provided in an embodiment of this application;

[0034] Figure 11 This is a structural diagram of a server provided in an embodiment of this application. Detailed Implementation

[0035] The embodiments of this application will now be described with reference to the accompanying drawings.

[0036] Human motion recognition has a wide range of applications, including security, human-computer interaction, object understanding, object (human) special effects, games and entertainment, film / short video production, and 3D modeling. For example, human motion recognition can be used to locate the 3D joint positions of human body joints, thereby enabling 3D modeling to simulate human movement, which can then be used in film and television production. It can also be used for object understanding, including understanding the movement of objects (such as the human body), such as whether the human is swinging its arms, dancing, or performing other movements. Furthermore, it can be used to add human body special effects based on motion recognition. It can also enable human-computer interaction and games and entertainment; for example, motion-sensing games achieve game interaction through human motion recognition, and so on. These are just a few examples; further details are omitted here.

[0037] It should be noted that with the widespread use of mobile terminals, people's lives and work are basically inseparable from mobile terminals, and real-time action recognition on mobile terminals has gradually become a demand. However, the action recognition methods provided by related technologies mainly use large models such as ResNet-50 to encode image features to realize 3D human action recognition. This involves large computational load, long computation time, and low action recognition efficiency, making it difficult to achieve real-time action recognition and difficult to apply to mobile terminals.

[0038] To address the aforementioned technical problems, this application provides an action recognition method. This method directly uses the two-dimensional joint point position information required for action recognition as input to the action recognition model, significantly reducing the computational load and time of the action recognition model, improving action recognition efficiency, and facilitating real-time action recognition. Simultaneously, the reduced computational load also greatly reduces the network structure complexity of the action recognition model, making it easier to implement action recognition based on lightweight networks, and more suitable for real-time action recognition on mobile terminals.

[0039] like Figure 1 As shown, Figure 1 An application scenario architecture diagram of an action recognition method is shown. This application scenario may include a terminal 101. Terminal 101 can be a mobile terminal or a fixed terminal; this embodiment primarily describes it as a mobile terminal. Terminal 101 can be, for example, a mobile phone, computer, intelligent voice interaction device, smart home appliance, vehicle terminal, aircraft, etc., but is not limited to these. This embodiment can be applied to various scenarios, including but not limited to cloud technology, artificial intelligence, smart transportation, assisted driving, and vehicle scenarios.

[0040] Taking terminal 101 as a mobile terminal, and mobile terminal as a mobile phone as an example, if the target object is photographed by the mobile phone (the target object is such as a human body, animal, etc., and this application embodiment mainly takes the human body as an example), and special effects animation is added to the target object in real time, such as adding the special effect animation of armor to the photographed target object, then it is necessary to perform motion recognition on the target object, and then add armor matching the current motion of the target object to the target object according to the three-dimensional joint point position information.

[0041] Specifically, when capturing an image frame of the target object to be identified, the mobile phone can first obtain the two-dimensional joint point position information of the target object in the image frame. Joint points can represent movable joints on the target object, such as the right heel, left heel, right knee, left knee, right hip, left hip, right wrist, left wrist, right elbow, left elbow, right shoulder, left shoulder, head, etc. The two-dimensional joint point position information, used to represent the position of the target object's joint points in the image frame, can be pre-extracted.

[0042] The mobile phone uses two-dimensional joint position information as input to the action recognition model, so that the action recognition model can predict three-dimensional joint position information based on the two-dimensional joint position information, thereby achieving action recognition. Since the input of the action recognition model is two-dimensional joint position information, rather than images or videos, the action recognition model does not need to extract two-dimensional joint position information from the large amount of information contained in images or videos through complex processing. This greatly reduces the computational load and computation time of the action recognition model, and also greatly reduces the complexity of the network structure of the action recognition model.

[0043] By predicting 3D joint position information using a motion recognition model, the mobile phone can generate a target feature vector from the feature generation module of the motion recognition model based on the 2D joint position information. Then, based on the target feature vector, the prediction module of the motion recognition model is used to predict the rotation and displacement parameters of each joint of the target object. Finally, based on these parameters, kinematic analysis (e.g., forward kinematic analysis) is performed using the kinematic analysis module of the motion recognition model to obtain the corresponding 3D joint position information. This 3D joint position information reflects the actions performed by the target object in 3D space. Based on this information, armor matching the target object's current action can be added to the target object. The display result after adding armor matching the target object's current action can be seen in [reference needed]. Figure 1 As shown in Figure 102.

[0044] It is understood that the methods provided in this application may involve artificial intelligence (AI). AI is the theory, methods, technology, and application system that uses digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use that knowledge to obtain optimal results. AI technology is a comprehensive discipline involving a wide range of fields, encompassing both hardware and software technologies. Fundamental AI technologies generally include sensors, dedicated AI chips, cloud computing, distributed storage, big data processing technology, operating / interactive systems, and mechatronics. AI software technologies mainly include computer vision technology, speech processing technology, natural language processing technology, as well as machine learning / deep learning, autonomous driving, and intelligent transportation.

[0045] The methods provided in this application specifically relate to Computer Vision (CV) technology. Computer vision is a science that studies how to enable machines to "see." More specifically, it refers to machine vision that uses cameras and computers to replace human eyes for target recognition, tracking, and measurement, and further performs image processing to make the computer-processed images more suitable for human observation or transmission to instruments for detection. Computer vision technology typically includes image processing, image recognition, image semantic understanding, image retrieval, OCR, video processing, video semantic understanding, video content / behavior recognition, 3D object reconstruction, 3D technology, virtual reality, augmented reality, simultaneous localization and mapping (SLAM), autonomous driving, intelligent transportation, and other technologies. It also includes common biometric recognition technologies such as facial recognition and fingerprint recognition. This application mainly relates to behavior recognition and 3D technology.

[0046] The methods provided in this application embodiment can also relate to Machine Learning (ML), a multidisciplinary field involving probability theory, statistics, approximation theory, convex analysis, algorithm complexity theory, and other disciplines. It specifically studies how computers can simulate or implement human learning behavior to acquire new knowledge or skills and reorganize existing knowledge structures to continuously improve their performance. Machine learning is the core of artificial intelligence and the fundamental way to endow computers with intelligence; its applications span all areas of artificial intelligence. Machine learning and deep learning typically include techniques such as artificial neural networks, belief networks, reinforcement learning, transfer learning, inductive learning, and instructional learning. For example, an action recognition model can be obtained by training based on machine learning.

[0047] Next, taking the mobile terminal action recognition method as an example, the action recognition method provided in the embodiments of this application will be described in detail with reference to the accompanying drawings.

[0048] See Figure 2 , Figure 2 A flowchart of an action recognition method is shown, the method comprising:

[0049] S201. When capturing an image frame to be identified by photographing the target object, obtain the two-dimensional joint position information of the target object in the image frame to be identified.

[0050] If a target object (such as a human body, animal, etc., in this embodiment of the application, the target object is mainly a human body) is photographed by a mobile terminal, and the target object is recognized in real time, the mobile terminal can obtain the two-dimensional joint point position information of the target object in the image frame to be recognized.

[0051] S202. Based on the two-dimensional joint position information, feature generation is performed using the feature generation module of the action recognition model to obtain the target feature vector.

[0052] The mobile terminal inputs two-dimensional joint position information into the motion recognition model, enabling the model to predict three-dimensional joint position information based on the two-dimensional joint position information, thus achieving motion recognition. The motion recognition model may include a feature generation module, a prediction module, and a kinematic analysis module. The motion recognition model can be pre-trained, and the training method for the motion recognition model will be described in detail later.

[0053] After the mobile terminal inputs the two-dimensional joint point location information into the action recognition model, it can generate features using the feature generation module based on the two-dimensional joint point location information to obtain the target feature vector. The two-dimensional joint point location information can be represented by X.

[0054] It is understandable that when performing action recognition on an image frame to be recognized, if the image frame to be recognized is not the first image frame, for example, if the previous image frame and the image frame to be recognized are in the same image frame sequence, and the previous image frame is before and adjacent to the image frame to be recognized in the image frame sequence, since actions are usually continuous, the action performed by the target object in the image frame to be recognized is usually not much different from the action performed in the previous image frame. Therefore, in order to ensure the stability of the prediction results, relevant information from the previous image frame can be effectively used to stabilize the prediction results when performing action recognition on the image frame to be recognized.

[0055] In this case, the feature generation module may include a first feature extraction module and a first feature fusion module. Here, the implementation of S202 may be as follows: first, obtain the feature extraction result of the two-dimensional joint position information; then, based on the feature extraction result of the two-dimensional joint position information, use the first feature extraction module to generate the first feature vector corresponding to the image frame to be identified; then, obtain the relevant information of the previous image frame, such as the first feature vector corresponding to the previous image frame; and then, through the first feature fusion module, fuse the first feature vector corresponding to the image frame to be identified with the first feature vector corresponding to the previous image frame to obtain the target feature vector.

[0056] The above method can effectively utilize the relevant information of the previous image frame to enhance the feature vector of the image frame to be identified, thereby obtaining a target feature vector containing richer information. This allows for action recognition of the target object in the image frame to be identified based on the richer information, thus stabilizing the prediction results.

[0057] In some cases, the feature generation module may include a multi-layer feature extraction module. Taking an example where the multi-layer feature extraction module consists of a second feature extraction module followed by a first feature extraction module, the feature vectors extracted by the feature extraction module increasingly reflect the positions of the 3D joints from the input to the output direction of the action recognition module. To achieve more accurate feature fusion, the feature generation module also includes a multi-layer feature fusion module. When the multi-layer feature extraction module consists of a second feature extraction module followed by a first feature extraction module, the multi-layer feature fusion module consists of a second feature fusion module followed by a first feature fusion module, with the second feature extraction module and the second feature fusion module preceding the first feature extraction module. In this case, the action recognition model can refer to... Figure 3 As shown. In Figure 3 In the action recognition model, a feature generation module 301, a prediction module 302, and a kinematic analysis module 303 may be included. The feature generation module 301 includes a second feature extraction module 3011, a second feature fusion module 3012, a first feature extraction module 3013, and a first feature fusion module 3014.

[0058] In this scenario, the method for obtaining the feature extraction result of the two-dimensional joint position information can be to extract features from the two-dimensional joint position information using a second feature extraction module to obtain a second feature vector corresponding to the image frame to be identified, and then determine the second feature vector corresponding to the image frame to be identified as the feature extraction result. Correspondingly, based on the feature extraction result of the two-dimensional joint position information, the method for generating the first feature vector corresponding to the image frame to be identified using the first feature extraction module can be to fuse the second feature vector corresponding to the image frame to be identified with the second feature vector corresponding to the previous image frame using a second feature fusion module to obtain a fused feature vector; and then encode the fused feature vector using the first feature extraction module to obtain the first feature vector.

[0059] It should be noted that the second feature extraction module 3011 is located before the first feature extraction module 3013. Compared to the first feature extraction module 3013, the second feature extraction module 3011 extracts early feature vectors. Therefore, the second feature extraction module 3011 can be called the early feature extraction module, denoted as Early Stage; the first feature extraction module 3013 can be called the late feature extraction module, denoted as Late Stage; correspondingly, the second feature fusion module 3012 can be called the early feature fusion module, denoted as Early Fusion; the first feature fusion module 3014 can be called the late feature fusion module, denoted as Late Fusion. The second feature extraction module 3011 extracts features from the two-dimensional joint position information, and the resulting second feature vector can be called the early feature vector of the image frame to be identified, denoted as... Where t is the frame number, indicating that the image frame to be identified is the t-th image frame; the second feature vector corresponding to the previous image frame can be called the early feature vector of the previous image frame, denoted as: Where t-1 is the frame number, indicating that the previous image frame is the (t-1)th image frame; the fused feature vector output by the second feature fusion module 3012 can be called the fused early feature vector, and can be expressed as: The first feature vector output by the first feature extraction module 3013 can be called the later feature vector of the image frame to be identified, and can be represented as follows: The first feature vector corresponding to the previous image frame can be called the later feature vector of the previous image frame, denoted as: The target feature vector output by the first feature fusion module 3014 can be called the fused post-feature vector, denoted as:

[0060] Compared with related technologies, the temporal network modules used in related technologies are designed for a pre-recorded video. The preceding and following frames of the image frame to be identified for action recognition are visible, which is not suitable for real-time action recognition. However, the embodiments of this application only use the preceding image frame, which is more suitable for real-time action recognition and more suitable for mobile terminals.

[0061] S203. Based on the target feature vector, the prediction module of the action recognition model is used to predict and obtain the action rotation parameters and action displacement parameters of each joint of the target object.

[0062] After extracting the target feature vector, the prediction module is used to predict the motion rotation and displacement parameters of each joint of the target object based on the target feature vector. In one possible implementation, since the motion rotation and displacement parameters are different parameters, their emphasis on the target feature vector may be slightly different. Therefore, the motion rotation and displacement parameters can be predicted separately based on different branches of the prediction module. For details, see [link to details]. Figure 3 As shown, the prediction module 302 may include a rotation parameter prediction module 3021 and a displacement prediction module 3022. The rotation parameter prediction module 3021 can be represented as a Quat Head, which can predict the rotation parameters Q of each joint of the target object based on the target feature vector. The displacement prediction module 3022 can be represented as a TransHead, which can predict the displacement parameters T of each joint of the target object based on the target feature vector.

[0063] S204. Based on the motion rotation parameters and the motion displacement parameters, perform kinematic analysis using the kinematic analysis module of the motion recognition model to obtain the three-dimensional joint position information of the corresponding joints.

[0064] See Figure 3 As shown, the mobile terminal can perform kinematic analysis using the kinematic analysis module 303 based on the motion rotation parameters and motion displacement parameters to obtain the three-dimensional joint position information of the corresponding joints. The kinematic analysis module can be represented as an FK layer. Kinematic analysis typically includes forward kinematic analysis and backward kinematic analysis. In this embodiment, the three-dimensional joint position information is mainly calculated through forward kinematic analysis. The three-dimensional joint position information can be represented by J.

[0065] In one possible implementation, the action recognition model may also include a discriminator network, such as... Figure 3 As shown, the discriminator network may include a first discriminator network 304. The first discriminator network 304 can determine the authenticity of an action based on the three-dimensional joint position information. The first discriminator network 304 can use D... J This indicates that, for example, the discriminator network may also include a second discriminator network 305, which can determine the authenticity of an action based on the action rotation parameters. The second discriminator network 305 can use D... Q This indicates that the discriminator network is mainly used during the training of the action recognition model. When using the action recognition model for action recognition, the discriminator network is not needed. Therefore, the discriminator network will be introduced in detail during the training of the action recognition model.

[0066] After obtaining the 3D joint point position information, the mobile terminal can generate an animation of the target object based on this information. This animation can be an animation driving the corresponding 3D model of the target object, or an animation adding human body effects to the target object. For example, if the human body effect is armor, when the mobile terminal displays the target object as... Figure 4 As shown in Figure (a), the diagram illustrating the addition of armor as a human body effect based on the three-dimensional joint position information obtained from motion recognition after motion recognition according to the embodiments of this application can be found in [reference needed]. Figure 4 As shown in Figure (b).

[0067] As can be seen from the above technical solution, when capturing an image frame of the target object to be recognized, the two-dimensional joint position information of the target object in the image frame is first obtained. This two-dimensional joint position information is then used as input to the action recognition model, allowing the model to predict the three-dimensional joint position information based on the two-dimensional joint position information, thus achieving action recognition. Since the input to the action recognition model is two-dimensional joint position information, rather than an image or video, there is no need for the action recognition model to extract the two-dimensional joint position information from the large amount of information contained in the image or video through complex processing. This significantly reduces the computational load and time of the action recognition model, as well as the complexity of its network structure. To predict the three-dimensional joint position information using the action recognition model, the feature generation module of the action recognition model can generate a target feature vector based on the two-dimensional joint position information. Then, based on the target feature vector, the prediction module of the action recognition model can predict the rotation and displacement parameters of each joint of the target object. Finally, based on the rotation and displacement parameters, the kinematic analysis module of the action recognition model can perform kinematic analysis (e.g., forward kinematic analysis) to obtain the corresponding three-dimensional joint position information. As can be seen, this scheme significantly reduces the computational load and time of the action recognition model by directly using the two-dimensional joint position information required for action recognition as input, thereby improving the efficiency of action recognition and facilitating real-time action recognition. Simultaneously, the reduced computational load also greatly reduces the network structure complexity of the action recognition model, making it easier to implement action recognition based on lightweight networks and more suitable for real-time action recognition on mobile terminals.

[0068] This application embodiment performs quantitative and qualitative evaluations on the provided action recognition method. The quantitative evaluation results are shown in Table 1.

[0069] Table 1

[0070]

[0071] In Table 1, MPJPE (Mean Per Joint Position Error), PVE (Per Vertex Error), and PA-MPJPE (Procrustes Aligned Mean Per Joint Position Error) are all quantitative evaluation metrics. PA-MPJPE is calculated by first performing a rigid transformation (e.g., translation, rotation, and scaling) on ​​the predicted output to align it to the corresponding ground truth value. As can be seen from Table 1, the method provided in this application embodiment has lower values ​​than related technologies in all evaluation metrics. Therefore, it can be seen that the action recognition method provided in this application embodiment has a significantly reduced computational load compared to the action recognition methods provided by related technologies, and shows a significant performance improvement on the synthetic dataset MOCAP and the real-world dataset 3DPW compiled for business evaluation.

[0072] The qualitative evaluation results can be found in Figure 5 and Figure 6 As shown, where, Figure 5 This is a schematic diagram based on the results of human motion estimation, tested through actual video recordings. Figure 5 It can be seen that the method provided in the embodiments of this application can accurately estimate human movements in videos. Figure 5 Image (a) is a video taken of a human body. Figure 5 Figure (b) shows the results of human motion estimation; additionally, videos and 3D characters can be randomly selected to test the driving effect. Figure 6 It can be seen that the motion recognition method provided in this application embodiment can be applied to human body effects such as 3D character driving. Figure 6 (a) is a video shot of a human body. Figure 6 The middle (b) image shows the human body effects driven by a 3D character.

[0073] Next, we will describe in detail the training method for the action recognition model. To train the action recognition model, we first construct its corresponding initial network model. The initial network model includes an initial feature generation module, an initial prediction module, and an initial kinematics analysis module. See [link to documentation]. Figure 7 As shown, the method includes:

[0074] S701. Obtain the historical position information of two-dimensional joints of historical objects in historical image frames.

[0075] In this embodiment, the historical position information of two-dimensional joints of historical objects in historical image frames can be used as training samples to train an action recognition model. Typically, the historical position information of two-dimensional joints used as training samples consists of the historical position information of two-dimensional joints of objects in multiple historical image frames, and multiple historical position information of two-dimensional joints can be input into the initial network model in batches.

[0076] S702. Based on the historical position information of the two-dimensional joints, feature generation is performed using the feature generation initial module to obtain the target historical feature vector.

[0077] S703. Based on the target historical feature vector, the prediction initialization module is used to make predictions to obtain the motion history rotation parameters and motion history displacement parameters of each joint of the historical object.

[0078] S704. Based on the motion history rotation parameters and the motion history displacement parameters, perform kinematic analysis using the kinematic analysis initialization module to obtain the three-dimensional joint point historical position information of the corresponding joint points.

[0079] It should be noted that the implementation of S701-S704 is similar to that of S201-S204 in the action recognition model, and will not be elaborated here.

[0080] S705. Construct a target loss function based on the historical position information of the three-dimensional joints.

[0081] S706. Optimize and adjust the model parameters of the initial network model according to the target loss function to obtain the action recognition model.

[0082] After obtaining the historical position information of three-dimensional joints through the initial module of kinematic analysis, in order to train the initial network model, a target loss function can be constructed based on the historical position information of three-dimensional joints. Then, the initial network model is optimized and adjusted according to the target loss function until the target loss function meets the preset conditions, and the training stops to obtain the action recognition model.

[0083] It should be noted that, in the embodiments of this application, the training of the action recognition model can be performed on a terminal or on a server, and this application does not limit this. The server can be a standalone server, an integrated server, or a cloud server, etc.

[0084] In one possible implementation, the action recognition model may further include a first discriminator network, thus the initial network model may also include a first discriminator network for determining the authenticity of the predicted 3D joint historical position information. Furthermore, the action recognition model may also include a second discriminator network, thus the initial network model may also include a second discriminator network for determining the authenticity of the action historical rotation parameters.

[0085] The embodiments of this application use a discriminator network, such as a first discriminator network and a second discriminator network, to determine the authenticity of an action, thereby further enhancing the smoothness of the action.

[0086] The action recognition model provided in this application embodiment is referred to Figure 3 As shown in Table 2, the specific network structure of the action recognition model is as follows:

[0087] Table 2

[0088]

[0089]

[0090] Where B is the number of samples in the video frame sequence, T is the number of frames, FC Block represents a fully connected module, FusionBlock represents a fusion module, BN (Batch Normalization) represents batch normalization (i.e., standardization), ReLU is the activation function, and GRU is a gated recurrent unit.

[0091] Figure 8 The network structures of Fusion Block and FC Block are shown. 801 is the network structure of Fusion Block. When Fusion Block is used as the second feature fusion module, the input of Fusion Block can be, for example, the second feature vector corresponding to the previous image frame and the second feature vector of the image frame to be identified, and the output can be the fused feature vector. 802 is the network structure of FC Block.

[0092] It should be noted that the specific design of the network structure defined in Table 2 is only an example. The network structure can be increased or decreased according to computing resources. For example, the design of FC Block and Fusion Block can appropriately increase the number of fully connected layers, or increase or decrease the number of output channels, and so on.

[0093] It should be noted that the target loss function is crucial for training the action recognition model. Whether the target loss function comprehensively constrains the action recognition results (such as 3D joint position information) will affect the accuracy of the trained action recognition model. Therefore, this application employs multiple loss functions to constrain the accuracy, stability, and rationality of the generated actions, so that more realistic and accurate action recognition results can be obtained using the action recognition model.

[0094] Based on this, the implementation of S705 can be as follows: an action recognition loss function, an action variation loss function, and an adversarial loss function can be constructed based on the historical position information of the three-dimensional joints. The action recognition loss function is used to measure the accuracy of action recognition, the action variation loss function is used to measure the stability of action variation between different image frames, and the adversarial loss function is used to measure the rationality of action recognition. A target loss function is constructed based on at least one of the action recognition loss function, the action variation loss function, and the adversarial loss function.

[0095] The target loss function can be expressed as:

[0096] L all =L action +L quat-velo +L gan

[0097] Among them, L all Let L be the target loss function. action Let L be the loss function for action recognition. quat-velo Let L be the loss function for action variation. gan To counteract the loss function.

[0098] In one possible implementation, based on Figure 2 In the corresponding embodiment, the action recognition model may also include a first discriminator network. Therefore, the initial network model may also include a first discriminator network. In this case, the adversarial loss function can be constructed based on the historical position information of the three-dimensional joints by using the first discriminator network to discriminate the historical position information of the three-dimensional joints to obtain a first discrimination result, and then constructing the adversarial loss function based on the historical position information of the three-dimensional joints and the first discrimination result.

[0099] To further enhance the smoothness of the motion, the motion recognition model may also include a second discriminator network. Therefore, the initial network model also includes a second discriminator network. Based on this, the embodiments of this application may also use the second discriminator network to discriminate the historical rotation parameters of the motion to obtain a second discrimination result. At this time, the adversarial loss function can be constructed based on the historical position information of the three-dimensional joints and the first discrimination result.

[0100] To assess the validity of action recognition, generative adversarial methods are employed to determine the authenticity of actions. The adversarial loss function can then be expressed as:

[0101]

[0102]

[0103] L gan =w gan (L gan-Q +L gan-J )

[0104] in, For L2 loss function, and These are the second discrimination result of the second discriminator network and the first discrimination result of the first discriminator network, respectively. pred Q represents the predicted motion history rotation parameters. gt J represents the true value of the rotation parameter of the motion. pred J represents the predicted historical position information of the three-dimensional joints. gt L represents the truth value of location information. gan L represents the adversarial loss function. gan-Q Let L represent the loss function of the second discriminator network. gan-J Let w represent the loss function of the first discriminator network. gan This represents the weight, which can be set according to actual needs.

[0105] In one possible implementation, the action recognition loss function can be constructed based on the historical position information of 3D joints by determining a first loss function based on the historical position information and the true value of the position information; determining a second loss function based on the historical rotation parameters of the action corresponding to the historical position information of 3D joints and the true value of the rotation parameters; determining a third loss function based on the historical displacement parameters of the action corresponding to the historical position information of 3D joints and the true value of the displacement parameters; and then weighted summing the first, second, and third loss functions to obtain the action recognition loss function.

[0106] To ensure the accuracy of action recognition, this application embodiment constrains the historical rotation parameters, historical displacement parameters, and historical position information of 3D joints to be close to their respective true values. Specifically, the action recognition loss function can be calculated using the following formula:

[0107]

[0108]

[0109]

[0110] L action =w quat L quat +w trans L trans +w joint L joint

[0111] Among them, L joint Let L be the first loss function. quat For the second loss function, L trans For the third loss function, Let Q be the L1 loss function, and R, O be the three representations of the motion history rotation parameters. The initial kinematic analysis module typically outputs a 6-dimensional representation O of the motion history rotation parameters, which can then be converted into a quaternion representation Q and a rotation matrix representation R. Here, Q... pred R pred O pred Q represents the rotation parameters of the action history obtained from different prediction methods. gt R gt O gt T represents the true values ​​of the motion rotation parameters in different representations. pred T represents the predicted motion history displacement parameters. gt J represents the true value of the displacement parameter. pred J represents the predicted historical position information of the three-dimensional joints. gt This represents the truth value of the location information. quat w trans w joint These represent the weights corresponding to each loss function, which can be set according to actual needs.

[0112] In one possible implementation, the motion change loss function can be constructed based on the historical position information of the three-dimensional joints by determining a fourth loss function based on the difference between the historical position information of the three-dimensional joints corresponding to any two historical image frames and the true value of the first difference; determining a fifth loss function based on the difference between the historical rotation parameters of the motion corresponding to the two historical image frames and the true value of the second difference; determining a sixth loss function based on the difference between the historical displacement parameters of the motion corresponding to the two historical image frames and the true value of the third difference; and then weighted summing the fourth, fifth, and sixth loss functions to obtain the motion change loss function.

[0113] Regarding motion stability, the embodiments of this application respectively constrain the changes in motion history rotation parameters, motion history displacement parameters, and changes in the historical position information of three-dimensional joints to be close to their respective true values. That is, the motion change loss function can be calculated by the following formula:

[0114]

[0115]

[0116]

[0117] L velo =w velo (L quat-velo +L trans-velo +L joint-velo )

[0118] Among them, L joint-velo For the fourth loss function, L quat-velo For the fifth loss function, L trans-velo For the sixth loss function, w velo The weights for each loss function can be set according to actual needs. This is the L1 loss function. These represent the differences between the motion history rotation parameters corresponding to any two historical image frames under different representation methods. These are the true values ​​of the second difference under different representations. To predict the difference between motion history displacement parameters corresponding to any two historical image frames, The true value of the third difference. To predict the difference between the historical position information of 3D key points corresponding to any two historical image frames, This is the true value of the first difference.

[0119] In one possible implementation, the weights mentioned above can be set to w. quat =10.0, w joint =20.0, w trans =15.0, w velo =5.0, w gan =0.1. Each weight can be adjusted according to actual needs, and the values ​​of each weight are not limited in this embodiment.

[0120] It should be noted that, based on the implementation methods provided in the above aspects, this application can be further combined to provide more implementation methods.

[0121] based on Figure 2 Corresponding to the action recognition method provided in the embodiments, this application also provides an action recognition device 900. See also Figure 9 The action recognition device 900 includes an acquisition unit 901, a generation unit 902, a prediction unit 903, and an analysis unit 904.

[0122] The acquisition unit 901 is used to acquire the two-dimensional joint position information of the target object in the image frame to be identified when the target object is photographed to obtain the image frame to be identified.

[0123] The generation unit 902 is used to generate a target feature vector by using the feature generation module of the action recognition model based on the two-dimensional joint position information.

[0124] The prediction unit 903 is used to predict the motion rotation parameters and motion displacement parameters of each joint of the target object based on the target feature vector using the prediction module of the motion recognition model.

[0125] The analysis unit 904 is used to perform kinematic analysis using the kinematic analysis module of the motion recognition model based on the motion rotation parameters and the motion displacement parameters, so as to obtain the three-dimensional joint position information of the corresponding joints.

[0126] In one possible implementation, the feature generation module includes a first feature extraction module and a first feature fusion module, and the generation unit 902 is specifically used for:

[0127] Obtain the feature extraction results of the two-dimensional joint position information;

[0128] Based on the feature extraction results of the two-dimensional joint point position information, the first feature extraction module is used to generate the first feature vector corresponding to the image frame to be identified.

[0129] The first feature fusion module fuses the first feature vector corresponding to the image frame to be identified with the first feature vector corresponding to the previous image frame to obtain the target feature vector. The previous image frame and the image frame to be identified are located in the same image frame sequence. In the image frame sequence, the previous image frame is before the image frame to be identified and is adjacent to the image frame to be identified.

[0130] In one possible implementation, the feature generation module further includes a second feature extraction module and a second feature fusion module. In the action recognition model, the second feature extraction module and the second feature fusion module are located before the first feature extraction module. The generation unit 902 is specifically used for:

[0131] The second feature extraction module extracts features from the two-dimensional joint position information to obtain the second feature vector corresponding to the image frame to be identified.

[0132] The second feature vector corresponding to the image frame to be identified is determined as the feature extraction result;

[0133] The second feature fusion module fuses the second feature vector corresponding to the image frame to be identified with the second feature vector corresponding to the previous image frame to obtain a fused feature vector.

[0134] The first feature vector is obtained by encoding the fused feature vector through the first feature extraction module.

[0135] In one possible implementation, the generating unit 902 is further configured to:

[0136] An animation of the target object is generated based on the three-dimensional joint position information.

[0137] In one possible implementation, the initial network model corresponding to the action recognition model includes a feature generation initial module, a prediction initial module, and a kinematic analysis initial module. The device further includes a training unit, which is used for:

[0138] Obtain the historical position information of two-dimensional joints of historical objects in historical image frames;

[0139] Based on the historical position information of the two-dimensional joints, feature generation is performed using the feature generation initial module to obtain the target historical feature vector.

[0140] Based on the target historical feature vector, the prediction initialization module is used to predict and obtain the motion history rotation parameters and motion history displacement parameters of each joint of the historical object.

[0141] Based on the motion history rotation parameters and the motion history displacement parameters, kinematic analysis is performed using the kinematic analysis initialization module to obtain the three-dimensional joint point historical position information of the corresponding joint points.

[0142] Construct a target loss function based on the historical position information of the three-dimensional joints;

[0143] The model parameters of the initial network model are optimized and adjusted according to the target loss function to obtain the action recognition model.

[0144] In one possible implementation, the training unit is specifically used for:

[0145] Based on the historical position information of the three-dimensional joints, a motion recognition loss function, a motion change loss function, and an adversarial loss function are constructed respectively. The motion recognition loss function is used to measure the accuracy of motion recognition, the motion change loss function is used to measure the stability of motion changes between different image frames, and the adversarial loss function is used to measure the rationality of motion recognition.

[0146] The target loss function is constructed based on at least one of the action recognition loss function, the action change loss function, and the adversarial loss function.

[0147] In one possible implementation, the initial network model further includes a first discriminator network, and the training unit is specifically used for:

[0148] The first discriminator network is used to discriminate the historical position information of the three-dimensional joints to obtain a first discrimination result;

[0149] The adversarial loss function is constructed based on the historical location information of the three-dimensional key points and the first discrimination result.

[0150] In one possible implementation, the initial network model further includes a second discriminator network, and the training unit is further used for;

[0151] The second discriminator network is used to discriminate the rotation parameters of the action history to obtain a second discrimination result;

[0152] The training unit is specifically used for:

[0153] The adversarial loss function is constructed based on the historical location information of the three-dimensional key points, the first discrimination result, and the second discrimination result.

[0154] In one possible implementation, the training unit is specifically used for:

[0155] Based on the historical position information and true values ​​of the position information of the three-dimensional joints, a first loss function is determined;

[0156] The second loss function is determined based on the motion history rotation parameters and the true values ​​of the motion rotation parameters corresponding to the historical position information of the three-dimensional joints.

[0157] The third loss function is determined based on the motion history displacement parameters and the true values ​​of the displacement parameters corresponding to the historical position information of the three-dimensional joints.

[0158] The action recognition loss function is obtained by weighted summation of the first loss function, the second loss function, and the third loss function.

[0159] In one possible implementation, the training unit is specifically used for:

[0160] The fourth loss function is determined based on the difference between the historical position information of the three-dimensional key points corresponding to two adjacent historical image frames and the true value of the first difference.

[0161] The fifth loss function is determined based on the difference between the motion history rotation parameters corresponding to two adjacent historical image frames and the true value of the second difference.

[0162] The sixth loss function is determined based on the difference between the motion history displacement parameters corresponding to two adjacent historical image frames and the true value of the third difference.

[0163] The action change loss function is obtained by weighted summation of the fourth loss function, the fifth loss function, and the sixth loss function.

[0164] This application also provides an electronic device for action recognition. This electronic device can be a terminal, with a smartphone as an example of a mobile terminal:

[0165] Figure 10 The diagram shown is a block diagram of a portion of the structure of a smartphone provided in an embodiment of this application. (Reference) Figure 10 The smartphone includes components such as: a radio frequency (RF) circuit 1010, a memory 1020, an input unit 1030, a display unit 1040, a sensor 1050, an audio circuit 1060, a Wi-Fi module 1070, a processor 1080, and a power supply 1090. The input unit 1030 may include a touch panel 1031 and other input devices 1032; the display unit 1040 may include a display panel 1041; and the audio circuit 1060 may include a speaker 1061 and a microphone 1062. It is understood that... Figure 10 The smartphone structure shown does not constitute a limitation on smartphones and may include more or fewer components than shown, or combine certain components, or have different component arrangements.

[0166] The memory 1020 can be used to store software programs and modules. The processor 1080 executes various functions and data processing of the smartphone by running the software programs and modules stored in the memory 1020. The memory 1020 may mainly include a program storage area and a data storage area. The program storage area may store the operating system, applications required for at least one function (such as sound playback function, image playback function, etc.), etc.; the data storage area may store data created according to the use of the smartphone (such as audio data, phonebook, etc.). In addition, the memory 1020 may include high-speed random access memory, and may also include non-volatile memory, such as at least one disk storage device, flash memory device, or other volatile solid-state storage device.

[0167] The processor 1080 is the control center of the smartphone, connecting various parts of the smartphone via various interfaces and lines. It performs various functions and processes data by running or executing software programs and / or modules stored in the memory 1020 and by accessing data stored in the memory 1020. Optionally, the processor 1080 may include one or more processing units; preferably, the processor 1080 may integrate an application processor and a modem processor, wherein the application processor mainly handles the operating system, user interface, and applications, and the modem processor mainly handles wireless communication. It is understood that the modem processor may not be integrated into the processor 1080.

[0168] In this embodiment, the processor 1080 in the smartphone can perform the following steps:

[0169] When capturing an image frame to be identified by photographing the target object, the two-dimensional joint point position information of the target object in the image frame to be identified is obtained;

[0170] Based on the two-dimensional joint position information, feature generation is performed using the feature generation module of the action recognition model to obtain the target feature vector.

[0171] Based on the target feature vector, the prediction module of the action recognition model is used to predict and obtain the action rotation parameters and action displacement parameters of each joint of the target object.

[0172] Based on the motion rotation parameters and the motion displacement parameters, kinematic analysis is performed using the kinematic analysis module of the motion recognition model to obtain the three-dimensional joint position information of the corresponding joints.

[0173] This application also provides a server; please refer to [link / reference]. Figure 11 As shown, Figure 11This is a structural diagram of a server 1100 provided in an embodiment of this application. The server 1100 can vary significantly due to different configurations or performance. It may include one or more central processing units (CPUs) 1122 (e.g., one or more processors) and a memory 1132, and one or more storage media 1130 (e.g., one or more mass storage devices) for storing application programs 1142 or data 1144. The memory 1132 and storage media 1130 may be temporary or persistent storage. The program stored in the storage media 1130 may include one or more modules (not shown in the diagram), each module including a series of instruction operations on the server. Furthermore, the CPU 1122 may be configured to communicate with the storage media 1130 and execute the series of instruction operations in the storage media 1130 on the server 1100.

[0174] Server 1100 may also include one or more power supplies 1126, one or more wired or wireless network interfaces 1150, one or more input / output interfaces 1158, and / or one or more operating systems 1141, such as Windows Server. TM Mac OS X TM Unix TM Linux TM FreeBSD TM etc.

[0175] In this embodiment, the steps that need to be executed by the central processing unit 1122 in the server 1100 can be based on Figure 11 The server architecture shown is implemented.

[0176] According to one aspect of this application, a computer-readable storage medium is provided for storing program code for performing the action recognition methods described in the foregoing embodiments.

[0177] According to one aspect of this application, a computer program product or computer program is provided, comprising computer instructions stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium and executes the computer instructions, causing the computer device to perform the methods provided in various optional implementations of the above embodiments.

[0178] The descriptions of the processes or structures corresponding to the above figures each have their own emphasis. For parts of a process or structure that are not described in detail, please refer to the relevant descriptions of other processes or structures.

[0179] The terms “first,” “second,” “third,” “fourth,” etc. (if present) in the specification and accompanying drawings of this application are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of this application described herein can be implemented, for example, in orders other than those illustrated or described herein. Furthermore, the terms “comprising” and “having,” and any variations thereof, are intended to cover a non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.

[0180] In the several embodiments provided in this application, it should be understood that the disclosed systems, apparatuses, and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be an indirect coupling or communication connection between apparatuses or units through some interfaces, and may be electrical, mechanical, or other forms.

[0181] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.

[0182] Furthermore, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.

[0183] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.

[0184] The above-described embodiments are only used to illustrate the technical solutions of this application, and are not intended to limit them. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of this application.

Claims

1. An action recognition method, characterized in that, The method includes: When capturing an image frame to be identified by photographing the target object, the two-dimensional joint point position information of the target object in the image frame to be identified is obtained; The two-dimensional joint position information is used as the input to the action recognition model, and the feature generation module of the action recognition model is used to generate features to obtain the target feature vector. Based on the target feature vector, the prediction module of the action recognition model is used to predict the motion rotation parameters and motion displacement parameters of each joint of the target object. The prediction module includes a rotation parameter prediction module and a displacement prediction module. The rotation parameter prediction module is used to predict the motion rotation parameters of each joint of the target object based on the target feature vector. The displacement prediction module is used to predict the motion displacement parameters of each joint of the target object based on the target feature vector. Based on the motion rotation parameters and the motion displacement parameters, kinematic analysis is performed using the kinematic analysis module of the motion recognition model to obtain the three-dimensional joint position information of the corresponding joints. The feature generation module includes a first feature extraction module and a first feature fusion module. The two-dimensional joint position information is used as input to the action recognition model, and the feature generation module of the action recognition model is used to generate features to obtain a target feature vector, including: Obtain the feature extraction results of the two-dimensional joint position information; Based on the feature extraction results of the two-dimensional joint point position information, the first feature extraction module is used to generate the first feature vector corresponding to the image frame to be identified. The first feature fusion module fuses the first feature vector corresponding to the image frame to be identified with the first feature vector corresponding to the previous image frame to obtain the target feature vector. The previous image frame and the image frame to be identified are located in the same image frame sequence. In the image frame sequence, the previous image frame is before the image frame to be identified and is adjacent to the image frame to be identified.

2. The method according to claim 1, characterized in that, The feature generation module further includes a second feature extraction module and a second feature fusion module. In the action recognition model, the second feature extraction module and the second feature fusion module are located before the first feature extraction module. The feature extraction result for obtaining the two-dimensional joint position information includes: The second feature extraction module extracts features from the two-dimensional joint position information to obtain the second feature vector corresponding to the image frame to be identified. The second feature vector corresponding to the image frame to be identified is determined as the feature extraction result; The step of generating a first feature vector corresponding to the image frame to be identified using the first feature extraction module based on the feature extraction results of the two-dimensional joint position information includes: The second feature fusion module fuses the second feature vector corresponding to the image frame to be identified with the second feature vector corresponding to the previous image frame to obtain a fused feature vector. The first feature vector is obtained by encoding the fused feature vector through the first feature extraction module.

3. The method according to any one of claims 1-2, characterized in that, The method further includes: An animation of the target object is generated based on the three-dimensional joint position information.

4. The method according to any one of claims 1-2, characterized in that, The initial network model corresponding to the action recognition model includes a feature generation initial module, a prediction initial module, and a kinematic analysis initial module. The method further includes: Obtain the historical position information of two-dimensional joints of historical objects in historical image frames; Based on the historical position information of the two-dimensional joints, feature generation is performed using the feature generation initial module to obtain the target historical feature vector. Based on the target historical feature vector, the prediction initialization module is used to predict and obtain the motion history rotation parameters and motion history displacement parameters of each joint of the historical object. Based on the motion history rotation parameters and the motion history displacement parameters, kinematic analysis is performed using the kinematic analysis initialization module to obtain the three-dimensional joint point historical position information of the corresponding joint points. Construct a target loss function based on the historical position information of the three-dimensional joints; The model parameters of the initial network model are optimized and adjusted according to the target loss function to obtain the action recognition model.

5. The method according to claim 4, characterized in that, The step of constructing the target loss function based on the historical position information of the three-dimensional joints includes: Based on the historical position information of the three-dimensional joints, a motion recognition loss function, a motion change loss function, and an adversarial loss function are constructed respectively. The motion recognition loss function is used to measure the accuracy of motion recognition, the motion change loss function is used to measure the stability of motion changes between different image frames, and the adversarial loss function is used to measure the rationality of motion recognition. The target loss function is constructed based on at least one of the action recognition loss function, the action change loss function, and the adversarial loss function.

6. The method according to claim 5, characterized in that, The initial network model also includes a first discriminator network, and the adversarial loss function is constructed based on the historical position information of the three-dimensional key points in the following ways: The first discriminator network is used to discriminate the historical position information of the three-dimensional joints to obtain a first discrimination result; The adversarial loss function is constructed based on the historical location information of the three-dimensional key points and the first discrimination result.

7. The method according to claim 6, characterized in that, The initial network model further includes a second discriminator network, and the method further includes: The second discriminator network is used to discriminate the rotation parameters of the action history to obtain a second discrimination result; The step of constructing the adversarial loss function based on the three-dimensional keypoint historical position information and the first discrimination result includes: The adversarial loss function is constructed based on the historical location information of the three-dimensional key points, the first discrimination result, and the second discrimination result.

8. The method according to claim 5, characterized in that, The methods for constructing the action recognition loss function based on the historical position information of the three-dimensional joints include: Based on the historical position information and true values ​​of the position information of the three-dimensional joints, a first loss function is determined; The second loss function is determined based on the motion history rotation parameters and the true values ​​of the motion rotation parameters corresponding to the historical position information of the three-dimensional joints. The third loss function is determined based on the motion history displacement parameters and the true values ​​of the displacement parameters corresponding to the historical position information of the three-dimensional joints. The action recognition loss function is obtained by weighted summation of the first loss function, the second loss function, and the third loss function.

9. The method according to claim 5, characterized in that, The methods for constructing the motion change loss function based on the historical position information of the three-dimensional joints include: The fourth loss function is determined based on the difference between the historical position information of the three-dimensional key points corresponding to two adjacent historical image frames and the true value of the first difference. The fifth loss function is determined based on the difference between the motion history rotation parameters corresponding to two adjacent historical image frames and the true value of the second difference. The sixth loss function is determined based on the difference between the motion history displacement parameters corresponding to two adjacent historical image frames and the true value of the third difference. The action change loss function is obtained by weighted summation of the fourth loss function, the fifth loss function, and the sixth loss function.

10. A motion recognition device, characterized in that, The device includes an acquisition unit, a generation unit, a prediction unit, and an analysis unit: The acquisition unit is used to acquire the two-dimensional joint point position information of the target object in the image frame to be identified when the target object is photographed to obtain the image frame to be identified; The generation unit is used to take the two-dimensional joint position information as input to the action recognition model, and use the feature generation module of the action recognition model to generate features to obtain the target feature vector. The prediction unit is used to predict the motion rotation parameters and motion displacement parameters of each joint of the target object based on the target feature vector using the prediction module of the motion recognition model. The prediction module includes a rotation parameter prediction module and a displacement prediction module. The rotation parameter prediction module is used to predict the motion rotation parameters of each joint of the target object based on the target feature vector. The displacement prediction module is used to predict the motion displacement parameters of each joint of the target object based on the target feature vector. The analysis unit is used to perform kinematic analysis using the kinematic analysis module of the motion recognition model based on the motion rotation parameters and the motion displacement parameters, and to obtain the three-dimensional joint position information of the corresponding joints. The feature generation module includes a first feature extraction module and a first feature fusion module, wherein the generation unit is specifically used for: Obtain the feature extraction results of the two-dimensional joint position information; Based on the feature extraction results of the two-dimensional joint point position information, the first feature extraction module is used to generate the first feature vector corresponding to the image frame to be identified. The first feature fusion module fuses the first feature vector corresponding to the image frame to be identified with the first feature vector corresponding to the previous image frame to obtain the target feature vector. The previous image frame and the image frame to be identified are located in the same image frame sequence. In the image frame sequence, the previous image frame is before the image frame to be identified and is adjacent to the image frame to be identified.

11. The apparatus according to claim 10, characterized in that, The feature generation module further includes a second feature extraction module and a second feature fusion module. In the action recognition model, the second feature extraction module and the second feature fusion module are located before the first feature extraction module. The generation unit is specifically used for: The second feature extraction module extracts features from the two-dimensional joint position information to obtain the second feature vector corresponding to the image frame to be identified. The second feature vector corresponding to the image frame to be identified is determined as the feature extraction result; The second feature fusion module fuses the second feature vector corresponding to the image frame to be identified with the second feature vector corresponding to the previous image frame to obtain a fused feature vector. The first feature vector is obtained by encoding the fused feature vector through the first feature extraction module.

12. The apparatus according to any one of claims 10-11, characterized in that, The generation unit is further configured to: An animation of the target object is generated based on the three-dimensional joint position information.

13. The apparatus according to any one of claims 10-11, characterized in that, The initial network model corresponding to the action recognition model includes an initial feature generation module, an initial prediction module, and an initial kinematics analysis module. The device further includes a training unit, which is used for: Obtain the historical position information of two-dimensional joints of historical objects in historical image frames; Based on the historical position information of the two-dimensional joints, feature generation is performed using the feature generation initial module to obtain the target historical feature vector. Based on the target historical feature vector, the prediction initialization module is used to predict and obtain the motion history rotation parameters and motion history displacement parameters of each joint of the historical object. Based on the motion history rotation parameters and the motion history displacement parameters, kinematic analysis is performed using the kinematic analysis initialization module to obtain the three-dimensional joint point historical position information of the corresponding joint points. Construct a target loss function based on the historical position information of the three-dimensional joints; The model parameters of the initial network model are optimized and adjusted according to the target loss function to obtain the action recognition model.

14. The apparatus according to claim 13, characterized in that, The training unit is specifically used for: Based on the historical position information of the three-dimensional joints, a motion recognition loss function, a motion change loss function, and an adversarial loss function are constructed respectively. The motion recognition loss function is used to measure the accuracy of motion recognition, the motion change loss function is used to measure the stability of motion changes between different image frames, and the adversarial loss function is used to measure the rationality of motion recognition. The target loss function is constructed based on at least one of the action recognition loss function, the action change loss function, and the adversarial loss function.

15. The apparatus according to claim 14, characterized in that, The initial network model further includes a first discriminator network, and the training unit is specifically used for: The first discriminator network is used to discriminate the historical position information of the three-dimensional joints to obtain a first discrimination result; The adversarial loss function is constructed based on the historical location information of the three-dimensional key points and the first discrimination result.

16. The apparatus according to claim 15, characterized in that, The initial network model further includes a second discriminator network, and the training unit is further used for: The second discriminator network is used to discriminate the rotation parameters of the action history to obtain a second discrimination result; The training unit is specifically used for: The adversarial loss function is constructed based on the historical location information of the three-dimensional key points, the first discrimination result, and the second discrimination result.

17. The apparatus according to claim 14, characterized in that, The training unit is specifically used for: Based on the historical position information and true values ​​of the position information of the three-dimensional joints, a first loss function is determined; The second loss function is determined based on the motion history rotation parameters and the true values ​​of the motion rotation parameters corresponding to the historical position information of the three-dimensional joints. The third loss function is determined based on the motion history displacement parameters and the true values ​​of the displacement parameters corresponding to the historical position information of the three-dimensional joints. The action recognition loss function is obtained by weighted summation of the first loss function, the second loss function, and the third loss function.

18. The apparatus according to claim 14, characterized in that, The training unit is specifically used for: The fourth loss function is determined based on the difference between the historical position information of the three-dimensional key points corresponding to two adjacent historical image frames and the true value of the first difference. The fifth loss function is determined based on the difference between the motion history rotation parameters corresponding to two adjacent historical image frames and the true value of the second difference. The sixth loss function is determined based on the difference between the motion history displacement parameters corresponding to two adjacent historical image frames and the true value of the third difference. The action change loss function is obtained by weighted summation of the fourth loss function, the fifth loss function, and the sixth loss function.

19. An electronic device for action recognition, characterized in that, The electronic device includes a processor and a memory: The memory is used to store program code and transmit the program code to the processor; The processor is configured to execute the method according to any one of claims 1-9 according to the instructions in the program code.

20. A computer-readable storage medium, characterized in that, The computer-readable storage medium is used to store program code that, when executed by a processor, causes the processor to perform the method according to any one of claims 1-9.

21. A computer program product, comprising a computer program, characterized in that, When executed by a processor, the computer program implements the method described in any one of claims 1-9.

Citation Information

Patent Citations

  • Motion recognition method and device

    CN111401318A

  • Action behavior recognition method and device, storage medium and terminal equipment

    CN113723185A

  • Action recognition method and device, equipment and storage medium

    CN113792712A