Human action recognition method, device, apparatus and storage medium

By using a human pose estimation model to identify key points and calculate relative position features during simulation training, and combining this with limb feature vectors for judgment, the real-time and accuracy problems of human motion recognition in existing technologies are solved, achieving fast and accurate motion recognition.

CN122435672APending Publication Date: 2026-07-21HEBEI JUNTAO TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-04-16
Publication Date
2026-07-21

AI Technical Summary

Technical Problem

Existing technologies struggle to meet the real-time and accuracy requirements of human motion recognition in simulation training, especially when monitoring and correcting trainees' postures in real time in high-risk, high-precision training scenarios, where the recognition results are poor.

Method used

By acquiring training images of trainees, key points are identified using a pre-set human pose estimation model. The relative positional features between key points are calculated and compared with a pre-set reference pose. The main key points are selected, and human motion recognition is achieved by combining limb feature vectors and similarity judgment.

Benefits of technology

It improves the real-time performance and accuracy of human motion recognition, enabling it to quickly and accurately identify the movements of trainees, adapt to individual differences, and provide overall motion recognition results.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122435672A_ABST
    Figure CN122435672A_ABST
Patent Text Reader

Abstract

The application provides a human action recognition method, device and equipment and a storage medium, and relates to the technical field of training action recognition. The method comprises the following steps: obtaining a training image of a trainee, identifying a human joint point in the training image based on a preset human posture estimation model, and obtaining each key point of the trainee in the training image; obtaining a relative position feature between each two key points in the training image as a target relative position feature according to the key points; calculating a distance between each target relative position feature and a relative position feature between corresponding key points in a preset reference posture, and determining a main key point according to the distance; comparing the main key point with a main key point in a standard posture corresponding to the training image, and obtaining a human action recognition result of the trainee in the training image according to a comparison result. The application can balance the real-time performance and accuracy of human action recognition.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of simulation training motion recognition technology, and in particular to a human motion recognition method, device, equipment and storage medium. Background Technology

[0002] In military simulation training, especially in high-risk and high-precision training fields such as grenade throwing, using extended reality (XR) technology to construct virtual training scenarios has become an important means to improve training safety and realism. In order to evaluate the effect of simulation training, it is usually necessary to acquire and identify human movements during the simulation training process.

[0003] Currently, conventional methods typically rely on computer vision technology to analyze training images, extract the trainee's human movements, and compare them with standard movements to achieve human movement recognition. However, while these methods perform well in scenarios with relatively simple movements or low requirements for recognition speed, in simulated training scenarios, it is usually necessary to monitor and correct the trainee's practice postures in real time. This places high demands on both real-time performance and accuracy, making it difficult for existing methods to meet practical needs in simulated training. Summary of the Invention

[0004] This invention provides a human motion recognition method, apparatus, device, and storage medium to solve the problem that conventional methods are unable to meet the real-time and accuracy requirements of human motion recognition during simulation training.

[0005] In a first aspect, embodiments of the present invention provide a human motion recognition method, including: Acquire training images of trainees, and identify human joints in the training images based on a preset human pose estimation model to obtain the trainees' key points in the training images. Based on each of the key points, the relative position features between every two key points in the training image are obtained and denoted as the target relative position features. Calculate the distance between the relative position features of each target and the corresponding key points in the preset reference pose, and determine the main key points based on the distance; The key points are compared with the key points in the standard pose corresponding to the training image, and the human motion recognition result of the trainee in the training image is obtained based on the comparison result.

[0006] In one possible implementation, before calculating the distance between the relative position features of each target and the corresponding key point in the preset reference pose, the method further includes: Determine whether the training image is the first frame image during the simulation training of the trainee; If the training image is the first frame image, the preset pose is determined as the preset reference pose; If the training image is not the first frame image, the standard pose corresponding to the previous training image preceding the training image is determined as the preset reference pose.

[0007] In one possible implementation, the key points are compared with the key points in the standard pose corresponding to the training image, and the human motion recognition result of the trainee in the training image is obtained based on the comparison result, including: Based on the main key points, extract the limb feature vector represented by the main key points, and denote it as the main target limb feature vector; Based on the main key points in the standard pose corresponding to the training image, extract the main standard limb feature vector corresponding to the training image; Calculate the similarity between the main target limb feature vector and the main standard limb feature vector, and denote it as the main similarity. The human motion recognition results of the trainees in the training images are obtained based on the main similarity.

[0008] In one possible implementation, obtaining the human action recognition result of the trainee in the training image based on the main similarity includes: The main similarity is compared with a first threshold and a second threshold, respectively, wherein the first threshold is greater than the second threshold; If the main similarity is greater than the first threshold, the human motion recognition result of the trainee in the training image is determined to be qualified. If the main similarity is greater than the second threshold and less than or equal to the first threshold, the main target limb feature vector is input into the Gaussian mixture model of the standard pose corresponding to the training image to obtain the pass probability of the main target limb feature vector, and the human action recognition result of the trainee in the training image is obtained according to the pass probability. If the primary similarity is less than or equal to the second threshold, then secondary key points are determined based on the distance, and the human motion recognition results of the trainee in the training image are obtained based on the secondary key points.

[0009] In one possible implementation, obtaining the human motion recognition result of the trainee in the training image based on the secondary key points includes: Based on the secondary key points, the limb feature vector represented by the secondary key points is extracted and denoted as the secondary target limb feature vector. Based on the secondary key points in the standard pose corresponding to the training image, extract the secondary standard limb feature vector corresponding to the training image; Calculate the similarity between the secondary target limb feature vector and the secondary standard limb feature vector, and denote it as the secondary similarity. The human motion recognition results of the trainees in the training images are obtained based on the secondary similarity.

[0010] In one possible implementation, after obtaining the human action recognition result of the trainee in the training image based on the comparison result, the method further includes: Determine whether the training image is the last frame image during the simulation training of the trainee; If the training image is the last frame of the training image during the simulation training, then the overall action recognition result of the trainee is obtained based on the human action recognition result corresponding to each frame of the training image.

[0011] In one possible implementation, before comparing the principal keypoints with the principal keypoints in the standard pose corresponding to the training image, the method further includes: Based on the relative position features of each target, the ratio between different bone lengths of the trainee is extracted; Determine whether the stated ratio exceeds the standard ratio range; If the ratio exceeds the standard ratio range, the standard pose corresponding to the training image is updated according to the ratio.

[0012] Secondly, embodiments of the present invention provide a human motion recognition device, comprising: The first processing module is used to acquire training images of trainees and identify human joints in the training images based on a preset human pose estimation model to obtain the various key points of the trainees in the training images. The second processing module is used to obtain the relative position features between every two key points in the training image based on each of the key points, and denoted as the target relative position features; The third processing module is used to calculate the distance between the relative position features of each target and the corresponding key points in the preset reference pose, and to determine the main key points based on the distance. The recognition module is used to compare the main key points with the key points of the standard pose corresponding to the training image, and obtain the human motion recognition result of the trainee in the training image based on the comparison result.

[0013] Thirdly, embodiments of the present invention provide an electronic device, including a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to implement the method described in the first aspect or any possible implementation thereof.

[0014] Fourthly, embodiments of the present invention provide a computer-readable storage medium storing a computer program that, when executed by a processor, implements the method described in the first aspect or any possible implementation thereof.

[0015] In this embodiment of the invention, training images of the trainee are first acquired, and human joints in the training images are identified based on a preset human pose estimation model to obtain various key points of the trainee in the training images. Then, based on each key point, the relative position features between every two key points in the training images are obtained and recorded as target relative position features. Furthermore, the distance between each target relative position feature and the corresponding key point in a preset reference pose is calculated, and the main key points are determined based on the distance. The main key points are then compared with the main key points in the standard pose corresponding to the training images, and the human motion recognition result of the trainee in the training images is obtained based on the comparison result. Thus, by selecting the main key points, the human motion recognition result of the trainee in the training images can be obtained faster and more accurately. Attached Figure Description

[0016] Figure 1 This is a flowchart illustrating the implementation of the human motion recognition method provided in this embodiment of the invention. Figure 2 This is a schematic diagram of the human motion recognition device provided in an embodiment of the present invention; Figure 3 This is a schematic diagram of an electronic device provided in an embodiment of the present invention. Detailed Implementation

[0017] The embodiments of the present invention will now be described in detail with reference to the accompanying drawings.

[0018] See Figure 1 The flowchart illustrating the implementation of the human motion recognition method provided in this embodiment of the invention is shown below: Step 101: Obtain training images of trainees and identify human joints in the training images based on a preset human pose estimation model to obtain the key points of trainees in the training images.

[0019] For example, taking grenade simulation training, the training images can be images captured during grenade simulation training, such as images of actions like holding the grenade, detonating it, throwing it, and crouching. The preset human pose estimation model can be a pre-trained human joint detection model, such as the human pose detection model using OpenPose.

[0020] Step 102: Based on each key point, obtain the relative position features between every two key points in the training image, and denot them as the target relative position features.

[0021] In this embodiment, the relative position feature between any two key points can be the relative distance between them. By calculating the relative distance between any two key points, the main key points can be quickly filtered out in the subsequent process.

[0022] Step 103: Calculate the distance between the relative position features of each target and the corresponding key points in the preset reference pose, and determine the main key points based on the distance.

[0023] In one embodiment, before calculating the distance between the relative position features of each target and the corresponding keypoint in the preset reference pose, the method further includes: Determine whether the training image is the first frame image when the trainee is conducting simulation training.

[0024] If the training image is the first frame, the preset pose will be set as the preset reference pose.

[0025] If the training image is not the first frame, the standard pose corresponding to the previous training image is determined as the preset reference pose.

[0026] The preset posture can be a standing posture, an outstretched arms posture, etc. In this embodiment, during the process of recognizing the human body movements of the trainees, a fixed preset posture is set for the first frame image as a preset reference posture, and the standard posture corresponding to the previous image is set for the images after the first frame image as the preset reference posture. By calculating the distance between the relative position features of each target and the corresponding key points in the preset reference posture, the key points with large posture changes, i.e., the main key points, can be accurately located, thus providing a foundation for quickly and accurately obtaining the human body movement recognition results.

[0027] For example, when calculating the distance between the relative position features of each target and the corresponding key points in the preset reference pose, and determining the main key points based on the distance, key points whose distance is greater than or equal to a set distance threshold can be determined as main key points.

[0028] For example, after determining the primary key points, key points whose distance is less than a set distance threshold but greater than zero can also be determined as secondary key points.

[0029] Step 104: Compare the main key points with the main key points in the standard pose corresponding to the training image, and obtain the human motion recognition result of the trainee in the training image based on the comparison result.

[0030] In this embodiment, by comparing the main key points with the main key points in the standard pose corresponding to the training image, the human motion recognition result of the trainee in the training image is obtained based on the comparison result. Only the comparison between the main key points can be performed, thereby improving the speed of human motion recognition.

[0031] In one embodiment, before comparing the principal keypoints with the principal keypoints in the standard pose corresponding to the training image, the method further includes: Based on the relative position features of each target, the ratio between different bone lengths of the trainee is extracted.

[0032] Determine whether the ratio exceeds the standard ratio range.

[0033] If the scale exceeds the standard scale range, the standard pose corresponding to the training image is updated according to the scale.

[0034] In this embodiment, before obtaining the human motion recognition result based on the main key points, considering the differences between different individuals, the standard pose corresponding to the training image is adaptively updated by extracting the ratio between different bone lengths of the trainees, so that the human motion recognition can be more accurate based on the updated standard pose.

[0035] In one embodiment, step 104 includes: Based on the main key points, extract the limb feature vectors represented by the main key points, and denote them as the main target limb feature vectors.

[0036] Based on the main key points in the standard pose corresponding to the training image, the main standard limb feature vectors corresponding to the training image are extracted.

[0037] Calculate the similarity between the main target limb feature vector and the main standard limb feature vector, and denote it as the main similarity.

[0038] The results of human motion recognition of trainees in training images are obtained based on the main similarity.

[0039] In one embodiment, obtaining the human motion recognition result of the trainee in the training image based on the main similarity includes: The main similarity is compared with a first threshold and a second threshold, respectively, with the first threshold being greater than the second threshold.

[0040] If the primary similarity is greater than the first threshold, the human motion recognition result of the trainee in the training image is determined to be qualified.

[0041] If the main similarity is greater than the second threshold and less than or equal to the first threshold, the main target limb feature vector is input into the Gaussian mixture model of the standard pose corresponding to the training image to obtain the pass probability of the main target limb feature vector. Based on the pass probability, the human motion recognition result of the trainee in the training image is obtained.

[0042] If the primary similarity is less than or equal to the second threshold, secondary key points are determined based on the distance, and the human motion recognition results of the trainees in the training images are obtained based on the secondary key points.

[0043] In this embodiment, considering individual differences in obtaining human action recognition results based on key points, after obtaining the similarity between the main target limb feature vector corresponding to the key points and the main standard limb feature vector, a first threshold and a second threshold are set. When the main similarity is greater than the first threshold, the human action recognition result of the trainee in the training image is directly determined as qualified. When the main similarity is greater than the second threshold but less than or equal to the first threshold, a Gaussian mixture model is used for further judgment. When the main similarity is less than or equal to the second threshold, secondary key points are combined for judgment. This improves the speed of human action recognition while maintaining accuracy.

[0044] In one embodiment, obtaining the human motion recognition result of the trainee in the training image based on secondary key points includes: Based on secondary key points, extract the limb feature vectors represented by the secondary key points, and denote them as secondary target limb feature vectors.

[0045] Based on the secondary key points in the standard pose corresponding to the training image, the secondary standard limb feature vector corresponding to the training image is extracted.

[0046] The similarity between the secondary target limb feature vector and the secondary standard limb feature vector is calculated and denoted as the secondary similarity.

[0047] The results of human motion recognition of trainees in training images are obtained based on secondary similarity.

[0048] For example, the human action recognition results of the trainees in the training images are obtained based on secondary similarity, that is, the secondary similarity is combined with the primary similarity. If the similarity after the combination is high, for example, exceeding the second threshold, the human action recognition result can be marked as qualified. If the similarity after the combination is still low, the human action recognition result is marked as unqualified.

[0049] In one embodiment, after obtaining the human motion recognition result of the trainee in the training image based on the comparison result, the method further includes: Determine whether the training image is the last frame image taken during the simulation training of the trainee.

[0050] If the training image is the last frame of the training process, then the overall action recognition result of the trainee is obtained based on the human action recognition result corresponding to each frame of the training image.

[0051] In this embodiment, based on the human motion recognition results of the trainee in the training images, the overall motion recognition result of the trainee is obtained according to the human motion recognition results corresponding to each frame of the training images.

[0052] This invention first acquires training images of the trainee and identifies human joints in the training images based on a preset human pose estimation model to obtain various key points of the trainee in the training images. Then, based on each key point, the relative position features between every two key points in the training images are obtained and denoted as target relative position features. Furthermore, by calculating the distance between each target relative position feature and the corresponding key point in a preset reference pose, the main key points are determined based on the distance. The main key points are then compared with the main key points in the standard pose corresponding to the training images, and the human motion recognition result of the trainee in the training images is obtained based on the comparison result. Therefore, by selecting the main key points, the human motion recognition result of the trainee in the training images can be obtained faster and more accurately.

[0053] It should be understood that the sequence number of each step in the above embodiments does not imply the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present invention.

[0054] The following are device embodiments of the present invention. For details not described in detail, please refer to the corresponding method embodiments described above.

[0055] Figure 2 A schematic diagram of the human motion recognition device provided in an embodiment of the present invention is shown. For ease of explanation, only the parts related to the embodiment of the present invention are shown, and are described in detail below: like Figure 2 As shown, the human motion recognition device includes: The first processing module 21 is used to acquire training images of trainees and identify human joints in the training images based on a preset human pose estimation model to obtain the various key points of the trainees in the training images.

[0056] The second processing module 22 is used to obtain the relative position features between every two key points in the training image based on each key point, which is denoted as the target relative position features.

[0057] The third processing module 23 is used to calculate the distance between the relative position features of each target and the corresponding key points in the preset reference pose, and to determine the main key points based on the distance.

[0058] The recognition module 24 is used to compare the main key points with the key points of the standard pose corresponding to the training image, and obtain the human motion recognition result of the trainee in the training image based on the comparison result.

[0059] In one possible implementation, the third processing module 23 is further configured to: Determine whether the training image is the first frame image when the trainee is conducting simulation training.

[0060] If the training image is the first frame, the preset pose will be set as the preset reference pose.

[0061] If the training image is not the first frame, the standard pose corresponding to the previous training image is determined as the preset reference pose.

[0062] In one possible implementation, the identification module 24 is specifically used for: Based on the main key points, extract the limb feature vectors represented by the main key points, and denote them as the main target limb feature vectors.

[0063] Based on the main key points in the standard pose corresponding to the training image, the main standard limb feature vectors corresponding to the training image are extracted.

[0064] Calculate the similarity between the main target limb feature vector and the main standard limb feature vector, and denote it as the main similarity.

[0065] The results of human motion recognition of trainees in training images are obtained based on the main similarity.

[0066] In one possible implementation, the identification module 24 is specifically used for: The main similarity is compared with a first threshold and a second threshold, respectively, with the first threshold being greater than the second threshold.

[0067] If the primary similarity is greater than the first threshold, the human motion recognition result of the trainee in the training image is determined to be qualified.

[0068] If the main similarity is greater than the second threshold and less than or equal to the first threshold, the main target limb feature vector is input into the Gaussian mixture model of the standard pose corresponding to the training image to obtain the pass probability of the main target limb feature vector. Based on the pass probability, the human motion recognition result of the trainee in the training image is obtained.

[0069] If the primary similarity is less than or equal to the second threshold, secondary key points are determined based on the distance, and the human motion recognition results of the trainees in the training images are obtained based on the secondary key points.

[0070] In one possible implementation, the identification module 24 is specifically used for: Based on secondary key points, extract the limb feature vectors represented by the secondary key points, and denote them as secondary target limb feature vectors.

[0071] Based on the secondary key points in the standard pose corresponding to the training image, the secondary standard limb feature vector corresponding to the training image is extracted.

[0072] The similarity between the secondary target limb feature vector and the secondary standard limb feature vector is calculated and denoted as the secondary similarity.

[0073] The results of human motion recognition of trainees in training images are obtained based on secondary similarity.

[0074] In one possible implementation, the identification module 24 is further configured to: Determine whether the training image is the last frame image taken during the simulation training of the trainee.

[0075] If the training image is the last frame of the training process, then the overall action recognition result of the trainee is obtained based on the human action recognition result corresponding to each frame of the training image.

[0076] In one possible implementation, the identification module 24 is further configured to: Based on the relative position features of each target, the ratio between different bone lengths of the trainee is extracted.

[0077] Determine whether the ratio exceeds the standard ratio range.

[0078] If the scale exceeds the standard scale range, the standard pose corresponding to the training image is updated according to the scale.

[0079] Figure 3 This is a schematic diagram of an electronic device provided in an embodiment of the present invention. For example... Figure 3As shown, the electronic device 3 of this embodiment includes a processor 30 and a memory 31. The memory 31 stores a computer program 32. When the processor 30 executes the computer program 32, it implements the steps in the various method embodiments described above. Alternatively, when the processor 30 executes the computer program 32, it implements the functions of each module / unit in the various device embodiments described above.

[0080] For example, computer program 32 may be divided into one or more modules / units, which are stored in memory 31 and executed by processor 30 to complete the present invention. The one or more modules / units may be a series of computer program instruction segments capable of performing a specific function, which describe the execution process of computer program 32 in electronic device 3.

[0081] Electronic device 3 may include, but is not limited to, processor 30 and memory 31. Those skilled in the art will understand that... Figure 3 This is merely an example of electronic device 3 and does not constitute a limitation on electronic device 3. It may include more or fewer components than shown, or combine certain components, or different components. For example, electronic device 3 may also include input / output devices, network access devices, buses, etc.

[0082] The processor 30 can be a central processing unit (CPU), or other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor can be a microprocessor or any conventional processor.

[0083] The memory 31 can be an internal storage unit of the electronic device 3, such as a hard disk or memory of the electronic device 3. The memory 31 can also be an external storage device of the electronic device 3, such as a plug-in hard disk, smart media card (SMC), secure digital (SD) card, flash card, etc., equipped on the electronic device 3. Furthermore, the memory 31 can include both internal and external storage units of the electronic device 3. The memory 31 is used to store the computer program 32 and other programs and data required by the electronic device 3. The memory 31 can also be used to temporarily store data that has been output or will be output.

[0084] For the sake of simplicity and clarity, only the above-described functional modules / units are used as examples. In practical applications, the functions described above can be assigned to different functional modules / units as needed. These modules / units can be implemented in hardware, software, or a combination of both.

[0085] This invention also provides a computer-readable storage medium storing a computer program. When the computer program is executed by a processor, it implements the methods described in the above-described method embodiments.

[0086] Computer programs include computer program code, which can be in the form of source code, object code, executable files, or certain intermediate forms. Computer-readable media can include: any entity or device capable of carrying computer program code, recording media, USB flash drives, portable hard drives, magnetic disks, optical disks, computer memory, read-only memory (ROM), random access memory (RAM), electrical carrier signals, telecommunication signals, and software distribution media, etc.

[0087] In the above embodiments, the descriptions of each embodiment have their own emphasis. Parts not detailed or described in a particular embodiment can be referred to in the relevant descriptions of other embodiments. Unless otherwise specified or in conflict with logic, the terminology and / or descriptions between different embodiments are consistent and can be referenced interchangeably. Technical features in different embodiments can be combined to form new embodiments based on their inherent logical relationships.

[0088] The above-described embodiments are only used to illustrate the technical solutions of the present invention, and are not intended to limit it. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention, and should all be included within the protection scope of the present invention.

Claims

1. A method for recognizing human motion, characterized in that, include: Acquire training images of trainees, and identify human joints in the training images based on a preset human pose estimation model to obtain the trainees' key points in the training images. Based on each of the key points, the relative position features between every two key points in the training image are obtained and denoted as the target relative position features. Calculate the distance between the relative position features of each target and the corresponding key points in the preset reference pose, and determine the main key points based on the distance; The key points are compared with the key points in the standard pose corresponding to the training image, and the human motion recognition result of the trainee in the training image is obtained based on the comparison result.

2. The human motion recognition method according to claim 1, characterized in that, Before calculating the distance between the relative position features of each target and the corresponding key point in the preset reference pose, the method further includes: Determine whether the training image is the first frame image during the simulation training of the trainee; If the training image is the first frame image, the preset pose is determined as the preset reference pose; If the training image is not the first frame image, the standard pose corresponding to the previous training image preceding the training image is determined as the preset reference pose.

3. The human motion recognition method according to claim 1, characterized in that, The key points are compared with the key points in the standard pose corresponding to the training image, and the human motion recognition result of the trainee in the training image is obtained based on the comparison result, including: Based on the main key points, extract the limb feature vector represented by the main key points, and denote it as the main target limb feature vector; Based on the main key points in the standard pose corresponding to the training image, extract the main standard limb feature vector corresponding to the training image; Calculate the similarity between the main target limb feature vector and the main standard limb feature vector, and denote it as the main similarity. The human motion recognition results of the trainees in the training images are obtained based on the main similarity.

4. The human motion recognition method according to claim 3, characterized in that, The human motion recognition results of the trainees in the training images are obtained based on the main similarity, including: The main similarity is compared with a first threshold and a second threshold, respectively, wherein the first threshold is greater than the second threshold; If the main similarity is greater than the first threshold, the human motion recognition result of the trainee in the training image is determined to be qualified. If the main similarity is greater than the second threshold and less than or equal to the first threshold, the main target limb feature vector is input into the Gaussian mixture model of the standard pose corresponding to the training image to obtain the pass probability of the main target limb feature vector, and the human action recognition result of the trainee in the training image is obtained according to the pass probability. If the primary similarity is less than or equal to the second threshold, then secondary key points are determined based on the distance, and the human motion recognition results of the trainee in the training image are obtained based on the secondary key points.

5. The human motion recognition method according to claim 4, characterized in that, Obtaining the human motion recognition results of the trainee in the training image based on the secondary key points includes: Based on the secondary key points, the limb feature vector represented by the secondary key points is extracted and denoted as the secondary target limb feature vector. Based on the secondary key points in the standard pose corresponding to the training image, extract the secondary standard limb feature vector corresponding to the training image; Calculate the similarity between the secondary target limb feature vector and the secondary standard limb feature vector, and denote it as the secondary similarity. The human motion recognition results of the trainees in the training images are obtained based on the secondary similarity.

6. The human motion recognition method according to claim 1, characterized in that, After obtaining the human motion recognition results of the trainees in the training images based on the comparison results, the method further includes: Determine whether the training image is the last frame image during the simulation training of the trainee; If the training image is the last frame of the training image during the simulation training, then the overall action recognition result of the trainee is obtained based on the human action recognition result corresponding to each frame of the training image.

7. The human motion recognition method according to claim 1, characterized in that, Before comparing the principal keypoints with the principal keypoints in the standard pose corresponding to the training image, the method further includes: Based on the relative position features of each target, the ratio between different bone lengths of the trainee is extracted; Determine whether the stated ratio exceeds the standard ratio range; If the ratio exceeds the standard ratio range, the standard pose corresponding to the training image is updated according to the ratio.

8. A human motion recognition device, characterized in that, include: The first processing module is used to acquire training images of trainees and identify human joints in the training images based on a preset human pose estimation model to obtain the various key points of the trainees in the training images. The second processing module is used to obtain the relative position features between every two key points in the training image based on each of the key points, and denoted as the target relative position features; The third processing module is used to calculate the distance between the relative position features of each target and the corresponding key points in the preset reference pose, and to determine the main key points based on the distance. The recognition module is used to compare the main key points with the key points of the standard pose corresponding to the training image, and obtain the human motion recognition result of the trainee in the training image based on the comparison result.

9. An electronic device, characterized in that, It includes a memory and a processor, the memory storing a computer program, and the processor executing the computer program to implement the method as described in any one of claims 1 to 7.

10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program that, when executed by a processor, implements the method as described in any one of claims 1 to 7.