Human motion similarity matching scoring method, device and readable storage medium

By performing similar transformation and weighted summing in the matching of human movement postures, the problem of human body orientation and key point weights in the prior art are solved, the accuracy of matching evaluation is improved, and the real-time cost and calculation amount are reduced.

CN114724013BActive Publication Date: 2025-05-13SHENZHEN MAXVISION TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210276645.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-03-21
Publication Date
2025-05-13
Estimated Expiration
2042-03-21

AI Technical Summary

Technical Problem

The prior art ignores the weight of the human body orientation and key points of reference movements and imitation movements in the matching of human body movement postures, resulting in large errors in matching evaluation, and methods based on depth information are costly and have large calculations.

Method used

By obtaining the imitator's action diagram, reference action diagram and human body stance diagram, the key points representing the human body skeleton are detected, the similarity transformation is carried out to unify the human body's orientation, the skeleton feature vector is calculated, and the similarity distances of the imitator's action and reference action are weighted and summed to evaluate the similarity of the action skeleton.

Benefits of technology

Effectively prevent different orientations of the human body from affecting matching evaluation, improve the accuracy of action similarity matching, and reduce real-time cost and calculation amount.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114724013B_ABST
    Figure CN114724013B_ABST
Patent Text Reader

Abstract

The present application discloses a human action similarity matching and scoring method, which includes: obtaining an imitation action graph F, a reference action graph I and a human standing graph J; detecting multiple human key points representing the human skeleton in the imitation action graph F, the reference action graph I and the human standing graph J; performing similarity transformation on the imitation action graph F according to the multiple human key points of the reference action graph I to obtain multiple new human key points of the imitation action graph F; obtaining multiple skeleton feature vectors representing the human skeleton of the imitation action graph F, the reference action graph I and the human standing graph J; taking the similarity distance between the skeleton feature vectors corresponding to the reference action graph I and the human standing graph J as the weight, weighted summing the similarity distance between the skeleton feature vectors corresponding to the imitation action graph F and the reference action graph I, and obtaining the action skeleton similarity of the imitation action graph F relative to the reference action graph I. The present application also provides a device and a readable storage medium.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of image processing, and more specifically, to a human motion similarity matching scoring method, device and readable storage medium. Background Art

[0002] Analyzing human posture in images is a common topic in computer vision and is the basis for applications such as motion recognition and posture matching. With the development of artificial intelligence technology and the innovation of hardware technology, posture matching has been widely used in human-computer interactive games, video action teaching, and action error correction.

[0003] In the prior art, the matching of human body motion postures mostly extracts the key point information of the reference motion and the imitation motion, then extracts the features and performs a similarity comparison; although this method is simple, it ignores the human body orientation of the reference motion and the imitation motion and the weight of the key point of each reference motion; the human body orientation of the reference motion and the imitation motion can be understood as the existence of different inclinations or directions between the imitator's body and the imitation motion body, for example, the imitation motion body is upright while the imitator's body is tilted to make the motion, or the imitation motion body makes the motion in a certain direction while the imitation motion body makes the motion in another direction; the key point of each reference motion can be understood as: relative to the standing motion, the human body limb movement position point of the reference motion is the key point; if these two key factors are ignored, directly comparing the position information of the key points of the imitation motion and the reference motion will bring great errors. In addition, there are some human motion posture matching methods based on bodyweight fitness assisted teaching methods based on human posture recognition. The fitness person is captured by a camera with depth information, and the detection key points of the fitness person are combined with the depth information conversion to generate a 3D human skeleton map. The motion range, joint angle and other features of the fitness action are extracted and compared with the features of the reference action to obtain the action similarity score. These methods combine the depth information of the image for similarity estimation, and require a camera with depth information to capture the image. The real-time cost is relatively high, and processing the depth information also requires additional computing power. Summary of the invention

[0004] In view of the prior art, the technical problem solved by the present application is to provide a human motion similarity matching method, device and readable storage medium, which can effectively prevent the different human body orientations in the reference action graph and the imitation action graph from affecting the judgment of the reference action similarity, and take into account the importance of key action points, thereby facilitating improving the accuracy of action similarity matching evaluation.

[0005] In order to solve the above technical problems, the present application provides a human motion similarity matching scoring method, which includes:

[0006] Obtain the imitator's action diagram F, the reference action diagram I and the human body standing diagram J;

[0007] Detect multiple human key points representing the human skeleton in the imitator action graph F, the reference action graph I and the human standing graph J, wherein the multiple human key point sets of the imitator action graph F, the reference action graph I and the human standing graph J are respectively denoted as V F 、V I and V J ;

[0008] According to the multiple human body key points of the reference action graph I, the imitator action graph F is similarly transformed to obtain a new multiple human body key points of the imitator action graph F. The new multiple human body key point set of the imitator action graph F is recorded as V F’ ;

[0009] According to V F’ 、V I and V J Obtaining a plurality of skeleton feature vectors representing the human skeleton of the imitator action graph F, the reference action graph I and the human standing graph J; and,

[0010] The similarity distance between the skeleton feature vectors corresponding to the reference action graph I and the human body standing graph J is used as the weight to perform weighted summation on the similarity distance between the skeleton feature vectors corresponding to the imitator action graph F and the reference action graph I to obtain the action skeleton similarity of the imitator action graph F relative to the reference action graph I.

[0011] In a possible implementation, the step of detecting a plurality of human key points representing a human skeleton of any one of the imitator action graph F, the reference action graph I, and the human standing graph J includes: detecting a human region using a YOLOX detection model, and then detecting 17 human key points representing a human skeleton in the human region using HRNet;

[0012] Among them, the 17 key points of the human body representing the human skeleton include the nose tip A v0 , Left Eye A v1 , right eye A v2 , left ear A v3 , right ear A v4 , Left Shoulder A v5 , right shoulder A v6 , left elbow A v7 , right elbow A v8 , left wrist A v9 , right wrist A v10 , Left thigh root A v11 , right thigh root A v12 , left knee Av13 , right knee A v14 , Left ankle A v15 and right ankle A v16 ; A is any image among the imitator action graph F, the reference action graph I and the human body standing graph J, A∈(F, J, I).

[0013] In a possible implementation, the step of performing a similarity transformation on the imitator action graph F according to the multiple human body key points of the reference action graph I to obtain a new multiple human body key points of the imitator action graph F includes:

[0014] According to the reference action diagram I, V F The 17 key points of the human body and the V of the imitator action graph F F Calculate the similarity transformation matrix H between the reference action graph I and the imitator action graph F based on the 17 human key points in the reference action graph I;

[0015] The imitator moves the V of Figure F F Each human key point in the similarity transformation matrix H is subjected to a similarity transformation to obtain 17 new human key points of the imitator's action graph F.

[0016] In a possible implementation, obtaining a plurality of skeleton feature vectors representing the human skeleton of any image A in the imitator action graph F, the reference action graph I, and the human standing graph J is to obtain 19 skeleton feature vectors representing the human skeleton,

[0017] The 19 skeleton feature vectors of any image A are sequentially composed of nose tip A v0 and left eye A v1 The eigenvector P A0 , from the nose tip A v0 and right eye A v2 The eigenvector P A1 , from left eye A v1 and right eye A v2 The eigenvector P A2 , from left eye A v1 and left ear A v3 The eigenvector P A3 , by right eye A v2 and right ear A v4 The eigenvector P A4 , from the nose tip A v0 and left shoulder A v5 The eigenvector P A5 , from the nose tip A v0 and right shoulder A v6 The eigenvector P A6 , from left shoulder Av5 and left elbow A v7 The eigenvector P A7 , from right shoulder A v6 and right elbow A v8 The eigenvector P A8 , from left elbow A v7 and left wrist A v9 The eigenvector P A9 , by the right elbow A v8 and right wrist Av 10 The eigenvector P A10 , by the left shoulder Av5 and the left thigh root A v11 The eigenvector P A11 , from right shoulder A v6 and right thigh A v12 The eigenvector P A12 , from left shoulder A v5 and right shoulder A v6 The eigenvector P A13 , from the root of the left thigh A v11 and right thigh A v12 The eigenvector P A14 , from the root of the left thigh A v11 and left knee A v13 The eigenvector P A15 , from the right thigh A v12 and right knee A v14 The eigenvector P A16 , from the left knee A v13 and left ankle A v15 The eigenvector P A17 and by the right knee A v14 and right ankle A v16 The eigenvector P A18 .

[0018] In a possible implementation, the steps of weighting and summing the similarity distances between the skeleton feature vectors corresponding to the imitator action graph F and the reference action graph I by taking the similarity distance between the skeleton feature vectors corresponding to the reference action graph I and the human body standing graph J as weights, and obtaining the action skeleton similarity of the imitator action graph F relative to the reference action graph I include:

[0019] Calculate the skeleton feature vector P of the reference action graph I Ii The skeleton feature vector P corresponding to the human body standing figure J Ji The similarity distance W between i :W i =1-cos <P Ii ·PJi >=1-P Ii ·P Ji / |P Ii |·|P Ji |;

[0020] For all similarity distances W i Perform normalization operation, similarity distance W i The normalized result is recorded as

[0021] Calculate the skeleton feature vector P of the imitator action graph F Fi The skeleton feature vector P corresponding to the reference action graph I Ii The similarity distance D i :D i =1-cos <P Fi ·P Ii >=1-P Fi ·P Ii / |P Fi |·|P Ii |;

[0022] With all the similarity distances W i is the weight for all similarity distances D i Perform weighted summation to obtain the action skeleton similarity M of the imitator action graph F relative to the reference action graph I:

[0023] Among them, P Ii is the i-th eigenvector of the reference action graph I, P Fi is the i-th feature vector of the imitation action graph F, P Ji is the i-th eigenvector of the human body stereogram J, i∈[0,18].

[0024] In a possible implementation, the human body action similarity matching scoring method further includes scoring the similarity of the imitator action graph F according to the action skeleton similarity of the imitator action graph F relative to the reference action graph I,

[0025] S=(dt-M)*(100-st) / dt+st-neg_score,

[0026] Among them, S represents the scoring result of the similarity of the imitator's action graph F; dt is used to control the matching scoring standard, and the smaller dt is, the higher the matching scoring standard is; st is used to control the minimum score, which can make the final score of S controlled within the range of 0 to 100; neg_score is the action reverse penalty item, which is used to punish the situation where the human body movements of the imitator's action graph F are far different from the human body movements of the reference action graph I.

[0027] In a possible implementation, dt takes a value of 2 and st takes a value of 50; when M is less than 0.4, neg_score takes a value of 1; when M is greater than or equal to 0.4 and less than or equal to 0.6, neg_score takes a value of 2; when M is greater than 0.6 and less than or equal to 0.8, M takes a value of 4; when M is greater than 0.8 and less than or equal to 1.0, neg_score takes a value of 6; when M is greater than 1, neg_score takes a value of 8.

[0028] The present application also provides a human motion similarity matching and scoring device, comprising:

[0029] Image acquisition unit: used to acquire the imitator's action image F, the reference action image I and the human body standing image J;

[0030] An image processing unit: connected to the image acquisition unit and used to execute the steps of the human action similarity matching and scoring method; and

[0031] Voice announcer: connected to the image processing unit, when the image processing unit completes the matching and scoring of the action similarity of the imitator's action graph F relative to the reference action graph I, the voice announcer is used to announce the score.

[0032] The beneficial effects of the human motion similarity matching scoring method, device and readable storage medium provided in the present application are: performing a similarity transformation on the imitation action graph F relative to the reference action graph I makes the human body in the imitation action graph F and the human body in the reference action graph I in the same orientation, thereby effectively preventing the matching evaluation of the imitation action and the reference action from being affected by the different angle positions of the human body in the imitation action graph F and the human body in the reference action graph I; at the same time, when calculating the action skeleton similarity of the imitator action graph F relative to the reference action graph I, the similarity distance between the skeleton feature vectors corresponding to the reference action graph I and the human standing graph J is considered as the weight for summing up, thereby taking into account the key position points of the human body limb movement in the reference action graph, and when evaluating the action skeleton similarity of the imitator action graph F relative to the reference action graph I, these key position points are given a larger weight relative to the non-moving parts of the human body, which is more conducive to correctly evaluating the similarity of the imitation action. BRIEF DESCRIPTION OF THE DRAWINGS

[0033] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative labor.

[0034] Figure 1 A flowchart of a human motion similarity matching and scoring method according to an embodiment of the present application;

[0035] Figure 2 A schematic diagram of 17 human key points and 19 skeleton feature vectors extracted from a human body standing image according to an embodiment of the present application;

[0036] Figure 3 This is a flow chart of the steps of obtaining the action skeleton similarity of the imitation action graph F relative to the reference action graph I in an embodiment of the present application. DETAILED DESCRIPTION

[0037] In order to make the technical problems, technical solutions and beneficial effects to be solved by the present application more clearly understood, the present application is further described in detail below in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.

[0038] It should be noted that when an element is referred to as being "fixed to" or "disposed on" another element, it can be directly on the other element or indirectly on the other element. When an element is referred to as being "connected to" another element, it can be directly connected to the other element or indirectly connected to the other element.

[0039] It should be understood that the terms "length", "width", "up", "down", "front", "back", "left", "right", "vertical", "horizontal", "top", "bottom", "inside", "outside", etc., indicating the orientation or position relationship, are based on the orientation or position relationship shown in the drawings, and are only for the convenience of describing the present application and simplifying the description, and do not indicate or imply that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and therefore should not be understood as a limitation on the present application.

[0040] In addition, the terms "first" and "second" are used for descriptive purposes only and should not be understood as indicating or implying relative importance or implicitly indicating the number of the indicated technical features. Thus, a feature defined as "first" or "second" may explicitly or implicitly include one or more of the features. In the description of this application, the meaning of "plurality" is two or more, unless otherwise clearly and specifically defined.

[0041] The human body motion similarity matching scoring method, device and readable storage medium of the present application are now described in detail with reference to the accompanying drawings.

[0042] Please refer to Figure 1 The human body motion similarity matching scoring method provided in the embodiment of the present application is used to evaluate and judge the similarity of the imitator's imitating motion relative to the reference motion, and comprises the following steps:

[0043] Step S100: Obtain the imitation action diagram F, the reference action diagram I and the human body standing diagram J.

[0044] Step S200: Detect multiple human key points representing the human skeleton in the imitation action graph F, the reference action graph I and the human standing graph J, wherein the multiple human key point sets of the imitation action graph F, the reference action graph I and the human standing graph J are respectively denoted as V F 、V I and V J .

[0045] Step S300: performing similarity transformation on the imitation action graph F according to the multiple human body key points of the reference action graph I to obtain multiple new human body key points of the imitation action graph F, wherein the new set of multiple human body key points of the imitation action graph F is denoted as V F’ .

[0046] Step S400: According to V F’ 、V I and V J A plurality of skeleton feature vectors representing the human skeleton of the imitation action graph F, the reference action graph I and the human standing graph J are obtained.

[0047] Step S500: Taking the similarity distance between the skeleton feature vectors corresponding to the reference action graph I and the human body standing graph J as the weight, perform weighted summation on the similarity distance between the skeleton feature vectors corresponding to the imitation action graph F and the reference action graph I to obtain the action skeleton similarity of the imitation action graph F relative to the reference action graph I.

[0048] In step S200, the step of detecting a plurality of key points of the human body representing the human skeleton of any image in the imitation action graph F, the reference action graph I and the human body standing graph J includes: using the YOLOX detection model to detect the human body region, and then using HRNet to detect 17 key points of the human body representing the human skeleton in the human body region. It is worth noting that the 17 key points of the human body representing the human skeleton include the nose tip A v0 , Left Eye A v1 , right eye A v2 , left ear A v3 , right ear Av4 , Left Shoulder A v5 , right shoulder A v6 , left elbow A v7 , right elbow A v8 , left wrist A v9 , right wrist A v10 , Left thigh root A v11 , right thigh root A v12 , left knee A v13 , right knee A v14 , Left ankle A v15 and right ankle A v16 ; Wherein, A is any image among the imitation action graph F, the reference action graph I and the human body standing graph J, A∈(F, J, I), that is, A belongs to any one of the images F, J and I.

[0049] like Figure 2 As shown, Figure 2 The figure shows 17 key points of the human body representing the human skeleton in a three-dimensional action diagram, among which: Figure 2 The positions from 0 to 16 in the figure represent the nose tip I of the three-dimensional action diagram respectively. v0 ,Left Eye I v1 , right eye I v2 , left ear I v3 、Right Ear I v4 , Left Shoulder I v5 、Right Shoulder I v6 , left elbow I v7 、Right elbow I v8 , left wrist I v9 , right wrist I v10 , Left thigh root I v11 , right thigh root I v12 , left knee I v13 、Right knee I v14 , left ankle I v15 and right ankle I v16 . The 17 key points of the human body representing the human skeleton detected in the imitation action graph F and the reference action graph I are compared with Figure 2 The positions of the human skeleton points indicated from 0 to 16 are consistent and will not be repeated here.

[0050] In the step S300, the steps of performing similarity transformation on the imitation action graph F according to the multiple human body key points of the reference action graph I to obtain multiple new human body key points of the imitation action graph F include:

[0051] Step S310: According to V of the reference action diagram I F The 17 key points of the human body and the V of the simulated action graph F FThe similarity transformation matrix H between the reference action graph I and the imitation action graph F is calculated based on the 17 human key points in the reference action graph I and the imitation action graph F.

[0052] Step S320: The V of the imitation action diagram F is F Each human key point in the image is transformed by similarity transformation matrix H to obtain 17 new human key points of the imitation action graph F.

[0053] It can be understood that in the step S300, the imitation action graph F is similarly transformed relative to the reference action graph, that is, the imitation action graph F is rotated, translated and scaled so that the human body in the imitation action graph F is in the same position as the human body in the imitation action graph F. This effectively prevents the problem of different angles and positions of the human body in the imitation action graph F and the human body in the reference action graph I from affecting the matching score results of the imitation action and the reference action.

[0054] In the step S400, the multiple skeleton feature vectors representing the human skeleton of any image A in the imitation action graph F, the reference action graph I and the human standing graph J are obtained to obtain 19 skeleton feature vectors representing the human skeleton, wherein the 19 skeleton feature vectors of any image A are sequentially from nose tip A to nose tip B. v0 and left eye A v1 The eigenvector P A0 , from the nose tip A v0 and right eye A v2 The eigenvector P A1 , from left eye A v1 and right eye A v2 The eigenvector P A2 , from left eye A v1 and left ear A v3 The eigenvector P A3 , by right eye A v2 and right ear A v4 The eigenvector P A4 , from the nose tip A v0 and left shoulder A v5 The eigenvector P A5 , from the nose tip A v0 and right shoulder A v6 The eigenvector P A6 , from left shoulder A v5 and left elbow A v7 The eigenvector P A7 , from right shoulder A v6 and right elbow A v8 The eigenvector P A8 , from left elbow A v7 and left wrist A v9 The eigenvector PA9 , by the right elbow A v8 and right wrist Av 10 The eigenvector P A10 , by the left shoulder Av5 and the left thigh root A v11 The eigenvector P A11 , from right shoulder A v6 and right thigh A v12 The eigenvector P A12 , from left shoulder A v5 and right shoulder A v6 The eigenvector P A13 , from the root of the left thigh A v11 and right thigh A v12 The eigenvector P A14 , from the root of the left thigh A v11 and left knee A v13 The eigenvector P A15 , from the right thigh A v12 and right knee A v14 The eigenvector P A16 , from the left knee A v13 and left ankle A v15 The eigenvector P A17 and by the right knee A v14 and right ankle A v16 The eigenvector P A18 Similarly, Figure 2 The lines connecting points 0-16 in the figure indicate the human body standing in Figure I. I0 To P I18 19 skeleton feature vectors.

[0055] It is worth noting that the eigenvector P A0 There is a key point nose tip A v0 The coordinates and key points of the left eye A v1 The difference in coordinates of, for example, the nose tip A v0 The coordinates are (x1, y1), left eye A v1 The coordinates of are (x2,y2), then P A0 =(x2-x1,y2-y2), and so on, the calculation of other eigenvectors is the same.

[0056] Reference Figure 3 In step S500, the similarity distance between the skeleton feature vectors corresponding to the reference action graph I and the human body standing graph J is used as a weight to perform weighted summation on the similarity distance between the skeleton feature vectors corresponding to the imitation action graph F and the reference action graph I, and the step of obtaining the action skeleton similarity of the imitation action graph F relative to the reference action graph I includes:

[0057] Step S510: Calculate the skeleton feature vector P of the reference action graph I Ii The skeleton feature vector P corresponding to the human body standing figure J Ji The similarity distance W between i :W i =1-cos <P Ii ·P Ji >=1-P Ii ·P Ji / |P Ii |·|P Ji |;

[0058] Step S520: For all similarity distances W i Perform normalization operation, similarity distance W i The normalized result is recorded as

[0059] Step S530: Calculate the skeleton feature vector P of the imitator's action graph F Fi The skeleton feature vector P corresponding to the reference action graph I Ii The similarity distance D i :D i =1-cos <P Fi ·P Ii >=1-P Fi ·P Ii / |P Fi |·|P Ii |;

[0060] Step S540: Using all the similarity distances W i is the weight for all similarity distances D i Perform weighted summation to obtain the action skeleton similarity M of the imitator action graph F relative to the reference action graph I:

[0061] Among them, P Ii is the i-th eigenvector of the reference action graph I, P Fi is the i-th feature vector of the imitation action graph F, P Ji is the i-th eigenvector of the human body stereogram J, i∈[0,18].

[0062] P Fi |、|P Ii | and |P Ji |are the modulus of the i-th eigenvector of the imitation action graph F, the modulus of the i-th eigenvector of the reference action graph I, and the modulus of the i-th eigenvector of the human body standing graph J, respectively.

[0063] In this embodiment, when the human body in the reference action graph I performs a corresponding reference action, some of the 19 skeleton feature vectors in the reference action graph I and the 19 skeleton feature vectors of the human body in the human body standing graph J will have some corresponding skeleton feature vectors that are different. For example, when the left hand of the person in the reference action graph I performs a motion gesture, the skeleton feature vector P on the left hand of the person in the reference action graph I is different. I7 , P I9 and the skeleton feature vector P on the left hand of the person in the human figure J J7 , P J9 There are differences. It can be understood that when a skeleton feature vector P in the reference action graph I Ii P corresponding to the human body in the human body standing figure J Ji When there is a difference in the skeleton feature vector, the larger the difference, the larger the value of , which means that the skeleton feature vector P in the reference action graph I Ii It is the most critical comparison feature when imitating. When evaluating the action matching of the imitator imitating the reference action, the skeleton feature vector P Ii More attention should be paid to it. Therefore, with all the similarity distances W i is the weight for all similarity distances D i A weighted summation is performed to obtain the skeleton feature vector similarity M of the imitation action graph F relative to the reference action graph I, thereby fully considering the importance of key points in the imitation action, that is, the importance of the moving limbs relative to the non-moving limbs in evaluating the similarity of human body movements, so as to more accurately and effectively evaluate the similarity of the imitation action.

[0064] Further references Figure 1 In this embodiment, the human body action similarity matching scoring method further includes step S600: scoring the similarity of the imitation action graph F according to the action skeleton similarity of the imitation action graph F relative to the reference action graph I:

[0065] S=(dt-M)*(100-st) / dt+st-neg_score,

[0066] Among them, S represents the scoring result of the similarity of the imitation action graph F; dt is used to control the matching scoring standard, and the smaller dt is, the higher the matching scoring standard is; st is used to control the minimum score, which can make the final score of S controlled within the range of 0 to 100; neg_score is the action reverse deduction item, which is used to punish the situation where the human body movements of the imitation action graph F are far different from the human body movements of the reference action graph I.

[0067] In this embodiment, dt takes a value of 2, and st takes a value of 50; when M is less than 0.4, neg_score takes a value of 1, when M is greater than or equal to 0.4 and less than or equal to 0.6, neg_score takes a value of 2; when M is greater than 0.6 and less than or equal to 0.8, M takes a value of 4, when M is greater than 0.8 and less than or equal to 1.0, neg_score takes a value of 6, and when M is greater than 1, neg_score takes a value of 8. It can be understood that the greater the similarity of the skeleton feature vector of the imitation action graph F relative to the reference action graph I, the higher the score S, which means that the imitator imitates more accurately.

[0068] In the human body motion similarity matching scoring method, the imitation action graph F is subjected to a similarity transformation relative to the reference action graph I so that the human body in the imitation action graph F and the human body in the reference action graph I are in the same position, thereby effectively preventing the matching evaluation of the imitation action and the reference action from being affected by the different angle positions of the human body in the imitation action graph F and the human body in the reference action graph I; at the same time, when calculating the action skeleton similarity of the imitator action graph F relative to the reference action graph I, the similarity distance between the skeleton feature vectors corresponding to the reference action graph I and the human body standing graph J is considered as the weight for summing, thereby taking into account the key position points of the human body limb movement in the reference action graph, and when evaluating the action skeleton similarity of the imitator action graph F relative to the reference action graph I, these key position points are given a larger weight relative to the non-moving parts of the human body, which is more conducive to correctly evaluating the similarity of the imitation action.

[0069] The embodiment of the present application also provides a human motion similarity matching scoring device using the above human motion similarity matching scoring method, which includes an image acquisition unit, an image processing unit and a voice announcer. The image processing unit is connected to the image acquisition unit and the voice announcer. The image acquisition unit is used to acquire the imitation action diagram F, the reference action diagram I and the human standing diagram J; the image processing unit is used to execute the steps of the human motion similarity matching scoring method; when the image processing unit completes the action similarity matching scoring of the imitation action diagram F relative to the reference action diagram I, the voice announcer is used to announce the score.

[0070] An embodiment of the present application further provides a readable storage medium, wherein the readable storage medium stores a computer program, and when the computer program is executed by a processor, the human body motion similarity matching scoring method is implemented.

[0071] The above description is only a preferred embodiment of the present application and is not intended to limit the present application. Any modifications, equivalent substitutions and improvements made within the spirit and principles of the present application should be included in the protection scope of the present application.

Claims

1. A human motion similarity matching scoring method, characterized in that: include: Obtain the imitator's action diagram F, the reference action diagram I and the human body standing diagram J; Detect multiple human key points representing the human skeleton in the imitator action graph F, the reference action graph I and the human standing graph J, wherein the multiple human key point sets of the imitator action graph F, the reference action graph I and the human standing graph J are respectively denoted as V F 、V I and V J ; According to the multiple human body key points of the reference action graph I, the imitator action graph F is similarly transformed to obtain a new multiple human body key points of the imitator action graph F. The new multiple human body key point set of the imitator action graph F is recorded as V F’ ; According to V F’ 、V I and V J Obtaining a plurality of skeleton feature vectors representing the human skeleton of the imitator action graph F, the reference action graph I and the human standing graph J; as well as Taking the similarity distance between the skeleton feature vectors corresponding to the reference action graph I and the human body standing graph J as the weight, weighted summation is performed on the similarity distance between the skeleton feature vectors corresponding to the imitator action graph F and the reference action graph I to obtain the action skeleton similarity of the imitator action graph F relative to the reference action graph I; The steps of weighting and summing the similarity distances between the skeleton feature vectors corresponding to the imitator's action graph F and the reference action graph I by taking the similarity distance between the skeleton feature vectors corresponding to the reference action graph I and the human body standing graph J as weights, and obtaining the action skeleton similarity of the imitator's action graph F relative to the reference action graph I include: Calculate the skeleton feature vector P of the reference action graph I Ii The skeleton feature vector P corresponding to the human body standing figure J Ji The similarity distance W between i :W i =1-cos <P Ii ·P Ji >=1-P Ii ·P Ji / |P Ii |·|P Ji |; For all similarity distances W i Perform normalization operation, similarity distance W i The normalized result is recorded as Calculate the skeleton feature vector P of the imitator action graph F Fi The skeleton feature vector P corresponding to the reference action graph I Ii The similarity distance D i :D i =1-cos <P Fi ·P Ii >=1-P Fi ·P Ii / |P Fi |·|P Ii |; With all the similarity distances W i is the weight for all similarity distances D i Perform weighted summation to obtain the action skeleton similarity M of the imitator action graph F relative to the reference action graph I: Among them, P Ii is the i-th eigenvector of the reference action graph I, P Fi is the i-th feature vector of the imitator action graph F, P Ji Determine the i-th eigenvector of the human body graph J, i∈[0,18].

2. The human motion similarity matching scoring method according to claim 1, characterized in that: The step of detecting a plurality of human key points representing the human skeleton of any one of the imitator action graph F, the reference action graph I and the human standing graph J comprises: using a YOLOX detection model to detect a human region, and then using HRNet to detect 17 human key points representing the human skeleton in the human region; Among them, the 17 key points of the human body representing the human skeleton include the nose tip A v0 , Left Eye A v1 , right eye A v2 , left ear A v3 , right ear A v4 , Left Shoulder A v5 , right shoulder A v6 , left elbow A v7 , right elbow A v8 , left wrist A v9 , right wrist A v10 , Left thigh root A v11 , right thigh root A v12 , left knee A v13 , right knee A v14 , Left ankle A v15 and right ankle A v16 ; A is any image among the imitator action graph F, the reference action graph I and the human body standing graph J, A∈(F, J, I).

3. The human motion similarity matching scoring method according to claim 2, characterized in that: The steps of performing similarity transformation on the imitator action graph F according to the multiple human body key points of the reference action graph I to obtain multiple new human body key points of the imitator action graph F include: According to the reference action diagram I, V F The 17 key points of the human body and the V of the imitator action graph F F Calculate the similarity transformation matrix H between the reference action graph I and the imitator action graph F based on the 17 human key points in the reference action graph I; The imitator moves the V of Figure F F Each human key point in the similarity transformation matrix H is subjected to a similarity transformation to obtain 17 new human key points of the imitator's action graph F.

4. The human motion similarity matching scoring method according to claim 2, characterized in that: Obtaining a plurality of skeleton feature vectors representing the human skeleton of any image A in the imitator action graph F, the reference action graph I and the human body standing graph J is to obtain 19 skeleton feature vectors representing the human skeleton, The 19 skeleton feature vectors of any image A are sequentially composed of nose tip A v0 and left eye A v1 The eigenvector P A0 , from the nose tip A v0 and right eye A v2 The eigenvector P A1 , from left eye A v1 and right eye A v2 The eigenvector P A2 , from left eye A v1 and left ear A v3 The eigenvector P A3 , by right eye A v2 and right ear A v4 The eigenvector P A4 , from the nose tip A v0 and left shoulder A v5 The eigenvector P A5 , from the nose tip A v0 and right shoulder A v6 The eigenvector P A6 , from left shoulder A v5 and left elbow A v7 The eigenvector P A7 , from right shoulder A v6 and right elbow A v8 The eigenvector P A8 , from left elbow A v7 and left wrist A v9 The eigenvector P A9 , by the right elbow A v8 and right wrist Av 10 The eigenvector P A10 , by the left shoulder Av5 and the left thigh root A v11 The eigenvector P A11 , from right shoulder A v6 and right thigh A v12 The eigenvector P A12 , from left shoulder A v5 and right shoulder A v6 The eigenvector P A13 , from the root of the left thigh A v11 and right thigh A v12 The eigenvector P A14 , from the left thigh A v11 and left knee A v13 The eigenvector P A15 , from the right thigh A v12 and right knee A v14 The eigenvector P A16 , from the left knee A v13 and left ankle A v15 The eigenvector P A17 and by the right knee A v14 and right ankle A v16 The eigenvector P A18 .

5. The human motion similarity matching scoring method according to claim 4, characterized in that: The human body action similarity matching scoring method further includes scoring the similarity of the imitator action graph F according to the action skeleton similarity of the imitator action graph F relative to the reference action graph I, S=(dt-M)*(100-st) / dt+st-neg_score, Among them, S represents the scoring result of the similarity of the imitator's action graph F; dt is used to control the matching scoring standard, and the smaller dt is, the higher the matching scoring standard is; st is used to control the minimum score, which can make the final score of S controlled within the range of 0 to 100; neg_score is the action reverse penalty item, which is used to punish the situation where the human body movements of the imitator's action graph F are far different from the human body movements of the reference action graph I.

6. The human motion similarity matching scoring method according to claim 5, characterized in that: The value of dt is 2, and the value of st is 50; when M is less than 0.4, the value of neg_score is 1, when M is greater than or equal to 0.4 and less than or equal to 0.6, the value of neg_score is 2; when M is greater than 0.6 and less than or equal to 0.8, the value of M is 4, when M is greater than 0.8 and less than or equal to 1.0, the value of neg_score is 6, and when M is greater than 1, the value of neg_score is 8.

7. A human motion similarity matching and scoring device, characterized in that: include: Image acquisition unit: used to acquire the imitator's action image F, the reference action image I and the human body standing image J; An image processing unit: connected to the image acquisition unit and used to execute the steps of the human motion similarity matching and scoring method according to any one of claims 1 to 6; as well as, Voice announcer: connected to the image processing unit, when the image processing unit completes the matching and scoring of the action similarity of the imitator's action graph F relative to the reference action graph I, the voice announcer is used to announce the score.

8. A readable storage medium, characterized in that: The readable storage medium stores a computer program, and when the computer program is executed by a processor, the human body motion similarity matching and scoring method according to any one of claims 1 to 6 is implemented.

Citation Information

Patent Citations

  • Camera-based human body action scoring method and system

    CN109829442A

  • Human body action contrastive analysis method based on image retrieval

    CN111046715A