Taijiquan action scoring method and device

By identifying and extracting the spatiotemporal and spatial feature vector sequences of human target parts in Tai Chi action videos and comparing them with the preset sequences, the problem of insufficient scoring accuracy in the prior art is solved, and a higher precision Tai Chi action score is achieved.

CN119992131APending Publication Date: 2025-05-13ZHUMADIAN PRESCHOOL TEACHERS COLLEGE

Patent Information

Application Number
CN202510057919.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-14
Publication Date
2025-05-13

AI Technical Summary

Technical Problem

The prior art cannot identify the subtle differences between the movement characteristics of human body parts under different force modes in Tai Chi movements, resulting in insufficient scoring accuracy.

Method used

By identifying the human movement moves in the video to be scored, the real-time spatio-temporal eigenvector sequence of the human target part in the move video is extracted, and the preset spatio-temporal eigenvector sequence is compared with the preset spatio-temporal eigenvector sequence is scored according to the similarity.

Benefits of technology

It improves the accuracy of Tai Chi movement scores, can more accurately identify and evaluate subtle differences in martial arts movements, and provides referees with a more reliable reference.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119992131A_ABST
    Figure CN119992131A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of action recognition, in particular to a shadowboxing action scoring method and device, and the method comprises the steps: recognizing one or more types formed by human body actions in a to-be-scored video based on the to-be-scored video containing shadowboxing actions, and carrying out the scoring of the shadowboxing actions according to the types, segmenting the to-be-scored video into a plurality of one or more portfolio videos corresponding to the portfolio; extracting a real-time spatio-temporal feature vector sequence about the force exerting mode of the target part of the human body in the portfolio video; comparing the real-time spatio-temporal feature vector sequence with a preset spatio-temporal feature vector sequence of the recruitment, and according to the similarity between the real-time spatio-temporal feature vector sequence and the preset spatio-temporal feature vector sequence, scoring the shadowboxing actions in the to-be-scored video; the spatial-temporal feature vector sequence comprises a time sequence of linkage features among a plurality of key nodes included in the human body target part based on the shown video. According to the method, the accuracy degree of shadowboxing action scoring can be improved, and more accurate reference is provided for martial arts action referees.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of action recognition, and in particular to a Tai Chi action scoring method and device. Background Art

[0002] Traditional martial arts is a traditional sport with a long history. It has a unique training effect of integrating fitness, self-defense and health preservation. It is a good fitness and lifelong sports. In the long-term evolution process, it has formed certain routines. Modern competitive sports have also introduced traditional martial arts related competitions. Martial arts action referees score according to the technical level of the athletes on the spot and the action specifications of each competition event. However, this scoring method relies on the experience of martial arts action referees and is greatly influenced by their subjective concepts.

[0003] In response to the above problems, the prior art provides a martial arts action recognition method based on deep learning, which analyzes the compliance of the action by extracting the motion characteristics of each major joint of the human body. For example, the patent document with application number CN202210052217.1 discloses a martial arts action recognition method based on human posture estimation, comprising the following steps: obtaining a standard teaching video; extracting a short video from the teaching video according to the martial arts action to obtain a single-frame character action picture; sending the single-frame character action picture to a human posture recognition network to obtain joint point data; standardizing and normalizing the joint point data; classifying the processed joint point data, saving and displaying the joint point data according to the classification results; the human posture recognition network is simpler to deploy in practice than the action recognition network, can effectively identify the actions performed by the characters in the video, can effectively assist the referee in scoring the martial arts actions, improve the accuracy of the scoring, and make the scoring more based on evidence. However, this recognition method simply displays the position and connection relationship of each joint of the human body. In the scoring scenarios of martial arts actions such as Tai Chi that emphasize force patterns, it can only identify whether the position of human body parts is compliant, but cannot identify the subtle differences between the movement characteristics of human body parts brought about by different force patterns, thus limiting the accuracy of the scoring. Summary of the invention

[0004] 1. Technical issues to be resolved

[0005] In view of the above-mentioned shortcomings and deficiencies of the prior art, the present invention provides a Tai Chi movement scoring method and device, which solves the technical problem that the prior art cannot recognize the subtle differences between the movement characteristics of human body parts brought about by different force modes of the human body, thereby limiting the accuracy of scoring.

[0006] (II) Technical solution

[0007] In order to achieve the above object, the main technical solutions adopted by the present invention include:

[0008] In a first aspect, an embodiment of the present invention provides a Tai Chi action scoring method, comprising:

[0009] Based on the video to be rated containing Tai Chi movements, one or more moves formed by human body movements in the video to be rated are identified, and according to the moves, the video to be rated is divided into a plurality of one or more move videos corresponding to the moves;

[0010] Extracting a real-time spatiotemporal feature vector sequence of a force pattern of a target part of a human body in the move video;

[0011] Comparing the real-time spatiotemporal feature vector sequence with the preset spatiotemporal feature vector sequence of the move, and scoring the Tai Chi moves in the video to be scored according to the similarity between the two;

[0012] The spatiotemporal feature vector sequence includes: linkage features between multiple key nodes included in the target part of the human body are based on the time sequence of the shown move video.

[0013] Optionally, extracting a real-time spatiotemporal feature vector sequence about a force pattern of a target part of a human body in the move video includes:

[0014] The first action recognition model is used to identify the positions of key nodes contained in the target part of the human body in a frame image of a move video, and the characteristic values ​​of the linkage features between the key nodes are determined according to the positions of the key nodes to form a real-time spatiotemporal feature vector. Based on the real-time spatiotemporal feature vector corresponding to each frame image in the move video, a real-time spatiotemporal feature vector sequence is formed.

[0015] Optionally, extracting a real-time spatiotemporal feature vector of a force pattern of a target part of a human body in the move video includes:

[0016] The target parts of the human body include: one or more of the wrist, waist, spine or ankles and knees.

[0017] Optionally, when the human body target part is the wrist, the key nodes included in the human body target part include: shoulder joint, elbow joint and wrist joint; in a frame of the move video, the position of the shoulder joint recognized by the first action recognition model is (x S ,y S ), the position of the elbow joint is (x E ,y E ), the position of the wrist joint is (x W ,y W );

[0018] The method of determining the characteristic values ​​of the linkage characteristics between the key nodes according to the positions of the key nodes to form a real-time spatiotemporal characteristic vector includes:

[0019] The real-time spatiotemporal feature vector of the wrist is expressed as:

[0020]

[0021] in, Indicates the spatial relative position characteristics of the shoulder and elbow; Indicates the spatial relative position characteristics of the elbow and wrist; Indicates the spatial relative position characteristics of the shoulder, elbow and wrist; Indicates the length constraint relationship of the shoulder, elbow and wrist; represents the rate of change of the relative angle between the shoulder, elbow and wrist, θ SEW (t) represents the relative angle between the shoulder, elbow and wrist in the current frame image, θ SEW (t-1) represents the relative angle between the shoulder, elbow and wrist in the previous frame image, Δt represents the time interval between two adjacent frames of images, v S 、v E 、v W Respectively represent the movement speed of the shoulder joint, elbow joint and wrist joint; a S 、a E 、a W Represent the movement acceleration of the shoulder joint, elbow joint and wrist joint respectively.

[0022] Optionally, when the human target part is the ankle and knee, the key nodes included in the human target part include: hip joint, knee joint and ankle joint; in a frame of the move video, the position of the hip joint recognized by the first action recognition model is (x P ,y P ), the position of the knee joint is (x K ,y K ), the position of the ankle joint is (x A ,y A );

[0023] The method of determining the characteristic values ​​of the linkage characteristics between the key nodes according to the positions of the key nodes to form a real-time spatiotemporal characteristic vector includes:

[0024] The real-time spatiotemporal feature vector of the ankle and knee is expressed as:

[0025]

[0026] in, Indicates the spatial relative position characteristics of the hip and knee; Indicates the spatial relative position characteristics of the knee and ankle; Indicates the spatial relative position characteristics of the hip, knee and ankle; Indicates the length constraint relationship of the hip, knee and ankle; represents the rate of change of the relative angle between the hip, knee and ankle, θ PKA (t) represents the relative angle between the hip, knee and ankle in the current frame image, θ PKA (t-1) represents the relative angle between the hip, knee and ankle in the previous frame image, Δt represents the time interval between two adjacent frames of images, v P 、v K 、v A Respectively represent the movement speed of the hip joint, knee joint and ankle joint; a P 、a K 、a A Represent the movement acceleration of the hip joint, knee joint and ankle joint respectively.

[0027] Optionally, when the human body target part is the waist and spine, the key nodes included in the human body target part include: left shoulder joint, right shoulder joint, left hip joint and right hip joint; in a frame image of the move video, the position of the left shoulder joint recognized by the first action recognition model is (x S1 ,y S1 ), the position of the right shoulder joint is (x S2 ,y S2 ), the position of the left hip joint is (x P1 ,y P1 ), the position of the right hip joint is (x P2 ,y P2 );

[0028] The method of determining the characteristic values ​​of the linkage characteristics between the key nodes according to the positions of the key nodes to form a real-time spatiotemporal characteristic vector includes:

[0029] The real-time spatiotemporal feature vector of the lumbar spine is expressed as:

[0030]

[0031] in, Indicates the rotational characteristics of the shoulder; Indicates the rotational characteristics of the hip; Indicates the spatial relative position characteristics between the left shoulder and the right hip; Indicates the spatial relative position characteristics between the right shoulder and the left hip; It represents the first angle between shoulder and hip; It represents the second angle between shoulder and hip; The speed of change of the first angle between the shoulder and the hip; The speed of change of the second angle between the shoulder and the hip; The acceleration of the change of the first angle between the shoulder and the hip; The acceleration of the change of the second angle between the shoulder and the hip.

[0032] Optionally, comparing the real-time spatiotemporal feature vector sequence with a preset spatiotemporal feature vector sequence of the move includes:

[0033] The preset spatiotemporal feature vector is: a spatiotemporal feature vector sequence obtained in advance based on a standard action video of the move.

[0034] Optionally, the real-time spatiotemporal feature vector sequence is compared with a preset spatiotemporal feature vector sequence of the move, and the Tai Chi moves are scored according to the similarity between the two, including:

[0035] A1. Extracting feature values ​​of the same feature in the real-time spatiotemporal feature vector sequence to form a first time series; extracting feature values ​​corresponding to the same feature in the preset spatiotemporal feature vector sequence to form a second time series;

[0036] A2. Calculate the DTW distance between the first time series and the second time series to obtain the DTW distance corresponding to each feature;

[0037] A3. Perform weighted summation of the DTW distances corresponding to all features to obtain the distance sum. Based on the distance sum, score the Tai Chi movements in the video to be scored.

[0038] Optionally, identifying one or more moves formed by human body movements in the video to be rated includes:

[0039] Using the second action recognition model to recognize one or more moves formed by human body movements in the video to be rated, and obtaining the name of each move and the starting time point and the ending time point of the move in the video with rating;

[0040] The first action recognition model or the second action recognition model is:

[0041] Pre-trained PoseNet model with adapted model parameters.

[0042] In a second aspect, an embodiment of the present invention provides a Tai Chi action scoring device, comprising:

[0043] A decomposition module is used to identify one or more moves formed by human body movements in the video to be rated based on the video to be rated containing Tai Chi movements, and according to the moves, divide the video to be rated into a plurality of one or more move videos corresponding to the moves;

[0044] A feature extraction module, used to extract a real-time spatiotemporal feature vector sequence about the force pattern of the target part of the human body in the move video;

[0045] A scoring module is used to compare the real-time spatiotemporal feature vector sequence with the preset spatiotemporal feature vector sequence of the move, and score the Tai Chi moves in the video to be scored according to the similarity between the two;

[0046] The spatiotemporal feature vector sequence includes: linkage features between multiple key nodes included in the target part of the human body are based on the time sequence of the shown move video.

[0047] (III) Beneficial effects

[0048] The Tai Chi action scoring method proposed in the present invention is based on a video to be scored containing Tai Chi actions, identifies one or more moves formed by human body movements in the video to be scored, and according to the moves, divides the video to be scored into multiple one or more move videos corresponding to the moves; extracts a real-time spatiotemporal feature vector sequence about the force mode of a target part of the human body in the move video; compares the real-time spatiotemporal feature vector sequence with a preset spatiotemporal feature vector sequence of the move, and scores the Tai Chi actions in the video to be scored according to the similarity between the two; the spatiotemporal feature vector sequence includes: linkage features between multiple key nodes included in the target part of the human body are based on the time sequence of the move video shown.

[0049] Based on the above scoring method, it is possible to extract features about the force pattern of the target part of the human body based on the spatiotemporal feature vector sequence containing the linkage features between key nodes, compare the real-time spatiotemporal feature vector sequence with the preset spatiotemporal feature vector sequence of the move, and score the Tai Chi movements in the video to be scored based on the similarity between the two, thereby improving the accuracy of Tai Chi movement scoring and providing a more accurate reference for martial arts action referees. BRIEF DESCRIPTION OF THE DRAWINGS

[0050] Figure 1 A schematic diagram of a process flow of a Tai Chi action scoring method provided in an embodiment;

[0051] Figure 2 A schematic diagram of key nodes of a human body part provided in an embodiment;

[0052] Figure 3 Schematic diagram of the architecture of a Tai Chi action scoring device provided in an embodiment. DETAILED DESCRIPTION

[0053] In order to better explain the present invention and facilitate understanding, the present invention is described in detail below through specific implementation modes in conjunction with the accompanying drawings.

[0054] The force generation mode of Tai Chi movements emphasizes the overall force generation of the whole body, rather than relying solely on the strength of local muscles. It requires starting from the feet, passing through the legs, waist, spine, and finally transmitting to the arms and other parts. For example, in Tai Chi's force generation movements such as "kicking" or "releasing", the force surges from the Yongquan point on the sole of the foot like a wave, passes through the muscles of the calves and thighs, and transmits the force to the upper limbs through the rotation and penetration of the waist and spine, and finally through the palms. This overall force generation method can make the force more complete and strong, avoiding the thinness and stiffness of local force. In addition, Tai Chi's force generation has a spiral characteristic. In the movement, the limbs often have a rotating movement to generate spiral force. For example, in the squeezing force of "Lan Que Wei", the arm is accompanied by internal rotation in the process of squeezing forward. This spiral force generation method can increase the penetration and stability of the force. Just like a screw rotating and drilling into an object, the spiral force can make the force act better on the opponent and is difficult to resist. In view of the above-mentioned special force pattern of Tai Chi, the present invention specifically constructs the spatiotemporal feature vectors of the force pattern of the target part of the human body, and scores the moves in the video to be scored by comparing the time series composed of the spatiotemporal feature vectors in the video to be scored with the spatiotemporal feature vectors of the standard action video, thereby improving the accuracy of Tai Chi action scoring and providing a more accurate reference for martial arts action referees.

[0055] In order to better explain the present invention, so as to facilitate understanding, exemplary embodiments of the present invention will be described in more detail below with reference to the accompanying drawings. Although exemplary embodiments of the present invention are shown in the accompanying drawings, it should be understood that the present invention can be implemented in various forms and should not be limited by the embodiments set forth herein. On the contrary, these embodiments are provided in order to enable a clearer and more thorough understanding of the present invention, and to be able to fully convey the scope of the present invention to those skilled in the art.

[0056] Embodiment 1

[0057] like Figure 1 As shown, the embodiment of the present invention provides a Tai Chi action scoring method, comprising:

[0058] S1. Based on a video to be rated that contains Tai Chi movements, one or more moves formed by human body movements in the video to be rated are identified, and according to the moves, the video to be rated is divided into a plurality of one or more move videos corresponding to the moves.

[0059] Specifically, depending on the school or routine, the Tai Chi movements may include: starting position, wild horse spreading its mane, white crane spreading its wings, embracing the peacock's tail, cloud hands, single whip, high horse exploring, right kick, double peaks piercing the ears, etc.

[0060] S2. extracting a real-time spatiotemporal feature vector sequence of the force pattern of the target part of the human body in the move video.

[0061] Specifically, the human target part may include: one or more of the wrist, waist, spine, ankle, and knee. The first action recognition model may be used to identify the position of the key nodes contained in the human target part in a frame image of the move video, and according to the position of the key nodes, the characteristic values ​​of the linkage characteristics between the key nodes are determined to form a real-time spatiotemporal feature vector, and based on the real-time spatiotemporal feature vector corresponding to each frame image in the move video, a real-time spatiotemporal feature vector sequence is formed.

[0062] Among them, for a move, the corresponding human target part may be one or more. For example, for the "White Crane Spreads Its Wings" move involving whole-body movement, the corresponding human target parts include wrists, waist, spine, ankles and knees; for another example, for the "Starting Position" move mainly involving leg movement, the corresponding human target parts only include ankles and knees.

[0063] S3. Compare the real-time spatiotemporal feature vector sequence with the preset spatiotemporal feature vector sequence of the move, and score the Tai Chi moves in the video to be scored according to the similarity between the two.

[0064] Specifically, the similarity between a real-time spatiotemporal feature vector sequence of a move and a preset spatiotemporal feature vector sequence can be calculated to determine the similarity between the two. The smaller the DTW distance, the higher the similarity between the two, and the higher the corresponding score.

[0065] In steps S2 and S3, the spatiotemporal feature vector sequence includes: linkage features between multiple key nodes included in the target part of the human body are based on the time sequence of the shown move video.

[0066] Preferably, the preset spatiotemporal feature vector is: a spatiotemporal feature vector sequence obtained in advance based on a standard action video of the move.

[0067] The method of obtaining the preset spatiotemporal feature vector sequence is substantially the same as that in step S2, except that the move video used is a standard action video. In addition, for one move, multiple standard action videos can be obtained, which are adjusted to the same duration or number of frames by inserting or extracting frames, and after obtaining the corresponding spatiotemporal feature vectors respectively, the average value is taken as the preset spatiotemporal feature vector, so that the obtained preset spatiotemporal feature vector is more representative.

[0068] The scoring method provided in this embodiment can extract features about the force pattern of the target part of the human body based on the spatiotemporal feature vector sequence containing the linkage features between key nodes, compare the real-time spatiotemporal feature vector sequence with the spatiotemporal feature vector sequence preset for the move, and score the Tai Chi movements in the video to be scored based on the similarity between the two, thereby improving the accuracy of Tai Chi movement scoring and providing a more accurate reference for martial arts action referees.

[0069] Embodiment 2

[0070] Based on the scoring method provided in the first embodiment, this embodiment specifically describes the spatiotemporal feature vector constructed in step S2 according to different target parts of the human body.

[0071] Specifically, the target part of the human body includes one or more of the wrist, waist, spine, ankle and knee. Figure 2 As shown, in this embodiment, the key nodes included in the target part of the human body are defined as: the wrist includes the left shoulder S1, left elbow E1 and left wrist W1 on the left side, and the right shoulder S2, right elbow E2 and right wrist W2 on the right side; the ankle and knee include the left hip P1, left knee K1 and left ankle A1 on the left side, and the right hip P2, right knee K2 and right ankle A2 on the right side; the lumbar spine includes the left hip P1, right hip P2, left shoulder S1 and right shoulder S2.

[0072] When the human body target part is the wrist, the key nodes included in the human body target part include: shoulder joint, elbow joint and wrist joint. Since the joint nodes of the left wrist and the right wrist of the human body are the same, the constructed spatiotemporal feature vector architecture is the same.

[0073] Specifically, in one frame of the move video, the position of the shoulder joint identified by the first action recognition model is (x S ,y S ), the position of the elbow joint is (x E ,y E ), the position of the wrist joint is (x W ,y W ).

[0074] The real-time spatiotemporal feature vector of the wrist is expressed as:

[0075]

[0076] in, Indicates the spatial relative position characteristics of the shoulder and elbow; (x W -x E ,y W -y E ) represents the spatial relative position characteristics of the elbow and wrist; Indicates the spatial relative position characteristics of the shoulder, elbow and wrist; Indicates the length constraint relationship of the shoulder, elbow and wrist; represents the rate of change of the relative angle between the shoulder, elbow and wrist, θ SEW (t) represents the relative angle between the shoulder, elbow and wrist in the current frame image, θ SEW (t-1) represents the relative angle between the shoulder, elbow and wrist in the previous frame image, Δt represents the time interval between two adjacent frames of images, v S 、v E 、v W Respectively represent the movement speed of the shoulder joint, elbow joint and wrist joint; a S 、a E 、a W Respectively represent the movement acceleration of the shoulder joint, elbow joint and wrist joint. Among them, Δt is related to the frame rate f of the move video, Δt = 1 / f; v S The displacement s of the shoulder joint in the current frame image compared to the shoulder joint in the previous frame image can be calculated S , through v S =s S / Δt determines the movement speed of the shoulder joint, v E 、v W Similarly; a S The moving speed v of the shoulder joint in the current frame image can be calculated S (t) Compared with the movement speed v of the shoulder joint in the previous frame S (t-1), through a S =[v S (t)-v S (t-1)] / Δt determines the movement acceleration of the shoulder joint, a E 、a W Same reason.

[0077] Based on the above-mentioned spatiotemporal feature vectors, for a move video, the real-time spatiotemporal feature vectors of the left wrist and the right wrist corresponding to each frame image contained therein can be arranged in chronological order to form the real-time spatiotemporal feature vectors of the wrist corresponding to the move.

[0078] In the real-time spatiotemporal feature vector of the above arm, in addition to extracting the velocity feature v of each independent key node S 、v E 、v W and acceleration characteristic a S 、a E 、a W In addition, the linkage features between multiple key nodes in the wrist area are also extracted. cosθ SEW , ω SEW, thereby obtaining the force conduction pattern between key nodes during wrist movement, obtaining the characteristics of wrist movement more accurately from multiple dimensions, and improving the accuracy of the final score.

[0079] Similar to the above-mentioned wrist, when the human body target part is the ankle and knee, the key nodes included in the human body target part include: hip joint, knee joint and ankle joint. Since the joint nodes of the left ankle and knee of the human body are the same as those of the right ankle and knee, the constructed spatiotemporal feature vector architecture is the same.

[0080] Specifically, in one frame of the move video, the position of the hip joint identified by the first action recognition model is (x P ,y P ), the position of the knee joint is (x K ,y K ), the position of the ankle joint is (x A ,y A ).

[0081] The real-time spatiotemporal feature vector of the ankle and knee is expressed as:

[0082]

[0083] in, Indicates the spatial relative position characteristics of the hip and knee; Indicates the spatial relative position characteristics of the knee and ankle; Indicates the spatial relative position characteristics of the hip, knee and ankle; Indicates the length constraint relationship of the hip, knee and ankle; represents the rate of change of the relative angle between the hip, knee and ankle, θ PKA (t) represents the relative angle between the hip, knee and ankle in the current frame image, θ PKA (t-1) represents the relative angle between the hip, knee and ankle in the previous frame image, Δt represents the time interval between two adjacent frames of images, v P 、v K 、v A Respectively represent the movement speed of the hip joint, knee joint and ankle joint; a P 、a K 、a A Represent the movement acceleration of the hip joint, knee joint and ankle joint respectively. P The displacement s of the hip joint in the current frame image compared to the hip joint in the previous frame image can be calculated P , through v P =s P / Δt determines the movement speed of the hip joint, v K 、v A Similarly; a PThe moving speed v of the hip joint in the current frame image can be calculated P (t) Compared with the movement speed v of the hip joint in the previous frame P (t-1), through a P =[v P (t)-v P (t-1)] / Δt determines the movement acceleration of the hip joint, a K 、a A Same reason.

[0084] In the above-mentioned real-time spatiotemporal feature vector of the knee and ankle, in addition to extracting the velocity feature v of each independent key node P 、v K 、v A and acceleration characteristic a P 、a K 、a A In addition, the linkage features between multiple key nodes in the knee and ankle area are also extracted. cosθ SEW , ω SEW , thereby obtaining the force conduction pattern between key nodes during knee and ankle movement, obtaining the characteristics of knee and ankle movement more accurately from multiple dimensions, and improving the accuracy of the final score.

[0085] When the human body target part is the lumbar spine, the key nodes included in the human body target part include: a left shoulder joint, a right shoulder joint, a left hip joint and a right hip joint.

[0086] Specifically, in one frame of the move video, the position of the left shoulder joint identified by the first action recognition model is (x S1 ,y S1 ), the position of the right shoulder joint is (x S2 ,y S2 ), the position of the left hip joint is (x P1 ,y P1 ), the position of the right hip joint is (x P2 ,y P2 );

[0087] The real-time spatiotemporal feature vector of the lumbar spine is expressed as:

[0088]

[0089] in, Indicates the rotational characteristics of the shoulder; Indicates the rotational characteristics of the hip; Indicates the spatial relative position characteristics between the left shoulder and the right hip; Indicates the spatial relative position characteristics between the right shoulder and the left hip; It represents the first angle between shoulder and hip; It represents the second angle between shoulder and hip; The speed of change of the first angle between the shoulder and the hip; The speed of change of the second angle between the shoulder and the hip; The acceleration of the change of the first angle between the shoulder and the hip; The acceleration of the change of the second angle between the shoulder and the hip.

[0090] Based on the above-mentioned spatiotemporal feature vectors, for a move video, the real-time spatiotemporal feature vectors of the waist and spine corresponding to each frame image contained therein can be arranged in chronological order to form the real-time spatiotemporal feature vectors of the waist and spine corresponding to the move.

[0091] In the real-time spatiotemporal feature vector of the lumbar spine, the linkage features between multiple key nodes in the lumbar spine are extracted. cosθ1, cosθ2, ω1, ω2, β1, β2, especially the four characteristic values ​​of ω1, ω2, β1, β2, which reflect the changing speed and acceleration of the first angle and the second angle, can well reflect the force transmission mode between key nodes during lumbar spine rotation, and more accurately obtain the characteristics of lumbar spine movement from multiple dimensions, thereby improving the accuracy of the final score.

[0092] Embodiment 3

[0093] Based on the first embodiment, this embodiment specifically describes step S3.

[0094] This embodiment provides a scoring method for Tai Chi movements. In the first embodiment, step S3 compares the similarity between the real-time spatiotemporal feature vector sequence and the preset spatiotemporal feature vector sequence by calculating the DTW algorithm. The DTW algorithm is used to measure the similarity between two time series under the condition that there may be expansion and contraction and distortion on the time axis, and still has good recognition ability for the similarity of movements performed by the human body at different speeds.

[0095] Step S3 specifically includes sub-steps A1 to A3, which are as follows:

[0096] A1. Extracting feature values ​​of the same feature in the real-time spatiotemporal feature vector sequence to form a first time series; extracting feature values ​​corresponding to the same feature in the preset spatiotemporal feature vector sequence to form a second time series.

[0097] For example, for the spatiotemporal feature vector sequence corresponding to the wrist, if the feature to be obtained is ω SEW , all ω in the real-time spatiotemporal feature vector sequence of the wrist can be SEW Take it as the first time series and express it as All ω in the preset spatiotemporal feature vector sequence of the wrist SEW Take it as the second time series and express it as

[0098] A2. Calculate the DTW distance between the first time series and the second time series to obtain the DTW distance corresponding to each feature.

[0099] Specifically, taking the first time series and the second time series For example, the calculation method of DTW distance is:

[0100] Construct an m×n distance matrix D, where D(i,j) represents T 1i and T 2j The Euclidean distance between them is then calculated, and a path from D(1,1) to D(m,n) is found so that the sum of the distances on the path is minimized. The minimum distance is the DTW distance between the first time series T1 and the second time series T2.

[0101] A3. Perform weighted summation of the DTW distances corresponding to all features to obtain the distance sum. Based on the distance sum, score the Tai Chi movements in the video to be scored.

[0102] Specifically, the weight of each feature can be assigned according to the importance of the feature, and more important features can be assigned larger weights.

[0103] When a move corresponds to multiple human target parts, the sum or average of the distances corresponding to the multiple human target parts can be used as the distance result. For the distance result, a normalized score can be performed by dividing the distance range. For example, when the distance result of a move is [1,10], the score is 100 points; when the distance result is (10,20], the score is 95 points; when the distance result is (20,30], the score is 90 points, and so on.

[0104] In another preferred implementation of this embodiment, in step S1, identifying one or more moves formed by human body movements in the video to be rated includes:

[0105] The second action recognition model is used to recognize one or more moves formed by human body actions in the video to be rated, and the name of each move and the starting time point and the ending time point of the move in the video with rating are obtained.

[0106] Based on the name of each move and the starting time point and the ending time point of the move in the scored video, the video to be scored can be divided into multiple move videos, each of which contains and only contains one move.

[0107] Preferably, in steps S1 and S2, the first action recognition model or the second action recognition model is: a pre-trained PoseNet model with adapted model parameters.

[0108] Specifically, the training data set of the PoseNet model consists of standard action videos and non-standard action videos with the names of the moves labeled, so as to improve the PoseNet model's ability to recognize moves with lower compliance levels.

[0109] Embodiment 4

[0110] Based on the scoring methods provided in the first and second embodiments, the present invention provides a Tai Chi action scoring device, including a decomposition module, a feature extraction module and a scoring module. Figure 3 As shown, the details are as follows:

[0111] The decomposition module is used to identify one or more moves formed by human body movements in the video to be scored based on the video to be scored containing Tai Chi movements, and according to the moves, divide the video to be scored into multiple one or more move videos corresponding to the moves.

[0112] The feature extraction module is used to extract a real-time spatiotemporal feature vector sequence about the force pattern of the target part of the human body in the move video.

[0113] The scoring module is used to compare the real-time spatiotemporal feature vector sequence with the preset spatiotemporal feature vector sequence of the move, and score the Tai Chi movements in the video to be scored based on the similarity between the two.

[0114] The spatiotemporal feature vector sequence includes: linkage features between multiple key nodes included in the target part of the human body are based on the time sequence of the shown move video.

[0115] The scoring device provided in this embodiment can extract features about the force pattern of the target part of the human body based on the spatiotemporal feature vector sequence containing the linkage features between key nodes, compare the real-time spatiotemporal feature vector sequence with the spatiotemporal feature vector sequence preset for the move, and score the Tai Chi movements in the video to be scored based on the similarity between the two, thereby improving the accuracy of the Tai Chi movement scoring and providing a more accurate reference for martial arts action referees.

[0116] The present invention is described with reference to the flowcharts and / or block diagrams of the methods, devices (systems) and computer program products according to the embodiments of the present invention. It should be understood that each process and / or box in the flowchart and / or block diagram, as well as the combination of the processes and / or boxes in the flowchart and / or block diagram, can be implemented by computer program instructions.

[0117] It should be noted that in the claims, any reference numerals placed between brackets shall not be construed as limiting the claims. The word "comprising" does not exclude the presence of components or steps not listed in the claims. The word "a" or "an" preceding a component does not exclude the presence of a plurality of such components. The invention may be implemented by means of hardware comprising several different components and by means of a suitably programmed computer. In the claims enumerating several means, several of these means may be embodied by the same hardware. The use of the words first, second, third, etc., is for convenience of expression only and does not indicate any order. These words may be understood as part of the component name.

[0118] In addition, it should be noted that, in the description of this specification, the description of the terms "one embodiment", "some embodiments", "embodiment", "example", "specific example" or "some examples" etc. means that the specific features, structures, materials or characteristics described in conjunction with the embodiment or example are included in at least one embodiment or example of the present invention. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials or characteristics described may be combined in any one or more embodiments or examples in a suitable manner. In addition, those skilled in the art may combine and combine the different embodiments or examples described in this specification and the features of the different embodiments or examples, unless they are contradictory.

[0119] Although the preferred embodiments of the present invention have been described, those skilled in the art may make other changes and modifications to these embodiments after knowing the basic creative concept. Therefore, the claims should be interpreted as including the preferred embodiments and all changes and modifications falling within the scope of the present invention.

[0120] Obviously, those skilled in the art can make various modifications and variations to the present invention without departing from the spirit and scope of the present invention. Thus, if these modifications and variations of the present invention fall within the scope of the claims of the present invention and their equivalents, the present invention should also include these modifications and variations.

Claims

1. A Tai Chi action scoring method, characterized in that: include: Based on the video to be rated containing Tai Chi movements, one or more moves formed by human body movements in the video to be rated are identified, and according to the moves, the video to be rated is divided into a plurality of one or more move videos corresponding to the moves; Extracting a real-time spatiotemporal feature vector sequence of a force pattern of a target part of a human body in the move video; Comparing the real-time spatiotemporal feature vector sequence with the preset spatiotemporal feature vector sequence of the move, and scoring the Tai Chi moves in the video to be scored according to the similarity between the two; The spatiotemporal feature vector sequence includes: linkage features between multiple key nodes included in the target part of the human body are based on the time sequence of the shown move video.

2. The scoring method according to claim 1, characterized in that: Extracting a real-time spatiotemporal feature vector sequence of the force pattern of the target part of the human body in the move video, including: The first action recognition model is used to identify the positions of key nodes contained in the target part of the human body in a frame image of a move video, and the characteristic values ​​of the linkage features between the key nodes are determined according to the positions of the key nodes to form a real-time spatiotemporal feature vector. Based on the real-time spatiotemporal feature vector corresponding to each frame image in the move video, a real-time spatiotemporal feature vector sequence is formed.

3. The scoring method according to claim 1, characterized in that: Extracting the real-time spatiotemporal feature vector of the force pattern of the target part of the human body in the move video includes: The target parts of the human body include: one or more of the wrist, waist, spine or ankles and knees.

4. The scoring method according to claim 3, characterized in that: When the human body target part is the wrist, the key nodes included in the human body target part include: shoulder joint, elbow joint and wrist joint; in a frame of the move video, the position of the shoulder joint recognized by the first action recognition model is (x S ,y S ), the position of the elbow joint is (x E ,y E ), the position of the wrist joint is (x W ,y W ); The method of determining the characteristic values ​​of the linkage characteristics between the key nodes according to the positions of the key nodes to form a real-time spatiotemporal characteristic vector includes: The real-time spatiotemporal feature vector of the wrist is expressed as: in, Indicates the spatial relative position characteristics of the shoulder and elbow; Indicates the spatial relative position characteristics of the elbow and wrist; Indicates the spatial relative position characteristics of the shoulder, elbow and wrist; Indicates the length constraint relationship of the shoulder, elbow and wrist; represents the rate of change of the relative angle between the shoulder, elbow and wrist, θ SEW (t) represents the relative angle between the shoulder, elbow and wrist in the current frame image, θ SEW (t-1) represents the relative angle between the shoulder, elbow and wrist in the previous frame image, Δt represents the time interval between two adjacent frames of images, v S 、v E 、v W Respectively represent the movement speed of the shoulder joint, elbow joint and wrist joint; a S 、a E 、a W Represent the movement acceleration of the shoulder joint, elbow joint and wrist joint respectively.

5. The scoring method according to claim 3, characterized in that: When the human target part is the ankle and knee, the key nodes included in the human target part include: hip joint, knee joint and ankle joint; in a frame of the move video, the position of the hip joint recognized by the first action recognition model is (x P ,y P ), the position of the knee joint is (x K ,y K ), the position of the ankle joint is (x A ,y A ); The method of determining the characteristic values ​​of the linkage characteristics between the key nodes according to the positions of the key nodes to form a real-time spatiotemporal characteristic vector includes: The real-time spatiotemporal feature vector of the ankle and knee is expressed as: in, Indicates the spatial relative position characteristics of the hip and knee; Indicates the spatial relative position characteristics of the knee and ankle; Indicates the spatial relative position characteristics of the hip, knee and ankle; Indicates the length constraint relationship of the hip, knee and ankle; represents the rate of change of the relative angle between the hip, knee and ankle, θ PKA (t) represents the relative angle between the hip, knee and ankle in the current frame image, θ PKA (t-1) represents the relative angle between the hip, knee and ankle in the previous frame image, Δt represents the time interval between two adjacent frames of images, v P 、v K 、v A Respectively represent the movement speed of the hip joint, knee joint and ankle joint; a P 、a K 、a A Represent the movement acceleration of the hip joint, knee joint and ankle joint respectively.

6. The scoring method according to claim 3, characterized in that: When the human body target part is the waist and spine, the key nodes included in the human body target part include: left shoulder joint, right shoulder joint, left hip joint and right hip joint; in a frame image of the move video, the position of the left shoulder joint recognized by the first action recognition model is (x S1 ,y S1 ), the position of the right shoulder joint is (x S2 ,y S2 ), the position of the left hip joint is (x P1 ,y P1 ), the position of the right hip joint is (x P2 ,y P2 ); The method of determining the characteristic values ​​of the linkage characteristics between the key nodes according to the positions of the key nodes to form a real-time spatiotemporal characteristic vector includes: The real-time spatiotemporal feature vector of the lumbar spine is expressed as: in, Indicates the rotational characteristics of the shoulder; Indicates the rotational characteristics of the hip; Indicates the spatial relative position characteristics between the left shoulder and the right hip; Indicates the spatial relative position characteristics between the right shoulder and the left hip; It represents the first angle between shoulder and hip; It represents the second angle between shoulder and hip; The speed of change of the first angle between the shoulder and the hip; The speed of change of the second angle between the shoulder and the hip; The acceleration of the change of the first angle between the shoulder and the hip; The acceleration of the change of the second angle between the shoulder and the hip.

7. The scoring method according to claim 1, characterized in that: Comparing the real-time spatiotemporal feature vector sequence with the spatiotemporal feature vector sequence preset for the move, including: The preset spatiotemporal feature vector is: a spatiotemporal feature vector sequence obtained in advance based on a standard action video of the move.

8. The scoring method according to claim 7, characterized in that: The real-time spatiotemporal feature vector sequence is compared with the preset spatiotemporal feature vector sequence of the move, and the Tai Chi moves are scored according to the similarity between the two, including: A1. Extracting feature values ​​of the same feature in the real-time spatiotemporal feature vector sequence to form a first time series; extracting feature values ​​corresponding to the same feature in the preset spatiotemporal feature vector sequence to form a second time series; A2. Calculate the DTW distance between the first time series and the second time series to obtain the DTW distance corresponding to each feature; A3. Perform weighted summation of the DTW distances corresponding to all features to obtain the distance sum. Based on the distance sum, score the Tai Chi movements in the video to be scored.

9. The scoring method according to claim 2, characterized in that: Identify one or more moves formed by human body movements in the video to be rated, including: Using the second action recognition model to recognize one or more moves formed by human body movements in the video to be rated, and obtaining the name of each move and the starting time point and the ending time point of the move in the video with rating; The first action recognition model or the second action recognition model is: Pre-trained PoseNet model with adapted model parameters.

10. A Tai Chi action scoring device, characterized in that: include: A decomposition module is used to identify one or more moves formed by human body movements in the video to be rated based on the video to be rated containing Tai Chi movements, and according to the moves, divide the video to be rated into a plurality of one or more move videos corresponding to the moves; A feature extraction module, used to extract a real-time spatiotemporal feature vector sequence about the force pattern of the target part of the human body in the move video; A scoring module is used to compare the real-time spatiotemporal feature vector sequence with the preset spatiotemporal feature vector sequence of the move, and score the Tai Chi moves in the video to be scored according to the similarity between the two; The spatiotemporal feature vector sequence includes: linkage features between multiple key nodes included in the target part of the human body are based on the time sequence of the shown move video.

Citation Information

Patent Citations

  • Wushu action recognition method based on human body posture estimation

    CN114419505A

Cited By

  • Parachute landing posture recognition method based on simulation platform

    CN120316525A