Motion posture difference evaluation method and device

Through the improved Transformer model and DTW loss function, the problems of uneven distribution of individual samples and matching errors in complex movements in existing posture evaluation methods are solved, unsupervised posture difference evaluation is realized, and the accuracy and efficiency of posture evaluation are improved.

CN120689927APending Publication Date: 2025-09-23福州立行体育产业有限公司
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510573450.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-06
Publication Date
2025-09-23

AI Technical Summary

Technical Problem

Existing posture assessment methods cannot accurately evaluate motion standards when processing individual samples with normal distribution, and the skeleton point detection model is prone to matching errors in complex and periodic motions, resulting in incorrect posture assessment results.

Method used

An improved Transformer model is used for context feature extraction, sparse convolutional layers are used to replace the linear layer weights of the standard Transformer, and a DTW-based loss function is constructed to calculate the pose sequence error to achieve unsupervised difference estimation.

Benefits of technology

It improves parameter efficiency, realizes the ability to perceive the time dimension of the action, can more realistically reflect the similarity of the action itself, and improves the accuracy and efficiency of posture evaluation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120689927A_ABST
    Figure CN120689927A_ABST
Patent Text Reader

Abstract

The invention discloses a motion posture difference evaluation method and device, and relates to the technical field of posture evaluation. According to the method, context feature extraction is carried out by depending on an autonomously designed motion attitude difference evaluation model, errors are calculated for an actual attitude sequence and a predicted attitude sequence by using a DTW-based loss function, and the model is updated through back propagation, so that the model is ensured to have a time-dimension motion perception capability. In the reasoning process, the motion standard degree is evaluated according to the difference between the prediction result and the actual motion posture by means of the motion posture difference evaluation model, and unsupervised difference estimation is achieved. According to the motion posture difference evaluation method and device provided by the invention, the parameter efficiency is improved, the calculation acceleration in the reasoning process is realized, the local offset in the time sequence can be tolerated, and the similarity of the motion ontology can be reflected more truly.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of posture assessment, and in particular to a method and device for assessing movement posture differences. Background Art

[0002] With the development of AI, the use of AI algorithms for posture assessment has become a growing trend. In sports training, the accuracy of a student's movements is crucial to their training safety and ability improvement. The current training method involves instructors providing one-on-one instruction and one-on-one correction. This approach is time-consuming and slow in correcting student training details, resulting in low training efficiency. There is an urgent need for an intelligent method to accurately assess the accuracy of a student's movements.

[0003] The existing method uses the average sampling method to extract frames in the time dimension to control the time to balance the time differences of different samples in the same action; the student's action in each frame is obtained through the skeleton point detection model, and DTW matching is performed with the standard teacher's action. The skeleton point similarity of the matching results is calculated to obtain the action accuracy of the matching frame.

[0004] The existing average sampling method for compressing the time dimension can only solve the time differences of relatively balanced speeds of different samples in the same type of action, but in reality individual samples are normally distributed rather than uniformly distributed; the method of using skeleton point detection results for DTW matching is prone to matching errors in complex movements, periodic movements and similar movements, resulting in incorrect posture assessment results; and the features obtained by this method do not contain contextual information, and cannot obtain accurate assessment results for longer time series. Summary of the Invention

[0005] The technical problem to be solved by the present invention is to provide a method and device for evaluating the difference in motion posture, which uses an improved Transformer model to extract contextual features, and uses the DTW loss function to calculate the error between the predicted result and the actual motion posture to evaluate the degree of action standardization, thereby realizing unsupervised difference estimation.

[0006] In a first aspect, the present invention provides a method for evaluating motion posture differences, comprising:

[0007] Model construction process: A motion posture difference assessment model based on the Transformer structure is constructed. The motion posture difference assessment model includes an encoder module, an attention mechanism module, and a decoder module. Compared with the standard Transformer structure, the encoder and decoder modules discard the position information encoding module. The attention mechanism module uses sparse convolutional layers to replace the three linear layer weights of the Transformer structure.

[0008] Model training process: A loss function is constructed based on dynamic time warping to calculate the DTW distance between the predicted pose sequence and the true pose sequence, thereby comparing the timing matching error between the predicted pose sequence and the true pose sequence; skeleton point features are extracted from the correct action video, and after preprocessing, the motion posture difference evaluation model is trained until completion;

[0009] Model inference process: The f frames of pre-motion video before the start of motion are input into the encoder module of the motion posture difference evaluation model, and the nth frame of motion is input into the decoder module. The predicted posture of the n+1th frame is obtained at the output end. The DTW distance between the predicted posture sequence and the true posture sequence is calculated, and the motion posture difference is evaluated based on the average DTW distance of multiple consecutive frames.

[0010] In a second aspect, the present invention provides a motion posture difference assessment device, comprising:

[0011] A model construction module is used to construct a motion posture difference assessment model based on the Transformer structure. The motion posture difference assessment model includes an encoder module, an attention mechanism module, and a decoder module. Compared with the standard Transformer structure, the encoder module and the decoder module abandon the position information encoding module. The attention mechanism module uses a sparse convolutional layer to replace the three linear layer weights of the Transformer structure.

[0012] A model training module is used to construct a loss function based on dynamic time warping to calculate the DTW distance between the predicted posture sequence and the true posture sequence, thereby comparing the timing matching error between the predicted posture sequence and the true posture sequence; extract skeleton point features from the correct action video, and train the motion posture difference evaluation model after preprocessing until completion;

[0013] The model inference module is used to input the f frames of pre-motion video before the start of motion into the encoder module of the motion posture difference evaluation model, and input the nth frame of motion start into the decoder module. The predicted posture of the n+1th frame is obtained at the output end, the DTW distance between the predicted posture sequence and the true posture sequence is calculated, and the motion posture difference is evaluated based on the average DTW distance of multiple consecutive frames.

[0014] The technical solutions provided in the embodiments of the present invention have at least the following technical effects:

[0015] The model relies on a self-designed motion posture difference assessment model for contextual feature extraction. Compared to the standard Transformer structure, the model's encoder and decoder modules discard the position information encoding module. The attention mechanism module uses sparse convolutional layers to replace the original three linear layer weights of the Transformer structure, improving parameter efficiency and accelerating the calculation of the inference process. The motion posture difference assessment model uses a DTW-based loss function to calculate the error between the actual posture sequence and the predicted posture sequence, and backpropagates the model to ensure that the model has the ability to perceive motion in the time dimension. During the inference process, the motion posture difference assessment model is used to evaluate the degree of motion standardization based on the difference between the predicted result and the actual motion posture, achieving unsupervised difference estimation and more realistically reflecting the similarity of the action itself.

[0016] The above description is only an overview of the technical solution of the present invention. In order to more clearly understand the technical means of the present invention, it can be implemented in accordance with the contents of the specification. In order to make the above and other purposes, features and advantages of the present invention more obvious and easy to understand, the specific implementation methods of the present invention are specifically listed below. BRIEF DESCRIPTION OF THE DRAWINGS

[0017] The present invention will be further described below with reference to the accompanying drawings and embodiments.

[0018] Figure 1 is an overall flow chart of the method in embodiment 1 of the present invention;

[0019] Figure 2 Schematic diagram of the structure of the motion posture difference evaluation model in the first embodiment of the present invention;

[0020] Figure 3 Schematic diagram of the structure of the encoder module in the first embodiment of the present invention;

[0021] Figure 4 Schematic diagram of the structure of the decoder module in the first embodiment of the present invention;

[0022] Figure 5 This is a schematic diagram of the structure of the device in Example 2 of the present invention. DETAILED DESCRIPTION

[0023] The embodiment of the present invention provides a method and device for evaluating motion posture differences, performs context feature extraction through an improved Transformer model, and calculates the error between the predicted result and the actual motion posture based on the DTW loss function to evaluate the degree of action standardization, thereby achieving unsupervised difference estimation.

[0024] The technical solution in the embodiment of the present invention has the following general ideas:

[0025] The system relies on a self-designed motion posture difference assessment model to extract contextual features. For prediction results, the error is calculated using a DTW-based loss function between the actual and predicted posture sequences, and the model is updated through backpropagation, ensuring the model's temporal motion perception capabilities. During inference, the motion posture difference assessment model evaluates the degree of movement standardization based on the discrepancy between the predicted results and the student's actual motion posture, achieving unsupervised difference estimation.

[0026] Example 1

[0027] This embodiment provides a method for evaluating motion posture differences. Figure 1 As shown, including;

[0028] S1. Model construction process: Construct a motion posture difference assessment model based on the Transformer structure, which includes an encoder module, an attention mechanism module, and a decoder module. Compared with the standard Transformer structure, the encoder module and the decoder module abandon the position information encoding module. The attention mechanism module uses a sparse convolutional layer to replace the original three linear layer weights of the Transformer structure.

[0029] S2. Model training process: construct a loss function based on dynamic time warping to calculate the DTW distance between the predicted posture sequence and the true posture sequence, thereby comparing the timing matching error between the predicted posture sequence and the true posture sequence; extract skeleton point features from the correct action video, and train the motion posture difference evaluation model after preprocessing until completion;

[0030] S3. Model inference process: The f frames of pre-motion video before the start of motion are input into the encoder module of the motion posture difference evaluation model, and the nth frame of motion start is input into the decoder module. The predicted posture of the n+1th frame is obtained at the output end. The DTW distance between the predicted posture sequence and the true posture sequence is calculated, and the motion posture difference is evaluated based on the average DTW distance of multiple consecutive frames.

[0031] In a specific embodiment, the implementation process is as follows:

[0032] 1. Model Building

[0033] 1.1 Skeleton Point Detection Model

[0034] The MotionBERT open source algorithm is used as the 3D human skeleton point detection model. Through this model, 25 corresponding 3D skeleton points are obtained for each frame of the image. For a skeleton point sequence S with a time length of f frames,

[0035] 1.2 Motion Posture Difference Evaluation Model

[0036] A motion posture difference evaluation model based on Transformer structure is proposed, such as Figure 2 As shown. The encoder of this model inputs the skeleton points of each frame, for example, input S1, S2, S3, S4, where S i Represents the skeleton point of the i-th frame, i∈[0,f-1]. The features extracted by the encoder are calculated by the attention mechanism and then sent to the decoder. The decoder inputs the skeleton points of the subsequent frame of the encoder. For example, if the input is S5, the decoder can predict the skeleton point S6 of the subsequent frame. Based on the predicted S6, it is sent to the decoder again to obtain the predicted S7, and so on.

[0037] The structure of the encoder module is as follows Figure 3 As shown in the figure, it consists of an attention module (Attention), a matrix addition module (Add), a normalization module (Norm), and a feedforward module (FeedForward). Compared to the standard Transformer structure, the position information encoding process is discarded to accelerate the calculation process of the inference process; the relative position relationship between frames is closely modeled by the Attention component.

[0038] The traditional Attention module uses three learnable weights W q 、W k and W v The Q, K, and V matrices are learned, and the attention value is calculated using the standard Attention formula, as shown in Equations (1) and (2).

[0039]

[0040] To accelerate model training and inference, this embodiment proposes replacing the original three linear layer weights with a sparse convolutional layer with a 1×1 convolution kernel, as shown in Equation (3). The number of parameters in this sparse convolutional layer is much smaller than that of three independent fully connected layers, improving parameter efficiency.

[0041] (Q,K,V)=SparseConv(C)(3)

[0042] Among them, C represents the input features.

[0043] The decoder side structure is as follows Figure 4 As shown in the figure, the core modifications are the same as the Attention part on the Encoder side.

[0044] 2. Loss Function

[0045] In order to enhance the model's ability to perceive the dynamics of action timing, the present invention constructs a loss function based on Dynamic Time Warping (DTW) to compare the timing matching error between the predicted posture sequence and the real posture sequence. The traditional frame-by-frame Euclidean distance loss has limitations when processing speed changes and inconsistent action rhythms. The DTW algorithm uses dynamic programming to find the optimal matching path between the two sequences, allowing the model to tolerate local offsets in timing, thereby more realistically reflecting the similarity of the action entity. Suppose the predicted skeleton point sequence is The actual skeleton point sequence is S = s1, s2, ..., s f , DTW loss is calculated by calculating the Euclidean distance between two frames The cumulative distance matrix D(i,j) is defined using the formula shown in formula (4).

[0046]

[0047] in, is the Euclidean distance between two frames, i represents the index value of the predicted sequence, j represents the index value of the real sequence, and the initial value is D(l,l) is the DTW distance between the predicted sequence and the true sequence. To achieve differentiable backpropagation, the existing soft-DTW is used to calculate the loss function and optimize the Pose-Transformer.

[0048] 3. Training and Inference Process

[0049] 3.1 Data Collection and Processing

[0050] By collecting a large number of correct actions, MotionBERT is used to extract skeleton point features, and the three-dimensional skeleton points are normalized to the [0,1] range on the xyz axis.

[0051] 3.2 Model Training

[0052] We used normalized skeletal point data as training data, splitting the training set into a validation set with a ratio of 7:3. We used the Pose-Transformer model for training, using the AdamW optimizer, a learning rate of 0.01, and 300 training epochs.

[0053] 3.3 Model Inference

[0054] Assume that the student starts to move in the nth frame, according to the student's pre-motion state in the f frames before the action starts S startThe encoder side of the motion posture difference evaluation model is sent to extract features, and the skeleton action of the nth frame is input to the decoder side, and the predicted posture P_pred of the n+1th frame is obtained at the output side. (n+1) Compared with the actual posture P_gt (n+1) When the average MED (MAED) of 25 consecutive frames is greater than the threshold (0.5), it is judged as abnormal motion, and the evaluation score is 1-MAED.

[0055] Based on the same inventive concept, this application also provides a device corresponding to the method in Example 1, see Example 2 for details.

[0056] Example 2

[0057] In this embodiment, a motion posture difference evaluation device is provided. Figure 5 As shown, including:

[0058] A model construction module is used to construct a motion posture difference assessment model based on the Transformer structure. The motion posture difference assessment model includes an encoder module, an attention mechanism module, and a decoder module. Compared with the standard Transformer structure, the encoder module and the decoder module abandon the position information encoding module. The attention mechanism module uses a sparse convolutional layer to replace the three linear layer weights of the Transformer structure.

[0059] A model training module is used to construct a loss function based on dynamic time warping to calculate the DTW distance between the predicted posture sequence and the true posture sequence, thereby comparing the timing matching error between the predicted posture sequence and the true posture sequence; extract skeleton point features from the correct action video, and train the motion posture difference evaluation model after preprocessing until completion;

[0060] The model inference module is used to input the f frames of pre-motion video before the start of motion into the encoder module of the motion posture difference evaluation model, and input the nth frame of motion start into the decoder module. The predicted posture of the n+1th frame is obtained at the output end, the DTW distance between the predicted posture sequence and the true posture sequence is calculated, and the motion posture difference is evaluated based on the average DTW distance of multiple consecutive frames.

[0061] In a specific embodiment, the attention mechanism module uses a sparse convolution layer with a convolution kernel of 1×1, and the formula is as follows:

[0062] (Q,K,V)=SparseConv(C)

[0063] Among them, C represents the input features.

[0064] Assume that the predicted skeleton point sequence is The actual skeleton point sequence is S = s1, s2, ..., s f , the formula of the loss function is as follows:

[0065]

[0066] in, is the Euclidean distance between two frames, i represents the index value of the predicted sequence, j represents the index value of the real sequence, and the initial value is

[0067] The motion posture difference is evaluated based on the average DTW distance of multiple consecutive frames. Specifically, in the model reasoning module, when the average MED of the accumulated 25 consecutive frames is greater than the threshold, it is determined that the action is abnormal.

[0068] Since the device described in the second embodiment of the present invention is used to implement the method of the first embodiment of the present invention, those skilled in the art will be able to understand the specific structure and variations of the device based on the method described in the first embodiment of the present invention, and therefore will not be described in detail here. All devices used in the method of the first embodiment of the present invention fall within the scope of protection of the present invention.

[0069] The present invention relies on a self-designed motion posture difference evaluation model to extract contextual features, wherein the encoder module and decoder module of the model abandon the position information encoding module compared to the standard Transformer structure; the attention mechanism module uses a sparse convolutional layer to replace the original three linear layer weights of the Transformer structure, which improves parameter efficiency and accelerates the calculation of the reasoning process. The motion posture difference evaluation model uses a DTW-based loss function to calculate the error between the actual posture sequence and the predicted posture sequence, and backpropagates to update the model, thereby ensuring that the model has the ability to perceive motion in the time dimension. During the reasoning process, the motion posture difference evaluation model is relied upon to evaluate the degree of motion standardization based on the difference between the following predicted results and the actual student's motion content, thereby achieving unsupervised difference estimation and more realistically reflecting the similarity of the motion entity.

[0070] It will be understood by those skilled in the art that embodiments of the present invention may be provided as methods, systems, or computer program products. Thus, the present invention may take the form of an entirely hardware embodiment, an entirely software embodiment, or an embodiment combining software and hardware. Furthermore, the present invention may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to magnetic disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0071] The present invention is described with reference to flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to embodiments of the present invention. It should be understood that each process and / or block in the flowcharts and / or block diagrams, as well as combinations of processes and / or blocks in the flowcharts and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowcharts and / or block diagrams. Figure 1 a process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.

[0072] These computer program instructions may also be stored in a computer readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 a process or multiple processes and / or boxes Figure 1 The function specified in one or more boxes.

[0073] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operational steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing the instructions executed on the computer or other programmable device for implementing the process. Figure 1 a process or multiple processes and / or boxes Figure 1 A step that specifies a function in one or more boxes.

[0074] Although the specific embodiments of the present invention are described above, those skilled in the art should understand that the specific embodiments described are merely illustrative and are not intended to limit the scope of the present invention. Equivalent modifications and changes made by those skilled in the art in accordance with the spirit of the present invention should be included within the scope of protection of the claims of the present invention.

Claims

1. A method for evaluating motion posture differences, characterized in that: include: Model construction process: A motion posture difference assessment model based on the Transformer structure is constructed. The motion posture difference assessment model includes an encoder module, an attention mechanism module, and a decoder module. Compared with the standard Transformer structure, the encoder and decoder modules discard the position information encoding module. The attention mechanism module uses sparse convolutional layers to replace the three linear layer weights of the Transformer structure. Model training process: A loss function is constructed based on dynamic time warping to calculate the DTW distance between the predicted pose sequence and the true pose sequence, thereby comparing the timing matching error between the predicted pose sequence and the true pose sequence; skeleton point features are extracted from the correct action video, and after preprocessing, the motion posture difference evaluation model is trained until completion; Model inference process: The f frames of pre-motion video before the start of motion are input into the encoder module of the motion posture difference evaluation model, and the nth frame of motion is input into the decoder module. The predicted posture of the n+1th frame is obtained at the output end. The DTW distance between the predicted posture sequence and the true posture sequence is calculated, and the motion posture difference is evaluated based on the average DTW distance of multiple consecutive frames.

2. The method according to claim 1, wherein: The attention mechanism module uses a sparse convolution layer with a convolution kernel of 1×1, and the formula is as follows: (Q,K,V)=SparseConv(C) Among them, C represents the input features.

3. The method according to claim 1, characterized in that Assume that the predicted skeleton point sequence is The actual skeleton point sequence is S = s1, s2, ..., s f , the formula of the loss function is as follows: in, is the Euclidean distance between two frames, i represents the index value of the predicted sequence, j represents the index value of the real sequence, and the initial value is 4. The method according to claim 1, wherein: During the model inference process, the motion posture difference is evaluated based on the average DTW distance of multiple consecutive frames. Specifically, when the average MED of the accumulated 25 consecutive frames is greater than a threshold, it is determined to be an abnormal motion.

5. A motion posture difference assessment device, characterized in that: include: A model construction module is used to construct a motion posture difference assessment model based on the Transformer structure. The motion posture difference assessment model includes an encoder module, an attention mechanism module, and a decoder module. Compared with the standard Transformer structure, the encoder module and the decoder module abandon the position information encoding module. The attention mechanism module uses a sparse convolutional layer to replace the three linear layer weights of the Transformer structure. A model training module is used to construct a loss function based on dynamic time warping to calculate the DTW distance between the predicted posture sequence and the true posture sequence, thereby comparing the timing matching error between the predicted posture sequence and the true posture sequence; extract skeleton point features from the correct action video, and train the motion posture difference evaluation model after preprocessing until completion; The model inference module is used to input the f frames of pre-motion video before the start of motion into the encoder module of the motion posture difference evaluation model, and input the nth frame of motion start into the decoder module. The predicted posture of the n+1th frame is obtained at the output end, the DTW distance between the predicted posture sequence and the true posture sequence is calculated, and the motion posture difference is evaluated based on the average DTW distance of multiple consecutive frames.

6. The device according to claim 5, characterized in that: The attention mechanism module uses a sparse convolution layer with a convolution kernel of 1×1, and the formula is as follows: (Q,K,V)=SparseConv(C) Among them, C represents the input features.

7. The device according to claim 5, characterized in that Assume that the predicted skeleton point sequence is The actual skeleton point sequence is S = s1, s2, ..., s f , the formula of the loss function is as follows: in, is the Euclidean distance between two frames, i represents the index value of the predicted sequence, j represents the index value of the real sequence, and the initial value is 8. The device according to claim 5, characterized in that: The motion posture difference is evaluated based on the average DTW distance of multiple consecutive frames. Specifically, in the model reasoning module, when the average MED of 25 consecutive frames is greater than a threshold, it is determined that the action is abnormal.