Action correction method based on dynamic time warping
Through the dynamic time warping algorithm, the user action and the standard action are dually aligned in time and space for evaluation, which solves the multi-dimensionality problem of action evaluation in the existing technology, realizes precise error positioning and quantitative correction guidance, and improves the accuracy of action evaluation and the pertinence of correction suggestions.
Patent Information
- Application Number
- CN202510933287.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-07
- Publication Date
- 2025-10-10
AI Technical Summary
Existing technologies in human motion assessment suffer from insufficient temporal and spatial alignment and a single-dimensional scoring mechanism, resulting in inaccurate error positioning and difficulty in providing targeted correction suggestions.
A method based on dynamic time warping is adopted. The user's skeleton point sequence is compared with the standard action template through the kinematic dynamic time warping algorithm. The time alignment similarity and spatial trajectory similarity are combined, and a weighted fusion algorithm is used to calculate the local composite evaluation score, and a quantized multidimensional deviation vector is generated to provide correction instructions.
It realizes multi-dimensional evaluation of user actions, accurately identifies error fragments, and provides high-precision correction guidance, which improves the accuracy and comprehensiveness of action evaluation and the pertinence and operability of correction suggestions.
Smart Images

Figure CN120766356A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of human motion recognition, and in particular to a motion correction method, device, equipment and computer-readable storage medium based on dynamic time regularization. Background Art
[0002] With the advancement of artificial intelligence and computer vision technologies, human motion recognition and assessment have found widespread application in fields such as sports training, rehabilitation medicine, and online education. In these applications, it is often necessary to compare the user's actual movements with a preset standard to identify and correct any errors. The technical foundation of this process lies in the ability to extract 2D or 3D human skeletal joint data from videos or images using pose estimation frameworks such as OpenPose or MediaPipe. This data forms the basis for subsequent motion analysis.
[0003] At present, the industry mainly uses two technical paths for motion evaluation. One is a keyframe-based comparison method, which extracts several key posture frames from user motion and standard motion, and compares the differences between these key frames. However, this method ignores the continuity and temporal dynamics of the motion, resulting in its insensitivity to inconsistent motion speeds or errors occurring in non-keyframe parts. The other method uses the dynamic time warping (DTW) algorithm to align the time axes of the user motion sequence and the standard motion sequence to deal with the problem of different motion durations. Although the DTW algorithm can handle speed changes, its evaluation in the spatial dimension may not be detailed enough. Relying solely on the single-dimensional information of the DTW path cost often cannot fully reflect the complex spatial differences of local limb motion trajectories.
[0004] In summary, existing technologies for motion assessment commonly suffer from insufficient temporal and spatial alignment, as well as a single-dimensional scoring mechanism. These issues directly lead to inaccurate error localization, making it difficult for the system to provide targeted, quantifiable corrective recommendations. Therefore, there is an urgent need for a new technical solution that can provide precise error localization and guidance. Summary of the Invention
[0005] The embodiment of the present application aims to provide a motion correction method based on dynamic time regularization, aiming to provide a technical solution that can perform dual alignment evaluation of user motion in terms of time dynamics and spatial trajectory, and on this basis, achieve accurate error positioning and quantitative correction guidance.
[0006] To achieve the above objectives, the present invention provides a motion correction method based on dynamic time warping, comprising:
[0007] Obtain the user skeleton point sequence corresponding to the user action and the standard skeleton point sequence corresponding to the standard action;
[0008] Applying a kinematic dynamic time warping algorithm to compare the user's skeleton point sequence with the standard motion template to determine the optimal warping path and time alignment similarity;
[0009] Extracting the joint point motion trajectories in the user's skeleton point sequence and the corresponding joint point motion trajectories in the standard skeleton point sequence based on the optimal regularized path, and calculating the spatial trajectory similarity;
[0010] Dividing the user skeleton point sequence into a plurality of sub-segments according to the optimal regularized path, and for each of the sub-segments, calculating a local composite evaluation score based on the local values of the temporal alignment similarity and the spatial trajectory similarity through a weighted fusion algorithm;
[0011] Determine the sub-segment whose local composite evaluation score is lower than a preset error threshold as an error action segment in the user skeleton point sequence;
[0012] For the error action segment, a quantized multi-dimensional deviation vector is calculated, and a correction instruction is generated according to the multi-dimensional deviation vector.
[0013] In one embodiment, obtaining a user skeleton point sequence corresponding to a user action includes:
[0014] Use visual sensors to collect video stream data containing user actions in real time;
[0015] Processing the video stream data frame by frame using a preset posture estimation algorithm to identify and extract the three-dimensional coordinates of multiple preset skeletal joints of the human body in each video frame;
[0016] A corresponding time stamp is added to the three-dimensional coordinates of each video frame to form the user skeleton point sequence.
[0017] In one embodiment, a kinematic dynamic time warping algorithm is applied to compare the user's skeleton point sequence with the standard motion template to determine the optimal warping path and time alignment similarity, including:
[0018] Based on the user skeleton point sequence and the standard skeleton point sequence, first-order and second-order difference operations are performed on the joint point position coordinates to respectively calculate the user velocity sequence, the user acceleration sequence, the standard velocity sequence, and the standard acceleration sequence;
[0019] The local cost d_kinematic between any frame in the user skeleton point sequence and any frame in the standard skeleton point sequence is calculated using the following formula:
[0020] d_kinematic = a*d_pos + b*d_vel + g*d_accel,
[0021] wherein d_pos is the Euclidean distance of the position coordinate vectors of the corresponding joints between two frames, d_vel is the Euclidean distance of the corresponding user speed sequence and the standard speed sequence between two frames, d_accel is the Euclidean distance of the corresponding user acceleration sequence and the standard acceleration sequence between two frames, a, b, g are preset weight coefficients;
[0022] constructing the cost matrix with the local cost d_kinematic as the matrix elements;
[0023] determining the optimal normalized path and the time alignment similarity based on the cost matrix.
[0024] In an embodiment, according to the optimal normalized path, the joint motion trajectories in the user skeleton point sequence and the corresponding joint motion trajectories in the standard skeleton point sequence are extracted, and a spatial trajectory similarity is calculated, including:
[0025] extracting one or more preset corresponding target joints in the three-dimensional coordinate point sets in the respective sequences from the user skeleton point sequence and the standard skeleton point sequence to form user joint motion trajectories and standard joint motion trajectories, respectively;
[0026] using the Fréchet distance algorithm to calculate the distance between the user joint motion trajectories and the standard joint motion trajectories, and taking the reciprocal or non-linear mapping value of the distance as the spatial trajectory similarity.
[0027] In an embodiment, based on the local values of the time alignment similarity and the spatial trajectory similarity, a local composite evaluation score is calculated by a weighted fusion algorithm, including:
[0028] extracting the corresponding local time alignment similarity and local spatial trajectory similarity for each of the sub-segments;
[0029] According to the preset action type of the standard action, a first weight and a second weight respectively used for the local time alignment similarity and the local spatial trajectory similarity are found and determined from a preset weight mapping table;
[0030] the product of the local time alignment similarity and the first weight, and the product of the local spatial trajectory similarity and the second weight are summed to obtain the local composite evaluation score.
[0031] In one embodiment, the preset error threshold is a dynamic threshold, and the dynamic threshold is determined according to a preset action type of the standard action.
[0032] In one embodiment, for the error action segment, calculating a quantized multi-dimensional deviation vector and generating a correction instruction based on the multi-dimensional deviation vector include:
[0033] In the error action segment, the deviation values of the user action and the standard action in terms of joint point positions, joint angles and motion trajectories are quantitatively calculated to form the multi-dimensional deviation vector;
[0034] The multi-dimensional deviation vector is input into a pre-trained action-language generative model to generate natural language text containing error diagnosis and improvement suggestions as the correction instruction.
[0035] To achieve the above objectives, the present application further provides a motion correction device based on dynamic time warping, comprising:
[0036] A data acquisition module, used to acquire a user skeleton point sequence corresponding to a user action and a standard skeleton point sequence corresponding to a standard action;
[0037] A time alignment module, configured to apply a kinematic dynamic time warping algorithm to compare the user's skeleton point sequence with the standard motion template to determine an optimal warping path and time alignment similarity;
[0038] A spatial analysis module is used to extract the joint point motion trajectory in the user's skeleton point sequence and the corresponding joint point motion trajectory in the standard skeleton point sequence based on the optimal regular path, and calculate the spatial trajectory similarity;
[0039] a score calculation module, configured to divide the user skeleton point sequence into a plurality of sub-segments according to the optimal regularized path, and for each of the sub-segments, calculate a local composite evaluation score based on the local values of the temporal alignment similarity and the spatial trajectory similarity using a weighted fusion algorithm;
[0040] an error positioning module, configured to determine a sub-segment whose local composite evaluation score is lower than a preset error threshold as an error action segment in the user skeleton point sequence;
[0041] The instruction generation module is used to calculate a quantized multi-dimensional deviation vector for the error action segment and generate a correction instruction according to the multi-dimensional deviation vector.
[0042] To achieve the above-mentioned purpose, an embodiment of the present application also proposes a motion correction device based on dynamic time regularization, including a memory, a processor, and a motion correction program based on dynamic time regularization stored in the memory and runnable on the processor. When the processor executes the motion correction program based on dynamic time regularization, it implements the motion correction method based on dynamic time regularization as described in any one of the above items.
[0043] To achieve the above-mentioned purpose, an embodiment of the present application also proposes a computer-readable storage medium, on which a motion correction program based on dynamic time regularization is stored. When the motion correction program based on dynamic time regularization is executed by a processor, the motion correction method based on dynamic time regularization as described in any one of the above items is implemented.
[0044] The motion correction method based on dynamic time warping of this application has at least the following beneficial effects:
[0045] 1. Improved the accuracy and comprehensiveness of action assessment
[0046] This application significantly improves the accuracy and breadth of user motion quality assessments through a parallel, multi-dimensional evaluation framework. Firstly, the kinematic dynamic time warping algorithm employed in this application incorporates not only a comparison of "position" when calculating similarity, but also innovatively incorporates considerations of "velocity" and "acceleration." This enables the evaluation to move beyond the "similarity" of static postures and delve into the "quality" of the movements, providing a comprehensive quantitative assessment of the smoothness, coherence, and force dynamics of the movements. Secondly, this application independently utilizes the Fréchet distance algorithm to precisely compare the geometric shapes of the spatial motion trajectories of key limbs. Finally, the system combines these two distinct evaluation results, focusing on "global dynamics" and "local trajectories," through a weighted fusion algorithm that dynamically adjusts weights based on the motion type. This multi-dimensional, complementary evaluation and fusion mechanism ensures that the final evaluation results are far more accurate and comprehensive than existing methods that rely on a single dimension.
[0047] 2. Achieved more accurate error positioning
[0048] The present application can locate specific error segments in user actions with higher precision and granularity. Unlike the prior art which usually only gives a general total score for the entire action, the present application first scientifically divides the user's complete action sequence into multiple sub-segments based on the optimal regular path. Then, it calculates a local composite evaluation score for each sub-segment. By comparing these local scores with a preset error threshold (which can also be dynamically adjusted according to the action type), the system can very accurately identify which specific sub-segments or sub-segments perform poorly. This shift from "overall evaluation" to "segmented quantitative evaluation" has enabled the accuracy of error location to go from the macro level to the micro action stage level.
[0049] 3. Provide more targeted and actionable corrective suggestions
[0050] Traditional feedback is usually "right / wrong" or a simple score, but this application adopts a two-step feedback mechanism of "diagnosis-generation". First, after determining the error action segment, the system will conduct an in-depth "quantitative diagnosis" on it and calculate a multi-dimensional deviation vector containing specific deviation values such as joint position, joint angle and motion trajectory. This vector is an accurate digital description of the error. Then, the system inputs this structured diagnostic data into a pre-trained action-language generation model. The model can "translate" cold and complex data into natural language text that is easy for users to understand and contains specific diagnoses and improvement suggestions. For example, it can generate clear and executable instructions such as "Your knee position is too forward 5 cm, please sit back your hips" instead of "Your squat action is not standard." This mechanism fundamentally improves the quality of correction instructions and greatly enhances user experience and training effects. BRIEF DESCRIPTION OF THE DRAWINGS
[0051] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on the structures shown in these drawings without paying any creative work.
[0052] Figure 1 This is a module structure diagram of an embodiment of a motion correction device based on dynamic time warping of the present invention;
[0053] Figure 2 FIG. 4 is a flow chart of an embodiment of a motion correction method based on dynamic time warping according to the present invention.
[0054] The purpose, features and advantages of the present invention will be further described with reference to the accompanying drawings and in conjunction with the embodiments. DETAILED DESCRIPTION
[0055] It should be understood that the specific embodiments described herein are only used to explain the present invention and are not intended to limit the present invention.
[0056] To better understand the above technical solutions, exemplary embodiments of the present disclosure will be described in more detail below with reference to the accompanying drawings. Although exemplary embodiments of the present disclosure are shown in the accompanying drawings, it should be understood that the present disclosure can be implemented in various forms and should not be limited by the embodiments described herein. Instead, these embodiments are provided to enable a more thorough understanding of the present disclosure and to fully convey the scope of the present disclosure to those skilled in the art.
[0057] It should be noted that in the claims, any reference signs placed between brackets shall not be construed as limiting the claims. The presence of "comprising" in the text does not exclude the presence of components or steps not listed in the claims. The quantifier "a" or "an" preceding a component does not exclude the presence of a plurality of such components. The present invention can be implemented by means of hardware comprising several different components and by means of a suitably programmed computer. In a unit claim that lists several means, several of these means may be embodied by the same hardware item. The use of "first", "second", and "third" etc. does not indicate any order and these words may be interpreted as names.
[0058] like Figure 1 As shown, Figure 1 It is a structural diagram of a server 1 (also called a motion correction device based on dynamic time regularization) in the hardware operating environment involved in the embodiment of the present invention.
[0059] The server of the embodiment of the present invention is a device with display function such as "Internet of Things devices", smart air conditioners, smart lights, smart power supplies with networking functions, AR / VR devices with networking functions, smart speakers, self-driving cars, PCs, smart phones, tablet computers, e-book readers, portable computers, etc.
[0060] like Figure 1 As shown, the server 1 includes: a memory 11 , a processor 12 and a network interface 13 .
[0061] The memory 11 includes at least one type of readable storage medium, including a flash memory, a hard disk, a multimedia card, a card-type memory (e.g., an SD or DX memory), a magnetic memory, a magnetic disk, an optical disk, etc. In some embodiments, the memory 11 may be an internal storage unit of the server 1, such as a hard disk of the server 1. In other embodiments, the memory 11 may also be an external storage device of the server 1, such as a plug-in hard disk equipped on the server 1, a smart media card (SMC), a secure digital (SD) card, a flash card, etc.
[0062] Furthermore, the memory 11 may include both an internal storage unit of the server 1 and an external storage device. The memory 11 can be used not only to store application software installed on the server 1 and various data, such as the code of the motion correction program 10 based on dynamic time warping, but also to temporarily store data that has been output or is about to be output.
[0063] In some embodiments, the processor 12 can be a central processing unit (CPU), a controller, a microcontroller, a microprocessor, or other data processing chip, used to run the program code stored in the memory 11 or process data, such as executing the motion correction program 10 based on dynamic time warping.
[0064] The network interface 13 may optionally include a standard wired interface or a wireless interface (such as a WI-FI interface), and is generally used to establish a communication connection between the server 1 and other electronic devices.
[0065] The network may be the Internet, a cloud network, a wireless fidelity (Wi-Fi) network, a personal area network (PAN), a local area network (LAN), and / or a metropolitan area network (MAN). Various devices in the network environment may be configured to connect to the communication network according to various wired and wireless communication protocols. Examples of such wired and wireless communication protocols may include, but are not limited to, at least one of the following: Transmission Control Protocol and Internet Protocol (TCP / IP), User Datagram Protocol (UDP), Hypertext Transfer Protocol (HTTP), File Transfer Protocol (FTP), ZigBee, EDGE, IEEE 802.11, Light Fidelity (Li-Fi), 802.16, IEEE 802.11s, IEEE 802.11g, multi-hop communication, wireless access point (AP), device-to-device communication, cellular communication protocol, and / or Bluetooth communication protocol, or a combination thereof.
[0066] Optionally, the server may further include a user interface, which may include a display and an input unit such as a keyboard. The optional user interface may also include a standard wired interface or a wireless interface. Optionally, in some embodiments, the display may be an LED display, a liquid crystal display, a touch-sensitive liquid crystal display, or an OLED (Organic Light-Emitting Diode) touchscreen. The display, which may also be referred to as a display screen or display unit, is used to display information processed in the server 1 and to display a visual user interface.
[0067] Figure 1 Only the server 1 having components 11-13 and the motion correction program 10 based on dynamic time warping is shown. It can be understood by those skilled in the art that Figure 1 The structure shown does not constitute a limitation on the server 1 , and the server 1 may include fewer or more components than shown in the figure, or combine certain components, or arrange the components differently.
[0068] In this embodiment, the processor 12 may be configured to call the motion correction program based on dynamic time warping stored in the memory 11 and perform the following operations:
[0069] Obtain the user skeleton point sequence corresponding to the user action and the standard skeleton point sequence corresponding to the standard action;
[0070] Applying a kinematic dynamic time warping algorithm to compare the user's skeleton point sequence with the standard motion template to determine the optimal warping path and time alignment similarity;
[0071] Extracting the joint point motion trajectories in the user's skeleton point sequence and the corresponding joint point motion trajectories in the standard skeleton point sequence based on the optimal regularized path, and calculating the spatial trajectory similarity;
[0072] Dividing the user skeleton point sequence into a plurality of sub-segments according to the optimal regularized path, and calculating a local composite evaluation score for each of the sub-segments based on the local values of the temporal alignment similarity and the spatial trajectory similarity using a weighted fusion algorithm;
[0073] Determine the sub-segment whose local composite evaluation score is lower than a preset error threshold as an error action segment in the user skeleton point sequence;
[0074] For the error action segment, a quantized multi-dimensional deviation vector is calculated, and a correction instruction is generated according to the multi-dimensional deviation vector.
[0075] Based on the hardware architecture of the motion correction device based on dynamic time warping described above, embodiments of the motion correction method based on dynamic time warping of the present application are proposed. The motion correction method based on dynamic time warping of the present application aims to provide a technical solution that can perform time dynamic and spatial trajectory dual alignment evaluation on user motion, and on this basis, realize accurate error positioning and quantitative correction guidance.
[0076] Reference Figure 2 , Figure 2 For an embodiment of the motion correction method based on dynamic time warping of the present application, the motion correction method based on dynamic time warping comprises the following steps:
[0077] S10, obtaining a user skeleton point sequence corresponding to a user motion and a standard skeleton point sequence corresponding to a standard motion. Here, the user skeleton point sequence is the result of the system real-time digitizing the physical motion being performed by the user, while the standard skeleton point sequence is a pre-stored reference data serving as a comparison benchmark, which can be derived from the recording of demonstration motions by professional trainers, or synthesized through algorithms. Both sequences are time sequence data structures containing continuous multi-frame posture information.
[0078] In some embodiments, the process of obtaining the user skeleton point sequence described above can be realized by the following steps S11 to S13:
[0079] In step S11, one or more visual sensors, such as a camera, are used to capture and collect continuous video images containing the motion being performed by the user in real time, forming video stream data. In step S12, the system applies a pre-set or trained posture estimation algorithm to analyze each frame of the video stream data obtained in S11. The function of this algorithm is to identify the human body in each frame of image and accurately locate the spatial position of a pre-set group of body joints, such as the x, y, z three-dimensional coordinates of thirty-two or more joints such as head, shoulder, elbow, wrist, hip, knee and ankle. In step S13, in order to record the time dynamic characteristics of the motion, the system adds an accurate time stamp to the three-dimensional coordinate data of each frame extracted in S12. The time stamp can be derived from the time information of the video stream itself or the system clock. Finally, the continuous multi-frame coordinate data with time stamp are combined to form the user skeleton point sequence for subsequent analysis.
[0080] Specifically, in one embodiment of the present invention, the entire data acquisition process begins with the activation of the visual sensor. Once the camera starts working, it continuously generates video images at a certain frame rate (for example, 30 frames per second). These images are fed into a posture estimation processing module in real time. Within this module, a high-performance posture estimation algorithm (such as a model based on deep learning, such as MediaPipe) quickly processes each frame of the image. Its task is to decode the human skeleton structure from the pixel data and output a set of three-dimensional spatial coordinates containing all predefined joint points. To ensure that the subsequent dynamic time warping algorithm can handle the speed and rhythm of the movement, each time a set of three-dimensional coordinates is generated, the system will immediately stamp it with the current timestamp and store this "posture snapshot" in a sequence. Over time, this sequence continues to grow, eventually forming a structured sequence of user skeletal points that fully records all the user's postures over a period of time and the time when they occurred.
[0081] As can be understood, the pose estimation algorithm converts continuous video image data into a structured sequence of discrete, time-stamped three-dimensional coordinate points. Thus, it successfully transforms an unstructured, high-dimensional visual signal into a standardized mathematical object that can be used by subsequent algorithms for precise quantitative comparison and timing analysis. Furthermore, because each frame in the sequence is accurately timestamped, it provides the fundamental information necessary for subsequent dynamic time warping algorithms to assess temporal dynamic characteristics such as motion rhythm and speed.
[0082] S20. Apply a kinematic dynamic time warping algorithm to compare the user's skeleton point sequence with the standard action template to determine the optimal warping path and time alignment similarity.
[0083] Specifically, when performing sequence comparison, the traditional DTW algorithm usually only compares the differences in the spatial positions of corresponding data points in the two sequences, and is unable to evaluate the intrinsic quality of the action, such as whether the action is smooth and whether the force is accurate. The kinematic dynamic time warping algorithm proposed in the embodiment of the present application achieves a comprehensive evaluation of the quality of action execution by introducing a comprehensive consideration of the speed and acceleration information of the action. This step will eventually generate two core outputs: the optimal warping path, which defines the optimal nonlinear mapping relationship between the user action timeline and the standard action timeline; and the time alignment similarity, which is a quantitative score that characterizes the overall degree of conformity of the user action with the standard action in terms of spatiotemporal dynamics.
[0084] In some embodiments, the process of applying the kinematic dynamic time warping algorithm can be implemented by following the steps S21 to S24:
[0085] First, in step S21, the system estimates the velocity sequence and acceleration sequence corresponding to each of the two sequences based on the user's skeleton point sequence and the standard action template by performing first-order and second-order difference operations on the joint position coordinates of each frame. Then, in step S22, a specific formula is used to calculate the local cost d_kinematic between any frame in the user sequence and any frame in the standard sequence. The formula is: d_kinematic = α·d_pos+β·d_vel+γ·d_accel. Among them, d_pos is the difference value of the joint position between the two frames, d_vel is the difference value of the joint velocity between the two frames, d_accel is the difference value of the joint acceleration between the two frames, and α, β, γ are weight coefficients that can be preset according to the focus of the evaluation task. Specifically, the technical details of the calculation of the position difference value, velocity difference value and acceleration difference value are as follows
[0086] Position difference value (d_pos): The system calculates the Euclidean distance between the coordinate vectors of all K corresponding joint points in two frames and sums or averages them. The formula can be expressed as:
[0087]
[0088] Velocity difference (d_vel): The system first calculates the velocity vector of each joint point in the player and the standard template by taking the first-order difference of the position coordinates between consecutive frames (i.e., (pos_t - pos_{t-1}) / Δt). Then, the system calculates the sum of the Euclidean distances between the corresponding joint point velocity vectors between the two frames.
[0089] Acceleration difference (d_accel): Similarly, the system calculates the acceleration vector for each joint by taking the first-order difference of the velocity vector. The sum of the Euclidean distances between the corresponding joint acceleration vectors between the two frames is then calculated.
[0090] Next, in step S23, the system constructs a two-dimensional cost matrix using the calculated local cost d_kinematic as a matrix element. Specifically, the system initializes a matrix of size N x M, where N is the length of the user's skeletal point sequence and M is the length of the standard motion template. The system then populates the cell in row i and column j of the cost matrix with the local cost calculated between frame i of the user sequence and frame j of the template sequence to obtain the cost matrix.
[0091] Finally, in step S24, the system uses a dynamic programming algorithm to find a path with the minimum cumulative cost from the starting point to the end point of the matrix based on the constructed cost matrix. This path is the optimal regular path, and its final cumulative cost value is normalized to obtain the time alignment similarity.
[0092] Specifically, the entire computational process begins with feature enhancement of the input data. Instead of relying solely on raw position coordinates, the system first computes corresponding velocity and acceleration sequences for both the user and standard action sequences through differential operations (step S21). Based on this, when comparing any two "pose snapshots" (i.e., frames), the degree of dissimilarity (i.e., local cost) between them is no longer a simple spatial distance, but instead consists of a weighted combination of three components (step S22): position coordinate difference (d_pos), which reflects pose accuracy; velocity difference (d_vel), which reflects the smoothness and coherence of the movement; and acceleration difference (d_accel), which reflects the force pattern and explosiveness of the movement. These local costs are used to populate a cost matrix (step S23), which forms the basis for subsequent analysis. Finally, a standard dynamic programming pathfinding algorithm finds a path on this "cost map" with the lowest total cost. This total cost represents the overall difference between the two action sequences, which is converted into a temporal alignment similarity score. The path itself records detailed frame alignment information (step S24).
[0093] For example, the algorithm's evaluation focus can be adjusted by adjusting the weight coefficients α, β, and γ, making it flexible and adaptable to different evaluation scenarios. When evaluating a rehabilitation exercise, the key requirement is smoothness and control of the movement. In this case, the weight β (corresponding to velocity) can be set to be greater than α (corresponding to position) and γ (corresponding to acceleration). In this configuration, even if the user's final posture position is accurate (d_pos is low), if any jitter or sudden velocity changes occur during the process (resulting in a high d_vel), the calculated local cost d_kinematic will still be high, accurately assessing the movement's poor quality. Conversely, when evaluating a martial arts movement, the key lies in the explosive force at the moment of impact. In this case, the weight γ (corresponding to acceleration) can be set to be much greater than α and β. This makes the algorithm highly sensitive to the sharpness of acceleration peaks, effectively distinguishing between powerful, professional movements and hesitant, imitative ones, even if the final positions and overall velocities are similar.
[0094] As can be understood, the calculation of local costs innovatively incorporates position information representing the spatial accuracy of the motion, velocity information representing the smoothness of the motion, and acceleration information representing the dynamic force generation of the motion. This elevates the evaluation capability of the dynamic time warping algorithm from a simple morphological similarity comparison to a higher dimension of comprehensive assessment of the quality of motion execution. Furthermore, because the weight coefficients α, β, and γ in this implementation can be adjusted based on the evaluation objective, this method can flexibly adapt to different types of motion assessment tasks, enabling targeted and more accurate quantitative assessments of both smooth rehabilitation movements and explosive sports movements.
[0095] S30 , extracting the joint point motion trajectories in the user's skeleton point sequence and the corresponding joint point motion trajectories in the standard skeleton point sequence based on the optimal regularized path, and calculating the spatial trajectory similarity.
[0096] In some embodiments, the above process of calculating the spatial trajectory similarity can be implemented by following the steps S31 to S32:
[0097] First, in step S31, the system extracts only the three-dimensional coordinate points of one or more predefined target joint points (for example, the right wrist joint point) in all frames of each sequence from the previously acquired user skeleton point sequence and standard motion template. These two sets of extracted three-dimensional coordinate points constitute two spatial curves representing the user's actual motion trajectory and the standard motion trajectory, respectively. Then, in step S32, the system uses the Fréchet distance algorithm to calculate the similarity between the two spatial curves. The smaller the calculated Fréchet distance value, the more similar the two trajectories are in geometric shape. The system will map this distance value to a standardized spatial trajectory similarity score by taking the inverse or other nonlinear functions.
[0098] Specifically, after obtaining the global time alignment relationship through S20, the system will focus on the most important body parts for the current action. For example, when evaluating a throwing action, the target joint point may be set to the wrist. The system will traverse the user's bone point sequence and the standard action template, and connect all the three-dimensional coordinate points related to the wrist in chronological order to obtain two three-dimensional space curves representing the actual trajectory and the standard trajectory. Then, instead of performing a simple point-by-point distance comparison on these two curves, the system applies the more robust Fréchet distance algorithm.
[0099] As can be understood, by ensuring temporal comparability of trajectories through the use of an optimal regularized path and employing the Fréchet distance algorithm to assess the geometric similarity of two spatial curves, an independent, refined quantitative assessment of the spatial geometric accuracy of the motion trajectory of a user's specific limb can be performed, independent of changes in the speed or rhythm of the user's movements. Furthermore, because this evaluation dimension complements the overall posture-based evaluation dimension in step S20, it provides another critical, differently focused input for generating a more comprehensive composite evaluation score in subsequent steps, thereby enhancing the comprehensiveness of the movement accuracy assessment.
[0100] S40. Divide the user skeleton point sequence into multiple sub-segments according to the optimal regularized path, and for each of the sub-segments, calculate a local composite evaluation score based on the local values of the time alignment similarity and the spatial trajectory similarity through a weighted fusion algorithm.
[0101] Specifically, step S40 performs a segmented, local evaluation of the entire action to calculate a local composite evaluation score. Unlike the macro-level evaluations of the complete action, such as temporal alignment similarity and spatial trajectory similarity, generated in the previous steps, this step aims to reduce the evaluation granularity to the internal segments of the action. This segmented evaluation method can more precisely identify the specific link in a long sequence of actions that has problems.
[0102] Specifically, the step of dividing the sub-segments according to the optimal planned path includes: first, when the standard action template is created, it can be semantically segmented in advance and divided into multiple stages with specific technical meanings. For example, a "shooting" action can be divided into four stages: "preparation", "take-off", "shooting" and "landing". Each stage corresponds to a starting frame index and an ending frame index in the standard action sequence. After the optimal regularized path is determined in step S20, the path establishes a precise correspondence between each frame of the user sequence and the standard sequence. Therefore, the system only needs to find the user action frames corresponding to the starting frames and ending frames of each preset stage in the standard template in the optimal regularized path, and it can divide the user's complete action sequence into corresponding multiple sub-segments in the same way and completely aligned in time.
[0103] In some embodiments, the above process of calculating the local composite evaluation score can be implemented by following steps S41 to S43:
[0104] First, in step S41, for each divided sub-segment, the system calculates its local "time alignment similarity" and "spatial trajectory similarity" respectively. The former comes from the cumulative cost of the corresponding path segment of the segment on the cost matrix, and the latter comes from the Fréchet distance of the corresponding trajectory of the segment. Then, in step S42, the system will query and determine a set of weight coefficients specifically for this type of action from a preset weight mapping table based on the preset action type of the standard action currently being evaluated. This set of weight coefficients includes a first weight for weighting the local time alignment similarity and a second weight for weighting the local spatial trajectory similarity. Finally, in step S43, the system multiplies the local time alignment similarity of the sub-segment with the first weight, multiplies the local spatial trajectory similarity with the second weight, and then adds the two products to obtain the final, local composite evaluation score of the sub-segment.
[0105] Specifically, through the above embodiments, when the present application comprehensively scores an action segment, the present method does not use a fixed fusion weight, but first identifies the preset type of the action in step S42, for example, is it a "balance action", "explosive action" or "rhythm action". Then, based on the type, the system will find the most suitable weight combination for evaluating this type of action from a pre-configured weight mapping table. For example, for actions that focus on posture stability, the weight of the spatial dimension will be higher; and for actions that focus on force rhythm, the weight of the time dimension will be higher. After determining the weight applicable to the current sub-segment, the system performs a weighted summation in step S43, and finally obtains a composite score that can intelligently reflect the core requirements of the action type.
[0106] For example, suppose a standard movement is labeled "balance," such as the yoga pose "Tree Pose." When evaluating a user's imitation of this movement, the system will find weights appropriate for "balance" in step S42. For example, a first weight of 0.8 is assigned to spatial trajectory similarity and a second weight of 0.2 to temporal alignment similarity. This is because, for balance movements, maintaining a stable body trajectory (spatial accuracy) is far more important than completing the movement quickly (temporal alignment). Therefore, even if the user's rhythm of completing the movement deviates slightly, as long as the body sway is small, the local composite evaluation score will still be high. Conversely, if another standard movement is labeled "explosive movement," such as a fast punch, the system will assign the opposite set of weights, for example, a first weight of 0.7 to temporal alignment similarity and a second weight of 0.3 to spatial trajectory similarity. In this case, the system will place greater emphasis on whether the user can complete the movement with the correct timing and explosiveness, and will be more tolerant of slight deviations in the trajectory of the punch.
[0107] As you can see, by segmenting the entire movement and assigning an independent composite score to the local performance of each segment, it is possible to transform a general assessment of a long movement into a refined evaluation of multiple shorter movements, thereby more accurately identifying the specific intervals where problems occur. Furthermore, since the composite score is calculated by dynamically adjusting weights based on movement type, the evaluation criteria are no longer fixed, but can be adaptively adjusted based on the core requirements of different movements, making the final evaluation results more targeted and reasonable.
[0108] S50: Determine the sub-segment whose local composite evaluation score is lower than a preset error threshold as an error action segment in the user skeleton point sequence.
[0109] Specifically, the core logic of this step is a threshold comparison process. The system traverses each sub-segment and compares its corresponding local composite evaluation score with a preset error threshold. When the score of a sub-segment is lower than the threshold, the system marks the sub-segment and determines it as an error action segment. A key aspect is that the preset error threshold here is not a fixed value applicable to all actions, but a dynamic threshold. Its specific value is determined by the preset action type of the standard action currently being evaluated, which makes the judgment standard more flexible and adaptable.
[0110] In a specific embodiment, the process of determining and applying this dynamic threshold is as follows: when it is necessary to evaluate the user's action, the system will first read the preset "action type" label in the standard action template it refers to. Then, the system uses the "action type" as an index to query in a pre-configured "action type-error threshold" mapping table to find the error threshold specifically used to evaluate this type of action. The mapping table can store multiple sets of correspondences, for example, "high-precision actions" correspond to a higher threshold, and "normal range actions" correspond to a lower threshold. After obtaining the dynamic error threshold that matches the current action type, the system uses this threshold to compare and judge the local composite evaluation scores of each sub-segment one by one.
[0111] As can be understood, the dynamic threshold mechanism—determining the threshold for determining whether an error constitutes an error based on the preset type of action being evaluated—avoids the "one-size-fits-all" problem associated with a single, fixed threshold. This allows the "rigor" of the evaluation to be tailored to the inherent difficulty and error tolerance of the action itself. Furthermore, this adaptive judgment standard significantly improves the accuracy and rationality of identifying error-prone action segments, reducing misjudgments of simple actions while ensuring sensitivity in evaluating demanding actions.
[0112] S60. Calculate a quantized multi-dimensional deviation vector for the error action segment, and generate a correction instruction based on the multi-dimensional deviation vector.
[0113] Specifically, the core idea of this step is not only to tell the user "what is wrong", but also to quantitatively tell the user "where the mistake is and how serious the mistake is", and finally convert these quantitative diagnostic data into humane and easy-to-understand improvement suggestions.
[0114] In a specific embodiment, the above process of generating the correction instruction can be implemented through steps S61 to S62.
[0115] First, in step S61, the system performs a frame-by-frame, multi-dimensional comparison of the user's motion and the standard motion within the identified error motion segment to quantify the specific deviation between the two, thereby forming a structured multi-dimensional deviation vector. This vector can contain data from multiple dimensions, such as: joint position deviation, i.e., how many centimeters a user's joint deviates from the standard position in space; joint angle deviation, i.e., how many degrees the flexion or extension angle of a user's joint (such as a knee or elbow) differs from the standard angle; and motion trajectory deviation, i.e., the degree to which the geometric shape of the motion trajectory of a user's limb (such as a hand or foot) differs from the standard trajectory.
[0116] Next, in step S62, after obtaining this multidimensional deviation vector containing precise, quantitative diagnostic information, the system then inputs it into a pre-trained, advanced motion-language generative model. This generative model functions similarly to a professional AI coach, trained to understand the kinematic meaning of each value in the multidimensional deviation vector. Based on the input deviation values, the model automatically and dynamically organizes the language to generate a fluent natural language text containing the error diagnosis and specific improvement suggestions. This AI-generated text is the final correction instruction output to the user.
[0117] For example, suppose the system identifies an error segment during the squat phase between seconds 3 and 5 while the user is performing a "deep squat." In step S61, the system calculates the multidimensional deviation vector corresponding to this segment: {joint: 'left knee', deviation type: 'position', deviation value: '+10cm_forward'}, {joint: 'hip', deviation type: 'angle', deviation value: '-15 degrees_insufficient'}. In step S62, this structured deviation vector is fed into the action-language generation model. After receiving the input, the model "translates" it into natural language, such as: "During the squat, please note that your left knee extends too far beyond your toes (approximately 10 cm) and your hips do not sit back enough (approximately 15 degrees). Please try to place more weight on your heels and actively sit back and downward, as if you were sitting on an invisible chair." This text is the final generated correction instruction.
[0118] As can be understood, by performing multi-dimensional quantification on detected errors, a multi-dimensional deviation vector containing specific position, angle, and trajectory deviation values is formed, enabling in-depth, data-driven diagnosis of errors, rather than simply stating "this part of the action is wrong." Furthermore, by feeding this quantified multi-dimensional deviation vector into an action-language generative model for processing, the complex, multi-dimensional data-based diagnostic results can be automatically and smoothly converted into specific, natural language corrective instructions that are easy for human users to understand and execute, greatly improving the effectiveness of feedback and user experience.
[0119] Based on the above embodiments, the motion correction method based on dynamic time warping of the present application has at least the following beneficial effects:
[0120] 1. Improved the accuracy and comprehensiveness of action assessment
[0121] This application significantly improves the accuracy and breadth of user motion quality assessments through a parallel, multi-dimensional evaluation framework. Firstly, the kinematic dynamic time warping algorithm employed in this application incorporates not only a comparison of "position" when calculating similarity, but also innovatively incorporates considerations of "velocity" and "acceleration." This enables the evaluation to move beyond the "similarity" of static postures and delve into the "quality" of the movements, providing a comprehensive quantitative assessment of the smoothness, coherence, and force dynamics of the movements. Secondly, this application independently utilizes the Fréchet distance algorithm to precisely compare the geometric shapes of the spatial motion trajectories of key limbs. Finally, the system combines these two distinct evaluation results, focusing on "global dynamics" and "local trajectories," through a weighted fusion algorithm that dynamically adjusts weights based on the motion type. This multi-dimensional, complementary evaluation and fusion mechanism ensures that the final evaluation results are far more accurate and comprehensive than existing methods that rely on a single dimension.
[0122] 2. Achieved more accurate error positioning
[0123] The present application can locate specific error segments in user actions with higher precision and granularity. Unlike the prior art which usually only gives a general total score for the entire action, the present application first scientifically divides the user's complete action sequence into multiple sub-segments based on the optimal regular path. Then, it calculates a local composite evaluation score for each sub-segment. By comparing these local scores with a preset error threshold (which can also be dynamically adjusted according to the action type), the system can very accurately identify which specific sub-segments or sub-segments perform poorly. This shift from "overall evaluation" to "segmented quantitative evaluation" has enabled the accuracy of error location to go from the macro level to the micro action stage level.
[0124] 3. Provide more targeted and actionable corrective suggestions
[0125] Traditional feedback is usually “right / wrong” or a simple score, while the present application adopts a “diagnosis-generation” two-step feedback mechanism. First, after determining the error motion segment, the system will conduct an in-depth “quantitative diagnosis” on it, calculating a multi-dimensional deviation vector containing specific deviation values such as joint position, joint angle and motion trajectory, etc. This vector is an accurate data-based description of the error. Then, the system inputs this structured diagnostic data into a pre-trained motion-language generation model. The model can “translate” the cold and complex data into natural language text that is easy for users to understand, containing specific diagnosis and improvement suggestions. For example, it can generate clear and executable instructions such as “your knee position is 5 cm too forward, please sit back on your hips”, instead of “your squat motion is not standard”. This mechanism fundamentally improves the quality of correction instructions, greatly enhancing user experience and training effectiveness.
[0126] In addition, the embodiment of the present application also proposes a motion correction device based on dynamic time warping, which comprises:
[0127] A data acquisition module is configured to acquire a user skeleton point sequence corresponding to a user motion and a standard skeleton point sequence corresponding to a standard motion.
[0128] A time alignment module is configured to apply a kinematic dynamic time warping algorithm to compare the user skeleton point sequence with the standard motion template to determine an optimal warping path and a time alignment similarity.
[0129] A spatial analysis module is configured to extract a joint motion trajectory in the user skeleton point sequence and a corresponding joint motion trajectory in the standard skeleton point sequence according to the optimal warping path, and calculate a spatial trajectory similarity.
[0130] A score calculation module is configured to divide the user skeleton point sequence into a plurality of sub-segments according to the optimal warping path, and calculate a local composite evaluation score for each sub-segment based on a local value of the time alignment similarity and the spatial trajectory similarity through a weighted fusion algorithm.
[0131] An error positioning module is configured to determine a sub-segment with a local composite evaluation score lower than a preset error threshold as an error motion segment in the user skeleton point sequence.
[0132] An instruction generation module is configured to calculate a quantitative multi-dimensional deviation vector for the error motion segment, and generate a correction instruction according to the multi-dimensional deviation vector.
[0133] Among them, the steps implemented by each functional module of the motion correction device based on dynamic time warping can refer to the various embodiments of the motion correction method based on dynamic time warping of the present invention, and will not be repeated here.
[0134] In addition, an embodiment of the present invention further provides a computer-readable storage medium, which can be any one of a hard disk, a multimedia card, an SD card, a flash memory card, an SMC, a read-only memory (ROM), an erasable programmable read-only memory (EPROM), a portable compact disk read-only memory (CD-ROM), a USB memory, or any combination thereof. The computer-readable storage medium includes a motion correction program 10 based on dynamic time warping. The specific implementation of the computer-readable storage medium of the present invention is substantially the same as the specific implementation of the motion correction method based on dynamic time warping and the server 1 described above, and will not be repeated here.
[0135] It will be understood by those skilled in the art that embodiments of the present invention may be provided as methods, systems, or computer program products. Thus, the present invention may take the form of an entirely hardware embodiment, an entirely software embodiment, or an embodiment combining software and hardware. Furthermore, the present invention may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to magnetic disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0136] The present invention is described with reference to flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to embodiments of the present invention. It should be understood that each process and / or block in the flowcharts and / or block diagrams, as well as combinations of processes and / or blocks in the flowcharts and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowcharts and / or block diagrams. Figure 1 a process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.
[0137] These computer program instructions may also be stored in a computer readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 a process or multiple processes and / or boxes Figure 1 The function specified in one or more boxes.
[0138] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operational steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing the instructions executed on the computer or other programmable device for implementing the process. Figure 1 a process or multiple processes and / or boxes Figure 1 A step that specifies a function in one or more boxes.
[0139] Although the preferred embodiments of the present invention have been described, those skilled in the art may make additional changes and modifications to these embodiments once they have learned the basic creative concept. Therefore, the appended claims are intended to be interpreted as including the preferred embodiments and all changes and modifications that fall within the scope of the present invention.
[0140] Obviously, those skilled in the art may make various changes and modifications to the present invention without departing from the spirit and scope of the present invention. Thus, if such changes and modifications fall within the scope of the claims and their equivalents, the present invention is intended to include such changes and modifications.
Claims
1. A motion correction method based on dynamic time warping, characterized in that: include: Obtain the user skeleton point sequence corresponding to the user action and the standard skeleton point sequence corresponding to the standard action; Applying a kinematic dynamic time warping algorithm to compare the user's skeleton point sequence with the standard motion template to determine the optimal warping path and time alignment similarity; Extracting the joint point motion trajectories in the user's skeleton point sequence and the corresponding joint point motion trajectories in the standard skeleton point sequence based on the optimal regularized path, and calculating the spatial trajectory similarity; Dividing the user skeleton point sequence into a plurality of sub-segments according to the optimal regularized path, and calculating a local composite evaluation score for each of the sub-segments based on the local values of the temporal alignment similarity and the spatial trajectory similarity using a weighted fusion algorithm; Determine the sub-segment whose local composite evaluation score is lower than a preset error threshold as an error action segment in the user skeleton point sequence; For the error action segment, a quantized multi-dimensional deviation vector is calculated, and a correction instruction is generated according to the multi-dimensional deviation vector.
2. The motion correction method based on dynamic time warping according to claim 1, characterized in that: Get the user skeleton point sequence corresponding to the user action, including: Use visual sensors to collect video stream data containing user actions in real time; Processing the video stream data frame by frame using a preset posture estimation algorithm to identify and extract the three-dimensional coordinates of multiple preset skeletal joints of the human body in each video frame; A corresponding time stamp is added to the three-dimensional coordinates of each video frame to form the user skeleton point sequence.
3. The motion correction method based on dynamic time warping according to claim 1, wherein: Applying a kinematic dynamic time warping algorithm to compare the user's skeleton point sequence with the standard motion template to determine the optimal warping path and time alignment similarity, including: Based on the user skeleton point sequence and the standard skeleton point sequence, first-order and second-order difference operations are performed on the joint point position coordinates to respectively calculate the user velocity sequence, the user acceleration sequence, the standard velocity sequence, and the standard acceleration sequence; The local cost d_kinematic between any frame in the user skeleton point sequence and any frame in the standard skeleton point sequence is calculated using the following formula: d_kinematic=α·d_pos+β·d_vel+γ·d_accel, Where d_pos is the Euclidean distance between the joint position coordinate vectors of the two frames, d_vel is the Euclidean distance between the user velocity sequence and the standard velocity sequence of the two frames, d_accel is the Euclidean distance between the user acceleration sequence and the standard acceleration sequence of the two frames, and α, β, and γ are preset weight coefficients. Constructing the cost matrix with the local cost d_kinematic as a matrix element; The optimal regularized path and the time-aligned similarity are determined based on the cost matrix.
4. The motion correction method based on dynamic time warping according to claim 1, wherein: Extracting the joint point motion trajectory in the user's skeleton point sequence and the corresponding joint point motion trajectory in the standard skeleton point sequence based on the optimal regularized path, and calculating the spatial trajectory similarity, including: Extracting one or more preset three-dimensional coordinate point sets of corresponding target joint points in the respective sequences from the user skeleton point sequence and the standard skeleton point sequence, and forming user joint point motion trajectories and standard joint point motion trajectories respectively; The Fréchet distance algorithm is used to calculate the distance between the user joint point motion trajectory and the standard joint point motion trajectory, and the reciprocal or nonlinear mapping value of the distance is used as the spatial trajectory similarity.
5. The motion correction method based on dynamic time warping according to claim 1, wherein: Based on the local values of the temporal alignment similarity and the spatial trajectory similarity, a local composite evaluation score is calculated by a weighted fusion algorithm, including: For each of the sub-segments, extracting the corresponding local time alignment similarity and local spatial trajectory similarity; According to a preset action type of the standard action, searching and determining a first weight and a second weight for the local time alignment similarity and the local spatial trajectory similarity, respectively, from a preset weight mapping table; The product of the local time alignment similarity and the first weight and the product of the local spatial trajectory similarity and the second weight are summed to obtain the local composite evaluation score.
6. The motion correction method based on dynamic time warping according to claim 1, wherein: The preset error threshold is a dynamic threshold, and the dynamic threshold is determined according to a preset action type of the standard action.
7. The motion correction method based on dynamic time warping according to claim 1, wherein: For the error action segment, a quantized multi-dimensional deviation vector is calculated, and a correction instruction is generated according to the multi-dimensional deviation vector, including: In the error action segment, the deviation values of the user action and the standard action in terms of joint point positions, joint angles and motion trajectories are quantitatively calculated to form the multi-dimensional deviation vector; The multi-dimensional deviation vector is input into a pre-trained action-language generative model to generate natural language text containing error diagnosis and improvement suggestions as the correction instruction.
8. A motion correction device based on dynamic time warping, characterized in that: include: A data acquisition module is used to obtain a user skeleton point sequence corresponding to a user action and a standard skeleton point sequence corresponding to a standard action; A time alignment module, configured to apply a kinematic dynamic time warping algorithm to compare the user's skeleton point sequence with the standard motion template to determine an optimal warping path and time alignment similarity; A spatial analysis module is used to extract the joint point motion trajectory in the user's skeleton point sequence and the corresponding joint point motion trajectory in the standard skeleton point sequence based on the optimal regular path, and calculate the spatial trajectory similarity; a score calculation module, configured to divide the user skeleton point sequence into a plurality of sub-segments according to the optimal regularized path, and for each of the sub-segments, calculate a local composite evaluation score based on the local values of the temporal alignment similarity and the spatial trajectory similarity using a weighted fusion algorithm; an error positioning module, configured to determine a sub-segment whose local composite evaluation score is lower than a preset error threshold as an error action segment in the user skeleton point sequence; The instruction generation module is used to calculate a quantized multi-dimensional deviation vector for the error action segment and generate a correction instruction according to the multi-dimensional deviation vector.
9. A motion correction device based on dynamic time warping, characterized in that: It includes a memory, a processor, and a motion correction program based on dynamic time regularization stored in the memory and executable on the processor. When the processor executes the motion correction program based on dynamic time regularization, it implements the motion correction method based on dynamic time regularization as described in any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that The computer-readable storage medium stores a motion correction program based on dynamic time warping, and when the motion correction program based on dynamic time warping is executed by the processor, the motion correction method based on dynamic time warping as described in any one of claims 1 to 7 is implemented.