3D exercise training auxiliary method and device based on single-view video

Through the 3D motion training auxiliary method of single-view video, the 3D mannequin model is constructed and the posture is compared using the depth estimation algorithm, which solves the high cost problem of relying on professional coaches in the existing technology, and realizes convenient and personalized sports training and feedback, which is suitable for athletes and rehabilitated patients.

CN120472531APending Publication Date: 2025-08-12FUZHOU UNIV
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202510547369.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-28
Publication Date
2025-08-12

AI Technical Summary

Technical Problem

The existing sports training, rehabilitation training and movement assessment systems rely on professional coaches, which are costly and inconvenient enough. Especially for ordinary users, the existing technology relies on sportswear with sensors still have problems such as high cost and inconvenient enough.

Method used

Through single-view video, the depth estimation calculation method is used to extract the two-dimensional and three-dimensional coordinates of the key points of the human body, and the 3D human body model is constructed, and compared with the standard motion posture model to generate feedback information to guide users to adjust their postures and reduce their dependence on special sports clothing.

Benefits of technology

It reduces user costs, improves the convenience of sports training, provides personalized feedback, and improves training results. It is suitable for athletes, rehabilitation patients and users who need action assessment.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120472531A_ABST
    Figure CN120472531A_ABST
Patent Text Reader

Abstract

The invention discloses a 3D exercise training assisting method and device based on a single-view video, and relates to the field of exercise training assistion.The method comprises the steps that two-dimensional coordinates of key points of a human body in a first current video frame are extracted; determining a depth value of each key point of the human body by using a depth estimation algorithm, and further obtaining a three-dimensional coordinate of each key point of the human body; constructing a first 3D human body model according to the three-dimensional coordinates of the key points of the human body in the first current video frame; comparing the first 3D human body model with the second 3D human body model to judge whether the current motion posture of the user accords with a standard motion posture or not; when it is judged that the current motion posture of the user does not conform to the standard motion posture, first feedback information is generated to guide the user to adjust the current motion posture of the user. The exercise training auxiliary cost of the user can be effectively reduced, and high-convenience exercise training auxiliary service is provided for the user.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of sports training assistance technology, and in particular to a 3D sports training assistance method and device based on single-view video. Background Art

[0002] Most existing sports training, rehabilitation training and action assessment systems rely on the guidance of professional coaches, which are costly and inconvenient for general users, and there is also the problem of uneven coaching levels. Prior art CN110109532A designs sportswear with sensors to collect motion parameters of various parts of the user (such as joints), such as speed, acceleration, displacement, etc., when the user exercises, and marks them on the user's human body model, and then compares them with the marked standard human body model to obtain the difference information between the user's real-time action and the standard action and output it for display, so that the above problems are improved to a certain extent. However, for general users, the above existing technologies still have the problems of high cost and inconvenience. Summary of the Invention

[0003] The purpose of this application is to provide a 3D motion training assistance method and device based on single-view video, which can effectively reduce the user's motion training assistance cost and provide users with highly convenient motion training assistance services.

[0004] To achieve the above objectives, this application provides the following solutions:

[0005] In a first aspect, the present application provides a 3D motion training assistance method based on a single-view video, comprising:

[0006] Extracting two-dimensional coordinates of key points of a human body in a first current video frame; the first current video frame is a frame in a user motion video; the user motion video is a video of a user exercising according to a sports training program captured by a monocular camera;

[0007] Determine the depth values of key points of the human body in the first current video frame by combining the first current video frame and a first preset number of consecutive historical video frames before the first current video frame using a depth estimation algorithm;

[0008] Determine the three-dimensional coordinates of each key point of the human body in the first current video frame based on the two-dimensional coordinates and depth values of each key point of the human body in the first current video frame;

[0009] Constructing a first 3D human body model based on the three-dimensional coordinates of key points of the human body in the first current video frame; the first 3D human body model is a model including the user's current motion posture;

[0010] Comparing the first 3D human body model with the second 3D human body model to determine whether the user's current motion posture conforms to the standard motion posture; the second 3D human body model is a model that includes the standard motion posture;

[0011] When it is determined that the user's current motion posture does not conform to the standard motion posture, first feedback information is generated; the first feedback information is used to guide the user to adjust the user's current motion posture.

[0012] In a second aspect, the present application provides a computer device comprising: a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the 3D motion training assistance method based on single-view video as described in the first aspect.

[0013] In a third aspect, the present application provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the 3D motion training auxiliary method based on single-view video described in the first aspect.

[0014] In a fourth aspect, the present application provides a computer program product, including a computer program, which, when executed by a processor, implements the 3D motion training assistance method based on single-view video described in the first aspect.

[0015] According to the specific embodiments provided in this application, this application has the following technical effects:

[0016] The present application provides a 3D motion training assistance method and device based on single-view video, the method comprising: extracting the two-dimensional coordinates of each key point of the human body in a first current video frame; determining the depth value of each key point of the human body in the first current video frame by using a depth estimation algorithm in combination with the first current video frame and a first preset number of consecutive historical video frames before the first current video frame; determining the three-dimensional coordinates of each key point of the human body in the first current video frame based on the two-dimensional coordinates and depth values of each key point of the human body in the first current video frame; constructing a first 3D human body model according to the three-dimensional coordinates of each key point of the human body in the first current video frame; comparing the first 3D human body model with a second 3D human body model to determine whether the user's current motion posture conforms to the standard motion posture; the second 3D human body model is a model that includes the standard motion posture; when it is determined that the user's current motion posture does not conform to the standard motion posture, generating first feedback information to guide the user to adjust the user's current motion posture. Through the above-mentioned solution of this application, users do not need to wear special sportswear (with sensors for collecting signals) when using the sports training assistance system. They only need a monocular camera that can shoot user sports videos. This not only reduces user costs, but also makes it more convenient to use. In addition, this application can provide personalized feedback for current non-standard sports postures, guide users to make appropriate adjustments, so as to improve the user's sports training effect. BRIEF DESCRIPTION OF THE DRAWINGS

[0017] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.

[0018] Figure 1 This is a diagram illustrating an application environment of a 3D sports training assistance method based on a single-view video in an embodiment of the present application;

[0019] Figure 2 A flowchart of a 3D motion training assistance method based on single-view video provided in one embodiment of the present application;

[0020] Figure 3 A schematic diagram of a 3D sports training assistance system based on single-view video provided in one embodiment of the present application;

[0021] Figure 4 A schematic diagram of the structure of a computer device provided in one embodiment of the present application. DETAILED DESCRIPTION

[0022] The following will be combined with the drawings in the embodiments of this application to clearly and completely describe the technical solutions in the embodiments of this application. Obviously, the embodiments described are only part of the embodiments of this application, not all of the embodiments. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.

[0023] In order to make the above-mentioned purposes, features and advantages of the present application more obvious and easy to understand, the present application is further described in detail below with reference to the accompanying drawings and specific implementation methods.

[0024] The 3D motion training auxiliary method based on single-view video provided in the embodiment of the present application can be applied to Figure 1 In the application environment shown, the terminal 102 communicates with the server 104 via a network. The data storage system can store data that the server 104 needs to process. The data storage system can be set up separately, integrated on the server 104, or placed on the cloud or other servers. Terminal 102 may send a user motion video to server 104. After receiving the user motion video, server 104 extracts the two-dimensional coordinates of each key point of the human body in a first current video frame; combines the first current video frame with a first predetermined number of consecutive historical video frames before the first current video frame, and uses a depth estimation algorithm to determine the depth value of each key point of the human body in the first current video frame; determines the three-dimensional coordinates of each key point of the human body in the first current video frame based on the two-dimensional coordinates and depth value of each key point of the human body in the first current video frame; constructs a first 3D human body model based on the three-dimensional coordinates of each key point of the human body in the first current video frame; compares the first 3D human body model with a second 3D human body model to determine whether the user's current motion posture conforms to a standard motion posture; the second 3D human body model is a model that includes the standard motion posture; and when it is determined that the user's current motion posture does not conform to the standard motion posture, generates first feedback information to guide the user to adjust the current user's motion posture. Server 104 may provide the generated first feedback information to terminal 102. In addition, in some embodiments, the 3D motion training assistance method based on single-view video can also be implemented independently by the server 104 or the terminal 102. For example, the terminal 102 can directly process the user motion video, or the server 104 can obtain the user motion video from the data storage system and process it.

[0025] Terminal 102 may include, but is not limited to, various desktop computers, laptops, smartphones, tablet computers, IoT devices, and portable wearable devices. IoT devices may include smart speakers, smart TVs, smart air conditioners, and smart car devices. Portable wearable devices may include smart watches, smart bracelets, and head-mounted devices. Server 104 may be implemented as a standalone server or a server cluster consisting of multiple servers, or may be a cloud server.

[0026] In an exemplary embodiment, Figure 2 As shown, this embodiment provides a 3D motion training auxiliary method based on single-view video, which is executed by a computer device, specifically a computer device such as a terminal or a server, or a terminal and a server. In this embodiment of the application, the method is applied to Figure 1 The server 104 in FIG. 1 is taken as an example to illustrate the process, including the following steps 201 to 206. In which:

[0027] Step 201: extract the two-dimensional coordinates of key points of the human body in a first current video frame; the first current video frame is a frame in a user motion video; the user motion video is a video of a user exercising according to a sports training item, shot by a monocular camera.

[0028] Step 202 : Determine the depth values of key points of the human body in the first current video frame by combining the first current video frame and a first preset number of continuous historical video frames before the first current video frame using a depth estimation algorithm.

[0029] Step 203 : determining the three-dimensional coordinates of each key point of the human body in the first current video frame based on the two-dimensional coordinates and depth values of each key point of the human body in the first current video frame.

[0030] Step 204 : constructing a first 3D human body model according to the three-dimensional coordinates of key points of the human body in the first current video frame; the first 3D human body model is a model including the current motion posture of the user.

[0031] Step 205 : Compare the first 3D human body model with the second 3D human body model to determine whether the user's current motion posture conforms to the standard motion posture; the second 3D human body model is a model that includes the standard motion posture.

[0032] Step 206 : When it is determined that the user's current motion posture does not conform to the standard motion posture, first feedback information is generated; the first feedback information is used to guide the user to adjust the user's current motion posture.

[0033] In which, the second 3D human body model is constructed based on the second current video frame in the professional demonstration video; the professional demonstration video is a video of a professional demonstrating the sports training item; the first current video frame and the second current video frame are in the same serial position in their respective video frame sequences.

[0034] In another exemplary embodiment, the single-view video-based 3D motion training assistance method further includes:

[0035] Step 207: Determine the motion trajectory of a key point of the human body based on the 3D human body models corresponding to the first current video frame and a second preset number of consecutive historical video frames before the first current video frame to obtain a first motion trajectory; the key point of the human body is determined based on the sports training item.

[0036] Step 208, using a dynamic time warping algorithm to determine the alignment information of the first motion trajectory and the second motion trajectory on the time axis; the second motion trajectory is the motion trajectory of a key point of the human body determined based on each 3D human body model corresponding to the second current video frame and a second preset number of consecutive historical video frames before the second current video frame.

[0037] Step 209 : Determine whether the user's motion speed meets the standard motion speed based on the alignment information of the first motion trajectory and the second motion trajectory on the time axis.

[0038] Step 210: When it is determined that the user's motion speed does not meet the standard motion speed, second feedback information is generated; the second feedback information is used to guide the user to adjust the user's motion speed.

[0039] As an optional implementation, step 205 specifically includes:

[0040] Step 205.01, calculating the Euclidean distance between each key point of the human body in the first 3D human body model and each key point of the human body in the second 3D human body model, and obtaining the offset distance of each key point of the human body in the first 3D human body model.

[0041] Step 205.02: for any key point, determine whether the offset distance of a key point of the human body in the first 3D human body model exceeds the offset distance threshold corresponding to the key point.

[0042] Step 205.03, wherein, when the offset distance of any key point of the human body in the first 3D human body model exceeds the offset distance threshold corresponding to any key point of the human body, it is determined that the user's current motion posture does not meet the standard motion posture.

[0043] As another optional implementation, step 205 specifically includes:

[0044] Step 205.11, calculating the difference between the bending angle of each human joint in the first 3D human body model and the bending angle of each human joint in the second 3D human body model, to obtain the bending angle error value of each human joint in the first 3D human body model.

[0045] Step 205.12: for any joint, determine whether the bending angle error value of a human joint in the first 3D human body model exceeds a bending angle error threshold corresponding to the joint.

[0046] Step 205.13, wherein, when the bending angle error value of any human joint in the first 3D human body model exceeds the bending angle error threshold corresponding to any human joint, it is determined that the user's current motion posture does not meet the standard motion posture.

[0047] As an optional embodiment, the first feedback information includes graphic prompt information, which is a first 3D human body model and a second 3D human body model superimposed on the display interface, and the parts corresponding to the key points in the user's current motion posture that do not conform to the standard motion posture are marked in the first 3D human body model with a color or line type different from other parts to highlight them.

[0048] As another optional implementation, the first feedback information includes voice prompt information, and the voice prompt information is voice information that guides the user to adjust the user's current movement posture.

[0049] As another optional implementation, the first feedback information includes text prompt information, and the text prompt information is text information that guides the user to adjust the user's current motion posture.

[0050] The present application also provides an application scenario that applies the above-mentioned 3D motion training assistance method based on single-view video. Specifically: The 3D motion training assistance method based on single-view video provided in this embodiment can be applied in a motion training assistance scenario. The motion training assistance scenario includes a user motion video acquisition link and a feedback information generation link; the user motion video enters the feedback information generation link from the user motion video acquisition link. The 3D motion training assistance method based on single-view video provided in this embodiment belongs to the feedback information generation link. Specifically, in the feedback information generation link, the two-dimensional coordinates of each key point of the human body in the first current video frame are extracted; the first current video frame and a first preset number of continuous historical video frames before the first current video frame are combined, and the depth value of each key point of the human body in the first current video frame is determined using a depth estimation algorithm; based on the two-dimensional coordinates and depth values of each key point of the human body in the first current video frame, the three-dimensional coordinates of each key point of the human body in the first current video frame are determined; according to the three-dimensional coordinates of each key point of the human body in the first current video frame, a first 3D human body model is constructed; the first 3D human body model is compared with the second 3D human body model to determine whether the user's current motion posture meets the standard motion posture; when it is determined that the user's current motion posture does not meet the standard motion posture, first feedback information is generated to guide the user to adjust the user's current motion posture.

[0051] In another exemplary embodiment, Figure 3 As shown, this application provides a 3D motion training system for implementing a 3D motion training assistance method based on single-view video. The system consists of two parts, hardware and software, specifically hardware modules and software modules. The hardware module includes a high-performance computer and a camera; the software module includes a posture estimation module, a 3D display module, a real-time feedback module, and a user interaction module.

[0052] The system can achieve 3D posture estimation through single-view video and provide real-time motion analysis and feedback. The system converts 2D posture data into a 3D skeleton model and a standard human body model (rendered from the 3D skeleton model) and displays it in 360°. The system can convert the skeleton model into a standard human body model for display, as well as provide real-time motion analysis and feedback, thereby helping users better understand and learn sports skills. The system is not only suitable for athlete training, but can also be used in fields such as rehabilitation and assessments that require standard movements. It provides movement correction and assessment functions through comparative analysis to help users quickly master the key points of the movements. By monitoring the trainees' movements in real time through cameras and comparing them with standard movements, real-time feedback and movement correction can be achieved, improving the real-time and effectiveness of training.

[0053] Existing sports training, rehabilitation, and movement assessment systems mostly rely on the guidance of professional coaches, making automation and real-time feedback difficult to achieve. While significant progress has been made in pose estimation and 3D display technologies, most solutions either rely on expensive hardware or have limitations in real-time processing and user interaction. In particular, 3D pose estimation from single-view video remains a significant challenge in terms of accuracy and practicality. This system addresses these shortcomings in existing technologies through the following approaches:

[0054] 3D pose estimation from single-view video: This leverages OpenPose's 2D keypoint extraction capabilities and converts them to 3D keypoints through depth estimation algorithms, reducing the need for additional hardware. This eliminates the need for multi-view video and depth sensors, lowering the system's hardware cost and complexity.

[0055] Efficient 3D display and real-time feedback: Three.js is used for efficient 3D skeletal model creation and rendering, enabling 360° display and providing high-quality visual effects. Optimized real-time feedback algorithms are used to improve the accuracy of motion correction.

[0056] User-friendly interface design: Design a simple and intuitive user interface to lower the usage threshold and make sports training, rehabilitation and movement assessment more convenient and efficient.

[0057] The overall architecture of the system is described below.

[0058] 1. Hardware module

[0059] 1.1. High-performance computer: As the core of the system, it runs all algorithm modules (software modules).

[0060] GPUs in high-performance computers provide parallel computing capabilities and support fast reasoning of deep learning models. CPUs in high-performance computers provide efficient data processing and real-time computing.

[0061] 1.2. Camera: collects user motion videos in real time, connects to a computer via USB or network interface, and transmits the video data to the posture estimation module in a high-performance computer.

[0062] 2. Software Module

[0063] 2.1. Posture Estimation Module

[0064] 1) 2D key point extraction: Use the OpenPose algorithm (model) to extract 2D key point data (2D coordinates) from video frames.

[0065] The OpenPose model can identify and extract the main 2D key points of the human body, such as joints (shoulders, elbows, knees), limb positions, and key nodes of the face and hands.

[0066] In addition, based on OpenPose, we can use model compression technologies such as pruning and quantization to reduce the number of model parameters and computational complexity, adopt a more efficient key point matching algorithm, and avoid unnecessary calculations.

[0067] Process: OpenPose takes video data captured by a camera as input, analyzes each frame, and outputs 2D keypoints containing joint coordinates. Before the data is fed into the model, it undergoes preprocessing, such as reducing image resolution and cropping irrelevant areas, to reduce computational effort.

[0068] Key point coverage:

[0069] Upper limbs: shoulders, elbows, wrists;

[0070] Lower limbs: hips, knees, ankles;

[0071] Core parts: head, neck, chest, and pelvis.

[0072] 2) 3D key point conversion: Using the depth estimation algorithm, the 2D key point data (2D coordinates) are converted into 3D coordinates and transmitted to the 3D display module.

[0073] That is, the 2D key points are converted into 3D key points through the depth estimation algorithm.

[0074] Since a single camera can only capture two-dimensional plane information, in order to obtain 3D key points, a depth estimation algorithm is needed to infer the depth information (Z-axis coordinate) of the key points.

[0075] You can complete the 2D-3D conversion by using either of the following two methods:

[0076] Method 1: Convolutional Neural Network (CNN) and Regression Model

[0077] Use pre-trained models such as AlphaPose, VNect, or the OpenPose 3D extension module to perform spatial inference on 2D keypoints.

[0078] Input: A sequence of 2D keypoints in a series of consecutive video frames.

[0079] Output: 3D coordinate sequence corresponding to the 2D key point sequence.

[0080] Method 2: Motion trajectory estimation and time series analysis

[0081] By using time series data (2D key point sequence) in the video, the depth change is estimated through algorithms such as recurrent neural networks (RNN) to better capture the dynamic characteristics of the action.

[0082] Input: A sequence of 2D keypoints in a series of consecutive video frames.

[0083] Output: 3D coordinate sequence corresponding to the 2D key point sequence.

[0084] In addition, in terms of depth estimation, a more accurate depth estimation model can be designed by combining prior knowledge, such as human bone length constraints, and adopting graph optimization technology to improve the accuracy and robustness of depth estimation.

[0085] 2.2, 3D display module

[0086] 1) 3D skeleton model creation: Create a 3D skeleton model based on the 3D key point data transmitted by the posture estimation module. This embodiment uses Three.js to create a 3D skeleton model.

[0087] 2) Keypoint mapping: Map the generated 3D keypoints to the 3D skeleton model in Three.js and render it as a standard 3D human body model.

[0088] The purpose of keypoint mapping is to associate the detected keypoint coordinates with the corresponding nodes in the 3D skeleton model. This process involves the following steps:

[0089] Coordinate transformation:

[0090] Since the detected key point coordinate system may be different from the 3D skeleton model coordinate system, coordinate conversion is required to convert the key point coordinates into the 3D skeleton model coordinate system.

[0091] Key points correspond to:

[0092] Determine the correspondence between the detected keypoints and the nodes in the 3D skeletal model. For example, the detected "right elbow" keypoint corresponds to the "right elbow" node in the model.

[0093] Model Update:

[0094] According to the mapped key point coordinates, the positions of the nodes in the 3D skeleton model are updated to achieve the reconstruction of the human body posture.

[0095] 3) 360° display and interaction: Provides functions such as 3D model rotation and zooming, allowing users to view action details from different angles.

[0096] The 3D display module transmits the current user's 3D posture data (i.e., the first 3D human body model), and the real-time feedback module compares and analyzes it with the standard action template (i.e., the second 3D human body model).

[0097] 2.3 Real-time feedback module

[0098] 1) Action Comparison: The user's movements (i.e., the first 3D human body model) transmitted by the 3D display module are compared with standard movement templates. This means comparing the user's real-time posture with pre-recorded standard movements. The system stores 3D movement sequences from professional athletes or standard demonstrations as reference templates for users to practice with.

[0099] The following explains the specific comparison method and the algorithm for quantifying the deviation between user actions and standard actions.

[0100] a. Key point position comparison: Compare the user's real-time key points with the key points of the standard action frame by frame.

[0101] Euclidean distance calculation: Calculate the offset distance of each key point in space.

[0102] For example, if the position difference between the user's knee joint and the standard knee joint in 3D space is ddd, the system records this error.

[0103] b. Motion trajectory comparison: Analyze whether the motion trajectory of key points conforms to the trajectory of standard actions.

[0104] Dynamic Time Warping (DTW) algorithm: Compares the entire action sequence to identify inconsistencies in the timeline, such as speed or rhythm of the action.

[0105] This comparison method is generally suitable for sports that require precise timing control.

[0106] c. Angle comparison: Calculate the difference between the joint angle change (such as knee flexion, arm lifting angle) and the standard value.

[0107] Joint angle error calculation: The system detects the angle change of each major joint (such as the elbow flexion angle) and compares it with the ideal angle in the standard movement.

[0108] The formula for joint angle error is: Δθ=|θ user -θ standard |, Δθ represents the joint angle error, θ user represents the user's joint angle, θ standard Indicates the joint angle in the standard action (i.e. standard value).

[0109] Furthermore, we can comprehensively evaluate user movements and train different motion error analysis models based on different sports by comprehensively considering multiple dimensions, such as key point errors, joint angle errors, and motion trajectory errors. For example, in yoga training, we can focus more on joint angle errors. In dance training, we can prioritize motion trajectory errors. In running, we can give higher weight to knee and ankle joint errors.

[0110] 2) Feedback information generation: Analyze the comparison results, generate feedback information, and provide real-time feedback to help users adjust or correct their actions.

[0111] The system uses a camera to capture the user's real-time motion data (in real-time mode) and compares and analyzes it with a standard motion template. Based on the error analysis results, the system provides multi-dimensional real-time feedback.

[0112] a. Feedback form:

[0113] Graphic feedback: The user's 3D human body model and standard motion model are superimposed on the interface, and the different parts are highlighted with different colors or dotted lines.

[0114] Voice prompts: Use voice to inform users of specific areas that need improvement (e.g., "Please raise your arms").

[0115] Text suggestions: The system interface updates with improvement suggestions in real time, such as "Keep your knees bent" and "Relax your shoulders".

[0116] b. Type of movement deviation and correction suggestions (i.e. feedback information)

[0117] Deviation Type:

[0118] Position Deviation: Detects the offset of the user's key points relative to the standard action. This is determined by comparing the key point positions.

[0119] Timing deviation: Analyzes the alignment between the time series of the user's completed actions and the time series of the standard actions. This is determined by comparing the user's movement trajectory.

[0120] Joint Angle Deviation: Calculates the difference in joint angles between the user's motion and the standard motion. This is determined through angle comparison.

[0121] For example, if the knee bending angle should be 45°, but the user only bends 30°, the angle deviation is large, and the system will suggest "bend the knee another 15°".

[0122] Example of corrective suggestions:

[0123] Position Adjustment: Prompts the user to “move forward a step” or “stand up straighter.”

[0124] Speed adjustment: If the user moves too fast or too slow, the system prompts "Adjust movement speed".

[0125] Posture Correction: If your shoulders or spine are not in the correct posture, the system will suggest "Keep your back straight."

[0126] 3) Dynamic adjustment and continuous feedback

[0127] Frame-by-frame comparison and feedback: The system analyzes the user's posture at every frame and provides new feedback as the posture changes.

[0128] Gradual correction: When the user approaches the standard action, the system provides encouraging prompts, such as "Good, keep it up for another 5 seconds." In other words, the feedback information includes not only corrective suggestions but also encouragement.

[0129] Intelligent fault tolerance: For beginners, the system automatically adjusts the feedback frequency and difficulty based on the user's proficiency to avoid learning pressure caused by over-correction.

[0130] Finally, the real-time feedback module transmits the generated feedback information to the user interaction module, which is displayed on the interface in real time. That is, the real-time feedback module generates feedback information (such as text, graphics or voice prompts) and displays it on the user interface through the user interaction module.

[0131] Advantages of real-time feedback:

[0132] Accurately locate problems: The system can quickly locate specific motion errors and provide targeted adjustment suggestions.

[0133] Immediacy: Feedback delay is controlled within 100 milliseconds to ensure that users can obtain corrective information in a timely manner.

[0134] User participation: Combining multiple feedback forms enhances user immersion and learning motivation.

[0135] Through this real-time feedback system, users can practice and adjust at the same time, continuously improve their movements during practice, shorten the learning curve and improve training effects.

[0136] 2.4 User Interaction Module

[0137] 1) Interface Design and Data Interaction: Users can upload videos (offline mode) or select real-time mode to view 3D models and provide feedback. A user-friendly interface allows users to upload videos, view 3D models, and receive feedback. Furthermore, users can rotate, zoom, and pan the view to observe detailed movements in the 3D display.

[0138] 2) Interaction with other modules: When a user adjusts the viewing angle or uploads a new video, the interface transmits the request to the relevant module (such as the 3D display module) for real-time display updates. This means that when a user uploads a video, adjusts the viewing angle, or rotates the 3D model through the interface, the interaction request is transmitted to the corresponding module for update.

[0139] The system embodiment has the following beneficial effects:

[0140] 1) Efficient pose estimation and 3D transformation

[0141] High efficiency: Using OpenPose for pose estimation has high computational efficiency. It can quickly extract 2D key points in real-time video and quickly convert them into 3D key points through depth estimation algorithms.

[0142] OpenPose algorithm: OpenPose is an efficient pose estimation algorithm that can quickly extract 2D key points in single-view videos and achieve real-time processing through an optimized deep learning model.

[0143] Depth estimation algorithm: By using an optimized depth estimation algorithm, 2D key point data is quickly converted into 3D key points, reducing computational complexity and time cost.

[0144] 2) High precision and visualization of 3D display

[0145] High precision: The 3D skeletal model and key point mapping created using Three.js can accurately display the user's movements and provide high-quality 3D visualization effects.

[0146] Strong intuitiveness: The 360° 3D display function enables users to view movement details from different angles, enhancing movement understanding and learning effects.

[0147] Effect source:

[0148] Three.js framework: Three.js is a powerful 3D rendering tool that can efficiently create and render 3D models and provide realistic visualization effects.

[0149] Key point mapping technology: Through a precise key point mapping algorithm, the 3D key point data extracted by OpenPose is accurately mapped to the 3D skeletal model, ensuring display accuracy.

[0150] 3) Real-time feedback and action correction

[0151] Strong real-time performance: The system can monitor and compare the user's movements with standard movements in real time while the user is exercising, providing immediate feedback and corrective suggestions.

[0152] Good correction effect: Through detailed feedback information, users can quickly identify and correct movement errors and improve training effects.

[0153] Effect source:

[0154] Real-time pose estimation and comparison algorithm: Leveraging the efficiency of OpenPose and depth estimation algorithms, the system can extract and process user motion data in real time and compare it with standard motions.

[0155] Feedback generation algorithm: Generates specific feedback information through sophisticated difference calculation and analysis algorithms to help users accurately adjust and improve their actions.

[0156] 4) User-friendly interactive interface

[0157] Easy to operate: The system provides a friendly user interface, allowing users to easily upload videos, view 3D models and provide feedback. The interface interaction is intuitive and easy to use.

[0158] Low learning cost: The system design reduces the user's learning cost, making it easy for non-professional users to use. It is widely applicable to multiple fields such as sports, rehabilitation and assessment.

[0159] Effect source:

[0160] Interaction design optimization: Through the optimization of human-computer interaction design, the system interface is concise and clear, the functional layout is reasonable, and the user experience is good.

[0161] Multi-function integration: It integrates multiple functions such as video uploading, 3D display, real-time feedback, etc. Users do not need to switch between multiple platforms, making it easy to use.

[0162] Advantages analysis of the system embodiment:

[0163] The system has significant advantages in efficiency, accuracy, real-time performance, and user-friendliness. These advantages are derived from the integration and optimization of multiple key technical points:

[0164] The OpenPose algorithm provides efficient 2D key point extraction, and the depth estimation algorithm provides efficient 3D conversion, improving the system's computational efficiency and real-time performance.

[0165] Three.js framework: High-quality 3D model creation and rendering, providing accurate motion display and visualization effects.

[0166] Real-time feedback algorithm: Accurate movement matching and feedback generation ensure that users can receive corrective suggestions immediately to improve training results.

[0167] User interface design: Simple and intuitive interface design and multi-function integration reduce the difficulty of use.

[0168] Through the organic combination and optimization of these technical points, the system can provide users with efficient, accurate, real-time and user-friendly sports training and motion analysis solutions. In particular, it provides an efficient, accurate and convenient training and analysis tool for athletes, rehabilitation patients and users who need movement assessment.

[0169] In an exemplary embodiment, a computer device is provided. The computer device may be a server or a terminal. The internal structure diagram thereof may be as follows: Figure 4 As shown. The computer device includes a processor, a memory, an input / output interface (Input / Output, abbreviated as I / O) and a communication interface. The processor, memory and input / output interface are connected through a system bus, and the communication interface is connected to the system bus through the input / output interface. The processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program and a database. The internal memory provides an environment for the operation of the operating system and computer program in the non-volatile storage medium. The database of the computer device is used to store user motion videos. The input / output interface of the computer device is used to exchange information between the processor and an external device. The communication interface of the computer device is used to communicate with an external terminal through a network connection. When the computer program is executed by the processor, a 3D motion training auxiliary method based on single-view video is implemented.

[0170] Those skilled in the art will understand that Figure 4 The structure shown in the figure is merely a block diagram of a portion of the structure related to the solution of the present application and does not constitute a limitation on the computer device to which the solution of the present application is applied. A specific computer device may include more or fewer components than shown in the figure, or combine certain components, or have a different component arrangement. In an exemplary embodiment, a computer device is provided, including a memory and a processor. The memory stores a computer program, and the processor implements the steps of the above-mentioned method embodiments when executing the computer program.

[0171] In an exemplary embodiment, a computer-readable storage medium is provided, storing a computer program. When the computer program is executed by a processor, the steps in the above-mentioned method embodiments are implemented.

[0172] In an exemplary embodiment, a computer program product is provided, including a computer program. When the computer program is executed by a processor, the steps in the above method embodiments are implemented.

[0173] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of relevant data must comply with relevant regulations.

[0174] Those skilled in the art will appreciate that all or part of the processes in the above-mentioned embodiment methods can be implemented by instructing the relevant hardware through a computer program, and the computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above-mentioned methods. Among them, any reference to memory, database or other media used in the embodiments provided in this application may include at least one of non-volatile and volatile memory. Non-volatile memory may include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetic random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory may include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM may be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM).

[0175] The databases involved in the various embodiments provided herein may include at least one of a relational database and a non-relational database. Non-relational databases may include, but are not limited to, distributed databases based on blockchains. The processors involved in the various embodiments provided herein may include, but are not limited to, general-purpose processors, central processing units, graphics processing units, digital signal processors, programmable logic units, data processing logic units based on quantum computing, and the like.

[0176] The technical features of the above embodiments can be combined arbitrarily. To make the description concise, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.

[0177] This document uses specific examples to illustrate the principles and implementation methods of this application. The description of the above examples is only intended to help understand the method and core concept of this application. At the same time, for those skilled in the art, based on the concept of this application, there may be changes in the specific implementation methods and application scope. In summary, the content of this specification should not be understood as limiting this application.

Claims

1. A 3D motion training auxiliary method based on single-view video, characterized in that: The 3D motion training auxiliary method based on single-view video includes: Extracting two-dimensional coordinates of key points of a human body in a first current video frame; the first current video frame is a frame in a user motion video; the user motion video is a video of a user exercising according to a sports training program captured by a monocular camera; Determine the depth values of key points of the human body in the first current video frame by combining the first current video frame and a first preset number of consecutive historical video frames before the first current video frame using a depth estimation algorithm; Determine the three-dimensional coordinates of each key point of the human body in the first current video frame based on the two-dimensional coordinates and depth values of each key point of the human body in the first current video frame; Constructing a first 3D human body model based on the three-dimensional coordinates of key points of the human body in the first current video frame; the first 3D human body model is a model including the user's current motion posture; Comparing the first 3D human body model with the second 3D human body model to determine whether the user's current motion posture conforms to the standard motion posture; the second 3D human body model is a model that includes the standard motion posture; When it is determined that the user's current motion posture does not conform to the standard motion posture, first feedback information is generated; the first feedback information is used to guide the user to adjust the user's current motion posture.

2. The 3D motion training auxiliary method based on single-view video according to claim 1, characterized in that: The second 3D human body model is constructed based on a second current video frame in a professional demonstration video; the professional demonstration video is a video of a professional demonstrating the exercise training item; the first current video frame and the second current video frame are at the same position in their respective video frame sequences; The 3D motion training auxiliary method based on single-view video also includes: determining a motion trajectory of a key point of a human body based on the first current video frame and the 3D human body models corresponding to a second preset number of consecutive historical video frames before the first current video frame, to obtain a first motion trajectory; the key point of the human body being determined based on the sports training item; Determining alignment information between the first motion trajectory and the second motion trajectory on a time axis using a dynamic time warping algorithm; the second motion trajectory is a motion trajectory of a key point of the human body determined based on 3D human body models corresponding to a second current video frame and a second preset number of consecutive historical video frames before the second current video frame; determining, based on alignment information of the first motion trajectory and the second motion trajectory on the time axis, whether the user's motion speed meets the standard motion speed; When it is determined that the user's motion speed does not meet the standard motion speed, second feedback information is generated; the second feedback information is used to guide the user to adjust the user's motion speed.

3. The 3D motion training auxiliary method based on single-view video according to claim 1, characterized in that: Comparing the first 3D human body model with the second 3D human body model to determine whether the user's current motion posture conforms to the standard motion posture, specifically including: Calculating the Euclidean distance between each key point of the human body in the first 3D human body model and each key point of the human body in the second 3D human body model to obtain an offset distance between each key point of the human body in the first 3D human body model; For any key point, determining whether an offset distance of a key point of the human body in the first 3D human body model exceeds an offset distance threshold corresponding to the key point; When the offset distance of any key point of the human body in the first 3D human body model exceeds the offset distance threshold corresponding to any key point of the human body, it is determined that the current motion posture of the user does not meet the standard motion posture.

4. The 3D motion training auxiliary method based on single-view video according to claim 1, characterized in that: Comparing the first 3D human body model with the second 3D human body model to determine whether the user's current motion posture conforms to the standard motion posture, specifically including: Calculating the difference between the bending angle of each joint of the human body in the first 3D human body model and the bending angle of each joint of the human body in the second 3D human body model to obtain the bending angle error value of each joint of the human body in the first 3D human body model; For any joint, determining whether a bending angle error value of a human joint in the first 3D human body model exceeds a bending angle error threshold corresponding to the joint; When the bending angle error value of any human joint in the first 3D human body model exceeds the bending angle error threshold corresponding to any human joint, it is determined that the user's current motion posture does not meet the standard motion posture.

5. The 3D motion training auxiliary method based on single-view video according to claim 1, characterized in that: The first feedback information includes graphic prompt information, which is a first 3D human body model and a second 3D human body model superimposed on the display interface. The parts corresponding to the key points in the user's current motion posture that do not conform to the standard motion posture are marked in the first 3D human body model with a color or line type different from other parts to highlight them.

6. The 3D motion training auxiliary method based on single-view video according to claim 1, characterized in that: The first feedback information includes voice prompt information, and the voice prompt information is voice information that guides the user to adjust the user's current movement posture.

7. The 3D motion training auxiliary method based on single-view video according to claim 1, characterized in that: The first feedback information includes text prompt information, and the text prompt information is text information that guides the user to adjust the user's current exercise posture.

8. A computer device comprising: A memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the 3D motion training assistance method based on single-view video according to any one of claims 1 to 7.

9. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the 3D motion training auxiliary method based on single-view video according to any one of claims 1 to 7 is implemented.

10. A computer program product comprising a computer program, characterized in that When the computer program is executed by a processor, the 3D motion training auxiliary method based on single-view video according to any one of claims 1 to 7 is implemented.

Citation Information

Patent Citations

  • Human body action comparison system based on human body posture obtaining system

    CN110109532A