A digital human reconstructed motion scoring method, system, device and medium
By extracting skeletal key points from user videos and standard videos, calculating the angle difference and rhythm score, and generating two-dimensional and three-dimensional digital human models, the high cost, low efficiency and lack of objectivity of traditional scoring methods are solved, and precise motion scoring and movement adjustment guidance are achieved.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- SUN YAT SEN UNIV
- Filing Date
- 2023-08-03
- Publication Date
- 2026-05-08
AI Technical Summary
Traditional methods for scoring human posture are costly and inefficient. The use of sensors can interfere with human activity, are complex to install, lack objectivity and accuracy, have high space requirements, cannot provide a score for the standard of posture for different parts of the human body, and lack a visual virtual human for users to compare.
The algorithm extracts skeleton key points from user videos and standard videos by object detection, calculates the difference in the angle between the lines connecting the key points, combines audio analysis to obtain a rhythm score, generates two-dimensional and three-dimensional digital human models for fusion, and provides a refined motion score.
It enables detailed evaluation of the movement of various body parts of the user, rhythm scoring, and provides intuitive 2D fusion video and dynamic 3D model for users to adjust their movements, reducing equipment costs and complexity, and improving the accuracy and objectivity of the scoring.
Smart Images

Figure CN117095457B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of motion posture evaluation technology, and in particular to a method, system, device and medium for digital human motion reconstruction scoring. Background Technology
[0002] Traditional methods for scoring human posture primarily rely on expensive motion capture equipment. This requires real-time monitoring and data collection of the subject, acquiring information such as acceleration and angular velocity. Algorithms then analyze and process this data to assess posture. This approach is costly and inefficient. The use of sensors and other motion capture equipment can interfere with human movement, affecting measurement accuracy. The installation and debugging of the equipment and sensors require specialized technicians, a complex process that increases operating costs. Furthermore, it demands significant spatial resources, requiring the scanning and reconstruction of the human form and trajectory. This process demands substantial computational resources and time, placing high demands on hardware. In addition, traditional scoring methods often rely on human observation and judgment of the virtual human's posture, lacking objectivity and accuracy.
[0003] Movement posture evaluation is an important health detection method with broad application prospects. Currently, there are many existing technologies similar to the technology of this invention, the most common of which are sensor-based movement posture evaluation technology and laser-based movement posture evaluation technology. After reviewing domestic patent databases and related academic research, the following existing technologies similar to the technology of this invention are selected: (1) a device and method for capturing human movement posture; (2) an ankle joint movement posture acquisition device; (3) a video-based movement posture recognition method and system; (4) a movement recognition method and wearable device; (5) a laser human body sensor.
[0004] Based on the above brief description of the existing technologies, we can see that the existing technologies have achieved human posture capture, but have not addressed the following issues: (1) The device needs to be closely attached to the body surface, which will cause some interference to human activities and affect the accuracy of the measurement. (2) The installation and debugging of the sensors are relatively complex and require professional technicians to operate, which increases the cost of use. (3) No posture standard score is given for each part of the human body. (4) No visual virtual human is generated for users to compare the standard of their movements and postures. (5) Scanning and reconstruction are carried out in a specific environment, which requires a relatively large space. Summary of the Invention
[0005] The purpose of this invention is to propose a digital human reconstruction motion scoring method, system, device, and medium to provide a detailed evaluation of the motion of various body parts of a user.
[0006] To achieve the above objectives, in a first aspect, embodiments of the present invention provide a digital human reconstruction motion scoring method, the method comprising:
[0007] User skeleton key points and standard skeleton key points are obtained from user video and standard video respectively according to the target detection algorithm. The user video is a user motion posture video and the standard video is a standard motion posture video.
[0008] Calculate the first included angle of the line connecting the key points of the user skeleton and the second included angle of the line connecting the key points of the standard skeleton respectively, and calculate the score of each key point based on the difference between the first included angle and the second included angle;
[0009] The total score of the exercise is obtained by multiplying the scores of each key point by the preset weights corresponding to each key point and summing them.
[0010] Furthermore, after obtaining the total sports score, the process also includes:
[0011] Audio is extracted from the user's video according to an audio extraction algorithm;
[0012] The audio track is analyzed using a deep neural network to obtain the replay time.
[0013] Calculate the first included angle of the line connecting the key points of the user skeleton at the retake time and the second included angle of the line connecting the key points of the standard skeleton at the retake time, and calculate the score of each key point at the retake time based on the difference between the first included angle and the second included angle.
[0014] The rhythm score is obtained by multiplying the scores of each key point at the retake moment by the preset weights corresponding to each key point and summing them.
[0015] Furthermore, the step of obtaining user skeleton key points and standard skeleton key points from user videos and standard videos respectively according to the object detection algorithm includes:
[0016] Human body regions in user videos and standard videos were identified using the Yolov3 algorithm.
[0017] The Openpose algorithm is used to obtain user skeleton key points and standard skeleton key points from the human body region.
[0018] Furthermore, after obtaining the total sports score, the process also includes:
[0019] Human body reconstruction is performed on the standard video using the Vibe algorithm to obtain a standard two-dimensional digital human video;
[0020] The standard two-dimensional digital human video is superimposed on the user video to obtain a two-dimensional fused video.
[0021] Furthermore, after obtaining the total sports score, the process also includes:
[0022] Human body reconstruction was performed on user videos and standard videos using the Vibe algorithm to establish user 3D digital human models and standard 3D digital human models.
[0023] The user's 3D digital human model and the standard 3D digital human model are fused to obtain a dynamic fusion model.
[0024] Furthermore, the step of establishing a user 3D digital human model and a standard 3D digital human model based on the user video and standard video respectively using the Vibe algorithm includes:
[0025] Adjust the height and body shape parameters of the SMPL model based on the user video and the standard video.
[0026] Furthermore, after obtaining the dynamic fusion model, the process includes:
[0027] The dynamic fusion model at each time step is arranged sequentially to obtain a time series model;
[0028] The time series models of multiple users are fused together in one scene to obtain the action models of multiple users.
[0029] Secondly, embodiments of the present invention provide a digital human motion reconstruction scoring system, the system comprising:
[0030] The key point acquisition module is used to acquire user skeleton key points and standard skeleton key points from user videos and standard videos respectively according to the target detection algorithm. The user video is a user motion posture video and the standard video is a standard motion posture video.
[0031] The key point score calculation module is used to calculate the first included angle of the line connecting key points of the user skeleton and the second included angle of the line connecting key points of the standard skeleton, and calculate the score of each key point based on the difference between the first included angle and the second included angle.
[0032] The sports scoring module is used to multiply the scores of each key point by the preset weights corresponding to each key point and sum them to obtain the total sports score.
[0033] Thirdly, embodiments of the present invention also provide a computer device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the steps of the above-described method.
[0034] Fourthly, embodiments of the present invention also provide a computer-readable storage medium having a computer program stored thereon, wherein the computer program, when executed by a processor, implements the steps of the above-described method.
[0035] This invention provides a method, system, device, and medium for motion scoring in digital human reconstruction. It involves obtaining key points of the user skeleton and key points of the standard skeleton from a user video and a standard video, respectively, based on a target detection algorithm. The user video is a video of the user's motion posture, and the standard video is a video of a standard motion posture. A first angle and a second angle are calculated between the lines connecting the key points of the user skeleton and the standard skeleton. A score is calculated for each key point based on the difference between the first and second angles. The scores of each key point are multiplied by a preset weight corresponding to that key point, and the results are summed to obtain a total motion score. This invention enables a detailed evaluation of the motion of various body parts of a user. Attached Figure Description
[0036] Figure 1 This is a flowchart illustrating a motion scoring method for digital human reconstruction provided in an embodiment of the present invention;
[0037] Figure 2 This is a circular diagram showing the fractions of each key point provided in the embodiments of the present invention;
[0038] Figure 3 This is a schematic diagram of the action model for selecting a specific user provided in an embodiment of the present invention;
[0039] Figure 4 This is a schematic diagram of an action model for selecting a specific moment provided in an embodiment of the present invention;
[0040] Figure 5 This is a system block diagram of a digital human motion scoring system provided in an embodiment of the present invention;
[0041] Figure 6 This is an internal structural diagram of the computer device in an embodiment of the present invention. Detailed Implementation
[0042] To make the objectives, technical solutions, and beneficial effects of this application clearer, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. Obviously, the embodiments described below are only part of the embodiments of the present invention and are used to illustrate the present invention, but are not intended to limit the scope of the present invention. All other embodiments obtained by those skilled in the art based on the embodiments of the present invention without creative effort are within the scope of protection of the present invention.
[0043] In one embodiment, such as Figure 1 As shown, a method for motion scoring in digital human reconstruction is provided, the method comprising:
[0044] S11. Obtain user skeleton key points and standard skeleton key points from user video and standard video respectively according to target detection algorithm. The user video is a user motion posture video and the standard video is a standard motion posture video.
[0045] In this embodiment, the uploaded original video is first preprocessed, such as by removing noise and filtering. The process of obtaining skeleton keypoints involves: identifying human body regions in the user video and standard video using the YOLOv3 algorithm; and obtaining user skeleton keypoints and standard skeleton keypoints from these human body regions using the OpenPose algorithm. Upon obtaining the skeleton keypoints, confidence estimation is performed, returning an array of the skeleton keypoint coordinates and their corresponding confidence estimates. In subsequent human movement, Kalman filtering is used to predict the detection box position, and adjustments are made based on the human detection results to continue tracking the human body region. After detection, based on the detection box position, OpenCV is used to scale each frame of the individual image to the same size, cropping and stitching them frame by frame. If the image is below a set threshold, such as approximately 85% of the total duration, it is discarded to avoid interference from users outside the classroom. In fact, if skeleton keypoints are identified without first extracting the human body region, irrelevant information may be misidentified as skeleton keypoints, reducing recognition efficiency. Therefore, in this embodiment, the human body region is extracted first, and then the key points of the skeleton are parsed from the human body region, which can improve the efficiency of obtaining the key points of the skeleton and reduce invalid extraction.
[0046] In this embodiment, face recognition and face matching can also be performed. Specifically, firstly, face data is associated with names and placed in a database through registration. After processing, the user-uploaded video becomes a segmented single-person video. The segmented single-person video is then frame-by-frame extracted to obtain frame images. The Insightface open-source library is used for feature point matching, extracting one frame every 0.5 seconds, and recognizing a face image in each frame. Next, the detected faces are corrected and aligned for subsequent feature extraction and matching. A deep learning model CNN is used to extract the feature vectors of the faces. These are compared with the face information in the database to obtain the name of the most matching face. The number of times the face name is returned for each frame is integrated and statistically analyzed, and the identity information of the person whose non-empty name appears most frequently after integration and statistical analysis is recorded as the face recognition result. Then, the Insightface open-source library is used to identify facial key points in the first image that matches the frame image and the face recognition result, thereby performing face segmentation, obtaining and storing it as an avatar image. In typical applications such as sports and dance, the database will contain a large amount of user data from multiple classes. This embodiment can distinguish and match information from different users through facial recognition. Additionally, user videos sometimes contain the movement postures of multiple people. This embodiment can separate each person's movement posture and match it with the corresponding user in the database, eliminating the need for manual segmentation and better reflecting real-world application scenarios.
[0047] S12. Calculate the first included angle of the line connecting the key points of the user skeleton and the second included angle of the line connecting the key points of the standard skeleton respectively, and calculate the score of each key point based on the difference between the first included angle and the second included angle.
[0048] In this embodiment, the scores for each key point are calculated using the following formula.
[0049]
[0050] Where, score is the score for each key point, dis is the difference between the first and second included angles, and α is defined according to the degree to which the joints of the human limbs can open and close: α 脖子 =90°, α 胯部 =60°, α 其余 =180°. The resulting scores for different body parts, such as... Figure 2 An example is shown below. Figure 2 The key point scores for each part of the user's body are displayed intuitively to the user in a pie chart format. Typically, different users have different physical conditions and athletic abilities, resulting in varying performance of different body parts. This embodiment can quantitatively evaluate the movement posture of each body part through key point scores, meeting the personalized exercise needs of users.
[0051] S13. Multiply the scores of each key point by the preset weights corresponding to each key point and sum them to obtain the total motion score;
[0052] To evaluate a user's movement posture from a rhythmic perspective, this embodiment extracts audio from the user's video using an audio extraction algorithm; it then performs track analysis on the audio using a deep neural network to obtain the stressed beat moments; it calculates the first angle between the lines connecting key points of the user's skeleton at the stressed beat moment and the second angle between the lines connecting key points of the standard skeleton at the stressed beat moment, and calculates the score for each key point at the stressed beat moment based on the difference between the first and second angles; finally, it multiplies each key point score at the stressed beat moment by a preset weight corresponding to each key point and sums them to obtain a rhythm score. Specifically, it can also evaluate based on rhythm types such as 4 / 4, 2 / 4, and 1 / 3 in the audio. Besides the fit of the movement posture, rhythm is also an important indicator for evaluating user movement. This embodiment achieves accurate evaluation of the user's movement posture at stressed beat moments through a rhythm score.
[0053] In this embodiment, a 2D fused video can be generated through the following process: Human body reconstruction is performed on the standard video using the Vibe algorithm to obtain a standard 2D digital human video; the standard 2D digital human video is then superimposed on the user video to obtain a 2D fused video. Specifically, using Vibe human body reconstruction technology, a 2D digital human RGB video with standard movements is obtained. This standard-movement digital human is then superimposed onto the pre-cut personal video, ensuring center alignment, to obtain a fused MP4 video of the standard virtual human and the user. In fact, digital scoring alone cannot provide users with sufficient information on how to adjust their posture. This embodiment allows users to directly view the non-overlapping parts of the limbs in the 2D fused video between the standard digital human and the personal video, intuitively observing any non-standard movements and making corresponding corrections.
[0054] In this embodiment, a dynamic fusion model can be generated through the following process: Human body reconstruction is performed on the user video and standard video using the Vibe algorithm to establish a user 3D digital human model and a standard 3D digital human model. Specifically, the height and body shape parameters of the SMPL model are adjusted based on the user video and the standard video. The user 3D digital human model and the standard 3D digital human model are then fused to obtain a dynamic fusion model. For the generated 3D model, a Savitzky-Golay filter is used to smooth the motion of each node. After obtaining the dynamic fusion model, the dynamic fusion model at each moment is sequentially arranged to obtain a time series model. The time series models of multiple users are then fused in a scene to obtain action models for multiple users. Corresponding guidance suggestions can also be given for different moments. Figure 3As shown, each column represents a time-series model of the entire class for a single student. When student 2 or student 3 is selected, the corresponding column of the user model is raised for easier observation; for example... Figure 4 As shown, each row represents the motion model of all trainees at a certain time point. When time 3 or time 7 is selected, the height of the corresponding row of user models is raised for easier observation. This embodiment can conveniently observe the motion posture of the same user throughout the entire process or at the same time. For example, the motion posture of all users can be selected at the same time to compare and identify those with the best posture performance and those that need improvement, so as to conduct an overall evaluation of the motion process.
[0055] Compared to existing technologies, this invention overlays and compares the reconstructed 3D human body model with the 3D human body model in a standard video, allowing trainees to quickly identify areas of non-standard movement. A group movement array for trainees is designed, allowing teachers to simultaneously view the movements of different students every second, and to select specific seconds or sequences of movements from specific trainees for observation. Users do not need to install or wear motion capture devices such as sensors; they only need to record a video using their mobile phone or other monocular camera's recording function, and select a standard video. This invention compares the skeletal sequence of the human body in the video with the skeletal sequence of the human body in the standard video, calculates the angular differences at each key point, obtains the corresponding angular scores, and derives the normalized scores for each key point. These scores are highly accurate, comprehensive, and computationally efficient. This invention can be applied to warm-up exercises, ball sports, aerobics, dance, rehabilitation training, yoga, and pre-employment training.
[0056] Based on the aforementioned method for motion scoring in digital human reconstruction, this invention also provides a system for motion scoring in digital human reconstruction, such as... Figure 5 As shown, the system includes:
[0057] The key point acquisition module 1 is used to acquire user skeleton key points and standard skeleton key points from user video and standard video respectively according to target detection algorithm. The user video is a user motion posture video and the standard video is a standard motion posture video.
[0058] Key point score calculation module 2 is used to calculate the first included angle of the line connecting key points of the user skeleton and the second included angle of the line connecting key points of the standard skeleton, and calculate the score of each key point based on the difference between the first included angle and the second included angle.
[0059] The sports scoring module 3 is used to multiply the scores of each key point by the preset weights corresponding to each key point and sum them to obtain the total sports score.
[0060] For specific limitations regarding a digital human motion reconstruction scoring system, please refer to the limitations of a digital human motion reconstruction scoring method described above, which will not be repeated here. Each module in the above system can be implemented entirely or partially through software, hardware, or a combination thereof. These modules can be embedded in or independent of the processor in a computer device, or stored in the memory of a computer device as software, so that the processor can call and execute the corresponding operations of each module.
[0061] In one embodiment, a computer device is provided, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the steps of the method described above.
[0062] Figure 6 This diagram illustrates the internal structure of a computer device in one embodiment, which may specifically be a terminal or a server. The computer device includes a processor, memory, a network interface, a display, and input devices connected via a system bus. The processor provides computing and control capabilities. The memory includes non-volatile storage media and internal memory. The non-volatile storage media stores the operating system and computer programs. The internal memory provides an environment for the operation of the operating system and computer programs stored in the non-volatile storage media. The network interface is used to communicate with external terminals via a network connection. The display screen may be a liquid crystal display (LCD) or an e-ink display. The input devices may be a touch layer covering the display screen, buttons, a trackball, or a touchpad mounted on the computer device's casing, or an external keyboard, touchpad, or mouse.
[0063] Those skilled in the art will understand that Figure 6 The structure shown is merely a block diagram of a portion of the structure related to the present application and does not constitute a limitation on the computer device to which the present application is applied. Specifically, the computing device may include more or fewer components than shown in the diagram, or combine certain components, or have the same component arrangement.
[0064] In one embodiment, a computer-readable storage medium is provided having a computer program stored thereon, which, when executed by a processor, implements the steps of the above method.
[0065] In summary, this invention provides a method, system, device, and medium for motion scoring in digital human reconstruction. The method includes: obtaining user skeleton key points and standard skeleton key points from a user video and a standard video respectively, based on a target detection algorithm; calculating a first angle and a second angle between the lines connecting the user skeleton key points and the standard skeleton key points; calculating a score for each key point based on the difference between the first and second angles; and multiplying each key point score by a preset weight corresponding to each key point and summing the results to obtain a total motion score. This invention enables a detailed evaluation of the motion of various body parts of a user.
[0066] The various embodiments in this specification are described in a progressive manner. For directly identical or similar parts of the embodiments, refer to each other. Each embodiment focuses on its differences from other embodiments. In particular, the system embodiments are basically similar to the method embodiments, so the description is relatively simple; relevant parts can be referred to the descriptions in the method embodiments. It should be noted that the technical features of the above embodiments can be combined arbitrarily. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as the combination of these technical features does not contradict each other, it should be considered within the scope of this specification.
[0067] The embodiments described above are merely preferred embodiments of this application, and while the descriptions are specific and detailed, they should not be construed as limiting the scope of the invention patent. It should be noted that those skilled in the art can make various improvements and substitutions without departing from the technical principles of this invention, and these improvements and substitutions should also be considered within the scope of protection of this application. Therefore, the scope of protection of this patent application should be determined by the scope of the claims.
Claims
1. A method for motion scoring in digital human reconstruction, characterized in that, The method includes: User skeleton key points and standard skeleton key points are obtained from user video and standard video respectively according to the target detection algorithm. The user video is a user motion posture video and the standard video is a standard motion posture video. Calculate the first included angle of the line connecting keypoints in the user skeleton and the second included angle of the line connecting keypoints in the standard skeleton, and calculate the score for each keypoint based on the difference between the first and second included angles; wherein, the calculation of the keypoint score is expressed as follows: in, Score for each key point. It is the difference between the first included angle and the second included angle. The degree of opening and closing of the joints of the human limbs; The total motion score is obtained by multiplying the scores of each key point by the preset weights corresponding to each key point and summing the results. After obtaining the total sports score, the process also includes: Audio is extracted from the user's video according to an audio extraction algorithm; The audio track is analyzed using a deep neural network to obtain the replay time. Calculate the first included angle of the line connecting the key points of the user skeleton at the retake time and the second included angle of the line connecting the key points of the standard skeleton at the retake time, and calculate the score of each key point at the retake time based on the difference between the first included angle and the second included angle. The rhythm score is obtained by multiplying the scores of each key point at the retake moment by the preset weights corresponding to each key point and summing them.
2. The method for motion scoring in digital human reconstruction according to claim 1, characterized in that, The step of obtaining user skeleton key points and standard skeleton key points from user videos and standard videos respectively according to the object detection algorithm includes: Human body regions in user videos and standard videos were identified using the Yolov3 algorithm. The Openpose algorithm is used to obtain user skeleton key points and standard skeleton key points from the human body region.
3. The method for motion scoring in digital human reconstruction according to claim 1, characterized in that, After obtaining the total sports score, the process also includes: Human body reconstruction is performed on the standard video using the Vibe algorithm to obtain a standard two-dimensional digital human video; The standard two-dimensional digital human video is superimposed on the user video to obtain a two-dimensional fused video.
4. The method for motion scoring in digital human reconstruction according to claim 1, characterized in that, After obtaining the total sports score, the process also includes: Human body reconstruction was performed on user videos and standard videos using the Vibe algorithm to establish user 3D digital human models and standard 3D digital human models. The user's 3D digital human model and the standard 3D digital human model are fused to obtain a dynamic fusion model.
5. The method for motion scoring in digital human reconstruction according to claim 4, characterized in that, The process of establishing a user's 3D digital human model and a standard 3D digital human model based on user video and standard video respectively using the Vibe algorithm includes: Adjust the height and body shape parameters of the SMPL model based on the user video and the standard video.
6. The method for motion scoring in digital human reconstruction according to claim 4, characterized in that, After obtaining the dynamic fusion model, the following is included: The dynamic fusion model at each time step is arranged sequentially to obtain a time series model; The time series models of multiple users are fused together in one scene to obtain the action models of multiple users.
7. A digital human motion reconstruction scoring system, characterized in that, The system includes: The key point acquisition module is used to acquire user skeleton key points and standard skeleton key points from user videos and standard videos respectively according to the target detection algorithm. The user video is a user motion posture video and the standard video is a standard motion posture video. The keypoint score calculation module is used to calculate the first included angle of the line connecting keypoints in the user skeleton and the second included angle of the line connecting keypoints in the standard skeleton, and to calculate the score of each keypoint based on the difference between the first included angle and the second included angle; wherein, the calculation of the keypoint score is expressed as follows: in, Score for each key point. It is the difference between the first included angle and the second included angle. The degree of opening and closing of the joints of the human limbs; The sports scoring module is used to multiply the scores of each key point by the preset weights corresponding to each key point and sum them to obtain the total sports score; After obtaining the total sports score, the process also includes: Audio is extracted from the user's video according to an audio extraction algorithm; The audio track is analyzed using a deep neural network to obtain the replay time. Calculate the first included angle of the line connecting the key points of the user skeleton at the retake time and the second included angle of the line connecting the key points of the standard skeleton at the retake time, and calculate the score of each key point at the retake time based on the difference between the first included angle and the second included angle. The rhythm score is obtained by multiplying the scores of each key point at the retake moment by the preset weights corresponding to each key point and summing them.
8. A computer device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the steps of the method according to any one of claims 1 to 6.
9. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1 to 6.
Citation Information
Patent Citations
Human motion rhythm comparison system and method based on attitude estimation
CN113255450A
Vision-based motion video fine analysis method and device
CN114550027A