Motion capture and feedback adjustment method and system for smart motion of meta universe

User action data is collected through multiple inertial measurement units, bone node mapping relationship is established, motion capture data is evaluated and optimized in real time, and motion correction instructions are generated using motion feature comparison, which solves the problems of insufficient motion capture accuracy and lack of real-time feedback in the intelligent movement of the metaverse, and improves the immersion and learning effect of users.

CN120123889AActive Publication Date: 2025-06-10HANGZHOU MOXI TECH DEV CO LTD

Patent Information

Application Number
CN202510616042.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-14
Publication Date
2025-06-10
Estimated Expiration
2045-05-14

AI Technical Summary

Technical Problem

The existing meta-universe intelligent motion technology lacks motion accuracy, especially in fast motion scenarios, bone point drift and position errors are prone to occur, resulting in the virtual image movements that do not match the user's actual movements, reducing the user's immersion and interactive experience. At the same time, the lack of effective real-time feedback and correction mechanisms has affected the learning effect and training efficiency of the metacosmic intelligent movement.

Method used

The user's actual action data is collected through multiple inertial measurement units, a mapping relationship with the preset human skeleton nodes is established, and the three-dimensional spatial position coordinates of the skeleton nodes are calculated. A multi-dimensional scoring matrix is ​​constructed based on the motion coherence and action standardization indicators, and bone point position data is evaluated and optimized, and optimized bone point position data is generated to build a user virtual image. Use local and global action features to compare with standard action features to generate angle, velocity and displacement adjustments, and convert them into motion correction instructions to guide users to adjust their motion posture in real time.

Benefits of technology

It improves the accuracy and stability of motion capture, solves the problem that traditional systems are susceptible to interference in complex environments, and enhances the user's immersion and interactive experience. Through real-time feedback and correction mechanisms, the learning effect and training efficiency of the metacosmic intelligent movement are improved, and personalized exercise guidance and feedback are provided to users.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120123889A_ABST
    Figure CN120123889A_ABST
Patent Text Reader

Abstract

The invention provides a motion capture and feedback adjustment method and system for smart motion of meta universe, and relates to the technical field of meta universe, and the method comprises the steps: collecting actual motion data of a user through a plurality of inertial measurement units, establishing a mapping relation with skeleton nodes, evaluating and optimizing the motion based on a multi-dimensional scoring matrix, and constructing a user virtual image. And performing time sequence analysis on the virtual motion data to obtain local and global motion features, comparing the local and global motion features with standard motion features to generate an adjustment amount, and converting the adjustment amount into a motion correction instruction to guide a user to adjust a posture, so that accurate capture and real-time feedback can be realized, and a motion learning effect is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of the metaverse, and particularly to a method and system for motion capture and feedback adjustment for intelligent motion in the metaverse. Background Art

[0002] The intelligent motion in the metaverse combines motion capture technology, virtual avatar construction, and a motion feedback system, enabling users to have an immersive motion experience in a virtual environment. Motion capture technology is the key to realizing this motion mode, and usually devices such as inertial measurement units, optical capture systems, or depth cameras are used to collect the motion data of users. In the virtual environment of the metaverse, users project the motion in the real world into the virtual world through motion capture technology to achieve human-computer interaction and immersive experience.

[0003] The existing technologies related to intelligent motion in the metaverse have the following defects and deficiencies: Firstly, the motion capture accuracy in the existing technologies is insufficient. Especially in fast-motion scenarios, bone point drift and position errors are likely to occur, resulting in the mismatch between the actions of the virtual avatar and the actual actions of the user, reducing the immersion and interaction experience of the user in the metaverse environment.

[0004] Secondly, the existing technologies lack an effective real-time feedback and correction mechanism. After the user completes an action in the virtual environment, it is often difficult to obtain targeted posture adjustment suggestions in a timely manner, making it difficult to achieve self-correction and continuous improvement of actions, and affecting the learning effect and training efficiency of intelligent motion in the metaverse. Summary of the Invention

[0005] Embodiments of the present invention provide a method and system for motion capture and feedback adjustment for intelligent motion in the metaverse, which can solve the problems in the existing technologies.

[0006] In the first aspect of the embodiments of the present invention, a method for motion capture and feedback adjustment for intelligent motion in the metaverse is provided, including: Collecting the actual motion data of the user during the motion through multiple inertial measurement units; establishing a mapping relationship between the actual motion data and preset human bone nodes, and calculating the three-dimensional spatial position coordinates of each bone node according to the mapping relationship to obtain bone point position data; Constructing a multi-dimensional scoring matrix based on the motion coherence index and the action standardness index, performing real-time evaluation and scoring on each bone node in the bone point position data, and calculating the action scores of each bone node; optimizing the positions of the bone nodes in the bone point position data with scores lower than the preset score threshold according to the action scores to generate optimized bone point position data; Construct a user virtual avatar using the optimized bone point position data, calculate the virtual action data of the user virtual avatar, and display the user virtual avatar in the metaverse virtual scene; Perform a temporal analysis on the virtual action data, calculate the continuity features between adjacent action frames using a short-term analysis method to obtain local action features, and extract the complete action sequence using a long-term analysis method to obtain global action features; Compare the local action features and global action features with the standard action features, generate specific angle adjustment amounts, speed adjustment amounts, and displacement adjustment amounts, convert the adjustment amounts into motion correction instructions, and guide the user to adjust the motion posture in real time.

[0007] Establish a mapping relationship between the actual action data and the preset human bone nodes, calculate the three-dimensional spatial position coordinates of each bone node according to the mapping relationship, and obtain the bone point position data including: The actual action data includes acceleration data and angular velocity data obtained through multiple inertial measurement units; perform coordinate system alignment on the actual action data, and uniformly convert the local coordinate systems of the multiple inertial measurement units to the global coordinate system; Construct a human bone chain structure, which includes a torso node, limb nodes, and joint nodes, and preset length constraints and angle constraints between each node; Calculate the initial positions and initial postures of the multiple inertial measurement units in the global coordinate system, and bind the initial positions and initial postures to the corresponding nodes in the human bone chain structure; calculate the displacement change amounts of each inertial measurement unit based on the acceleration data, and calculate the attitude change amounts of each inertial measurement unit based on the angular velocity data; According to the displacement change amounts and attitude change amounts, combined with the length constraints and angle constraints, calculate the three-dimensional spatial position coordinates of each node in the human bone chain structure.

[0008] Construct a multi-dimensional scoring matrix based on the motion coherence index and the action standardness index, perform real-time evaluation and scoring on each bone node in the bone point position data, and calculate the action scores of each bone node including: Sample the bone point position data at a fixed time interval to generate a continuous action data sequence, and the action data sequence records the motion trajectories of each bone node changing with time; perform segmentation processing on the action data sequence, and divide the action data sequence into multiple overlapping action segments with a preset time window length, and the overlap rate of adjacent action segments is 50%; Calculate the motion feature parameters of each bone node in each action segment. The motion feature parameters include the position change rate, the speed change rate, and the acceleration change rate. Take the weighted sum result of the position change rate, the speed change rate, and the acceleration change rate as the coherence score of the action segment; Retrieve a standard action sequence that matches the current action segment from a pre-annotated standard action library. The standard action sequence contains the standard motion trajectories of each bone node in an ideal state; Compare the actual motion trajectories of each bone node in the current action segment with the standard motion trajectories, calculate the trajectory deviation value, and calculate the standardness score of the action segment based on the trajectory deviation value; Map the coherence score and the standardness score to different dimensions of a multi-dimensional scoring matrix respectively, set dynamic weight coefficients for different dimensions, and the weight coefficients are adaptively adjusted according to the action type; Multiply the scores of each dimension in the multi-dimensional scoring matrix by the corresponding weight coefficients and sum them to obtain the action score of each bone node in the current action segment.

[0009] Perform temporal analysis on the virtual action data, and use a short-time analysis method to calculate the continuity features between adjacent action frames to obtain local action features, including: Sample the virtual action data at a fixed frame rate, form an action analysis window with adjacent N action frames, and the number of overlapping frames between adjacent action analysis windows is M; calculate the position change amount of bone nodes between adjacent action frames within each action analysis window, including the spatial displacement vector and the angle change vector; Construct a temporal correlation matrix between action frames based on the position change amount. The matrix elements of the temporal correlation matrix represent the motion correlation of corresponding bone nodes between adjacent frames; perform time-domain decomposition on the temporal correlation matrix to extract the speed feature, the acceleration feature, and the angular velocity feature between adjacent action frames; Calculate the motion consistency coefficient of each bone node within the action analysis window. The motion consistency coefficient is used to characterize the coherence degree of the local action segment; identify abnormal motion frames within the action analysis window according to the motion consistency coefficient, and mark the action frames with continuity feature values lower than the preset continuity threshold as abnormal frames; Based on the speed feature, the acceleration feature, and the angular velocity feature, combined with the motion consistency coefficient, generate a feature vector representing the local action feature; associate the feature vector of the local action feature with the feature vector of the adjacent action analysis window to construct a complete local action feature sequence.

[0010] Use a long-time analysis method to extract the complete action sequence to obtain global action features, including: Performing temporal analysis on the virtual action data to obtain action sequence data, where the action sequence data includes three-dimensional coordinate information, velocity information, and acceleration information of each bone node from the start to the end of the action; dividing the action sequence data in the time dimension into a preparation stage, an acceleration stage, a stable stage, a deceleration stage, and an end stage according to the change trend of the bone node motion state; Extracting features from the action data of each stage, calculating the average velocity, maximum acceleration, displacement amplitude, and angle change range of the bone nodes in each stage to generate stage feature vectors; analyzing the change in motion state between adjacent stages, calculating the velocity continuity and position smoothness of the bone nodes at the stage transition moment to generate transition feature vectors; Concatenating the stage feature vectors and the transition feature vectors into global action features.

[0011] The method further includes: Constructing a two-dimensional temporal feature map based on the stage feature vectors and the transition feature vectors; Performing Fourier transform on the two-dimensional temporal feature map to obtain the frequency spectrum distribution, calculating the dynamic motion energy of each bone node, calculating the instantaneous motion energy of each frame according to the velocity vector and acceleration vector, and generating an energy time series curve; Aligning the energy time series curve and the two-dimensional temporal feature map in the time dimension, calculating the temporal correlation between the feature intensity and the energy peak, and establishing a force-feature mapping relationship; Analyzing the coordination degree between the force change and the motion feature conversion in the action sequence based on the force-feature mapping relationship to generate a coordination evaluation index.

[0012] Comparing the local action features and global action features with the standard action features to generate specific angle adjustment amounts, velocity adjustment amounts, and displacement adjustment amounts, and converting the adjustment amounts into motion correction instructions includes: Retrieving the standard action features matching the current action type from the action standard library, comparing the local action features with the standard action features at the frame level, calculating the pose deviation between adjacent frames of each bone node to obtain real-time correction parameters; comparing the global action features with the standard action features at the sequence level, analyzing the overall deviation trend during the action execution process to obtain trend correction parameters; Calculating the angle change amount that each bone node needs to adjust based on the real-time correction parameters, where the angle change amount is used to correct the joint bending degree and rotation direction; calculating the velocity adjustment amount of the bone nodes based on the trend correction parameters, where the velocity adjustment amount is used to optimize the action rhythm and motion smoothness; Calculate the displacement adjustment amount of the bone nodes in the three-dimensional space by combining the angle change amount and the speed adjustment amount to ensure that the adjusted motion trajectory meets the standard requirements; Perform a priority ranking on the angle change amount, the speed adjustment amount, and the displacement adjustment amount, establish a multi-level correction strategy to avoid conflicts between multiple adjustment instructions; convert the sorted adjustment amounts into specific motion correction instructions, where the motion correction instructions include an adjustment direction, an adjustment amplitude, and an execution timing.

[0013] In a second aspect of the embodiments of the present invention, there is provided an action capture and feedback adjustment system for metaverse intelligent motion, including: A first unit for collecting actual action data of a user during exercise through a plurality of inertial measurement units; establishing a mapping relationship between the actual action data and preset human bone nodes, and calculating the three-dimensional space position coordinates of each bone node according to the mapping relationship to obtain bone point position data; A second unit for constructing a multi-dimensional scoring matrix based on motion coherence indicators and action standardization indicators, performing real-time evaluation and scoring on each bone node in the bone point position data, and calculating the action scores of each bone node; optimizing the positions of the bone nodes in the bone point position data with scores lower than a preset score threshold according to the action scores to generate optimized bone point position data; A third unit for constructing a user virtual image by using the optimized bone point position data, calculating virtual action data of the user virtual image, and displaying the user virtual image in a metaverse virtual scene; A fourth unit for performing timing analysis on the virtual action data, calculating local action features of the continuity features between adjacent action frames by using a short-time analysis method, and extracting global action features of a complete action sequence by using a long-time analysis method; A fifth unit for comparing the local action features and the global action features with standard action features, generating specific angle adjustment amounts, speed adjustment amounts, and displacement adjustment amounts, converting the adjustment amounts into motion correction instructions, and guiding the user to adjust the motion posture in real time.

[0014] In a third aspect of the embodiments of the present invention, there is provided an electronic device, including: A processor; A memory for storing instructions executable by the processor; Wherein, the processor is configured to call the instructions stored in the memory to execute the method described above.

[0015] In a fourth aspect of the embodiments of the present invention, there is provided a computer-readable storage medium, on which computer program instructions are stored, and when the computer program instructions are executed by a processor, the method described above is implemented.

[0016] The beneficial effects of this application are as follows: By mapping actual motion data to human skeleton nodes and calculating spatial position coordinates, the user's motion posture can be accurately captured, and real-time evaluation and optimization can be carried out based on a multi-dimensional scoring matrix, improving the accuracy and stability of motion capture and solving the problem that traditional motion capture systems are vulnerable to interference in complex environments.

[0017] Using the optimized bone point position data to construct a user virtual image and display it in the metaverse virtual scene realizes the seamless connection between the physical world and the virtual world, enhances the user's immersion and interaction experience, and makes the motion performance of the virtual image more smooth and natural, conforming to the laws of human motion.

[0018] By performing time series analysis on the virtual motion data and comparing it with the standard motion characteristics, generating specific adjustment amounts and converting them into motion correction instructions, it can guide the user to adjust the motion posture in real time, improve the motion effect, effectively solve the problem of difficult to accurately guide the user's motion details in traditional motion teaching, and provide personalized motion guidance and feedback for the user. Brief Description of the Drawings

[0019] Figure 1 It is a schematic flow chart of the method for motion capture and feedback adjustment for metaverse intelligent motion in an embodiment of the present invention; Figure 2 It is a schematic flow chart for generating motion correction instructions. Detailed Embodiments

[0020] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.

[0021] The technical solutions of the present invention will be described in detail below with specific embodiments. These specific embodiments can be combined with each other, and the same or similar concepts or processes may not be repeated in some embodiments.

[0022] Refer to Figure 1 and Figure 2 The method for motion capture and feedback adjustment for metaverse intelligent motion in an embodiment of the present invention includes: Collect the actual motion data of the user during exercise through multiple inertial measurement units; establish a mapping relationship between the actual motion data and the preset human bone nodes, calculate the three-dimensional spatial position coordinates of each bone node according to the mapping relationship, and obtain the bone point position data; Construct a multi-dimensional scoring matrix based on the motion coherence index and the action standard index, perform real-time evaluation and scoring on each bone node in the bone point position data, and calculate the action scores of each bone node; optimize the positions of the bone nodes in the bone point position data with scores lower than the preset score threshold according to the action scores, and generate the optimized bone point position data; Construct a user virtual image using the optimized bone point position data, calculate the virtual motion data of the user virtual image, and display the user virtual image in the metaverse virtual scene; Perform temporal analysis on the virtual motion data, calculate the continuity features between adjacent action frames using the short-time analysis method to obtain local action features, and extract the complete action sequence using the long-time analysis method to obtain global action features; Compare the local action features and global action features with the standard action features, generate specific angle adjustment amounts, speed adjustment amounts, and displacement adjustment amounts, convert the adjustment amounts into motion correction instructions, and guide the user to adjust the motion posture in real time.

[0023] In an alternative embodiment, establishing a mapping relationship between the actual motion data and the preset human bone nodes, and calculating the three-dimensional spatial position coordinates of each bone node according to the mapping relationship to obtain the bone point position data includes: The actual motion data includes acceleration data and angular velocity data obtained through multiple inertial measurement units; perform coordinate system alignment on the actual motion data, and uniformly convert the local coordinate systems of the multiple inertial measurement units to the global coordinate system; Construct a human bone chain structure, which includes a torso node, limb nodes, and joint nodes, and preset length constraints and angle constraints between each node; Calculate the initial positions and initial postures of the multiple inertial measurement units in the global coordinate system, and bind the initial positions and initial postures to the corresponding nodes in the human bone chain structure; calculate the displacement change amounts of each inertial measurement unit based on the acceleration data, and calculate the attitude change amounts of each inertial measurement unit based on the angular velocity data; According to the displacement change amounts and attitude change amounts, combined with the length constraints and angle constraints, calculate the three-dimensional spatial position coordinates of each node in the human bone chain structure.

[0024] Exemplarily, the collected raw data consists of two parts: acceleration data, denoted as ax, ay, az, with the unit of m / s²; angular velocity data, denoted as wx, wy, wz, with the unit of rad / s. Since each IMU has its own local coordinate system, coordinate system alignment processing needs to be performed first. The specific implementation process is as follows: Select the Earth coordinate system as the global reference coordinate system, where the direction of gravity is the negative direction of the Z-axis, the front of the human body is the positive direction of the Y-axis, and the right side is the positive direction of the X-axis. For each inertial measurement unit IMU, initial attitude calibration is pre-performed, and the initial quaternion q_initial is recorded. This quaternion describes the rotation relationship of the IMU local coordinate system relative to the global coordinate system. All subsequent measurement data are coordinate-transformed through this quaternion to convert the acceleration and angular velocity in the local coordinate system to the global coordinate system. For example, an IMU measures an acceleration of (2.1, -0.5, 9.8) m / s² in the local coordinate system, and the data in the global coordinate system obtained through quaternion rotation transformation may be (0.2, 1.8, 9.6) m / s².

[0025] Next, a human skeletal chain structure model is constructed. This model consists of multiple nodes and bone segments connecting the nodes. According to standard human anatomical characteristics, the following node types are defined: 1. Trunk nodes: including the center of the head, neck, center of the chest, waist, and center of the pelvis; 2. Limb nodes: including the left / right shoulder, left / right elbow, left / right wrist, left / right hip joint, left / right knee, and left / right ankle; 3. Joint nodes: key points connecting the above nodes, such as the shoulder joint and elbow joint.

[0026] At the same time, to make the reconstructed human skeleton conform to physiological characteristics, two types of constraint conditions are preset: - Length constraint: Based on the body data of the subject, the length of each bone segment is determined. For example, for an adult male with a height of 175 cm, the upper arm length is about 28 cm, and the forearm length is about 26 cm; - Angle constraint: Angle limits are set according to the range of human joint activities. For example, the flexion and extension angle range of the elbow joint is 0° - 145°, and the abduction angle range of the shoulder joint is 0° - 180°.

[0027] Initial pose determination is a crucial part of the capture process. The subject is required to stand with arms hanging naturally and face forward. At this time, static state data is collected for about 3 seconds to determine the initial positions and poses of each IMU in the global coordinate system. The vertical direction is determined by the gravity alignment method, and the forward direction is determined by a predefined forward pose. For example, the initial position of the chest IMU may be (0, 0, 0.7) m, indicating a height of 0.7 meters relative to the origin of the ground coordinates.

[0028] When binding the IMU to the bone nodes, the corresponding relationships are established according to the actual placement positions of the IMUs on the human body. For example, the right upper arm IMU is associated with the right shoulder and right elbow nodes, and the waist IMU is associated with the lumbar vertebra node. A mapping table is established to record the relative position relationships between each IMU and the corresponding bone nodes.

[0029] The core of motion capture is to update the bone node positions in real time. For each frame of data (assuming a time interval of dt, with a typical value of 16.7 ms), the following processing is performed: First, the acceleration data is processed to remove the influence of gravity and obtain the linear acceleration generated by human motion. The quadratic integration method is used to calculate the displacement. The specific steps are as follows: According to the acceleration a of the current frame and the velocity v of the previous frame, calculate the velocity v' of the current frame as v' = v + a × dt; then calculate the displacement change ds = (v + v') / 2 × dt according to the average velocity. For example, if an IMU measures a horizontal acceleration of 2 m / s² and the time interval is 16.7 ms, the single-frame velocity increment is approximately 0.033 m / s, and the corresponding displacement change is approximately 0.28 mm.

[0030] When processing the angular velocity data, the attitude quaternion is updated by integrating the angular velocity. For example, if the IMU measures an angular velocity of (0.1, 0.2, -0.3) rad / s and the time interval is 16.7 ms, the rotation increment can be calculated and applied to the current attitude quaternion to obtain the updated attitude.

[0031] Considering the problem of integral error accumulation, this method uses a sliding window filter and zero velocity update technology for correction. When it is detected that the IMU is in an approximately stationary state (the changes in acceleration and angular velocity are less than the preset thresholds, such as 0.1 m / s² and 0.05 rad / s), it is considered that the velocity of the IMU is zero, and the integral calculation is reset.

[0032] Finally, based on the updated results of the position and orientation, combined with the bone constraint conditions, the inverse kinematics algorithm is used to calculate the three-dimensional spatial positions of each node in the bone chain. When multiple IMUs affect the same bone node, a confidence-based weighted average method is adopted to determine the final position. For example, the elbow joint position is affected by both the upper arm IMU and the forearm IMU, and weights can be assigned according to the measurement confidence of the two. Assuming the weights are 0.6 and 0.4 respectively, the final position is the weighted average of the calculation results of the two.

[0033] Through the above technical implementation, this method can achieve accurate human motion capture, and the accuracy of bone point positions can reach the centimeter level, which is applicable to various application scenarios such as virtual reality, motion analysis, and rehabilitation training.

[0034] In an alternative embodiment, a multi-dimensional scoring matrix is constructed based on the motion coherence index and the action standardization index to perform real-time evaluation and scoring on each bone node in the bone point position data. Calculating the action scores of each bone node includes: Sampling the bone point position data at fixed time intervals to generate a continuous action data sequence, and the action data sequence records the motion trajectories of each bone node over time; segmenting the action data sequence, and dividing the action data sequence into multiple overlapping action segments with a preset time window length, and the overlap rate of adjacent action segments is 50%; Calculating the motion feature parameters of each bone node in each action segment, where the motion feature parameters include the position change rate, the speed change rate, and the acceleration change rate, and taking the weighted sum result of the position change rate, the speed change rate, and the acceleration change rate as the coherence score of the action segment; Retrieving a standard action sequence matching the current action segment from a pre-annotated standard action library, and the standard action sequence contains the standard motion trajectories of each bone node in an ideal state; Comparing the actual motion trajectories of each bone node in the current action segment with the standard motion trajectories, calculating the trajectory deviation value, and calculating the standardization score of the action segment based on the trajectory deviation value; Mapping the coherence score and the standardization score to different dimensions of the multi-dimensional scoring matrix respectively, setting dynamic weight coefficients for different dimensions, and the weight coefficients are adaptively adjusted according to the action type; Multiplying the scores of each dimension in the multi-dimensional scoring matrix by the corresponding weight coefficients and summing them to obtain the action score of each bone node in the current action segment.

[0035] This embodiment provides a method for constructing a multi-dimensional scoring matrix based on the motion coherence index and the action standardization index to perform real-time evaluation and scoring on each bone node in the bone point position data.

[0036] First, the system obtains the skeletal point position data when the user performs an action through sensors or motion capture devices, including at least 15 key skeletal nodes, such as the head, shoulders, elbows, wrists, hips, knees, and ankles. The position of each skeletal node is represented by three-dimensional coordinates (x, y, z).

[0037] The system samples at fixed time intervals. For example, it collects skeletal point position data every 20 milliseconds, generating a continuous sequence of action data. Suppose the user performs a squatting action that lasts for 5 seconds. The system will collect approximately 250 frames of skeletal point position data, forming a complete action sequence.

[0038] Next, the system segments the action data sequence. Using a preset time window length, such as 2 seconds (i.e., 100 frames) as a complete action segment, and setting the overlap rate of adjacent action segments to 50%. Therefore, the first action segment contains the data from frames 1 - 100, the second action segment contains the data from frames 51 - 150, and so on. Through this overlapping segmentation method, the system can capture the continuous changes in the action and ensure that key action frames are not missed.

[0039] For each action segment, the system calculates the motion feature parameters of each skeletal node. Taking the right knee joint as an example, first, the position change rate is calculated: the system extracts the position change amount between every two adjacent frames of the right knee joint in this action segment, accumulates the position change amounts of all frames, and then divides by the total number of frames to obtain the average position change rate. For example, the average position change rate of the right knee joint in a certain action segment is 0.05 m / frame.

[0040] The calculation method of the speed change rate is: first, calculate the instantaneous speed between every two adjacent frames, then calculate the difference between adjacent instantaneous speeds (speed change amount), accumulate all speed change amounts and divide by the total number of frames minus 2 to obtain the average speed change rate. For example, the average speed change rate of the right knee joint in the same action segment is 0.003 m / frame².

[0041] Similarly, the acceleration change rate is obtained by calculating the difference between adjacent acceleration values and then averaging. For example, the average acceleration change rate of the right knee joint in this action segment is 0.0005 m / frame³.

[0042] The system sets different weights to calculate the coherence score. Suppose the weight of the position change rate is 0.3, the weight of the speed change rate is 0.4, and the weight of the acceleration change rate is 0.3. For the above example of the right knee joint, the coherence score is: 0.05×0.3 + 0.003×0.4 + 0.0005×0.3 = 0.01665. The score range is mapped to the 0 - 100 interval through normalization processing.

[0043] Meanwhile, the system retrieves a matching standard action sequence from a pre-annotated standard action library. This standard action library contains various common action types, such as squats, lunges, planks, etc. Each action is demonstrated by professional athletes and manually annotated. The system uses the dynamic time warping algorithm to compare the similarity between the current action segment and the standard action sequence, and selects the standard action with the highest similarity as a reference.

[0044] The system compares the actual movement trajectories of each bone node in the current action segment with the standard movement trajectories. Taking the right knee joint as an example, the Euclidean distance between its actual position and the standard position in each frame is calculated, and the distances of all frames are accumulated and divided by the total number of frames to obtain the average trajectory deviation value. For example, the average trajectory deviation value of the right knee joint in a certain action segment is 0.08 meters.

[0045] The calculation of the standardness score adopts an inverse proportion relationship. The smaller the deviation value, the higher the score. The specific calculation method is: Standardness score = 100 - Deviation value × Coefficient, where the coefficient is adjusted according to the importance of different bone nodes. For example, the importance coefficient of the knee joint in the squat action is 1.5, then the standardness score of the right knee joint is 100 - 0.08×1.5×100 = 88 points.

[0046] The system constructs a multi-dimensional scoring matrix and maps the coherence score and the standardness score to different dimensions respectively. For the squat action, the weight of the standardness dimension is usually higher. For example, the weight of the coherence dimension is 0.4, and the weight of the standardness dimension is 0.6. The system will adaptively adjust these weights according to different action types. For example, for dance-like actions, the weight of the coherence dimension may be increased to 0.6.

[0047] Finally, the system multiplies the scores of each dimension in the multi-dimensional scoring matrix by the corresponding weights and sums them up to obtain the comprehensive score of each bone node in the current action segment. Taking the above-mentioned right knee joint as an example, if its coherence score (after normalization) is 85 points and the standardness score is 88 points, then the comprehensive score is 85×0.4 + 88×0.6 = 86.8 points.

[0048] The system also sets specific scoring focuses for different types of actions. For example, in the squat action, the system assigns higher scoring weights to the knee joint and the hip joint, which may account for 60% of the total score; while in the plank action, the system pays more attention to the stability of the spine and shoulders, and the scoring weights of the bone nodes in these parts can reach 65%.

[0049] Through this multi-dimensional and adaptive weight scoring mechanism, the system can comprehensively and objectively evaluate the quality of the user's actions and provide targeted improvement suggestions.

[0050] In an alternative embodiment, temporal analysis is performed on the virtual action data, and local action features are obtained by calculating the continuity features between adjacent action frames using a short-time analysis method, including: Sample the virtual action data at a fixed frame rate, form an action analysis window with N adjacent action frames, and the number of overlapping frames between adjacent action analysis windows is M; calculate the position change amount of the skeletal nodes between adjacent action frames within each action analysis window, including the spatial displacement vector and the angular change vector; Based on the position change amount, construct a temporal correlation matrix between action frames, and the matrix elements of the temporal correlation matrix represent the motion correlation of the corresponding skeletal nodes between adjacent frames; perform time-domain decomposition on the temporal correlation matrix to extract the speed feature, acceleration feature, and angular velocity feature between adjacent action frames; Calculate the motion consistency coefficient of each skeletal node within the action analysis window, and the motion consistency coefficient is used to characterize the coherence degree of the local action segment; identify abnormal motion frames within the action analysis window according to the motion consistency coefficient, and mark the action frames with continuity feature values lower than the preset continuity threshold as abnormal frames; Based on the speed feature, acceleration feature, and angular velocity feature, combined with the motion consistency coefficient, generate a feature vector representing the local action feature; associate the feature vector of the local action feature with the feature vector of the adjacent action analysis window to construct a complete local action feature sequence.

[0051] Sample the virtual action data at a fixed frame rate. In this embodiment, the action data of the virtual character is sampled at a fixed frame rate of 30 frames per second. Form an action analysis window with N adjacent action frames, and the number of overlapping frames between adjacent action analysis windows is M. Specifically, N can be set to 60 frames, that is, each action analysis window contains 2 seconds of action data; M can be set to 15 frames, that is, the adjacent windows overlap 0.5 seconds of action data. This design of overlapping windows can ensure the continuity of action analysis and avoid analysis discontinuities at the window boundaries.

[0052] Assume that the skeletal model of the virtual character contains K nodes (for example, K = 23, including key nodes such as the head, neck, shoulders, elbows, wrists, chest, waist, hips, knees, and ankles). For the i-th frame and the (i + 1)-th frame, calculate the spatial displacement vector ΔP(i, j) and the angular change vector ΔA(i, j) of each skeletal node j. The spatial displacement vector represents the position change of the skeletal node in three-dimensional space and contains three components of x, y, and z; the angular change vector represents the Euler angle change of the skeletal node relative to its parent node and contains the rotation angle changes around the x-axis, y-axis, and z-axis.

[0053] For example, for the change of the right elbow node between the 10th frame and the 11th frame, the possible spatial displacement vector is (0.05 meters, 0.02 meters, -0.03 meters), and the angle change vector is (2 degrees, -1 degrees, 0.5 degrees).

[0054] The temporal correlation matrix C between action frames is constructed based on the position change. For a window consisting of N action frames, the size of the temporal correlation matrix C is (N-1)×K×K. The matrix elements C(i,j, l ) indicates that between the i-th frame and the i+1-th frame, the bone node j and the bone node l The correlation is calculated by calculating the motion correlation between node j and node l The similarity between the displacement vector and the angle change vector is used to determine the motion trend of the two nodes. A high similarity indicates that the motion trends of the two nodes are consistent, while a low similarity indicates that the motion trends are different.

[0055] The temporal correlation matrix is ​​decomposed in the time domain to extract the velocity features, acceleration features, and angular velocity features between adjacent action frames. The velocity feature V(i,j) is obtained by dividing the displacement vector ΔP(i,j) by the inter-frame time interval (e.g. 1 / 30 second). The acceleration feature A(i,j) is calculated by dividing the velocity difference between two adjacent frames by the inter-frame time interval. The angular velocity feature ω(i,j) is calculated by dividing the angle change vector ΔA(i,j) by the inter-frame time interval.

[0056] In actual cases, the velocity characteristics of the right elbow node at a certain moment may be (1.5 m / s, 0.6 m / s, -0.9 m / s), the acceleration characteristics may be (0.3 m / s², -0.1 m / s², 0.2 m / s²), and the angular velocity characteristics may be (60 degrees / s, -30 degrees / s, 15 degrees / s).

[0057] Next, calculate the motion consistency coefficient R of each skeletal node in the motion analysis window. The motion consistency coefficient is used to characterize the coherence of the local motion segment, and it is calculated by analyzing the stability of the changes in the velocity direction and acceleration direction in the window. Specifically, for skeletal node j, its motion consistency coefficient R(j) can be represented by the average cosine value of the angle between the velocity vectors of adjacent frames in the window. The closer the R(j) value is to 1, the more consistent the motion direction of the node in the window and the more coherent the motion; the closer the R(j) value is to 0 or negative, the more drastic the change in motion direction and the possible incoherent motion.

[0058] Identify abnormal motion frames within the action analysis window based on the motion consistency coefficient. Set a preset continuous threshold T = 0.65. When the motion consistency coefficient R(j) of any key bone node j in a certain frame i is lower than T, mark this frame as an abnormal frame. Abnormal frames may represent mutations, jitters, or unnatural transitions in the action and require special attention or correction in subsequent processing.

[0059] In a specific example, when analyzing the window data of a running action, it is found that the motion consistency coefficient of the left knee joint in the 35th frame is only 0.42, which is lower than the threshold of 0.65. The system marks it as an abnormal frame. Further inspection reveals that the angle of the left knee joint in this frame has suddenly changed abnormally, which does not conform to the normal running pattern.

[0060] Based on the extracted velocity features, acceleration features, and angular velocity features, combined with the motion consistency coefficient, generate a feature vector F that characterizes the local action features. For each action analysis window, the feature vector F includes the average velocity, maximum velocity, average acceleration, maximum acceleration, average angular velocity, maximum angular velocity, and the motion consistency coefficients of all bone nodes within the window. For example, for a character with 23 bone nodes, the dimension of the feature vector F can reach hundreds of dimensions, comprehensively describing the action features within the window.

[0061] Associate the feature vectors of the local action features with the feature vectors of adjacent action analysis windows to construct a complete local action feature sequence S. Since there is an overlapping area between adjacent windows, it is necessary to smoothly fuse the features in the overlapping area, and a weighted average method can be used, where the weights are related to the positions of the samples in the overlapping area. In this way, the finally obtained feature sequence S can coherently represent the local dynamic characteristics of the entire virtual action.

[0062] In an alternative implementation, the global action features obtained by using a long-term analysis method to extract the complete action sequence include: Perform temporal analysis on the virtual action data to obtain action sequence data, which contains the three-dimensional coordinate information, velocity information, and acceleration information of each bone node from the start to the end of the action; according to the change trend of the motion state of the bone nodes, divide the action sequence data into a preparation stage, an acceleration stage, a stable stage, a deceleration stage, and an end stage in the time dimension; Extract features from the action data of each stage, calculate the average velocity, maximum acceleration, displacement amplitude, and angle change range of the bone nodes in each stage to generate stage feature vectors; analyze the change in the motion state between adjacent stages, calculate the velocity continuity and position smoothness of the bone nodes at the stage transition moment to generate transition feature vectors; Concatenate the stage feature vectors and the transition feature vectors into global action features.

[0063] In this embodiment, global features of a complete action sequence are extracted through long-time sequence analysis, mainly including steps such as action sequence data processing, action stage division, feature extraction, and global feature construction.

[0064] Temporal analysis is performed on virtual action data to obtain action sequence data. In this embodiment, a depth camera is used to collect human action information, and 20 key bone nodes are extracted through a human bone recognition algorithm, including positions such as the head, neck, shoulders, elbows, wrists, spine, hips, knees, and ankles. The system records the three-dimensional coordinates (x, y, z) of each bone node at a sampling rate of 30 frames per second to form an original coordinate sequence. For a specific bone node, such as the right wrist, its coordinate at time t is denoted as (x_t, y_t, z_t).

[0065] To calculate the speed information, the system performs a difference operation on the coordinate changes between two adjacent frames. For example, the speed v_t of the right wrist at time t is calculated as the difference between the coordinates of the current frame and the previous frame divided by the sampling time interval. In a specific example, if the coordinate of the right wrist at the 10th frame is (0.45m, 1.2m, 0.3m), the coordinate at the 11th frame is (0.48m, 1.25m, 0.32m), and the sampling interval is 0.033 seconds, then the speed is calculated as ((0.48 - 0.45) / 0.033, (1.25 - 1.2) / 0.033, (0.32 - 0.3) / 0.033) = (0.91m / s, 1.52m / s, 0.61m / s).

[0066] Acceleration information is obtained by performing a second-order difference on the speed sequence. For example, the acceleration a_t of the right wrist at time t is the difference between the current speed and the previous speed divided by the sampling time interval. Based on the above case, if the speed calculated from the 9th frame to the 10th frame is (0.76m / s, 1.21m / s, 0.45m / s), then the acceleration at the 10th frame is ((0.91 - 0.76) / 0.033, (1.52 - 1.21) / 0.033, (0.61 - 0.45) / 0.033) = (4.55m / s², 9.39m / s², 4.85m / s²).

[0067] According to the change trend of the motion state of the bone nodes, the action sequence data is divided into five stages in the time dimension. First, the speed curve of each node is calculated, and then the stage division is performed by setting speed thresholds. In this embodiment, for the throwing action, the speed thresholds are set at four points: 15%, 40%, 60%, and 30% of the maximum speed, which are used as the demarcation points for each stage respectively.

[0068] The preparation phase is defined as the time period from the start of the movement to when the speed first exceeds 15% of the maximum speed. For example, if the maximum speed of the right wrist in a certain movement sequence is 3.6 m / s, then when the speed first exceeds 0.54 m / s, it is marked as the end of the preparation phase. In a specific case, the preparation phase lasts from the 1st frame to the 18th frame, and the right wrist mainly completes the transition from rest to initial movement.

[0069] The acceleration phase is the time period from the end of the preparation phase to when the speed reaches 40% of the maximum value. Taking the above example, the end point of the acceleration phase is when the speed reaches 1.44 m / s, corresponding to the 32nd frame in the sequence. In this phase, the speed of the right wrist rises from 0.54 m / s to 1.44 m / s, and the displacement is about 12 cm.

[0070] The stable phase is the time period from the end of the acceleration phase to when the speed reaches 60% of the maximum value, corresponding to the speed rising from 1.44 m / s to 2.16 m / s, which occurs between the 32nd frame and the 45th frame. In this phase, the displacement of the right wrist is about 25 cm, and the angle changes by about 15 degrees.

[0071] The deceleration phase is the time period from the end of the stable phase to when the speed drops to 30% of the maximum speed, that is, from the 45th frame to the 68th frame, and the speed drops from 2.16 m / s to 1.08 m / s. In this phase, the displacement of the right wrist is about 30 cm, and the acceleration changes significantly.

[0072] The ending phase is from the end of the deceleration phase to when the movement completely stops, that is, from the 68th frame to the 85th frame, and the speed drops from 1.08 m / s to nearly 0 m / s. In this phase, the displacement is about 8 cm, and the movement gradually stops and stabilizes.

[0073] Feature extraction is performed on the movement data of each phase. The average speed of the skeletal nodes in each phase is calculated. For example, the average speed of the right wrist in the preparation phase is 0.32 m / s, 0.92 m / s in the acceleration phase, 1.75 m / s in the stable phase, 1.58 m / s in the deceleration phase, and 0.42 m / s in the ending phase.

[0074] The maximum acceleration in each phase is calculated. The maximum acceleration of the right wrist in the preparation phase is 2.7 m / s², 8.5 m / s² in the acceleration phase, 6.2 m / s² in the stable phase, -9.1 m / s² in the deceleration phase, and -4.3 m / s² in the ending phase.

[0075] The displacement amplitude is calculated. Taking the right wrist as an example, the displacement amplitude in the preparation phase is 5 cm, 12 cm in the acceleration phase, 25 cm in the stable phase, 30 cm in the deceleration phase, and 8 cm in the ending phase.

[0076] Calculate the range of angular change. In the preparation stage, the range of elbow joint angle change is from 5 degrees to 25 degrees, from 25 degrees to 65 degrees in the acceleration stage, from 65 degrees to 80 degrees in the stable stage, from 80 degrees to 45 degrees in the deceleration stage, and from 45 degrees to 30 degrees in the ending stage.

[0077] The above calculation results form the feature vectors of each stage. For example, the feature vector of the preparation stage is [0.32 m / s, 2.7 m / s², 5 cm, 20 degrees], and the others are constructed similarly.

[0078] Analyze the change of motion state between adjacent stages, calculate the velocity continuity and position smoothness of the bone nodes at the moment of stage transition, and generate the transition feature vector. The velocity continuity is obtained by calculating the average value of the velocity differences of the three frames before and after the junction of adjacent stages. For example, at the transition from the preparation stage to the acceleration stage, the velocity continuity value is 0.08 m / s, indicating a smooth transition.

[0079] The position smoothness is obtained by calculating the jitter degree of the position coordinates of the three frames before and after the junction. For example, the position smoothness value at the junction from the preparation stage to the acceleration stage is 0.5 cm, indicating a stable position change. Example of the feature vector of each conversion point: preparation-acceleration conversion point [0.08 m / s, 0.5 cm], acceleration-stable conversion point [0.12 m / s, 0.8 cm], stable-deceleration conversion point [0.15 m / s, 1.2 cm], deceleration-ending conversion point [0.07 m / s, 0.6 cm].

[0080] Finally, concatenate the stage feature vector and the transition feature vector to form the global action feature. For the above throwing action example, the global action feature is [0.32 m / s, 2.7 m / s², 5 cm, 20 degrees, 0.92 m / s, 8.5 m / s², 12 cm, 40 degrees, 1.75 m / s, 6.2 m / s², 25 cm, 15 degrees, 1.58 m / s, -9.1 m / s², 30 cm, 35 degrees, 0.42 m / s, -4.3 m / s², 8 cm, 15 degrees, 0.08 m / s, 0.5 cm, 0.12 m / s, 0.8 cm, 0.15 m / s, 1.2 cm, 0.07 m / s, 0.6 cm], a 28-dimensional feature vector, which comprehensively describes the timing characteristics and stage transition characteristics of the entire action.

[0081] This global action feature can be used in application scenarios such as action recognition, action quality assessment, and action similarity comparison, providing a more comprehensive action timing representation than single-frame features.

[0082] In an alternative embodiment, the method further includes: Construct a two-dimensional time-series feature map based on the phase feature vector and the conversion feature vector; Perform a Fourier transform on the two-dimensional time-series feature map to obtain a spectral distribution, calculate the dynamic motion energy of each skeletal node, calculate the instantaneous motion energy of each frame according to the velocity vector and the acceleration vector, and generate an energy time-series curve; Align the energy time-series curve with the two-dimensional time-series feature map in the time dimension, calculate the temporal correlation between the feature intensity and the energy peak, and establish a force-feature mapping relationship; Based on the force-feature mapping relationship, analyze the coordination degree between the force change and the motion feature conversion in the action sequence, and generate a coordination evaluation index.

[0083] This embodiment provides a method for evaluating action coordination based on time-series features and dynamic energy analysis. This method realizes the objective evaluation of the coordination of human action sequences by constructing a two-dimensional time-series feature map, performing spectral analysis, establishing a force-feature mapping relationship, and generating a coordination evaluation index. The following are the detailed technical implementation steps: Construct a two-dimensional time-series feature map based on the obtained phase feature vector and conversion feature vector. The horizontal axis of this feature map represents the time dimension, specifically corresponding to the frame index or timestamp of the action sequence; the vertical axis represents the feature dimension, corresponding to different feature components; the pixel value in the figure represents the feature intensity. In specific implementation, arrange the feature vectors at each time point in rows to form a matrix structure. For example, for an action sequence with 120 frames and a feature dimension of 64, the size of the constructed two-dimensional time-series feature map is 64×120. To enhance the visualization effect, it can be presented in the form of a heat map, where areas with high feature intensity are displayed in warm colors (such as red), and low areas are displayed in cold colors (such as blue).

[0084] Perform a Fourier transform on the constructed two-dimensional time-series feature map to obtain a spectral distribution. The specific steps are as follows: perform a fast Fourier transform on the time series of each feature dimension separately to obtain the corresponding frequency components and amplitudes; summarize the spectral results of all feature dimensions to form a spectral distribution matrix. For example, for 64-dimensional features, 64 spectral distribution curves can be obtained. By analyzing the spectral distribution, the main periodic patterns in the action can be identified. In practice, taking a dance movement as an example, if there are obvious peaks in the spectrum in the range of 2 - 4 Hz, it means that the action repeats at a frequency of 2 - 4 times per second.

[0085] Extract the three-dimensional coordinate sequences of 25 key skeletal nodes from the motion capture data; calculate the position differences between adjacent frames to obtain velocity vectors; further calculate the differences of the velocity vectors to obtain acceleration vectors. Based on the velocity vectors and acceleration vectors, calculate the instantaneous motion energy for each frame. The specific calculation method is as follows: for the t-th frame, first calculate the weighted sum of the squared velocities and squared accelerations of all skeletal nodes as the instantaneous energy value E(t) of this frame. The weights can be set according to the importance of the skeletal nodes. Usually, the weights of the trunk part are higher (such as 0.4), and the weights of the limb ends are lower (such as 0.2). Connect the instantaneous energy values of all frames to form an energy time series curve.

[0086] Taking a modern dance performance of a professional dancer as an example, the energy curve of its 120-frame motion sequence shows obvious fluctuation characteristics, and the energy peaks appear at the action transition points and the moments with strong explosive power. The specific data shows that the energy value of the static posture is about 5 - 10 units, the medium-intensity actions are about 30 - 50 units, and the energy peaks of high-intensity actions such as jumps can reach 80 - 100 units.

[0087] Align the energy time series curve with the two-dimensional time series feature map in the time dimension, calculate the time series correlation between the feature intensity and the energy peak, and establish a force - feature mapping relationship. The specific implementation is as follows: normalize the energy curve so that its value range is between 0 and 1; similarly, normalize the time series data of each feature dimension in the feature map; calculate the cross-correlation function between the energy curve and the time series data of each feature dimension to determine the time delay relationship and the correlation coefficient; select the feature dimensions with correlation coefficients exceeding the threshold (such as 0.7), and these dimensions are considered to be closely related to the change of action force; pair these high-correlation feature dimensions with the corresponding energy change patterns to form a force - feature mapping table.

[0088] Taking the above modern dance performance as an example, the analysis shows that the correlation coefficients between the 3rd, 7th, 15th, and 22nd dimensional features and the energy curve are 0.81, 0.75, 0.89, and 0.72 respectively. These features mainly correspond to the feature expressions of trunk torsion, arm extension, leg flexion and extension, and jumping actions. Through the force - feature mapping table, when the energy value is in the range of 75 - 85, the 15th dimensional feature value should be in the range of 0.8 - 0.9, indicating the standard feature response of high-intensity leg actions.

[0089] Based on the force - feature mapping relationship, analyze the coordination degree between the force change and the motion feature conversion in the action sequence, and generate a coordination evaluation index. The specific method is as follows: for each frame, according to the current energy value, query the ideal feature response value from the force - feature mapping table; calculate the Euclidean distance between the actual feature value and the ideal value as the coordination deviation of a single frame; average the coordination deviations of the entire sequence to obtain the average coordination deviation index; convert the average deviation value into a coordination score from 0 to 100, and the smaller the deviation, the higher the score.

[0090] In practical applications, when comparing and analyzing the same dance moves of professional dancers and beginners, the average coordination deviation of professional dancers is 0.12, which is converted into a coordination score of 92; while the average deviation of beginners is 0.31, and the score is 73. This indicates that this method can effectively distinguish the movement coordination of performers at different levels and provides an objective quantitative basis for the assessment of movement quality.

[0091] Through the above method, not only can the overall coordination be evaluated, but also the specific uncoordinated time periods and corresponding movement characteristics can be located, providing precise feedback for movement training.

[0092] In an alternative embodiment, comparing the local movement characteristics and global movement characteristics with standard movement characteristics to generate specific angle adjustment amounts, speed adjustment amounts, and displacement adjustment amounts, and converting the adjustment amounts into movement correction instructions includes: Retrieving standard movement characteristics matching the current movement type from the movement standard library, comparing the local movement characteristics with the standard movement characteristics at the frame level, calculating the pose deviation of each bone node between adjacent frames to obtain real-time correction parameters; comparing the global movement characteristics with the standard movement characteristics at the sequence level, analyzing the overall deviation trend during the execution of the movement to obtain trend correction parameters; Calculating the angle change amount that each bone node needs to adjust based on the real-time correction parameters, where the angle change amount is used to correct the joint bending degree and rotation direction; calculating the speed adjustment amount of the bone node based on the trend correction parameters, where the speed adjustment amount is used to optimize the movement rhythm and movement smoothness; Combining the angle change amount and the speed adjustment amount, calculating the displacement adjustment amount of the bone node in three-dimensional space to ensure that the adjusted movement trajectory meets the standard requirements; Prioritizing the angle change amount, speed adjustment amount, and displacement adjustment amount, establishing a multi-level correction strategy to avoid conflicts between multiple adjustment instructions; converting the sorted adjustment amounts into specific movement correction instructions, where the movement correction instructions include adjustment direction, adjustment amplitude, and execution timing.

[0093] The system retrieves standard movement characteristics matching the current movement type from the movement standard library. The system maintains a movement standard library, which stores the characteristic data of various standard movements. For example, for the yoga tree pose, the standard library stores the key bone node position data of this pose, including that the hip joint angle should be 90° ± 5° when the left leg is lifted, the knee straightness of the supporting leg should be greater than 175°, and the angle between the two arms when lifted should be 180° ± 10°, etc. When it is detected that the user is performing the tree pose, the system matches the corresponding standard movement characteristics through the movement recognition algorithm.

[0094] The system makes frame-level comparisons between local action features and standard action features. The system collects data for each frame of the user's current action and analyzes the postures of each bone node between adjacent frames. Specifically, taking the hip joint as an example, if the standard tree pose requires a hip joint angle of 90°, and the user's actual angle is 75°, the system records the joint angle deviation as 15°. The system performs similar deviation calculations for 20 major joint points throughout the body and saves these deviation data as real-time correction parameters. This parameter set includes data such as the angle differences, position offsets, and rotation errors of each joint.

[0095] The system makes sequence-level comparisons between global action features and standard action features. The system collects the user's action data for 30 consecutive frames (about 1 second) and analyzes its overall movement trend. For example, when analyzing the user's squatting action, the system detects that the user's squatting speed gradually slows down from 0.3 m / s at the beginning to 0.1 m / s, while the standard squat requires a uniform squatting speed of 0.25 m / s. The system records this speed change trend as a trend correction parameter for subsequent optimization of the action rhythm.

[0096] Based on the real-time correction parameters, the system calculates the angle change amounts that each bone node needs to adjust. Taking the action of raising the arm as an example, if the user's elbow joint flexion is 150°, while the standard action requires 170°, the system generates an angle adjustment amount of increasing by 20°. For joints with multi-dimensional rotation, such as the shoulder joint, the system calculates the angle adjustment amounts for the three degrees of freedom respectively. For example, the forward lift direction needs to increase by 15°, the abduction direction needs to decrease by 8°, and the internal rotation needs to decrease by 10°. These angle change amounts are directly used to correct the joint flexion degree and rotation direction.

[0097] Based on the trend correction parameters, the system calculates the speed adjustment amounts of the bone nodes. Taking the running action as an example, if the standard action requires an arm swing frequency of 1.2 times per second, and the user's actual frequency is 0.8 times per second, the system generates a speed adjustment instruction to increase the arm swing speed by 50%. The system also analyzes the smoothness of the connection between actions and generates corresponding speed adjustment amounts for the parts with pauses to ensure the coherence of the actions.

[0098] The system combines the angle change amounts and speed adjustment amounts to calculate the displacement adjustment amounts of the bone nodes in three-dimensional space. Taking the squat start action as an example, if the user's knee joint angle needs to be adjusted forward by 15° and the speed needs to be increased by 20% at the same time, the comprehensive calculation shows that the displacement adjustment amount of the center point of the knee joint in the forward direction is 7 cm, and the displacement adjustment amount in the upward direction is 5 cm. The system performs similar displacement calculations for all key nodes in the body to ensure that the adjusted movement trajectory meets the standard requirements.

[0099] The system prioritizes the angular change amount, speed adjustment amount, and displacement adjustment amount to establish a multi-level correction strategy. The priority sorting adopts the following rules: First, correct large deviations that may cause safety risks, such as hyperextension of the knee joint; second, correct key deviations that affect the action effect, such as insufficient hip joint angle; finally, correct aesthetic deviations, such as uneven arm arcs. For example, in the squatting motion, if the system detects that the knee joint angle is too small (65°) with a safety risk and the center of gravity leans forward too much at the same time, it will first issue a knee joint angle adjustment command, and then a center of gravity position adjustment command.

[0100] The system converts the sorted adjustment amounts into specific motion correction commands. Taking the yoga warrior pose as an example, if it is detected that the external rotation of the right leg is insufficient and the forward inclination speed is too fast, the correction commands generated by the system are: "Rotate the right leg outward by 15 degrees (adjustment direction), with a moderate amplitude (adjustment amplitude), and complete it in the next breathing cycle (execution timing)" and "Slow down the forward inclination speed of the torso by 40% (adjustment direction and amplitude), and execute immediately (execution timing)". These commands can be transmitted to the user through voice prompts, visual images, or tactile feedback, etc.

[0101] Through the above detailed steps, the system can accurately compare the user's actions with the standard actions, generate targeted adjustment commands, help the user effectively correct action errors, and improve the exercise effect and safety. This method is especially applicable to scenarios that require precise action guidance, such as fitness training, rehabilitation therapy, and dance learning.

[0102] The action capture and feedback adjustment system for metaverse intelligent movement in the embodiment of the present invention includes: The first unit is used to collect the actual action data of the user during the movement through multiple inertial measurement units; establish a mapping relationship between the actual action data and the preset human body bone nodes, and calculate the three-dimensional spatial position coordinates of each bone node according to the mapping relationship to obtain the bone point position data; The second unit is used to construct a multi-dimensional scoring matrix based on the motion coherence index and the action standardness index, evaluate and score each bone node in the bone point position data in real time, and calculate the action scores of each bone node; optimize the positions of the bone nodes in the bone point position data with scores lower than the preset score threshold according to the action scores to generate optimized bone point position data; The third unit is used to construct a user virtual image by using the optimized bone point position data, calculate the virtual action data of the user virtual image, and display the user virtual image in the metaverse virtual scene; A fourth unit, configured to perform a timing analysis on the virtual action data, calculate the continuity features between adjacent action frames by using a short-time analysis method to obtain local action features, and extract a complete action sequence by using a long-time analysis method to obtain global action features; A fifth unit, configured to compare the local action features and the global action features with standard action features, generate specific angular adjustment amounts, speed adjustment amounts, and displacement adjustment amounts, convert the adjustment amounts into motion correction instructions, and guide a user to adjust a motion posture in real time.

[0103] In a third aspect of the embodiments of the present invention, there is provided an electronic device, including: A processor; A memory for storing instructions executable by the processor; Wherein, the processor is configured to call the instructions stored in the memory to execute the method described above.

[0104] In a fourth aspect of the embodiments of the present invention, there is provided a computer-readable storage medium, on which computer program instructions are stored, and when the computer program instructions are executed by a processor, the method described above is implemented.

[0105] The present invention may be a method, an apparatus, a system, and / or a computer program product. The computer program product may include a computer-readable storage medium, on which computer-readable program instructions for executing various aspects of the present invention are uploaded.

[0106] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements on some or all of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention.

Claims

1. A motion capture and feedback adjustment method for metaverse intelligent motion, characterized in that: include: The actual motion data of the user during exercise is collected through multiple inertial measurement units; Establishing a mapping relationship between the actual motion data and preset human skeleton nodes, calculating the three-dimensional spatial position coordinates of each skeleton node according to the mapping relationship, and obtaining skeleton point position data; A multi-dimensional scoring matrix is ​​constructed based on the motion coherence index and the action standardization index, each bone node in the bone point position data is evaluated and scored in real time, and the action score of each bone node is calculated; According to the action score, positions of the skeleton nodes in the skeleton point position data whose scores are lower than a preset score threshold are optimized to generate optimized skeleton point position data; Using the optimized skeletal point position data to construct a user virtual image, calculating virtual action data of the user virtual image, and displaying the user virtual image in a metaverse virtual scene; Performing time series analysis on the virtual motion data, calculating the continuity features between adjacent motion frames using a short-time analysis method to obtain local motion features, and extracting the complete motion sequence using a long-time analysis method to obtain global motion features; The local motion features and the global motion features are compared with the standard motion features to generate specific angle adjustment amounts, speed adjustment amounts and displacement adjustment amounts, and the adjustment amounts are converted into motion correction instructions to guide the user to adjust the motion posture in real time.

2. The method according to claim 1, characterized in that Establishing a mapping relationship between the actual motion data and the preset human skeleton nodes, calculating the three-dimensional spatial position coordinates of each skeleton node according to the mapping relationship, and obtaining the skeleton point position data includes: The actual motion data includes acceleration data and angular velocity data obtained by a plurality of inertial measurement units; coordinate system alignment is performed on the actual motion data, and the local coordinate systems of the plurality of inertial measurement units are converted into a global coordinate system; Constructing a human skeleton chain structure, wherein the human skeleton chain structure includes a trunk node, limb nodes and joint nodes, and presetting length constraints and angle constraints between each node; Calculating the initial positions and initial postures of the multiple inertial measurement units in the global coordinate system, and binding the initial positions and initial postures to corresponding nodes in the human skeleton chain structure; calculating the displacement change of each inertial measurement unit based on the acceleration data, and calculating the posture change of each inertial measurement unit based on the angular velocity data; According to the displacement change and the posture change, combined with the length constraint and the angle constraint, the three-dimensional spatial position coordinates of each node in the human skeleton chain structure are calculated.

3. The method according to claim 1, characterized in that A multi-dimensional scoring matrix is ​​constructed based on the motion coherence index and the action standardization index, and each bone node in the bone point position data is evaluated and scored in real time. The action score of each bone node is calculated, including: The skeletal point position data are sampled at fixed time intervals to generate a continuous action data sequence, wherein the action data sequence records the motion trajectory of each skeletal node over time; the action data sequence is segmented to divide the action data sequence into a plurality of overlapping action segments according to a preset time window length, and the overlapping rate of adjacent action segments is 50%; Calculating motion characteristic parameters of each skeletal node in each action segment, the motion characteristic parameters including position change rate, velocity change rate and acceleration change rate, and taking a weighted sum of the position change rate, velocity change rate and acceleration change rate as the coherence score of the action segment; Retrieving a standard action sequence matching the current action clip from a pre-annotated standard action library, wherein the standard action sequence includes a standard motion trajectory of each skeletal node under an ideal state; Compare the actual motion trajectory of each skeletal node in the current action segment with the standard motion trajectory, calculate the trajectory deviation value, and calculate the standard score of the action segment based on the trajectory deviation value; Mapping the coherence score and the standardization score to different dimensions of a multi-dimensional scoring matrix respectively, setting dynamic weight coefficients for different dimensions, and the weight coefficients are adaptively adjusted according to the action type; The scores of each dimension in the multi-dimensional scoring matrix are multiplied by the corresponding weight coefficients and the sum is calculated to obtain the action score of each skeletal node in the current action segment.

4. The method according to claim 1, characterized in that: The virtual action data is subjected to time series analysis, and the continuity features between adjacent action frames are calculated using a short-time analysis method to obtain local action features, including: The virtual motion data is sampled at a fixed frame rate, and N adjacent motion frames are grouped into a motion analysis window, where the number of overlapping frames between adjacent motion analysis windows is M; and the position change of the bone nodes between adjacent motion frames in each motion analysis window is calculated, including a spatial displacement vector and an angle change vector. Based on the position change, a temporal correlation matrix between action frames is constructed, wherein the matrix elements of the temporal correlation matrix represent the motion correlation of corresponding bone nodes between adjacent frames; the temporal correlation matrix is ​​decomposed in the time domain to extract velocity features, acceleration features and angular velocity features between adjacent action frames; Calculate the motion consistency coefficient of each skeletal node in the motion analysis window, the motion consistency coefficient is used to characterize the coherence of the local motion segment; identify abnormal motion frames in the motion analysis window according to the motion consistency coefficient, and mark the motion frames whose continuity feature values ​​are lower than a preset continuity threshold as abnormal frames; Based on the velocity features, acceleration features and angular velocity features, combined with the motion consistency coefficient, a feature vector representing the local motion features is generated; the feature vector of the local motion features is associated with the feature vector of the adjacent motion analysis window to construct a complete local motion feature sequence.

5. The method according to claim 1, characterized in that The global action features obtained by extracting the complete action sequence using the long-term analysis method include: Performing time series analysis on the virtual action data to obtain action sequence data, the action sequence data including three-dimensional coordinate information, speed information and acceleration information of each skeletal node from the start to the end of the action; dividing the action sequence data into a preparation stage, an acceleration stage, a stabilization stage, a deceleration stage and an end stage in the time dimension according to a change trend of the motion state of the skeletal node; Extract features from the motion data of each stage, calculate the average speed, maximum acceleration, displacement amplitude and angle change range of the bone nodes in each stage, and generate the stage feature vector; analyze the changes in motion state between adjacent stages, calculate the speed continuity and position smoothness of the bone nodes at the stage transition moment, and generate the conversion feature vector; The stage feature vector and the conversion feature vector are concatenated into a global action feature.

6. The method according to claim 5, characterized in that The method further comprises: Based on the stage feature vector and the conversion feature vector, construct a two-dimensional time series feature graph; Performing Fourier transformation on the two-dimensional time series feature graph to obtain spectrum distribution, calculating the dynamic motion energy of each bone node, calculating the instantaneous motion energy of each frame according to the velocity vector and the acceleration vector, and generating an energy time series curve; Aligning the energy time series curve with the two-dimensional time series feature graph in the time dimension, calculating the time series correlation between the feature intensity and the energy peak, and establishing a strength-feature mapping relationship; Based on the force-feature mapping relationship, the coordination degree between the force change and the motion feature conversion in the action sequence is analyzed to generate a coordination evaluation index.

7. The method according to claim 1, characterized in that Comparing the local motion features and the global motion features with the standard motion features to generate specific angle adjustment amounts, speed adjustment amounts, and displacement adjustment amounts, and converting the adjustment amounts into motion correction instructions includes: Retrieve standard action features that match the current action type from the action standard library, perform frame-level comparison between the local action features and the standard action features, calculate the posture deviation of each skeletal node between adjacent frames, and obtain real-time correction parameters; perform sequence-level comparison between the global action features and the standard action features, analyze the overall deviation trend during the action execution process, and obtain trend correction parameters; Based on the real-time correction parameters, the angle change amount that needs to be adjusted for each bone node is calculated, and the angle change amount is used to correct the degree of joint bending and the direction of rotation; based on the trend correction parameters, the speed adjustment amount of the bone node is calculated, and the speed adjustment amount is used to optimize the action rhythm and movement smoothness; Combining the angle change and speed adjustment, the displacement adjustment of the skeletal node in the three-dimensional space is calculated to ensure that the adjusted motion trajectory meets the standard requirements; Prioritize the angle change, speed adjustment and displacement adjustment, establish a multi-level correction strategy, and avoid conflicts between multiple adjustment instructions; convert the sorted adjustment amounts into specific motion correction instructions, which include adjustment direction, adjustment amplitude and execution timing.

8. A motion capture and feedback adjustment system for Metaverse Intelligent Movement, used to implement the method as described in any one of claims 1 to 7, characterized in that: include: The first unit is used to collect actual motion data of the user during the movement through multiple inertial measurement units; Establishing a mapping relationship between the actual motion data and preset human skeleton nodes, calculating the three-dimensional spatial position coordinates of each skeleton node according to the mapping relationship, and obtaining skeleton point position data; The second unit is used to construct a multi-dimensional scoring matrix based on the motion coherence index and the action standardization index, to evaluate and score each bone node in the bone point position data in real time, and to calculate the action score of each bone node; According to the action score, positions of the skeleton nodes in the skeleton point position data whose scores are lower than a preset score threshold are optimized to generate optimized skeleton point position data; A third unit is used to construct a user virtual image using the optimized skeleton point position data, calculate virtual action data of the user virtual image, and display the user virtual image in a metaverse virtual scene; The fourth unit is used to perform time series analysis on the virtual action data, calculate the continuity features between adjacent action frames using a short-time analysis method to obtain local action features, and extract the complete action sequence using a long-time analysis method to obtain global action features; The fifth unit is used to compare the local motion features and the global motion features with the standard motion features, generate specific angle adjustment amounts, speed adjustment amounts and displacement adjustment amounts, convert the adjustment amounts into motion correction instructions, and guide the user to adjust the motion posture in real time.

9. An electronic device, characterized in that: include: processor; a memory for storing processor-executable instructions; The processor is configured to call the instructions stored in the memory to execute the method described in any one of claims 1 to 7.

10. A computer-readable storage medium having computer program instructions stored thereon, characterized in that: When the computer program instructions are executed by a processor, the method according to any one of claims 1 to 7 is implemented.

Citation Information

Patent Citations

  • Device and method for improving motion guiding efficiency

    CN104834384A

  • VR-based sports action scoring method and system, electronic equipment and storage medium

    CN116259111A

  • Human body posture recognition system based on artificial intelligence

    CN117671738A

  • Method and system for managing hazardous chemical substances in laboratory

    CN119476961A

  • Real-time rendering optimization method and system in meta universe scene building engine

    CN119941956A

Cited By

  • Personalized training plan generation method and system in element universe smart sports

    CN120126680A

  • Method and System for Generating Personalized Training Plans in the Metaverse Smart Sports

    CN120126680B

  • Eight-section brocade action detail capture and comprehensive evaluation system based on artificial intelligence

    CN121459426A

  • Baibuejinzhenqiangxianxiedaocaijihezonghepingjiaxitong

    CN121459426B

  • Virtual reality feedback processing method, device and system for cancer postoperative rehabilitation training

    CN122415815A