Motion analysis system based on artificial intelligence

By designing a sports analysis system based on artificial intelligence, it solves the problem that traditional methods are difficult to accurately capture athletes' key details and cannot provide personalized training guidance, and high-precision collection and intelligent analysis of sports data are achieved, and personalized sports optimization suggestions are provided to help athletes break through the bottleneck of performance.

CN120180321AActive Publication Date: 2025-06-20HUNAN EDUCATION AUDIO-VISUAL ELECTRONIC PUBLISHING HOUSE CO LTD

Patent Information

Application Number
CN202510239303.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-03
Publication Date
2025-06-20
Estimated Expiration
2045-03-03

AI Technical Summary

Technical Problem

Traditional sports analysis methods are difficult to accurately capture the key details of athletes during exercise, and cannot provide personalized training guidance, and cannot meet the athletes' unique physical conditions and exercise habits.

Method used

A motion analysis system based on artificial intelligence is designed, and through multi-source data acquisition module, spatiotemporal preprocessing module, feature fusion module, motion pattern recognition module, abnormal detection module and feedback generation module, all-round, high-precision acquisition and intelligent analysis of motion data are realized.

Benefits of technology

The system can accurately record and analyze every step of the athlete's movement, provide personalized sports optimization suggestions and abnormal handling strategies, helping athletes break through the bottleneck of performance and meet personalized improvement needs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120180321A_ABST
    Figure CN120180321A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of artificial intelligence, and discloses an artificial intelligence-based motion analysis system which comprises a multi-source data acquisition module, a space-time preprocessing module, a feature fusion module, a motion mode recognition module, an anomaly detection module and a feedback generation module. The multi-source data acquisition module synchronously acquires data at a multi-band sampling rate by means of an inertial measurement unit, an optical sensor and a depth camera array, generates an original motion data set and stores the original motion data set to the distributed edge nodes; the space-time preprocessing module uses a multi-scale dynamic time warping algorithm and adaptive Kalman filtering to calibrate noise reduction; the feature fusion module extracts and fuses key spatio-temporal features; the motion mode recognition module and the anomaly detection module accurately judge a motion mode according to the fused feature tensor and position anomaly; the feedback generation module generates a personalized executable instruction set according to the results, and provides powerful support for the athletes to dig potential and break through bottlenecks.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of artificial intelligence, and specifically to a motion analysis system based on artificial intelligence. Background Art

[0002] In the field of sports competitions, the competition in events has become increasingly fierce, and athletes' requirements for the refinement of their own training have reached an unprecedented level. Taking the track and field sprint event as an example, factors such as the force of each step of the athlete during the 100-meter dash, the amplitude and frequency of the arm swing, and the angle of body lean all subtly and precisely affect the final competition result. The traditional method of relying on coaches' on-site visual observation and rough video analysis after the event is difficult to accurately capture these fleeting key motion details. On the one hand, it is difficult for coaches to simultaneously pay attention to the motion performances of multiple key parts of the athlete's body in an instant; on the other hand, simple video playback lacks the ability of in-depth data mining and cannot quantitatively analyze the advantages and disadvantages of each motion parameter. Moreover, different athletes have unique physical conditions and motion habits, and the general training guidance mode can no longer meet the needs of personalized improvement. There is an urgent need for a system that can collect athletes' motion data comprehensively and with high precision, intelligently analyze the motion patterns, and then provide exclusive optimization suggestions to help athletes tap their potential and break through the performance bottleneck. Summary of the Invention

[0003] The purpose of the present invention is to provide a motion analysis system based on artificial intelligence to solve the problems raised in the above background art.

[0004] To achieve the above purpose, the present invention provides the following technical solution: A motion analysis system based on artificial intelligence, the system includes:

[0005] Including a multi-source data acquisition module, a spatio-temporal preprocessing module, a feature fusion module, a motion pattern recognition module, an anomaly detection module, and a feedback generation module;

[0006] The multi-source data acquisition module is used to obtain the original multi-modal data required for motion analysis. Specifically, it synchronously collects data by deploying inertial measurement units, optical sensors, and depth camera arrays, generates an original motion data set, and transmits the original motion data set to the spatio-temporal preprocessing module;

[0007] The synchronous data acquisition specifically refers to collecting motion sequence data at different sampling rates, including multi-band data of 30 frames per second, 60 frames per second, and 120 frames per second, and generating spatio-temporal synchronous data through dynamic timestamp alignment, which is stored in a distributed edge node;

[0008] The spatio-temporal preprocessing module is used to perform spatio-temporal calibration and noise reduction on the original data. Specifically, it aligns heterogeneous sensor data through a multi-scale dynamic time warping algorithm and uses an adaptive Kalman filter to eliminate motion noise, generating optimized motion data and transmitting it to the feature fusion module;

[0009] The feature fusion module is used to extract and fuse spatio-temporal features. Specifically, through a three-dimensional residual network optimized by an attention mechanism based on spatio-temporal convolution, it extracts skeletal joint trajectory features, motion speed distribution features, and pose angle change features respectively, generating a fused feature tensor and inputting it into the motion pattern recognition module and the anomaly detection module;

[0010] The motion pattern recognition module is used to classify and recognize motion patterns. Specifically, based on the fused feature tensor, it adopts a multi-task graph convolutional network combined with an adversarial generation training model to output the probability distribution of motion categories and key action labels;

[0011] The anomaly detection module is used to detect abnormal behaviors during the motion process. Specifically, based on a hybrid model of a variational autoencoder and an isolation forest algorithm, it performs reconstruction error analysis and distribution shift detection on the fused feature tensor, generating an anomaly score and a localization result;

[0012] The feedback generation module is used to generate motion optimization suggestions and anomaly handling strategies. Specifically, according to the motion pattern recognition results and anomaly detection data, it generates a personalized feedback plan through a decision tree model driven by reinforcement learning and encodes it into an executable instruction set.

[0013] Preferably, the original multi-modal data includes skeletal joint coordinate data, accelerometer data, gyroscope angular velocity data, depth image sequences, and RGB video stream data; the spatio-temporal synchronization data further includes sensor calibration parameters and environmental light compensation information.

[0014] Preferably, the multi-scale dynamic time warping algorithm specifically includes a time series slicing operator, a local similarity metric matrix construction operator, and a global path optimization operator;

[0015] The time series slicing operator is used to divide heterogeneous sensor data into segments by windows and calculate the similarity between segments through dynamic programming;

[0016] The local similarity metric matrix construction operator specifically uses the dynamic time warping distance combined with the cosine similarity to generate a multi-modal alignment cost matrix;

[0017] The global path optimization operator specifically solves the minimum cumulative cost path based on the branch and bound algorithm to achieve frame-level alignment of cross-sensor data.

[0018] Preferably, the three-dimensional residual network includes parallel spatio-temporal convolutional branches, where the first branch uses dilated convolution to extract long-term temporal dependence features, the second branch optimizes the weights of local motion features through a channel attention mechanism, and the third branch uses atrous spatial pyramid pooling to capture multi-scale spatial context information.

[0019] Preferably, the multi-task graph convolutional network jointly trains an adversarial generation model, which specifically includes a skeletal topology graph construction module, a dynamic adjacency matrix learning module, an adversarial discriminant sub-network, and a multi-head classifier;

[0020] The skeletal topology graph construction module is used to construct an initial joint connection graph based on human anatomy priors and dynamically adjust the edge weights through learnable parameters;

[0021] The dynamic adjacency matrix learning module specifically uses a gated graph attention mechanism to iteratively update the relationships between nodes and generate a spatio-temporal enhanced graph representation;

[0022] The adversarial discriminant sub-network is used to constrain the manifold consistency between the feature distribution and the real motion data;

[0023] The multi-head classifier specifically outputs the motion category probability, action phase label, and energy consumption estimate in parallel.

[0024] Preferably, the working process of the anomaly detection module includes:

[0025] Model the feature distribution of normal motion patterns through a variational autoencoder and calculate the reconstruction error of the input data;

[0026] Use the isolation forest algorithm to recursively partition the high-dimensional feature space and detect abnormal subspaces that deviate from the mainstream distribution;

[0027] Fuse the reconstruction error probability and subspace anomaly degree to generate a comprehensive anomaly score and associated joint indices.

[0028] Preferably, the reinforcement learning-driven decision tree model specifically includes a state encoder, a policy search tree, and a reward function optimizer;

[0029] The state encoder is used to map the motion pattern recognition result and anomaly detection data into a low-dimensional state vector;

[0030] The policy search tree specifically explores the action space through Monte Carlo tree search and generates candidate feedback paths;

[0031] The reward function optimizer dynamically adjusts the policy weights according to the motion efficiency improvement degree, anomaly correction speed, and physiological load constraint.

[0032] Preferably, the personalized feedback solution includes motion posture correction parameters, training intensity adjustment suggestions, abnormal movement avoidance strategies, and real-time biomechanical load warnings.

[0033] Preferably, the executable instruction set is encoded by the edge computing node into at least one of the following forms: tactile feedback pulse sequence, augmented reality visualization annotation, robotic exoskeleton drive signal, and voice prompt instruction stream.

[0034] Compared with the prior art, the beneficial effects of the present invention are as follows:

[0035] The multi-source data acquisition module uses inertial measurement units, optical sensors, and depth camera arrays to synchronously collect data, which can comprehensively capture various information of athletes during the movement process, and is no longer limited to the one-sidedness of the traditional coach's visual observation. Whether it is the force of each step of a 100-meter sprinter when pushing off the ground, the details of the arm swing, or the changes in body posture, they can be accurately recorded in the form of an original motion data set, without missing any key moments, providing a solid foundation for subsequent accurate analysis.

[0036] Collect motion sequence data at different sampling rates (multi-band data of 30 frames / second, 60 frames / second, and 120 frames / second), and generate spatio-temporal synchronous data through dynamic timestamp alignment, which is stored in the distributed edge node to ensure the integrity and accuracy of the data, meet the needs of fine analysis of high-speed sports events, and effectively solve the problem that it is difficult to capture fleeting key motion details in the traditional way.

[0037] The spatio-temporal preprocessing module uses the multi-scale dynamic time warping algorithm to align heterogeneous sensor data, enabling information from different types of sensors to be accurately matched in the spatio-temporal dimension, avoiding data chaos and deviation. At the same time, adaptive Kalman filtering is used to eliminate motion noise, making the optimized motion data cleaner, excluding interference for subsequent feature extraction and analysis, greatly improving the usability and reliability of the data, and overcoming the drawback of the simple video playback lacking the ability of in-depth data mining.

[0038] The feature fusion module uses a three-dimensional residual network optimized by an attention mechanism based on spatio-temporal convolution to extract bone joint point trajectory features, motion speed distribution features, and attitude angle change features respectively, and generates a fused feature tensor. This deep fusion method can organically integrate multiple key dimension information during the athlete's movement process, discover the deep-seated motion pattern associations hidden behind the data, so as to accurately grasp the athlete's motion state and provide rich and targeted basis for personalized training guidance.

[0039] Based on the fused feature tensor, the motion pattern recognition module adopts a multi-task graph convolutional network to jointly train an adversarial generation model, and outputs the probability distribution of motion categories and key action labels. This means that the system can not only accurately determine the motion pattern category of the athlete, but also precisely locate the key actions. Whether it is the starting, accelerating, sprinting stages in sprinting or the specific action links in other complex sports events, they can all be clearly identified, providing clear improvement directions for coaches and athletes, and changing the status quo that the general training guidance mode cannot meet the personalized improvement needs.

[0040] The anomaly detection module, based on a hybrid model of variational autoencoder and isolation forest algorithm, conducts reconstruction error analysis and distribution shift detection on the fused feature tensor, generating anomaly scores and location results. During the training or competition of athletes, it can promptly detect abnormal behaviors such as action deformation, abnormal force application, and rhythm disorder. Whether it is caused by fatigue, injury or technical action mistakes, the problems can be detected immediately, avoiding the continuous deterioration of abnormal situations and affecting the performance of athletes. The feedback generation module, according to the motion pattern recognition results and anomaly detection data, generates a personalized feedback plan through a decision tree model driven by reinforcement learning and encodes it into an executable instruction set. This enables each athlete to obtain exclusive training optimization suggestions based on their unique physical conditions and exercise habits. For example, for sprinters with problems such as insufficient arm swing amplitude and uneven pushing force, precise adjustment plans are given to help athletes tap their potential and break through the performance bottleneck. Brief Description of the Drawings

[0041] Figure 1 It is the working principle diagram of the motion analysis system described in the present invention;

[0042] Figure 2 It is the flowchart of multi-scale alignment of heterogeneous sensor data;

[0043] Figure 3 It is the flowchart of motion pattern recognition. Detailed Embodiments

[0044] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.

[0045] Please refer to Figures 1-3 , the present invention provides a technical solution: a motion analysis system based on artificial intelligence, and the system includes:

[0046] Multi-source Data Acquisition Module: To obtain the original multi-modal data required for motion analysis, the system deploys inertial measurement units, optical sensors, and depth camera arrays. These devices collect motion sequence data at different sampling rates, covering multi-band data of 30 frames per second, 60 frames per second, and 120 frames per second. During acquisition, through dynamic timestamp alignment technology, spatio-temporal synchronous data is generated and stored in distributed edge nodes. The original multi-modal data includes skeletal joint point coordinate data, accelerometer data, gyroscope angular velocity data, depth image sequences, and RGB video stream data. The spatio-temporal synchronous data further includes sensor calibration parameters and environmental light compensation information to ensure the accuracy and integrity of the data.

[0047] Spatio-temporal Preprocessing Module: This module is responsible for spatio-temporal calibration and noise reduction of the collected raw data. The multi-scale dynamic time warping algorithm is used to align heterogeneous sensor data. This algorithm includes a temporal slicing operator, a local similarity metric matrix construction operator, and a global path optimization operator. At the same time, the adaptive Kalman filter algorithm is used to eliminate motion noise, thereby generating optimized motion data and transmitting it to the feature fusion module.

[0048] Feature Fusion Module: This module uses a three-dimensional residual network optimized by an attention mechanism based on spatio-temporal convolution to extract skeletal joint point trajectory features, motion speed distribution features, and pose angle change features respectively, generating a fused feature tensor. The three-dimensional residual network includes parallel spatio-temporal convolution branches. The first branch uses dilated convolution to extract long temporal dependence features, the second branch optimizes the weights of local motion features through a channel attention mechanism, and the third branch uses atrous spatial pyramid pooling to capture multi-scale spatial context information. The fused feature tensor is then input into the motion pattern recognition module and the anomaly detection module.

[0049] Motion Pattern Recognition Module: Based on the fused feature tensor generated by the feature fusion module, this module uses a multi-task graph convolutional network combined with an adversarial generation training model to achieve the classification and recognition of motion patterns. This model includes a skeletal topology graph construction module, a dynamic adjacency matrix learning module, an adversarial discriminant sub-network, and a multi-head classifier. The final output is the probability distribution of motion categories and key action labels.

[0050] Anomaly Detection Module: Based on a hybrid model of variational autoencoder and isolation forest algorithm, this module conducts reconstruction error analysis and distribution shift detection on the fused feature tensor. The specific workflow is as follows: First, the variational autoencoder is used to model the feature distribution of normal motion patterns and calculate the reconstruction error of the input data; then the isolation forest algorithm is used to recursively partition the high-dimensional feature space to detect abnormal sub-spaces deviating from the mainstream distribution; finally, the reconstruction error probability and sub-space anomaly degree are fused to generate a comprehensive anomaly score and associated joint point indices.

[0051] Feedback Generation Module: Based on the motion pattern recognition results and anomaly detection data, this module generates personalized feedback solutions through a decision tree model driven by reinforcement learning. This model includes a state encoder, a policy search tree, and a reward function optimizer. The personalized feedback solutions include motion posture correction parameters, training intensity adjustment suggestions, abnormal movement avoidance strategies, and real-time biomechanical load warnings. The feedback generation module encodes the personalized feedback solutions into an executable instruction set, which can be encoded into at least one of a tactile feedback pulse sequence, an augmented reality visualization annotation, a robotic exoskeleton drive signal, and a voice prompt instruction stream by an edge computing node.

[0052] The present invention will be further described below in conjunction with Embodiments 1 to 5:

[0053] Embodiment 1:

[0054] This embodiment elaborates in detail the specific working process of the multi-source data acquisition module, as well as the characteristics of the original multi-modal data and spatio-temporal synchronization data, providing a basis for subsequent data analysis and processing to ensure that the collected data is accurate, comprehensive, and meets the system processing requirements.

[0055] In an actual application scenario, taking the training monitoring of athletes as an example, the inertial measurement units of the multi-source data acquisition module are installed at key parts of the athlete's body, such as the wrists, ankles, knees, and waist. These inertial measurement units are built-in with accelerometers and gyroscopes. The accelerometer is used to collect the acceleration data of each part of the athlete during movement. Its principle is based on Newton's second law F = ma, and the acceleration a is calculated by measuring the inertial force F, with the unit of m / s 2 . The gyroscope is used to measure the angular velocity data. Its measurement principle is based on the law of conservation of angular momentum, and the output angular velocity data has the unit of rad / s.

[0056] Optical sensors are arranged around the training venue to capture the bone joint point coordinate data of the athletes using the principle of optical imaging. The depth camera array collects depth image sequences and RGB video stream data from different angles. The depth image sequences can accurately obtain the distance information between each part of the athlete's body and the camera, and the RGB video stream data provides rich color and texture information for a more comprehensive analysis of the athlete's motion posture.

[0057] When collecting data, due to the differences in working frequencies and performance of different sensors, motion sequence data needs to be collected at different sampling rates. Taking the multi-band data collection at 30 frames per second, 60 frames per second, and 120 frames per second as an example, the low sampling rate (30 frames per second) data is suitable for capturing the general trend of motion, while the high sampling rate (120 frames per second) data can accurately record the details of fast action changes. Through the dynamic timestamp alignment technology, accurate time stamps are added to each data sample to achieve the spatio-temporal synchronization of data at different sampling rates. In this process, the synchronized data also records the sensor calibration parameters, which are used to calibrate the measurement errors of the sensors to ensure the accuracy of the data. At the same time, the environmental light compensation information is also recorded to eliminate the influence of environmental light changes on data collection. For example, when the light brightness changes in the training venue, the quality of the depth image and RGB video stream data is ensured to be stable.

[0058] Embodiment 2:

[0059] This embodiment focuses on the multi-scale dynamic time warping algorithm in the spatio-temporal preprocessing module, and details the working principles and operation processes of its various operators to achieve the effective alignment of heterogeneous sensor data, laying a foundation for subsequent accurate feature extraction and analysis.

[0060] The multi-scale dynamic time warping algorithm in the spatio-temporal preprocessing module is the key to achieving the alignment of heterogeneous sensor data. When processing heterogeneous data from inertial measurement units, optical sensors, and depth camera arrays, the time series slicing operator is first used. Assume that the heterogeneous sensor data sequences are X = {x1, x2, …, x m} and Y = {y1, y2, …, y n}. The time series slicing operator divides these data into segments according to a window. Let the window size be w, then the segment divided from the data sequence X is X i = {x i , x i+1 , …, x i+w-1} (i = 1, 2, …, m - w + 1), and similarly, the segment Y j is divided from Y. The similarity between segments is calculated through dynamic programming. The core idea of dynamic programming is to decompose a complex problem into multiple sub-problems and save the solutions of sub-problems to avoid repeated calculations. The formula for calculating the similarity d(X i , Y j ) between segments X i and Y j is:

[0061]

[0062] where, ‖·‖ represents a certain distance metric, such as the Euclidean distance.

[0063] Next, the local similarity metric matrix construction operator comes into play. This operator uses the dynamic time warping distance combined with cosine similarity to generate a multi-modal alignment cost matrix. The dynamic time warping distance (DTW distance) is used to measure the similarity of two time series in terms of time. Its calculation process is to find an optimal time alignment path such that the sum of the distances between the two time series along this path is minimized. Let two time series A = {a1, a2, …, a p} and B = {b1, b2, …, b q}, the calculation of the DTW distance DTW(A, B) is based on the recursive formula:

[0064]

[0065] where A p-1 = {a1, a2, …, a p-1}, B q-1 and so on. At the same time, combined with the cosine similarity cosine(A, B), its formula is:

[0066]

[0067] Combining the two generates a multi-modal alignment cost matrix C, where C ij represents the alignment cost between segment X i and Y j .

[0068] Finally, the global path optimization operator solves the minimum cumulative cost path based on the branch and bound algorithm. The branch and bound algorithm is an algorithm for searching for the optimal solution on the solution space tree. It improves the search efficiency by continuously dividing the solution space (branching) and using the bound function to prune the subtrees that cannot contain the optimal solution (bounding). In this algorithm, by solving the minimum cumulative cost path, frame-level alignment of cross-sensor data is achieved, enabling heterogeneous sensor data to be accurately matched in time and space, providing a reliable data basis for subsequent data analysis.

[0069] Example 3:

[0070] This example details the structure and working principle of the three-dimensional residual network in the feature fusion module, demonstrating how it effectively extracts skeletal joint point trajectory features, motion speed distribution features, and pose angle change features, providing high-quality fused feature tensors for motion pattern recognition and anomaly detection.

[0071] The core of the feature fusion module is a three-dimensional residual network optimized by an attention mechanism based on spatio-temporal convolution. This network contains parallel spatio-temporal convolution branches, and each branch undertakes different feature extraction tasks.

[0072] The first branch uses dilated convolutions to extract long-term temporal dependence features. Dilated convolutions add a dilation rate parameter to ordinary convolutions. Assuming the size of an ordinary convolution kernel is k×k, the convolution kernel of dilated convolutions is spatially expanded at a dilation rate of d. In the time dimension, for the input time series data T = {t1, t2, …, t n}, the dilated convolution operation can be expressed as:

[0073]

[0074] where w j is the convolution kernel weight and y i is the convolution output result. By setting an appropriate dilation rate, the first branch can capture motion dependence relationships within a relatively long time range. For example, when analyzing the motion patterns of long-distance runners, it can capture the characteristics of the pace rhythm changes within several seconds.

[0075] The second branch optimizes the weights of local motion features through a channel attention mechanism. The channel attention mechanism analyzes the channel dimension of the feature map and adaptively adjusts the weights of each channel. For the input feature map (C is the number of channels, H is the height, and W is the width), first, global average pooling and global max pooling are performed on the feature map in the spatial dimension to obtain two one-dimensional vectors z avg and z max , and their calculations are respectively:

[0076]

[0077] Then, these two vectors are respectively passed through a shared multi-layer perceptron (MLP), and the outputs are added to obtain the attention weight vector a. Then, a weighted operation is performed on the original feature map F in the channel dimension to optimize the weights of local motion features and highlight key motion features. For example, when analyzing the shooting action of a basketball player, more attention is paid to the motion features of the arm joints.

[0078] The third branch uses atrous spatial pyramid pooling to capture multi-scale spatial context information. Atrous spatial pyramid pooling samples spatial information at multiple scales through convolution operations with different dilation rates. Assuming the dilation rates are r1, r2, r3 respectively, for the input feature map F, the feature maps after convolutions with different dilation rates are F1, F2, F3 respectively. Then, these feature maps are concatenated and convolved to obtain a fused feature map, so as to capture multi-scale spatial context information from local to global. For example, when analyzing the movements of a dancer, it can not only pay attention to the subtle movement changes of the body's local parts but also grasp the overall dance posture and spatial position relationship. Finally, the features extracted by the three branches are fused to generate a fused feature tensor, providing rich feature information for subsequent motion pattern recognition and anomaly detection.

[0079] Example 4:

[0080] In this example, each component and its working principle of the multi-task graph convolutional network joint adversarial generation training model in the motion pattern recognition module are analyzed in detail, showing how it accurately classifies and recognizes motion patterns, and outputs the motion category probability distribution and key action labels.

[0081] The multi-task graph convolutional network joint adversarial generation training model of the motion pattern recognition module consists of a skeletal topology graph construction module, a dynamic adjacency matrix learning module, an adversarial discriminant sub-network, and a multi-head classifier.

[0082] The skeletal topology graph construction module constructs an initial joint connection graph based on prior knowledge of human anatomy. There are specific connection relationships between human skeletal joints. For example, the shoulder joint connects the upper arm bone and the collarbone, and the elbow joint connects the upper arm bone and the forearm bone, etc. The initial joint connection graph constructed according to these relationships can be expressed as G=(V, E), where V represents the set of joints, and E represents the set of connection edges between joints. The edge weights are dynamically adjusted by learnable parameters. Let the edge e ij connect joints v i and v j , and the edge weight w ij can be adjusted by the following formula:

[0083]

[0084] where is the initial edge weight, and Δw ij is the weight adjustment amount obtained through learning. In this way, the connection weights between joints can be adaptively optimized according to different motion data, better reflecting the mutual relationship between joints during the motion process.

[0085] The dynamic adjacency matrix learning module uses the gated graph attention mechanism to iteratively update the relationships between nodes. The gated graph attention mechanism determines the influence degree of each node on other nodes by calculating the attention weights of the nodes. For node v i , its attention weight α ij is calculated based on the following formula:

[0086]

[0087] where h i and h j are the feature representations of nodes v i and v j respectively, W1 is the weight matrix, and N i is the set of neighbors of node v iThe set of neighboring nodes, and LeakyReLU is an activation function. Through the gating mechanism, the attention weights are screened and adjusted to generate a spatio-temporal enhanced graph representation, which can more accurately capture the dynamic relationship changes of joint points during movement.

[0088] The adversarial discriminator sub-network is used to constrain the manifold consistency between the feature distribution and the real motion data. During the training process, the generator (the main part of the multi-task graph convolutional network) tries to generate samples as close as possible to the feature distribution of the real motion data, while the discriminator tries to distinguish between the generated samples and the real samples. Through adversarial training, the feature distribution generated by the generator becomes more similar to the manifold structure of the real motion data, improving the accuracy of motion pattern recognition.

[0089] The multi-head classifier outputs the motion category probability, action phase label, and energy consumption estimate in parallel. For the input fused feature tensor, different branches of the multi-head classifier perform classification and regression operations respectively. Taking motion category classification as an example, assuming there are K motion categories, through the linear transformation and Softmax function of the classifier, the probability distribution P(y = k|x) of each motion category is output:

[0090]

[0091] where, W k and b k are the weights and biases of the classifier, and x is the input feature tensor. In this way, a comprehensive analysis and recognition of the motion pattern are realized, providing rich motion information for users.

[0092] Example 5:

[0093] This example details the working process of the anomaly detection module and the operation mechanism of the reinforcement learning-driven decision tree model in the feedback generation module, demonstrating how the system detects motion anomalies and generates personalized feedback solutions to ensure the safety and effectiveness of the motion.

[0094] The anomaly detection module works based on a hybrid model of a variational autoencoder and an isolation forest algorithm. First, the variational autoencoder models the feature distribution of normal motion patterns. The variational autoencoder consists of an encoder and a decoder. The encoder maps the input fused feature tensor x to the latent space z, and its mapping relationship can be expressed as z = Encoder(x), usually implemented through a neural network, such as a multi-layer perceptron. The decoder then reconstructs the vector z in the latent space into the output By minimizing the reconstruction error, such as the mean squared error (MSE) loss function:

[0095]

[0096] where, N is the number of samples, xi and are the original feature and the reconstructed feature of the i-th sample respectively. During the training process, the variational autoencoder learns the feature distribution of normal motion patterns. When abnormal data is input, a large reconstruction error will be generated.

[0097] Next, the isolation forest algorithm is used to recursively partition the high-dimensional feature space. The isolation forest algorithm is based on the assumption that in the data space, points far from other data points are more likely to be outliers. The algorithm constructs multiple isolation trees. For each data point, the path length of the point in the tree is calculated. The shorter the path length, the more isolated the point is and the more likely it is to be an outlier. Let the path length of the data point x in the isolation tree T be h T (x), and the calculation of its outlier score S(x) is based on the following formula:

[0098]

[0099] where E(h T (x)) is the average value of the path lengths of the data point x in multiple isolation trees, and c(n) is a correction coefficient related to the number of samples n.

[0100] Finally, the reconstruction error probability and the subspace outlier degree are fused to generate a comprehensive outlier score and the associated joint point index. Assume the reconstruction error probability is P recon (x), and the subspace outlier degree is A(x). The comprehensive outlier score R(x) can be calculated by weighted summation:

[0101] R(x) = w1P recon (x) + w2A(x)

[0102] where w1 and w2 are weight coefficients, which can be determined through experiments or other optimization methods according to the actual situation. Through this comprehensive outlier score, it is possible to accurately judge whether the motion data is abnormal and locate the joint point where the abnormality occurs.

[0103] The feedback generation module generates a personalized feedback plan through a decision tree model driven by reinforcement learning. The model includes a state encoder, a policy search tree, and a reward function optimizer.

[0104] The state encoder maps the motion pattern recognition result and the anomaly detection data to a low-dimensional state vector. Assume the motion pattern recognition result is M and the anomaly detection data is A. The state encoder converts them into a low-dimensional state vector s through a neural network or other means:

[0105] s = Encoder state (M, A)

[0106] where Encoder stateRepresents the mapping function of the state encoder. In this way, complex motion data and anomaly information are converted into a form suitable for processing by the decision tree model.

[0107] The policy search tree explores the action space through Monte Carlo tree search to generate candidate feedback paths. Monte Carlo tree search is a heuristic search algorithm based on random simulation. It finds the optimal policy by continuously simulating the game or decision-making process. In the feedback generation module, the motion scenario is regarded as a decision-making process, and each decision point corresponds to different feedback actions. For example, when an abnormal motion posture is detected, the decision points may include actions such as adjusting the angle of a certain joint and changing the motion speed. Monte Carlo tree search finds the optimal candidate feedback path by simulating different action sequences multiple times and calculating the rewards of each action sequence.

[0108] The reward function optimizer dynamically adjusts the policy weights according to the degree of improvement in motion efficiency, the speed of anomaly correction, and the physiological load constraint. Let the degree of improvement in motion efficiency be E improve , the speed of anomaly correction be V correct , and the physiological load be L physio , and the reward function R total can be expressed as:

[0109] R total = w3E improve + w4V correct - w5L physio

[0110] where w3, w4, and w5 are weight coefficients. The reward function optimizer dynamically adjusts these weight coefficients according to the actual motion data and goals, so that the generated feedback scheme can not only effectively improve motion efficiency, quickly correct abnormal actions, but also avoid excessive physiological load on athletes.

[0111] The personalized feedback scheme includes motion posture correction parameters, training intensity adjustment suggestions, abnormal action avoidance strategies, and real-time biomechanical load warnings. For example, when an abnormal action of knee valgus is detected during running, the motion posture correction parameters may include specific values for adjusting the knee angle; the training intensity adjustment suggestions may recommend increasing or decreasing the running speed or distance according to the current exercise intensity and the physical condition of the athlete; the abnormal action avoidance strategy can be to provide a demonstration of alternative correct actions or remind the athlete to avoid certain incorrect actions; the real-time biomechanical load warning monitors the force conditions of various parts of the athlete's body and issues a warning message in a timely manner when the load exceeds the safety threshold, such as through a voice prompt command stream "Your knee is under excessive stress, please adjust your running posture".

[0112] These personalized feedback solutions are finally encoded into executable instruction sets by edge computing nodes, such as tactile feedback pulse sequences, which, through tactile feedback devices installed on the athlete's body, remind the athlete to adjust their movements with pulses of different frequencies and intensities; augmented reality visualization annotations, which display adjustment prompts for the movement posture and annotations for abnormal movements on the augmented reality device worn by the athlete; mechanical exoskeleton drive signals, for scenarios where mechanical exoskeletons are used to assist movement, driving the mechanical exoskeleton to adjust the athlete's movement posture according to the feedback solution; voice prompt instruction streams, which directly inform the athlete of the actions to be taken with concise and clear voice messages. Through these executable instruction sets, athletes can obtain feedback in a timely manner and adjust their movement behaviors, improving the movement effect and safety.

[0113] It should be noted that in this article, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the terms "include", "comprise" or any other variant thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements not only includes those elements, but also includes other elements not expressly listed, or further includes elements inherent to such process, method, article or device.

[0114] Although the embodiments of the present invention have been shown and described, for those of ordinary skill in the art, it can be understood that various changes, modifications, substitutions and variations can be made to these embodiments without departing from the principles and spirit of the present invention. The scope of the present invention is defined by the appended claims and their equivalents.

Claims

1. A motion analysis system based on artificial intelligence, characterized in that: It includes a multi-source data acquisition module, a spatiotemporal preprocessing module, a feature fusion module, a motion pattern recognition module, an anomaly detection module and a feedback generation module; The multi-source data acquisition module is used to acquire the original multimodal data required for motion analysis, specifically by deploying an inertial measurement unit, an optical sensor and a depth camera array for synchronous data acquisition, generating an original motion data set, and transmitting the original motion data set to a spatiotemporal preprocessing module; The synchronous data acquisition specifically refers to acquiring motion sequence data at different sampling rates, including multi-band data at 30 frames / second, 60 frames / second, and 120 frames / second, and generating spatiotemporal synchronous data through dynamic timestamp alignment, and storing it in distributed edge nodes; The spatiotemporal preprocessing module is used to perform spatiotemporal calibration and noise reduction on the raw data, specifically aligning heterogeneous sensor data through a multi-scale dynamic time warping algorithm, and using an adaptive Kalman filter to eliminate motion noise, generate optimized motion data, and transmit it to the feature fusion module; The feature fusion module is used to extract and fuse spatiotemporal features. Specifically, the feature fusion module extracts the skeletal joint trajectory features, motion speed distribution features and posture angle change features through a three-dimensional residual network optimized based on the spatiotemporal convolution attention mechanism, generates a fused feature tensor, and inputs it into the motion pattern recognition module and the anomaly detection module. The motion pattern recognition module is used to classify and recognize motion patterns, specifically, based on the fused feature tensor, a multi-task graph convolutional network is used to jointly generate adversarial training models to output motion category probability distribution and key action labels; The anomaly detection module is used to detect abnormal behavior during movement, specifically a hybrid model based on a variational autoencoder and an isolation forest algorithm, which performs reconstruction error analysis and distribution offset detection on the fused feature tensor to generate anomaly scores and positioning results; The feedback generation module is used to generate motion optimization suggestions and exception handling strategies. Specifically, based on the motion pattern recognition results and exception detection data, a personalized feedback plan is generated through a decision tree model driven by reinforcement learning, and the plan is encoded into an executable instruction set.

2. The motion analysis system based on artificial intelligence according to claim 1, characterized in that: The original multimodal data includes skeletal joint coordinate data, accelerometer data, gyroscope angular velocity data, depth image sequence and RGB video stream data; the spatiotemporal synchronization data further includes sensor calibration parameters and ambient light compensation information.

3. The motion analysis system based on artificial intelligence according to claim 2, characterized in that: The multi-scale dynamic time warping algorithm specifically includes a time series slicing operator, a local similarity measurement matrix construction operator and a global path optimization operator; The time series slicing operator is used to divide the heterogeneous sensor data into segments according to windows and calculate the similarity between the segments through dynamic programming; The local similarity measurement matrix construction operator specifically uses dynamic time warping distance combined with cosine similarity to generate a multimodal alignment cost matrix; The global path optimization operator specifically solves the minimum cumulative cost path based on a branch and bound algorithm to achieve frame-level alignment of cross-sensor data.

4. The motion analysis system based on artificial intelligence according to claim 3, characterized in that: The three-dimensional residual network includes parallel spatiotemporal convolution branches, wherein the first branch uses dilated convolution to extract long temporal dependency features, the second branch optimizes local motion feature weights through a channel attention mechanism, and the third branch uses dilated spatial pyramid pooling to capture multi-scale spatial context information.

5. The motion analysis system based on artificial intelligence according to claim 4, characterized in that: The multi-task graph convolutional network joint adversarial generation training model specifically includes a skeleton topology graph construction module, a dynamic adjacency matrix learning module, an adversarial discriminant subnetwork and a multi-head classifier; The skeleton topology construction module is used to construct an initial joint point connection diagram based on human anatomy priors and dynamically adjust edge weights through learnable parameters; The dynamic adjacency matrix learning module specifically utilizes a gated graph attention mechanism to iteratively update the relationship between nodes and generate a spatiotemporal enhanced graph representation; The adversarial discriminant subnetwork is used to constrain the manifold consistency of feature distribution and real motion data; The multi-head classifier specifically outputs motion category probabilities, action stage labels and energy consumption estimates in parallel.

6. The motion analysis system based on artificial intelligence according to claim 5, characterized in that: The workflow of the anomaly detection module includes: The characteristic distribution of normal motion patterns is modeled through a variational autoencoder, and the reconstruction error of the input data is calculated; The isolation forest algorithm is used to recursively segment the high-dimensional feature space to detect abnormal subspaces that deviate from the mainstream distribution; The reconstruction error probability and subspace anomaly degree are fused to generate a comprehensive anomaly score and the associated joint point index.

7. The motion analysis system based on artificial intelligence according to claim 6, characterized in that: The reinforcement learning driven decision tree model specifically includes a state encoder, a policy search tree and a reward function optimizer; The state encoder is used to map the motion pattern recognition results and the abnormality detection data into a low-dimensional state vector; The strategy search tree specifically explores the action space through Monte Carlo tree search to generate candidate feedback paths; The reward function optimizer dynamically adjusts the strategy weight according to the degree of improvement in movement efficiency, the speed of abnormal correction and the physiological load constraint.

8. The motion analysis system based on artificial intelligence according to claim 7, characterized in that: The personalized feedback program includes movement posture correction parameters, training intensity adjustment suggestions, abnormal movement avoidance strategies and real-time biomechanical load warnings.

9. The motion analysis system based on artificial intelligence according to claim 8, characterized in that: The executable instruction set may be encoded into at least one of the following forms through an edge computing node: a tactile feedback pulse sequence, an augmented reality visualization annotation, a mechanical exoskeleton drive signal, and a voice prompt instruction stream.

Citation Information

Patent Citations

  • Gait anomaly early recognition and risk early warning method and device

    CN111700620A

  • Lip language recognition method based on hybrid three-dimensional residual gating circulation unit

    CN116884412A

  • Human body behavior recognition and data acquisition system based on artificial intelligence

    CN118430070A

  • Athlete throwing action analysis and training method and equipment based on deep learning

    CN119068558A

  • Predictive maintenance management method and system for industrial robot

    CN119295060A

Cited By

  • Remote senile health assessment state prediction system based on Kalman optimization tree

    CN120432172A

  • Intelligent AI construction method and system for track and field event referees

    CN120673483A

  • Athletic event judging intelligent AI construction method and system

    CN120673483B

  • Running data authenticity evaluation method and system based on biological feature recognition

    CN121350919A

  • Intelligent physical education data management method and system

    CN121544127A