Agile motion capture and instruction feedback method and system suitable for sports competition
By using multimodal data fusion and real-time modeling technology, the accuracy and feedback issues of motion capture systems in complex motion scenarios have been solved, enabling high-precision personalized feedback and training optimization, and reducing the risk of injury.
Patent Information
- Application Number
- CN202510917653.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-03
- Publication Date
- 2025-11-18
AI Technical Summary
Existing motion capture systems lack accuracy in complex motion scenarios, suffer from data processing delays and limited feedback methods, and have insufficient multimodal data fusion, making it difficult to meet the demands for high precision and real-time feedback.
Multimodal data is collected using ultra-wideband technology, inertial sensors, and cameras. Data fusion and modeling are performed using graph neural networks and spatiotemporal graph convolutional networks, and real-time feedback is provided by combining edge computing and large language models.
It improves motion capture accuracy, provides personalized real-time feedback, enhances multimodal data fusion capabilities, helps athletes correct movement problems in a timely manner, optimizes training effects, and reduces the risk of injury.
Smart Images

Figure CN120977487A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the field of image processing, and in particular to a method and system for agile motion capture and instruction feedback suitable for sports competitions. BACKGROUND
[0002] In sports competitions, motion capture technology is widely used in athlete action analysis, training optimization, referee assistance, and audience experience enhancement, etc. Human motion data is recorded through sensors, cameras or wearable devices. Existing motion capture systems mainly have the following problems: 1. Limited motion capture accuracy. On the one hand, traditional motion capture systems usually rely on cameras or basic sensors, which are affected by visual obstructions or environmental interference in complex motion scenarios, resulting in insufficient accuracy. On the other hand, current positioning technologies (such as GPS or inertial sensors) are often not accurate enough in indoor or dense environments, making it difficult to meet the needs of high-precision training and competition.
[0003] 2. Delay in data processing and real-time feedback. Data processing usually requires a long delay, especially in multi-sensor and multi-user scenarios, making it difficult to provide immediate feedback, which is unacceptable for dynamic and rapidly changing sports competitions. Many systems lack a deep understanding of complex motion patterns and cannot provide effective and personalized feedback based on the real-time performance of athletes.
[0004] 3. Single feedback mode. Current feedback methods mainly use basic text or chart forms, which cannot provide more interactive, immediate and intelligent guidance for athletes. There are certain limitations in the personalization and intelligence of feedback, making it difficult to provide precise guidance based on the training progress and technical characteristics of different athletes.
[0005] 4. Insufficient multi-modal data fusion. Existing systems often process different types of data (such as visual data, inertial data, and location data) separately, lack efficient fusion strategies, and result in low data utilization and inability to accurately capture complex motion patterns.
[0006] Therefore, there is an urgent need to provide a method for agile motion capture and instruction feedback suitable for sports competitions. SUMMARY
[0007] To solve the above problems, the present application provides a method for agile motion capture and instruction feedback suitable for sports competitions, which can improve motion capture accuracy, provide real-time personalized feedback, and enhance multi-modal data fusion capabilities.
[0008] According to a first aspect of the present application, a method for agile motion capture and instruction feedback suitable for sports competitions is provided, comprising: S1, real-time collection of athlete position information and capture of athlete action details, multi-modal fusion of position information and action details; S2, sending the data signal of the multi-modal fused position information and action details to the edge computing node after beamforming for processing, and then synchronizing the data; S3, modeling according to the synchronized motion data, analyzing the athlete's movement pattern and predicting future action changes; S4, training feedback according to the analysis result.
[0009] In the above scheme, step S1 includes: Determine the position information of the athlete by transmitting and receiving pulse signals through ultra-wideband technology; Determine the acceleration, angular velocity and direction change of the athlete according to the inertial sensor; Capture the athlete's posture according to the camera; Multi-modal fusion of position information, inertial sensor data and camera capture data to form a real-time data set.
[0010] In the above scheme, step S2 includes: S21, adjust the data weighting coefficient of the multi-modal fused position information and action details to concentrate in the set direction for receiving, and send the data signal to the edge computing node; S22, clean and denoise the data, and then synchronize through timestamp.
[0011] In the above scheme, step S3 includes: S31, model the human joints based on graph neural network and real-time motion data; S32, based on graph attention mechanism, weight and fuse position information and action details to form joint features and adjust the human joint model; S33, based on spatiotemporal graph convolution network, establish the time series data of human action, capture the time change and spatial relationship of each joint in the human joint model and predict the future action change.
[0012] In the above scheme, step S4 includes: Convert the athlete's action data and analysis result through a large language model, and generate natural language feedback to athletes for improvement suggestions.
[0013] In the above scheme, the feedback output mode in step S4 includes: Voice feedback, text feedback and real-time data visualization feedback.
[0014] In the above scheme, it also includes: Based on the large language model conversion of the action data and the content of the analysis result of the athletes, a training plan is generated.
[0015] In the above-mentioned scheme, further comprising: According to the movement data and action analysis of the athletes, real-time tactical adjustment is carried out, and injury warning is carried out according to the deviation between the action of the athletes and the standard action.
[0016] According to the second aspect of the technical scheme of the present application, a system suitable for agile action capture and instruction feedback in sports competition is provided, which is used to realize the method suitable for agile action capture and instruction feedback in sports competition in any of the above-mentioned schemes, and the system comprises: The acquisition module is used for real-time acquisition of the position information of the athletes and capture of the action details of the athletes, and the position information and the action details are fused in multiple modes; The processing module is used for sending the data signals of the position information and the action details fused in multiple modes to the edge computing node for processing after beamforming, and then synchronizing the data; The analysis module is used for modeling according to the synchronized movement data, analyzing the movement mode of the athletes and predicting the future action change; The feedback module is used for training feedback according to the analysis result.
[0017] According to the third aspect of the technical scheme of the present application, an electronic device is provided, which comprises: A memory stores executable instructions; The processor runs the executable instructions in the memory to realize the method in any of the above-mentioned schemes.
[0018] The beneficial effects of the present application are: The method and system suitable for agile action capture and instruction feedback in sports competition disclosed in the present application capture the action of the athletes, help the athletes to find problems in time and correct them, optimize the action quality, improve the competitive performance and reduce the risk of injury; in the competition, the performance and action data of the athletes are monitored in real time, and the coach is assisted to make tactical adjustment; according to the historical data and training progress of the athletes, a personalized training plan is generated to ensure the pertinence and effect of the training and improve the comprehensive competitive ability of the athletes. BRIEF DESCRIPTION OF DRAWINGS
[0019] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings needed in the embodiments or prior art description. Obviously, the drawings in the following description only represent some of the embodiments of the present application, and for those skilled in the art, other drawings can also be obtained from the structures shown in the drawings without creative labor.
[0020] Figure 1 The flow chart of the method for agile motion capture and instruction feedback suitable for sports competition provided by the embodiments of the present application.
[0021] The implementation, functional features and advantages of the present application will be further described with reference to the embodiments and the accompanying drawings. DETAILED DESCRIPTION
[0022] The exemplary embodiments will be described in detail herein with reference to the accompanying drawings. When the following description refers to the drawings, the same numbers in different drawings represent the same or similar elements unless otherwise indicated. The implementations described in the following exemplary embodiments do not represent all implementations consistent with the present disclosure. Instead, they are merely examples of apparatuses and methods consistent with some aspects of the present disclosure as detailed in the appended claims.
[0023] The terms "first", "second", and the like in the description and claims of the present disclosure are used for distinguishing between similar objects and do not necessarily have a particular order or a particular spatial or chronological precedence. It is to be understood that the data so designated are interchangeable under appropriate circumstances such that the embodiments of the present disclosure described herein can operate in other sequences than the one illustrated or described herein.
[0024] In addition, the terms "comprise", "comprising", "include", "including", and their conjugates, denote an open-ended inclusion, such that the processes, methods, articles, or apparatuses that "comprise", "comprising", "include", "including" something lack limitation to the scope of the enclosed process, method, article, or apparatus and include additional or other processes, methods, articles, or apparatuses that lack expressly including the enclosed limitation.
[0025] Plural, including two or more.
[0026] And / or, it should be understood that the term "and / or" used in the present disclosure only describes the association relationship of the associated objects, and means that there can be three relationships. For example, A and / or B can represent the three cases of A alone, A and B together, and B alone.
[0027] As Figure 1As shown, one embodiment of the technical solution of the present application provides a method for agile motion capture and instruction feedback suitable for sports competition, comprising: S1, real-time collection of position information of athletes and capture of action details of athletes, multi-modal fusion of position information and action details; S2, after the multi-modal fusion of the position information and the action details data signal, the data signal is sent to the edge computing node for processing after beamforming, and then the data is synchronized; S3, modeling according to the synchronized motion data, analyzing the athlete's movement mode and predicting the future action change; S4, according to the analysis result, training feedback.
[0028] In step S1, the position information of the athletes is determined by transmitting and receiving pulse signals through ultra-wideband technology, multiple base stations are arranged in the sports field, and the athletes wear UWB tags. Further, the position of each location needs to be determined by at least three base stations for triangulation positioning. By measuring the distance from each base station to the athlete's tag, the position of the athlete is solved. Because the UWB signal has the characteristics of ultra-wideband, especially in indoor environment, it can realize centimeter-level positioning, which ensures accurate tracking of athletes.
[0029] According to the inertial sensor including accelerometer, gyroscope and magnetometer, the acceleration, angular velocity and direction change of the athlete are determined.
[0030] According to the camera, the posture of the athlete is captured, and the image captured by the multi-view camera is used to recover the position of the athlete in three-dimensional space according to the triangulation method. Further, a deep learning algorithm can be used to calculate the relative position between the athlete and the camera, further improving the accuracy.
[0031] The obtained position information, acceleration, angular velocity, direction change and athlete posture data are multi-modal fused, and the multi-modal fusion method includes graph attention mechanism fusion and deep neural network fusion.
[0032] Step S2 includes: S21, adjusting the data weighting coefficient of the multi-modal fused position information and action details to concentrate in the set direction for receiving, and sending the data signal to the edge computing node; S22, cleaning and denoising the data, and then synchronizing through timestamp.
[0033] In step S21, the beamforming technology is used to control the signal radiation direction and receiving direction, so that the signal can be concentrated in a certain direction, thereby improving the signal strength and transmission efficiency and reducing unnecessary interference. After beamforming, the data is transmitted to the edge computing device for edge computing, preferably using 5G transmission to provide low-latency and high-reliability communication services.
[0034] In step S22, edge computing is preferably used to clean and denoise the data. Through data cleaning, noise in the sensor, missing data filling, and outlier detection are removed to improve the efficiency of data analysis. Kalman filtering is used to smooth the motion data and reduce random noise in angular velocity and acceleration. Common denoising methods also include mean filtering and median filtering.
[0035] Further, edge computing also has preliminary data analysis and processing functions. For example, simple motion analysis based on sensor data, anomaly detection, etc. The edge computing device can process the data in real time through deep learning or machine learning models to analyze the athlete's action state in advance and provide support for subsequent action prediction and feedback generation.
[0036] Since multiple sensors collect data at different timestamps, the data of these sensors needs to be time-synchronized. This includes timestamp-based synchronization, and other methods such as network time protocol synchronization can also be used.
[0037] Step S3 includes: S31, based on graph neural network and real-time motion data, modeling human joints; S32, based on graph attention mechanism, weighting and fusing position information and action details to form joint features, adjusting the human joint model; S33, based on spatio-temporal graph convolution network to establish time series data of human action, capture the time variation and spatial relationship of each joint in the human joint model and predict future action changes.
[0038] In step S31, the graph neural network (GNN) is used to model the spatio-temporal relationship of human joints based on real-time motion data to analyze the athlete's action pattern. The graph neural network dynamically updates the topological relationship of joint nodes and edges through the graph structure, and identifies the action quality and deviation of the athlete in real time. The joints of the human body can be regarded as a node, and the topological relationship between nodes is the connection between joints. The core idea of the graph neural network is to propagate information through the graph structure, and the information transmission and update from one node to adjacent nodes. For each node, it will be updated according to the characteristics of the neighbor nodes and the connection relationship. The update formula of the graph neural network is: (1) where, is the feature representation of node v at the l-th layer, is the set of neighbor nodes of node v, AGGREGATE is an aggregation function used to integrate the information of neighbor nodes.
[0039] In each layer, the representation of a node is updated based on the information of its neighbor nodes, with the specific update formula being: (2) where, is an activation function, is a weight matrix, is a bias term.
[0040] In human motion analysis, the spatio-temporal relationship of joints is modeled through dynamic graph neural networks, and the state of each joint changes over time. Graph neural networks capture these changes by dynamically updating the connections and features of nodes.
[0041] In step S32, the weighted fusion is performed through the graph attention mechanism, which strengthens more important information by assigning different weights to each neighbor node. It can adaptively assign different importance to each neighbor node, thereby improving the fusion effect and the accuracy of action recognition.
[0042] For node v, the feature h u is updated through the graph attention mechanism as follows: (3) where, is the attention coefficient between node v and node u, which is usually calculated by the following formula: (4) where a is a learnable weight, W is a transformation matrix of node features, || represents a feature concatenation operation, and LeakyReLU is an activation function.
[0043] In step S33, the spatio-temporal graph convolutional network combines graph convolution and temporal convolution. It captures the spatial relationship between joints by performing convolution operations on the graph structure, and captures the temporal features of actions by performing convolution operations in the time dimension.
[0044] Specifically, it includes: spatial convolution, for node v, its feature The update formula at time t is: (5) where Ws is a spatial convolution kernel that captures the spatial relationship between joints.
[0045] Temporal convolution is used to capture the temporal changes of actions, with the formula being: (6) wherein, is a weighting coefficient between time steps, is a set of time steps for node v.
[0046] By training the spatio-temporal graph convolution model, the system can predict future movement changes of the athlete based on current movement data. For example, the system can predict the movement deviation or inaccurate posture that an athlete may have in the next few seconds.
[0047] Step S4 includes: The action data of the athlete and the analysis results are converted by the large language model, and natural language feedback athlete improvement suggestions are generated. Among them, the feedback methods include voice feedback, picture-text feedback and real-time data visualization feedback.
[0048] Based on the action analysis and spatio-temporal graph convolution model prediction, the system generates natural language feedback through the large language model to provide real-time action improvement suggestions to the athlete. These feedbacks can be presented to the athlete in the form of voice, image or augmented reality (AR) to help them optimize their actions. For example: if the system detects that the athlete's take-off angle is low, the large language model will generate feedback similar to "your take-off angle is low, suggest adjusting to 30°" and present it to the athlete through voice or picture-text. Through the above analysis and feedback generation process, the system can accurately identify the athlete's action problems and provide real-time and personalized improvement suggestions to improve the athlete's training effect and competition performance.
[0049] Voice feedback provides real-time guidance to athletes through intelligent voice assistants such as Alexa, Siri, etc. This feedback method is very suitable for athletes in high-intensity training, who do not need to stop to check the screen to get immediate suggestions or corrections.
[0050] Picture-text feedback presents the athlete's action data through charts, dynamic graphs, and augmented reality (AR) technology. AR technology can superimpose a virtual joint model on the athlete's real-time action, allowing the athlete to visually see the deviation and improvement direction of the action. Charts and dynamic graphs help coaches and athletes observe the trend of action data changes, facilitating further analysis and decision-making. The generation of picture-text feedback combines data visualization technology, augmented reality technology and graph neural network analysis results. Through the display of 3D type, action trajectory and deviation graph, the system compares the athlete's movement posture with the standard action and marks the deviation.
[0051] Real-time data visualization displays the trends of athletes' motion data and key indicators (such as speed, acceleration, and angle). This data can help athletes and coaches make quick decisions about the real-time training status, identify potential problems, and adjust training strategies in a timely manner.
[0052] The invention also includes: making real-time tactical adjustments based on the athlete's motion data and movement analysis, and providing injury warnings based on the deviation between the athlete's movements and standard movements.
[0053] Personalized training plans are generated based on athletes' historical training data, real-time feedback, and movement analysis results, automatically adjusting training goals and continuously optimizing the process. Because each athlete has different exercise habits, fitness levels, and training goals, training plans are tailored to their individual needs, allowing for targeted adjustments and training.
[0054] Specifically, by analyzing historical training data, the system identifies athletes' strengths and weaknesses. For example, if an athlete has low accuracy in a particular movement, the system will increase the intensity or frequency of training for that movement. When generating personalized training plans, reinforcement learning algorithms are used to self-adjust the plans, ensuring they are dynamically optimized based on the athlete's progress. Furthermore, after each training session, the training plan is adjusted based on feedback and the athlete's actual performance.
[0055] In addition, during competitions, coaches and athletes can use the system to monitor athletes' performance in real time. The system can provide dynamic tactical adjustment suggestions based on athletes' real-time action data and the competition environment.
[0056] Specifically, by analyzing the athlete's position, speed, acceleration, and the position of opposing defenders, the system can identify weaknesses in the opponent's defense and provide real-time tactical suggestions to coaches or athletes. Through inertial sensors, heart rate sensors, and other data on physical condition and replacement suggestions, the system monitors the athlete's energy expenditure, analyzes the trend of declining physical strength in real time, and provides replacement suggestions when physical strength drops to a certain level. By monitoring the athlete's heart rate, acceleration, and exercise load, the system uses a physical fitness analysis model to assess the athlete's physical condition. For example: "The opponent's defense is weaker on the right side, so we suggest increasing the frequency of our attacks on the right side." "Athlete A's physical condition has declined, and it is recommended to replace him." By monitoring and analyzing athletes' movements in real time, the system can promptly identify movement patterns that may lead to injury and issue warnings to help athletes avoid improper movements.
[0057] The system can monitor the athlete's movements in real time, identify potential injury risks by analyzing the athlete's movement patterns. Especially in high-intensity competition or training scenarios, non-standard movement patterns may cause athletes to be injured, the system can issue warnings in advance to help athletes avoid dangerous movements, thereby effectively preventing sports injuries. Including: action deviation monitoring, joint angle and load analysis, and load prediction of athlete movements.
[0058] Action deviation monitoring: based on graph neural networks to model the spatio-temporal relationship of athlete joints, by comparing the differences between the actual movements of the athlete and the standard movements, abnormal behaviors are detected in time.
[0059] Joint angle and load analysis: by analyzing the angle changes and load of the athlete's joints, especially the angle deviation of high-risk parts such as the knee joint and shoulder joint, to detect whether there are abnormal angle changes, and then assess the risk of injury.
[0060] Load prediction of athlete movements: through the information captured by the inertial sensor, it predicts whether the load assessment of the athlete's movements will cause excessive fatigue or injury.
[0061] Further, all the training data of the athletes, the action analysis results and the feedback will be uploaded to the cloud in real time for storage and backup. These data provide valuable information when reviewing after the game, coaches and athletes can view long-term tracking data to find the progress and shortcomings of athletes in training, and optimize subsequent training.
[0062] Cloud storage ensures the reliable preservation of all data and can be accessed for analysis and review at any time. The system uses big data analysis technology to mine and model all the training data of the athletes to form a long-term performance analysis report of the athletes. By aggregating the data of each training, the system can generate more accurate long-term training goals and improvement suggestions for athletes.
[0063] According to the second aspect of the technical scheme of the application, a system suitable for agile motion capture and instruction feedback in sports competition is provided, which is used to implement the method of any one of the above schemes. The system includes: The acquisition module is used for real-time acquisition of the position information of the athlete and capture of the action details of the athlete, and multi-modal fusion of the position information and the action details; The processing module is used for sending the data signals of the multi-modal fused position information and action details to the edge computing node for processing after beamforming, and then synchronizing the data; The analysis module is used for modeling according to the synchronized motion data, analyzing the movement patterns of the athletes and predicting future changes in movements; The feedback module is configured to provide training feedback based on the analysis result.
[0064] According to a third aspect of the present application, an electronic device is provided, comprising: a memory storing executable instructions; a processor configured to execute the executable instructions in the memory to implement the method according to any one of the above aspects.
[0065] It should be noted that, in this document, the terms "comprising", "including", or any other variant thereof are intended to cover a non-exclusive inclusion, such that processes, methods, articles, or devices that comprise a list of elements do not include only those elements recited, but also other elements that are not expressly listed or inherent to such processes, methods, articles, or devices. Without further limitation, an element preceded by "comprises a" does not, without more constraints, foreclose the existence of additional identical elements in the process, method, article, or device that includes the element.
[0066] The above-mentioned embodiment numbers of the present application are only for description, and do not represent the advantages or disadvantages of the embodiments.
[0067] From the above description of the embodiments, those skilled in the art can clearly understand that the above-mentioned implementation methods can be realized by means of software and necessary general hardware platforms, of course, they can also be realized by hardware, but in many cases the former is a better implementation method. Based on such understanding, the technical solutions of the present application can be embodied in the form of a software product, which is stored in a storage medium (such as a ROM / RAM, a magnetic disk, or an optical disk) and includes a number of instructions for making a terminal (which can be a mobile phone, a computer, a server, an air conditioner, or a network device) execute the methods described in the various embodiments of the present application.
[0068] The embodiments of the present application are described above in conjunction with the accompanying drawings, but the present application is not limited to the above-mentioned specific embodiments, which are only illustrative and not limiting. Those skilled in the art can make many forms under the inspiration of the present application without departing from the purpose of the present application and the scope of the claims, which are all within the protection of the present application.
Claims
1. A method for agile motion capture and instruction feedback suitable for sports competitions, characterized in that, include: S1. Real-time acquisition of athlete's position information and capture of athlete's movement details, and multimodal fusion of position information and movement details; S2. The multimodal fusion of position information and action details data signals is beamformed and sent to the edge computing node for processing, and then the data is synchronized. S3. Model the motion data after synchronization, analyze the athlete's motion patterns and predict future motion changes; S4. Provide training feedback based on the analysis results.
2. The method for agile motion capture and command feedback applicable to sports competitions according to claim 1, characterized in that, Step S1 includes: The athlete's location information is determined by transmitting and receiving pulse signals using ultra-wideband technology. The athlete's acceleration, angular velocity, and changes in direction are determined using inertial sensors; Based on the athlete's posture captured by the camera; Location information, inertial sensor data, and camera-captured data are fused in a multimodal manner to form a real-time dataset.
3. The method for agile motion capture and instruction feedback in sports competitions according to claim 1, characterized in that, Step S2 includes: S21. Adjust the weighting coefficients of the position information and action details after multimodal fusion to concentrate the reception in the set direction and send the data signal to the edge computing node; S22. Clean and denoise the data, and then synchronize it using timestamps.
4. The method for agile motion capture and command feedback applicable to sports competitions according to claim 1, characterized in that, Step S3 includes: S31. Based on graph neural networks and real-time motion data, model human joints; S32. Based on the graph attention mechanism, position information and action details are weighted and fused to form joint features, which are then used to adjust the human joint model. S33. Establish time series data of human motion based on spatiotemporal graph convolutional network, capture the temporal changes and spatial relationships of each joint in the human joint model, and predict future motion changes.
5. The method for agile motion capture and command feedback applicable to sports competitions according to claim 1, characterized in that, Step S4 includes: The large language model transforms athletes' motion data and analysis results, and generates natural language feedback and suggestions for athletes to improve.
6. The method for agile motion capture and command feedback applicable to sports competitions according to claim 1, characterized in that, The feedback output methods in step S4 include: Voice feedback, text and image feedback, and real-time data visualization feedback.
7. The method for agile motion capture and instruction feedback in sports competitions according to claim 5, characterized in that, Also includes: Based on the large language model, the athlete's motion data and analysis results are transformed to generate a training plan.
8. The method for agile motion capture and instruction feedback in sports competitions according to claim 4, characterized in that, Also includes: Based on athletes' athletic data and movement analysis, tactical adjustments are made in real time, and injury warnings are issued based on deviations between athletes' movements and standard movements.
9. A system for agile motion capture and command feedback suitable for sports competitions, characterized in that, The system is used to implement the agile motion capture and instruction feedback method for sports competitions as described in any one of claims 1-8, the system comprising: The acquisition module is used to collect athletes' location information in real time, capture the details of their movements, and perform multimodal fusion of location information and movement details. The processing module is used to send the multimodal fusion location information and motion details data signals to the edge computing node for processing after beamforming, and then synchronize the data; The analysis module is used to model based on synchronized motion data, analyze athletes' movement patterns, and predict future movement changes; The feedback module is used to provide training feedback based on the analysis results.
10. An electronic device, characterized in that, The electronic device includes: Memory, which stores executable instructions; A processor that executes the executable instructions in the memory to implement the method of any one of claims 1-8.