Taekwondo leg-out track correction method and system based on multi-modal data processing
By using a multimodal data processing system that combines visual and inertial sensors to collect and fuse data, multimodal correction guidance signals are generated, solving the problem of real-time and accurate correction of leg movements in Taekwondo training and improving training effectiveness and safety.
Patent Information
- Application Number
- CN202511754969.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-11-26
- Publication Date
- 2026-03-06
AI Technical Summary
Existing technologies cannot provide real-time, precise, and intelligent assistance for correcting the leg movements of Taekwondo practitioners. In particular, the lack of personalized, multi-sensory movement correction methods under conventional training conditions makes it difficult to correct errors in the movements of beginners and increases the risk of sports injuries.
A multimodal data processing method is adopted, which synchronously collects multimodal data of the trainee's leg movements through visual sensing unit and inertial sensing unit, performs time synchronization and data fusion, generates three-dimensional spatial posture data, compares it with standard movement model, calculates posture deviation, generates multimodal correction guidance signal, and drives visual, auditory and tactile feedback units to provide real-time correction feedback.
It enables real-time, precise, and personalized movement correction for Taekwondo trainees, improving learning efficiency and training immersion, reducing the risk of sports injuries, and simulating comprehensive coaching guidance.
Smart Images

Figure CN121617550A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of Taekwondo training technology, and in particular to a method and system for correcting Taekwondo leg trajectory based on multimodal data processing. Background Technology
[0002] In the Taekwondo training system, the standardization of leg techniques (such as side kicks and back kicks) is a fundamental reflection of technical level, directly affecting striking power, speed, accuracy, and the risk of sports injuries. Beginners, whose proprioception is not yet established, are highly susceptible to incorrect force application patterns, deviations in movement trajectories, and insufficient joint angles. Once these errors become ingrained, they are not only difficult to correct and affect competitive performance, but also significantly increase the risk of acute and chronic injuries to the knee and ankle joints. Therefore, timely and precise movement correction during training is crucial for ensuring training effectiveness, improving technical level, and preventing sports injuries.
[0003] Currently, mainstream correction methods primarily rely on human coaching and limited intelligent assistive technologies. Human correction mainly depends on the coach's visual observation and experience, using verbal cues, demonstrations, or simple teaching aids like lines and limiters for physical restriction. While this method offers flexibility, it is limited by the coach's attention, experience, and energy, making it difficult to provide frequent, individualized, and immediate feedback to each student in group classes. Furthermore, the correction standards are highly subjective and difficult to quantify. Existing intelligent assistive technologies, such as motion recognition software based on mobile apps or motion capture systems from advanced laboratories, have significant limitations. The former is limited by the accuracy and algorithms of ordinary cameras, resulting in large recognition errors and severe feedback delays (only providing retrospective review, not real-time guidance), failing to provide synchronous correction during movement execution. The latter involves expensive equipment, demanding environmental requirements, and complex operation, making it completely unsuitable for the general training scenarios in conventional dojos. Moreover, existing technologies generally focus on visual feedback, lacking utilization of crucial proprioceptive channels such as touch, failing to simulate the tactile guidance experience of a coach's "hands-on" correction, and hindering the efficient establishment of correct muscle memory. In conclusion, none of the existing technical solutions can provide a personalized movement correction method for Taekwondo beginners that can be implemented in real time, accurately, and incorporating multiple sensory channels under normal training conditions.
[0004] Therefore, how to provide real-time, precise, and intelligent assistance and correction for the leg movements of Taekwondo practitioners is a technical problem that urgently needs to be solved. Summary of the Invention
[0005] In view of this, the present invention provides a method and system for correcting the leg trajectory of Taekwondo practitioners based on multimodal data processing, in order to solve the problem that the existing technology cannot provide real-time and accurate intelligent assistance correction for the leg movements of Taekwondo practitioners.
[0006] The technical solution adopted in this invention is: In a first aspect, the present invention provides a method for correcting the trajectory of a Taekwondo kick based on multimodal data processing, the method comprising: Based on the kinematic characteristics of the preset target Taekwondo leg movements, a standard movement model is constructed. The standard movement model includes multiple sets of standard parameters of the leg movements in the visual image dimension and the limb kinematic dimension. By using a visual sensing unit deployed in the training environment and an inertial sensing unit worn on the trainee's limbs, multimodal raw data is simultaneously collected when the trainee performs the target Taekwondo leg technique. The multimodal raw data includes a continuous image sequence acquired by the visual sensing unit and limb joint motion data acquired by the inertial sensing unit. Based on the multimodal raw data, the received continuous image sequence and the limb joint motion data are time-synchronized and data-fused to generate fused posture data of the trainee's limbs in three-dimensional space. The fused posture data is compared with each set of standard parameters in the standard motion model to calculate multi-dimensional posture deviation information, and based on the posture deviation information, a corresponding multimodal correction guidance signal is generated. Based on the multimodal correction guidance signal, the multimodal feedback execution unit is driven to provide the trainee with real-time corrective feedback adapted to the posture deviation information.
[0007] Preferably, the construction of a standard movement model based on the kinematic characteristics of a preset target Taekwondo leg technique includes: Obtain multiple preset learning stage sequences corresponding to different skill levels, the learning stage sequences including beginner stage, advancement stage and competition stage; A biomechanical analysis was performed on the target Taekwondo leg technique, and the continuous movement process of the leg technique was analyzed into several key posture segments arranged in sequence. For each of the key posture segments, a multi-dimensional standard motion parameter set corresponding to the key posture segment is defined, and a corresponding benchmark value and allowable deviation tolerance range are set for each parameter in the standard motion parameter set. Each stage in the learning stage sequence is associated and mapped with the target configuration of the multi-dimensional standard action parameter set; wherein, the target configuration includes at least the setting of the width of the deviation tolerance range and the setting of the activation priority of different dimension parameters in the multi-dimensional standard action parameter set; In response to the trainee's skill level assessment results, the skill level assessment results are mapped to the corresponding learning stage, and a personalized movement model adapted to the current trainee is generated from the standard movement parameter set as the standard movement model based on the target configuration associated with the learning stage.
[0008] Preferably, the multimodal raw data of the trainee performing the target Taekwondo leg technique is simultaneously collected by a visual sensing unit deployed in the training environment and an inertial sensing unit worn on the trainee's limbs. The multimodal raw data includes a continuous image sequence acquired by the visual sensing unit and limb joint motion data acquired by the inertial sensing unit, including: Before the training of the target Taekwondo leg techniques begins, the internal parameters and external spatial position of the visual sensing unit are calibrated, and the inertial sensing unit is initially zero-biased calibrated. The ambient light intensity of the preset training site is monitored in real time, and the acquisition parameters of the visual sensing unit are dynamically adjusted based on the ambient light intensity to ensure that the imaging quality of the continuous image sequence meets the preset recognition requirements. In response to the signal that the training of the target Taekwondo leg technique has started, a synchronous acquisition command is sent to the visual sensing unit and the inertial sensing unit to acquire the continuous image sequence and the limb joint motion data respectively with the same time base. The acquired continuous image sequence and limb joint motion data are preprocessed, and the integrity of the preprocessed multimodal raw data is verified. The continuous image sequence that has passed integrity verification is packaged with the limb joint motion data to generate the multimodal raw data stream.
[0009] Preferably, the step of performing time synchronization and data fusion processing on the received continuous image sequence and the limb joint motion data based on the multimodal raw data to generate fused posture data of the trainee's limbs in three-dimensional space includes: Based on the synchronization timestamp carried by the synchronization acquisition command, a temporal correspondence is established between the continuous image sequence and the limb joint motion data. The two-dimensional image space perceived by the visual sensing unit and the three-dimensional motion space measured by the inertial sensing unit are registered in the same world coordinate system to obtain data after unified spatiotemporal reference processing. Visual modal features and inertial modal features are extracted from the data after unified processing of the spatiotemporal reference. The visual modal features include the two-dimensional pixel coordinates of the joints identified in each frame of the image, and the inertial modal features include the angular velocity vector and acceleration vector of the limb. The visual modal features and the inertial modal features at the same time point are fused to obtain the fused multimodal features; The fused multimodal features are input into a preset human kinematics model. By solving the model, the three-dimensional spatial posture sequence of the trainee's limbs in the world coordinate system is obtained. The three-dimensional spatial attitude sequence is smoothed and filtered to obtain the fused attitude data.
[0010] Preferably, the step of fusing the visual modal features and the inertial modal features at the same time point to obtain the fused multimodal features includes: The visual modal features and the inertial modal features at the same time point are combined to construct a multimodal feature vector; Based on the confidence level and image sharpness of key point recognition in the image frame associated with the multimodal feature vector, the first uncertainty measure corresponding to the multimodal feature vector is quantified; Based on the predicted values of sensor signal noise level and cumulative error corresponding to the inertial mode features, the second uncertainty measure corresponding to the inertial mode feature vector is quantified. Based on the quantified first uncertainty measure and second uncertainty measure, corresponding fusion weights are assigned to the visual modal features and the inertial modal features, and the fusion weights are used to perform weighted calculations on the visual modal features and the inertial modal features to obtain the fused multimodal features.
[0011] Preferably, the step of comparing the fused posture data with each set of standard parameters in the standard motion model to calculate multi-dimensional posture deviation information, and generating corresponding multimodal correction guidance signals based on the posture deviation information, includes: The real-time motion parameters representing each key posture segment in the fused posture data are compared one by one with the corresponding standard parameters in the standard motion model to calculate the real-time deviation value of each real-time motion parameter. Based on the preset deviation priority rules, all calculated real-time deviation values are filtered and sorted to determine the key deviation item that is currently most prioritized for correction. According to the preset mapping relationship, the key deviation items are mapped to the predefined correction strategy library to generate targeted correction guidance content. The correction strategy library is built based on preset knowledge. The mapping relationship defines the best correction prompts corresponding to different categories of key deviation items. The categories include at least movement trajectory deviation, joint angle deviation, movement rhythm deviation and body center of gravity deviation. The correction guidance content is mapped into instruction data applicable to different feedback channels to generate the multimodal correction guidance signal, wherein the multimodal correction guidance signal includes a graphics rendering signal for driving the visual feedback unit, a speech synthesis signal for driving the auditory feedback unit, and a tactile vibration signal for driving the tactile feedback unit.
[0012] Preferably, the step of filtering and sorting all calculated real-time deviation values according to a preset deviation priority rule to determine the key deviation item that is currently most prioritized for correction includes: For each real-time deviation value, a predefined influence weight coefficient is defined for the real-time motion parameter corresponding to the target leg movement's overall effectiveness. The influence weight coefficient is assigned a value based on the biomechanical importance of the real-time motion parameter in the movement kinetic chain. For each real-time deviation value, a comprehensive priority score is generated by weighting the deviation degree corresponding to the real-time deviation value and the influence weight coefficient. Based on the comprehensive priority score, all real-time deviation values are sorted in descending order to generate a dynamic deviation item priority sequence. The deviation item with the highest ranking from the deviation item priority sequence is selected as the key deviation item.
[0013] Preferably, the step of mapping the correction guidance content into instruction data applicable to different feedback channels to generate the multimodal correction guidance signal includes: The corrective guidance content is semantically parsed to determine the corrective information category conveyed by the corrective guidance content, wherein the corrective information category includes at least spatial trajectory information, movement rhythm information and body posture information; Based on the correction information category and the predefined channel mapping rules, one or more target feedback channels are assigned to the correction guidance content. The target feedback channels include visual channels, auditory channels and tactile channels. According to the data specifications corresponding to the target feedback channel, the correction guidance content is converted into corresponding device-executable signals, wherein the device-executable signals include graphics rendering signals corresponding to the visual channel, audio synthesis signals corresponding to the auditory channel, and tactile driving signals corresponding to the tactile channel; The multimodal correction guidance signal is obtained by combining the graphic rendering signal, the audio synthesis signal and the tactile driving signal generated for the same correction event.
[0014] Secondly, embodiments of the present invention also provide a Taekwondo leg trajectory correction system based on multimodal data processing, comprising: at least one processor, at least one memory, and computer program instructions stored in the memory, wherein when the computer program instructions are executed by the processor, the method of the first aspect described above is implemented.
[0015] In summary, the beneficial effects of the present invention are as follows: This invention provides a method and system for correcting the trajectory of Taekwondo leg movements based on multimodal data processing. The method includes: constructing a standard movement model based on the kinematic characteristics of a preset target Taekwondo leg movement, wherein the standard movement model includes multiple sets of standard parameters of the leg movement in both the visual image dimension and the limb kinematic dimension; and simultaneously acquiring multimodal raw data of the trainee performing the target Taekwondo leg movement through a visual sensing unit deployed in the training environment and an inertial sensing unit worn on the trainee's limbs, wherein the multimodal raw data includes a continuous image sequence acquired by the visual sensing unit and data acquired by the inertial sensing unit. The invention acquires limb joint motion data from the multimodal raw data; based on the received continuous image sequence and the limb joint motion data, it performs time synchronization and data fusion processing on the received data to generate fused posture data of the trainee's limbs in three-dimensional space; it compares the fused posture data with each set of standard parameters in the standard movement model to calculate multi-dimensional posture deviation information, and generates corresponding multimodal correction guidance signals based on the posture deviation information; based on the multimodal correction guidance signals, it drives a multimodal feedback execution unit to provide the trainee with real-time correction feedback adapted to the posture deviation information. This invention does not rely on a single data source, but rather comprehensively utilizes visual sensing units deployed in the environment and inertial sensing units worn on the limbs to simultaneously acquire continuous image sequences and limb joint motion data. This combination effectively overcomes the limitations of single technologies: visual sensors may lose targets during rapid movement or occlusion, while inertial sensors, although capable of capturing motion at high frequencies, suffer from cumulative errors. By fusing these two types of spatiotemporally synchronized data, more reliable and accurate fused posture data is generated, laying a solid foundation for subsequent precise evaluation. Secondly, objectivity and personalization of correction are achieved through intelligent decision-making based on standardized models. It does not rely on the coach's subjective experience for judgment, but rather pre-constructs the kinematic characteristics of excellent movements into a quantifiable standard movement model. This model includes multiple sets of standard parameters in visual images and kinematic dimensions. The generated fused posture data is automatically compared with the standard model to calculate multi-dimensional posture deviation information. This process makes movement evaluation objective, accurate, and measurable. Finally, the multi-channel real-time feedback execution completely solves the problems of real-time performance and multi-sensory coordination. It does not stop at simple error prompts, but generates multimodal correction guidance signals based on posture deviation information and drives the multimodal feedback execution unit to provide real-time feedback to the trainee. This method, which combines visual, auditory, and tactile feedback, simulates a coach's comprehensive guidance, enabling beginners to more efficiently develop correct proprioception and focus their attention on the most critical movement aspects, thereby greatly improving learning efficiency and training immersion. Attached Figure Description
[0016] To more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings used in the embodiments of the present invention will be briefly introduced below. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort, and these are all within the protection scope of the present invention.
[0017] Figure 1 This is a schematic diagram of the overall working process of the Taekwondo leg trajectory correction method based on multimodal data processing in Embodiment 1 of the present invention; Figure 2 This is a schematic diagram of the process for constructing a standard action model in Embodiment 1 of the present invention; Figure 3 This is a schematic diagram of the process for generating multimodal correction guidance signals in Embodiment 1 of the present invention; Figure 4 This is a schematic diagram of the Taekwondo leg trajectory correction system based on multimodal data processing in Embodiment 2 of the present invention. Detailed Implementation
[0018] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. It should be noted that, in this document, relational terms such as "first" and "second" are used merely to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. In the description of the present invention, it should be understood that the terms "center," "upper," "lower," "front," "rear," "left," "right," "vertical," "horizontal," "top," "bottom," "inner," and "outer," etc., indicate orientations or positional relationships based on the orientations or positional relationships shown in the accompanying drawings, and are only for the convenience of describing this application and simplifying the description, and do not indicate or imply that the referred device or element must have a specific orientation, or be constructed and operated in a specific orientation, and therefore should not be construed as a limitation of the present invention. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Unless otherwise specified, an element defined by the phrase "comprising..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element. Where there is no conflict, embodiments of the present invention and the various features thereof can be combined with each other, all of which are within the scope of protection of the present invention.
[0019] Example 1 Please see Figure 1Embodiment 1 of the present invention discloses a method for correcting the trajectory of Taekwondo kicks based on multimodal data processing, the method comprising: Based on the kinematic characteristics of the preset target Taekwondo leg movements, a standard movement model is constructed. The standard movement model includes multiple sets of standard parameters of the leg movements in the visual image dimension and the limb kinematic dimension. Specifically, the standard movement model is established based on a systematic analysis of the kinematic characteristics of specific Taekwondo leg techniques. This model decomposes the movement into continuous key posture segments and defines a multi-dimensional set of standard parameters for each segment. These parameters encompass biomechanical characteristics such as spatial trajectory, joint angles, and movement timing, forming a quantified benchmark for movement standards. Model construction requires combining motion capture data from professional athletes with sports biomechanical analysis, generating representative parameter benchmark values and their tolerance ranges through statistical modeling. This digital modeling method transforms the qualitative standards that rely on subjective experience in traditional teaching into objective and calculable movement paradigms, providing a precise reference system for subsequent automated posture assessment. Its technological advantage lies in achieving quantitative unification of movement standards, effectively solving the problem of inconsistent evaluation standards caused by individual coaching differences.
[0020] By using a visual sensing unit deployed in the training environment and an inertial sensing unit worn on the trainee's limbs, multimodal raw data is simultaneously collected when the trainee performs the target Taekwondo leg technique. The multimodal raw data includes a continuous image sequence acquired by the visual sensing unit and limb joint motion data acquired by the inertial sensing unit. Specifically, multimodal data acquisition is achieved through the collaborative work of heterogeneous sensing units. The visual sensing unit captures continuous image sequences of the trainee's movements using optical imaging principles, while the inertial sensing unit directly measures the angular velocity and acceleration parameters of limb movements via a microelectromechanical system (MEMS). The two types of sensor data are acquired via hardware synchronization signals to ensure consistency of the spatiotemporal reference. The data acquisition process must consider factors such as lighting conditions and electromagnetic interference in the training environment, employing adaptive exposure control and signal filtering techniques to ensure data quality. This multi-source complementary acquisition strategy preserves the spatial information of visual data while also possessing the high-frequency dynamic characteristics of inertial data, providing rich and reliable raw data support for subsequent fusion processing.
[0021] Based on the multimodal raw data, the received continuous image sequence and the limb joint motion data are time-synchronized and data-fused to generate fused posture data of the trainee's limbs in three-dimensional space. Specifically, the data fusion process comprises three core stages: time synchronization, spatial registration, and feature fusion. Time synchronization aligns sensor data with different sampling rates using interpolation algorithms; spatial registration transforms the visual and inertial coordinate systems into a unified world coordinate system; and feature fusion employs an adaptive weighting algorithm to comprehensively process visual keypoint information and inertial motion vectors. The fusion algorithm dynamically evaluates the confidence level of each sensor's data, automatically increasing the weight of inertial data in cases of visual occlusion or motion blur, and prioritizing visual positioning results otherwise. This intelligent fusion mechanism significantly improves the accuracy and robustness of attitude estimation, and the generated six-DOF attitude data accurately reflects the limb's motion state in three-dimensional space.
[0022] The fused posture data is compared with each set of standard parameters in the standard motion model to calculate multi-dimensional posture deviation information, and based on the posture deviation information, a corresponding multimodal correction guidance signal is generated. Specifically, attitude deviation analysis is achieved by comparing real-time attitude data with a standard model across multiple dimensions. The system calculates parameters such as trajectory deviation, angle error, and temporal differences at each joint point and performs a comprehensive evaluation based on preset weighting coefficients. For identified significant deviations, the feedback mapping engine automatically generates targeted correction schemes based on the deviation type. For example, trajectory deviations trigger spatial guidance signals, angle deviations generate attitude correction commands, and temporal deviations generate rhythm control prompts. This hierarchical feedback strategy ensures the accuracy and effectiveness of the correction guidance.
[0023] Based on the multimodal correction guidance signal, the multimodal feedback execution unit is driven to provide the trainee with real-time corrective feedback adapted to the posture deviation information.
[0024] Specifically, the multimodal feedback system achieves coordinated guidance by driving visual, auditory, and tactile sensory channels in parallel. The visual channel presents virtual guidance on movement trajectories and visual cues for deviations; the auditory channel provides voice guidance and rhythmic references; and the tactile channel transmits spatial orientation and force information through vibration feedback. The signals from each channel undergo rigorous time synchronization processing to ensure a high degree of consistency in multisensory stimulation. This immersive feedback mechanism can simultaneously activate multiple sensory pathways in the trainee, significantly improving the learning efficiency of the neuromuscular system and forming a complete corrective loop from perception to execution.
[0025] Preferably, please refer to Figure 2 The construction of a standard movement model based on the kinematic characteristics of a preset target Taekwondo leg technique includes: Obtain multiple preset learning stage sequences corresponding to different skill levels, the learning stage sequences including beginner stage, advancement stage and competition stage; Specifically, the learning phase sequence is divided based on the maturity gradient of the trainee's motor skills, typically including key developmental stages such as the basic movement solidification period, the movement quality optimization period, and the competitive performance improvement period. Each stage corresponds to different training objectives; for example, the basic stage focuses on the correctness of the movement trajectory, the advanced stage emphasizes the speed and fluidity of the movement, and the competitive stage pursues explosive power and stability. The pre-design of the phase sequence needs to combine the laws of motor skill formation with expert teaching experience, and use quantitative indicators to clarify the critical thresholds of each stage. This tiered mechanism provides a structured framework for the subsequent construction of personalized models, ensuring that the training content matches the trainee's actual ability.
[0026] A biomechanical analysis was performed on the target Taekwondo leg technique, and the continuous movement process of the leg technique was analyzed into several key posture segments arranged in sequence. Specifically, biomechanical analysis employs principles of movement anatomy and dynamics to decompose continuous movements into key postural segments with clearly defined functional boundaries. Taking a side kick as an example, it can be analyzed into four segments: knee lift initiation, hip rotation for power generation, shin strike, and leg retraction for return. Each segment corresponds to a specific joint motion sequence, weight transfer pattern, and muscle activation timing. The analysis process requires combining high-speed motion capture and electromyography (EMG) signal analysis to identify the core elements affecting movement efficiency. This step transforms the complex overall movement into quantifiable unit modules, laying the foundation for parametric modeling.
[0027] For each of the key posture segments, a multi-dimensional standard motion parameter set corresponding to the key posture segment is defined, and a corresponding benchmark value and allowable deviation tolerance range are set for each parameter in the standard motion parameter set. Specifically, the multi-dimensional parameter set covers three categories of indicators: spatial, temporal, and dynamic. Spatial parameters include joint angles and limb displacement trajectories; temporal parameters involve segment duration and movement rhythm; and dynamic parameters include angular velocity and ground reaction force. Benchmark values are determined through statistical analysis of movement data collected from elite athletes, while tolerance ranges are dynamically set based on teaching experience and safety requirements. For example, the tolerance for joint angles can be ±10° in the initial learning stage, narrowing to ±3° in the competitive stage. The establishment of this parameter set enables a refined description of movement standards.
[0028] Each stage in the learning stage sequence is associated and mapped with the target configuration of the multi-dimensional standard action parameter set; wherein, the target configuration includes at least the setting of the width of the deviation tolerance range and the setting of the activation priority of different dimension parameters in the multi-dimensional standard action parameter set; Specifically, the target configuration includes a tolerance range adjustment strategy and parameter activation priority rules. In the initial stage, a tolerance is adopted and the focus is on basic spatial parameters. In the advanced stage, the tolerance is gradually tightened and higher-order parameters such as time and dynamics are enabled. The mapping relationship is implemented through decision trees or rule engines. For example, when the trainee enters the advancement stage, the system automatically reduces the trajectory tolerance by 30% and enables the motion velocity parameter. This dynamic mapping mechanism enables the model to adapt.
[0029] In response to the trainee's skill level assessment results, the skill level assessment results are mapped to the corresponding learning stage, and a personalized movement model adapted to the current trainee is generated from the standard movement parameter set as the standard movement model based on the target configuration associated with the learning stage.
[0030] Specifically, the personalized model is generated based on skill assessment results. Assessment data can come from historical training records or specialized tests, and a pattern matching algorithm maps trainees to corresponding learning stages. The system extracts a subset of parameters that meet the criteria from the master parameter set based on the target configuration corresponding to each stage, forming customized evaluation standards. For example, trajectory requirements are relaxed for beginners but safety angles are strictly limited, while for advanced learners, both speed and accuracy are assessed. This model truly realizes the teaching principle of individualized instruction.
[0031] Preferably, the multimodal raw data of the trainee performing the target Taekwondo leg technique is simultaneously collected by a visual sensing unit deployed in the training environment and an inertial sensing unit worn on the trainee's limbs. The multimodal raw data includes a continuous image sequence acquired by the visual sensing unit and limb joint motion data acquired by the inertial sensing unit, including: Before the training of the target Taekwondo leg techniques begins, the internal parameters and external spatial position of the visual sensing unit are calibrated, and the inertial sensing unit is initially zero-biased calibrated. Specifically, before the data acquisition process begins, the visual sensing unit needs to undergo intrinsic parameter calibration to correct lens distortion, and extrinsic parameter calibration to determine its spatial pose in the global coordinate system of the training area. Simultaneously, the inertial sensing unit needs initial zero-bias calibration, estimating and compensating for the inherent biases of the gyroscope and accelerometer through static sampling. This step aims to establish a precise sensor measurement benchmark, eliminate the impact of device errors on data quality, and provide accurate transformation relationships for the spatial alignment and fusion of multi-source data. The calibration process must use dedicated calibration materials or preset calibration actions to ensure that all sensing units operate under a unified measurement standard.
[0032] The ambient light intensity of the preset training site is monitored in real time, and the acquisition parameters of the visual sensing unit are dynamically adjusted based on the ambient light intensity to ensure that the imaging quality of the continuous image sequence meets the preset recognition requirements. Specifically, changes in lighting conditions at the training site directly affect the imaging quality of the visual sensing unit. The system monitors ambient illuminance in real time using a light intensity sensor and dynamically adjusts camera exposure time, gain, and white balance parameters based on a preset illuminance-parameter mapping table. For example, it automatically reduces exposure in strong light to prevent overexposure and increases gain in low light to ensure image signal-to-noise ratio. This mechanism ensures that continuous image sequences with clear outlines and appropriate contrast can be obtained under different lighting conditions, providing stable visual input for subsequent keypoint recognition.
[0033] In response to the signal that the training of the target Taekwondo leg technique has started, a synchronous acquisition command is sent to the visual sensing unit and the inertial sensing unit to acquire the continuous image sequence and the limb joint motion data respectively with the same time base. Specifically, when the system detects that the trainee has entered a ready posture or receives a manual start signal, the central controller sends a synchronization acquisition command based on a global clock to all sensing units. The vision unit captures image sequences at a fixed frame rate, while the inertial unit records triaxial angular velocity and acceleration data at a higher sampling rate. All data packets are embedded with timestamps accurate to the millisecond level, and hardware synchronization circuitry ensures that the visual frames and inertial data packets are strictly aligned on the time axis. This hard synchronization mechanism effectively avoids the random delays that may exist in software triggering, providing a foundation for timing consistency for multimodal data fusion.
[0034] The acquired continuous image sequence and limb joint motion data are preprocessed, and the integrity of the preprocessed multimodal raw data is verified. Specifically, the raw data undergoes a preprocessing pipeline to improve usability. Visual image sequences are first denoised and enhanced, while inertial data is low-pass filtered to eliminate high-frequency noise. After preprocessing, the system performs automated quality checks: checking for motion blur or occlusion in the image sequences and verifying whether the inertial data exceeds its range or exhibits abnormal drift. The verification algorithm identifies abnormal segments by analyzing the statistical distribution of data features, marking or triggering re-acquisition of data segments that do not meet the integrity threshold. This step acts as a gatekeeper for data quality, ensuring the integrity and reliability of the data input to subsequent fusion modules.
[0035] The continuous image sequence that has passed integrity verification is packaged with the limb joint motion data to generate the multimodal raw data stream.
[0036] Specifically, verified visual image sequences and inertial motion data are paired by timestamp and encapsulated into a structured multimodal data stream. The data stream uses a standardized container format, containing a metadata header and synchronization data blocks. The metadata records information such as acquisition time, sensor ID, and calibration parameters; visual frames within data blocks are associated with their corresponding inertial measurement values via indexes. This streaming encapsulation not only ensures the spatiotemporal consistency of the multimodal data but also provides a unified data interface for subsequent processing modules, supporting efficient data access and parsing.
[0037] Preferably, the step of performing time synchronization and data fusion processing on the received continuous image sequence and the limb joint motion data based on the multimodal raw data to generate fused posture data of the trainee's limbs in three-dimensional space includes: Based on the synchronization timestamp carried by the synchronization acquisition command, a temporal correspondence is established between the continuous image sequence and the limb joint motion data. The two-dimensional image space perceived by the visual sensing unit and the three-dimensional motion space measured by the inertial sensing unit are registered in the same world coordinate system to obtain data after unified spatiotemporal reference processing. Specifically, the core of this step is to establish temporal and spatial consistency between visual and inertial data. Time synchronization is based on the hardware timestamp embedded during acquisition, using linear interpolation or polynomial fitting algorithms to align visual frames and inertial data packets with different sampling rates to a unified time axis. Spatial registration is achieved through coordinate system transformation. The two-dimensional pixel coordinates in the visual sensing unit's coordinate system, combined with camera calibration parameters, are transformed to the world coordinate system through perspective transformation; simultaneously, the three-dimensional motion data in the carrier coordinate system measured by the inertial sensing unit is rotated and translated to the same world coordinate system through an initial alignment matrix. This process solves the fundamental problem of inconsistent spatiotemporal reference systems of heterogeneous sensor data, providing operable standardized data for subsequent fusion.
[0038] Visual modal features and inertial modal features are extracted from the data after unified processing of the spatiotemporal reference. The visual modal features include the two-dimensional pixel coordinates of the joints identified in each frame of the image, and the inertial modal features include the angular velocity vector and acceleration vector of the limb. Specifically, feature vectors representing motion characteristics are extracted from spatiotemporally unified data. Visual modal feature extraction utilizes deep learning joint detection algorithms to identify the two-dimensional pixel coordinates of key joints such as the hip, knee, and ankle in each frame of the image, and calculates their relative displacement and velocity features. Inertial modal features are directly extracted from the inertial measurement unit data stream, using triaxial angular velocity and triaxial acceleration vectors, and, if necessary, calculating higher-order motion features such as jerk through differential calculation. The feature extraction process requires normalization to eliminate the influence of dimensions, forming multi-dimensional feature vectors representing the motion state at the same moment.
[0039] The visual modal features and the inertial modal features at the same time point are fused to obtain the fused multimodal features; Specifically, feature fusion employs an uncertainty-based adaptive weighting algorithm. First, the quality score of visual features is evaluated, with reliability quantified by keypoint detection confidence and image sharpness. Simultaneously, the signal-to-noise ratio and drift error of inertial features are assessed. Then, fusion weights are dynamically assigned based on the real-time quality scores of each modality feature: the weight of visual features is increased when their quality is high, and the weight of inertial features is increased when severe motion causes image blurring. Finally, the weighted feature vectors are fused using Kalman filtering, fully utilizing the absolute positional accuracy of vision and the high-frequency dynamic characteristics of inertia to generate optimally estimated multimodal features.
[0040] The fused multimodal features are input into a preset human kinematics model. By solving the model, the three-dimensional spatial posture sequence of the trainee's limbs in the world coordinate system is obtained. Specifically, the fused multimodal features are input into a parameterized human kinematics model. This model defines biomechanical rules such as length constraints on limb segments and joint degrees of freedom restrictions. By solving the inverse kinematics problem, the optimal joint angle configuration that satisfies both the multimodal feature observations and the model constraints is calculated. The solution process typically employs gradient descent or particle swarm optimization algorithms to minimize the residuals between the observed and model predicted values while ensuring real-time performance. The final output is a continuous and smooth sequence of six-DOF joint poses, accurately reconstructing the limb's motion trajectory in three-dimensional space.
[0041] The three-dimensional spatial attitude sequence is smoothed and filtered to obtain the fused attitude data.
[0042] Specifically, the inverted 3D posture sequence undergoes post-processing optimization. Kalman filters or Butterworth low-pass filters are used to eliminate high-frequency noise while preserving realistic motion details. For abnormal jitter that the kinematic model cannot fully constrain, outlier detection and correction are performed using prior knowledge of motion continuity. Filter parameters are adaptively adjusted according to the amplitude of the movement to ensure the smoothness and realism of the posture data during both rapid kicks and slow transitions. The fused posture data after optimization possesses high accuracy and robustness, providing reliable input for subsequent deviation analysis.
[0043] Preferably, the step of fusing the visual modal features and the inertial modal features at the same time point to obtain the fused multimodal features includes: The visual modal features and the inertial modal features at the same time point are combined to construct a multimodal feature vector; Specifically, this step aims to transform heterogeneous sensor data into a unified mathematical expression. The system concatenates visual modal features (such as the two-dimensional pixel coordinate set of key points) extracted at the same time with inertial modal features (such as angular velocity and acceleration vectors) to form a high-dimensional multimodal feature vector. This vector serves as the input to the fusion algorithm, and its dimension is equal to the sum of the dimensions of the visual features and the inertial features. The first part of the vector can be assigned to the visual features, and the second part corresponds to the inertial features, thus establishing a standardized data structure. This vectorized expression facilitates subsequent matrix-based fusion algorithms and forms the basis for numerical computation.
[0044] Based on the confidence level and image sharpness of key point recognition in the image frame associated with the multimodal feature vector, the first uncertainty measure corresponding to the multimodal feature vector is quantified; Specifically, the uncertainty of visual features mainly stems from image quality and limitations of the recognition algorithm. The system quantifies the first uncertainty measure by analyzing the sharpness indicators of image frames (such as gradient magnitude variance) and the confidence scores output by the keypoint recognition algorithm. For example, under conditions of motion blur or insufficient lighting, image sharpness decreases, and the confidence of keypoint localization decreases, significantly increasing the uncertainty of visual features. The quantization process normalizes multiple visual quality indicators into a single scalar value; the larger this value, the lower the reliability of the visual feature at the current moment.
[0045] Based on the predicted values of sensor signal noise level and cumulative error corresponding to the inertial mode features, the second uncertainty measure corresponding to the inertial mode feature vector is quantified. Specifically, the uncertainty of inertial characteristics is mainly caused by sensor noise and integral drift effects. The system quantifies a second uncertainty measure by analyzing the noise variance of the inertial measurement unit signal (which can be estimated through static sampling) and the cumulative error based on motion model predictions. For example, during high-speed motion, the accelerometer signal has a large dynamic range, and the impact of noise is relatively reduced; while in stationary or low-speed phases, noise becomes the main source of error. The quantification model comprehensively considers the time accumulation effect and changes in motion state to dynamically evaluate the reliability of inertial characteristics.
[0046] Based on the quantified first uncertainty measure and second uncertainty measure, corresponding fusion weights are assigned to the visual modal features and the inertial modal features, and the fusion weights are used to perform weighted calculations on the visual modal features and the inertial modal features to obtain the fused multimodal features.
[0047] Specifically, based on the uncertainty measure obtained in the first two steps, the system assigns adaptive fusion weights to visual and inertial features. The weight allocation follows the principle of "the smaller the uncertainty, the greater the weight," typically using the reciprocal of the uncertainty measure for normalization. Then, a weighted average is applied to the visual and inertial features to obtain the fused multimodal features. This adaptive mechanism ensures that visual localization is prioritized when visual conditions are good, while inertial measurement is relied upon when severe motion causes image blurring, thus achieving complementary advantages.
[0048] Preferably, please refer to Figure 3 The step of comparing the fused posture data with each set of standard parameters in the standard motion model to calculate multi-dimensional posture deviation information, and generating corresponding multimodal correction guidance signals based on the posture deviation information, includes: The real-time motion parameters representing each key posture segment in the fused posture data are compared one by one with the corresponding standard parameters in the standard motion model to calculate the real-time deviation value of each real-time motion parameter. Specifically, the core of this step is to finely compare the trainee's real-time performance with the ideal standard. The system extracts real-time motion parameters representing each key postural phase from the fused posture data, such as joint angles, limb end-effector trajectory coordinates, and duration of each movement phase. These real-time parameters are compared one by one with the corresponding preset benchmark values and tolerance ranges in the standard movement model. The comparison algorithm calculates the real-time deviation value of each parameter, which is usually expressed as absolute error or relative error. For example, in a side kick, the target knee joint angle for the shin strike phase is 150 degrees; if the measured value is 140 degrees, the deviation value is -10 degrees. This process achieves a quantitative assessment of movement quality, transforming the abstract concept of "non-standard movement" into specific, quantifiable numerical indicators.
[0049] Based on the preset deviation priority rules, all calculated real-time deviation values are filtered and sorted to determine the key deviation item that is currently most prioritized for correction. Specifically, after obtaining multi-dimensional deviation values, the system does not treat all deviations equally. Instead, it intelligently filters and sorts them according to a preset deviation priority rule. This rule comprehensively considers the severity of the deviation (e.g., exceeding the safety tolerance range), its impact weight on overall movement efficiency (e.g., deviations in core components take precedence over secondary components), and the urgency of correction (e.g., the key objective of the current training phase). Through a weighted scoring algorithm, the system comprehensively evaluates all real-time deviation values and sorts them in descending order to identify the key deviations that currently require the most intervention. For example, in the initial learning stage, maintaining the stability of the body's center of gravity may have a higher priority than pursuing kick height.
[0050] According to the preset mapping relationship, the key deviation items are mapped to the predefined correction strategy library to generate targeted correction guidance content. The correction strategy library is built based on preset knowledge. The mapping relationship defines the best correction prompts corresponding to different categories of key deviation items. The categories include at least movement trajectory deviation, joint angle deviation, movement rhythm deviation and body center of gravity deviation. Specifically, after identifying key deviations, the system maps them to a predefined correction strategy library to generate targeted correction guidance. This strategy library, built upon knowledge of sports biomechanics and teaching experience, establishes a mapping relationship from "deviation category" to "optimal correction prompt." For example, for a detected "trajectory deviation" (such as a side kick not going in a straight line), the strategy library might map a voice prompt like "eject the lower leg in a straight line"; for a "joint angle deviation" (such as insufficient hip rotation), it might map a prompt like "adjust the hip forward." This step transforms numerical deviation information into operational instructions with clear guidance that can be understood and executed by the trainee.
[0051] The correction guidance content is mapped into instruction data applicable to different feedback channels to generate the multimodal correction guidance signal, wherein the multimodal correction guidance signal includes a graphics rendering signal for driving the visual feedback unit, a speech synthesis signal for driving the auditory feedback unit, and a tactile vibration signal for driving the tactile feedback unit.
[0052] Specifically, abstract corrective guidance content is transformed into cross-sensory instruction data that can drive specific hardware devices—that is, synthesized multimodal corrective guidance signals. The system converts the guidance content in parallel into signal formats suitable for different feedback channels, based on the nature of the content. For example, for guidance content related to trajectory correction, it simultaneously generates graphic rendering signals for the visual feedback unit (such as overlaying virtual guide lines on the display), speech synthesis signals for the auditory feedback unit (such as playing "Please kick along a straight line"), and tactile vibration signals for the tactile feedback unit (such as applying directional vibrations to specific muscle groups in the leg). This multisensory collaborative feedback mechanism can more effectively guide trainees to perceive and correct errors.
[0053] Preferably, the step of filtering and sorting all calculated real-time deviation values according to a preset deviation priority rule to determine the key deviation item that is currently most prioritized for correction includes: For each real-time deviation value, a predefined influence weight coefficient is defined for the real-time motion parameter corresponding to the target leg movement's overall effectiveness. The influence weight coefficient is assigned a value based on the biomechanical importance of the real-time motion parameter in the movement kinetic chain. Specifically, the core of this step is to establish a parameter importance evaluation system based on biomechanical principles. The influence weight coefficient is a quantitative indicator used to measure the contribution of a specific motion parameter to overall movement effectiveness. Its assignment is based on the biomechanical importance of the motion parameter in the movement kinetic chain, which needs to be determined through movement biomechanical analysis. For example, in the kinetic chain of a side kick, the rotation angle and speed of the hip joint are the hub of power transmission, and their weight coefficients will be assigned higher values; while the subtle angle of the ankle joint may have a relatively small impact on the final striking effect, and its weight coefficient will be lower. Achieving this process requires combining literature research, expert experience, and experimental data to assign a relative weight to each parameter. The purpose of this step is to transform the coach's implicit experience (such as "hip rotation is more important") into explicit rules that the system can calculate, providing a scientific basis for subsequent intelligent decision-making, thereby ensuring that the system's correction focus remains on the key links affecting movement quality.
[0054] For each real-time deviation value, a comprehensive priority score is generated by weighting the deviation degree corresponding to the real-time deviation value and the influence weight coefficient. Specifically, after obtaining the weighting coefficients and real-time deviation values, the system generates a comprehensive priority score through weighted calculation. This score is the core indicator for quantifying the urgency of deviation correction. The calculation process is not simply multiplying the deviation value by the weights, but rather considering the non-linear impact of the degree of deviation. For example, a deviation that slightly deviates from the target on a key parameter may have a higher priority than a deviation that significantly deviates on a minor parameter. In implementation, a weighted summation or a more complex utility function model can be used to combine the degree of deviation and the impact weighting coefficients to form a unified score. The purpose of this step is to integrate multi-dimensional, heterogeneous deviation information (one concerning "how much was wrong," and the other concerning "where and how important the mistake was") into a comparable single-dimensional indicator, thereby solving the core decision-making problem in teaching: "which error to correct first," and giving the system's corrective recommendations a clear priority.
[0055] Based on the comprehensive priority score, all real-time deviation values are sorted in descending order to generate a dynamic deviation item priority sequence. The deviation item with the highest ranking from the deviation item priority sequence is selected as the key deviation item.
[0056] Specifically, the system sorts and makes decisions based on comprehensive priority scores. It sorts all real-time deviation values in descending order of their comprehensive priority scores, generating a dynamic sequence of deviation priority. This sequence is updated in real time and accurately reflects the relative severity of all deviations in the current action. The system then selects the highest-ranking deviation from this sequence, identifying it as the most critical deviation to be corrected. This process requires an efficient sorting algorithm to ensure rapid output under real-time requirements. The purpose of this step is to transform the calculated scores into explicit execution instructions, providing precise input for subsequent correction strategy mapping. Its beneficial effects include optimized allocation of correction resources, avoiding information overload and attention distraction caused by simultaneously correcting multiple errors in trainees, aligning with the effective teaching principle of "correcting one critical error at a time," and greatly improving the efficiency and effectiveness of correction.
[0057] Preferably, the step of mapping the correction guidance content into instruction data applicable to different feedback channels to generate the multimodal correction guidance signal includes: The corrective guidance content is semantically parsed to determine the corrective information category conveyed by the corrective guidance content, wherein the corrective information category includes at least spatial trajectory information, movement rhythm information and body posture information; Specifically, the core of this step is to structurally understand the abstract corrective instructions and categorize them into specific information categories. The semantic parsing process requires analyzing the keywords and intentions within the instructions. For example, if the instruction is "Please kick your lower leg in a straight line," its core is describing the movement path of the limb's end, and therefore it is categorized as "spatial trajectory information." If the instruction is "Increase the kicking speed," it involves the temporal characteristics of the movement and is categorized as "movement rhythm information." "Keep your upper body stable" explicitly points to the relative positional relationships of the joints, belonging to "body posture information." This classification process forms the basis for subsequent selection of the optimal feedback channel. Its purpose is to transform ambiguous language instructions into machine-processable information types with clear semantic labels, thus laying the foundation for accurate channel mapping.
[0058] Based on the correction information category and the predefined channel mapping rules, one or more target feedback channels are assigned to the correction guidance content. The target feedback channels include visual channels, auditory channels and tactile channels. Specifically, after clarifying the category of corrective information, the system assigns one or more target feedback channels most suitable for conveying that type of information, based on predefined channel mapping rules. These mapping rules are based on the matching degree between the characteristics of different sensory channels and the information type. For example, spatial trajectory information has strong spatial attributes and highly matches the presentation capabilities of the visual channel; therefore, it is preferentially mapped to the visual channel, graphically displaying trajectory deviations and the ideal path. Movement rhythm information has time-series characteristics, matching the rhythmic transmission advantage of the auditory channel; therefore, it is often mapped to the auditory channel, providing cues through sound effects of beat or rate changes. Body posture information is directly related to tactile perception, especially information about joint angles and body balance; mapping it to the tactile channel provides intuitive proprioceptive feedback through vibration patterns of specific body parts. This step achieves intelligent decision-making from "what to correct" to "which sense to use for correction."
[0059] According to the data specifications corresponding to the target feedback channel, the correction guidance content is converted into corresponding device-executable signals, wherein the device-executable signals include graphics rendering signals corresponding to the visual channel, audio synthesis signals corresponding to the auditory channel, and tactile driving signals corresponding to the tactile channel; Specifically, the corrective instructions mapped to a specific channel are converted into physical signals that the hardware device for that channel can directly recognize and execute. This is an encoding and translation process. For the visual channel, the system needs to convert spatial trajectory information into instructions that the graphics rendering engine can understand, such as generating a virtual guide line superimposed on the real-time screen, highlighting error areas with arrows, or creating 3D animation demonstrations. For the auditory channel, rhythms or simple instruction text need to be converted into audio synthesis signals, which could be specific cue tones, speech synthesis segments, or audio streams with a specific rhythm. For the tactile channel, adjustments to body posture (such as "hip push") need to be encoded into tactile actuation signals to control the vibration pattern, intensity, and duration at specific locations (such as a vibrator worn on the hip) to simulate the feeling of being pushed or prompted. The key to this step is adhering to the data specifications of each hardware interface to ensure that the generated signals can be executed accurately.
[0060] The multimodal correction guidance signal is obtained by combining the graphic rendering signal, the audio synthesis signal and the tactile driving signal generated for the same correction event.
[0061] Specifically, the final step is not simply packaging signals from different channels, but ensuring that the graphical rendering signal, audio synthesis signal, and tactile actuation signal generated for the same corrective event are highly synchronized in time and coordinated in space, combining into a unified multimodal corrective guidance signal. The system needs to bind a unified timestamp to this signal data packet to ensure that visual cues, audio cues, and tactile vibrations act on the trainee at the same moment, avoiding cognitive confusion caused by signal delays and asynchrony. For example, when prompting "hip rotation," the rotating arrow on the screen, the "hip rotation" voice in the ear, and the instantaneous vibration in the hip should occur simultaneously. This precise combination creates an immersive corrective experience, greatly enhancing the efficiency and depth of guidance information delivery through multi-sensory synergistic stimulation, helping trainees quickly form correct muscle memory and proprioception.
[0062] Through these four meticulous steps, the system successfully transforms the correction strategy derived from data analysis into an intelligent feedback mechanism that can efficiently guide trainees and closely aligns with human multi-sensory cognitive habits.
[0063] Example 2 In addition, combined Figure 1 The Taekwondo leg trajectory correction method based on multimodal data processing described in Embodiment 1 of the present invention can be implemented by an electronic device. Figure 4 A schematic diagram of the hardware structure of the electronic device provided in Embodiment 3 of the present invention is shown.
[0064] Electronic devices may include processors and memory storing computer program instructions.
[0065] Specifically, the processor may include a central processing unit (CPU), an application-specific integrated circuit (ASIC), or one or more integrated circuits that can be configured to implement embodiments of the present invention.
[0066] The memory may include a large-capacity storage device for data or instructions. For example, and not limitingly, the memory may include a hard disk drive (HDD), a floppy disk drive, flash memory, an optical disk drive, a magneto-optical disk drive, magnetic tape, or a Universal Serial Bus (USB) drive, or a combination of two or more of these. Where appropriate, the memory may include removable or non-removable (or fixed) media. Where appropriate, the memory may be internal or external to a data processing device. In a particular embodiment, the memory is a non-volatile solid-state memory. In a particular embodiment, the memory includes a read-only memory (ROM). Where appropriate, the ROM may be a mask-programmed ROM, a programmable ROM (PROM), an erasable PROM (EPROM), an electrically erasable PROM (EEPROM), an electrically rewritable ROM (EAROM), or flash memory, or a combination of two or more of these.
[0067] The processor reads and executes computer program instructions stored in the memory to implement any of the Taekwondo kick trajectory correction methods based on multimodal data processing in the above embodiments.
[0068] In one example, the electronic device may also include a communication interface and a bus. For example, Figure 4 As shown, the processor, memory, and communication interface are connected via a bus and communicate with each other.
[0069] The communication interface is mainly used to enable communication between various modules, devices, units and / or equipment in the embodiments of the present invention.
[0070] A bus, including hardware, software, or both, couples components of the device together. For example, and not limitingly, a bus may include an Accelerated Graphics Port (AGP) or other graphics bus, an Enhanced Industry Standard Architecture (EISA) bus, a Front Side Bus (FSB), HyperTransport (HT) interconnect, an Industry Standard Architecture (ISA) bus, an Infinite Bandwidth Interconnect, a Low Pin Count (LPC) bus, a memory bus, a Microchannel Architecture (MCA) bus, a Peripheral Component Interconnect (PCI) bus, a PCI-Express (PCI-X) bus, a Serial Advanced Technology Attachment (SATA) bus, a Video Electronics Standards Association Local (VLB) bus, or other suitable buses, or combinations of two or more of these. Where appropriate, a bus may include one or more buses. While specific buses are described and illustrated in embodiments of the invention, the invention contemplates any suitable bus or interconnect.
[0071] In summary, the embodiments of the present invention provide a method and system for correcting the trajectory of Taekwondo kicks based on multimodal data processing.
[0072] It should be clarified that the present invention is not limited to the specific configurations and processes described above and shown in the figures. For the sake of brevity, detailed descriptions of known methods are omitted here. In the above embodiments, several specific steps are described and shown as examples. However, the method process of the present invention is not limited to the specific steps described and shown. Those skilled in the art can make various changes, modifications, and additions, or change the order of steps, after understanding the spirit of the present invention.
[0073] The functional blocks shown in the above-described structural diagram can be implemented as hardware, software, firmware, or a combination thereof. When implemented in hardware, they can be, for example, electronic circuits, application-specific integrated circuits (ASICs), appropriate firmware, plug-ins, function cards, etc. When implemented in software, the elements of this invention are programs or code segments used to perform the required tasks. The programs or code segments can be stored on a machine-readable medium or transmitted over a transmission medium or communication link via data signals carried in a carrier wave. "Machine-readable medium" can include any medium capable of storing or transmitting information. Examples of machine-readable media include electronic circuits, semiconductor memory devices, ROM, flash memory, erasable ROM (EROM), floppy disks, CD-ROMs, optical disks, hard disks, fiber optic media, radio frequency (RF) links, etc. Code segments can be downloaded via computer networks such as the Internet, intranets, etc.
[0074] The user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, data stored, data displayed, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties. Furthermore, the collection, use and processing of the relevant data shall comply with the relevant laws, regulations and standards of the relevant locality, and corresponding operation entry points shall be provided for the user to choose to authorize or refuse.
[0075] It should also be noted that the exemplary embodiments mentioned in this invention describe methods or systems based on a series of steps or apparatus. However, this invention is not limited to the order of the steps described above; that is, the steps can be performed in the order mentioned in the embodiments, or in a different order, or several steps can be performed simultaneously.
[0076] The above description is merely a specific embodiment of the present invention. Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the specific working processes of the systems, modules, and units described above can be referred to the corresponding processes in the foregoing method embodiments, and will not be repeated here. It should be understood that the protection scope of the present invention is not limited thereto. Any person skilled in the art can easily conceive of various equivalent modifications or substitutions within the technical scope disclosed in the present invention, and these modifications or substitutions should all be covered within the protection scope of the present invention.
Claims
1. A taekwondo kicking trajectory correction method based on multi-modal data processing, characterized in that, The method comprises: based on the preset kinematic characteristics of the target taekwondo leg action, a standard action model is constructed, the standard action model comprises a plurality of sets of standard parameters of the leg action in the visual image dimension and the limb kinematics dimension; through the visual sensing unit deployed in the training environment and the inertial sensing unit worn on the limbs of the trainer, the multi-modal raw data of the trainer performing the target taekwondo leg action is synchronously collected, the multi-modal raw data comprises a continuous image sequence acquired by the visual sensing unit and limb joint motion data acquired by the inertial sensing unit; according to the multi-modal raw data, the received continuous image sequence and the limb joint motion data are subjected to time synchronization and data fusion processing, and fused posture data of the trainer's limbs in three-dimensional space is generated; the fused posture data is compared with each set of standard parameters in the standard action model, multi-dimensional posture deviation information is calculated, and corresponding multi-modal correction guidance signals are generated based on the posture deviation information; according to the multi-modal correction guidance signals, a multi-modal feedback execution unit is driven to provide real-time correction feedback to the trainer that is adapted to the posture deviation information.
2. The taekwondo leg-out trajectory correction method based on multi-modal data processing according to claim 1, wherein, The method comprises: a plurality of preset learning stage sequences corresponding to different skill levels are acquired, the learning stage sequences comprise a beginner stage, a promotion stage and a competition stage; biomechanical analysis is performed on the target taekwondo leg action, and the continuous motion process of the leg action is analyzed into a plurality of sequentially arranged key posture links; for each key posture link, a set of multi-dimensional standard action parameters corresponding to the key posture link is defined, and a corresponding reference value and an allowable deviation tolerance range are set for each parameter in the set of standard action parameters; each stage in the learning stage sequence is associated and mapped with a target configuration of the set of multi-dimensional standard action parameters; wherein the target configuration at least comprises a width setting of the deviation tolerance range and an enablement priority setting of different dimension parameters in the set of multi-dimensional standard action parameters; in response to the skill level evaluation result of the trainer, the skill level evaluation result is mapped to the corresponding learning stage, and according to the target configuration associated with the learning stage, an individualized action model adapted to the current trainer is generated from the set of standard action parameters as the standard action model.
3. The taekwondo leg-out trajectory correction method based on multi-modal data processing according to claim 1, characterized in that, The method comprises: before the target taekwondo leg action training starts, the internal parameters and the external space position of the visual sensing unit are calibrated, and the inertial sensing unit is subjected to initial zero offset calibration; Real-time monitoring of the preset training field environment light intensity, based on the ambient light intensity, dynamically adjust the acquisition parameters of the visual sensing unit, to ensure that the imaging quality of the continuous image sequence meet the pre-set identification requirements; In response to the signal of the start of the target taekwondo leg action training, send a synchronous acquisition instruction to the visual sensing unit and the inertial sensing unit, to obtain the continuous image sequence and the limb joint motion data at the same time reference; Pretreatment of the collected continuous image sequence and limb joint motion data, and integrity check of the pretreated multi-modal raw data; The continuous image sequence and the limb joint motion data that pass the integrity check are packaged to generate the multi-modal raw data stream.
4. The taekwondo leg-out trajectory correction method based on multi-modal data processing according to claim 3, characterized in that, The time synchronization and data fusion processing of the received continuous image sequence and limb joint motion data based on the multi-modal raw data includes: Based on the synchronization timestamp carried by the synchronous acquisition instruction, the time corresponding relationship between the continuous image sequence and the limb joint motion data is established, and the two-dimensional image space perceived by the visual sensing unit and the three-dimensional motion space measured by the inertial sensing unit are registered to the same world coordinate system to obtain the data after space-time reference unified processing; From the space-time reference unified processing data, visual modal features and inertial modal features are extracted respectively, the visual modal features include the recognized joint node two-dimensional pixel coordinates in each frame of image, and the inertial modal features include the angular velocity vector and acceleration vector of the limb; Fuse the visual modal features and the inertial modal features at the same time node to obtain the fused multi-modal features; Input the fused multi-modal features into a preset human kinematics model, and inversely solve the three-dimensional space posture sequence of the trainer's limb in the world coordinate system through model solving; Smooth filtering processing is performed on the three-dimensional space posture sequence to obtain the fused posture data.
5. The taekwondo leg-out trajectory correction method based on multi-modal data processing according to claim 4, wherein, The fusion of the visual modal features and the inertial modal features at the same time node to obtain the fused multi-modal features includes: Combine the visual modal features and the inertial modal features at the same time node to form a multi-modal feature vector; Based on the confidence of the joint node recognition in the image frame associated with the multi-modal feature vector and the image clarity, the first uncertainty measure corresponding to the multi-modal feature vector is quantified; Based on the predicted value of the sensor signal noise level and the cumulative error of the inertial modal features, the second uncertainty measure corresponding to the inertial modal feature vector is quantified; According to the first uncertainty measure and the second uncertainty measure quantified, the visual modal features and the inertial modal features are assigned with corresponding fusion weights, and the visual modal features and the inertial modal features are calculated according to the fusion weights to obtain the fused multi-modal features.
6. The taekwondo leg-out trajectory correction method based on multi-modal data processing according to any one of claims 1-5, characterized in that, The comparing the fused post-fusion attitude data with each set of standard parameters in the standard action model, calculating multi-dimensional attitude deviation information, and generating corresponding multi-modal correction guidance signals based on the attitude deviation information comprises: The real-time motion parameters representing each key attitude link in the fused post-fusion attitude data are compared one by one with the corresponding standard parameters in the standard action model, and the real-time deviation value of each real-time motion parameter is calculated. According to the preset deviation priority rule, all the calculated real-time deviation values are filtered and sorted to determine the key deviation item that needs to be corrected most urgently. According to a preset mapping relationship, the key deviation item is mapped to a predefined correction strategy library to generate targeted correction guidance content. The correction strategy library is constructed based on preset knowledge, and the mapping relationship defines the best correction prompt corresponding to different categories of key deviation items. The categories include at least motion trajectory deviation, joint angle deviation, action rhythm deviation, and body center of gravity deviation. The correction guidance content is mapped to instruction data suitable for different feedback channels to generate the multi-modal correction guidance signals. The multi-modal correction guidance signals include graphical rendering signals for driving a visual feedback unit, speech synthesis signals for driving an auditory feedback unit, and tactile vibration signals for driving a tactile feedback unit.
7. The taekwondo leg-out trajectory correction method based on multi-modal data processing according to claim 6, wherein, The filtering and sorting of all the calculated real-time deviation values according to the preset deviation priority rule to determine the key deviation item that needs to be corrected most urgently comprises: For each real-time motion parameter corresponding to a real-time deviation value, the influence weight coefficient of the real-time motion parameter on the overall performance of the target leg method action is predefined. The influence weight coefficient is assigned based on the biomechanical importance of the real-time motion parameter in the action power chain. For each real-time deviation value, the deviation degree corresponding to the real-time deviation value and the influence weight coefficient are integrated to generate a comprehensive priority score through weighted calculation. According to the comprehensive priority score, all the real-time deviation values are sorted in descending order to generate a dynamic deviation item priority sequence. The highest ranked deviation item in the deviation item priority sequence is selected as the key deviation item.
8. The taekwondo leg-out trajectory correction method based on multi-modal data processing according to claim 6, wherein, The mapping of the correction guidance content to instruction data suitable for different feedback channels to generate the multi-modal correction guidance signals comprises: Semantic analysis of the correction guidance content is performed to determine the correction information categories conveyed by the correction guidance content. The correction information categories include at least spatial trajectory information, action rhythm information, and body posture information. According to the correction information categories and a predefined channel mapping rule, one or more target feedback channels are assigned to the correction guidance content. The target feedback channels include visual, auditory, and tactile channels. According to the data specifications corresponding to the target feedback channels, the correction guidance content is converted into corresponding device executable signals. The device executable signals include graphical rendering signals for the visual channel, audio synthesis signals for the auditory channel, and tactile driving signals for the tactile channel. The graphical rendering signal, the audio synthesis signal and the haptic drive signal generated for the same correction event are combined to obtain the multi-modal correction guidance signal. 9.A taekwondo leg-out trajectory correction system based on multi-modal data processing, characterized by, Comprise: at least one processor, at least one memory, and computer program instructions stored in the memory that, when executed by the processor, implement the method of any one of claims 1-8.