Standardized real-time monitoring and correction method and system for martial arts routine teaching

By using a two-stage temporal alignment algorithm based on image semantic attention and a deep learning model, the problem of misaligned action frames in martial arts teaching has been solved, achieving precise temporal alignment and personalized error correction of martial arts movements, thus improving the accuracy and efficiency of teaching error correction.

CN122290211APending Publication Date: 2026-06-26WUHAN SPORTS UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
WUHAN SPORTS UNIV
Filing Date
2026-04-20
Publication Date
2026-06-26

Smart Images

  • Figure CN122290211A_ABST
    Figure CN122290211A_ABST
Patent Text Reader

Abstract

This invention discloses a standardized real-time monitoring and error correction method and system for martial arts routine teaching, relating to the field of image recognition technology. The method includes: S1. Acquiring information on students' learning stages and training objectives, retrieving corresponding standard movement image models and image feature parameters from a multi-dimensional martial arts standard movement image database, and dynamically adjusting the weights of evaluation indicators based on image recognition; S2. Acquiring movement image data through a multi-modal image acquisition array, and after image preprocessing, using a two-stage temporal alignment algorithm based on image semantic attention to achieve accurate temporal alignment between the student's movement image sequence and the standard movement image model; S3. Calculating movement deviations based on image feature matching, extracting error precursor image features using a deep learning image recognition model, and triggering early warnings; S4. Outputting multi-modal error correction feedback based on image overlay, pushing a group common error image analysis report to the teacher's end, generating and dynamically adjusting personalized correction schemes based on the deviation diagnosis results of image recognition, forming a closed-loop error correction.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of image recognition technology, and in particular to a method and system for real-time monitoring and error correction of standardized martial arts routine teaching. Background Technology

[0002] Traditional algorithms are mostly designed based on fixed frame rates and uniform motion scenarios. Their matching logic is rigid and lacks the ability to adapt to changes in motion speed. However, martial arts movements have distinct characteristics of fast power generation, slow recovery, and alternating movement and stillness. The power generation moment has a large amplitude and high intra-frame information density, while the recovery phase is smooth with subtle inter-frame changes. When traditional algorithms perform frame matching at a uniform time interval, misalignment between the student's action frames and the standard action frames is very likely to occur. This causes subsequent calculations of joint spatial position deviations and posture contour similarity analysis to lose a precise temporal reference, directly leading to distorted motion error diagnosis results. This makes it difficult to provide a reliable basis for error correction in martial arts teaching. Therefore, it is necessary to propose a standardized real-time monitoring and error correction method and system for martial arts routine teaching. Summary of the Invention

[0003] The purpose of this invention is to address the shortcomings of existing technologies where student action frames and standard action frames are misaligned, resulting in a lack of accurate temporal reference for subsequent calculations of joint spatial position deviations and posture contour similarity analysis. This directly leads to distorted action error diagnosis results and makes it difficult to provide a reliable basis for error correction in martial arts teaching. Therefore, this invention proposes a standardized real-time monitoring and error correction method and system for martial arts routine teaching.

[0004] To achieve the above objectives, the present invention adopts the following technical solution: Standardized real-time monitoring and error correction methods for martial arts routine teaching include: S1. Obtain information on the student's learning stage and training objectives, retrieve the corresponding standard movement image model and image feature parameters from the martial arts multidimensional standard movement image database, and dynamically adjust the weight of the evaluation indicators based on image recognition. S2. Motion image data is acquired through a multimodal image acquisition array. After image preprocessing, a two-stage temporal alignment algorithm based on image semantic attention is used to achieve accurate temporal alignment between the trainee's motion image sequence and the standard motion image model. S3. Calculate the action deviation based on image feature matching, classify the error type through the image recognition technology, distinguish the deviation attribute by combining historical action image data, and use a deep learning image recognition model to extract the error precursor image features and trigger an early warning. S4. Output multimodal error correction feedback based on image overlay, push the group common error image analysis report to the teacher's end, generate the deviation diagnosis results based on image recognition and dynamically adjust the personalized correction scheme to form a closed-loop error correction.

[0005] The above technical solution further includes: The training objectives include at least one of the following: martial arts exam preparation, martial arts grading certification, intangible cultural heritage martial arts inheritance, and professional competitive training. The image recognition-based evaluation metrics specifically include three categories: image key point matching degree, action posture image similarity, and action temporal image consistency. The weights of the three evaluation metrics are dynamically adjusted according to the student's learning stage. The initial learning stage focuses on the weight of the image joint matching degree and the similarity of the action posture image, while the advanced learning stage focuses on the weight of the consistency of the action temporal image and the accuracy of the posture detail image.

[0006] The multimodal image acquisition array consists of multiple synchronous high-speed industrial cameras with a sampling frame rate of ≥120fps and image anti-motion blur processing function. It is used to acquire multi-angle continuous image sequences of trainees' movements and extract two-dimensional or three-dimensional joint coordinates and posture contour features through human posture image recognition algorithms. The image preprocessing includes image denoising, image enhancement, image scaling, and region of interest extraction to ensure the clarity of human motion regions in the image, provide high-quality data support for subsequent image feature extraction and matching, and ensure accurate synchronization between motion image data and standard motion image models.

[0007] The two-stage temporal alignment algorithm based on image semantic attention specifically includes: The first stage is global image semantic alignment, which extracts global pose image features of the trainee action image sequence and the standard action image model through a convolutional neural network, calculates the image semantic similarity, and achieves coarse-grained temporal alignment. The second stage involves precise alignment of core node images. Based on the attention mechanism, the image feature vectors of key joint points in the core movements of martial arts are focused. The image matching accuracy is optimized through adaptive dynamic weight iteration to achieve fine-grained temporal calibration. The dual-stage temporal alignment algorithm can automatically adapt to the non-uniform speed characteristics of martial arts movements. Through image frame synchronization calibration, it ensures that the temporal consistency error of the temporal alignment is ≤5ms, and the alignment accuracy is not affected by changes in movement speed.

[0008] The deep learning image recognition model is trained using a self-supervised contrastive learning strategy combined with a spatial attention mechanism. The spatial attention mechanism accurately focuses on the image feature regions of key segments in the transmission of force in martial arts movements. The training dataset includes the standard movement image sequence, the keyframe image features, and the student's historical training error movement image dataset from the martial arts multidimensional standard movement image database. The deep learning image recognition model can capture error precursor image features through feature changes in consecutive image frames and trigger warning signals.

[0009] The personalized correction training program is built based on the deviation diagnosis results of image recognition and includes three core modules: image-guided basic morphological correction module, image temporal comparison-based strength enhancement module, and image frame synchronization-based rhythm control module. The basic morphology correction module identifies joint point matching deviations based on image recognition and designs stretching and shaping training guided by image overlay. The force enhancement module designs specific training for force transmission based on feature difference analysis of the key frame images of the action. The rhythm control module designs beat-following training by comparing the frame synchronization of the trainee's motion image sequence with the standard motion image sequence; The personalized correction training program can dynamically adjust the training intensity and duration based on the training effect feedback from image recognition.

[0010] The AR visualization overlay display in the multimodal error correction feedback is implemented based on image recognition registration technology, specifically including: The image recognition registration algorithm accurately overlays the student's real-time motion image with the standard motion image model on the same screen. The image recognition technology is used to extract the coordinates of key joint points and posture contours of the two. After calculating the degree of image feature deviation, the image is visualized and labeled to intuitively present the position, angle and range of the motion deviation.

[0011] A standardized real-time monitoring and error correction system for martial arts routine teaching that implements the method includes an image benchmark matching module, an image acquisition and alignment module, an image recognition deviation diagnosis and prediction module, an image feedback intervention module, and an interaction module.

[0012] The multidimensional standard martial arts movement image database is integrated into the image benchmark matching module. It stores the three-dimensional dynamic image sequence of multiple core martial arts movements, the key frame image features, the joint point image coordinates and other quantitative parameters, and supports fast retrieval and matching based on image features.

[0013] The image acquisition and alignment module includes an edge image calculation unit, which is responsible for real-time preprocessing of the motion image data, pose reconstruction based on the image recognition, and time alignment calculation, ensuring low latency in image recognition and error correction feedback. The edge image computing unit works in collaboration with the cloud server, which is responsible for storing the student's historical action image data, training and updating the deep learning image recognition model, and performing big data analysis of group error images. The interactive module includes a student-side AR image display device, a teacher-side monitoring screen, and an image data synchronization unit. The image data synchronization unit can synchronize the student's training image sequence, the image recognition error diagnosis report, and the image data of the correction plan execution progress to the cloud server in real time.

[0014] The present invention has the following beneficial effects: In this invention, a two-stage temporal alignment algorithm based on image semantic attention is adopted. It completes coarse-grained global image semantic alignment through a convolutional neural network, and then focuses on the feature vectors of core joints by relying on the attention mechanism. Combined with adaptive dynamic weight iterative optimization, fine-grained calibration is achieved. This ensures that the temporal alignment time consistency error is ≤5ms, and the alignment accuracy is not affected by changes in motion speed. It completely avoids the frame misalignment problem caused by the uniform frame interval matching of traditional algorithms, and provides accurate temporal reference for subsequent calculation of joint spatial position deviation and posture contour similarity analysis. Attached Figure Description

[0015] Figure 1 This is a schematic diagram of the overall process of the standardized real-time monitoring and error correction method and system for martial arts routine teaching proposed in this invention. Detailed Implementation

[0016] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0017] Example

[0018] like Figure 1 As shown, the standardized real-time monitoring and error correction method for martial arts routine teaching proposed in this invention includes: S1. Obtain information on the student's learning stage and training objectives, retrieve the corresponding standard movement image model and image feature parameters from the martial arts multidimensional standard movement image database, and dynamically adjust the weight of the evaluation indicators based on image recognition. S2. Motion image data is acquired through a multimodal image acquisition array. After image preprocessing, a two-stage temporal alignment algorithm based on image semantic attention is used to achieve accurate temporal alignment between the trainee's motion image sequence and the standard motion image model. S3. Calculate action deviation based on image feature matching, classify error types through image recognition technology, distinguish deviation attributes by combining historical action image data, and use a deep learning image recognition model to extract error precursor image features and trigger an early warning. S4. Output multimodal error correction feedback based on image overlay, push the group common error image analysis report to the teacher's end, generate and dynamically adjust personalized correction schemes based on the deviation diagnosis results of image recognition, and form a closed-loop error correction.

[0019] Furthermore, before conducting martial arts movement monitoring and correction, two key pieces of information are collected: the student's current learning stage (e.g., beginner level, advanced level) and training objectives (e.g., martial arts exam preparation, rank certification, intangible cultural heritage transmission). Based on this, standard movement image models and corresponding image feature parameters matching the stage and objectives are retrieved from the multi-dimensional standard movement image database of martial arts. At the same time, the weight ratio of various evaluation indicators based on image recognition technology is flexibly adjusted according to the student's learning stage, so as to establish a standardized reference and evaluation system that meets the actual needs of the students for subsequent movement collection, comparison and correction. Among them, the martial arts multidimensional standard movement image database stores three-dimensional dynamic image sequences of martial arts movements, key frame image features, joint point image coordinates and other quantitative parameters. It can be classified and archived according to the student's learning stage and training objectives, and supports fast retrieval and matching based on image features. It can provide standardized references for subsequent image recognition, temporal alignment, deviation diagnosis and other processes.

[0020] First, relying on a multimodal image acquisition array composed of multiple synchronous high-speed industrial cameras, a multi-angle continuous image sequence of the student's martial arts movements is acquired. Simultaneously, the two-dimensional or three-dimensional joint coordinates and posture contour features of the human body in the image are extracted through a human posture image recognition algorithm. Preprocessing operations are performed on the acquired motion image data, including image denoising, image enhancement, image scaling, and region of interest extraction, to filter out invalid interference information and ensure the clarity of the human motion area in the image, providing high-quality data support for subsequent processing. A two-stage temporal alignment algorithm based on image semantic attention is adopted. First, the global pose image features of the student's action image sequence and the standard action image model are extracted by the convolutional neural network. The semantic similarity of the images is calculated to achieve coarse-grained temporal alignment. Then, based on the attention mechanism, the image feature vectors of key joint points of the core martial arts movement are focused. The image matching accuracy is optimized by adaptive dynamic weight iteration to complete fine-grained temporal calibration. Finally, the precise temporal alignment of the student's action image sequence and the standard action image model is achieved. Moreover, the algorithm can automatically adapt to the non-uniform speed characteristics of martial arts movements and ensure that the temporal consistency error of the temporal alignment is ≤5ms.

[0021] First, the temporally aligned sequence of trainee motion images is matched with the standard motion image model for image feature matching. The specific deviation between the two is calculated for core feature dimensions such as key joint coordinates, posture contours, and motion timing. Then, through image recognition technology, errors are classified into different types based on the characteristics of the deviation, such as deviation in movement form, deviation in timing and rhythm, and deviation related to force transmission. Subsequently, the trainee's historical motion image data was retrieved for comparison and analysis to distinguish whether the current deviation was an accidental random deviation or a long-standing habitual deviation. A deep learning image recognition model trained with a self-supervised contrastive learning strategy combined with spatial attention mechanism was used to capture error precursor image features such as joint offset trends and posture distortion precursors from continuous trainee motion image frames. Once a feature change that meets the warning conditions is identified, a warning signal is immediately triggered to provide a preliminary basis for subsequent error correction intervention. First, multimodal error correction feedback is output, centered on image overlay. Image recognition registration technology precisely overlays real-time motion images of trainees with standard motion image models on the same screen, marking the location and degree of deviations in key joints and posture contours, allowing trainees to intuitively perceive motion problems. Simultaneously, motion deviation data from all trainees is collected, common group errors are filtered and extracted, and a visualized image analysis report is generated and pushed to the teacher for targeted, intensive guidance. Then, based on the motion deviation diagnosis results from the initial image recognition, a personalized correction plan is generated, including image-guided basic form correction, image-time sequence comparison-based power enhancement, and image-frame synchronized rhythm control modules. Furthermore, the training intensity and duration of the plan are dynamically adjusted based on the motion improvement effects of image recognition feedback during subsequent training. Finally, a closed-loop error correction process is formed: "motion acquisition—deviation diagnosis—feedback correction—plan optimization—retraining," continuously improving the standardized error correction effect of martial arts routine teaching.

[0022] Training objectives include at least one of the following: martial arts exam preparation, martial arts grading certification, intangible cultural heritage martial arts inheritance, and professional competitive training. The evaluation metrics based on image recognition specifically include three categories: image key point matching degree, action posture image similarity, and action temporal image consistency. The weights of the three evaluation metrics are dynamically adjusted according to the learning stage of the learners. The initial learning stage focuses on the weight of image key point matching degree and action pose image similarity, while the advanced learning stage focuses on the weight of action temporal image consistency and pose detail image accuracy.

[0023] Furthermore, the specific scope of the student training goal information that needs to be obtained in step S1 is clarified. That is, the student's training goal must be determined from at least one of the four directions: martial arts exam preparation, martial arts rank certification, intangible cultural heritage martial arts inheritance, and professional competitive training. The system will retrieve the standard movement image model and feature parameters that match the goal from the multi-dimensional standard movement image database of martial arts based on the specific training goal selected by the student (such as only martial arts rank certification, or simultaneously including intangible cultural heritage martial arts inheritance and professional competitive training). For example, for the goal of intangible cultural heritage martial arts inheritance, the system will retrieve the movement standard that is more in line with the original appearance of traditional routines, and for the goal of professional competitive training, the system will retrieve the movement standard that conforms to the technical specifications of the competition. This will provide a precise goal-oriented basis for the subsequent dynamic adjustment of the weight of the evaluation index based on image recognition. Three core evaluation indicators based on image recognition were clearly defined for quantitatively judging the standardization of students' martial arts movements. Among them, the image joint point matching degree is used to accurately compare the coordinate offset of key joint points in the student's movement image and the standard movement image. Image similarity of movement postures is used to measure the visual similarity between a trainee's overall movement posture and a standard movement posture; Motion sequence image consistency is used to evaluate the frame-level synchronization matching between trainee motion image sequences and standard motion image sequences; At the same time, the weights of the three evaluation indicators are not fixed values, but rather need to be dynamically adjusted according to the current learning stage of the learners. For students in the initial learning stage, the focus is on increasing the weight of image joint matching degree and motion posture image similarity, thereby strengthening the assessment of the standardization of basic motion forms. For trainees at the advanced learning stage, the emphasis is placed on increasing the weight of consistency in movement sequence images and accuracy in posture detail images. This highlights the assessment of advanced training points such as control of movement rhythm and details of force transmission, thereby constructing a scientific evaluation system that meets the training needs of different learning stages and providing accurate judgment criteria for subsequent movement deviation diagnosis.

[0024] The multimodal image acquisition array consists of multiple synchronous high-speed industrial cameras with a sampling frame rate of ≥120fps. It has image anti-motion blur processing capabilities and is used to acquire multi-angle continuous image sequences of trainees' movements. Two-dimensional or three-dimensional joint coordinates and posture contour features are extracted through human posture image recognition algorithms. Image preprocessing includes image denoising, image enhancement, image scaling, and region of interest extraction to ensure the clarity of human motion regions in the image, providing high-quality data support for subsequent image feature extraction and matching, and ensuring accurate synchronization between motion image data and standard motion image models.

[0025] Furthermore, action image data acquisition is carried out by first relying on a multimodal image acquisition array composed of multiple synchronous high-speed industrial cameras. The camera sampling frame rate of this array is no less than 120fps, and it has image anti-motion blur processing function, which can effectively capture multi-angle continuous image sequences during the martial arts movements of the students. Then, the human posture image recognition algorithm is used to extract the coordinates of two-dimensional or three-dimensional joints and posture contour features of the human body from the acquired image sequence; The acquired raw motion image data undergoes preprocessing operations including image denoising, image enhancement, image scaling, and region of interest extraction. These processes filter out invalid interference information in the images and improve the clarity of human motion areas, thereby providing high-quality data support for subsequent stages of image feature extraction and matching with standard motion image models. At the same time, it ensures that the acquired student motion image data can be accurately synchronized with the retrieved standard motion image model.

[0026] The two-stage temporal alignment algorithm based on image semantic attention specifically includes: The first stage is global image semantic alignment, which extracts global pose image features between the trainee action image sequence and the standard action image model through a convolutional neural network, calculates the image semantic similarity, and achieves coarse-grained temporal alignment. The second stage involves precise alignment of core node images. Based on the attention mechanism, the image feature vectors of key joint points in the core movements of martial arts are focused. The image matching accuracy is optimized through adaptive dynamic weight iteration to achieve fine-grained temporal calibration. The two-stage temporal alignment algorithm can automatically adapt to the non-uniform speed characteristics of martial arts movements. Through image frame synchronization calibration, it ensures that the temporal consistency error of the temporal alignment is ≤5ms, and the alignment accuracy is not affected by changes in movement speed.

[0027] Furthermore, the two-stage temporal alignment algorithm based on image semantic attention completes the precise temporal alignment of the trainee action image sequence and the standard action image model in two progressive stages. The first stage performs global image semantic alignment, using a convolutional neural network to extract global pose image features of the trainee action image sequence and the standard action image model respectively, and completes coarse-grained temporal alignment by calculating the image semantic similarity between the two, thus initially matching the overall temporal framework of the action sequence. The second stage involves precise alignment of core node images. It relies on the attention mechanism to focus on the feature vectors of key joint images of core martial arts movements, constructs an adaptive dynamic weight iterative optimization model, and continuously adjusts the matching weights of core features to improve alignment accuracy and complete fine-grained temporal calibration. The entire two-stage temporal alignment algorithm can automatically adapt to the non-uniform speed characteristics of martial arts movements, which involve fast retraction and release and a combination of movement and stillness. Through the image frame synchronization calibration mechanism, it ensures that the temporal consistency error of the temporal alignment is ≤5ms, and the alignment accuracy will not decrease due to changes in movement speed.

[0028] The formula for the two-stage temporal alignment algorithm based on image semantic attention is as follows: Global pose image feature extraction formula: ,in Given a sequence of motion images as input, For convolutional neural network feature extractors, This is the output global pose image feature matrix; Semantic similarity calculation (coarse alignment) formula: ,in For the features of trainee action image sequence, Features of standard motion image models, The cosine similarity is used to perform coarse-grained temporal matching. Attention weight calculation formula: ,in , The first and second movements of the student and the standard movement, respectively. Key features of key nodes For feature matching scoring function, For the first Attention weights for each core node; Fine-grained alignment loss function (iterative optimization) formula: By minimizing this loss function, adaptive weight iterative optimization is completed, achieving accurate time series calibration. These calculation formulas are called sequentially in a two-stage temporal alignment logical order, progressing from global feature extraction to fine-grained precision optimization, to achieve accurate temporal matching between the student's action image sequence and the standard action image model. The specific usage, combined with martial arts action image processing scenarios, is as follows: Global pose image feature extraction formula This is the basic input layer formula of the algorithm. When used, it involves two types of image sequences: a continuous image sequence of the student's martial arts movements. Standard martial arts movement image model sequence Input the pre-trained convolutional neural network respectively Output the corresponding global pose feature matrix. (Student characteristics) and (Standard features) These two feature matrices contain key semantic information such as the overall posture and contour of the action, providing a data foundation for subsequent similarity calculation; Among them, the continuous image sequence of students' martial arts movements and standard martial arts movement image model sequence The acquisition method is based on the separate implementation of image acquisition and database retrieval processes in the entire monitoring and error correction method, as detailed below: Continuous image sequence of students' martial arts movements The images are directly acquired through a multimodal image acquisition array, which consists of multiple synchronous high-speed industrial cameras. These cameras have a sampling frame rate of ≥120fps and are equipped with anti-motion blur processing capabilities. During the student's martial arts routine movements, the cameras simultaneously capture the student's real-time actions from multiple angles, generating a continuous stream of motion images. After preprocessing operations such as image denoising, image enhancement, image scaling, and region of interest extraction, background interference is filtered out, and details of the human body's movement areas are enhanced. Finally, a continuous image sequence of the student's martial arts movements is obtained, which can be used for subsequent feature extraction. ; Standard martial arts movement image model sequence This is achieved through the S1 step process. The system first obtains the student's learning stage and training objectives, and then uses this as a matching basis to retrieve a sequence of standard movement images that matches the student's learning stage and training objectives from the multi-dimensional standard movement image database of martial arts integrated into the image benchmark matching module. This database pre-stores multi-angle continuous images of various standard routine movements performed by professional martial arts practitioners, and has completed the calibration of key joints, posture contours, and other features. After retrieval, a sequence of standard martial arts movement image models for comparison can be obtained. ; The above calculation formula is called sequentially in a two-stage temporal alignment logical order, progressing from global feature extraction to fine-grained precision optimization, to achieve accurate temporal matching between the student's action image sequence and the standard action image model. The specific usage, combined with martial arts action image processing scenarios, is as follows: Global pose image feature extraction formula This is the basic input layer formula of the algorithm. When used, it involves two types of image sequences: a continuous image sequence of the student's martial arts movements. Standard martial arts movement image model sequence Input the pre-trained convolutional neural network respectively (Previous technology), output the corresponding global pose feature matrix (Student characteristics) and (Standard features) These two feature matrices contain key semantic information such as the overall posture and contour of the action, providing a data foundation for subsequent similarity calculation; Semantic similarity calculation formula Used for the first stage of coarse-grained timing alignment, the result obtained in the previous step and Substitute the values ​​into the formula to calculate the cosine similarity, which ranges from [−1, 1]. The closer the value is to 1, the more similar the global semantic features of the two are. The algorithm will use the similarity value as a basis to perform preliminary matching of the frame order of the student's action image sequence. For example, it will match the feature frames of the student's punching action with the feature frames of the standard punching action to complete the alignment of the overall action temporal framework and avoid global temporal misalignment caused by the speed of the action. Attention weight calculation formula Applied to precise alignment of core nodes in the second stage, it is used by first starting from... and Extract feature vectors of key joints in core martial arts movements (such as features of key parts like hands, waist, and legs). , ),pass The function calculates the matching score for each pair of features, and then... Normalization yields attention weights The higher the weight value, the stronger the importance of the joint point to the timing of the movement (for example, the knee and ankle joints in a martial arts kicking movement will have a higher weight than other parts), thereby achieving focus on the core movement nodes. in, The specific operational logic is as follows: first, feature matching and scoring are performed on each core joint of martial arts. An exponential calculation is performed to avoid negative scores; then, the exponential score of a single key point is divided by the sum of the exponential scores of all key points to obtain the attention weight of that key point. Such calculations allow core joints with higher matching degrees (such as hands, waist, and legs in martial arts movements) to receive higher weights, thereby highlighting the impact of key parts on the accuracy of movement matching in subsequent fine-grained temporal alignment loss calculations. Fine-grained alignment loss function This is the iterative optimization loss function formula for accurate alignment of core node images in the second stage of a two-stage temporal alignment algorithm based on image semantic attention. It incorporates the attention weights obtained in the previous step. The total loss value is obtained by multiplying the difference between the features of the corresponding core joints and summing the results. The algorithm continuously iterates and adjusts the feature matching parameters using optimization algorithms such as gradient descent to minimize the loss value. At this point, the image frames of the student's core joints are precisely matched with the corresponding frames of the standard movements, ultimately ensuring that the timing alignment error is ≤5ms and adapting to the non-uniform speed characteristics of martial arts movements. The two-stage temporal alignment algorithm based on image semantic attention sets a termination condition for loss convergence, which avoids the waste of computational resources caused by excessive algorithm iteration and prevents alignment accuracy defects caused by insufficient iteration, thus ensuring a balance between running efficiency and matching accuracy. Achieving a time consistency error of ≤5ms ensures accurate matching between the core joint image frames of the trainee's movements and the corresponding frames of the standard movements, providing a high-precision time reference for subsequent calculation of deviation data such as joint offset and posture similarity, and significantly improving the accuracy of movement deviation diagnosis. It adapts to the non-uniform speed characteristics of martial arts movements. It addresses the frame misalignment problem that traditional timing alignment algorithms often encounter when handling variable-speed movements, taking into account the characteristics of fast power generation, slow recovery, and alternating movement in martial arts movements. This ensures stable and accurate timing matching for both fast long punches and slow Tai Chi movements.

[0029] The deep learning image recognition model is trained using a self-supervised contrastive learning strategy combined with a spatial attention mechanism. The spatial attention mechanism accurately focuses on the image feature regions of key segments in the transmission of force in martial arts movements. The training dataset includes standard movement image sequences, keyframe image features, and a dataset of historical training errors from students, all from a multidimensional standard movement image database of martial arts. Deep learning image recognition models can capture error precursor image features by detecting feature changes in consecutive image frames and trigger warning signals.

[0030] Furthermore, a training dataset is first constructed, which integrates standard action image sequences acquired from multiple angles and with joint point calibration from the multi-dimensional standard action image database of martial arts, key frame image features marked with force transmission characteristics, and a dataset of historical training error action images of students covering typical problems such as deviation of force exertion points, interruption of force transmission, and deformation of action posture. Subsequently, training was conducted on the deep learning image recognition model. A self-supervised contrastive learning strategy was adopted, which could drive the model to learn discriminative representations of action features without manual annotation by constructing positive and negative sample pairs (setting images of the same standard action from different angles as positive samples and images of standard and incorrect actions as negative samples). At the same time, a spatial attention mechanism was integrated, which assigned higher feature extraction weights to key segments of power transmission in martial arts movements, such as the shoulder-elbow-wrist power chain and the hip-knee-ankle support chain, so that the model could accurately focus on the image features of the above core areas. After the model training converges and performance verification is completed, it is deployed to the martial arts teaching real-time monitoring system. Continuous motion image frames during the student's training process are input in real time. The model extracts and tracks the feature change patterns of the continuous image frames, compares them with standard motion feature thresholds, and accurately captures error precursor image features such as abnormal shoulder posture before exertion and deviation of leg support angle. It immediately triggers warning signals so that the coach or system can intervene in time to correct the motion deviation. Those skilled in the art can complete the training, deployment, and application of the model according to the above process.

[0031] The personalized correction training program is built based on the deviation diagnosis results of image recognition and includes three core modules: image-guided basic morphological correction module, image temporal comparison-based strength enhancement module, and image frame synchronization-based rhythm control module. The basic morphology correction module identifies joint point matching deviations from image recognition and designs stretching and shaping training guided by image overlay. The force enhancement module is designed to conduct specific training for force transmission based on feature difference analysis of key frame images of movements. The rhythm control module designs beat-following training by comparing the frame synchronization of the trainee's action image sequence with the standard action image sequence; Personalized correction training programs can dynamically adjust training intensity and duration based on the training effect feedback from image recognition.

[0032] Furthermore, based on the motion deviation diagnosis results output by the deep learning image recognition model, including joint matching deviation data, motion key frame feature difference information, and temporal rhythm deviation parameters between the trainee and the standard motion sequence, a personalized correction training program is constructed. The program integrates three core modules: image-guided basic form correction module, image temporal comparison force enhancement module, and image frame synchronization rhythm control module. The basic posture correction module addresses the positional mismatch of core joints such as the shoulder, elbow, hip, and knee identified in the diagnosis. It overlays the student's movement image with the standard movement image in real time, marking the location and value of the deviation. Based on this, it designs corresponding static stretching and movement shaping training for the affected areas to help students correct their basic posture. The power enhancement module is based on the power transmission feature difference analysis results of key frame images of the movement. It focuses on the characteristic deviations of joint angle and power transmission chain at the moment of force exertion. It designs special power transmission training such as shoulder and elbow power connection and hip and knee support transmission. At the same time, it visually displays the difference in power transmission between the trainee and the standard movement through time-series comparison images. The rhythm control module compares the trainee's action image sequence with the standard action image sequence in real time, marking the frame intervals where the trainee's movements are too fast or too slow. Based on this, beat-following training is designed to guide the trainee to match the standard action rhythm. During the training process, the system continuously collects trainee's training action images and uses image recognition models to provide feedback on training effect data. Based on this, the system dynamically adjusts the training intensity and duration of each module, forming a closed-loop training process of "deviation diagnosis - scheme formulation - training implementation - effect feedback - scheme optimization".

[0033] AR visualization overlay display in multimodal error correction feedback is achieved based on image recognition registration technology, specifically including: The image recognition registration algorithm accurately overlays the real-time motion images of trainees with standard motion image models on the same screen. The image recognition technology is used to extract the coordinates of key joint points and posture contours of the two. After calculating the degree of image feature deviation, the images are visualized and labeled to intuitively present the position, angle and range of the motion deviation.

[0034] Furthermore, firstly, standard movement image models that match the student's learning stage and training objectives are called from the martial arts multidimensional standard movement image database, and at the same time, real-time movement images of the student are acquired through a multimodal image acquisition array; Then, the image recognition registration algorithm is started to extract the two-dimensional or three-dimensional coordinates and overall posture contour feature vectors of the core joints such as shoulder, elbow, wrist, hip, knee and ankle in the real-time action image of the trainee and the standard action image model. Based on the spatial mapping relationship of the joint coordinates, the trainee image and the standard image are accurately registered and superimposed on the same screen. Then, based on the registered coordinate data, the degree of feature deviation between the two is calculated, specifically including the deviation position of the joints, the angle difference, and the deviation range of the posture contour, and these quantified deviation data are converted into visual annotation information; Finally, the standard movement outline is superimposed on the student's real-time movement screen in a semi-transparent form through the AR display interface. Different colors (such as red to highlight the deviation area) are used to mark the deviation position, and the deviation angle value and deviation range area are displayed simultaneously, so that students and coaches can intuitively and in real time grasp the details of the movement deviation. Those skilled in the art can complete the deployment and implementation of the AR visualization overlay display function according to the above process.

[0035] A standardized real-time monitoring and error correction system for martial arts routine teaching includes an image benchmark matching module, an image acquisition and alignment module, an image recognition deviation diagnosis and prediction module, an image feedback intervention module, and an interaction module.

[0036] Furthermore, the image benchmark matching module first retrieves the corresponding standard action image sequence, key frame feature data, and force transmission annotation information from the pre-made multi-dimensional standard action image database of martial arts, based on the learning stage and training goal selected by the student (such as martial arts exam preparation, rank certification, intangible cultural heritage inheritance, etc.), to complete the accurate matching of the standard benchmark. Next, the image acquisition and alignment module acquires real-time martial arts action images of trainees through multiple synchronous high-speed industrial cameras (frame rate ≥120fps, with anti-motion blur function). After preprocessing such as denoising, enhancement, and region of interest extraction, the module calls a two-stage temporal alignment algorithm based on image semantic attention to accurately register the trainee action image sequence with the standard action image sequence in the same frame, achieving a time consistency error alignment of ≤5ms, which is suitable for the non-uniform speed characteristics of martial arts actions with fast force exertion and slow recovery. The image recognition deviation diagnosis and prediction module loads a deep learning model trained by self-supervised contrastive learning combined with spatial attention mechanism. It inputs the aligned student action image features and compares them with standard feature data to diagnose problems such as joint matching deviation, force transmission defects, and movement rhythm imbalance. At the same time, it captures the feature change patterns of continuous image frames to identify error precursors and trigger warning signals. Based on the deviation diagnosis results, the image feedback intervention module automatically generates a personalized correction training plan that includes three core modules: image-guided basic form correction, image time sequence comparison-based strength enhancement, and image frame synchronization-based rhythm control. At the same time, it calls the AR visualization overlay function to overlay the standard movement outline with the trainee's real-time movement on the same screen through image recognition registration technology, marking the deviation position, angle and range, and realizing multimodal error correction feedback. The interaction module provides a human-computer interaction interface for trainees and coaches, supporting operations such as setting training goals, retrieving standard movements, viewing warning information, and adjusting training plans. At the same time, it provides real-time feedback on training effect data, driving the system to dynamically optimize alignment algorithm parameters and training plan intensity.

[0037] The multidimensional standard martial arts movement image database is integrated into the image benchmark matching module. It stores quantified parameters such as three-dimensional dynamic image sequences, key frame image features, and joint point image coordinates of multiple core martial arts movements, and supports fast retrieval and matching based on image features.

[0038] Furthermore, a multi-dimensional standard martial arts movement image database is integrated into the image benchmark matching module as core data support. It pre-stores 3D dynamic image sequences of multiple core martial arts movements performed by professional martial artists, keyframe image features marking force transmission characteristics, and quantified parameters such as 2D or 3D image coordinates of 18 core human joints. The database is structured and archived according to martial arts routine type, learning stage, and training goal. It also includes a built-in fast retrieval engine based on image features. When the interaction module receives instructions from the student regarding the learning stage and training goal (such as rank certification or intangible cultural heritage transmission), the image benchmark matching module immediately calls this integrated database. Through the retrieval engine, it compares the student's requirements with the movement feature tags and joint parameter features in the database, quickly retrieving and matching the corresponding standard movement data. Simultaneously, this standard movement data is transmitted to the image acquisition and alignment module, providing a unified benchmark for the subsequent temporal alignment and feature comparison of the student's real-time movement images. The entire process requires no additional external database access, and the retrieval and matching logic is clear and the parameters are traceable. Those skilled in the art can deploy and implement this system based on the aforementioned database construction specifications and module integration methods.

[0039] The image acquisition and alignment module includes an edge image calculation unit, which is responsible for real-time preprocessing of motion image data, pose reconstruction based on image recognition, and timing alignment calculation, ensuring low latency in image recognition and error correction feedback. The edge image computing unit works in collaboration with the cloud server, which is responsible for storing students' historical action images, training and updating deep learning image recognition models, and performing big data analysis of group error images. The interactive module includes an AR image display device for students, a monitoring screen for teachers, and an image data synchronization unit. The image data synchronization unit can synchronize students' training image sequences, image recognition error diagnosis reports, and image data on the progress of correction schemes to the cloud server in real time.

[0040] Furthermore, the edge image computing unit built into the image acquisition and alignment module immediately performs preprocessing operations such as image denoising, enhancement, and region of interest extraction after receiving the real-time action images of trainees transmitted by the multimodal image acquisition array. Simultaneously, it performs human pose reconstruction and temporal alignment calculation based on image recognition (calling a two-stage temporal alignment algorithm based on image semantic attention). By leveraging local computing power at the edge to rapidly process data, cloud transmission delays are avoided, ensuring low-latency requirements for image recognition and error correction feedback. The edge image computing unit and the cloud server form a collaborative working mode. The edge end synchronizes the processed real-time action feature data of the trainees and error precursor information to the cloud. The cloud server is responsible for storing the trainees' historical action image data, and at the same time, it carries out the training, iterative update of the deep learning image recognition model and the big data analysis of group error images to explore the patterns of high-frequency error actions and feed back to optimize the model. The optimized model is then synchronously sent to the edge computing unit to improve the recognition accuracy. The student-side AR image display device in the interaction module is used to present AR visualization overlay images and personalized correction solutions. The teacher-side monitoring screen displays the student's training status, deviation diagnosis results, and early warning information in real time. The image data synchronization unit synchronizes the student's training image sequence, image recognition error diagnosis report, and correction solution execution progress image data to the cloud server in real time, realizing data interconnection and interoperability between the edge, cloud, student, and teacher ends. This ensures the smoothness of real-time teaching monitoring and error correction, and provides data support for model optimization and group teaching analysis. Those skilled in the art can complete the deployment of each unit and data flow configuration based on the above division of labor and collaborative logic.

[0041] Although embodiments of the invention have been shown and described, it will be understood by those skilled in the art that various changes, modifications, substitutions and alterations can be made to these embodiments without departing from the principles and spirit of the invention, the scope of which is defined by the appended claims and their equivalents.

Claims

1. A standardized real-time monitoring and error correction method for martial arts routine teaching, characterized in that: include: S1. Obtain information on the student's learning stage and training objectives, retrieve the corresponding standard movement image model and image feature parameters from the martial arts multidimensional standard movement image database, and dynamically adjust the weight of the evaluation indicators based on image recognition. S2. Motion image data is acquired through a multimodal image acquisition array. After image preprocessing, a two-stage temporal alignment algorithm based on image semantic attention is used to achieve accurate temporal alignment between the trainee's motion image sequence and the standard motion image model. S3. Calculate the action deviation based on image feature matching, classify the error type through the image recognition technology, distinguish the deviation attribute by combining historical action image data, and use a deep learning image recognition model to extract the error precursor image features and trigger an early warning. S4. Output multimodal error correction feedback based on image overlay, push the group common error image analysis report to the teacher's end, generate the deviation diagnosis results based on image recognition and dynamically adjust the personalized correction scheme to form a closed-loop error correction.

2. The method according to claim 1, characterized in that, The training objectives include at least one of the following: martial arts exam preparation, martial arts grading certification, intangible cultural heritage martial arts inheritance, and professional competitive training. The image recognition-based evaluation metrics specifically include three categories: image key point matching degree, action posture image similarity, and action temporal image consistency. The weights of the three evaluation metrics are dynamically adjusted according to the student's learning stage. The initial learning stage focuses on the weight of the image joint matching degree and the similarity of the action posture image, while the advanced learning stage focuses on the weight of the consistency of the action temporal image and the accuracy of the posture detail image.

3. The method according to claim 1, characterized in that, The multimodal image acquisition array consists of multiple synchronous high-speed industrial cameras with a sampling frame rate of ≥120fps and image anti-motion blur processing function. It is used to acquire multi-angle continuous image sequences of trainees' movements and extract two-dimensional or three-dimensional joint coordinates and posture contour features through human posture image recognition algorithms. The image preprocessing includes image denoising, image enhancement, image scaling, and region of interest extraction to ensure the clarity of human motion regions in the image, provide high-quality data support for subsequent image feature extraction and matching, and ensure accurate synchronization between motion image data and standard motion image models.

4. The method according to claim 1, characterized in that, The two-stage temporal alignment algorithm based on image semantic attention specifically includes: The first stage is global image semantic alignment, which extracts global pose image features of the trainee action image sequence and the standard action image model through a convolutional neural network, calculates the image semantic similarity, and achieves coarse-grained temporal alignment. The second stage involves precise alignment of core node images. Based on the attention mechanism, the image feature vectors of key joint points in the core movements of martial arts are focused. The image matching accuracy is optimized through adaptive dynamic weight iteration to achieve fine-grained temporal calibration. The dual-stage temporal alignment algorithm can automatically adapt to the non-uniform speed characteristics of martial arts movements. Through image frame synchronization calibration, it ensures that the temporal consistency error of the temporal alignment is ≤5ms, and the alignment accuracy is not affected by changes in movement speed.

5. The method according to claim 1, characterized in that, The deep learning image recognition model is trained using a self-supervised contrastive learning strategy combined with a spatial attention mechanism. The spatial attention mechanism accurately focuses on the image feature regions of key segments in the transmission of force in martial arts movements. The training dataset includes the standard movement image sequence, the keyframe image features, and the student's historical training error movement image dataset from the martial arts multidimensional standard movement image database. The deep learning image recognition model can capture error precursor image features through feature changes in consecutive image frames and trigger warning signals.

6. The method according to claim 1, characterized in that, The personalized correction training program is built based on the deviation diagnosis results of image recognition and includes three core modules: image-guided basic morphological correction module, image temporal comparison-based strength enhancement module, and image frame synchronization-based rhythm control module. The basic morphology correction module identifies joint point matching deviations based on image recognition and designs stretching and shaping training guided by image overlay. The force enhancement module designs specific training for force transmission based on feature difference analysis of the key frame images of the action. The rhythm control module designs beat-following training by comparing the frame synchronization of the trainee's motion image sequence with the standard motion image sequence; The personalized correction training program can dynamically adjust the training intensity and duration based on the training effect feedback from image recognition.

7. The method according to claim 1, characterized in that, The AR visualization overlay display in the multimodal error correction feedback is implemented based on image recognition registration technology, specifically including: The image recognition registration algorithm accurately overlays the student's real-time motion image with the standard motion image model on the same screen. The image recognition technology is used to extract the coordinates of key joint points and posture contours of the two. After calculating the degree of image feature deviation, the image is visualized and labeled to intuitively present the position, angle and range of the motion deviation.

8. A standardized real-time monitoring and error correction system for martial arts routine teaching, implementing the method of any one of claims 1-7, characterized in that, It includes an image benchmark matching module, an image acquisition and alignment module, an image recognition deviation diagnosis and prediction module, an image feedback intervention module, and an interaction module.

9. The system according to claim 8, characterized in that, The multidimensional standard martial arts movement image database is integrated into the image benchmark matching module. It stores the three-dimensional dynamic image sequence of multiple core martial arts movements, the key frame image features, the joint point image coordinates and other quantitative parameters, and supports fast retrieval and matching based on image features.

10. The system according to claim 8, characterized in that, The image acquisition and alignment module includes an edge image calculation unit, which is responsible for real-time preprocessing of the motion image data, pose reconstruction based on the image recognition, and time alignment calculation, ensuring low latency in image recognition and error correction feedback. The edge image computing unit works in collaboration with the cloud server, which is responsible for storing the student's historical action image data, training and updating the deep learning image recognition model, and performing big data analysis of group error images. The interactive module includes a student-side AR image display device, a teacher-side monitoring screen, and an image data synchronization unit. The image data synchronization unit can synchronize the student's training image sequence, the image recognition error diagnosis report, and the image data of the correction plan execution progress to the cloud server in real time.