Physical training action posture recognition and correction system based on computer vision
By constructing a mathematical model for physiological state perception, dynamic environment adaptation, and temporal action recognition modules, combined with a multi-factor collaborative correction module, the problem of posture recognition error accumulation in existing technologies is solved, achieving high-precision posture correction.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-12-11
- Publication Date
- 2026-03-13
AI Technical Summary
Existing sports training posture recognition and correction systems fail to effectively consider the impact of changes in the trainee's physiological state, dynamic interference in the training environment, and the accumulation of time sequence movement errors. This results in a high misjudgment rate in posture recognition during fatigue training, complex environments, or long-term training scenarios. Furthermore, the correction instructions do not match the trainee's actual movement deviation requirements, thus failing to meet the needs of high-precision and personalized sports training.
The system constructs a physiological state perception module, a dynamic environment adaptation module, a temporal action recognition module, and an initial calibration module. Through a multi-factor collaborative correction module combined with a mathematical model, it calculates the basic error coefficients, temporal cumulative error values, and final attitude correction compensation amount for attitude recognition, and outputs attitude correction commands.
It achieves accurate posture recognition and correction in complex training scenarios, adapts to changes in trainee status and environmental interference, and meets the needs of high-precision sports training.
Smart Images

Figure CN121662290A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the fields of computer vision and sports training technology, specifically a sports training movement posture recognition and correction system based on computer vision. Background Technology
[0002] Existing sports training posture recognition and correction systems generally rely on the analysis of single computer vision image features, failing to consider the progressive impact of changes in the trainee's physiological state, dynamic interference in the training environment, and the accumulation of sequential movement errors. In actual training scenarios, muscle fatigue can lead to abnormal movement amplitude and force exertion; changes in lighting and background movement can interfere with the accuracy of image feature extraction; and small deviations in a single frame of movement during continuous training can gradually accumulate, forming significant posture errors. These factors interact, resulting in high posture recognition misjudgment rates in existing systems under fatigue training, complex environments, or long-term training scenarios. Furthermore, the correction instructions do not match the trainee's actual movement deviation requirements, failing to meet the needs of high-precision, personalized sports training.
[0003] Based on the above problems, there is an urgent need for a technical solution that can coordinate multiple key influencing factors and achieve progressive error correction. Summary of the Invention
[0004] The purpose of this invention is to address the shortcomings of existing technologies by proposing a computer vision-based sports training posture recognition and correction system. This system includes a physiological state perception module, a dynamic environment adaptation module, a temporal movement recognition module, a multi-factor collaborative correction module, and an initial calibration module. These modules establish data transmission connections with the multi-factor collaborative correction module. The initial calibration module acquires baseline physiological parameters and baseline motion parameters for standard movements in a non-fatigue state, and sets baseline environmental parameters for the standard training environment. The physiological state perception module collects real-time electromyographic signal amplitude, muscle exertion level, and joint range of motion of the trainee. The dynamic environment adaptation module collects the illumination spectrum characteristics and background dynamic interference speed of the training environment. The temporal movement recognition module extracts the single-frame posture deviation angle, inter-frame time interval, number of movement frames, and cumulative training time of the trainee's movements based on computer vision. The multi-factor collaborative correction module receives the aforementioned real-time data and baseline parameters, and sequentially calculates the basic posture recognition error coefficient, temporal cumulative error value, and final posture correction compensation amount through three progressively linked mathematical models, outputting a posture correction command.
[0005] Preferably, the baseline physiological parameters include baseline electromyographic signal amplitude and standard joint range of motion; the baseline motion parameters include inter-frame baseline time interval and baseline force level; and the baseline environmental parameters include baseline illumination spectrum value. The initial calibration module determines the baseline electromyographic signal amplitude and standard joint range of motion by collecting physiological data of the trainee in a static state, determines the inter-frame baseline time interval and baseline force level by analyzing the temporal data of standard movements, and determines the baseline illumination spectrum value by detecting illumination data of an interference-free training environment.
[0006] More preferably, the physiological state sensing module includes an electromyography (EMG) sensor and a joint range of motion sensor. The EMG sensor is fitted to the trainee's muscle groups to collect real-time EMG signal amplitude and convert it into muscle force level. The joint range of motion sensor is set at the trainee's key joints to collect joint range of motion.
[0007] More preferably, the dynamic environment adaptation module includes an optical sensor and an image processor. The optical sensor is set facing the training environment to collect spectral feature values of illumination. The image processor receives the environmental image collected synchronously by the optical sensor, separates the background area from the trainee area through an image segmentation algorithm, and calculates the moving speed of the moving target in the background area as the background dynamic interference speed.
[0008] More preferably, the temporal action recognition module includes a visual acquisition device and a joint point analysis unit. The visual acquisition device captures the trainee's action video in real time and splits it into continuous action frames. The joint point analysis unit extracts the coordinates of the trainee's key joint points in each action frame, calculates the single-frame posture deviation angle based on the deviation between the key joint point coordinates and the standard joint point coordinates, records the time difference between adjacent action frames as the inter-frame time interval, accumulates the number of action frames as the action frame count, and records the time from the start of the visual acquisition device to the current action frame as the cumulative training time.
[0009] More preferably, the attitude recognition basic error coefficient is obtained by the attitude recognition basic error coefficient calculation formula, which is: ; in The amplitude of the real-time electromyographic signal is measured in volts. The reference electromyographic signal amplitude is measured in volts. is a characteristic value of the illumination spectrum, with dimensions of candela per square meter; The standard value for the light spectrum is expressed in candela per square meter. The background dynamic disturbance velocity is expressed in meters per second. , , These are preset coordination coefficients, all of which are dimensionless.
[0010] More preferably, the time-series cumulative error value is obtained through a time-series cumulative error value calculation formula, which is: ; in The basic error coefficient for attitude recognition is dimensionless. The attitude deviation angle for a single frame is expressed in degrees. The number of action frames is dimensionless. The time interval between frames is measured in seconds. This is the inter-frame reference time interval, measured in seconds. This represents the level of muscle exertion, and is dimensionless. The base force level is dimensionless.
[0011] More preferably, the final attitude correction compensation amount is obtained through the final attitude correction compensation amount calculation formula, which is as follows: ; in This represents the cumulative time-series error, measured in degrees. The basic error coefficient for attitude recognition is dimensionless. The range of motion of a joint is measured in degrees. Standard joint range of motion, dimensionless; The preset time decay coefficient is measured in seconds. The total training time is measured in seconds.
[0012] More preferably, the preset synergistic adjustment coefficient is obtained by fitting training samples. The initial calibration module collects at least 50 sets of physiological data, environmental data, and action data from different trainees as training samples, and fits the training samples using a multiple linear regression algorithm to determine the coefficient. , , The specific values are then stored in the multi-factor collaborative correction module.
[0013] More preferably, the posture correction instructions include visual instructions and auditory instructions. The multi-factor collaborative correction module determines the direction parameters of the visual instructions and the speech parameters of the auditory instructions based on the final posture correction compensation amount. The visual instructions and auditory instructions are output synchronously to guide the trainee to adjust their movement posture.
[0014] Technical Effects: The core inventive technology of this invention lies in constructing a three-factor collaborative acquisition mechanism of physiological, environmental, and temporal factors and a progressive linkage mathematical model, combined with an initial calibration module to obtain accurate benchmark parameters. Through multi-module collaborative data transmission, the basic error coefficient, temporal cumulative error value, and correction compensation amount are calculated sequentially, effectively solving the problem of insufficient recognition and correction accuracy caused by the single reliance on image features in existing technologies. This enables accurate posture recognition and correction in complex training scenarios, adapting to changes in trainee state and environmental interference, and meeting the needs of high-precision sports training. Attached Figure Description
[0015] Figure 1 This is a connection block diagram of the computer vision-based sports training motion posture recognition and correction system of this application. Detailed Implementation
[0016] To make the objectives, technical solutions, and advantages of this invention clearer, the invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the invention.
[0017] Traditional technical solutions have the following technical problems: Existing sports training posture recognition and correction systems rely solely on computer vision image features for analysis, without considering the progressive effects of changes in the trainee's physiological state, dynamic interference in the training environment, and the accumulation of time-series movement errors. This results in a high misjudgment rate in posture recognition during fatigue training, complex environments, or long-term training scenarios, and the correction instructions do not match the trainee's actual movement deviation requirements, thus failing to meet the needs of high-precision training.
[0018] Based on this, please refer to Figure 1 This embodiment provides a computer vision-based sports training posture recognition and correction system, including a physiological state perception module, a dynamic environment adaptation module, a temporal movement recognition module, a multi-factor collaborative correction module, and an initial calibration module. The physiological state perception module, dynamic environment adaptation module, temporal movement recognition module, and initial calibration module establish data transmission connections with the multi-factor collaborative correction module. The initial calibration module acquires the baseline physiological parameters of the trainee in a non-fatigue state and the baseline motion parameters of the standard movement, and sets the baseline environmental parameters of the standard training environment. The physiological state perception module collects the trainee's real-time electromyographic signal amplitude, muscle force level, and joint range of motion. The dynamic environment adaptation module collects the illumination spectrum characteristics and background dynamic interference speed of the training environment. The temporal movement recognition module extracts the trainee's single-frame posture deviation angle, inter-frame time interval, number of movement frames, and cumulative training time based on computer vision. The multi-factor collaborative correction module receives the above real-time data and baseline parameters, and calculates the posture recognition basic error coefficient, temporal cumulative error value, and final posture correction compensation amount sequentially through three types of progressively linked mathematical models, and outputs posture correction instructions.
[0019] In this technical solution, the connection between modules uses Ethernet or USB 3.0 bus for data transmission, ensuring a data transmission rate of no less than 100Mbps to meet real-time requirements. The initial calibration module performs a calibration process after system startup, requiring the trainee to remain still for 30 seconds. Electromyography (EMG) signals and joint range of motion data are collected during this period and averaged as baseline physiological parameters. Simultaneously, a standard movement video is played, and the temporal data of the video is collected and analyzed, including frame intervals and force levels, as baseline motion parameters. Then, 10 minutes of illumination data is collected in a training environment with stable lighting and no movement of people, and the averaged data is taken. The mean value serves as the baseline environmental parameter. The electromyography (EMG) sensors in the physiological state sensing module are fitted to the trainee's upper arm, thigh, and other major muscle groups, detecting muscle electrical activity signals and outputting real-time EMG signal amplitudes. These amplitudes are then mapped to muscle exertion levels of 1-5. Joint range sensors are installed at key joints such as the shoulder, elbow, and knee, detecting joint rotation angles and outputting joint range data. The optical sensors in the dynamic environment adaptation module face the main light source of the training environment, collecting light intensity at different wavelengths and integrating it into a light spectrum feature value. The image processor receives the sensor data from the same sources. The environmental images acquired step-by-step are used to separate the background and trainee areas using image segmentation algorithms. Then, optical flow is used to calculate the pixel movement distance of moving targets within the background area, which is then converted into the actual dynamic speed of the background interference based on the image scale. The temporal motion recognition module uses a 2K resolution industrial camera, fixed 3 meters in front of the trainee, to capture motion video at 25fps and break it down into continuous motion frames. The joint point analysis unit uses a human pose estimation algorithm to extract the coordinates of 25 key joint points of the trainee in each frame, comparing the actual coordinates of each joint point with the standard motion joint point coordinates. The single-frame attitude deviation angle is obtained through geometric calculation, and the timestamp difference between adjacent frames is recorded as the inter-frame time interval. The number of action frames is counted starting from the video start frame, and the cumulative training time is obtained from the camera start time. After receiving the data transmitted by each module, the multi-factor collaborative correction module first substitutes it into the first type of mathematical model to calculate the basic error coefficient of attitude recognition, then substitutes it into the second type of mathematical model based on the coefficient and the action timing data to calculate the timing cumulative error value, and finally combines the joint range of motion and training time into the third type of mathematical model to calculate the final attitude correction compensation amount. Finally, the corresponding attitude correction command is generated based on the compensation amount.
[0020] Traditional technical solutions have the following technical problems: the existing system lacks a clear mechanism for obtaining reference parameters, which leads to a lack of unified reference standards in the subsequent error calculation and attitude correction process, thus affecting the calculation accuracy and correction effect.
[0021] Based on this, the baseline physiological parameters include baseline electromyographic signal amplitude and standard joint range of motion; the baseline motion parameters include inter-frame baseline time interval and baseline force level; and the baseline environmental parameters include baseline illumination spectrum value. The initial calibration module determines the baseline electromyographic signal amplitude and standard joint range of motion by collecting physiological data of the trainee in a static state, determines the inter-frame baseline time interval and baseline force level by analyzing the temporal data of standard movements, and determines the baseline illumination spectrum value by detecting illumination data of an interference-free training environment.
[0022] In this technical solution, when the initial calibration module collects physiological data of the trainee in a static state, it prompts the trainee to maintain a standard standing posture for 30 seconds. During this period, the electromyography (EMG) sensor continuously collects EMG signals, removes outliers, and takes the average value as the baseline EMG signal amplitude. The joint range of motion sensor records the maximum range of motion of each joint when the trainee slowly moves each key joint to its maximum range of motion and takes the median value as the standard joint range of motion. When analyzing the standard movement timing data, the initial calibration module plays a preset standard movement video with a duration of 10 seconds. The module extracts the movement frame sequence of the video, calculates the time interval between adjacent frames, and takes the average value as the baseline time interval between frames. At the same time, by analyzing the muscle exertion of the standard movement in the video and combining it with the range of EMG signal amplitude, the baseline exertion level corresponding to the standard movement is determined, which is usually set to level 3. When detecting interference-free training environment lighting data, the initial calibration module ensures that there are no people walking around or strong light blocking factors in the training environment. It controls the optical sensor to continuously collect 10 minutes of lighting spectrum data at a sampling frequency of 1Hz. After removing the maximum and minimum values, the average value is taken as the baseline value of the lighting spectrum to ensure the accuracy of subsequent environmental data comparison.
[0023] Traditional technical solutions have the following technical problems: the structure of existing physiological state sensing modules is vague, and the sensor placement and data acquisition methods are unclear, which makes it impossible to accurately obtain physiological state-related data of trainees and affects the reliability of subsequent error analysis.
[0024] Based on this, the physiological state sensing module includes an electromyography (EMG) sensor and a joint range of motion sensor. The EMG sensor is fitted to the trainee's muscle groups to collect real-time EMG signal amplitude and convert it into muscle force level. The joint range of motion sensor is set at the trainee's key joints to collect joint range of motion.
[0025] In this technical solution, the electromyography (EMG) sensor uses a surface EMG sensor. Its electrodes are attached to the skin surface of the trainee's main muscle groups, such as the biceps brachii in the upper arm and the quadriceps femoris in the thigh, ensuring good contact between the electrodes and the skin to reduce signal interference. The sensor sampling frequency is set to 2000Hz, enabling precise capture of electrical signal changes during muscle contraction. The output real-time EMG signal amplitude range is 0-5V. The internal signal processing unit divides this amplitude range into five intervals, corresponding to muscle exertion levels 1-5. An amplitude of 0-1V corresponds to level 1, 1-2, etc. V corresponds to level 2, 2-3V corresponds to level 3, 3-4V corresponds to level 4, and 4-5V corresponds to level 5. The joint range of motion sensor adopts a Hall effect angle sensor. Its fixed end is installed on one side of the limb of the joint, while the movable end is connected to the other side of the limb of the joint. It can detect angle changes synchronously with joint rotation. The sensor's measurement range is 0-180 degrees, and the measurement accuracy is 0.1 degrees. It can output the range of motion data of key joints such as shoulder, elbow, and knee joints in real time. The data is transmitted to the multi-factor collaborative correction module via wired means, with a transmission delay of less than 50 milliseconds.
[0026] Traditional technical solutions have the following technical problems: the existing dynamic environment adaptation module does not clearly define the functions of core components and data processing methods, and cannot effectively capture changes in lighting and background interference in the training environment, resulting in the inability to accurately quantify the impact of environmental factors on posture recognition.
[0027] Based on this, the dynamic environment adaptation module includes an optical sensor and an image processor. The optical sensor is set facing the training environment to collect spectral feature values of illumination. The image processor receives the environmental image collected synchronously by the optical sensor, separates the background area from the trainee area through an image segmentation algorithm, and calculates the moving speed of the moving target in the background area as the background dynamic interference speed.
[0028] In this technical solution, the optical sensor uses a miniature spectrometer, which detects the visible light band of 380-780nm and has a sampling frequency of 1Hz. The sensor's lens is oriented towards the main light source of the training environment, such as a window or overhead light, enabling it to collect light intensity data at different wavelengths. This data is then integrated by an internal algorithm into a light spectral feature value representing the current lighting state, measured in candela per square meter. The image processor uses an embedded processor, specifically an ARM Cortex-A9, which receives 1920×1080 resolution environmental images synchronously acquired by the optical sensor. The image frame rate is 25fps, consistent with the visual acquisition device. The processor first uses the MaskR-CNN image segmentation algorithm to accurately segment the trainee region and background region in the image based on the trained model, obtaining an image region containing only the background. Then, it uses optical flow to calculate the movement vector of each pixel in the background region. Based on the pixel movement distance and the time interval between adjacent frames, it calculates the pixel movement speed. Combining this with the image scale, it converts the pixel movement speed into the actual background dynamic interference speed, measured in meters per second. If there is no moving target in the background region, the background dynamic interference speed is set to 0.
[0029] Traditional technical solutions have the following technical problems: existing time-series action recognition modules do not clearly define the processing flow of action videos and the calculation method of key parameters, resulting in incomplete and inaccurate extraction of action-related data, and failing to provide reliable data support for time-series error calculation.
[0030] Based on this, the temporal action recognition module includes a visual acquisition device and a joint point analysis unit. The visual acquisition device captures the trainee's action video in real time and splits it into continuous action frames. The joint point analysis unit extracts the coordinates of the trainee's key joint points in each action frame, calculates the single-frame posture deviation angle based on the deviation between the key joint point coordinates and the standard joint point coordinates, records the time difference between adjacent action frames as the inter-frame time interval, accumulates the number of action frames as the action frame count, and records the time from the start of the visual acquisition device to the current action frame as the cumulative training time.
[0031] In this technical solution, the visual acquisition device uses a 2K resolution (2560×1440) industrial camera equipped with a fixed-focus lens with a focal length of 8mm. The camera is fixed 3 meters directly in front of the trainee, with the lens at shoulder height. It captures real-time video of the trainee's full-body movements at a frame rate of 25fps. The video data is transmitted to the joint point analysis unit via an HDMI interface. Simultaneously, the device's built-in timestamp module adds a precise timestamp to each frame. The joint point analysis unit uses the OpenPose human pose estimation algorithm based on deep learning. This algorithm can extract the three-dimensional coordinates of 25 key joint points from each frame, including the head, neck, shoulders, elbows, wrists, hips, knees, and ankles. It then retrieves the standard joint points for the corresponding movements from the system's preset standard movement database. The coordinates are calculated by substituting the actual coordinates of each key joint point with the standard coordinates into the Euclidean distance formula to calculate the coordinate deviation. Then, based on the coordinate deviation and the geometric relationship between the joint points, the single-frame attitude deviation angle corresponding to that joint point is calculated using trigonometric functions. The average deviation angle of all joint points is taken as the final single-frame attitude deviation angle of the action in that frame. The inter-frame time interval is obtained by calculating the difference between the timestamps of two adjacent frames. If the timestamp difference is abnormal, the average difference between the timestamps of the preceding and following frames is used for correction. The action frame count starts from the first frame image after the visual acquisition device is started. The counter increments by 1 for each frame image received, updating the action frame count in real time. The cumulative training time is obtained by subtracting the timestamp of the visual acquisition device's startup from the timestamp of the current frame image, in seconds, accurate to two decimal places.
[0032] Traditional technical solutions have the following technical problems: existing posture recognition error calculations lack a basic error coefficient model that integrates physiological state and environmental interference, resulting in inaccurate initial error quantification and failing to provide a reliable basis for subsequent error accumulation and correction compensation.
[0033] Based on this, the attitude recognition basic error coefficient is obtained through the attitude recognition basic error coefficient calculation formula, which is as follows: ; in The amplitude of the real-time electromyographic signal is measured in volts. The reference electromyographic signal amplitude is measured in volts. is a characteristic value of the illumination spectrum, with dimensions of candela per square meter; The standard value for the light spectrum is expressed in candela per square meter. The background dynamic disturbance velocity is expressed in meters per second. , , The preset coordinated adjustment coefficients are all dimensionless. In this technical solution, The electromyography (EMG) sensor in the physiological state sensing module collects data in real time at a frequency consistent with the sensor's sampling frequency of 2000Hz. Each calculation... The average of the most recent 100 sampling points is taken as the average value. The value of should be chosen to avoid the influence of instantaneous signal fluctuations; The baseline electromyographic signal amplitude acquired by the initial calibration module in the trainee's resting state is typically in the range of 0.3-0.8 volts, with the specific value varying depending on individual differences among trainees. The optical sensor of the dynamic environment adaptation module collects data once per second, which is directly used as the input value in the formula. If data is missing during the acquisition process, the previous acquisition value is used. The value is temporarily replaced; The reference value of the light spectrum collected by the initial calibration module in an interference-free environment is typically 400-600 candela per square meter, with the specific value depending on the standard lighting conditions of the training environment. The image processor of the dynamic environment adaptation module calculates once per frame, in meters per second. If there are no moving targets in the background, then... The value is set to 0. If there are multiple moving targets in the background, the maximum value of their velocities is taken as the threshold. ; , , The data from 50 sets of samples collected by the initial calibration module under different trainees and environmental conditions were fitted using a multiple linear regression algorithm. The actual measured posture recognition error was used as the dependent variable during the fitting process. , , Using the least squares method to find the optimal coefficient values for the independent variable, the optimal values are typically obtained. The value is 0.6. The value is 0.3. The value is set to 0.1 to ensure the accuracy of the formula calculation. It can accurately reflect the combined impact of physiological state and environmental interference on posture recognition error; in the formula calculation process, it first calculates separately. and The ratios are then multiplied by... and After summing, the combined influence of physiological and environmental factors is obtained, followed by calculation. Substituting the product into the exponential function The nonlinear influence term of the background interference is obtained, and finally the two terms are multiplied together to obtain the basic error coefficient of attitude recognition. , The value range is usually 1-3. The larger the value, the greater the interference of the current physiological state and environmental conditions on posture recognition.
[0034] Traditional technical solutions have the following technical problems: existing time-accumulated error calculation does not consider the synergistic effect of the basic error coefficient and the time-series motion parameters, and only simply accumulates the error of a single frame, resulting in inaccurate quantification of time-accumulated error and failing to reflect the true accumulation law of motion error over time.
[0035] Based on this, the time series cumulative error value is obtained through the time series cumulative error value calculation formula, which is as follows: ; in The basic error coefficient for attitude recognition is dimensionless. The attitude deviation angle for a single frame is expressed in degrees. The number of action frames is dimensionless. The time interval between frames is measured in seconds. This is the inter-frame reference time interval, measured in seconds. This represents the level of muscle exertion, and is dimensionless. The reference force level is dimensionless. In this technical solution, Using the most recently calculated basic error coefficients for attitude recognition, if in During frame action If a change occurs, then the action corresponding to each frame is taken. Substitute the values into the calculations respectively; The first time sequence action recognition module calculates the first time sequence action recognition module. The single-frame pose deviation angle of a frame action typically ranges from 0 to 10 degrees. If a certain frame's... If the angle exceeds 15 degrees, it is considered an abnormal frame, and the preceding and following frames are used. The average value is replaced; This is the current cumulative number of motion frames, which increases in real time as the motion capture process progresses. Typically, it increases by 25 frames per second of training, corresponding to a frame rate of 25 fps. For the first Frame and the The time interval between frame actions is calculated by the timing action recognition module based on the timestamp difference between two frames. Under normal circumstances... It should be close to the inter-frame reference time interval. ,like and If the deviation exceeds 20%, then... Perform smoothing; The inter-frame reference time interval determined by the initial calibration module corresponds to the inter-frame time of the standard action, and its value is 1 / frame rate, i.e., at a frame rate of 25fps. It takes 0.04 seconds; The first data collected by the physiological state sensing module The muscle exertion level corresponding to the frame action is taken as a value of 1-5 and is directly substituted into the formula as a dimensionless parameter. The baseline force level determined for the initial calibration module is typically level 3, representing the normal force level of a standard movement; during the formula calculation process, the following is calculated first: and The ratio of the two values, multiplied and then the square root is taken to obtain the comprehensive correction coefficient for timing and force application. This coefficient reflects the influence of the speed of movement and the magnitude of force on error accumulation; the faster the movement, the greater the error. The smaller the size, the greater the force exerted. The larger the value, the larger the correction factor; then the correction factor is compared with... , Multiply to get the first... The actual error contribution value of the frame action; finally, from frame 1 to frame 2. The cumulative timing error value is obtained by summing all the error contributions of the frame. , The dimension of is degrees. The larger the value, the more serious the cumulative attitude error, and the stronger the correction force required.
[0036] Traditional technical solutions have the following technical problems: the current calculation of final posture correction compensation amount does not take into account the influence of joint range of motion and cumulative training time, and only determines the compensation amount based on the time-series cumulative error value, which leads to a mismatch between the correction command and the trainee's actual movement ability and training state, affecting the correction effect.
[0037] Based on this, the final attitude correction compensation amount is obtained through the final attitude correction compensation amount calculation formula, which is as follows: ; in This represents the cumulative time-series error, measured in degrees. The basic error coefficient for attitude recognition is dimensionless. The range of motion of a joint is measured in degrees. Standard joint range of motion, dimensionless; The preset time decay coefficient is measured in seconds. The total training time is measured in seconds. In this technical solution, The latest result obtained through the formula for calculating the time series cumulative error value is directly used as the basis for calculating the compensation amount; Adoption and Calculation Consistent attitude recognition baseline error coefficients ensure consistency in the impact of errors; The physiological state sensing module collects the current range of motion data of key joints, taking the average range of motion of major joints such as the shoulder, elbow, and knee. If a certain joint's range of motion is... Less than standard joint range of motion If the joint movement is determined to be restricted to 50% or less, then the joint's range of motion is considered limited, and only that joint's... Perform calculations; The standard joint range of motion determined for the initial calibration module represents the normal range of motion of the trainee's joints, and the range of motion for different joints... Different values are used; for example, the shoulder joint is usually 90 degrees and the elbow joint is 120 degrees. The preset time decay coefficient was obtained by fitting a large amount of training data and is set to 0.02 per second. This coefficient reflects the effect of cumulative training time on the decay of correction strength. The longer the training time, the more obvious the decay, thus avoiding overcorrection caused by long-term training. The cumulative training time recorded by the temporal action recognition module starts from the moment the visual acquisition device is started, in seconds, accurate to one decimal place; in the formula calculation process, first calculate... The ratio, multiplied by Adding 1 gives the joint range of motion correction term, which reflects the impact of joint mobility on the correction compensation. The greater the joint range of motion (the closer it is to the standard value), the larger the correction term, and the corresponding increase in the correction compensation amount; then calculate... The product of these terms, plus 1, is taken as the square root, and the reciprocal is used to obtain the training duration decay term. This decay term decreases as training time increases, preventing excessive correction due to muscle fatigue after prolonged training. Finally, the time-series accumulated error value is calculated. Multiplying the joint range of motion correction term and the training duration decay term by the final posture correction compensation amount yields the final amount. , The dimension of is degrees, and its value directly determines the adjustment range of the posture correction command, ensuring that the correction command can not only make up for the accumulated error, but also conform to the trainee's current joint mobility and training status.
[0038] Traditional technical solutions have the following technical problems: the existing preset collaborative adjustment coefficients lack a scientific way of determination and are mostly set by empirical values, which makes it impossible for the coefficients to accurately match different trainees and training environments, affecting the calculation accuracy of the basic error coefficients of posture recognition.
[0039] Based on this, the preset collaborative adjustment coefficient is obtained by fitting training samples. The initial calibration module collects at least 50 sets of physiological data, environmental data, and action data from different trainees as training samples, and fits the training samples using a multiple linear regression algorithm to determine the coefficient. , , The specific values are then stored in the multi-factor collaborative correction module.
[0040] In this technical solution, when the initial calibration module collects training samples, it selects 50 trainees of different ages, genders, and training levels as sample subjects. Each trainee performs five common sports training movements under three different environmental conditions—normal lighting without interference, strong light interference, and background movement interference—such as squats, bench presses, pull-ups, running, and jumping. Each movement is performed three times, with each performance lasting 10 seconds, resulting in 50 × 3 × 5 × 3 = 2250 sets of raw data. From these, 50 sets of data without abnormalities and with complete data are selected as the final training samples. The physiological data contained in each training sample is the real-time electromyography signal amplitude. and the amplitude of the reference electromyographic signal Environmental data are characteristic values of the light spectrum. Illumination spectral reference values and background dynamic interference speed The motion data are the actual measured posture recognition error values. The error is obtained by manually annotating the deviation between the standard movement and the actual movement; in the modeling process of the multiple linear regression algorithm, the actual measured posture recognition error value is used. As the dependent variable, with , , As the independent variable, first assume The initial value is 0.1; construct a linear regression model. ,in , , The model parameters are solved by the least squares method, which improves the accuracy of the model calculation. With actual measurement The mean square error between them is minimized; after fitting, for , , The values were verified by substituting the verification samples into the formula for calculation. If the error between the calculated value and the actual measured value is less than 5%, the coefficients are considered valid; otherwise, the sample size and algorithm parameters are readjusted for fitting. , , The values are stored in binary format in the Flash memory of the multi-factor collaborative correction module, with storage addresses 0x00010000-0x00010008. Each time the system starts, it reads the coefficient values from this address and loads them into memory for subsequent calculation of the basic error coefficients of attitude recognition.
[0041] Traditional technical solutions have the following technical problems: the output form of existing posture correction instructions is singular, mostly using only visual or auditory instructions, which makes it impossible for trainees to receive correction information quickly and accurately, affecting the timeliness and accuracy of movement adjustment.
[0042] Based on this, the posture correction instructions include visual instructions and auditory instructions. The multi-factor collaborative correction module determines the direction parameters of the visual instructions and the speech parameters of the auditory instructions according to the final posture correction compensation amount. The visual instructions and auditory instructions are output synchronously to guide the trainee to adjust the movement posture.
[0043] In this technical solution, when the multi-factor collaborative correction module determines the visual command direction parameters, it will adjust the compensation amount based on the final posture correction. The size and corresponding deviation of the joint are used to generate direction parameters, such as... If the deviation is 5 degrees and corresponds to a shoulder joint offset, then the direction parameter is set to 5 degrees upward from the shoulder joint. Visual commands are transmitted via HDMI to the display screen in front of the trainee, presented as arrows. The direction of the arrow corresponds to the deviation direction, and the length of the arrow is... The size is proportional to the direction, and the deviation joint and adjustment angle are marked next to the arrow, such as shoulder +5°; when determining the auditory command voice parameters, the module's built-in speech synthesis unit will generate the corresponding speech text according to the direction parameters, such as "Please adjust the shoulder joint upward by 5 degrees", and then convert the text into a speech signal through TTS text-to-speech technology. The speech rate of the speech signal is set to 120 words per minute and the volume is set to 70 decibels to ensure that the trainee can hear clearly; the synchronous output of visual and auditory commands is achieved through a time synchronization mechanism. The module adds the same timestamp when generating both commands. After the display screen and speaker receive the command, they output it simultaneously at the time corresponding to the timestamp, with an output delay difference of less than 100 milliseconds to avoid confusion for the trainee due to command asynchrony; in addition, if If the value exceeds 10 degrees, the module will flash an arrow in the visual instructions and increase the volume to 80 decibels in the auditory instructions to enhance the prompting effect. If the degree is less than 2 degrees, only visual instructions are output to reduce unnecessary auditory interference and ensure that the trainee focuses on the execution of the movement.
[0044] The above are merely preferred embodiments of the present invention and are not intended to limit the present invention in any other way. Any person skilled in the art may make changes or modifications to the above-disclosed technical content to create equivalent embodiments that can be applied to other fields. However, any simple modifications, equivalent changes, and modifications made to the above embodiments based on the technical essence of the present invention without departing from the scope of the present invention shall still fall within the protection scope of the present invention.
Claims
1. A computer vision-based sports training movement posture recognition and correction system, characterized in that, The system includes a physiological state perception module, a dynamic environment adaptation module, a temporal action recognition module, a multi-factor collaborative correction module, and an initial calibration module. These modules establish data transmission connections with the multi-factor collaborative correction module. The initial calibration module acquires baseline physiological parameters and baseline motion parameters for standard movements in a non-fatigue state, and sets baseline environmental parameters for the standard training environment. The physiological state perception module collects the trainee's real-time electromyographic signal amplitude, muscle force level, and joint range of motion. The dynamic environment adaptation module collects the illumination spectrum characteristics and background dynamic interference speed of the training environment. The temporal action recognition module extracts the trainee's single-frame posture deviation angle, inter-frame time interval, number of action frames, and cumulative training time based on computer vision. The multi-factor collaborative correction module receives the aforementioned real-time data and baseline parameters, and sequentially calculates the basic posture recognition error coefficient, temporal cumulative error value, and final posture correction compensation amount through three progressively linked mathematical models, outputting a posture correction command.
2. The computer vision-based sports training movement posture recognition and correction system according to claim 1, characterized in that, The baseline physiological parameters include baseline electromyographic signal amplitude and standard joint range of motion; the baseline motion parameters include inter-frame baseline time interval and baseline force level; and the baseline environmental parameters include baseline illumination spectrum value. The initial calibration module determines the baseline electromyographic signal amplitude and standard joint range of motion by collecting physiological data of the trainee in a static state, determines the inter-frame baseline time interval and baseline force level by analyzing the time sequence data of standard movements, and determines the baseline illumination spectrum value by detecting illumination data of an interference-free training environment.
3. The computer vision-based sports training movement posture recognition and correction system according to claim 1, characterized in that, The physiological state sensing module includes an electromyography (EMG) sensor and a joint range of motion sensor. The EMG sensor is fitted to the trainee's muscle groups to collect real-time EMG signal amplitude and convert it into muscle force level. The joint range of motion sensor is set at the trainee's key joints to collect joint range of motion.
4. The computer vision-based sports training movement posture recognition and correction system according to claim 1, characterized in that, The dynamic environment adaptation module includes an optical sensor and an image processor. The optical sensor is oriented towards the training environment to collect spectral feature values of illumination. The image processor receives the environmental image synchronously collected by the optical sensor, separates the background area from the trainee area through an image segmentation algorithm, and calculates the moving speed of the moving target in the background area as the background dynamic interference speed.
5. The computer vision-based sports training movement posture recognition and correction system according to claim 1, characterized in that, The temporal motion recognition module includes a visual acquisition device and a joint point analysis unit. The visual acquisition device captures the trainee's motion video in real time and breaks it down into continuous motion frames. The joint point analysis unit extracts the coordinates of the trainee's key joint points in each motion frame, calculates the single-frame posture deviation angle based on the deviation between the key joint point coordinates and the standard joint point coordinates, records the time difference between adjacent motion frames as the inter-frame time interval, accumulates the number of motion frames as the motion frame count, and records the time from the start of the visual acquisition device to the current motion frame as the cumulative training time.
6. The computer vision-based sports training movement posture recognition and correction system according to claim 2, characterized in that, The attitude recognition basic error coefficient is obtained through the attitude recognition basic error coefficient calculation formula, which is as follows: ; in The amplitude of the real-time electromyographic signal is measured in volts. The reference electromyographic signal amplitude is measured in volts. is a characteristic value of the illumination spectrum, with dimensions of candela per square meter; The standard value for the light spectrum is expressed in candela per square meter. The background dynamic disturbance velocity is expressed in meters per second. , , These are preset coordination coefficients, all of which are dimensionless.
7. The computer vision-based sports training movement posture recognition and correction system according to claim 6, characterized in that, The time series cumulative error value is obtained through the time series cumulative error value calculation formula, which is: ; in The basic error coefficient for attitude recognition is dimensionless. The attitude deviation angle for a single frame is expressed in degrees. The number of action frames is dimensionless. The time interval between frames is measured in seconds. This is the inter-frame reference time interval, measured in seconds. This represents the level of muscle exertion, and is dimensionless. The base force level is dimensionless.
8. The computer vision-based sports training movement posture recognition and correction system according to claim 7, characterized in that, The final attitude correction compensation amount is obtained through the final attitude correction compensation amount calculation formula, which is as follows: ; in This represents the cumulative time-series error, measured in degrees. The basic error coefficient for attitude recognition is dimensionless. The range of motion of a joint is measured in degrees. Standard joint range of motion, dimensionless; The preset time decay coefficient is measured in seconds. The total training time is measured in seconds.
9. The computer vision-based sports training movement posture recognition and correction system according to claim 6, characterized in that, The preset coordination adjustment coefficient is obtained by fitting training samples. The initial calibration module collects at least 50 sets of physiological data, environmental data, and movement data from different trainees as training samples, and fits the training samples using a multiple linear regression algorithm to determine the coefficient. , , The specific values are then stored in the multi-factor collaborative correction module.
10. The computer vision-based sports training movement posture recognition and correction system according to claim 1, characterized in that, The posture correction instructions include visual instructions and auditory instructions. The multi-factor collaborative correction module determines the direction parameters of the visual instructions and the speech parameters of the auditory instructions based on the final posture correction compensation amount. The visual instructions and auditory instructions are output synchronously to guide the trainee to adjust their movement posture.