A method and system for detecting and evaluating the state of motion based on machine vision

By capturing motion video streams in real time to obtain light entropy and light threshold, defining infrared and visible spectra, acquiring joint coordinates and facial orientation angles, and constructing a physiological-motor coupling model, this approach solves the problems of fluctuating joint coordinate positioning accuracy and deviation in motion risk assessment in traditional methods, enabling personalized exercise recommendations and closed-loop health monitoring.

CN120853997BActive Publication Date: 2026-05-05江苏华郢智能技术有限公司
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
江苏华郢智能技术有限公司
Filing Date
2025-07-15
Publication Date
2026-05-05

AI Technical Summary

Technical Problem

Traditional machine vision-based motion state detection and assessment methods lack a dynamic weight coupling mechanism between illumination entropy and target biological features during spectral fusion, leading to fluctuations in joint coordinate positioning accuracy. Furthermore, the physiological-motion coupling model cannot drive incremental optimization through real-time user feedback data, causing motion risk assessment results to deviate from the true physiological state.

Method used

By capturing motion video streams in real time, obtaining light entropy and light threshold values, defining infrared and visible spectra, acquiring joint coordinate data and facial orientation angles, and combining physiological parameters and exercise intensity data, a physiological-motor coupling model is constructed to predict the exercise risk index and generate personalized suggestions, and push exercise status reports in real time.

Benefits of technology

It achieves stability in joint positioning accuracy and accuracy in motion risk assessment, provides personalized exercise suggestions, dynamically optimizes exercise safety and training science, and forms a closed-loop exercise health monitoring system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120853997B_ABST
    Figure CN120853997B_ABST
Patent Text Reader

Abstract

This invention discloses a motion state detection and evaluation method and system based on machine vision, belonging to the field of computer vision technology. The method includes: real-time capture of motion video streams to obtain illumination entropy values ​​and illumination thresholds; defining infrared and visible spectra based on the illumination thresholds to obtain joint coordinate data and facial orientation angles; extracting physiological parameters based on the facial orientation angles and simultaneously obtaining motion intensity data by combining the rate of change of joint coordinates; constructing a physiological-motion coupling model and outputting a physiological state matrix based on the physiological parameters and motion intensity data; predicting a motion risk index and assessing the motion risk level based on the physiological state matrix, generating personalized suggestions; and pushing a motion state report to the user terminal in real time based on the motion risk level and personalized suggestions, and recording feedback commands. This invention forms a closed-loop exercise health monitoring system by dynamically optimizing exercise safety and training scientificity, achieving personalized exercise risk management.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of computer vision technology, and in particular to a motion state detection and evaluation method and system based on machine vision. Background Technology

[0002] Machine vision-based motion state detection and assessment technology plays a crucial role in today's information security field, primarily providing the technological foundation for sports and health monitoring through non-contact sensing methods. Integrating cutting-edge fields such as multispectral imaging, biomechanical analysis, and temporal modeling, it is widely applied in smart fitness equipment, rehabilitation training guidance, and competitive sports performance optimization. With the development of wearable devices and edge computing, real-time motion state assessment has become a core module of smart health systems, playing a key role in improving the scientific nature of training and reducing the risk of sports injuries. Current technological systems have established a standardized processing framework from environmental perception to risk assessment, forming a deep cross-integration of computer vision and sports science.

[0003] In the field of machine vision-based motion state detection and assessment, traditional detection and assessment methods lack a dynamic weight coupling mechanism between illumination entropy and target biometric features during spectral fusion, leading to fluctuations in joint coordinate positioning accuracy under complex lighting conditions. Furthermore, the parameters of the physiological-motor coupling model are statically fixed after training, making it impossible to drive incremental optimization through real-time user feedback data. When an individual's motion pattern suddenly becomes abnormal, the risk assessment result deviates from the true physiological state, increasing the risk and necessitating manual calibration to compensate for detection errors. Summary of the Invention

[0004] In view of the aforementioned existing problems, the present invention is proposed.

[0005] Therefore, this invention provides a machine vision-based motion state detection and evaluation method to solve the problems of joint positioning accuracy fluctuations and sudden abnormal adaptation obstacles.

[0006] To solve the above-mentioned technical problems, the present invention provides the following technical solution:

[0007] In a first aspect, the present invention provides a motion state detection and evaluation method based on machine vision, which includes: capturing motion video streams in real time and obtaining illumination entropy values ​​and illumination thresholds;

[0008] Based on the illumination threshold, infrared and visible spectra are defined to obtain joint coordinate data and facial orientation angles;

[0009] Physiological parameters are extracted based on facial orientation angle, and motion intensity data are obtained by combining the rate of change of joint coordinates.

[0010] Construct a physiological-motor coupling model and output a physiological state matrix based on physiological parameters and exercise intensity data;

[0011] Based on the physiological state matrix, predict the exercise risk index and assess the exercise risk level to generate personalized recommendations;

[0012] Based on the exercise risk level and personalized recommendations, the user terminal pushes exercise status reports in real time and records feedback commands.

[0013] As a preferred embodiment of the machine vision-based motion state detection and evaluation method of the present invention, the steps of real-time capture of motion video stream and acquisition of illumination entropy value and illumination threshold are as follows:

[0014] Based on the motion video stream, initial spectral data is obtained, and grayscale image data is obtained using the weighted grayscale conversion method.

[0015] Based on grayscale image data, the sub-block entropy matrix is ​​obtained through image information entropy quantization rules, and the illumination entropy value is generated according to the weighted average formula.

[0016] Based on the illumination entropy value, the moving target region is extracted by the inter-frame difference method, and the illumination threshold is obtained by combining the joint positioning feedback optimization method.

[0017] As a preferred embodiment of the machine vision-based motion state detection and evaluation method of the present invention, the steps of defining infrared and visible spectra based on illumination thresholds and obtaining joint coordinate data and facial orientation angles are as follows:

[0018] By comparing the light entropy value with the light threshold, the image enhancement mode is dynamically adjusted to obtain a joint heatmap; by comparing and analyzing the light entropy value with the light threshold, the infrared spectrum and the visible spectrum are defined, and a spectral mode selection command is output.

[0019] Based on the spectral mode selection instruction, the spectral entropy weight is dynamically obtained through the photophysiological collaborative sensing method, and the joint thermal map region segmentation and spectral differential enhancement are performed according to the motion feature-guided spectral collaborative enhancement method, outputting a dual-spectral enhanced image.

[0020] Based on bispectral enhanced images and spectral entropy weights, joint coordinate data and facial orientation angles are obtained by combining posture-physiological joint feature extraction method with joint heatmaps.

[0021] As a preferred embodiment of the machine vision-based motion state detection and evaluation method of the present invention, the steps for extracting physiological parameters based on facial orientation angle and obtaining motion intensity data by combining the rate of change of joint coordinates are as follows:

[0022] Physiological parameters are obtained based on facial orientation angle by dynamic ROI positioning combined with subpixel-level optical flow tracing.

[0023] Based on joint coordinate data, the rate of change of joint coordinates is obtained through the time difference method. Combined with physiological parameters, exercise intensity data is obtained through a physiological-mechanical coupling algorithm.

[0024] As a preferred embodiment of the machine vision-based motion state detection and evaluation method of the present invention, the steps of constructing a physiological-motor coupling model and outputting a physiological state matrix based on physiological parameters and motion intensity data are as follows:

[0025] By combining the temporal characteristics of physiological parameters with the dynamic features of exercise intensity data, and using a heterogeneous feature extractor, a two-stream neural network architecture is obtained.

[0026] Based on loading physiological parameters into a two-stream neural network architecture, physiological feature vectors are obtained through one-dimensional temporal convolution operations. At the same time, motion intensity data is input into a bidirectional LSTM for temporal modeling to obtain motion feature vectors.

[0027] By concatenating physiological feature vectors and motion feature vectors, a dynamic coupling weight matrix is ​​obtained based on a cross-attention mechanism;

[0028] Based on the dynamic coupling weight matrix, a physiological-motor coupling model is constructed by performing dual-stream feature fusion through feature reweighting.

[0029] A multi-layer neural network was used to train the physiological-motor coupling model. By combining physiological parameters and exercise intensity data, the physiological state matrix was obtained through the trained physiological-motor coupling model.

[0030] As a preferred embodiment of the machine vision-based motion state detection and evaluation method of the present invention, the steps of predicting the motion risk index and evaluating the motion risk level based on the physiological state matrix, and generating personalized suggestions, are as follows:

[0031] Based on the physiological state matrix, the first-level processing of the three-level assessment method uses a gradient boosting decision tree to predict the exercise risk index.

[0032] Based on the sports risk index, the second-level processing of the three-level assessment method uses fuzzy logic decision trees to classify sports risk levels.

[0033] Based on the level of exercise risk, the three-level assessment method uses a rule engine to match personalized solutions, simultaneously integrates user historical data and suggestions, and outputs personalized exercise recommendations.

[0034] As a preferred embodiment of the machine vision-based motion state detection and evaluation method of the present invention, the steps of the user terminal pushing motion state reports in real time and recording feedback commands based on motion risk levels and personalized suggestions are as follows:

[0035] Based on exercise risk levels and personalized exercise suggestions, an augmented reality engine generates exercise status reports and pushes them to the user's terminal in real time.

[0036] After receiving the motion status report, the user performs multimodal interactive operations, records feedback instructions, and synchronizes them to the database for physiological-motor coupling model optimization.

[0037] Secondly, this invention provides a machine vision-based motion state detection and evaluation system, comprising a data acquisition module, a feature extraction module, a parameter tracking module, a model training module, a risk assessment module, and an optimization feedback module. The data acquisition module is used to capture motion video streams in real time and obtain illumination entropy values ​​and illumination thresholds. The feature extraction module is used to define infrared and visible spectra based on the illumination thresholds and obtain joint coordinate data and facial orientation angles. The parameter tracking module is used to extract physiological parameters based on the facial orientation angle and simultaneously obtain motion intensity data by combining the rate of change of joint coordinates. The model training module is used to construct a physiological-motor coupling model and output a physiological state matrix based on the physiological parameters and motion intensity data. The risk assessment module is used to predict a motion risk index and assess the motion risk level based on the physiological state matrix, generating personalized suggestions. The optimization feedback module is used to push motion state reports to the user terminal in real time based on the motion risk level and personalized suggestions, and record feedback commands.

[0038] Thirdly, the present invention provides a computer device including a memory and a processor, wherein the memory stores a computer program, wherein when the computer program is executed by the processor, it implements any step of the motion state detection and evaluation method based on machine vision as described in the first aspect of the present invention.

[0039] Fourthly, the present invention provides a computer-readable storage medium having a computer program stored thereon, wherein: when the computer program is executed by a processor, it implements any step of the motion state detection and evaluation method based on machine vision as described in the first aspect of the present invention.

[0040] The beneficial effects of this invention are as follows: by constructing a physiological-motor coupling model, it provides a high-dimensional biomechanical state representation for exercise risk assessment, accurately quantifies fatigue index and movement accuracy; through a three-level assessment method, it achieves multi-level refined decision-making on exercise risk, generates customized exercise suggestions for users, dynamically optimizes exercise safety and training scientificity, forms a closed-loop exercise health monitoring system, and realizes personalized exercise risk management. Attached Figure Description

[0041] To more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings used in the following description of the embodiments will be briefly introduced. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0042] Figure 1 This is a flowchart of a motion state detection and evaluation method based on machine vision.

[0043] Figure 2 This is a schematic diagram of a machine vision-based motion state detection and evaluation system.

[0044] Figure 3 This is a flowchart for calculating the light entropy value and obtaining the light threshold.

[0045] Figure 4 A flowchart for constructing a physiological-motor coupling model. Detailed Implementation

[0046] To make the above-mentioned objects, features and advantages of the present invention more apparent and understandable, the specific embodiments of the present invention will be described in detail below with reference to the accompanying drawings.

[0047] Many specific details are set forth in the following description in order to provide a full understanding of the invention. However, the invention may also be practiced in other ways different from those described herein, and those skilled in the art can make similar extensions without departing from the spirit of the invention. Therefore, the invention is not limited to the specific embodiments disclosed below.

[0048] Secondly, the term "one embodiment" or "embodiment" as used herein refers to a specific feature, structure, or characteristic that may be included in at least one implementation of the present invention. The phrase "in one embodiment" appearing in different places in this specification does not necessarily refer to the same embodiment, nor is it a single or selective embodiment that is mutually exclusive with other embodiments.

[0049] Reference Figures 1-4 As one embodiment of the present invention, this embodiment provides a motion state detection and evaluation method based on machine vision, including the following steps:

[0050] S1. Real-time capture of motion video streams to obtain illumination entropy and illumination threshold;

[0051] Based on the motion video stream, initial spectral data is obtained, and grayscale image data is obtained using the weighted grayscale conversion method.

[0052] Furthermore, by synchronously acquiring motion video streams using a dual-spectrum camera array and implementing cross-spectral frame alignment based on hardware timing calibration, initial spectral data containing infrared and visible spectra are obtained from the video. The visible spectrum segment uses a high-frame-rate CMOS sensor to capture RGB three-channel data, while the infrared spectrum segment uses a low-light-enhanced sensor to acquire thermal radiation intensity data. Nonlinear mapping transformation is performed on the infrared thermal radiation intensity data to generate an infrared thermogram. Simultaneously, a weighted grayscale conversion method is applied to the visible spectrum, assigning corresponding visual perception weight coefficients to the RGB three-channel data components of each pixel to generate visible light grayscale components and a visible light image. The infrared thermogram is then bilinearly interpolated to a fixed resolution (e.g., 1080p), and then weighted and fused with the visible light grayscale components to obtain a single-channel grayscale image data sequence.

[0053] Based on grayscale image data, the sub-block entropy matrix is ​​obtained through image information entropy quantization rules, and the illumination entropy value is generated according to the weighted average formula.

[0054] It should be noted that the image information entropy quantization rule evaluates the impact of the lighting environment on information carrying capacity by analyzing the dispersion of grayscale distribution in local areas of an image. The rule divides the grayscale image into fixed-size sub-blocks (e.g., 8×8 pixels), reads the grayscale values ​​of the pixels in each sub-block, discretizes the resulting integer levels to obtain grayscale levels, and counts the frequency of each grayscale level within each sub-block. Based on information theory principles, the entropy value of that region is obtained. The more uniform the grayscale distribution (e.g., a single background), the lower the entropy value; the more drastic the grayscale changes (e.g., strong shadows), the higher the entropy value. The entropy values ​​of all sub-blocks form an entropy matrix, reflecting the lighting disorder in different areas of the image, providing a quantitative basis for the joint localization feedback optimization method.

[0055] Furthermore, based on grayscale image data (taking 1920×1080 resolution as an example), the image is divided into fixed-size sub-blocks (e.g., 8×8 pixels) according to image information entropy quantization rules. The frequency of each grayscale level and the total number of sub-blocks are counted to obtain the information entropy value of the corresponding sub-block, thus obtaining the sub-block entropy value matrix. Subsequently, a weighted average formula is applied to generate a global illumination entropy value. Taking the image center coordinates as the origin, the Euclidean distance of each sub-block center is obtained (e.g., edge sub-blocks are 960 pixels from the center). Based on the distance decay principle, an exponential decay function is obtained for weight allocation. Finally, the entropy values ​​of all sub-blocks are weighted and averaged to output the illumination entropy value, thereby quantifying the degree of ambient lighting disorder: low entropy values ​​(e.g., <3.0) indicate uniform lighting environments (e.g., indoor constant light scenes), and high entropy values ​​(e.g., >5.4) indicate strongly interfering scenes (e.g., outdoor tree shadows or backlighting). The weighted average formula is:

[0056]

[0057] Where H is the global illumination entropy value, N is the total number of sub-blocks, and w(d i H represents the weight value of the current data point. i Let d be the illumination entropy value of the i-th sub-block. i Let be the Euclidean distance from the center point of the i-th sub-block to the center of the image.

[0058] Based on the illumination entropy value, the moving target region is extracted by the inter-frame difference method, and the illumination threshold is obtained by combining the joint positioning feedback optimization method.

[0059] It should be noted that the inter-frame difference sensitivity parameter, which is generated based on the illumination entropy value and is corrected in real time by joint positioning feedback, is set as the dynamic threshold. The moving target region refers to the connected region extracted from consecutive video frames by the inter-frame difference method, in which pixel changes are caused by human movement. Specifically, the pixel difference between two consecutive grayscale images is obtained, and the region with a pixel difference value greater than the dynamic threshold is marked as a moving pixel. After morphological processing (dilation + erosion), the marked region generates a connected moving target region.

[0060] Furthermore, a dynamic threshold for the inter-frame difference method is set based on the illumination entropy value. By extracting the moving target region (the connected region where human movement causes pixel changes, such as the outline of a runner occupying 35% of the image) and combining it with joint coordinate offset feedback (such as increasing the dynamic threshold by 0.5 for every 10 pixels when the knee joint is offset by 32 pixels), the offset distance is obtained. The illumination threshold is defined by the linear correction relationship between the offset distance and the dynamic threshold.

[0061] S2. Based on the illumination threshold, define the infrared spectrum and the visible spectrum, and obtain joint coordinate data and facial orientation angle;

[0062] By comparing the light entropy value with the light threshold, and based on the infrared feature selective enhancement method, the image enhancement mode is dynamically adjusted to obtain the joint heat map.

[0063] Furthermore, by combining the light entropy value and the light threshold, the infrared thermal map generation process is dynamically adjusted through infrared feature selective enhancement. When the light entropy value is higher than the light threshold, the infrared spectrum-dominated mode is activated to perform logarithmic enhancement processing on the infrared thermal radiation intensity data, thereby improving the abnormal contrast of temperature differences in the joint region. When the light entropy value is lower than the light threshold, the visible spectrum-dominated mode is adopted to perform local nonlinear contrast enhancement on the joint hot areas, while maintaining the basic radiation mapping in non-hot areas, thus achieving differentiated enhancement of thermal features. The joint thermal map is output by weighted fusion of spatial coordinates, and the pixel values ​​in the joint thermal map map map the temperature distribution on the joint surface, forming a visual representation of biomechanical features.

[0064] By comparing and analyzing the light entropy value and the light threshold, the infrared spectrum and the visible spectrum are defined, and the spectral mode selection command is output.

[0065] Furthermore, based on the quantitative comparison analysis of light entropy and light threshold, when the light entropy is higher than the light threshold, it is determined that there is strong backlight or shadow interference in the environment, triggering the infrared spectrum dominance mode, increasing the weight of the infrared spectrum and decreasing the weight of the visible spectrum, so as to penetrate optical noise with thermal radiation characteristics; when the entropy is lower than the light threshold, the visible spectrum dominance mode is activated, increasing the weight of the visible spectrum and decreasing the weight of the infrared spectrum, preserving the advantage of texture details, and outputting the spectral mode selection command through physiological signal confidence verification and motion state analysis.

[0066] Based on the spectral mode selection instruction, the spectral entropy weight is dynamically obtained through the photophysiological collaborative sensing method, and the joint thermal map region segmentation and spectral differential enhancement are performed according to the motion feature-guided spectral collaborative enhancement method, outputting a dual-spectral enhanced image.

[0067] It should be noted that the photophysiological co-sensing method uses physiological-optical dual-source reliability assessment to dynamically optimize spectral fusion weights, drive joint thermogram enhancement guided by motion features, and finally generate a dual-spectral image with both thermodynamic accuracy and texture detail. Its core value lies in overcoming the limitations of environmental interference on the extraction of biomechanical features.

[0068] Furthermore, based on the spectral mode selection instruction, the spectral entropy weights are dynamically obtained through photophysiological collaborative sensing. At the same time, combined with the motion feature-guided spectral collaborative enhancement method, the joint thermogram is segmented by morphological gradient to extract the high-temperature joint region. In the segmented region, the infrared thermogram is enhanced by unsharpening masking, and the visible light image is subjected to adaptive histogram equalization. In the non-joint region, the visible light image is contrast stretched, and the infrared image is enhanced while maintaining the baseline. Finally, the enhancement results of the two bands are weighted and fused based on the spectral entropy weights to generate a bispectral enhanced image.

[0069] Based on bispectral enhanced images and spectral entropy weights, joint coordinate data and facial orientation angles are obtained by combining posture-physiological joint feature extraction method with joint heatmaps.

[0070] It should be noted that the posture-physiology joint feature extraction method addresses the limitations of traditional methods in separating and extracting posture features (joint movements) and physiological features (facial micro-expressions). It improves the accuracy of biomechanical state representation through multi-source data collaboration, corrects joint movement trajectories by utilizing changes in facial hemodynamics (such as microvascular pulsation in the cheek), and filters facial optical flow noise based on joint coordinate stability (such as knee joint movement smoothness > 0.9), thereby improving the signal-to-noise ratio of physiological parameters.

[0071] Furthermore, based on bispectral enhanced images and spectral entropy weights, joint coordinate extraction is performed by combining posture-physiology joint feature extraction with joint heatmaps. Temperature peak coordinates are located on the joint heatmaps. Subpixel-level joint coordinates are output by spatial registration of infrared and visible light features of bispectral images and motion trajectory smoothing constraints. At the same time, bispectral texture gradient features are extracted in the facial ROI region, and facial orientation angles are synthesized by weighted clustering of optical flow vectors.

[0072] S3. Extract physiological parameters based on facial orientation angle, and obtain motion intensity data by combining the rate of change of joint coordinates;

[0073] Physiological parameters are obtained based on facial orientation angle by dynamic ROI positioning combined with subpixel-level optical flow tracing.

[0074] It should be noted that physiological parameters refer to three types of biological indicators obtained through facial microvascular motion feature analysis: heart rate variability (a time-domain / frequency-domain indicator reflecting the autonomic nervous system's ability to regulate heart rhythm), respiratory rate (number of respiratory cycles per unit time), and blood oxygen saturation trend (the relative change trend of the proportion of oxyhemoglobin in the blood).

[0075] Furthermore, based on the facial orientation angle, the detection area (ROI) is dynamically adjusted according to the facial orientation angle. For example, when the yaw angle is 12.3°, a 200×200 pixel tilted ROI area is generated with the nose tip coordinates (320, 240) as the center (tilt angle is synchronized with yaw angle) to focus on microvascular sensitive areas such as the nasal alae and cheekbones. An improved Lucas-Kanade optical flow algorithm is applied within the ROI to track the displacement of feature points on both sides of the nasal alae. Physiological parameters are extracted through optical flow vector temporal analysis, including the output of heart rate variability from optical flow amplitude spectrum analysis in the nasal alae region, the analysis of respiratory rate from optical flow periodic fluctuations in the cheekbone region, and the estimation of blood oxygen saturation trend from the red light / infrared light reflectance ratio in the nasal tip region.

[0076] Based on joint coordinate data, the rate of change of joint coordinates is obtained through the time difference method. Combined with physiological parameters, exercise intensity data is obtained through a physiological-mechanical coupling algorithm.

[0077] It should be noted that the physiological-mechanical coupling algorithm addresses the shortcomings of traditional exercise assessment, which involves the separate analysis of mechanical data (such as joint velocity) and physiological data (such as heart rate variability), by integrating joint kinematic characteristics (mechanical load) and physiological response indicators (metabolic load). The aim is to establish a dynamic mapping model of mechanical-physiological parameters, output exercise intensity data, and quantify the comprehensive biomechanical load of the human body during exercise. The exercise intensity data is a quantitative value of exercise load generated by integrating joint kinematic characteristics and physiological response indicators, which is a comprehensive indicator reflecting the energy consumption and biomechanical load of the human body during exercise.

[0078] Furthermore, based on joint coordinate data, the rate of change of joint displacement was obtained using the time difference method. Taking a 30fps video frame interval of 0.033 seconds as an example, the displacement difference between frame 1 and frame 2 was 5.83 pixels, with an instantaneous velocity of 176.7 pixels / second; the average joint velocity was 180 pixels / second (after motion trajectory smoothing). Simultaneously, physiological parameters were combined, including a heart rate variability of 56 milliseconds (reflecting the state of autonomic nervous system regulation) and a respiratory rate of 15 breaths / minute (characterizing metabolic intensity). Multi-source data fusion was performed using a physiological-mechanical coupling algorithm, weighting joint motion velocity (mechanical load) and physiological response (metabolic load) to synthesize motion intensity data.

[0079] S4. Construct a physiological-motor coupling model and output a physiological state matrix based on physiological parameters and exercise intensity data;

[0080] By combining the temporal characteristics of physiological parameters with the dynamic features of exercise intensity data, and using a heterogeneous feature extractor, a two-stream neural network architecture is obtained.

[0081] It should be noted that temporal characteristics refer to the periodicity and spectral distribution characteristics of physiological parameters (heart rate variability and respiratory rate), which are derived from the sub-pixel optical flow tracking results of facial microvascular pulsation; dynamic characteristics refer to the spatial variation pattern of motion intensity data (joint acceleration and trajectory curvature), which are derived from the kinematic parameters derived by the joint coordinate time difference method.

[0082] Furthermore, the temporal characteristics of physiological parameters are extracted into physiological feature vectors through a one-dimensional temporal convolutional layer; the dynamic characteristics of motion intensity data (such as the 30 pixel / second acceleration generated by the sudden increase of knee joint speed from 180 pixels / second to 240 pixels / second, and the sharp turn sign with a radius of curvature of motion trajectory <50 pixels) are derived from the kinematic parameters derived by the joint coordinate temporal difference method (such as the MET value of 0.95 corresponding to the peak acceleration). The motion feature vectors are extracted by a bidirectional long short-term memory neural network (LSTM). The physiological feature vectors and motion feature vectors are concatenated and then input into a cross-attention mechanism to dynamically allocate weights (such as 0.6 for physiological features and 0.4 for motion features) to generate a coupling matrix of a two-stream neural network architecture. Cross-modal information interaction is achieved through the feature fusion gating mechanism of a multi-source heterogeneous feature extractor, and residual skip connections are used to compensate for information loss, thus constructing a two-stream neural network architecture.

[0083] Based on loading physiological parameters into a two-stream neural network architecture, physiological feature vectors are obtained through one-dimensional temporal convolution operations. At the same time, motion intensity data is input into a bidirectional LSTM for temporal modeling to obtain motion feature vectors.

[0084] It should be noted that one-dimensional temporal convolution is a convolutional neural network layer specifically designed for processing time series data. It extracts local temporal features by sliding convolution kernels, adopts an end-to-end joint training mode, and is based on multiple physiological time series of motion scenes. It minimizes cross-entropy loss and temporal consistency regularization terms through the Adam optimizer, dynamically optimizes the convolution kernel weights, and enables the network to adaptively extract frequency domain features of physiological signals.

[0085] Furthermore, physiological parameters are loaded into the physiological branch of the two-stream neural network architecture. Through one-dimensional temporal convolution operation, features of the frequency band fluctuations of heart rate variability (HRV) and respiratory rate are extracted and output as physiological feature vectors. Simultaneously, exercise intensity data is input into the bidirectional LSTM of the exercise branch. Through temporal modeling, dynamic mutation features are captured and output as 64-dimensional motion feature vectors.

[0086] By concatenating physiological feature vectors and motion feature vectors, a dynamic coupling weight matrix is ​​obtained based on a cross-attention mechanism;

[0087] Furthermore, by concatenating physiological feature vectors and motion feature vectors to form a fused feature tensor, a dynamic coupling weight matrix is ​​calculated based on a cross-attention mechanism. A fully connected layer maps the two types of features to the same dimensional space, obtaining the query vector projection of the physiological feature vector and the key vector projection of the motion feature vector. The attention score is then calculated using the following formula:

[0088]

[0089] Where A is the attention score, Q is the query vector projection of the physiological feature vector, K is the key vector projection of the motion feature vector, and k is the dimension scaling factor of the feature projection space.

[0090] The dynamic coupling weight matrix is ​​generated by dynamically assigning weights based on the attention score (e.g., 0.6 for physiological features and 0.4 for motor features).

[0091] Based on the dynamic coupling weight matrix, a physiological-motor coupling model is constructed by performing dual-stream feature fusion through feature reweighting.

[0092] Furthermore, based on the dynamic coupling weight matrix, physiological weight scaling is applied to the physiological feature vector to generate weighted physiological features, and physiological weight scaling is applied to the motion feature vector to generate weighted motion features. The weighted physiological features and weighted motion features are cascaded to form a fusion tensor, which is input into a fully connected layer to generate a high-dimensional fused physiological feature vector and motion feature vector, quantifying the physiological-motor coupling state. Finally, the state matrix of the physiological-motor coupling model is output through a three-layer neural network to construct the physiological-motor coupling model.

[0093] A multi-layer neural network was used to train the physiological-motor coupling model. By combining physiological parameters and exercise intensity data, the physiological state matrix was obtained through the trained physiological-motor coupling model.

[0094] It should be noted that a multilayer neural network refers to a hierarchical computational structure that includes an input layer, at least one hidden layer, and an output layer. It achieves a complex mapping from high-dimensional features to the target output through the cascaded transformation of nonlinear activation functions and weight matrices.

[0095] Furthermore, a multi-layer neural network is used to train the physiological-motor coupling model to generate an optimized model parameter set. Simultaneously, based on physiological parameters and motion intensity data, the current frame offset is predicted according to the continuous knee joint trajectory using a joint motion compensation method. Coordinate compensation is performed using a linear regression model trained with historical joint coordinate sequences. After compensation, the coordinates are normalized to zero mean and unit variance to form a standardized tensor. Combined with the optimized model parameter set, feature enhancement is performed using a dynamic sharpening feature decoding method. The sharpening operator of the convolution kernel is used to improve the feature edge response. Matrix multiplication is performed in combination with the parameter set to obtain the physiological state matrix.

[0096] S5. Based on the physiological state matrix, predict the exercise risk index and assess the exercise risk level, and generate personalized suggestions.

[0097] Based on the physiological state matrix, the first-level processing of the three-level assessment method uses a gradient boosting decision tree to predict the exercise risk index.

[0098] Furthermore, based on the physiological state matrix, the first stage of the three-level assessment method is performed. The dimensional features of the physiological state matrix are input, and the splitting threshold is obtained according to the critical parameters controlling node splitting in the Gradient Boosting Decision Tree (GBDT). Weighted voting is performed through the GBDT model composed of decision trees, assigning weights to the fatigue index (high fatigue significantly increases risk) and metabolic load (e.g., risk surges when the value > 0.9). The splitting threshold is dynamically adjusted based on user baseline data (e.g., age 30, BMI 22). After multiple rounds of residual fitting (e.g., 300 iterations), the exercise risk index is obtained. Based on the exercise risk index, the second stage of the three-level assessment method uses a fuzzy logic decision tree to classify exercise risk levels.

[0099] Furthermore, based on the sports risk index, a second-level processing step of the three-level assessment method is performed. First, the risk index is fuzzified, and a risk interval is defined using a triangular membership function. A risk threshold value is obtained (e.g., a risk index of 0.92 defines high risk, with a risk threshold value of 0.9). The membership function parameters are dynamically adjusted based on the user's physiological baseline (e.g., age 30, BMI 22) (e.g., the risk threshold value is reduced to 0.65 when the age is >40). The fuzzified membership vector is input into the decision tree rule base, and the risk level value is obtained by defuzzification using the centroid method. Finally, discrete risk levels (e.g., low risk 1, medium risk 2, high risk 3) are output, and corresponding early warning strategies are triggered.

[0100] Based on the level of exercise risk, the three-level assessment method uses a rule engine to match personalized solutions, simultaneously integrates user historical data and suggestions, and outputs personalized exercise recommendations.

[0101] It should be noted that the user's historical data suggestions are derived from the user's long-term sports and health records (including exercise habits, physiological response patterns, injury records and rehabilitation plans), and are continuously synchronized to the local encrypted database through a sports and health data service platform (such as Huawei Sports and Health SDK).

[0102] Furthermore, based on the exercise risk level (e.g., high risk level 3), a three-level assessment method is applied. Personalized solutions are matched through a rule engine, and suggestions are retrieved from the user's historical data (e.g., past sports injury records, recovery cycle preferences, and risk aversion tendencies). Combined with real-time physiological-motor coupling status (e.g., fatigue index 0.87, joint stability 0.75), corresponding strategies in the rule base are activated. Personalized parameters from the user's historical data are simultaneously integrated (e.g., the user refuses squats due to a previous meniscus injury), and the details of the solution are dynamically adjusted (e.g., replacing squats with straight leg raises) to obtain structured exercise suggestions.

[0103] S6. Based on the exercise risk level and personalized suggestions, the user terminal pushes exercise status reports in real time and records feedback commands.

[0104] Based on exercise risk levels and personalized exercise suggestions, an augmented reality engine generates exercise status reports and pushes them to the user's terminal in real time.

[0105] Furthermore, based on the exercise risk level (e.g., high risk level 3) and personalized exercise suggestions, the exercise risk level and suggestion content are analyzed through the augmented reality engine (AR Engine), driving the dynamic rendering of the three-dimensional skeletal model, simultaneously integrating real-time physiological-motor data, and combining the user's historical exercise patterns (e.g., records of patellar pain attacks within 12 months) to generate a risk trend curve. Finally, the report is projected to the user's terminal through the spatial anchoring method.

[0106] After receiving the motion status report, the user performs multimodal interactive operations, records feedback instructions, and synchronizes them to the database for physiological-motor coupling model optimization.

[0107] Furthermore, after receiving the exercise status report, the user can perform multimodal interactive operations, such as the voice command "reduce intensity", the gesture of swiping down to close the AR interface, and the watch touch click "accept suggestion". The feedback instructions are encapsulated in JSON format and synchronized to the cloud database in real time. Combined with the user's historical data, the physiological-motor coupling model is incrementally optimized by dynamically lowering the risk threshold and increasing the joint stability weight coefficient.

[0108] This embodiment also provides a machine vision-based motion state detection and evaluation system, including: a data acquisition module, a feature extraction module, a parameter tracking module, a model training module, a risk assessment module, and an optimization feedback module; the data acquisition module is used to capture motion video streams in real time and obtain illumination entropy values ​​and illumination thresholds; the feature extraction module is used to define infrared and visible spectra based on illumination thresholds and obtain joint coordinate data and facial orientation angles; the parameter tracking module is used to extract physiological parameters based on facial orientation angles and obtain motion intensity data by combining the rate of change of joint coordinates; the model training module is used to construct a physiological-motor coupling model and output a physiological state matrix based on physiological parameters and motion intensity data; the risk assessment module is used to predict the motion risk index and assess the motion risk level based on the physiological state matrix and generate personalized suggestions; the optimization feedback module is used to push motion state reports to the user terminal in real time based on the motion risk level and personalized suggestions, and record feedback instructions.

[0109] This embodiment also provides a computer device applicable to the motion state detection and evaluation method based on machine vision, including: a memory and a processor; the memory is used to store computer-executable instructions, and the processor is used to execute the computer-executable instructions to implement the motion state detection and evaluation method based on machine vision as proposed in the above embodiment.

[0110] The computer device can be a terminal, comprising a processor, memory, communication interface, display screen, and input devices connected via a system bus. The processor provides computing and control capabilities. The memory includes non-volatile storage media and internal memory. The non-volatile storage media stores the operating system and computer programs. The internal memory provides an environment for the operation of the operating system and computer programs stored in the non-volatile storage media. The communication interface is used for wired or wireless communication with external terminals; wireless communication can be achieved through Wi-Fi, carrier networks, NFC (Near Field Communication), or other technologies. The display screen can be an LCD screen or an e-ink screen. The input devices can be a touch layer covering the display screen, buttons, a trackball, or a touchpad on the computer device's casing, or an external keyboard, touchpad, or mouse.

[0111] This embodiment also provides a storage medium storing a computer program, which, when executed by a processor, implements the motion state detection and evaluation method based on machine vision as proposed in the above embodiments. The storage medium can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as Static Random Access Memory (SRAM), Electrically Erasable Programmable Read-Only Memory (EEPROM), Erasable Programmable Read Only Memory (EPROM), Programmable Red-Only Memory (PROM), Read-Only Memory (ROM), magnetic storage, flash memory, magnetic disk, or optical disk.

[0112] In summary, this invention achieves personalized sports risk management by: constructing a physiological-motor coupling model to provide a high-dimensional biomechanical state representation for sports risk assessment, accurately quantifying fatigue index and movement accuracy; using a three-level assessment method to achieve multi-level refined decision-making on sports risk, generating customized sports suggestions for users, dynamically optimizing sports safety and training scientificity, forming a closed-loop sports health monitoring system, and realizing personalized sports risk management.

[0113] It should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit it. Although the present invention has been described in detail with reference to preferred embodiments, those skilled in the art should understand that modifications or equivalent substitutions can be made to the technical solutions of the present invention without departing from the spirit and scope of the technical solutions of the present invention, and all such modifications or substitutions should be covered within the scope of the claims of the present invention.

Claims

1. A motion state detection and evaluation method based on machine vision, characterized in that: include, Real-time capture of motion video streams, obtaining illumination entropy values ​​and illumination thresholds; Based on the illumination threshold, infrared and visible spectra are defined to obtain joint coordinate data and facial orientation angles; Physiological parameters are extracted based on facial orientation angle, and motion intensity data are obtained by combining the rate of change of joint coordinates. Construct a physiological-motor coupling model and output a physiological state matrix based on physiological parameters and exercise intensity data; Based on the physiological state matrix, predict the exercise risk index and assess the exercise risk level to generate personalized recommendations; Based on the exercise risk level and personalized suggestions, the user terminal pushes exercise status reports in real time and records feedback commands; The steps for defining infrared and visible spectra based on illumination thresholds, and obtaining joint coordinate data and facial orientation angles are as follows: By comparing the light entropy value with the light threshold, the image enhancement mode is dynamically adjusted to obtain a joint heatmap. By comparing and analyzing the light entropy value and the light threshold, the infrared spectrum and the visible spectrum are defined, and the spectral mode selection command is output. Based on the spectral mode selection command, the spectral entropy weight is dynamically obtained through physiological and optical dual-source reliability assessment, and joint heatmap region segmentation and spectral differential enhancement are performed to output a dual-spectral enhanced image. Based on bispectral enhanced images and spectral entropy weights, joint coordinate data and facial orientation angles are obtained by combining posture-physiological joint feature extraction method with joint heatmaps. The steps for constructing the physiological-motor coupling model, which outputs a physiological state matrix based on physiological parameters and exercise intensity data, are as follows: By combining the temporal characteristics of physiological parameters with the dynamic features of exercise intensity data, and using a heterogeneous feature extractor, a two-stream neural network architecture is obtained. Based on loading physiological parameters into a two-stream neural network architecture, physiological feature vectors are obtained through one-dimensional temporal convolution operations. At the same time, motion intensity data is input into a bidirectional LSTM for temporal modeling to obtain motion feature vectors. By concatenating physiological feature vectors and motion feature vectors, a dynamic coupling weight matrix is ​​obtained based on a cross-attention mechanism; Based on the dynamic coupling weight matrix, dual-stream features are fused to construct a physiological-motor coupling model; A multi-layer neural network was used to train the physiological-motor coupling model. By combining physiological parameters and exercise intensity data, the physiological state matrix was obtained through the trained physiological-motor coupling model.

2. The motion state detection and evaluation method based on machine vision as described in claim 1, characterized in that: The steps for real-time capture of motion video streams and acquisition of illumination entropy and illumination threshold are as follows: Based on the motion video stream, initial spectral data is obtained, and grayscale image data is obtained using the weighted grayscale conversion method. Based on grayscale image data, obtain the sub-block entropy matrix, and generate illumination entropy values ​​according to the weighted average formula; Based on the light entropy value, the moving target region is extracted, and the light threshold is obtained by combining the joint positioning feedback optimization method.

3. The motion state detection and evaluation method based on machine vision as described in claim 1, characterized in that: The steps for extracting physiological parameters based on facial orientation angle and obtaining motion intensity data by combining the rate of change of joint coordinates are as follows. Physiological parameters are obtained based on facial orientation angle by dynamic ROI positioning combined with subpixel-level optical flow tracing. Based on joint coordinate data, the rate of change of joint coordinates is obtained. Combined with physiological parameters, exercise intensity data is obtained through a physiological-mechanical coupling algorithm.

4. The motion state detection and evaluation method based on machine vision as described in claim 1, characterized in that: The steps for predicting the exercise risk index and assessing the exercise risk level based on the physiological state matrix, and generating personalized suggestions, are as follows: Based on the physiological state matrix, the first-level processing of the three-level assessment method uses a gradient boosting decision tree to predict the exercise risk index. Based on the sports risk index, the second-level processing of the three-level assessment method uses a fuzzy logic decision tree to classify sports risk levels. Based on the level of exercise risk, the three-level assessment method uses a rule engine to match personalized solutions, simultaneously integrates user historical data and suggestions, and outputs personalized exercise recommendations.

5. The motion state detection and evaluation method based on machine vision as described in claim 4, characterized in that: Based on the exercise risk level and personalized suggestions, the user terminal pushes exercise status reports in real time and records feedback commands. The steps are as follows: Based on exercise risk levels and personalized exercise suggestions, an augmented reality engine generates exercise status reports and pushes them to the user's terminal in real time. After receiving the motion status report, the user performs multimodal interactive operations, records feedback instructions, and synchronizes them to the database for physiological-motor coupling model optimization.

6. A machine vision-based motion state detection and evaluation system, based on the machine vision-based motion state detection and evaluation method according to any one of claims 1 to 5, characterized in that: include, The data acquisition module is used to capture motion video streams in real time and obtain illumination entropy values ​​and illumination thresholds. The feature extraction module is used to define infrared and visible spectra based on illumination thresholds to obtain joint coordinate data and facial orientation angles. The parameter tracking module is used to extract physiological parameters based on the facial orientation angle, and at the same time, combine the rate of change of joint coordinate data to obtain exercise intensity data; The model training module is used to build a physiological-motor coupling model and output a physiological state matrix based on physiological parameters and exercise intensity data. The risk assessment module is used to predict the exercise risk index and assess the exercise risk level based on the physiological state matrix, and generate personalized suggestions. The feedback module has been optimized to push real-time exercise status reports to user terminals based on exercise risk levels and personalized suggestions, and to record feedback commands. The steps for defining infrared and visible spectra based on illumination thresholds, and obtaining joint coordinate data and facial orientation angles are as follows: By comparing the light entropy value with the light threshold, the image enhancement mode is dynamically adjusted to obtain a joint heatmap. By comparing and analyzing the light entropy value and the light threshold, the infrared spectrum and the visible spectrum are defined, and the spectral mode selection command is output. Based on the spectral mode selection command, the spectral entropy weight is dynamically obtained through physiological and optical dual-source reliability assessment, and joint heatmap region segmentation and spectral differential enhancement are performed to output a dual-spectral enhanced image. Based on bispectral enhanced images and spectral entropy weights, joint coordinate data and facial orientation angles are obtained by combining posture-physiological joint feature extraction method with joint heatmaps. The steps for constructing the physiological-motor coupling model, which outputs a physiological state matrix based on physiological parameters and exercise intensity data, are as follows: By combining the temporal characteristics of physiological parameters with the dynamic features of exercise intensity data, and using a heterogeneous feature extractor, a two-stream neural network architecture is obtained. Based on loading physiological parameters into a two-stream neural network architecture, physiological feature vectors are obtained through one-dimensional temporal convolution operations. At the same time, motion intensity data is input into a bidirectional LSTM for temporal modeling to obtain motion feature vectors. By concatenating physiological feature vectors and motion feature vectors, a dynamic coupling weight matrix is ​​obtained based on a cross-attention mechanism; Based on the dynamic coupling weight matrix, dual-stream features are fused to construct a physiological-motor coupling model; A multi-layer neural network was used to train the physiological-motor coupling model. By combining physiological parameters and exercise intensity data, the physiological state matrix was obtained through the trained physiological-motor coupling model.

7. A computer device comprising a memory and a processor, wherein the memory stores a computer program, characterized in that: When the processor executes the computer program, it implements the steps of the machine vision-based motion state detection and evaluation method according to any one of claims 1 to 5.

8. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by the processor, it implements the steps of the motion state detection and evaluation method based on machine vision as described in any one of claims 1 to 5.

Citation Information

Patent Citations

  • Joint angle estimation method for dynamically fusing multimode information

    CN118244901A