Barbell fitness action real-time detection and evaluation method
Through deep learning algorithms and depth camera technology, combined with object detection and point cloud processing, high-precision barbell fitness movement recognition and evaluation are achieved, solving the problems of low recognition accuracy and high computational volume in the existing technology, and improving the scientificity and practicality of fitness training.
Patent Information
- Application Number
- CN202510478378.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-16
- Publication Date
- 2025-06-24
AI Technical Summary
The prior art has problems in the recognition and evaluation of fitness movements with low recognition accuracy, poor adaptability, large calculation amount and poor real-time performance, especially in complex, fast speed and obstructed human movement recognition.
Deep learning algorithm combined with depth cameras is used to collect action process videos in real time through the depth camera, use the target detection network model to identify the barbell area, extract point cloud data, and fit the barbell center straight line through the least squares method to obtain the real-time motion space trajectory of the barbell, and combine the bone point detection model for action evaluation.
It realizes high-precision and high-accuracy barbell fitness movement detection and evaluation, and can identify and calculate changes in a variety of fitness movements in real time, including motion cycle, speed and abnormal movement detection, improving the scientificity and practicality of fitness training.
Smart Images

Figure CN120198463A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of artificial intelligence and machine learning, and in particular to the technical field of visual real-time detection and evaluation of fitness movements. Background Art
[0002] As the fitness industry is booming, people are paying more and more attention to the effectiveness and scientificity of fitness. Accurate fitness movement recognition and quantitative analysis of training data can help fitness enthusiasts develop reasonable training plans and prevent sports injuries. The existing technologies mainly include the following two methods.
[0003] The first is a method based on traditional sensors: using sensors such as IMU to identify movements by measuring physical quantities such as acceleration and angular velocity generated when the human body moves. However, fitness enthusiasts need to wear multiple sensors, which is inconvenient to use, and the location and wearing method of the sensors can easily affect the accuracy of the data. For example, in complex barbell training movements, the sensor may deviate from the optimal position due to body shaking, resulting in data deviation.
[0004] For example, Chinese patent application No. 202410826147.X discloses a wearable fitness motion recognition and index calculation method and apparatus, which uses wearable devices to collect motion data of target users, and inputs the data into a motion recognition module to determine whether the user is currently in motion, and to determine the category of motion, and calculate the corresponding index. This method uses wearable sensor equipment for data collection. Fitness exercise is a type of exercise that requires high flexibility and strength. Wearable sensors can easily interfere with the exercise process and affect the exercise effect. In addition, fitness exercise is a multi-part and multi-dimensional exercise process of the body. To obtain accurate recognition results, more parts and quantities of wear are required, and the cost is also higher.
[0005] The second method is based on three-dimensional vision: using a depth camera to collect data on moving human bodies, combining neural networks and point cloud processing algorithms to identify and calculate human features, thereby obtaining motion-related information. Recognition algorithms based on neural networks usually focus on bone point detection, but this method has poor adaptability to complex, fast, and occluded human movements, and it is difficult to accurately identify the exact spatial position of bones. In fitness scenes with changing light or complex backgrounds, its recognition accuracy will drop significantly. In addition, methods based on depth data have difficulties in processing point cloud data volume. The process of real-time analysis of point clouds requires huge computational effort, and simple downsampling will affect recognition accuracy, making it difficult to achieve both accuracy and speed.
[0006] As disclosed in a speed and force feedback system based on a depth camera with Chinese Patent Application No. 202211414614.5, it includes an image acquisition module, a human body capture module, a motion monitoring module, and a speed and force calculation module. The human body capture module is used to efficiently locate 16 key points of the human body, and takes the Exc-Pose algorithm as the core to capture core technical indicators such as the posture, speed, force, and power generated by athletes during physical training. This method collects data through a depth camera and processes it to achieve contactless fitness action recognition. However, this method simply focuses on calculating parameters such as action recognition and speed through the human bone point recognition method. Due to the deep learning algorithm itself being vulnerable to environmental interference, it is weak in robustness, and the target features of bone points are small, making it easy to lose the target under high-speed actions, resulting in errors in algorithm calculation. In addition, the action recognition model requires very strong recognition ability to accurately predict the results for complex data recognition. There are differences in the exercise habits, body characteristics, and action styles of different users, and the model may be difficult to adapt to all users' situations. Finally, this method lacks an abnormal determination part and does not have the ability to discriminate abnormal actions (such as doing a squat action as a deadlift action), so there are deficiencies in practicality.
[0007] In summary, most of the existing technologies are limited to the recognition and calculation of a single method. Summary of the Invention
[0008] The purpose of the present invention is to provide a real-time detection and evaluation method for barbell fitness actions applicable to various actions of barbell fitness, which combines a deep learning algorithm and a depth camera, can effectively identify and process, and has a detection and evaluation effect with high precision and high accuracy.
[0009] To achieve the above object, the technical solution of the present invention is: a method for real-time detection and evaluation of barbell fitness movements. The steps of the detection and evaluation method are as follows: S1. Determine the type of barbell fitness movement, and obtain color images and depth images of different time frames in the action process video in real time through a depth camera; S2. Input the data of each frame of color image into the target detection network model for prediction to obtain the coordinates of the barbell rod area of each frame; S3. Synchronously map the coordinates of the barbell rod area of each frame to the corresponding area of the depth image of the corresponding frame to obtain the depth information of the barbell rod area of each frame, and extract the point cloud data of the barbell rod area of each frame; S4. Fit the extracted point cloud data of the barbell rod area by the least square method to fit out the center line of the barbell rod, and obtain the real-time motion space trajectory of the barbell through the center line of the barbell rod; S5. Calculate the change information of the barbell rod action in each action cycle in the real-time motion space trajectory according to the analysis of the state of a complete action cycle of the barbell of the preset corresponding type; S6. Analyze and evaluate the barbell rod action change information obtained in step S5 with the pre-created action stratification condition library for determining the qualification of a barbell rod action cycle of the corresponding type to obtain the real-time detection and evaluation information of the barbell fitness movement.
[0010] Further, in step S1, determining the type of barbell fitness movement is to identify and determine the type of barbell fitness movement through the data collected by the depth camera or to set the type of barbell fitness movement for the user; and / or, during the process of collecting data by the depth camera in step S1, a calibration operation will be performed to obtain the true three-dimensional world coordinates (x w , y w , z w ) of the points in the image. x w is calculated through the depth value d and the horizontal field of view FOV x of the depth camera. z w is obtained by decomposing the downward tilt angle θ and the depth value d of the depth camera into the component perpendicular to the ground, and y w is obtained by decomposing the component along the ground; and / or, in step S2, the target detection network model adopts the YOLOv7 target detection network model which is trained by a training data set and converted to the ONNX format CUDA acceleration model after training; and / or, in step S4, obtaining the real-time motion space trajectory of the barbell through the center line of the barbell rod is to calculate the center point of the center line of the barbell rod and obtain the change of the center point trajectory as the real-time motion space trajectory of the barbell; and / or, in step S6, the skeleton point detection model adopts the openpose skeleton point detection model optimized by the spatio-temporal consistency constraint algorithm.
[0011] Further, the calculation method of x w , y w , z w is as follows: x w= d·tan(Δα), wherein, d represents the depth value in the depth image, Δα is the angular offset of each pixel in the depth image in the horizontal direction, W represents the image width, u represents the abscissa of the current point before calibration, y w = d·sin(θ), z w = h - d·cos(θ), h represents the camera mounting height.
[0012] Furthermore, in step S5, by analyzing the Z-axis coordinates of the center point of the real-time motion space trajectory, the high points and low points of the action changes in different directions are located and compared, and each action cycle in the real-time motion space trajectory is confirmed by two consecutive high points and one low point or two consecutive low points and one high point. The upper and lower limits, height information, the moving distance of the barbell rod between two adjacent frames, average speed, maximum speed, and / or the number of actions are obtained through the action cycle.
[0013] Furthermore, the analysis and evaluation in step S6 also include the participation of the user's bone point information in the analysis. The user's bone point information is detected by a bone point detection model for the color image and depth image in step S1, and the bone point coordinates in each frame are extracted and obtained.
[0014] Furthermore, the bone point detection model adopts an openpose bone point detection model optimized by an algorithm based on spatio-temporal consistency constraints.
[0015] Furthermore, the optimization method of the openpose bone point detection model includes designing a temporal smoothing constraint loss function and adding a temporal smoothing term in the training stage to force the limitation of the continuity of the bone trajectory, and predicting and correcting the dynamic trajectory of the bone points based on the Kalman filter algorithm to obtain the estimated values of the optimized bone point positions and speeds.
[0016] Furthermore, the formula of the loss function is In this loss function, λ is the weight coefficient of the smoothing term; J is the number of bone points output by the openpose standard; represents the three-dimensional coordinates of the j-th bone point in the t-th frame.
[0017] Furthermore, the prediction and correction of the dynamic trajectory of the bone points based on the Kalman filter algorithm include a prediction stage, an observation stage, and a correction stage;
[0018] Prediction stage: First, define the bone point state vector X t = [p x , p y , p z , v x , v y , v z T , indicating the motion state of the bone point in three-dimensional space, p x , p y , p z indicating the coordinate sum of the bone point in three-dimensional space, v x , v y , v z indicating the velocity component, describing the change law between the states of the bone point in the front and back frames through the state transition matrix F, and the expression is: where, Δt represents the time interval between adjacent frames, I3 is a 3×3 identity matrix, the upper left block I3 represents the transmission of maintaining the position component, the upper right block Δt·I3 represents the cumulative contribution of velocity to position, and the lower right block I3 represents the transmission of maintaining the velocity component. Predict the bone point state vector of each frame based on the state transition matrix F, and according to the previous frame state X t-1 , predict the current frame state X t|t-1 , and the formula is X t|t-1 = F·X t-1 ;
[0019] Observation stage: Detect the position of the bone point in the current frame through the OpenPose algorithm to obtain the observation value where z t is the three-dimensional coordinate of the bone point output by OpenPose, and are the observation values in the x-axis direction of the bone point, the observation values in the y-axis direction of the bone point, and the observation values in the z-axis direction of the bone point respectively, and c t represents the confidence of the bone point detection output by openpose. When c t <the preset threshold, there is occlusion or misdetection of the bone point. When c t ≥ the preset threshold, the result is credible;
[0020] Calibration stage: Dynamically adjust the weights of the predicted value and the observed value according to the detection confidence c t , and calculate the optimal estimated value. Specifically, define the observation noise covariance matrix where, σ obs is the standard deviation of the observation position noise, indicating the error range of the detection algorithm. At the same time, establish the Kalman gain K t to determine the weight ratio of the predicted value and the observed value. Through the Kalman gain K t = P t|t-1 ·H T ·(H·P t|t-1 ·H T + R) -1 Map the bone point state vector x t to the dimension of the actual observation, representing the observation matrix, P t|t-1is the predicted covariance matrix, which represents the uncertainty of predicting the current state based on historical data. It is a 6×6 matrix. The diagonal elements represent the predicted error variances of each state component, and the off-diagonal elements represent the correlations between components. Its update formula is P t|t-1 = F·P t-1 ·F T ; According to the Kalman gain K t perform weighted fusion of the predicted value and the observed value of the skeleton point state. The formula is In practice, if the detection confidence of a certain skeleton point in the current frame by the openpose model is high, the state data of this skeleton point directly adopts the output observed value z of the model t ; If the model detection confidence is low, then through a correction algorithm, the correction result of the skeleton point is taken as the state data of this skeleton point; Through a preset database of the maximum reasonable distance between joints, if the detected value of the distance between joints exceeds the range, the interpolation correction method is started to perform forced error correction on the skeleton point position.
[0021] Furthermore, the interpolation correction method is started to perform forced error correction on the skeleton point position and is also sorted according to the importance from the center of the human body to the limbs from the inside to the outside. Here, the interpolation correction method between the two skeleton points from the elbow to the wrist is taken as an example. The interpolation correction methods between the joints of other skeleton points can be inferred by analogy. The joint correction formula between the two skeleton points from the elbow to the wrist:
[0022]
[0023] In the formula, P wrist is the three-dimensional coordinate of the wrist skeleton point predicted by the openpose model, and P elbow is the three-dimensional coordinate of the elbow skeleton point predicted by the openpose model, is the three-dimensional coordinate of the wrist position predicted by the Kalman filter, and d max is the maximum distance limit preset between these two skeleton points.
[0024] Furthermore, compare the coordinates of the skeleton points in the same frame with the coordinates of the center point of the barbell rod, and make a comprehensive judgment based on the action stratification condition library to eliminate non-standard actions. Here, the comprehensive judgment is taken as an example for the type of barbell squat in the bell fitness exercise type. The comprehensive judgment conditions for other types can be inferred by analogy; The comprehensive judgment conditions for the type of barbell squat include: 1) The z-axis height of the center point of the barbell rod in a complete action cycle shall not be lower than the height of the elbow skeleton point, that is, P center (t).z ≥ C elb (t).z, where P center (t).z is the coordinate of the center point of the barbell rod on the z-axis, and C elb(t).z is the coordinate of the elbow bone point on the z-axis; 2) The z-axis height of the center point of the barbell bar shall not be lower than the preset height during a complete movement cycle, that is, P center (t).z ≥ preset height; 3) The height of the wrist bone point shall not be lower than the height of the hip bone point, that is, C wri (t).z > C hip (t).z, where C wri (t).z is the coordinate of the wrist bone point on the z-axis, C hip (t).z is the coordinate of the hip bone point on the z-axis; 4) The center point of the barbell bar must be kept behind the nose bone point of the human body, that is, P center (t).y > C nose (t).y, P center (t).y represents the y-axis coordinate of the center point of the barbell bar, C nose (t).y is the y-axis coordinate of the nose bone point of the human body; Set a tolerance space, and set several ranges of abnormal frames for each bone point according to the probability of occurrence. When different conditions are all met within the abnormal frames of several ranges, the abnormal condition of the action is determined to pass.
[0025] By adopting the above technical solution, the beneficial effects of the present invention are as follows: The above technical solution of the present invention aims at the requirements of barbell fitness equipment in terms of intelligence and specialization, and proposes a method for identifying and evaluating barbell fitness actions. This method collects the action trajectories and body postures of the human body and the barbell bar during the movement process through a depth camera, combines 3D point cloud data for processing and extraction, and performs real-time calculation and determination, and can realize functions such as the number recognition, average speed calculation, maximum speed calculation, and abnormal action detection of at least 11 actions including barbell squats, deadlifts, bench presses, etc.
[0026] Specifically, in order to ensure the real-time performance of action detection in the above method, the algorithm trains and recognizes the barbell rod area based on the object detection model, uses the internal parameter matrix of the depth camera to map the coordinates of the barbell area in the color image to the corresponding position in the depth image, extracts the local point cloud and fits the barbell rod straight line, realizing the detection of the barbell's spatial position while significantly reducing the computing resources consumed by point cloud processing to improve the operation speed. In addition, in high-intensity fitness actions such as squats and deadlifts, the human body moves at a high speed (the peak speed of some actions > 1.2 m / s) and joint occlusion is frequent. Traditional bone point detection algorithms (such as OpenPose) are prone to problems such as spatial jitter, bone point jumps caused by limb occlusion, and discontinuous trajectories. In the present invention, a method for enhancing the stability of bone points under spatio-temporal consistency constraints is further proposed to effectively suppress bone point jumps and enhance the smoothness of the motion trajectory. Finally, combining human anatomy knowledge, an abnormal action determination mechanism based on multi-dimensional data fusion is proposed, which can realize the recognition and elimination of incorrect actions. In addition, compared with other methods, the method of the present invention can first perform non-contact recognition, eliminate the interference of contact sensors on the movement of fitness personnel, realize the intelligent transformation of barbell fitness exercises, and bring new possibilities and higher efficiency to fields such as sports training and fitness guidance. Through methods such as local point cloud extraction and GPU acceleration, the experiment can reduce the point cloud computing volume by at least 90%, effectively improve the processing speed, significantly reduce the algorithm computing volume, and the experimental processing speed can reach 20 fps, ensuring the real-time performance of the motion recognition process. Secondly, the OpenPose bone point detection neural network improved based on spatio-temporal consistency constraints obtains important features, can more accurately identify human motion feature points, effectively improve the recognition accuracy of human motion bone points, and is used to establish an effective action abnormality determination mechanism.
[0027] In summary, through the innovative integration of multiple technologies, it is possible to achieve intelligent barbell fitness recognition and evaluation with high precision, high accuracy, and low latency, helping fitness personnel obtain real-time, accurate, and reliable motion tracking results during exercise. Brief Description of the Drawings
[0028] Figure 1 It is a flowchart of a method for real-time detection and evaluation of barbell fitness actions involved in the present invention.
[0029] Figure 2 It is a schematic diagram of the hardware and spatial coordinate system involved in the present invention.
[0030] Figure 3 It is a graph showing the relationship between the barbell height and time of the squat action involved in the present invention.
[0031] Figure 4 It is a flowchart of the improvement of the spatio-temporal consistency constraint of the OpenPose model involved in the present invention.
[0032] Figure 5It is a schematic diagram of openpose skeleton point detection related to the present invention. Detailed implementation manners
[0033] In order to further explain the technical solution of the present invention, the present invention will be elaborated in detail below through specific embodiments.
[0034] A real-time detection and evaluation method for barbell fitness movements disclosed in this embodiment, the specific method flow steps are as Figure 1 shown, and the following is a detailed description. The steps of the real-time detection and evaluation method are as follows:
[0035] S1. Determine the type of barbell fitness movement, and obtain color images and depth images of different time frames in the action process video in real time through a depth camera.
[0036] Since the barbell fitness movement itself belongs to a movement with relatively complex types and flexibility, it is necessary to perform comprehensive detection in order to obtain a detection result with high accuracy and strong real-time performance. In response to this problem, the technical solution of the present invention proposes a fitness real-time detection method based on a depth three-dimensional camera. In this step, determining the type of barbell fitness movement is determined on a pre-constructed motion data acquisition system, which can be to identify the type of barbell fitness movement through the data collected by the depth camera, or to set the type of barbell fitness movement for the user. For example, Figure 1 as shown in the flowchart, the user first performs initialization settings to determine the type and then enters the acquisition, detection, and evaluation. The depth camera (or three-dimensional camera) is a hardware device of the pre-constructed motion data acquisition system, and its installation and setting method is to install and set it in such a way that the camera's field of view can cover the motion space (the motion space where the fitness person does barbell movements). As Figure 2 shown, it is fixedly installed at the upper part on one side of the motion space and the lens is fixed with a certain downward oblique shooting angle. The motion space in the figure for the test implementation includes a fitness rack, and the depth camera is installed on the cross beam at the upper rear of the fitness rack and shoots obliquely downward at 20° - 30°. The spatial coordinate axes of the constructed space coordinate system are such that the z-axis is perpendicular to the ground upward, the y-axis is away from the depth camera direction and parallel to the ground, and the x-axis corresponds to the length direction of the barbell rod and is parallel to the ground.
[0037] The depth camera obtains the action process video in real time, including simultaneously collecting the color video stream V color and the depth video stream V depth , and splitting them by frame F. The expression is as follows:
[0038] V color ={F color (t1), F color (t2), F color (t3), …, F color (t n)}
[0039] V depth = {F depth (t1), F depth (t2), F depth (t3), …, F depth (t n )}
[0040] where t1, t2, t3 … t n represent the time points of the 1st frame, 2nd frame, 3rd frame … nth frame respectively, and n is the total number of frames captured from the video during the entire movement process. F color represents the color image extracted from the color video stream, and F depth represents the depth image extracted from the depth video stream; it is used for subsequent calculations and data processing.
[0041] To ensure the accuracy and integrity of the data, the acquisition process of the depth camera will perform a calibration operation to eliminate errors caused by factors such as spatial position and viewing angle. The core of the calibration operation is to establish a mapping relationship from depth data to the world coordinate system by precisely measuring the installation height h and the downward tilt angle θ of the camera, so that the acquired data can accurately reflect the positions and motion states of the barbell and the human body in three-dimensional space. For a point (u, v) in the depth image with its corresponding depth value d, since the downward tilt angle θ of the depth camera does not affect the horizontal direction, the calibration operation finally obtains the true three-dimensional world coordinates (x w , y w , z w ) of the point in the image. Its horizontal coordinate x w can be directly calculated from the depth value d of the depth camera and the horizontal field of view FOV x , and the calculation formula is x w = d · tan(Δα), where Δα is the angular offset of each pixel in the depth image in the horizontal direction, W represents the image width, and u represents the abscissa of the current point before calibration. In the vertical direction, since the depth value d is measured along the optical axis when the depth camera tilts downward, the downward tilt angle θ and the depth value d need to be decomposed into components perpendicular to the ground to obtain z w and components along the ground to obtain y w . Decomposing into components perpendicular to the ground can correct the vertical height change caused by the tilt of the depth camera, and decomposing into components along the ground is to calculate the actual distance of the object on the ground plane, that is, the vertical distance of the object from the depth camera along the ground direction. The calculation formulas are: z w = h - d · cos(θ), y w = d · sin(θ), and the subsequent calculated coordinate values are world coordinates.
[0042] S2. Input the data of each frame of color image into the YOLOv7 object detection network model for prediction to obtain the barbell rod region coordinates of each frame.
[0043] For the construction of the YOLOv7 object detection network model, first, collect multiple (3000 in the experimental example) motion color image data containing barbell rods to establish a training dataset. These images cover different barbell types, lighting conditions, background environments, and different motion stages. To construct a high-quality training dataset, accurately annotate the barbell rod targets in these images. The annotation information includes the bounding box coordinates, colors, categories, etc. of the barbell rods. In this way, establish a training dataset, where represents different samples in the dataset. Use the established training dataset to train the YOLOv7 object detection network model. After completing the training of the YOLOv7 object detection network model, convert it into a CUDA-accelerated model in ONNX format. The main advantage of this conversion is that it can make full use of the powerful parallel computing ability of CUDA and significantly improve the recognition and inference speed of the model.
[0044] For the real-time acquired color image frame F color (t), input it into the YOLOv7 object detection network model for prediction to obtain the prediction information of the barbell rod region coordinates of each frame, that is, the real-time barbell rod region C bb (t) coordinates, which are represented by the upper left corner coordinates and the lower right corner coordinates of the region. The expression is C bb (t) = ((x1, z1), (x2, z2)), so as to realize the real-time positioning of the barbell rod in each frame of color image.
[0045] S3. Synchronously map the barbell rod region coordinates of each frame to the corresponding region of the depth image of the corresponding frame to obtain the depth information of the barbell rod region of each frame, and extract the point cloud data of the barbell rod region of each frame.
[0046] Since there is a spatial mapping relationship between the color image and the depth image, it is necessary to accurately map the real-time barbell rod region C color (t) obtained from the real-time acquired color image frame F bb (t) coordinates to the corresponding region of the real-time acquired depth image frame F depth (t). Here, it can be realized through the coordinate mapping function built in the depth camera to obtain the depth information of the corresponding barbell rod region, and further extract the real-time barbell rod region point cloud P bb (t). These points have three-dimensional coordinates and constitute the spatial information representation of the barbell rod in the depth image. By mapping and extracting the point cloud, small-area fast calculation can be realized. Compared with the full-image region point cloud extraction method, the calculation amount can be reduced by 90%, effectively improving the real-time processing speed.
[0047] S4. Fit the point cloud data of the barbell rod area extracted by the least squares method to fit the center line of the barbell rod, and obtain the real-time motion space trajectory of the barbell through the center line of the barbell rod.
[0048] Further obtain the exact spatial position of the barbell rod. Fit the real-time barbell rod area point cloud P bb (t) data by the least squares method. The purpose of the least squares method is to find a straight line that best represents the distribution of these point clouds as the center line L of the barbell rod center . By solving the linear equations, find the straight line equation L that satisfies the minimization of the sum of squared errors center : Ax + By + Cz + D = 0, where A, B, and C are the coefficients of the straight line equation. When actually calculating the trajectory of the barbell rod, it can be further simplified to represent the trajectory change of the barbell by calculating the change of the center point P center (t) = (x, y, z) (i.e., the real-time center point of the barbell rod) trajectory change, that is, obtain the center point trajectory change through the center point of the center line of the barbell rod as the real-time motion space trajectory of the barbell;
[0049] S5. According to the analysis and calculation of the state of a complete motion cycle of the barbell of the preset corresponding type, obtain the barbell rod motion change information (including the upper and lower limits of the barbell rod motion change, the moving distance of the barbell rod between adjacent time frames, and the height information of the barbell rod, etc.) in each motion cycle of the real-time motion space trajectory.
[0050] The barbell fitness exercise types on the pre-constructed motion data acquisition system may include common barbell squats, front squats, jump squats, deadlifts, power cleans, hang power cleans, high pulls, high grabs, military presses, bench presses, incline bench presses (including these 11 types in the test examples), etc.
[0051] The system analyzes the real-time motion space trajectory of the barbell according to the determined motion type (or the motion type determined by user settings, which includes the preset state of a complete motion cycle of the barbell). Specifically, by analyzing the Z-axis coordinate positioning of the center point of the real-time motion space trajectory and comparing to obtain the high and low points of the motion changes in different directions, and confirm each motion cycle in the real-time motion space trajectory with two consecutive high points and one low point or two consecutive low points and one high point. Obtain the upper and lower limits, height information, moving distance of the barbell rod between two adjacent frames, average speed, maximum speed, and / or the number of motions through the motion cycle.
[0052] Taking the barbell squat as an example below, within a complete cycle, the barbell rod will go through three states: high point, low point, and high point (as shown in the actual demonstration Figure 3 ), and the system analyzes the center point P in the real-time motion space trajectory of the barbellcenter Locate these three key points according to the z-axis coordinate of (t), and monitor and compare these z-axis coordinates in real time. When the z value is the largest, it is recorded as the high point, and when the z value is the smallest, it is recorded as the low point, so as to determine the heights h high1 and h high2 of the two high points and the corresponding time points t high1 and t high2 and the height h low of a low point and the corresponding time point t low , so as to confirm the movement cycle T of a barbell squat, T = t high2 - t high1 , obtain the upper and lower limits and height information of the barbell bar within a movement cycle T, and initially consider that this movement is a standard squat movement, and the number of movements N f can be initially recorded, and the value is incremented by 1.
[0053] Furthermore, by calculating the average speed and maximum speed of the movement based on the movement cycle T of the barbell squat, the calculation includes calculating the moving distance d of the barbell bar between every two frames k , where k is the time stamps corresponding to the adjacent frame time points t k and t k+1 , y represents the y-axis coordinate of the center point of the barbell bar in the world coordinate system, and z represents the z-axis coordinate of the center point of the barbell bar in the world coordinate system. Since the movement process of the fitness movement is a work process perpendicular to the ground, the moving distance determination in the x-axis direction is not performed when calculating the parameters. Calculate the average speed v a of the movement within a movement cycle T, and the calculation formula is where e is the last frame of this movement cycle. Further, by comparing the moving distances d of the barbell bar between all adjacent two frames within a movement cycle T k , find the moving distance d of the barbell bar with the largest displacement k , calculate the corresponding moving distance d k of the maximum speed v m , and the calculation formula is v m = max(d) / (t k+1 - t k ).
[0054] S6. Detect the color image and depth image in step S1 through the openpose skeleton point detection model optimized by the spatio-temporal consistency constraint algorithm, and extract the skeleton point coordinates in each frame.
[0055] In this embodiment, since only the real-time motion space trajectory of the barbell is used to judge the movement cycle in S5, and factors such as movement standardization are not considered, the number of movements N fIts effectiveness still needs to be further verified. In this embodiment, the spatial position of the moving skeletal points and the barbell rod are combined in real time in this step to determine the standardization of the action.
[0056] The OpenPose neural network is a deep learning algorithm widely used in human skeletal point detection and can be applied to the detection of human skeletal points in barbell fitness movements. However, due to the fast movement speed and large movement amplitude during barbell fitness, the following problems are likely to occur: 1. Spatial jitter: The skeletal point coordinates jump frame by frame, unable to reflect the continuous movement trajectory; 2. Occlusion failure: When a local limb is occluded, the algorithm relying on a single-frame image cannot recover the missing skeletal points; 3. Temporal discontinuity: The skeletal point trajectories within the action cycle are not coherent, resulting in misjudgment of action counting (such as splitting a single squat into two half-squat actions). Therefore, a spatio-temporal consistency enhancement algorithm is designed to fuse the motion smoothness in the time dimension, the Kalman filter algorithm, and the skeletal topological relationship in the space dimension to construct a multi-level stability enhancement mechanism to solve the problems of skeletal point jitter and occlusion failure under high-speed movement.
[0057] The specific process is as Figure 4 shown. First, to suppress the sudden change of skeletal point coordinates between adjacent frames, during the training stage of the OpenPose model, the continuity of the skeletal trajectory is forcibly restricted, and a temporal smoothing constraint loss function is designed to add a temporal smoothing term. The loss function formula is In this loss function, λ = 0.3 is the weight coefficient of the smoothing term; J = 18 is the number of skeletal points output by the OpenPose standard; represents the three-dimensional coordinates of the j-th skeletal point in the t-th frame. Through the temporal smoothing loss function, the change stability of adjacent skeletal points in the model prediction results can be increased, and the mutation probability can be reduced.
[0058] Furthermore, based on the Kalman filter algorithm, the dynamic trajectory prediction and correction of skeletal points are performed. The Kalman filter is a method for estimating the state of a dynamic system. Through the cycle of "prediction - observation - correction" (that is, this method is divided into three stages: skeletal point position prediction, obtaining the actual detection results of the current frame, and skeletal point correction), the estimated values of the skeletal point position and speed are gradually optimized.
[0059] Prediction stage:
[0060] First, define the skeletal point state vector X t =[p x , p y , p z , v x , v y , v z T , representing the motion state of the skeletal point in three-dimensional space, including the coordinates p x , p y , p z and the velocity component v x , v y , v z , where T is the transformation symbol.
[0061] The state transition matrix F describes the change law of the skeletal point state between the previous and current frames (i.e., from frame t - 1 to frame t. Assuming the skeletal point moves at a constant speed (the speed change can be ignored in a short time), the state transition matrix is:
[0062] Among them, I3 is a 3×3 identity matrix, Δt represents the time interval between adjacent frames (unit: s). The upper left block I3 represents the transmission of maintaining the position component, the upper right block Δt·I3 represents the cumulative contribution of velocity to position (displacement = velocity × time), and the lower right block I3 represents the transmission of maintaining the velocity component.
[0063] Predict the skeletal point state vector of each frame based on the state transition matrix F. According to the previous frame state X t-1 , predict the current frame state X t|t-1 , and the formula is X t|t-1 = F·X t-1 .
[0064] Observation stage:
[0065] Detect the skeletal point positions of the current frame through the OpenPose algorithm to obtain the observation values: Among them, z t is the three-dimensional coordinate of the skeletal point output by OpenPose, including which are the observation values in the x-axis direction of the skeletal point, the observation values in the y-axis direction of the skeletal point, and the observation values in the z-axis direction of the skeletal point respectively. c t represents the confidence of the skeletal point detection output by openpose (range 0 - 1). The higher the value, the more reliable the detection. It is defined that when the skeletal point confidence c t < 0.7, there is occlusion or misdetection of the skeletal point. When the confidence c t ≥ 0.7, the result is credible.
[0066] Correction stage:
[0067] According to the detection confidence c t , dynamically adjust the weights of the predicted value and the observed value, and calculate the optimal estimated value. Since there are errors in the observed values (such as algorithm misdetection or environmental noise), it is necessary to define the observation noise covariance matrix Among them, σ obs is the standard deviation of the observation position noise, representing the error range of the detection algorithm, and the value is 10mm.
[0068] At the same time, establish the Kalman gain K tTo determine the weight ratio of the predicted value and the observed value.
[0069] Kalman gain K t = P t|t-1 · H T · (H · P t|t-1 · H T + R) -1 , where the function of the observation matrix H is to map the skeleton point state vector x t to the dimension of the actual observation. Since in this scheme, the skeleton point state vector X t includes position and velocity (6 - dimensional), but the observed value z t only includes position (3 - dimensional). Therefore, the H matrix is designed to be 3 rows and 6 columns, indicating extracting the first 3 dimensions (position components) from the 6 - dimensional state P t|t-1 is the predicted covariance matrix, which represents the uncertainty of predicting the current state based on historical data. It is a 6×6 matrix. The diagonal elements represent the predicted error variances of each state component (position and velocity), and the non - diagonal elements represent the correlations between components. Its update formula is P t|t-1 = F · P t-1 · F T .
[0070] According to the Kalman gain formula, when the prediction uncertainty is much greater than the observation noise (HPH T >> R), then K t approaches 0, and the system trusts the observed value more; if the observation noise is larger (R >> HPH T ), then K t approaches 0, and the system trusts the predicted value more. Finally, according to the Kalman gain K t , a weighted fusion of the predicted value and the observed value of the skeleton point state is performed, and the formula is In practice, if the confidence of the openpose model for a certain skeleton point in the current frame is high (c t ≥0.7), then the state data of this skeleton point directly adopts the output observed value z t ; if the confidence of the model detection is low (c t <0.7), then through the correction algorithm, the correction result of the skeleton point is taken as the state data of this skeleton point.
[0071] Further, set up a correction measure for bone topology rules. According to human anatomy, preset a database of the maximum reasonable distances between joints (e.g., the elbow-wrist distance dmax = 400 mm. Based on the statistical average length of the adult upper arm from elbow to wrist of about 300 mm, a 30% margin is reserved). If the detected value exceeds the range, start the interpolation correction method to forcefully correct the position of the bone points. The correction formula is as follows (taking the interpolation correction method between the two bone points from elbow to wrist as an example, and the interpolation correction methods between other bone points can be inferred by analogy)
[0072]
[0073] The three-dimensional coordinates of the wrist bone point position predicted by the openpose model, P elbow are the three-dimensional coordinates of the elbow bone point position predicted by the openpose model is the three-dimensional coordinate of the wrist position predicted by the Kalman filter, d max is the maximum distance limit preset between these two bone points. According to the above formula, forced error correction can be achieved when the corresponding positions of the bone points do not conform to the actual situation, and they are sorted according to the importance from the center of the human body to the outside of the limbs. For example, the wrist is farther from the center of the human body than the elbow, so the probability of mutation and error is greater, and it is corrected
[0074] Through the above method, a set of bone point coordinates C skeleton (t) = {(x1, y1, z1), (x2, y2, z2), … (x 18 , y 18 , z 18 )} containing 18 points such as the wrists, elbows, shoulders, knees, and ankles of the body is obtained. The schematic diagram of the bone points is as Figure 5 shown
[0075] S7. Analyze and evaluate the barbell rod movement change information obtained in step S5 and the bone point coordinates obtained in step S6 with the pre-created action stratification condition library for determining the qualification of one action cycle of the corresponding type of barbell rod to obtain real-time detection and evaluation information of the barbell fitness action
[0076] By comparing the bone point coordinates of the same frame with the center point coordinates of the barbell rod and making a comprehensive judgment based on the action stratification condition library, evaluation can be achieved, and non-standard actions can be excluded. Here, the barbell fitness movement type takes the type of barbell squat as an example for comprehensive judgment, and other types can be inferred by analogy
[0077] Taking the squat action as an example, the human body bone point coordinates C skeleton (t) output from each frame of the color image are compared with the center coordinates P centerCompare (t) = (x, y, z) and make a comprehensive judgment based on the pre-created action stratification condition library for determining the eligibility of a barbell rod in one movement cycle. Eliminate non-standard actions. The judgment conditions are as follows:
[0078] 1) The z-axis height of the center point of the barbell rod in a complete movement cycle shall not be lower than the height of the elbow bone point, that is, P center (t).z ≥ C elb (t).z, where P center (t).z is the coordinate of the center point of the barbell rod on the z-axis, and C elb (t).z is the coordinate of the elbow bone point on the z-axis.
[0079] 2) The z-axis height of the center point of the barbell rod in a complete movement cycle shall not be lower than the preset height, that is, P center (t).z ≥ preset height.
[0080] 3) The height of the wrist bone point shall not be lower than the height of the hip bone point, that is, C wri (t).z > C hip (t).z, where C wri (t).z is the coordinate of the wrist bone point on the z-axis, and C hip (t).z is the coordinate of the hip bone point on the z-axis.
[0081] 4) The center point of the barbell rod must be kept behind the human nose tip bone point, that is, P center (t).y > C nose (t).y, P center (t).y represents the y-axis coordinate of the center point of the barbell rod, and C nose (t).y is the y-axis coordinate of the human nose tip bone point.
[0082] All of the above four conditions need to be met. To prevent occasional bone point jumps or point cloud fluctuations, a certain tolerance space is set. Each bone point is set with several ranges (such as 1 frame - 3 frames at most) of abnormal frames according to the probability of occurrence. When different conditions are all met within the abnormal frames, the abnormal condition determination of this action passes.
[0083] By synthesizing the above information, the possible number of actions N output in step S4 can be determined. f When all the determination conditions are met, the value of the true number of actions N r is incremented by 1, and the average speed v of this action is output a and the maximum speed v m and other parameters; when any one of the conditions is not met, the system can output an action abnormal signal and continue with action recognition without updating the true number of actions N r, continue with the action recognition evaluation of the next movement cycle.
[0084] The main technical features of the above technical solution of the present invention include: 1. By synchronously acquiring human skeleton points and spatial point cloud data through RGB-D data, multi-modal data collaborative analysis is realized. Combining deep learning algorithms and point cloud processing algorithms, comprehensive monitoring of the positions and movement states of the barbell and the human body in three-dimensional space is achieved, solving the problem that traditional monitoring methods cannot accurately obtain spatial information. In addition, a local point cloud extraction method is proposed. After locating the barbell area through YOLOv7, only the corresponding area in the depth image (accounting for <10% of the whole image) is used for point cloud calculation. Compared with the traditional full-image point cloud processing scheme, it can accurately focus on the primary target for processing, and the calculation time is reduced by 90%.
[0085] 2. The present invention proposes an improved method for enhancing the spatio-temporal consistency of openpose, which integrates the motion smoothness in the time dimension, the Kalman filtering algorithm, and the skeletal topological relationship in the space dimension to construct a multi-level stability enhancement mechanism, solving the problems of jitter and occlusion failure of the openpose model in identifying skeleton points under high-speed movement. Through the above method, high-precision skeleton point information can be extracted, and the problems of missing and jumping of skeleton points caused by high-intensity movements, body occlusions, etc. can be effectively reduced.
[0086] 3. An action anomaly determination library is established. By analyzing the spatial information of the human skeleton point coordinates and the barbell rod center coordinates frame by frame in the movement cycle, the action is comprehensively determined according to the preset conditions, accurately judging the standardization and effectiveness of the action, and outputting the number of actual actions and speed parameters. Realize the accurate detection of barbell fitness actions and speed and strength feedback, which can play a very good guiding role in the actual fitness exercise process.
[0087] The above embodiments and diagrams do not limit the product form and style of the present invention. Any appropriate changes or modifications made by those of ordinary skill in the art shall be regarded as not departing from the patent scope of the present invention.
Claims
1. A real-time detection and evaluation method for barbell fitness movements, characterized in that: The steps of the detection and evaluation method are as follows: S1. Determine the type of barbell fitness exercise, and obtain color images and depth images of different time frames in the action process video in real time through a depth camera; S2, inputting the data of each frame of color image into the target detection network model for prediction, and obtaining the coordinates of the barbell area of each frame; S3, synchronously mapping the barbell area coordinates of each frame to the corresponding area of the depth image of the corresponding frame, obtaining the depth information of the barbell area of each frame, and extracting the point cloud data of the barbell area of each frame; S4, fitting the extracted barbell area point cloud data by the least square method, fitting the barbell center line, and obtaining the real-time motion space trajectory of the barbell through the barbell center line; S5, analyzing and calculating a complete action cycle state of a barbell of a preset corresponding type to obtain the barbell movement change information in each action cycle in the real-time motion space trajectory; S6. Analyze and evaluate the barbell bar movement change information obtained in step S5 and the pre-created action hierarchical condition library of corresponding types for determining the eligibility of a barbell bar movement cycle, so as to obtain real-time detection and evaluation information of the barbell fitness movement.
2. A barbell fitness movement real-time detection and evaluation method as claimed in claim 1, characterized in that: In step S1, determining the barbell fitness exercise type is to identify and determine the barbell fitness exercise type through data collected by the depth camera or to set and determine the barbell fitness exercise type for the user; And / or, in the process of depth camera acquisition in step S1, a calibration operation is performed to obtain the real three-dimensional world coordinates (x w ,y w ,z w ), x w Through the depth value d and horizontal field of view FOV of the depth camera x It is calculated that z is obtained by decomposing the downward tilt angle θ of the depth camera and the depth value d into components perpendicular to the ground w and the component along the ground gives y w ; And / or, the target detection network model in step S2 adopts a YOLOv7 target detection network model that has been trained with a training data set and converted into a CUDA acceleration model in ONNX format after training; And / or, in step S4, obtaining the real-time spatial trajectory of the barbell's motion through the barbell's center line is to calculate the center point of the barbell's center line, and obtain the center point trajectory change as the real-time spatial trajectory of the barbell's motion.
3. A barbell fitness movement real-time detection and evaluation method as claimed in claim 2, characterized in that: The x w ,y w ,z w The calculation method is as follows: w =d·tan(Δα), Where d represents the depth value in the depth image, Δα is the horizontal angle offset of each pixel in the depth image, W represents the image width, u represents the horizontal coordinate of the current point before calibration, and y w =d·sin(θ),z w =hd·cos(θ), where h represents the camera installation height.
4. A barbell fitness action real-time detection and evaluation method as claimed in claim 1, 2 or 3, characterized in that: In step S5, the Z-axis coordinate positioning of the center point of the real-time motion space trajectory is analyzed and compared to obtain the high points and low points of the movement changes in different directions, and each movement cycle in the real-time motion space trajectory is confirmed by two consecutive high points and one low point or two consecutive low points and one high point, and the upper and lower limits, height information, the barbell movement distance of two adjacent frames, the average speed, the maximum speed and / or the number of movements are obtained through the movement cycle.
5. A barbell fitness action real-time detection and evaluation method as claimed in claim 4, characterized in that: The analysis and evaluation in step S6 also includes the user's skeleton point information participating in the analysis. The user's skeleton point information is detected by a skeleton point detection model on the color image and depth image of step S1 to extract the coordinates of the skeleton points in each frame.
6. A barbell fitness movement real-time detection and evaluation method as claimed in claim 5, characterized in that: The skeleton point detection model adopts the openpose skeleton point detection model optimized by a spatiotemporal consistency constraint algorithm.
7. A barbell fitness movement real-time detection and evaluation method as claimed in claim 6, characterized in that: The optimization method of the openpose skeleton point detection model includes designing a temporal smoothing constraint loss function and adding a temporal smoothing term in the training stage to force the continuity of the skeleton trajectory to be restricted, and performing skeleton point dynamic trajectory prediction and correction based on the Kalman filter algorithm to obtain the estimated values of the optimized skeleton point position and speed.
8. A barbell fitness movement real-time detection and evaluation method as claimed in claim 7, characterized in that: The loss function formula is: In this loss function, λ is the weight coefficient of the smoothing term; J is the number of skeleton points output by Openpose standard; Represents the three-dimensional coordinates of the jth bone point in the tth frame; And / or, the prediction and correction of the dynamic trajectory of the skeleton points based on the Kalman filter algorithm includes a prediction stage, an observation stage and a correction stage; Prediction stage: First define the skeleton point state vector X t =[p x ,p y ,p z ,v x ,v y ,v z ] T , represents the motion state of the skeleton point in three-dimensional space, p x ,p y ,p z Represents the coordinates of the bone point in three-dimensional space and v x ,v y ,v z Represents the velocity component, and the state transfer matrix F is used to describe the change law between the previous and next frames of the skeleton point state. The expression is: Among them, Δt represents the time interval between adjacent frames, I3 is a 3×3 unit matrix, the upper left block I3 represents the transfer of the position component, the upper right block Δt·I3 represents the cumulative contribution of velocity to position, and the lower right block I3 represents the transfer of the velocity component. The state vector of the skeleton point of each frame is predicted based on the state transfer matrix F, and the state vector of the skeleton point of each frame is predicted based on the state X of the previous frame. t-1 , predict the current frame state X t|t-1 , the formula is X t|t-1 =F·X t-1 ; Observation phase: Use the OpenPose algorithm to detect the position of the skeleton points in the current frame and obtain the observation value where z t is the 3D coordinates of the skeleton points output by OpenPose, and are the x-axis observation value of the skeleton point, the y-axis observation value of the skeleton point, and the z-axis observation value of the skeleton point, respectively. t Indicates the confidence of the skeleton point detection output by OpenPose. t When the value is less than the preset threshold, there is occlusion of bone points or false detection. When ct is greater than or equal to the preset threshold, the result is reliable. Correction phase: According to the detection confidence ct, dynamically adjust the weights of the predicted value and the observed value, and calculate the optimal estimate. Specifically, define the observation noise covariance matrix Among them, σ obs is the standard deviation of the observed position noise, indicating the error range of the detection algorithm, and establishing the Kalman gain K t To determine the weight ratio of the predicted value and the observed value, the Kalman gain K t =P t|t-1 ·H T ·(H·P t|t-1 ·H T +R) -1 Map the skeleton point state vector xt to the actual observed dimension, Tabular observation matrix, P t|t-1 is the prediction covariance matrix, which represents the uncertainty of predicting the current state based on historical data. It is a 6×6 matrix. The diagonal elements represent the prediction error variance of each state component, and the non-diagonal elements represent the correlation between components. Its update formula is P t|t-1 =F·P t-1 ·F T ; According to the Kalman gain K t The predicted value and observed value of the skeleton point state are weighted fused, and the formula is: In practice, if the openpose model detection confidence of a certain bone point in the current frame is high, the state data of the bone point directly uses the output observation value z of the model. t ; If the model detection confidence is low, the correction algorithm is used to obtain the correction result of the bone point As the state data of the bone point; by presetting the maximum reasonable distance database between joints, if the detection value of the distance between joints exceeds the range, the interpolation correction method is started to force error correction of the bone point position.
9. A barbell fitness movement real-time detection and evaluation method as claimed in claim 8, characterized in that: The interpolation correction method is started to perform forced correction of the skeletal point positions and is also sorted from the center of the human body to the limbs in order of importance from the inside to the outside. The interpolation correction method takes the inter-joint interpolation correction method of the two skeletal points from the elbow to the wrist as an example, and the inter-joint interpolation correction method of other skeletal points is analogous. The inter-joint correction formula of the two skeletal points from the elbow to the wrist is: Where P wrist is the three-dimensional coordinate of the wrist bone point position predicted by the openpose model, P elbow The three-dimensional coordinates of the elbow bone point predicted by the openpose model. is the three-dimensional coordinate of the wrist position predicted by Kalman filtering, d max The maximum distance limit between these two bone points is preset.
10. A barbell fitness movement real-time detection and evaluation method as claimed in claim 8, characterized in that: By comparing the coordinates of the skeleton points in the same frame with the coordinates of the center point of the barbell, and making a comprehensive judgment based on the action hierarchical condition library, non-standard actions are eliminated. Here, the barbell fitness exercise type is comprehensively judged by taking the barbell squat type as an example, and other types are analogous; The comprehensive judgment criteria for the type of barbell squat include: 1) The z-axis height of the barbell center point coordinate in a complete movement cycle must not be lower than the height of the elbow bone point, that is, P center (t).z≥C elb (t).z, where P center (t).z is the coordinate of the center point of the barbell on the z-axis, C elb (t).z is the coordinate of the elbow bone point on the z-axis; 2) The z-axis height of the barbell center point coordinate in a complete movement cycle must not be lower than the preset height, that is, P center (t).z≥preset height; 3) The height of the wrist bone point must not be lower than the height of the hip bone point, that is, C wri (t).z>C hip (t).z, where C wri (t).z is the coordinate of the wrist bone point on the z axis, C hip (t).z is the coordinate of the hip bone point on the z-axis; 4) The center point of the barbell must be kept behind the tip of the nose bone, that is, P center (t).y>C nose (t).y,P center (t).y represents the y-axis coordinate of the center point of the barbell, C nose (t).y y-axis coordinate of the human nose tip bone point; The error tolerance space is set. For each skeleton point, a certain range of abnormal frames is set according to the probability of occurrence. When different conditions are met within a certain range of abnormal frames, the abnormal condition of the action is judged as passed.
Citation Information
Patent Citations
Speed and force feedback system based on depth camera
CN115937895A
Wearable fitness action recognition and index calculation method and device
CN119296164A
Intelligent barbell-oriented motion state recognition method
CN113577651A
Self-weight fitness auxiliary coach system, method and terminal based on human body posture recognition
CN113762133A
KR20220110383A
Cited By
Open vocabulary multi-target tracking method based on confidence adjustment and wavelet convolution
CN120807587A