A computer vision-based intelligent barbell movement monitoring method and system

By integrating multimodal feature analysis of RGB images, depth images, and IMU data, intelligent barbell motion monitoring was achieved, solving the problem of insufficient safety during training and improving the safety and standardization of user exercise.

CN122135430APending Publication Date: 2026-06-02QINGDAO QUANJIA SPORTS TECHNOLOGY CO LTD

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
QINGDAO QUANJIA SPORTS TECHNOLOGY CO LTD
Filing Date
2026-02-24
Publication Date
2026-06-02

AI Technical Summary

Technical Problem

During individual training, especially chest bench press exercises, excessive pressure can lead to decreased sensitivity to the external environment, making it easy for dangerous situations to occur. Current technology is insufficient to effectively monitor and prevent these dangerous situations.

Method used

By acquiring RGB image sequences, depth image sequences, and IMU data from multiple shooting angles, multimodal features are extracted and fused. A dual-stream analysis network is used for deep fusion and time modeling to determine the physically reasonable motion state that meets kinematic, dynamic, and energy constraints, thereby realizing intelligent barbell motion monitoring.

Benefits of technology

It improves safety during training by automatically monitoring the user's movement status to ensure that movement parameters conform to physiological and physical laws, thereby reducing the occurrence of dangers.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122135430A_ABST
    Figure CN122135430A_ABST
Patent Text Reader

Abstract

This invention relates to a computer vision-based intelligent barbell motion monitoring method and system, belonging to the field of image processing technology. The method mainly includes: acquiring RGB image sequences and depth image sequences from multiple shooting angles, as well as IMU data corresponding to the barbell and wearable device used by the user; fusing the RGB image sequences, depth image sequences, and IMU data to obtain multimodal features; inputting the multimodal features into a two-stream analysis network to obtain enhanced spatiotemporal features; the enhanced spatiotemporal features are feature representations obtained after deep fusion and time modeling through the two-stream network; determining physically reasonable motion states through the enhanced spatiotemporal features, wherein the physically reasonable motion state is a set of motion parameters that simultaneously satisfy kinematic, dynamic, and energy constraints; and monitoring the user's motion state based on the sequence of physically reasonable motion states.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of image processing technology, and in particular to an intelligent barbell motion monitoring method and system based on computer vision. Background Technology

[0002] Due to people's direct need for health, most people choose to train on their own for various inconveniences. However, when training on their own, the intensity of the training often makes it difficult to consider certain risk factors, especially during chest bench press exercises. Due to excessive pressure, people often become insensitive to the external environment, which can lead to dangerous situations and put themselves in danger. Summary of the Invention

[0003] The present invention aims to provide a computer vision-based intelligent barbell motion monitoring method and system to address the shortcomings of the existing technology. The technical problem to be solved by the present invention is achieved through the following technical solution.

[0004] This invention provides a computer vision-based intelligent barbell motion monitoring method, the method comprising: Acquire RGB image sequences, depth image sequences, and IMU data corresponding to the barbell and wearable device used by the user from multiple shooting angles; Multimodal features are obtained by fusing the RGB image sequence, the depth image sequence, and the IMU data; The multimodal features are input into a two-stream analysis network to obtain enhanced spatiotemporal features; the enhanced spatiotemporal features are feature representations obtained after deep fusion and temporal modeling through the two-stream network. The enhanced spatiotemporal features are used to determine a physically reasonable sequence of motion states, wherein the physically reasonable motion states are a set of motion parameters that simultaneously satisfy kinematic, dynamic and energy constraints. The user's motion state is monitored based on the physically reasonable motion state sequence.

[0005] In an optional embodiment, the step of fusing the RGB image sequence, the depth image sequence, and the IMU data to obtain multimodal features includes: RGB visual features are extracted from the RGB image sequence, including human 3D key points, barbell 2D detection box and semantic segmentation mask; Depth visual features are extracted from the depth image sequence, including dense depth field, barbell 3D spatial coordinates, and ground plane parameters; IMU motion features are extracted from the IMU data, including acceleration sequences, angular velocity sequences, and attitude quaternion sequences. Multimodal features are obtained by fusing the RGB visual features, the depth visual features, and the IMU motion features.

[0006] In an optional embodiment, the step of fusing the RGB visual features, the depth visual features, and the IMU motion features to obtain multimodal features includes: The weights corresponding to the RGB visual features, the depth visual features, and the IMU motion features are dynamically determined based on the current scene. The multimodal features are obtained by weighting the RGB visual features, the depth visual features, and the IMU motion features.

[0007] In an optional embodiment, the step of inputting the multimodal features into a two-stream analysis network to obtain enhanced spatiotemporal features includes: The multimodal features are input into the spatial flow branch and the temporal flow branch of the two-stream analysis network; The spatial feature matrix corresponding to the multimodal features is obtained through the spatial flow branch; the spatial feature matrix is ​​used to represent the spatial relationship between the human body and the barbell at a given time point. The time feature matrix corresponding to the multimodal features is obtained through the time flow branch; the time feature matrix is ​​used to represent the change process of motion over time. The enhanced spatiotemporal features are determined based on the spatial feature matrix and the temporal feature matrix.

[0008] In an optional embodiment, determining the enhanced spatiotemporal features based on the spatial feature matrix and the temporal feature matrix includes: Calculate the attention weight matrices corresponding to the spatial feature matrix and the temporal feature matrix, respectively; The enhanced spatiotemporal features are obtained by weighting the spatial feature matrix and the temporal feature matrix according to the attention weight matrix.

[0009] In an optional embodiment, determining the physically plausible sequence of motion states through the enhanced spatiotemporal features includes: The enhanced spatiotemporal features are input into a fully connected network to obtain an initial motion state sequence, which includes a human joint angle sequence and a barbell three-dimensional position sequence. By applying kinematic, dynamic, and energy constraints to the initial motion state sequence, a physically reasonable motion state sequence is obtained.

[0010] In an optional embodiment, the step of inputting the enhanced spatiotemporal features into a fully connected network to obtain an initial motion state sequence includes: The enhanced spatiotemporal features are input into a fully connected network, and the input layer in the fully connected network inputs the enhanced spatiotemporal features into the first prediction module and the second prediction module of the fully connected network, respectively. The human joint angle sequence is obtained through the first prediction module in the fully connected network; the three-dimensional position sequence of the barbell is obtained through the second prediction module in the fully connected network.

[0011] In an optional embodiment, monitoring the user's motion state based on the physically reasonable motion state sequence includes: The user's motion state is monitored based on the motion state, dynamic state, and energy state in the physically reasonable motion state.

[0012] In an optional embodiment, monitoring the user's motion state based on the motion state, dynamic state, and energy state in the physically reasonable motion state includes: Extract the data corresponding to the motion state, dynamic state, and energy state from the physically reasonable motion state; Based on the data corresponding to the motion state, dynamic state, and energy state in the physically reasonable motion state, calculate the standardization score, safety score, and motion performance score respectively. The user's exercise status is monitored based on the normative score, the safety score, and the exercise performance score.

[0013] This invention provides a computer vision-based intelligent barbell motion monitoring system, the system comprising: The acquisition module is used to acquire RGB image sequences, depth image sequences, and IMU data corresponding to the barbell and wearable device used by the user from multiple shooting angles. The fusion module is used to fuse the RGB image sequence, the depth image sequence, and the IMU data to obtain multimodal features; The prediction module is used to input the multimodal features into a two-stream analysis network to obtain enhanced spatiotemporal features; the enhanced spatiotemporal features are feature representations obtained after deep fusion and temporal modeling through the two-stream network. The determination module is used to determine a physically reasonable sequence of motion states through the enhanced spatiotemporal features, wherein the physically reasonable motion states are a set of motion parameters that simultaneously satisfy kinematic, dynamic and energy constraints. The monitoring module is used to monitor the user's motion state based on the physically reasonable motion state sequence.

[0014] The embodiments of the present invention have the following advantages: This invention provides a computer vision-based intelligent barbell motion monitoring method and system. First, it acquires RGB image sequences and depth image sequences from multiple shooting angles, as well as IMU data corresponding to the barbell and wearable device used by the user. Then, it fuses the RGB image sequences, depth image sequences, and IMU data to obtain multimodal features. These multimodal features are input into a two-stream analysis network to obtain enhanced spatiotemporal features. The enhanced spatiotemporal features are feature representations obtained after deep fusion and time modeling using the two-stream network. A physically reasonable motion state sequence is determined using the enhanced spatiotemporal features. This physically reasonable motion state is a set of motion parameters that simultaneously satisfy kinematic, dynamic, and energy constraints. Finally, the user's motion state is monitored based on the physically reasonable motion state sequence. Thus, this application enables the determination of a user's physically reasonable motion state sequence based on RGB image sequences, depth image sequences, and IMU data, facilitating automatic monitoring of the user's motion state and improving user safety during exercise. Attached Figure Description

[0015] Figure 1 This is a flowchart of an intelligent barbell motion monitoring method based on computer vision provided in an embodiment of the present invention; Figure 2 This is a schematic diagram of the structure of an intelligent barbell motion monitoring system based on computer vision provided in an embodiment of the present invention. Detailed Implementation

[0016] It should be noted that, unless otherwise specified, the embodiments and features described in this application can be combined with each other. The present invention will now be described in detail with reference to the accompanying drawings and embodiments.

[0017] Please see Figure 1 This invention provides a computer vision-based intelligent barbell motion monitoring method, specifically comprising steps S101-S105: S101 acquires RGB image sequences, depth image sequences, and IMU data corresponding to the barbell and wearable device used by the user from multiple shooting angles.

[0018] In this embodiment, the shooting angles may include a primary view (a side view perpendicular to the plane of the barbell's movement trajectory), used to clearly record the horizontal (forward and backward) and vertical (up and down) displacement of the barbell; a secondary view (facing the trainee directly in front), used to check whether the barbell is level, whether the body is tilted to the left or right, and whether the knees or elbows are buckling inward; and an auxiliary view (top / oblique), used to shoot from above or at a 45-degree angle. Different shooting angles provide different biomechanical information, and through multi-angle RGB image sequences and depth image sequences, the user's movement state can be realistically reflected, thereby improving the accuracy of motion monitoring in subsequent steps.

[0019] For example, in squats, the primary viewpoint can be used to analyze squat depth, center of gravity trajectory, knee-hip linkage, and whether the barbell moves in a straight line. The secondary viewpoint can be used to detect knee valgus and trunk twisting. In bench presses, the primary viewpoint is used to analyze the barbell's chest contact point, the pushing trajectory (whether it's vertical or has a slight arc), and scapular stability. The secondary viewpoint, positioned directly above or behind the head, is used to observe the barbell's left-right balance and arm symmetry. In deadlifts, the primary viewpoint is used to analyze the starting hip position, back angle, whether the barbell moves close to the legs, and the locked position. The secondary viewpoint (45 degrees diagonally forward) is used to assist in observing the overall posture of the body.

[0020] IMU (Inertial Measurement Unit) data can be obtained by attaching it to the ends or center of the barbell, or through wearable devices (such as arm straps or belts) worn by the user. IMU data includes triaxial acceleration, triaxial angular velocity, and magnetic field direction.

[0021] S102, multimodal features are obtained by fusing the RGB image sequence, the depth image sequence and the IMU data.

[0022] In one optional embodiment provided in this application, the step of fusing the RGB image sequence, the depth image sequence, and the IMU data to obtain multimodal features includes: S1021, Extract RGB visual features from the RGB image sequence, the RGB visual features including human 3D key points, barbell 2D detection box and semantic segmentation mask.

[0023] In this model, 3D key points of the human body (such as shoulders, elbows, hips, knees, and ankles) represent the spatial positions of human joints, the 2D bounding box of the barbell represents the range of the barbell in the image, and the semantic segmentation mask classifies the various parts of the scene (distinguishing between people, barbells, and background). This embodiment can extract RGB visual features using 3D convolutional neural networks or 2D CNN+LSTM to obtain the RGB visual features of each frame or segment.

[0024] S1022, Extract depth visual features from the depth image sequence, the depth visual features including dense depth field, barbell 3D spatial coordinates, and ground plane parameters.

[0025] In this embodiment, the dense depth field represents the 3D geometry of the scene, and the ground plane parameters represent the ground equations. Precise 3D coordinates can be obtained through depth completion and point cloud processing.

[0026] S1023, Extract IMU motion features from the IMU data, the IMU motion features including acceleration sequence, angular velocity sequence, and attitude quaternion sequence.

[0027] Acceleration sequences are used to represent linear motion changes, angular velocity sequences are used to represent rotational motion changes, and attitude quaternion sequences are used to represent device orientation. This implementation can use a 1D convolutional neural network or a recurrent neural network to extract IMU motion features.

[0028] S1024, multimodal features are obtained by fusing the RGB visual features, the depth visual features, and the IMU motion features.

[0029] Specifically, the process of fusing the RGB visual features, depth visual features, and IMU motion features to obtain multimodal features includes: dynamically determining the weights corresponding to the RGB visual features, depth visual features, and IMU motion features based on the current scene; and calculating the multimodal features by weighting the RGB visual features, depth visual features, and IMU motion features. The dynamics of the current scene may include lighting conditions, occlusion levels, and movement speed. When lighting is poor, the RGB weights are reduced and the depth weights are increased; when occlusion is severe, the visual weights are reduced and the IMU weights are increased; and during rapid movement, the IMU weights are increased (due to their high sampling rate).

[0030] In this embodiment, since RGB and depth images are usually synchronized, and the IMU has a higher sampling rate, it is necessary to downsample the IMU features or upsample the image features to align the RGB visual features, depth visual features, and IMU motion features on the time axis. Therefore, this embodiment first aligns the features of different modalities in time, and then adjusts and normalizes the dimensions of each modality's features to facilitate fusion into multimodal features. If the features of each modality are aligned, three feature sequences are obtained: an RGB visual feature sequence, a depth visual feature sequence, and an IMU motion feature sequence. For each time step t, the three feature vectors are concatenated to form a multimodal feature vector. The concatenated feature vector may have a high dimension, which can be reduced and fused using a fully connected layer to obtain the fused multimodal features.

[0031] S103, the multimodal features are input into a two-stream analysis network to obtain enhanced spatiotemporal features; the enhanced spatiotemporal features are feature representations obtained after deep fusion and time modeling through a two-stream network.

[0032] In one optional embodiment provided in this application, the step of inputting the multimodal features into a two-stream analysis network to obtain enhanced spatiotemporal features includes: S1031 inputs multimodal features into the spatial flow branch and temporal flow branch in the two-stream analysis network.

[0033] S1032, the spatial feature matrix corresponding to the multimodal feature is obtained through the spatial flow branch; the spatial feature matrix is ​​used to represent the spatial relationship between the human body and the barbell at a given time point.

[0034] The spatial flow branch can use 1D convolutional layers with kernels sliding along the temporal direction but acting on the entire feature dimension. Thus, the output of each layer remains a sequence, but the features of each frame are enhanced, allowing feature fusion within a local temporal window to capture short-term temporal changes in spatial features. The features output by the spatial flow branch primarily contain spatial information, such as human pose, barbell position, and spatial geometric relationships. It focuses on the spatial structure within each frame or time step.

[0035] S1033, the time feature matrix corresponding to the multimodal feature is obtained through the time flow branch; the time feature matrix is ​​used to represent the change process of motion over time.

[0036] The temporal branch can use a bidirectional LSTM (Bi-LSTM) to capture long-term temporal dependencies. The multimodal feature sequence is input into the Bi-LSTM to obtain the hidden states at each time step; these hidden states contain long-term contextual information. The features output by the temporal branch mainly contain temporal information, such as the rhythm of the action, speed changes, and stage transitions.

[0037] S1034, Determine the enhanced spatiotemporal features based on the spatial feature matrix and the temporal feature matrix.

[0038] Specifically, determining the enhanced spatiotemporal features based on the spatial feature matrix and the temporal feature matrix includes: calculating the attention weight matrices corresponding to the spatial feature matrix and the temporal feature matrix respectively; and performing a weighted calculation on the spatial feature matrix and the temporal feature matrix based on the attention weight matrices to obtain the enhanced spatiotemporal features.

[0039] In this embodiment, the spatial feature matrix S, It can include spatial structural information (human posture, joint angles, relative positions of body parts), geometric relationship information (spatial relationship, distance, and angle between the human body and the barbell), trajectory information (spatial motion trajectory of key points), static features, etc.; temporal feature matrix T, It can include temporal evolution information (action rhythm, velocity changes, acceleration patterns), phase information (action stages (centrifugal, centripetal, transition)), repetition patterns (periodicity and repetitive characteristics of the action), and dynamic information (power and energy change trends), etc. Here, T is the number of time steps. In this embodiment, the spatial feature matrix and the temporal feature matrix are fused into enhanced spatiotemporal features.

[0040] Specifically, this embodiment employs two cross-attention modules: using spatial features as queries and temporal features as keys and values, it calculates the degree of attention that spatial features have to temporal features, thus obtaining the attention features. Using time features as queries and spatial features as keys and values, the degree of attention given to spatial features by time features is calculated, thus obtaining the attention features. Then, an attention weight matrix is ​​generated for the two attention features, and the spatial feature matrix and the temporal feature matrix are weighted according to the attention weight matrix to obtain the enhanced spatiotemporal features.

[0041] This embodiment uses scaled dot product attention, and the specific formula is as follows:

[0042] Where Q is the query matrix, K is the key matrix, V is the value matrix, and d is the feature dimension (d=512). Specifically, the similarity matrix is ​​obtained by calculating the dot product of Q and K: (Shape is T×T), scaling Perform softmax normalization on each row of sim to obtain the attention weight matrix (T×T). Multiply the attention weight matrix by V (i.e., T') to obtain the output. (Shape is T×512).

[0043] For example, suppose T=3, d=512, suppose d=4, suppose there are three time steps, and the feature vector dimension of each time step is 4.

[0044] S' = [ [1,0,0,0], # Spatial features at time step 1

[0045] [0,1,0,0], # Spatial features of time step 2

[0046] [0,0,1,0] ] # Spatial features of time step 3

[0047] T' = [ [0,0,1,0], # Time characteristics of time step 1

[0048] [0,0,0,1], # Time characteristics of time step 2

[0049] [1,0,0,0] ] # Time characteristics of time step 3

[0050] calculate The process is as follows: Calculate the similarity matrix : # The dot product of the first time step of S' and each time step of T' in the first row # Second line # Third line

[0051] Divide by

[0052] Softmax normalization: First row: exp(0)=1, exp(0)=1, exp(0.5)=1.6487, summing to 3.6487, normalized to [0.274, 0.274, 0.452]; Second row: all zeros, each position after softmax is 1 / 3≈0.333; Third row: exp(0.5)=1.6487, exp(0)=1, exp(0)=1, summing to 3.6487, normalized to [0.452, 0.274, 0.274], therefore the attention weight matrix is:

[0053] Multiply the attention weight by V (i.e., T') to get : First line

[0054] Second line

[0055] Third line The results represent the aggregation of information about temporal features (values) by focusing on temporal features (keys) for each spatial feature (as a query).

[0056] S104, determine a physically reasonable sequence of motion states through the enhanced spatiotemporal features, wherein the physically reasonable motion states are a set of motion parameters that simultaneously satisfy kinematic, dynamic and energy constraints.

[0057] In barbell exercise monitoring, a physically reasonable motion state includes: the angles of the human joints are within a reasonable physiological range, the trajectory of the barbell's motion conforms to the laws of dynamics (such as the speed and acceleration of the barbell matching the force conditions during the lifting process), and the interaction between the human posture and the barbell conforms to kinematic constraints (such as the contact points between the hands and the barbell being relatively fixed, and the barbell not passing through the body during the exercise).

[0058] In this embodiment, the initial motion state (including human joint angles, barbell position, velocity, acceleration, etc.) can be decoded from the enhanced spatiotemporal features. Then, physical constraints and kinematic constraints are defined, and an optimization problem is constructed to minimize the difference between the decoded state and the observed data while satisfying the constraints. Finally, the optimization problem is solved to obtain a physically reasonable motion state.

[0059] Specifically, the enhanced spatiotemporal features are mapped to motion states through a decoder neural network, and the enhanced spatiotemporal features are... Where T is the number of time steps and D is the feature dimension, the decoder can be represented as: , The initial motion state obtained from decoding includes the following variables for each time step t: Human joint angles: J is the number of joints; Barbell position: The coordinates of the barbell in three-dimensional space; Barbell speed: The velocity vector of the barbell; Barbell acceleration: , the acceleration vector of the barbell.

[0060] Joint angle constraints mean that each joint angle should be within the physiological range, i.e. ; The dynamic constraint is that, according to Newton's second law, the acceleration of the barbell should be proportional to the net force. In the vertical direction, the net force on the barbell is:

[0061] Where m is the mass of the barbell, and g is the acceleration due to gravity. The force exerted by the human body on the barbell can be estimated from the body's posture and the position of the barbell. .

[0062] Kinematic constraints require that the positions of the human joints and the barbell satisfy a kinematic chain relationship. For example, the position of the hand joints should be within a certain distance from the barbell (because the hand grips the barbell). Let the position of the hand joints be... ,but .

[0063] The motion smoothness constraint requires that the motion state change smoothly over time, meaning that the state change between adjacent time steps should not be too large. For example, , .

[0064] Define an objective function containing data terms and a smoothing term, and add the constraints mentioned above to ensure that the optimized motion state S is consistent with the initial decoded state. Get as close as possible.

[0065] In one optional embodiment provided in this application, determining the physically plausible motion state sequence through the enhanced spatiotemporal features includes: S1041, The enhanced spatiotemporal features are input into a fully connected network to obtain an initial motion state sequence, which includes a human joint angle sequence and a barbell three-dimensional position sequence.

[0066] Specifically, the step of inputting the enhanced spatiotemporal features into a fully connected network to obtain an initial motion state sequence includes: inputting the enhanced spatiotemporal features into a fully connected network; the input layer of the fully connected network inputs the enhanced spatiotemporal features into the first prediction module and the second prediction module of the fully connected network, respectively; obtaining a human joint angle sequence through the first prediction module of the fully connected network; and obtaining a barbell three-dimensional position sequence through the second prediction module of the fully connected network.

[0067] S1042, apply kinematic constraints, dynamic constraints, and energy constraints to the initial motion state sequence to obtain a physically reasonable motion state sequence.

[0068] In this embodiment, the physically reasonable sequence of motion states is obtained by applying physical constraints to the initial sequence of motion states, but it is not completely detached from the original data. The objective function of optimization typically includes two terms: one is to get as close as possible to the original estimate (data term), and the other is to satisfy the physical constraints (regularization term or constraint term). Specifically, regarding kinematic constraints: the angle of each joint must be within the physiological range, the state change between adjacent time steps cannot be too large (to ensure smoothness), and the relative position between joints must conform to the fact that the length of the human skeleton is fixed; regarding dynamic constraints: the movement of the barbell and various parts of the human body must be generated by forces and torques and conform to the dynamic equations; when standing on both feet, the ground reaction force must pass within the support polygon formed by the feet to maintain balance, and the joint torques must be within the physiological limits; regarding energy constraints: the change in mechanical energy must be equal to the work done by the muscles minus the dissipated energy.

[0069] For example, if a user's actual movement is a squat where their knees valgus significantly and the barbell trajectory moves forward, the optimization process will handle the following: Knee valgus angle: If the knee valgus angle is within the physiological range (e.g., ≤15°), the optimization will retain this feature; Barbell trajectory: Ensure the trajectory is smooth and continuous, but will not force it back to an ideal straight line; Joint angle: Ensure it does not exceed physiological limits, but will not adjust it to a standard angle.

[0070] Physical constraint optimization doesn't mask all errors; instead, it corrects those that clearly violate physical laws. For example, if the original estimate shows a sudden change in a joint angle due to occlusion or other reasons, optimization will smooth it out and make it conform to the joint angle range. However, user technical errors (such as knee valgus or back flexion) will not be forcibly corrected during optimization as long as they are within physically reasonable limits. This is because the optimization goal is only to ensure the physical reasonableness of the motion state, not to standardize the movement. In other words, the optimized motion state sequence still retains the user's actual motion characteristics, only removing those obviously impossible states (such as joint dislocation or violations of Newton's laws).

[0071] S105, monitor the user's motion state based on the physically reasonable motion state sequence.

[0072] In this embodiment, monitoring the user's motion state based on the physically reasonable motion state sequence includes: monitoring the user's motion state based on the motion state, dynamic state, and energy state in the physically reasonable motion state.

[0073] Specifically, monitoring the user's motion state based on the motion state, dynamic state, and energy state in the physically reasonable motion state includes: extracting data corresponding to the motion state, dynamic state, and energy state from the physically reasonable motion state; calculating a standardization score, a safety score, and a motion performance score based on the data corresponding to the motion state, dynamic state, and energy state in the physically reasonable motion state; and monitoring the user's motion state based on the standardization score, the safety score, and the motion performance score.

[0074] This embodiment provides a computer vision-based intelligent barbell motion monitoring method. First, it acquires RGB image sequences and depth image sequences from multiple shooting angles, as well as IMU data corresponding to the barbell and wearable device used by the user. Then, it fuses the RGB image sequences, depth image sequences, and IMU data to obtain multimodal features. These multimodal features are input into a two-stream analysis network to obtain enhanced spatiotemporal features. These enhanced spatiotemporal features are feature representations obtained after deep fusion and time modeling using the two-stream network. A physically reasonable motion state sequence is determined using these enhanced spatiotemporal features. This physically reasonable motion state is a set of motion parameters that simultaneously satisfy kinematic, dynamic, and energy constraints. Finally, the user's motion state is monitored based on this physically reasonable motion state sequence. Therefore, this application enables the determination of a user's physically reasonable motion state sequence based on RGB image sequences, depth image sequences, and IMU data, facilitating automatic monitoring of the user's motion state and thus improving the safety of user movement.

[0075] It should be understood that the sequence number of each step in the above embodiments does not imply the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present invention.

[0076] In one embodiment, a computer vision-based intelligent barbell motion monitoring system is provided. For example... Figure 2 As shown, the system includes: The acquisition module 21 is used to acquire RGB image sequences, depth image sequences, and IMU data corresponding to the barbell and wearable device used by the user from multiple shooting angles. The fusion module 22 is used to fuse the RGB image sequence, the depth image sequence, and the IMU data to obtain multimodal features; Prediction module 23 is used to input the multimodal features into a two-stream analysis network to obtain enhanced spatiotemporal features; the enhanced spatiotemporal features are feature representations obtained after deep fusion and time modeling through the two-stream network; The determination module 24 is used to determine a physically reasonable sequence of motion states through the enhanced spatiotemporal features, wherein the physically reasonable motion states are a set of motion parameters that simultaneously satisfy kinematic, dynamic and energy constraints; Monitoring module 25 is used to monitor the user's motion state based on the physically reasonable motion state sequence.

[0077] In an optional embodiment, the fusion module 22 is specifically used for: RGB visual features are extracted from the RGB image sequence, including human 3D key points, barbell 2D detection box and semantic segmentation mask; Depth visual features are extracted from the depth image sequence, including dense depth field, barbell 3D spatial coordinates, and ground plane parameters; IMU motion features are extracted from the IMU data, including acceleration sequences, angular velocity sequences, and attitude quaternion sequences. Multimodal features are obtained by fusing the RGB visual features, the depth visual features, and the IMU motion features.

[0078] In an optional embodiment, the fusion module 22 is specifically used for: The weights corresponding to the RGB visual features, the depth visual features, and the IMU motion features are dynamically determined based on the current scene. The multimodal features are obtained by weighting the RGB visual features, the depth visual features, and the IMU motion features.

[0079] In an optional embodiment, the prediction module 23 is specifically used for: The multimodal features are input into the spatial flow branch and the temporal flow branch of the two-stream analysis network; The spatial feature matrix corresponding to the multimodal features is obtained through the spatial flow branch; the spatial feature matrix is ​​used to represent the spatial relationship between the human body and the barbell at a given time point. The time feature matrix corresponding to the multimodal features is obtained through the time flow branch; the time feature matrix is ​​used to represent the change process of motion over time. The enhanced spatiotemporal features are determined based on the spatial feature matrix and the temporal feature matrix.

[0080] In an optional embodiment, the prediction module 23 is specifically used for: Calculate the attention weight matrices corresponding to the spatial feature matrix and the temporal feature matrix, respectively; The enhanced spatiotemporal features are obtained by weighting the spatial feature matrix and the temporal feature matrix according to the attention weight matrix.

[0081] In an optional embodiment, the determining module 24 is specifically used for: The enhanced spatiotemporal features are input into a fully connected network to obtain an initial motion state sequence, which includes a human joint angle sequence and a barbell three-dimensional position sequence. By applying kinematic, dynamic, and energy constraints to the initial motion state sequence, a physically reasonable motion state sequence is obtained.

[0082] In an optional embodiment, the determining module 24 is specifically used for: The enhanced spatiotemporal features are input into a fully connected network, and the input layer in the fully connected network inputs the enhanced spatiotemporal features into the first prediction module and the second prediction module of the fully connected network, respectively. The human joint angle sequence is obtained through the first prediction module in the fully connected network; the three-dimensional position sequence of the barbell is obtained through the second prediction module in the fully connected network.

[0083] In an optional embodiment, the monitoring module 25 is specifically used for: The user's motion state is monitored based on the motion state, dynamic state, and energy state in the physically reasonable motion state.

[0084] In an optional embodiment, the monitoring module 25 is specifically used for: Extract the data corresponding to the motion state, dynamic state, and energy state from the physically reasonable motion state; Based on the data corresponding to the motion state, dynamic state, and energy state in the physically reasonable motion state, calculate the standardization score, safety score, and motion performance score respectively. The user's exercise status is monitored based on the normative score, the safety score, and the exercise performance score.

[0085] It should be noted that the above detailed descriptions are exemplary and intended to provide further explanation of this application. Unless otherwise specified, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this application pertains.

[0086] Specific limitations regarding the computer vision-based intelligent barbell motion monitoring system can be found in the limitations of the computer vision-based intelligent barbell motion monitoring method described above, and will not be repeated here. Each module in the aforementioned device can be implemented entirely or partially through software, hardware, or a combination thereof. These modules can be embedded in the processor of the computer device in hardware form or independent of it, or stored in the memory of the computer device in software form, so that the processor can call and execute the corresponding operations of each module.

[0087] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the above-described division of functional units and modules is used as an example. In practical applications, the above functions can be assigned to different functional units and modules as needed, that is, the internal structure of the system can be divided into different functional units or modules to complete all or part of the functions described above.

[0088] The above embodiments are only used to illustrate the technical solutions of the present invention, and are not intended to limit it. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention, and should all be included within the protection scope of the present invention.

Claims

1. A computer vision-based intelligent barbell motion monitoring method, characterized in that, The method includes: Acquire RGB image sequences, depth image sequences, and IMU data corresponding to the barbell and wearable device used by the user from multiple shooting angles; Multimodal features are obtained by fusing the RGB image sequence, the depth image sequence, and the IMU data; The multimodal features are input into a two-stream analysis network to obtain enhanced spatiotemporal features; the enhanced spatiotemporal features are feature representations obtained after deep fusion and temporal modeling through the two-stream network. The enhanced spatiotemporal features are used to determine a physically reasonable sequence of motion states, wherein the physically reasonable motion states are a set of motion parameters that simultaneously satisfy kinematic, dynamic and energy constraints. The user's motion state is monitored based on the physically reasonable motion state sequence.

2. The method according to claim 1, characterized in that, The process of fusing the RGB image sequence, the depth image sequence, and the IMU data to obtain multimodal features includes: RGB visual features are extracted from the RGB image sequence, including human 3D key points, barbell 2D detection box and semantic segmentation mask; Depth visual features are extracted from the depth image sequence, including dense depth field, barbell 3D spatial coordinates, and ground plane parameters; IMU motion features are extracted from the IMU data, including acceleration sequences, angular velocity sequences, and attitude quaternion sequences. Multimodal features are obtained by fusing the RGB visual features, the depth visual features, and the IMU motion features.

3. The method according to claim 2, characterized in that, The process of fusing the RGB visual features, the depth visual features, and the IMU motion features to obtain multimodal features includes: The weights corresponding to the RGB visual features, the depth visual features, and the IMU motion features are dynamically determined based on the current scene. The multimodal features are obtained by weighting the RGB visual features, the depth visual features, and the IMU motion features.

4. The method according to claim 3, characterized in that, The step of inputting the multimodal features into a two-stream analysis network to obtain enhanced spatiotemporal features includes: The multimodal features are input into the spatial flow branch and the temporal flow branch of the two-stream analysis network; The spatial feature matrix corresponding to the multimodal features is obtained through the spatial flow branch; the spatial feature matrix is ​​used to represent the spatial relationship between the human body and the barbell at a given time point. The time feature matrix corresponding to the multimodal features is obtained through the time flow branch; the time feature matrix is ​​used to represent the change process of motion over time. The enhanced spatiotemporal features are determined based on the spatial feature matrix and the temporal feature matrix.

5. The method according to claim 3, characterized in that, Determining the enhanced spatiotemporal features based on the spatial feature matrix and the temporal feature matrix includes: Calculate the attention weight matrices corresponding to the spatial feature matrix and the temporal feature matrix, respectively; The enhanced spatiotemporal features are obtained by weighting the spatial feature matrix and the temporal feature matrix according to the attention weight matrix.

6. The method according to any one of claims 1-5, characterized in that, The process of determining a physically plausible sequence of motion states through the enhanced spatiotemporal features includes: The enhanced spatiotemporal features are input into a fully connected network to obtain an initial motion state sequence, which includes a human joint angle sequence and a barbell three-dimensional position sequence. By applying kinematic, dynamic, and energy constraints to the initial motion state sequence, a physically reasonable motion state sequence is obtained.

7. The method according to claim 6, characterized in that, The step of inputting the enhanced spatiotemporal features into a fully connected network to obtain the initial motion state sequence includes: The enhanced spatiotemporal features are input into a fully connected network, and the input layer in the fully connected network inputs the enhanced spatiotemporal features into the first prediction module and the second prediction module of the fully connected network, respectively. The human joint angle sequence is obtained through the first prediction module in the fully connected network; the three-dimensional position sequence of the barbell is obtained through the second prediction module in the fully connected network.

8. The method according to any one of claims 1-5, characterized in that, The monitoring of the user's motion state based on the physically reasonable motion state sequence includes: The user's motion state is monitored based on the motion state, dynamic state, and energy state in the physically reasonable motion state.

9. The method according to claim 8, characterized in that, The monitoring of the user's motion state based on the motion state, dynamic state, and energy state in the physically reasonable motion state includes: Extract the data corresponding to the motion state, dynamic state, and energy state from the physically reasonable motion state; Based on the data corresponding to the motion state, dynamic state, and energy state in the physically reasonable motion state, calculate the standardization score, safety score, and motion performance score respectively. The user's exercise status is monitored based on the normative score, the safety score, and the exercise performance score.

10. A computer vision-based intelligent barbell motion monitoring system, characterized in that, The system includes: The acquisition module is used to acquire RGB image sequences, depth image sequences, and IMU data corresponding to the barbell and wearable device used by the user from multiple shooting angles. The fusion module is used to fuse the RGB image sequence, the depth image sequence, and the IMU data to obtain multimodal features; The prediction module is used to input the multimodal features into a two-stream analysis network to obtain enhanced spatiotemporal features; the enhanced spatiotemporal features are feature representations obtained after deep fusion and temporal modeling through the two-stream network. The determination module is used to determine a physically reasonable sequence of motion states through the enhanced spatiotemporal features, wherein the physically reasonable motion states are a set of motion parameters that simultaneously satisfy kinematic, dynamic and energy constraints. The monitoring module is used to monitor the user's motion state based on the physically reasonable motion state sequence.