Intelligent training device for basketball dribbling rhythm quantitative detection and action correction
By combining an IMU (Inertial Measurement Unit) and a surface electromyography (EMG) sensor with a multimodal feature extraction model, the problem of not being able to quantify rhythm and joint angle deviation in traditional basketball dribbling training is solved, enabling real-time visual feedback and personalized training guidance.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-01-07
- Publication Date
- 2026-03-31
AI Technical Summary
Traditional basketball dribbling training cannot quantify dribbling rhythm and joint angle deviations, lacks real-time visual feedback, resulting in training effects that vary from person to person, and corrective suggestions that lack scientific basis.
Multimodal data acquisition is performed using an IMU (Inertial Measurement Unit) and a surface electromyography (EMG) sensor. Feature extraction and evaluation are then performed using a CNN-LSTM-Attention multimodal fusion model to generate real-time visualization and voice feedback, supporting personalized training and long-term progress tracking.
It enables precise quantitative detection and real-time movement correction of basketball dribbling actions, improving the scientific nature of training and personalized feedback.
Smart Images

Figure CN121754868A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of intelligent training device technology, and in particular to an intelligent training device for quantitative detection and motion correction of basketball dribbling rhythm. Background Technology
[0002] Traditional basketball dribbling training relies on coaches' visual observation, which can only qualitatively judge whether the movement is standard. It cannot quantify key indicators such as dribbling rhythm (temporal dynamic characteristics), joint angle deviation, and muscle force intensity, resulting in training effects that vary from person to person and correction suggestions that lack scientific basis. Existing training equipment either only collects video image data or only collects motion data from sensors. Single-modal data cannot comprehensively reflect whether the dribbling movement is standard, and feedback from traditional training is mostly summative guidance after training, so users cannot perceive movement deviations in real time during dribbling. Some smart devices only provide feedback through voice prompts, lacking visual movement comparison, making it difficult for users to quickly locate deviations in joints and rhythm. Summary of the Invention
[0003] The purpose of this invention is to provide an intelligent training device for quantitative detection and movement correction of basketball dribbling rhythm. It is equipped with an IMU inertial measurement unit and a surface electromyography sensor to achieve accurate quantitative detection of basketball dribbling movements. It is equipped with a multimodal feature extraction and evaluation module to improve the accuracy and reliability of movement evaluation. It is equipped with a feedback generation module to achieve real-time and intuitive movement correction and feedback. It is equipped with an interaction and management module to support personalized training and long-term progress tracking.
[0004] To achieve the above objectives, the present invention provides a smart training device for quantitative detection and movement correction of basketball dribbling rhythm, including a wearable device, a training device and a basketball. The wearable device is equipped with an IMU inertial measurement unit and a surface electromyography sensor. The training device includes a housing and a control system. The control system is installed inside the housing and includes a multi-source data acquisition module, a data preprocessing module, a multimodal feature extraction and evaluation module, a feedback generation module and an interaction and management module. The multi-data acquisition module includes a video acquisition unit, a sensor acquisition unit, and a data transmission unit; the video acquisition unit acquires user motion images and video data through a camera, and acquires frame data of teaching and competition videos; the sensor acquisition unit reads the acceleration, angular velocity, and electromyography signals of the wearable device; The data preprocessing module transforms the raw data into identifiable, standardized time-series features, eliminating noise and heterogeneous data differences; The multimodal feature extraction and evaluation module completes feature extraction and action evaluation by constructing a CNN-LSTM-Attention multimodal fusion model; The feedback generation module transforms the evaluation results into video and voice feedback that users can perceive; The interaction and management module enables users to interact with the control system, while also managing the standard action library and user history data.
[0005] Preferably, a display screen is mounted on the surface of the housing, control buttons are mounted on the housing at the bottom of the display screen, and speakers are mounted on the side panels of the housing on both sides of the display screen.
[0006] Preferably, the data preprocessing module extracts key points of the human skeleton based on the OpenPose algorithm to generate a time series of joint coordinates, as shown below: ; in, The frame number, For the number of joints, For the first Frame time The x-coordinate of each joint, For the first Frame time The ordinates of the joints, where ; The data preprocessing module performs Kalman filtering on the acceleration and angular velocity data to eliminate motion noise, performs frequency domain analysis on the electromyography (EMG) signals to extract force features, and aligns the time axes of the EMG signals, acceleration, and angular velocity data with the time axes of the video frames to generate a synchronized sensor time sequence, as shown below: ; in, For the first Frame time sensor acquisition Axial acceleration components, For the first Frame time sensor acquisition Axial acceleration components, For the first Frame time sensor acquisition Axial acceleration components, For the first Frame time sensor acquisition axial angular velocity components; The time series sequences of joint coordinates and sensor time series are normalized to eliminate the influence of individual differences such as user height and weight. The normalization calculation formula is shown below: ; in, The original variable values to be normalized. This represents the minimum value of the original variable in the corresponding dataset. This represents the maximum value of the original variable in the corresponding dataset.
[0007] Preferably, the multimodal feature extraction and evaluation module takes preprocessed joint coordinate time series and sensor time series as input, and extracts joint spatial topological features using CNN. LSTM extracts temporal dynamic features of actions, including the temporal dynamic features of joint coordinate sequences. Temporal dynamic characteristics of sensor time series The attention mechanism assigns higher weights to key action phases, ultimately generating a multimodal fusion feature vector. , To fuse feature dimensions; Attention mechanisms through attention weights Multimodal features are integrated to enhance the feature weights of key action phases. The calculation formula is shown below: ; Among them, the attention weights satisfy the constraints. ; The labeled multimodal feature data of standard actions are used to cluster similar action features and disperse dissimilar action features through contrastive loss; the formula for calculating contrastive loss is as follows: ; in, As an action category identifier, hour, and They are similar actions. hour, and It's an unusual action. and These are the feature vectors corresponding to two actions. for and The square of the Euclidean distance, The inter-class distance threshold; The standard action feature vectors extracted by the multimodal fusion model are stored and clustered to form a structured standard action library, and user action features are calculated. With standard feature center The cosine similarity is used to determine whether an action is standard, as shown below: ; like If it is a standard movement, it is considered a non-standard movement; otherwise, it is considered a non-standard movement. The preset threshold; Standard feature centers are generated by clustering multiple sets of standard features for the same action using the K-means clustering algorithm, as shown below: ; in, Multiple sets of standard features for the same action The number of action subclasses; For non-standard movements, calculate the quantization deviation of the joints and sensors, where the quantization deviation of the joints is the user's joint angle. Compared with standard joint angle The deviation value, the quantization deviation of the sensor is the user sensor data. Data range compared to standard sensors The deviation.
[0008] Preferably, the feedback generation module includes a video feedback submodule and a voice feedback submodule; the video feedback submodule synchronously overlays the user's standard motion skeleton diagram with the standard motion skeleton diagram, marks the standard joints in green and the deviated joints in red, and displays the deviation angle value in real time, generating a comparison video file and projecting it on the display screen; The voice feedback submodule generates natural language guidance text based on the quantized deviation value, and then converts the natural language guidance text into voice prompts that are played in real time by the speaker.
[0009] Therefore, the present invention employs the above-mentioned intelligent training device for quantitative detection and action correction of basketball dribbling rhythm, which sets up an IMU inertial measurement unit and a surface electromyography sensor to achieve accurate quantitative detection of basketball dribbling action, sets up a multimodal feature extraction and evaluation module to improve the accuracy and reliability of action evaluation, sets up a feedback generation module to achieve real-time and intuitive action correction and feedback, and sets up an interaction and management module to support personalized training and long-term progress tracking.
[0010] The technical solution of the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. Attached Figure Description
[0011] Figure 1 This is a schematic diagram of the structure of an intelligent training device for quantitative detection and motion correction of basketball dribbling rhythm according to the present invention.
[0012] Figure Labels 1. Training equipment; 11. Display screen; 12. Control buttons; 13. Speaker; 14. Camera; 2. Basketball; 3. Wearable device. Detailed Implementation
[0013] The technical solution of the present invention will be further described below with reference to the accompanying drawings and embodiments.
[0014] Unless otherwise defined, the technical or scientific terms used in this invention shall have the ordinary meaning understood by one of ordinary skill in the art to which this invention pertains. The terms "first," "second," and similar terms used in this invention do not indicate any order, quantity, or importance, but are merely used to distinguish different components. Terms such as "comprising" or "including" mean that the element or object preceding the word encompasses the elements or objects listed following the word and their equivalents, without excluding other elements or objects. Terms such as "connected" or "linked" are not limited to physical or mechanical connections, but can include electrical connections, whether direct or indirect. Terms such as "upper," "lower," "left," and "right" are used only to indicate relative positional relationships; when the absolute position of the described object changes, the relative positional relationship may also change accordingly.
[0015] Example 1 like Figure 1 As shown, the present invention provides a smart training device for quantitative detection and movement correction of basketball dribbling rhythm, including a wearable device 3, a training device 1 and a basketball 2. The wearable device 3 is equipped with an IMU inertial measurement unit and a surface electromyography sensor. The training device 1 includes a housing and a control system. A display screen 11 is installed on the surface of the housing. A control button 12 is installed on the housing at the bottom of the display screen 11. Speakers 13 are installed on the side panels of the housing on both sides of the display screen 11.
[0016] The control system is installed inside the housing and includes a multi-source data acquisition module, a data preprocessing module, a multi-modal feature extraction and evaluation module, a feedback generation module, and an interaction and management module. The multi-source data acquisition module includes a video acquisition unit, a sensor acquisition unit, and a data transmission unit. The video acquisition unit acquires user motion images and video data through camera 14, and acquires frame data of teaching and competition videos. The sensor acquisition unit reads the acceleration, angular velocity, and electromyography signals of the wearable device.
[0017] The data preprocessing module transforms the raw data into identifiable, standardized temporal features, eliminating noise and heterogeneous data differences. Based on the OpenPose algorithm, the module extracts key points of the human skeleton to generate a temporal sequence of joint coordinates, as shown below: ; in, The frame number, For the number of joints, For the first Frame time The x-coordinate of each joint, For the first Frame time The ordinates of the joints, where .
[0018] The data preprocessing module performs Kalman filtering on the acceleration and angular velocity data to eliminate motion noise, performs frequency domain analysis on the electromyography (EMG) signals to extract force features, and aligns the time axes of the EMG signals, acceleration, and angular velocity data with the time axes of the video frames to generate a synchronized sensor time sequence, as shown below: ; in, For the first Frame time sensor acquisition Axial acceleration components, For the first Frame time sensor acquisition Axial acceleration components, For the first Frame time sensor acquisition Axial acceleration components, For the first Frame time sensor acquisition Angular velocity component of the shaft.
[0019] The time series sequences of joint coordinates and sensor time series are normalized to eliminate the influence of individual differences such as user height and weight. The normalization calculation formula is shown below: ; in, The original variable values to be normalized. This represents the minimum value of the original variable in the corresponding dataset. This represents the maximum value of the original variable in the corresponding dataset.
[0020] The multimodal feature extraction and evaluation module completes feature extraction and action evaluation by constructing a CNN-LSTM-Attention multimodal fusion model. The module takes preprocessed temporal sequences of joint coordinates and sensor temporal sequences as input and extracts joint spatial topological features using a CNN. LSTM extracts temporal dynamic features of actions, including the temporal dynamic features of joint coordinate sequences. Temporal dynamic characteristics of sensor time series The attention mechanism assigns higher weights to key action phases, ultimately generating a multimodal fusion feature vector. , To fuse feature dimensions.
[0021] Attention mechanisms through attention weights Multimodal features are integrated to enhance the feature weights of key action phases. The calculation formula is shown below: ; Among them, the attention weights satisfy the constraints. .
[0022] The labeled multimodal feature data of standard actions are used to cluster similar action features and disperse dissimilar action features through contrastive loss; the formula for calculating contrastive loss is as follows: ; in, As an action category identifier, hour, and They are similar actions. hour, and It's an unusual action. and These are the feature vectors corresponding to two actions. for and The square of the Euclidean distance, This is the inter-class distance threshold.
[0023] The standard action feature vectors extracted by the multimodal fusion model are stored and clustered to form a structured standard action library, and user action features are calculated. With standard feature center The cosine similarity is used to determine whether an action is standard, as shown below: ; like If it is a standard movement, it is considered a non-standard movement; otherwise, it is considered a non-standard movement. This is a preset threshold.
[0024] Standard feature centers are generated by clustering multiple sets of standard features for the same action using the K-means clustering algorithm, as shown below: ; in, Multiple sets of standard features for the same action Let be the number of action subclasses.
[0025] For non-standard movements, calculate the quantization deviation of the joints and sensors, where the quantization deviation of the joints is the user's joint angle. Compared with standard joint angle The deviation value, the quantization deviation of the sensor is the user sensor data. Data range compared to standard sensors The deviation.
[0026] The feedback generation module transforms the evaluation results into user-perceptible video and voice feedback. The feedback generation module includes a video feedback submodule and a voice feedback submodule. The video feedback submodule synchronously overlays the user's standard motion skeleton diagram with the standard motion skeleton diagram, marking standard joints in green and deviating joints in red, and displays the deviation angle value in real time, generating a comparison video file that is projected onto the display screen. The voice feedback submodule generates natural language guidance text based on the quantified deviation value, and converts the natural language guidance text into voice prompts that are played in real time by the speaker.
[0027] The interaction and management module enables users to interact with the control system, while also managing the standard action library and user history data.
[0028] When using the intelligent training device for basketball dribbling rhythm quantification detection and motion correction provided by this invention, the control system's multi-source data acquisition module collects instructional videos, professional competition videos, camera image data, and standard motion data collected by sensors. After processing by the data preprocessing module, the multi-modal feature extraction and evaluation module constructs a CNN-LSTM-Attention multi-modal fusion model to extract features and generate a standard motion library. During user use, the multi-source data acquisition module collects motion data, which is then processed by the data preprocessing module. The multi-modal feature extraction and evaluation module extracts features and calculates the cosine similarity between the user's motion features and the standard feature centers of the standard motion library to evaluate whether the motion is standard. Feedback is then generated based on the calculated quantification deviation of the joints and sensors and played back via a display screen and speaker for real-time guidance and correction.
[0029] Therefore, the present invention employs the above-mentioned intelligent training device for quantitative detection and action correction of basketball dribbling rhythm, which sets up an IMU inertial measurement unit and a surface electromyography sensor to achieve accurate quantitative detection of basketball dribbling action, sets up a multimodal feature extraction and evaluation module to improve the accuracy and reliability of action evaluation, sets up a feedback generation module to achieve real-time and intuitive action correction and feedback, and sets up an interaction and management module to support personalized training and long-term progress tracking.
[0030] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and not to limit them. Although the present invention has been described in detail with reference to preferred embodiments, those skilled in the art should understand that modifications or equivalent substitutions can still be made to the technical solutions of the present invention, and these modifications or equivalent substitutions cannot cause the modified technical solutions to deviate from the spirit and scope of the technical solutions of the present invention.
Claims
1. A smart training device for quantitative detection and motion correction of basketball dribbling rhythm, characterized in that: The application relates to a basketball training system, which comprises a wearable device, a training device and a basketball, wherein the wearable device is internally provided with an IMU inertial measurement unit and a surface electromyography sensor, the training device comprises a shell and a control system, the control system is installed in the shell, and the control system comprises a multi-source data acquisition module, a data preprocessing module, a multi-modal feature extraction and evaluation module, a feedback generation module and an interaction and management module. The multi-source data acquisition module comprises a video acquisition unit, a sensor acquisition unit and a data transmission unit; the video acquisition unit collects user action images and video data through a camera and collects frame data of teaching and competition videos; the sensor acquisition unit reads acceleration, angular velocity and electromyography signals of the wearable device; The data preprocessing module converts original data into identifiable standardized time sequence features and eliminates noise and heterogeneous data differences; The multi-modal feature extraction and evaluation module completes feature extraction and action evaluation through a CNN-LSTM-Attention multi-modal fusion model; The feedback generation module converts the evaluation results into video and voice feedbacks that can be perceived by users; The interaction and management module realizes interactive control of the control system by users and simultaneously manages a standard action library and user historical data.
2. The basketball dribbling rhythm quantitative detection and motion correction intelligent training device according to claim 1, characterized in that: A display screen is mounted on the surface of the shell, control buttons are mounted on the shell at the bottom of the display screen, and loudspeakers are mounted on the shell side plates on the two sides of the display screen.
3. The basketball dribbling rhythm quantitative detection and motion correction intelligent training device according to claim 2, characterized in that: The data preprocessing module extracts human skeleton key points to generate joint coordinate time sequences based on an OpenPose algorithm, and the joint coordinate time sequences are as follows: ; wherein is the frame number, is the joint number, is the first frame is the first frame is the horizontal coordinate of the first frame is the vertical coordinate of the first frame is the vertical coordinate of the first frame is the vertical coordinate of the first ; The data preprocessing module performs Kalman filtering on acceleration and angular velocity data to eliminate motion noise, performs frequency domain analysis on electromyography signals to extract force features, aligns the time axes of the electromyography signals, acceleration and angular velocity data with the video frame time axis, and generates synchronous sensor time sequences, as follows: ; wherein, is the frame when the sensor collects axis acceleration components, is the frame when the sensor collects axis acceleration components, is the frame when the sensor collects axis acceleration components, is the frame when the sensor collects axis angular velocity components; The joint coordinate time sequences and the sensor time sequences are normalized to eliminate the influence of individual differences such as user height and weight, and the normalization calculation formula is as follows: ; wherein, is the original variable value to be normalized, is the minimum value of the original variable value in the corresponding data set, is the maximum value of the original variable value in the corresponding data set.
4. The basketball dribbling rhythm quantitative detection and motion correction intelligent training device according to claim 3, characterized in that: The multi-modal feature extraction and evaluation module inputs the pre-processed joint coordinate time sequence and sensor time sequence, extracts joint space topology features through CNN , extracts action time sequence dynamic features through LSTM, including joint coordinate time sequence time dynamic features and sensor time sequence time dynamic features , attention mechanism assigns higher weights to key action stages, and finally generates a multi-modal fusion feature vector , is the fusion feature dimension The attention mechanism passes through attention weights Fusion multi-modal features, strengthen the feature weight of key action stage, the calculation formula is shown as follows: ; wherein the attention weights satisfy the constraint ; The multi-modal feature data of the labeled standard action are clustered through a contrast loss, and the feature data of different actions are far away from each other; the calculation formula of the contrast loss is as follows: ; wherein, is an action category identifier, when, and are same-class actions, when, and are different-class actions, and are feature vectors corresponding to the two actions, is and is a square of the Euclidean distance, is an inter-class distance threshold value; The standard action feature vector extracted by the multi-modal fusion model is stored and clustered to form a structured standard action library, and a user action feature is calculated Cosine similarity with the standard feature center Determine whether the action is standard, as shown below: ; If , the standard action is determined, otherwise the non-standard action is determined, is a preset threshold value; The standard feature center is generated by clustering a plurality of standard features of the same action through a K-means clustering algorithm, and the specific process is as follows: ; wherein, a plurality of sets of standard features for the same action, is the number of action sub-classes; For non-standard motions, calculate the quantization bias of the joints and sensors, where the quantization bias of a joint is the difference between the user joint angle and the standard joint angle , and the quantization bias of a sensor is the difference between the user sensor data and the standard sensor data range .
5. The basketball dribbling rhythm quantitative detection and action correction intelligent training device according to claim 4, characterized in that: The feedback generation module comprises a video feedback submodule and a voice feedback submodule; The video feedback submodule synchronously superimposes and displays the user standard action skeleton graph and the standard action skeleton graph, labels standard joints with green and deviation joints with red, and displays deviation angle values in real time to generate a comparison video file and display the comparison video file on the display screen; The voice feedback submodule generates natural language guidance texts based on quantitative deviation values, converts the natural language guidance texts into voice reminders and plays the voice reminders in real time through the loudspeakers.