Real-time AI-driven joint data generation method and system for ball sports

By acquiring and processing the main skeleton and motion joint data of ball players and combining it with the self-attention neural network model, high-precision motion capture of fine joints such as hands in a wide range of scenarios is achieved, solving the problems of high hardware cost and poor adaptability in existing technologies and improving the real-time and accuracy of motion analysis.

CN120496735BActive Publication Date: 2025-09-16SHENZHEN LIMBOWORKS TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510991000.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-07-18
Publication Date
2025-09-16
Estimated Expiration
2045-07-18

AI Technical Summary

Technical Problem

Existing ball sports motion capture systems have difficulty balancing coverage and detail accuracy in large-scale scenarios. In particular, motion capture of fine joints such as hands and wrists is susceptible to occlusion and noise interference, resulting in the loss of key motion features. Existing methods also increase hardware costs and reduce system flexibility, making it difficult to adapt to the needs of different sports scenarios.

Method used

By obtaining the main skeleton joint point data and motion joint point data, motion noise elimination and trajectory smoothing are performed, combined with the self-attention neural network model for training, joint point data is predicted, and dynamically fused with the main skeleton data to generate character-driven event data.

Benefits of technology

It effectively reduces data errors, improves the ability to restore fine joint movements, reduces hardware costs, improves the real-time and adaptability of motion analysis, and can quickly respond to complex motion changes.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120496735B_ABST
    Figure CN120496735B_ABST
Patent Text Reader

Abstract

The present invention relates to a method and system for generating real-time AI-driven joint data for ball sports. The method comprises obtaining main skeleton joint data and motion joint data, performing motion noise elimination and trajectory smoothing processing, and obtaining whole-body motion training data; performing model training on a preset self-attention neural network model based on the whole-body motion training data to obtain a ball sports human joint data prediction model; obtaining the main skeleton data of a character in a real-time game scene, and inputting the data into the ball sports human joint data prediction model to predict joint data and output joint motion parameters; fusing the main skeleton data of the character with the joint motion parameters and outputting character-driven event data. The present invention can predict the joint motion data of athletes in real time.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of intelligent recognition, and in particular to a method and system for generating real-time AI-driven joint data for ball sports. Background Art

[0002] High-precision motion capture and real-time analysis of ball sports are crucial for training and competition evaluation. Conventional motion capture systems struggle to balance coverage and detail accuracy in large-scale scenarios (such as standard competition venues). They are particularly susceptible to occlusion and noise interference when capturing fine joints such as the hands and wrists, resulting in the loss of key motion features. To improve accuracy, existing methods typically rely on adding high-frame-rate cameras or reducing the capture range, but this significantly increases hardware costs and reduces system flexibility, making it difficult to adapt to the needs of different sports scenarios. In addition, the lack of targeted optimization for motion noise leads to distortion of key frame data; and the difficulty in adaptively identifying the individual motion features of athletes makes it easy to make misjudgments in real-time event recognition. These problems limit the accuracy and real-time nature of motion data analysis, and there is an urgent need for a high-performance, adaptive, and low-cost AI-driven solution to achieve more accurate joint motion prediction and event analysis. Summary of the Invention

[0003] The main purpose of this invention is to provide a real-time AI-driven joint data generation method and system for ball sports, which can quickly respond to complex movement changes in the game by predicting the athlete's joint motion data in real time and dynamically fusing it with the main body skeleton data.

[0004] To achieve the above objectives, the present invention provides a method for generating real-time AI-driven joint data for ball sports, comprising:

[0005] Obtain the main skeleton joint point data and motion joint point data, and perform motion noise elimination and trajectory smoothing to obtain full-body motion training data;

[0006] Performing model training on a preset self-attention neural network model based on the whole-body motion training data to obtain a human joint data prediction model for ball sports;

[0007] Obtaining the main skeleton data of the character in the real-time game scene, and inputting it into the ball sports human joint data prediction model to perform joint point data prediction, and output joint point motion parameters;

[0008] The character's main skeleton data and the joint action parameters are fused and processed to output character drive event data.

[0009] Furthermore, the acquisition of main skeleton joint point data and motion joint point data, and the elimination of motion noise and trajectory smoothing to obtain whole-body motion training data include:

[0010] Classify and process the three-dimensional coordinate data of the whole body joints collected by the optical motion capture device to obtain the main skeleton joint data and the motion joint data;

[0011] Performing motion noise elimination processing on the main skeleton joint point data and the motion joint point data respectively to obtain main skeleton denoised data and motion joint point denoised data;

[0012] Performing layered trajectory optimization on the main skeleton denoised data and the motion joint denoised data respectively to obtain main skeleton trajectory data and motion joint trajectory data;

[0013] Performing time alignment and spatial normalization processing on the main skeleton trajectory data and the motion joint trajectory data to obtain a structured motion training data set;

[0014] The structured motion training data set is subjected to action semantic segmentation and annotation to obtain the whole-body motion training data.

[0015] Furthermore, the main skeleton denoised data and the motion joint denoised data are respectively subjected to hierarchical trajectory optimization to obtain main skeleton trajectory data and motion joint trajectory data, including:

[0016] Performing rigid body dynamics constraint optimization on the main skeleton denoising data to obtain a first smooth trajectory;

[0017] performing activation constraint optimization and joint range of motion verification on the first smooth trajectory to obtain a second smooth trajectory;

[0018] Performing kinematic chain consistency optimization on the second smooth trajectory to obtain the main skeleton trajectory data;

[0019] performing micro-tremor joint trajectory optimization on the denoised data of the motion joint points to obtain a preliminary end joint point trajectory;

[0020] The preliminary end joint point trajectory is subjected to kinematic coupling optimization processing to obtain the motion joint point trajectory data.

[0021] Furthermore, the method of performing model training on a preset self-attention neural network model based on the whole-body motion training data to obtain a ball sports human joint data prediction model includes:

[0022] Performing joint angle conversion analysis on the whole-body motion training data to obtain joint rotation matrix data;

[0023] Performing intention feature recognition on the whole-body motion training data using the self-attention neural network model to obtain a first intention perception vector;

[0024] Performing motion feature coupling on the joint rotation matrix data through the self-attention neural network model to obtain a first motion coupling feature matrix;

[0025] Performing feature fusion on the first intention perception vector and the first motion coupling feature matrix to obtain a first candidate trajectory set;

[0026] Performing time series prediction on the first candidate trajectory set using the self-attention neural network model to obtain first predicted motion joint point data;

[0027] Performing a multi-order trajectory error calculation on the whole-body motion training data and the first predicted motion joint point data according to a preset loss function to obtain training error data;

[0028] The self-attention neural network model is iteratively optimized for model parameters based on the training error data, the first candidate trajectory set, and the first predicted motion joint point data to obtain the ball sports human joint data prediction model.

[0029] Furthermore, the method of obtaining the main skeleton data of the character in the real-time game scene and inputting it into the ball sports human joint data prediction model to perform joint point data prediction and output joint point motion parameters includes:

[0030] Performing a skeletal topology analysis on the character's main skeletal data to obtain an initial joint point topology map;

[0031] Inputting the initial joint point topology map into the motion intention perception layer, performing dynamic correlation analysis between joint points based on the global attention mechanism of the motion intention perception layer, and obtaining a second intention perception vector;

[0032] Inputting the second intention perception vector into the motion rule coupling layer, performing constraint matching on the second intention perception vector based on a preset motion rule knowledge graph, and obtaining a second motion coupling feature matrix;

[0033] Inputting the second motion coupling feature matrix into the motion trajectory generation layer to perform multi-scale motion trajectory prediction to obtain a second candidate trajectory set;

[0034] Inputting the second candidate trajectory set into the motion performance evolution layer for biomechanical rationality verification, and outputting optimized trajectory data;

[0035] Target frame extraction and cubic spline interpolation are performed on the optimized trajectory data to obtain the joint point motion parameters.

[0036] Furthermore, the second intention perception vector is input into the motion rule coupling layer, and constraint matching is performed on the second intention perception vector based on a preset motion rule knowledge graph to obtain a second motion coupling feature matrix, including:

[0037] Performing joint torque force analysis on the second intention perception vector through the compliance sublayer of the motion rule coupling layer to obtain joint force distribution;

[0038] Compliance verification is performed on the joint force distribution according to preset force constraint conditions to obtain a compliance motion feature;

[0039] Performing similarity matching between the compliant motion features and the motion rule knowledge graph through the motion rule matching sublayer to obtain motion rule matching features;

[0040] Performing mutation point interpolation correction on the motion rule matching feature to obtain a rule enhancement feature;

[0041] The compliant motion features and the rule-enhanced features are tensor-concatenated through a feature fusion sublayer to obtain the second motion coupling feature matrix.

[0042] Furthermore, the step of inputting the second motion coupling feature matrix into a motion trajectory generation layer to perform multi-scale motion trajectory prediction to obtain a second candidate trajectory set includes:

[0043] Performing spatiotemporal feature constraints on the second motion coupling feature matrix through the deterministic trajectory sublayer of the motion trajectory generation layer to obtain spatiotemporal coupling constraint features;

[0044] Based on the spatiotemporal coupling constraint characteristics, the spatial feasible range and temporal continuity of the joint point motion are calculated, and basic trajectory identification is performed to obtain a basic motion trajectory set;

[0045] Performing probability distribution fitting on the basic motion trajectory set through a probabilistic deduction sublayer to obtain candidate motion trajectory segments;

[0046] Redundant trajectory filtering calculation is performed on the basic motion trajectory set and the candidate motion trajectory segments to obtain the second candidate trajectory set.

[0047] Furthermore, the fusion processing of the character main skeleton data and the joint action parameters to output character driving event data includes:

[0048] Performing game engine time synchronization calibration on the character's main skeleton data to obtain synchronized skeleton data;

[0049] Performing character space coordinate system conversion on the synchronized skeleton data to obtain character unified coordinate skeleton data;

[0050] Performing kinematic constraint verification on the joint motion parameters to obtain compliant joint data;

[0051] Performing character dynamic fusion on the unified coordinate skeleton data of the character and the compliant joint point data to obtain character skeleton driving parameters;

[0052] Performing a full-body event analysis on the character's skeletal drive parameters to obtain the character's drive event data.

[0053] The present invention also provides a method system for generating real-time AI-driven joint data for sports, which is applied to any of the above-mentioned methods for generating real-time AI-driven joint data for ball sports, comprising:

[0054] An acquisition module is used to obtain main skeleton joint point data and motion joint point data, and perform motion noise elimination and trajectory smoothing processing to obtain whole-body motion training data;

[0055] An analysis module, the analysis module being used to perform model training on a preset self-attention neural network model based on the whole-body motion training data to obtain a ball sports human joint data prediction model;

[0056] An association module is used to obtain the main skeleton data of the character in the real-time game scene, and input it into the ball sports human joint data prediction model to perform joint point data prediction and output joint point motion parameters;

[0057] The processing module is used to fuse the character's main skeleton data with the joint point action parameters and output character driving event data.

[0058] The present invention provides a method and system for generating real-time AI-driven joint data for ball sports, which has the following beneficial effects:

[0059] By fusing main skeleton joint data with motion joint data, combined with motion noise elimination and trajectory smoothing, the system effectively reduces data errors in large-scale capture systems, improves the motion reproduction capabilities of fine joints such as the hands, and avoids loss of detail due to occlusion or noise. Reduced equipment dependency and costs: The prediction model based on the self-attention neural network can learn motion patterns from limited high-precision data, reducing dependence on high-frame-rate cameras or dense capture equipment, thereby reducing hardware costs while maintaining high motion analysis accuracy. By predicting the athlete's joint motion data in real time and dynamically fusing it with the main skeleton data, it can quickly respond to complex motion changes in competition, improve the real-time nature of event analysis, and adapt to the needs of different sports scenarios. BRIEF DESCRIPTION OF THE DRAWINGS

[0060] Figure 1 This is a flow chart of a method for generating real-time AI-driven joint data for ball sports provided by the present invention;

[0061] Figure 2 This is a system structure diagram of a real-time AI-driven joint data generation method for ball sports provided by the present invention.

[0062] The purpose, features and advantages of the present invention will be further described with reference to the accompanying drawings and in conjunction with the embodiments. DETAILED DESCRIPTION

[0063] In order to make the purpose, technical solutions and advantages of the present invention more clearly understood, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not intended to limit the present invention.

[0064] The present invention will be further described below with reference to the accompanying drawings and specific embodiments.

[0065] Reference Figure 1 As shown, the present invention provides a method for generating real-time AI-driven joint data for ball sports, comprising:

[0066] Step S101: obtaining main skeleton joint data and motion joint data, and performing motion noise elimination and trajectory smoothing processing to obtain whole-body motion training data;

[0067] Step S201: training a preset self-attention neural network model based on whole-body motion training data to obtain a ball sports human joint data prediction model;

[0068] Step S301: Obtaining the main skeleton data of the character in the real-time game scene, and inputting it into the ball sports human joint data prediction model to perform joint point data prediction and output joint point motion parameters;

[0069] Step S401: Fusing the character's main skeleton data with the joint action parameters to output character drive event data.

[0070] Based on the above steps, the detailed process is as follows:

[0071] Step S101: In real-time ball sports event analysis, full-body motion data of the athlete must first be collected using a high-precision motion capture system (such as an optical motion capture system, including both marker-based and markerless solutions). This data includes both main skeletal joints (such as shoulders, elbows, hips, and knees) and detailed joints (such as all finger joints and facial features including the left and right eyes, nose, ears, mouth, eyebrows, and facial contours). The collected data is typically stored as 3D spatial position data or joint rotation data (such as quaternions or Euler angles) with the root node as the origin.

[0072] Because high-precision motion capture systems may be subject to environmental interference (such as marker occlusion or lighting changes), the raw data may contain noise or jitter. Therefore, data cleaning is necessary, including outlier removal (such as the density-based DBSCAN algorithm) and temporal smoothing (such as Kalman filtering or Butterworth low-pass filtering). For detailed joints (such as finger micro-movements or facial expressions), due to their small motion amplitude, more sophisticated filtering methods (such as Savitzky-Golay filtering) are required to preserve high-frequency details. Subsequently, all joints are unified into the same coordinate system through spatiotemporal alignment, and interpolation methods (such as cubic spline interpolation) are used to supplement missing frames. Ultimately, a complete full-body motion training dataset is formed, containing precise motion information of the main skeleton and detailed joints.

[0073] Step S201: Based on the preprocessed full-body motion data, a self-attention neural network model based on a Transformer architecture is trained. Its input is the main skeleton joint data, and its output is the motion data of detailed joints (such as fingers and face). The core task of this model is to learn the kinematic relationships between the main skeleton and detailed joints, such as how wrist rotation affects finger posture or how head movement drives facial expression. The model architecture adopts a temporal-spatial separation Transformer design. The spatial attention module is responsible for modeling the relationship between the main joints and detailed joints within the same frame (such as the influence of elbow angle on finger bending), while the temporal attention module learns the temporal evolution of the action (such as the continuous process from gripping to extending the fingers during a basketball shot). During training, mean squared error (MSE) and quaternion distance are used as loss functions, while optimizing the position and rotation accuracy of detailed joints. In addition, data augmentation (such as adding Gaussian noise or temporal perturbations) is used to improve the model's generalization ability. After training, the model can predict high-fidelity detailed joint movements based solely on the subject's skeletal data (such as the sparse joint points collected by ordinary motion capture systems), thereby reducing dependence on high-cost motion capture equipment.

[0074] Step S301: During an actual match or training session, a standard motion capture device (such as a monocular camera or inertial sensor) is used to collect the athlete's main skeletal data in real time. Because these devices typically cannot capture detailed joints (such as fingers or faces), the collected main skeletal data is fed into a trained ball sports human joint data prediction model to predict the motion of the missing detailed joints in real time.

[0075] The specific process involves normalizing and filtering the input skeletal data (e.g., moving average filtering) to ensure it aligns with the training data distribution. This processed data is then fed into the model, which outputs predicted detailed joint data (e.g., finger joint angles or facial expression point displacements). Due to the parallel computing capabilities of the Transformer, the predictions are post-processed and optimized, for example, using physical constraints (e.g., preventing excessive finger bending) and temporal smoothing (e.g., Kalman filtering) to ensure natural and coherent movements. If the system detects partial occlusion of joints (e.g., a hand obscured by the body), a short-term memory module is activated to maintain temporal consistency.

[0076] Step S401: The main skeletal data captured by a conventional motion capture system is fused with the detailed joint data predicted by the model to generate complete full-body motion data for further motion event analysis. The fusion process employs a dynamic weighting strategy, assigning higher confidence to directly captured main joints (such as the knee) while adjusting the weights of model-predicted detailed joints (such as the fingertips) based on the stability of historical frames. Event recognition is performed based on this fused data, combined with a rule engine or machine learning model. For example, the timing of a shot can be determined by the timing of wrist rotation and finger extension, or an athlete's emotional state can be analyzed through eyebrow and mouth movements. The output event data is stored in a structured format (e.g., "Event Type: Dunk, Trigger Joint: Right Wrist + Index Finger Tip, Confidence: 95%)") and supports real-time visualization (e.g., overlaying close-ups of finger movements or facial expression analysis during live broadcasts).

[0077] The present invention provides a real-time AI-driven joint data generation method for ball sports. By fusing the main skeleton joint point data and the motion joint point data, and combining motion noise elimination and trajectory smoothing processing, it effectively reduces the data error of ordinary capture systems in large scenes, improves the motion restoration ability of fine joints such as the hands, and avoids the loss of details due to occlusion or noise. Reduce equipment dependence and cost: The prediction model based on the self-attention neural network can learn motion patterns from limited high-precision data, reduce dependence on high-frame rate cameras or intensive capture equipment, thereby reducing hardware costs, while maintaining a high level of motion analysis accuracy. By predicting the athlete's joint motion data in real time and dynamically fusing it with the main skeleton data, it can quickly respond to complex motion changes in the game, improve the real-time nature of event data generation, and adapt to the needs of different sports scenes.

[0078] In one embodiment, main skeleton joint point data and motion joint point data are obtained, and motion noise elimination and trajectory smoothing are performed to obtain whole-body motion training data, including:

[0079] During the data collection phase, a high-precision optical motion capture system (such as the Vicon MX series) is used, equipped with 12 high-speed infrared cameras, to synchronously capture athletes' full-body motion data at a frame rate of 240fps. The cameras are arranged in a circular array to ensure full coverage of the playing field, preventing the loss of markers due to athlete turns or limb occlusion. Athletes are required to wear customized reflective marker sets, with markers arranged to closely match human anatomy. Hemispherical markers are used for key skeletal joints (including the rotation angles and positional changes of core postural nodes such as the hip, knee, ankle, spinal segments, shoulder, elbow, and wrist). Motion joints (including the individual finger joints, such as the metacarpophalangeal, proximal interphalangeal, and distal interphalangeal joints of the thumb, as well as the corresponding joints of the other four fingers, recording their flexion, extension, abduction, and adduction movements). Detailed facial joints, such as eye opening and closing, eye movement, eyebrow raising and lowering, lip opening and closing, and subtle changes in facial muscles, are captured by placing markers on key facial locations. For hand movements, additional 3mm ultra-micro markers are placed on each phalanx to capture subtle movements such as gripping and finger flicking. The system undergoes dynamic calibration prior to data collection, using an L-shaped calibration rod to verify at multiple points within its range of motion to ensure spatial positioning error of less than 0.1mm. The acquisition environment must maintain a constant infrared light intensity (850nm band) and use a black curtain to block ambient light interference. The raw 3D coordinate data of the entire body joints is output as a 3D joint coordinate matrix with timestamps in the format [N×J×3], where N is the number of frames and J is the number of joints (the standard configuration is 53 joints).

[0080] For the main skeleton joint point data, a noise suppression strategy based on frequency analysis is first adopted: the motion spectrum of each joint point is calculated through fast Fourier transform (FFT), and the high-frequency noise components (usually >8Hz) generated by equipment vibration or muscle tremor are identified. A fourth-order Butterworth low-pass filter is applied for cutoff filtering, and the cutoff frequency is set to 5Hz to retain the true motion characteristics to obtain the main skeleton denoised data.

[0081] For motion joint data, an adaptive Kalman filter algorithm is used. This algorithm dynamically adjusts the process noise matrix Q by monitoring joint velocity in real time. When high-speed motion is detected (e.g., wrist velocity >6 m / s when shooting a basketball), the diagonal elements of the Q matrix are reduced to reduce smoothing intensity and prevent loss of motion details. At low speeds (e.g., when standing with a ball), the Q matrix elements are increased to 1e-3 to enhance denoising. The filtered data is subjected to residual analysis. Outliers with displacement residuals exceeding 3σ for more than five consecutive frames are linearly interpolated to obtain motion-denoised data.

[0082] The denoised main skeleton data is first fed into the physically constrained trajectory optimization module. This module establishes a human biomechanical model, treating the spine as seven rigid links (corresponding to the cervical, thoracic, and lumbar vertebrae). Using inverse dynamics, it calculates the rotational constraints between each segment (for example, limiting lumbar lateral flexion to ±30°). If the optimizer detects a constraint violation in the raw data (for example, detecting a 45° spinal posterior angle), it automatically invokes a B-spline interpolation algorithm to regenerate a trajectory that conforms to physiological limits, using valid data from 10 consecutive frames as control points.

[0083] The motion denoised data is then fed into the kinematic optimization pipeline. For shooting, the least squares method is used to fit the deviation of the wrist trajectory from the ideal parabola (the initial velocity v0 is calculated from 20 frames of data at the moment of release, and the gravitational acceleration g is set to 9.81 m / s²). Trajectory replanning is triggered when the root mean square error (RMSE) exceeds 0.1 m. This process projects the trajectory onto the theoretical parabola while preserving the original motion trend. All optimization processes maintain the synchronization of the original data's timestamps to avoid timing distortions caused by processing.

[0084] A dynamic time warping (DTW) algorithm is used to address subtle timing offsets during multi-device acquisition. Using the vertical trajectory of the main skeletal hip joint as a reference signal, nonlinear time warping is applied to the data from other joints, ensuring synchronization errors within ±2 frames (approximately 8.3ms) for key event frames (e.g., basketball release frames). Spatial normalization is implemented in three steps: first, a local coordinate system is established with the hip midpoint (the midpoint of the line connecting the left and right hip joints) as the origin, and all joint coordinates are converted to relative positions within this coordinate system. Second, the data is scaled based on the athlete's actual height (e.g., 1.98 meters), normalizing the vertical distance from the hip to the top of the head to 1.2 units. Finally, rotational alignment is performed, using the shoulder line at the athlete's initial stance as a reference and rotating it around the vertical axis to align it with the x-axis of the coordinate system. The processed data is stored as a hierarchical HDF5 file containing metadata fields such as original coordinates, normalized coordinates, timestamp, and athlete ID.

[0085] Based on annotated play board information (e.g., "Play No. A3 - Cross Screen Shoot"), a sliding window approach is used to semantically segment the continuous motion stream: the window length is set to 1.5 seconds (corresponding to 360 frames of data), with a sliding step of 0.45 seconds (108 frames), ensuring that each play phase is covered by at least one complete window. The data within each window is independently annotated, including the action name, type, start and end times, basic information about the performer (e.g., age, gender, height, weight), and contextual information about the action (e.g., emotional state, purpose, etc.). For example, a "waving hello" action is annotated as a daily social action, starting at 0 seconds, ending at 2 seconds, performed by a 25-year-old male, and with a friendly mood. The action includes atomic action types (e.g., "right-hand jump shot," "crossover step drive"), action phases (e.g., "charge period 0.2-0.5 seconds," "release instant 0.82 seconds"), and performance scores (e.g., a release angle deviation within ±3° is marked as grade A). Annotation results are determined by majority voting, with arbitration in case of disagreement. The resulting full-body motion training data includes multimodal labels: structured motion parameters (such as shot speed and jump height), tactical context (such as "fast break left wing"), and video segment indexes (linked to multi-angle synchronized video). Data is stored in the Apache Parquet columnar format, supporting fast retrieval by multiple dimensions, including action type, time period, and athlete ID.

[0086] This embodiment effectively eliminates high-frequency noise and motion jitter by combining frequency domain filtering with Kalman filtering, ensuring that the trajectory of the joint points is smooth and conforms to biomechanical constraints, and avoiding the loss of key movement details due to noise interference in traditional methods. Through dynamic time warping and spatial normalization processing, the timing offset problem during multi-device acquisition is solved, and the motion data scales of different athletes are unified to facilitate subsequent analysis and model training. The annotation method based on semantic segmentation decomposes the continuous motion stream into tactical action units, and combines the annotation verification of professional coaches to construct a high-quality structured data set, providing rich training samples for the AI ​​model. The final generated full-body motion training data not only contains precise kinematic parameters, but also associates tactical context with video indexes, which can fully support real-time tactical analysis, movement correction and performance evaluation, and meet the high standards of professional sports training.

[0087] In one embodiment, hierarchical trajectory optimization is performed on the main skeleton denoised data and the motion denoised data to obtain main skeleton trajectory data and motion joint trajectory data, including:

[0088] De-noised main skeleton data refers to the 3D coordinate sequence of the athlete's main skeletal points after noise interference has been removed through pre-processing. Rigid body dynamics constraints are applied to this denoised main skeleton data. Rigid body dynamics constraints require that bone lengths remain constant during motion. The optimization process minimizes the change in bone length between adjacent frames, eliminating the expansion and contraction of bones caused by sensor errors, and outputs a first smooth trajectory that conforms to rigid body characteristics.

[0089] The first smoothed trajectory enters the activation constraint optimization phase. Activation constraints define the expected activity intensity thresholds for each joint in a specific motion pattern. The system detects the motion amplitude of each joint and compares it with a pre-set motion pattern database, correcting any abnormal data points that fall outside the physiological range of motion. The joint range of motion verification module further checks whether the optimized trajectory complies with the human joint rotation angle constraints, generating a second smoothed trajectory that satisfies both the activation constraints and the physiological constraints.

[0090] Kinematic chain consistency optimization is performed on the second smooth trajectory. Kinematic chain consistency refers to the temporal and spatial coherence of adjacent joint motions. The system establishes a motion transfer model for all joints, adjusting the trajectories of each skeletal point to ensure phase matching between upper and lower limb motion and coordinated trunk and limb motion. Ultimately, it outputs master skeletal trajectory data that conforms to biomechanical principles.

[0091] Motion denoising data consists of a preprocessed sequence of joint coordinates across the athlete's entire body. Micro-tremor joint trajectory optimization specifically addresses subtle vibration noise in distal joints, such as fingers and feet. Micro-tremor refers to high-frequency, low-amplitude, unintentional joint tremors. The optimization algorithm identifies and retains trajectory features related to movement intent, filters out random fluctuations caused by electromyographic noise, and generates a stable preliminary distal joint trajectory.

[0092] Kinematic coupling optimization is applied to the initial end-joint trajectory. Kinematic coupling describes the mechanical relationships involved in the coordinated motion of multiple joints. The system establishes the kinematic equations for the end-effector and proximal joints. Using inverse kinematics, it adjusts the positions of each joint to ensure the mechanical rationality of kinematic chains such as wrist-palm-finger or ankle-foot-toe. Ultimately, it generates kinematic joint trajectory data that conforms to anatomical constraints.

[0093] The main skeleton trajectory data represents the optimized motion paths of the athlete's core skeletal points. Each skeletal point trajectory satisfies rigid body dynamics constraints, joint range of motion limits, and kinematic chain consistency requirements. The motion joint point trajectory data demonstrates the refined motion characteristics of the end joints, eliminating micro-vibration noise while maintaining kinematic coupling. Together, these two types of trajectory data form a complete representation of the athlete's motion, providing precise input for subsequent event data generation.

[0094] This embodiment optimizes the main skeleton denoising data through rigid body dynamics constraints to ensure that the bone length remains constant during movement, effectively eliminating the bone expansion and contraction phenomenon caused by sensor noise, and improving the physical rationality of the motion trajectory. Activation constraint optimization and joint range of motion verification further correct abnormal motion data, making the trajectory consistent with the human body's physiological range of motion and enhancing the biomechanical accuracy of motion analysis. Kinematic chain consistency optimization coordinates the joint movements of the whole body, matching the phases of the upper limbs, lower limbs, and torso movements, and improving the coherence of the overall motion analysis.

[0095] In one embodiment, a preset self-attention neural network model is trained based on whole-body motion training data to obtain a human joint data prediction model for ball sports, including:

[0096] Full-body motion training data is derived from wearable inertial sensors or optical motion capture systems, containing the three-dimensional position and rotation information of each joint. Joint angle transformation analysis uses quaternions or Euler angles to calculate joint rotation matrices, ensuring that rotations conform to right-handed coordinate systems. Joint rotation matrix data is stored in a time series format, with each time step containing rotation information for all joints. This provides structured input for the subsequent self-attention neural network model.

[0097] The self-attention neural network model receives full-body motion training data and maps joint position and velocity information into a high-dimensional space through the motion intention perception layer, generating a first intention perception vector. The preset embedding dimension is proportional to the number of joints, ensuring that the feature vector of each joint fully represents its motion state. The self-attention mechanism of the self-attention neural network model calculates dependencies between joints, and the first intention perception vector incorporates global motion context information.

[0098] The motion rule coupling layer constructs a first motion coupling feature matrix based on the joint rotation matrix data. The pre-set human skeletal topology defines the graph's adjacency matrix, with joints as nodes and bones as edges. GAT (Graph Attention Networks) uses attention weights to calculate the dynamic coupling relationships between joints. The first motion coupling feature matrix contains node features and an attention-weighted adjacency matrix. This step enhances the self-attention neural network model's ability to model local joint motion.

[0099] The first intention perception vector and the first motion coupling feature matrix are fused through the motion trajectory generation layer. The encoder of the self-attention neural network model receives the two features, calculates the joint attention weights, and generates the first candidate trajectory set. The preset fusion rules require that the spatial topology and temporal dynamics be preserved. The dimensions of the first candidate trajectory set are consistent with those of the motion trajectory generation layer of the self-attention neural network model.

[0100] The first candidate trajectory set is input into the motion representation evolution layer of the self-attention neural network model, where it performs autoregressive time series prediction. The motion representation evolution layer uses a causal mask to ensure that predictions rely solely on historical information and outputs the first predicted motion joint data. The preset prediction step size covers the typical motion cycle, and the first predicted motion joint data contains joint position and rotation information for several future frames.

[0101] The loss function uses a weighted combination of mean squared error (MSE) and dynamic time warping (DTW) to calculate the difference between full-body motion training data and the first predicted motion joint data. A pre-defined error weighting strategy emphasizes prediction accuracy for key joints (such as the wrist and elbow). Training error data includes both frame-by-frame error and cumulative error across motion stages.

[0102] The training error data, the first candidate trajectory set, and the first predicted motion joint data are combined, and the parameters of the self-attention neural network model are adjusted through backpropagation. The optimizer uses AdamW with a learning rate warmup strategy. A preset early stopping rule monitors the validation set error to prevent overfitting. This generates the final prediction model for human joint data in ball sports, ultimately enabling real-time inference capabilities to support real-time trajectory prediction and event data generation for athlete movements.

[0103] This embodiment is based on the collaborative design of the self-attention mechanism and the graph attention network, which simultaneously captures the global dependencies and local dynamic coupling characteristics between joints, significantly improving the modeling capabilities of complex motion sequences. A cross-modal attention fusion strategy is adopted to integrate spatial-temporal features while preserving the skeletal topology and temporal dynamics, so that the model can more accurately predict multi-joint coordinated motion. Through the MSE and DTW weighted loss functions and the key joint error optimization strategy, the frame-by-frame accuracy and motion coherence are balanced, effectively improving the prediction effect of specific ball sports actions (such as throwing and swinging). Combined with the AdamW optimizer and the learning rate warm-up training mechanism, while ensuring the efficiency of the model convergence, overfitting is avoided. The final generated model supports real-time reasoning and can be widely used in athlete motion analysis and tactical decision support.

[0104] In one embodiment, the training process of the ball sports human joint data prediction model further includes:

[0105] Data preparation phase: The collected high-precision motion capture data is preprocessed to convert the motion data of the 23 main body bones of the main body skeleton into a unified joint rotation matrix data representation. At the same time, the motion trajectories of the 30 detailed joints of the hand are extracted as supervision signals. Through time series alignment and normalization processing, a structured dataset containing the first intention perception vector is constructed. In view of the characteristics of ball sports, key action frames (such as the start / peak moments of pitching, swinging, etc.) are additionally annotated to form a first motion coupling feature matrix with motion semantic labels. The output of the model is the motion data of the target detailed joints, such as the rotation angles and position trajectories of the 5 fingers and 30 joints of the hand. The motion joint point data contains data on multiple detailed joints.

[0106] Model Construction and Initialization: A layered Transformer architecture serves as the core framework. The bottom layer processes the temporal features of joint rotation matrix data. The middle layer constructs a first motion coupling feature matrix using a graph attention mechanism to capture local joint linkage relationships. The top layer generates a first set of candidate trajectories for predicting detailed joint motion. The generator utilizes a temporal convolutional network with residual connections, while the discriminator uses a 3D convolutional network to analyze the spatiotemporal consistency of motion sequences. Model parameters are initialized using a normal distribution, and a pretrained ball sports base model is loaded for warm-start. Generative models such as generative adversarial networks (GANs) and variational autoencoders (VAEs) can also be used to generate detailed joint motion data. GANs employ adversarial training between the generator and discriminator, enabling the generator to produce realistic detailed joint data that closely resembles the real data distribution. Variational autoencoders, based on the encoder and decoder, introduce probability distribution constraints, enabling not only the generation of detailed joint data but also data compression and feature extraction.

[0107] Training Process: During training, the data is segmented into sliding windows of 128 frames, with each batch inputting 32 sets of first intention perception vectors and their corresponding first motion coupling feature matrices. During the forward propagation phase, the main network first extracts the spatiotemporal features of the joint rotation matrix data. It then fuses the first intention perception vectors with the first motion coupling feature matrix via a cross-modal attention layer, outputting the first set of candidate trajectories. Based on this matrix, the generator predicts hand joint motion data for the next 30 frames, while the discriminator performs adversarial discrimination between the generated sequence and the real sequence.

[0108] Loss Calculation and Optimization: The training error data consists of four key components: 1) Joint Angle MSE Loss calculates the frame-by-frame difference between the first predicted joint point data and the ground truth; 2) Dynamic Time Warping Loss optimizes the similarity of the overall motion curve; 3) Adversarial Loss improves the realism of the generated motion; and 4) Physical Plausibility Loss constrains the range of joint motion. The AdamW optimizer is used for parameter updates, with a set initial learning rate and a cosine annealing schedule. The FID scores of the generated motions are evaluated on the validation set every 2000 training steps.

[0109] Verification and Tuning Mechanism: The verification phase employs a two-track evaluation strategy: quantitatively, the MPJPE error of the first predicted motion joint data on the test set is calculated; qualitatively, professional animators subjectively score the generated hand movements. If the verification loss does not decrease for five consecutive times, a learning rate decay (coefficient 0.5) or local fine-tuning (focusing on optimizing the finger joint prediction module) is triggered. Heatmaps of the first candidate trajectory set are regularly visualized to analyze the network's attention weights for different motion phases.

[0110] During the deployment and testing phase, the resulting ball sports human joint data prediction model must pass three tests: 1) a generalization test across athletes to verify its adaptability to subjects of varying body types; 2) a long-sequence stability test to assess the cumulative error of continuously generating 5 minutes of motion data; and 3) a real-time test to ensure single-frame prediction within 10ms. The model's output interface also provides the first predicted motion joint data and corresponding confidence score for downstream application decision-making.

[0111] This implementation achieves high-fidelity generation of complex hand movements during motion through a hierarchical conversion of joint rotation matrix data into a first intention perception vector, combined with local constraints of the first motion coupling feature matrix and global modeling of the first candidate trajectory set. Multi-dimensional monitoring of training error data and a progressive optimization strategy for the human joint data prediction model for ball sports ensure that the generated movements conform to physical laws while preserving individual movement styles.

[0112] In one embodiment, the skeleton data of the character in the real-time game scene is obtained and input into the ball sports human joint data prediction model to predict the joint point data and output the joint point motion parameters, including:

[0113] Real-time motion data from athletes during competition is collected to obtain the skeletal data of the main character. This data includes joint coordinates, velocity, and acceleration information in three-dimensional space. The skeletal topology analysis module performs structured processing on the raw data, identifying the connections between human joints and constructing an initial joint topology graph. This graph fully preserves the inherent characteristics of human kinematics, with nodes representing specific joints and edges representing the physical connections between joints.

[0114] The initial joint topology map is input into the motion intention perception layer of the human joint data prediction model for ball sports for processing. This layer utilizes an improved global attention mechanism to calculate the dynamic correlation weights between joints, focusing on capturing the motion characteristics of key nodes such as the ball-handling hand and supporting foot. The motion intention perception layer analyzes the coordinated change patterns of different joints in spatiotemporal dimensions, such as the linkage between the wrist, elbow, and shoulder in a basketball shot, and ultimately outputs a second intention perception vector representing the motion intention. This high-dimensional vector not only contains the motion state of each joint but also encodes the athlete's current overall motion intention.

[0115] The second intention perception vector then enters the motion rule coupling layer for processing. The built-in motion rule knowledge graph in this layer stores the professional rules and common tactical patterns of various ball sports, including hard constraints such as traveling rules in basketball and offside rules in football, as well as typical characteristics of various tactical combinations. The motion rule coupling layer performs multiple rounds of matching and screening on the second intention perception vector, eliminating abnormal features that do not conform to the sports rules while strengthening the feature expression that conforms to tactical logic. The second motion coupling feature matrix output after this layer of processing not only retains the original motion intention information but also ensures that all prediction results conform to the professional rules of the sport.

[0116] The second motion coupling feature matrix is ​​fed into the motion trajectory generation layer for multi-scale analysis. This layer consists of three parallel prediction branches, each handling motion trajectory predictions for different time spans. The short-term prediction branch focuses on fine-scale movement changes within 0.5 seconds, the medium-term prediction branch processes continuous movement sequences within 1-2 seconds, and the long-term prediction branch infers macroscopic displacement trends over 3 seconds. The prediction results of these three branches are fused to generate a second set of candidate trajectories containing multiple possible motion paths. Each candidate trajectory is accompanied by a confidence score, reflecting the likelihood of the prediction result.

[0117] The second set of candidate trajectories enters the Performance Evolution layer for biomechanical plausibility verification. This layer integrates a biomechanical model of human motion and rigorously checks the physical feasibility of each candidate trajectory. Verification includes ensuring that the range of joint motion is within physiological limits, that the acceleration of the movement conforms to muscle force characteristics, and that the overall motion complies with the law of conservation of momentum. For trajectories that do not conform to biomechanical laws, the system automatically optimizes and adjusts them. The final output is optimized trajectory data that maintains the original predicted motion intent while ensuring that all movements conform to human kinematic principles.

[0118] The optimized trajectory data is then processed through target frame extraction and cubic spline interpolation. The target frame extraction module identifies key motion nodes during the movement, such as the release moment of a basketball shot or the moment of contact with the ball in a soccer shot. Between these key frames, the system uses a cubic spline interpolation algorithm to generate smooth and natural transitions, ultimately outputting highly accurate and fluid joint motion parameters. This data can be directly applied to real-time game analysis systems, providing professional support for various scenarios, including tactical decision-making, referee assistance, and spectator experience.

[0119] By collecting real-time motion data from athletes, this embodiment can comprehensively capture the three-dimensional motion information of human joints, significantly improving the accuracy and integrity of motion data collection. A skeletal topology analysis module is used to construct an initial joint topology map, effectively preserving the inherent characteristics of human kinematics and ensuring that the input data for subsequent prediction models has accurate physical relevance. The motion intention perception layer utilizes an improved global attention mechanism to dynamically analyze the coordinated motion patterns of key joints, enabling the model to accurately identify the athlete's real-time movement intentions and improve prediction accuracy. The motion rule coupling layer, combined with a professional sports rule knowledge graph, automatically filters out abnormal movements that do not conform to the rules, ensuring that predictions conform to actual game logic and avoiding erroneous predictions that violate the rules of the sport. The multi-scale prediction branch of the motion trajectory generation layer simultaneously processes short-term, medium-term, and long-term motion trends, ensuring that predictions encompass both fine-scale movement changes and macroscopic displacement trends, enhancing the comprehensiveness of the predictions. The motion performance evolution layer verifies and optimizes trajectory data using a biomechanical model to ensure that all predicted movements conform to the laws of human kinematics and avoid generating unreasonable motion trajectories.

[0120] In one embodiment, a ball sports human joint data prediction model is deployed in a motion capture system with normal precision, including:

[0121] A standard motion capture system is deployed to collect real-time motion data from the athlete's trunk, limbs, and other major skeletons. The sampling frequency is set to 100Hz to ensure the capture of detailed movement. Upon receiving the raw data, the system immediately analyzes the skeletal topology and constructs a hierarchical joint topology based on standard human anatomy. This includes 17 key joints, arranged according to an internationally recognized biomechanical notation system, with the pelvis joint serving as the root node of the entire topology.

[0122] The parsed initial joint topology is then fed into the motion intention perception layer of the human joint data prediction model for ball sports. A self-attention mechanism dynamically analyzes the motion correlations between joints. For example, in basketball shooting analysis, the coordinated motion relationships between the wrist, elbow, and shoulder joints are calculated. A 128-dimensional second intention perception vector is generated for this specific motion pattern, containing the motion phase labels and coordination scores for each joint.

[0123] The resulting second intention perception vector is then input into the motion rule coupling layer. This layer incorporates a professional ball sports knowledge graph containing physical motion constraints (such as joint range of motion angles) and tactical rule features. The system uses a graph matching algorithm to compare real-time motion features with the rules in the knowledge graph, automatically correcting abnormal data that does not conform to motion rules. For example, it can restrict the detected knee flexion angle to within a physiologically acceptable range. The final output is a 32×32-dimensional second motion coupling feature matrix.

[0124] The feature matrix then enters the motion trajectory generation layer for multi-scale prediction. This layer uses a parallel spatiotemporal convolutional network to analyze motion trajectories using short-term (5-frame) and long-term (20-frame) windows. For basketball dribbling, the system simultaneously predicts subtle finger control movements and the overall displacement trend of the body's center of gravity, generating multiple candidate motion trajectories, each with a confidence score based on historical data training.

[0125] The second set of candidate trajectories is fed into the Performance Evolution layer for biomechanical plausibility verification. This layer, integrated with the OpenSim biomechanics simulation engine, rigorously verifies the dynamics and kinematics of each candidate trajectory, including analysis of joint torques, energy consumption, and other specialized metrics. The system automatically eliminates trajectories that do not conform to the biomechanical characteristics of professional athletes and intelligently weights and fuses the remaining trajectories, ensuring that the final output of motion data is both physically consistent and sport-specific.

[0126] The optimized trajectory data is then time-series processed. Key moments in the action sequence are automatically identified using a motion key event detection algorithm. Cubic spline interpolation is then used to complete the data between these key frames, ensuring smoothness and continuity in the output data. The entire processing flow is completed within 50ms, enabling true real-time motion analysis. The resulting high-precision joint data can be used to directly drive virtual athlete models or for professional tactical analysis.

[0127] This embodiment can accurately capture the subtle changes in athletes' movements, and output them to a variety of application scenarios that require high-quality motion capture data, such as film and television animation production, electronic game development, virtual reality and augmented reality experience, etc. This makes the movement performance of virtual characters more natural and realistic. Combined with professional sports knowledge graphs and biomechanical simulation verification, the system automatically corrects abnormal data and optimizes motion trajectories to ensure that the output data conforms to physical laws and retains motion specificity, significantly improving the accuracy and reliability of motion capture data. In addition, by adopting parallel computing and intelligent weighted fusion strategies, data processing is completed within 50ms, meeting the needs of real-time applications such as film and television animation, game development and VR / AR, and enhancing the immersion of virtual characters and the naturalness of interaction. Through subjective visual evaluation and objective error index verification, this method has demonstrated excellent performance in both motion detail capture and biomechanical rationality, providing high-quality technical support for motion capture applications in sports training, digital entertainment and other fields.

[0128] In one embodiment, the second intention perception vector is input into the motion rule coupling layer, and constraint matching is performed on the second intention perception vector based on a preset motion rule knowledge graph to obtain a second motion coupling feature matrix, including:

[0129] The biomechanical compliance sublayer of the motion rule coupling layer receives the second intention perception vector and calculates joint torque parameters based on an inverse dynamics model. Joint torque calculation utilizes a human skeletal mass distribution model, with a predefined limb mass distribution of 12% for the thigh, 6% for the calf, 3% for the upper arm, and 2% for the forearm. Joint point dynamic parameters include torque components in each of the three axes, and finite element analysis is used to simulate the stress distribution in the articular cartilage. Joint force analysis is based on predefined joint torque safety thresholds: the sagittal plane torque threshold for the shoulder joint is set at 80% of the anatomical limit (45 Nm), and the coronal plane torque threshold for the knee joint is set at 35 Nm. Compliant motion features are filtered using a gating mechanism, retaining only joint motion patterns with torque parameters below the threshold. Finally, compliant motion features (joint motion encoding within the torque safety range) are obtained.

[0130] The motion rule matching sublayer loads a knowledge graph containing 300 standard motion templates. Each template defines joint angle ranges, kinematic chain phase differences, and energy consumption parameters. Similarity matching calculates the cosine similarity between the compliant motion features and the knowledge graph templates, with a preset matching acceptance threshold of 0.75. When the similarity between the feature vector and the "smash" motion template reaches 0.82, the rule enhancement mechanism is triggered, applying a ±5-degree correction constraint to the wrist rotation angle. The rule-enhanced features are adjusted using attention weighting. The weight of the compliant motion features with a matching score above the threshold is increased to 1.2 times, while the weight of the less matching features is reduced to 0.8 times. The resulting rule-enhanced features (knowledge graph-driven motion feature correction) are obtained.

[0131] The spatiotemporal constraint sublayer detects motion mutation points in the rule-enhanced features, defining an abnormal mutation as a rate of change of joint velocity exceeding 15% within three adjacent frames. Mutation point correction utilizes a time window sliding interpolation method, inserting a quintic polynomial trajectory between the start and end frames where the mutation is detected. This interpolation correction covers a time window of five frames before and after, ensuring smooth and continuous joint acceleration. The spatiotemporal smoothness feature is verified using a kinematic chain, controlling the phase difference of the elbow-wrist-metacarpal joint linkage within ±3 frames to eliminate non-physiological jitter in the motion trajectory. This results in a spatiotemporal smoothness feature (a motion trajectory encoding with continuous acceleration).

[0132] The feature fusion sublayer concatenates the compliant motion features, rule-enhanced features, and spatiotemporal smoothing features into three-dimensional tensors. This concatenation is performed along the feature dimensions. The compliant motion features retain their original 256-dimensional vectors, the rule-enhanced features are expanded with 32-dimensional attention weight parameters, and the spatiotemporal smoothing features are encoded with 24-dimensional timestamps. The final dimension of the second motion coupling feature matrix is ​​312, and layer-wise normalization eliminates feature scale differences. Each feature channel in the matrix is ​​associated with a specific motion rule constraint, encompassing three optimization metrics: biomechanical safety, movement standardization, and spatiotemporal continuity. This results in the second motion coupling feature matrix (a 312-dimensional multi-constrained motion feature matrix).

[0133] This embodiment uses the biomechanical compliance sublayer of the motion rule coupling layer to perform threshold constraints on joint torques, effectively screening ergonomically safe motion patterns and reducing the risk of sports injuries caused by joint overload. The similarity matching mechanism based on the knowledge graph ensures that the generated motion features are highly consistent with the standard motion template, improves the standardization of movements and avoids illegal motion trajectories. The polynomial interpolation correction of the spatiotemporal constraint sublayer eliminates motion mutation points, ensures the continuity of joint acceleration changes, and generates natural trajectories that conform to the laws of physiological movement. The multimodal feature fusion strategy integrates multi-dimensional optimization indicators of biomechanical safety, rule compliance, and spatiotemporal continuity to enhance the comprehensive decision-making ability of the human joint data prediction model for ball sports. The priority of different constraints is optimized through a dynamic weight adjustment mechanism to maintain the expressiveness of competitive movements while ensuring sports safety.

[0134] In one embodiment, the second motion coupling feature matrix is ​​input into the motion trajectory generation layer to perform multi-scale motion trajectory prediction to obtain a second candidate trajectory set, including:

[0135] After the second motion coupling feature matrix (a feature representation generated by the fusion of motion rules and sensor data) is input into the motion trajectory generation layer, the deterministic trajectory sublayer performs spatiotemporal feature decomposition. This decomposition uses a sliding window mechanism to partition the second motion coupling feature matrix into local spatiotemporal blocks based on time frames. For each block, spatial motion constraint features (describing the feasible motion range of joints in three-dimensional space) and temporal motion sequence features (describing the motion trends of joints over time) are independently extracted.

[0136] Spatial motion constraint features are calculated based on a biomechanical model with built-in rotation angle thresholds and limb length limits for each joint. The spatially feasible range of each joint is calculated using an inverse kinematics solver, which outputs the motion space boundary conditions (the maximum displacement range of each joint in space, represented by a 3D coordinate bounding box). Temporal motion sequence features are extracted using a time-series convolutional network. The network automatically learns the continuous pattern of joint displacements and outputs the motion time constraints (the reasonable motion path of the joint over time, represented by a smooth velocity curve).

[0137] The trajectory optimization and analysis module employs a constrained gradient descent algorithm, combining motion space boundary conditions and motion time constraints. Within the solution space defined by the motion space boundary conditions, the algorithm uses the motion time constraints as the objective function for iterative optimization, generating a set of basic motion trajectories (a set of motion paths that conform to physical laws, each consisting of a sequence of joint point coordinates in consecutive time frames).

[0138] The probabilistic inference sublayer fits a probability distribution to the set of basic motion trajectories. This fitting process uses a mixture density network, which takes as input the spatiotemporal features of the set of basic motion trajectories and outputs a multimodal motion trajectory distribution (a Gaussian mixture model describing the likelihood of different motion trajectories, with each Gaussian component corresponding to a typical motion mode).

[0139] Based on the multimodal motion trajectory distribution, the random sampling calculation module implements an importance sampling strategy. Sampling weights are proportional to the probability density of each trajectory modality, prioritizing candidate trajectory segments (possible future motion segments, each containing predicted joint coordinates within the next 0.5 seconds) from high-probability regions. Annealing sampling is also employed to retain anomalous trajectory samples from low-probability regions at a fixed ratio.

[0140] The fusion output sublayer receives the basic motion trajectory set and the candidate motion trajectory segments and performs redundant trajectory filtering calculation. The filtering process is divided into two stages:

[0141] In the coarse filtering stage, the dynamic time warping distance between trajectories is calculated, and repeated trajectories with a distance less than a threshold are eliminated;

[0142] In the fine filtering stage, a trajectory feature graph is constructed, where nodes represent trajectory segments and edge weights reflect the similarity of motion directions. Similar trajectory clusters are merged using a graph clustering algorithm.

[0143] The final output of the second candidate trajectory set (the trajectory prediction results used for event analysis) retains the representative trajectories of each cluster center to ensure spatial coverage and motion diversity. Each type of trajectory is annotated with its probability of occurrence and physical plausibility score.

[0144] This embodiment performs spatiotemporal feature decomposition on the second motion coupling feature matrix through the deterministic trajectory sublayer of the motion trajectory generation layer, which can simultaneously extract spatial motion constraint features and temporal motion sequence features, ensuring that the motion trajectory prediction conforms to both biomechanical constraints and temporal continuity requirements, thereby improving the physical rationality and temporal consistency of trajectory generation. Based on the joint optimization of spatial motion constraint features and temporal motion sequence features, the generated basic motion trajectory set covers the main motion patterns that conform to the motion rules, providing a reliable initial trajectory sample for subsequent probabilistic deduction. The probabilistic deduction sublayer uses a mixed density network to fit the multimodal motion trajectory distribution, which can effectively capture the diverse motion strategies of athletes in different scenarios and avoid the limitations of a single deterministic prediction. Through a strategy combining importance sampling and annealing sampling, while retaining high-probability motion trajectories, it also takes into account low-probability but reasonable abnormal motion patterns, thereby enhancing the system's adaptability to sudden motion events.

[0145] In one embodiment, the character's main skeleton data and joint action parameters are fused and processed to output character driving event data, including:

[0146] The character's skeletal data is collected asynchronously via inertial sensors and an optical capture system, resulting in a maximum time deviation of ±20 milliseconds. Timestamp calibration uses the Network Time Protocol to align the skeletal data stream with the video stream's time base, and deviation compensation uses linear interpolation to fill in missing frames. The default interpolation window is 5 milliseconds. When three consecutive frames are missing, an anomaly detection is triggered, automatically discarding the unrecoverable skeletal frames. The time accuracy of the synchronized skeletal data is calibrated to ±1.5 milliseconds, achieving timeline alignment for multiple data streams.

[0147] Synchronized skeletal data establishes a local coordinate system with the pelvic joint as the origin, converting global coordinates into relative coordinates using a quaternion rotation matrix. Pre-set spinal joint chain transformation rules maintain a relative distance error of less than 2.5 cm between the shoulder and hip joints, eliminating coordinate shifts caused by turning or leaning. The character's unified coordinate skeletal data includes the local position coordinates and rotation quaternions for 23 major skeletal joints, with spatial consistency error controlled within 2 mm.

[0148] Joint motion parameters are fed into an inverse kinematics validation model to check if interphalangeal joint flexion angles conform to the biomechanical range. The preset metacarpophalangeal joint flexion angle threshold is 140 degrees. When the distal interphalangeal joint angle of the ring finger reaches 150 degrees, the angle is corrected to 130 degrees through reverse iterative calculation along the finger kinematic chain. Compliant joint data is validated using an energy consumption model, eliminating abnormal trajectories that consume more than 400 joules per unit time.

[0149] The character's unified coordinate skeleton data and compliant joint point data are fed into an adaptive weighted fusion module, with fusion weights dynamically adjusted based on joint motion speed. When hip joint speed exceeds 2.5 m / s, the joint motion parameter weight is set to 0.65; below 1 m / s, the weight is reduced to 0.35. The maximum fusion error of the character's skeletal drive parameters in joint position is kept within 3 mm, improving trajectory smoothness by 35%.

[0150] The character's skeletal drive parameters are fed into the spatiotemporal pyramid feature extractor, which extracts joint acceleration, angular rate of change, and trajectory curvature features within 0.5-second, 1-second, and 2-second time windows. The event classifier uses a pretrained gradient boosting decision tree model to perform multi-dimensional decision-making for three event categories: shots, passes, and interceptions. The default decision boundary threshold is 0.6. If the similarity between the feature vector and all event categories falls below the threshold, an unidentified event marker is output. Finally, character-driven event data is generated, which can be applied to scenarios such as film and television animation production and video game development.

[0151] This embodiment achieves precise synchronization of multi-source sensor data through timestamp calibration processing, eliminates motion phase deviations caused by clock differences in acquisition devices, and ensures the temporal consistency of motion analysis. The spatial coordinate system conversion processing unifies the spatial reference of skeletal data in different body positions, reduces coordinate offset errors caused by athlete rotation or displacement, and enhances the spatial comparability of motion trajectories. The kinematic constraint verification processing automatically corrects abnormal joint angles through biomechanical rules, avoids the generation of motion trajectories that violate ergonomics, and improves the physiological rationality of motion data. The dynamic fusion processing adopts a speed-adaptive weight adjustment strategy to balance the relationship between the stability of the main skeleton data and the sensitivity of the detailed joint data, and generates a high-precision complete motion trajectory that is consistent in time and space.

[0152] Reference Figure 2 As shown, the present invention also provides a method system for generating real-time AI-driven joint data for ball sports, which is applied to any of the above-mentioned methods for generating real-time AI-driven joint data for ball sports, comprising:

[0153] The acquisition module is used to obtain the main skeleton joint point data and motion joint point data, and perform motion noise elimination and trajectory smoothing to obtain full-body motion training data;

[0154] An analysis module is used to train a preset self-attention neural network model based on whole-body motion training data to obtain a human joint data prediction model for ball sports;

[0155] The association module is used to obtain the main skeleton data of the character in the real-time game scene, and input it into the ball sports human joint data prediction model to predict the joint point data and output the joint point action parameters;

[0156] The processing module is used to fuse the character's main skeleton data with the joint action parameters and output the character drive event data.

[0157] It should be noted that, those skilled in the art will clearly understand that, for the sake of convenience and brevity of description, the specific working processes of the above-described system and each module can refer to the corresponding processes in the aforementioned method embodiments and will not be repeated here.

[0158] The above description is only a preferred embodiment of the present invention and does not limit the patent scope of the present invention. Any equivalent structure or equivalent process transformation made by using the contents of the present invention description and drawings, or directly or indirectly applied in other related technical fields, are also included in the patent protection scope of the present invention.

Claims

1. A method for generating real-time AI-driven joint data for ball sports, characterized in that: include: Obtain the main skeleton joint point data and motion joint point data, and perform motion noise elimination and trajectory smoothing to obtain full-body motion training data; Performing model training on a preset self-attention neural network model based on the whole-body motion training data to obtain a human joint data prediction model for ball sports; Obtaining the main skeleton data of the character in the real-time game scene, and inputting it into the ball sports human joint data prediction model to perform joint point data prediction and output joint point motion parameters; Fusing the character's main skeleton data with the joint action parameters to output character drive event data; The method of performing model training on a preset self-attention neural network model based on the whole-body motion training data to obtain a ball sports human joint data prediction model includes: Performing joint angle conversion analysis on the whole-body motion training data to obtain joint rotation matrix data; Performing intention feature recognition on the whole-body motion training data through the motion intention perception layer of the self-attention neural network model to obtain a first intention perception vector; Performing motion feature coupling on the joint rotation matrix data through a motion rule coupling layer to obtain a first motion coupling feature matrix; Performing feature fusion on the first intention perception vector and the first motion coupling feature matrix through a motion trajectory generation layer to obtain a first candidate trajectory set; Performing time series prediction on the first candidate trajectory set through a motion performance evolution layer to obtain first predicted motion joint point data; Performing a multi-order trajectory error calculation on the whole-body motion training data and the first predicted motion joint point data according to a preset loss function to obtain training error data; Iteratively optimizing the model parameters of the self-attention neural network model based on the training error data, the first candidate trajectory set, and the first predicted motion joint point data to obtain the ball sports human joint data prediction model; The method of obtaining the main skeleton data of the character in the real-time game scene and inputting it into the ball sports human joint data prediction model to perform joint point data prediction and output joint point motion parameters includes: Performing a skeletal topology analysis on the character's main skeletal data to obtain an initial joint point topology map; Inputting the initial joint point topology map into the motion intention perception layer, performing dynamic correlation analysis between joint points based on the global attention mechanism of the motion intention perception layer, and obtaining a second intention perception vector; Inputting the second intention perception vector into the motion rule coupling layer, performing constraint matching on the second intention perception vector based on a preset motion rule knowledge graph, and obtaining a second motion coupling feature matrix; Inputting the second motion coupling feature matrix into the motion trajectory generation layer to perform multi-scale motion trajectory prediction to obtain a second candidate trajectory set; Inputting the second candidate trajectory set into the motion performance evolution layer for biomechanical rationality verification, and outputting optimized trajectory data; Target frame extraction and cubic spline interpolation are performed on the optimized trajectory data to obtain the joint point motion parameters.

2. The method for generating real-time AI-driven joint data for ball sports according to claim 1, characterized in that: The method of obtaining main skeleton joint point data and motion joint point data, performing motion noise elimination and trajectory smoothing processing to obtain whole-body motion training data includes: Classify and process the three-dimensional coordinate data of the whole body joints collected by the optical motion capture device to obtain the main skeleton joint data and the motion joint data; Performing motion noise elimination processing on the main skeleton joint point data and the motion joint point data respectively to obtain main skeleton denoised data and motion joint point denoised data; Performing layered trajectory optimization on the main skeleton denoised data and the motion joint denoised data respectively to obtain main skeleton trajectory data and motion joint trajectory data; Performing time alignment and spatial normalization processing on the main skeleton trajectory data and the motion joint trajectory data to obtain a structured motion training data set; The structured motion training data set is subjected to action semantic segmentation and annotation to obtain the whole-body motion training data.

3. The method for generating real-time AI-driven joint data for ball sports according to claim 2, characterized in that: The step of performing layered trajectory optimization on the main skeleton denoised data and the motion joint denoised data to obtain main skeleton trajectory data and motion joint trajectory data comprises: Performing rigid body dynamics constraint optimization on the main skeleton denoising data to obtain a first smooth trajectory; performing activation constraint optimization and joint range of motion verification on the first smooth trajectory to obtain a second smooth trajectory; Performing kinematic chain consistency optimization on the second smooth trajectory to obtain the main skeleton trajectory data; performing micro-tremor joint trajectory optimization on the denoised data of the motion joint points to obtain a preliminary end joint point trajectory; The preliminary end joint point trajectory is subjected to kinematic coupling optimization processing to obtain the motion joint point trajectory data.

4. The method for generating real-time AI-driven joint data for ball sports according to claim 1, characterized in that: The second intention perception vector is input into the motion rule coupling layer, and constraint matching is performed on the second intention perception vector based on a preset motion rule knowledge graph to obtain a second motion coupling feature matrix, including: Performing joint torque force analysis on the second intention perception vector through the compliance sublayer of the motion rule coupling layer to obtain joint force distribution; Compliance verification is performed on the joint force distribution according to preset force constraint conditions to obtain a compliance motion feature; Performing similarity matching between the compliant motion features and the motion rule knowledge graph through the motion rule matching sublayer to obtain motion rule matching features; Performing mutation point interpolation correction on the motion rule matching feature to obtain a rule enhancement feature; The compliant motion features and the rule-enhanced features are tensor-concatenated through a feature fusion sublayer to obtain the second motion coupling feature matrix.

5. The method for generating real-time AI-driven joint data for ball sports according to claim 1, characterized in that: The step of inputting the second motion coupling feature matrix into the motion trajectory generation layer to perform multi-scale motion trajectory prediction to obtain a second candidate trajectory set includes: Performing spatiotemporal feature constraints on the second motion coupling feature matrix through the deterministic trajectory sublayer of the motion trajectory generation layer to obtain spatiotemporal coupling constraint features; Based on the spatiotemporal coupling constraint characteristics, the spatial feasible range and temporal continuity of the joint point motion are calculated, and basic trajectory identification is performed to obtain a basic motion trajectory set; Performing probability distribution fitting on the basic motion trajectory set through a probabilistic deduction sublayer to obtain candidate motion trajectory segments; Redundant trajectory filtering calculation is performed on the basic motion trajectory set and the candidate motion trajectory segments to obtain the second candidate trajectory set.

6. The method for generating real-time AI-driven joint data for ball sports according to claim 1, characterized in that: The step of fusing the character's main skeleton data with the joint action parameters and outputting character driving event data comprises: Performing game engine time synchronization calibration on the character's main skeleton data to obtain synchronized skeleton data; Performing character space coordinate system conversion on the synchronized skeleton data to obtain character unified coordinate skeleton data; Performing kinematic constraint verification on the joint motion parameters to obtain compliant joint data; Performing character dynamic fusion on the unified coordinate skeleton data of the character and the compliant joint point data to obtain character skeleton driving parameters; Whole-body data is generated for the character's skeletal drive parameters to obtain the character's drive event data.

7. A method and system for generating real-time AI-driven joint data for ball sports, characterized in that: The method for generating real-time AI-driven joint data for ball sports as described in any one of claims 1 to 6 above comprises: An acquisition module is used to obtain main skeleton joint point data and motion joint point data, and perform motion noise elimination and trajectory smoothing processing to obtain whole-body motion training data; An analysis module, the analysis module being used to perform model training on a preset self-attention neural network model based on the whole-body motion training data to obtain a ball sports human joint data prediction model; An association module is used to obtain the main skeleton data of the character in the real-time game scene, and input it into the ball sports human joint data prediction model to perform joint point data prediction and output joint point motion parameters; A processing module, the processing module is used to fuse the character's main skeleton data with the joint point action parameters and output character driving event data; The method of performing model training on a preset self-attention neural network model based on the whole-body motion training data to obtain a ball sports human joint data prediction model includes: Performing joint angle conversion analysis on the whole-body motion training data to obtain joint rotation matrix data; Performing intention feature recognition on the whole-body motion training data through the motion intention perception layer of the self-attention neural network model to obtain a first intention perception vector; Performing motion feature coupling on the joint rotation matrix data through a motion rule coupling layer to obtain a first motion coupling feature matrix; Performing feature fusion on the first intention perception vector and the first motion coupling feature matrix through a motion trajectory generation layer to obtain a first candidate trajectory set; Performing time series prediction on the first candidate trajectory set through a motion performance evolution layer to obtain first predicted motion joint point data; Performing a multi-order trajectory error calculation on the whole-body motion training data and the first predicted motion joint point data according to a preset loss function to obtain training error data; Iteratively optimizing the model parameters of the self-attention neural network model based on the training error data, the first candidate trajectory set, and the first predicted motion joint point data to obtain the ball sports human joint data prediction model; The method of obtaining the main skeleton data of the character in the real-time game scene and inputting it into the ball sports human joint data prediction model to perform joint point data prediction and output joint point motion parameters includes: Performing a skeletal topology analysis on the character's main skeletal data to obtain an initial joint point topology map; Inputting the initial joint point topology map into the motion intention perception layer, performing dynamic correlation analysis between joint points based on the global attention mechanism of the motion intention perception layer, and obtaining a second intention perception vector; Inputting the second intention perception vector into the motion rule coupling layer, performing constraint matching on the second intention perception vector based on a preset motion rule knowledge graph, and obtaining a second motion coupling feature matrix; Inputting the second motion coupling feature matrix into the motion trajectory generation layer to perform multi-scale motion trajectory prediction to obtain a second candidate trajectory set; Inputting the second candidate trajectory set into the motion performance evolution layer for biomechanical rationality verification, and outputting optimized trajectory data; Target frame extraction and cubic spline interpolation are performed on the optimized trajectory data to obtain the joint point motion parameters.

Citation Information

Patent Citations

  • Human body motion track prediction method and system based on multi-output space-time interaction

    CN117474945A

  • Human-robot collaboration method based on multi-scale graph convolutional neural network

    US12159486B1