A method, system, device and medium for predicting joint angles of infant crawling

By combining multi-joint collaboration theory and cross-attention network model with spatiotemporal dual discriminator GAN network, the joint angles of infants crawling are generated and predicted, solving the problems of data scarcity and large individual variability, achieving high-precision joint angle prediction, and improving the safety and effectiveness of crawling training.

CN120600226BActive Publication Date: 2025-11-11NANCHANG HANGKONG UNIVERSITY
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511099770.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-08-07
Publication Date
2025-11-11
Estimated Expiration
2045-08-07

AI Technical Summary

Technical Problem

Existing technologies lack data on infant crawling movements, and individual movement variability is high, making it impossible to accurately predict joint angles during crawling. Furthermore, traditional rehabilitation training devices lack initiative, posing safety hazards and poor effectiveness.

Method used

By adopting the multi-joint collaboration theory, we acquire three-dimensional coordinate data of infants crawling, generate augmented data using randomly sampled noise vectors, construct a joint collaboration and cross-attention network model (JoCrCAPNet), decompose real data into posture collaboration matrix and temporal recruitment matrix, and combine spatiotemporal dual discriminator GAN network to generate simulated data that conforms to biomechanical laws, thereby achieving high-precision joint angle prediction.

Benefits of technology

It solves the problem of scarce infant and toddler movement data, achieves high-precision joint angle prediction, enhances the initiative and safety of crawling training, and improves the effectiveness of rehabilitation training.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120600226B_ABST
    Figure CN120600226B_ABST
Patent Text Reader

Abstract

This invention provides a method, system, device, and medium for predicting joint angles during infant crawling, belonging to the field of medical artificial intelligence. The method includes: acquiring joint coordinates of infant crawling using a 3D motion capture system, converting them into real joint angles, generating simulated data using a spatiotemporal dual discriminator, and mixing them to form an extended dataset; extracting a posture coordination matrix and a temporal recruitment matrix through principal component analysis, projecting the real and simulated data into low-dimensional coordination coefficients, and constructing hybrid coordination features; employing dual encoders to capture the local dynamic features of joint angles and the global motion patterns of coordination coefficients respectively, fusing multimodal features through a cross-attention mechanism, and ultimately achieving accurate prediction of future joint angles. This innovative combination of data augmentation and biomechanical collaborative modeling significantly alleviates the problem of scarce infant data, improves prediction accuracy, and reduces training risks. It is applicable to infant rehabilitation training, intelligent rehabilitation engineering, and clinical motor function reconstruction.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of medical artificial intelligence, specifically relating to a method, system, device, and medium for predicting joint angles during infant crawling. Background Technology

[0002] Crawling is of great value for the motor development of normal infants and the functional rehabilitation of children with fish hole disorder. Crawling involves not only the coordination of flexion and extension of large joints such as the elbows and knees, but also the coordination of senses such as vision and touch, thus playing a positive role in promoting brain development. Sufficient crawling experience also lays the foundation for infants' subsequent upright walking and helps improve their spatial cognitive abilities. Furthermore, for infants with motor development problems such as cerebral palsy, early targeted crawling training can simultaneously stimulate the activity of multiple joints and muscle groups. Through regular, repetitive alternating limb movements, it helps children rebuild correct motor control patterns. Therefore, crawling can effectively improve children's limb coordination, promote the recovery and reconstruction of neuromuscular function, and thus enhance the overall rehabilitation effect.

[0003] Therefore, in recent years, many rehabilitation training devices have emerged that primarily function and function as crawling training. These devices typically use motor-driven passive methods to guide patients in crawling movements, thereby achieving the effect of motor function rehabilitation training. However, in practical applications, this method of control based on predefined crawling patterns ignores the initiative and enthusiasm of infants and young children in crawling movements. Crawling training devices may exhibit problems such as abnormal motor coordination and dragging of the wearer, which cannot guarantee the safety of infants and young children, and the rehabilitation training effect is also poor. Improper operation may even cause further muscle damage to the patient.

[0004] On the other hand, although deep learning technology has made significant progress in time series prediction such as joint angle changes, research on joint angle prediction of infant crawling movement based on deep learning is still in its early stages due to the scarcity of infant movement data and the large variability of individual movement. Summary of the Invention

[0005] To address the issues of scarce infant movement data, high individual movement variability, and inability to accurately predict joint angles during crawling, as mentioned above, this invention provides a method, system, computer device, and storage medium for predicting joint angles during infant crawling based on multi-joint coordination theory. To achieve the above objectives, this invention provides the following technical solution.

[0006] A method for predicting joint angles during infant crawling includes:

[0007] The three-dimensional coordinate data of the limb joints of infants crawling are obtained and converted into real joint angle data V_real; three-dimensional enhanced simulated crawling coordinate data is generated using randomly sampled noise vectors and converted into simulated joint angle data V_gen; V_real and V_gen are mixed to obtain the joint angle extended dataset V_mixed.

[0008] Decompose V_real to obtain the time-based fundraising matrix. H and pose coordination matrix W ;use W Projecting V_real yields the true low-dimensional cooperability coefficient sequence C_real; using W Projecting V_gen yields the simulated low-dimensional cooperative coefficient sequence C_gen. H C_gen is enhanced to generate a time-enhanced simulation covariance coefficient sequence C_gen_enhanced; C_real and C_gen_enhanced are mixed to form an extended covariance coefficient sequence C_mixed; the JoCrCAPNet model is used to predict infant crawling data to obtain the joint angle data of future time steps of infant crawling.

[0009] A joint coordination and cross-attention network model, JoCrCAPNet, was constructed. V_mixed and C_mixed were input into the JoCrCAPNet model for training, resulting in a model that predicts the joint angle data of future time steps of infants and toddlers while crawling.

[0010] Preferably, the step of acquiring the three-dimensional coordinate data of the limb joints of an infant crawling and converting the three-dimensional coordinate data into real joint angle data V_real specifically includes:

[0011] Noise in the 3D coordinate data is eliminated, and missing data in the 3D coordinate data is interpolated and filled to obtain complete crawling data;

[0012] Based on anatomical calibration of joint centers, and with standard crawling posture as a reference, a local coordinate system is established at each joint center of the complete crawling data;

[0013] The local coordinate system is normalized according to the proportion of infant limb length to obtain a normalized joint coordinate system, and the relative rotation matrix of the normalized joint coordinate system is constructed.

[0014] The true joint angle data V_real is calculated using the relative rotation matrix of adjacent normalized joint coordinate systems.

[0015] Preferably, the randomly sampled noise vector is processed by a spatiotemporal dual-discriminator GAN network to generate three-dimensional enhanced simulated crawling coordinate data. The spatiotemporal dual-discriminator GAN network is composed of a generator, a spatial discriminator, and a temporal discriminator connected in sequence; including:

[0016] The generator generates 3D simulated crawling coordinate data based on randomly sampled noise vectors;

[0017] The real 3D coordinate data and the 3D simulated crawling coordinate data are input into the spatial discriminator and the temporal discriminator respectively. The spatial discriminator extracts the spatial features of the real 3D coordinate data and the 3D simulated crawling coordinate data, and the temporal discriminator extracts the temporal features of the real 3D coordinate data and the 3D simulated crawling coordinate data.

[0018] The parameters of the spatial discriminator and the temporal discriminator are optimized alternately by utilizing the spatial and temporal features of real 3D coordinate data and 3D simulated crawling coordinate data. The generator parameters are then adjusted in reverse by using the spatial and temporal discriminators, ultimately enabling the spatiotemporal dual discriminator GAN to generate 3D enhanced simulated crawling coordinate data that conforms to biomechanical laws.

[0019] Preferably, the spatial discriminator consists of an input layer, a first convolutional layer, a ReLU layer, a fully connected layer, and an output layer connected in sequence, wherein the first convolutional layer includes a Conv1d layer and a ReLU layer; the temporal discriminator consists of an input layer, a second convolutional layer, a ReLU layer, a fully connected layer, and an output layer connected in sequence, wherein the second convolutional layer includes a Conv1d layer, a ReLU layer, and a MaxPooling layer.

[0020] Preferably, the decomposition of V_real yields the time-based recruitment matrix. H and pose coordination matrix W This includes: decoupling the V_real into a time-based recruitment matrix using principal component analysis (PCA). H and pose coordination matrix W, Among them, the time-based fundraising matrix H Describes the temporal dynamics of crawling motion, recording when other joints are activated for coordinated crawling; posture coordination matrix. W Describe the spatial coordination structure of crawling motion and record the spatial association patterns of joint movements.

[0021] Preferably, the V_mixed and C_mixed models are trained using the joint coordination and cross-attention network model JoCrCAPNet to obtain the joint angle data of the infant's future time steps during crawling. The joint coordination and cross-attention network model JoCrCAPNet consists of a coordination coefficient encoder, a joint angle encoder, an LSTM layer, a dual-path cross-attention module, an LTMCell layer, and a fully connected layer connected sequentially. The joint angle encoder is used to capture local joint angle features from V_mixed, and the coordination coefficient encoder is used to capture global motion features from C_mixed.

[0022] Preferably, the local joint angle features and global motion features are cross-fused by the dual-path cross-attention module to obtain fused features; the fused features are processed by the LTMCell layer and the fully connected layer to predict the joint angle data of the infant's future time step during crawling; wherein, the dual-path cross-attention module is composed of an input module, an attention calculation module and a fusion and output module connected in sequence.

[0023] This invention also provides a joint angle prediction system for infant crawling, comprising:

[0024] The dataset construction module is used to obtain the three-dimensional coordinate data of the limb joints of infants and toddlers during crawling, and convert the three-dimensional coordinate data into real joint angle data V_real; it uses randomly sampled noise vectors to generate three-dimensional enhanced simulated crawling coordinate data, and converts the three-dimensional enhanced simulated crawling coordinate data into simulated joint angle data V_gen; it mixes V_real and V_gen to obtain the joint angle extended dataset V_mixed.

[0025] The data processing module is used to decompose V_real to obtain the time-based fundraising matrix. H and pose coordination matrix W ;use W Projecting V_real yields the true low-dimensional cooperability coefficient sequence C_real; using W Projecting V_gen yields the simulated low-dimensional cooperative coefficient sequence C_gen. H Enhance C_gen to generate a time-enhanced simulation covariance coefficient sequence C_gen_enhanced; mix C_real with C_gen_enhanced to form an extended covariance coefficient sequence C_mixed.

[0026] The model training and application module is used to construct the joint coordination and cross-attention network model JoCrCAPNet. V_mixed and C_mixed are input into the JoCrCAPNet model for training to obtain a model that predicts the joint angle data of the future time step of an infant's crawling. By using the JoCrCAPNet model to predict the crawling data of an infant, the joint angle data of the future time step of the infant's crawling can be obtained.

[0027] The present invention also provides a computer device, including a memory, a processor, and a computer program stored in the memory, wherein the processor executes the computer program to implement the steps of any one of the methods for predicting joint angles during infant crawling.

[0028] The present invention also provides a computer-readable storage medium storing a computer program that, when loaded by a processor, is capable of executing any of the steps in the method for predicting joint angles during infant crawling.

[0029] The method for predicting joint angles during infant crawling provided by this invention has the following beneficial effects:

[0030] This invention addresses the scarcity of infant movement data by collecting three-dimensional coordinate data of infant crawling, generating augmented data using randomly sampled noise vectors, and converting it into a joint angle dataset. It decomposes real data to obtain a posture coordination matrix and a time recruitment matrix, projecting real and simulated data into low-dimensional coordination coefficient sequences to construct a hybrid coordination feature set. By extracting local dynamic features of joint angles and global motion patterns of coordination coefficients, it achieves high-precision, multi-joint coordinated prediction of future joint angles during infant crawling. This accurate prediction of changes in limb joint angles during infant crawling helps improve the active training effect and application safety of crawling devices. Attached Figure Description

[0031] To more clearly illustrate the embodiments and design schemes of the present invention, the accompanying drawings required for this embodiment will be briefly described below. The drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0032] Figure 1 This is a flowchart of a method for predicting joint angles during infant crawling, according to an embodiment of the present invention.

[0033] Figure 2 Joint reflex marker location diagram of this invention embodiment;

[0034] Figure 3Illustrated diagram of joint angle calculation during infant crawling according to an embodiment of the present invention;

[0035] Figure 4 The spatiotemporal dual discriminator GAN structure diagram of this invention embodiment;

[0036] Figure 5 This is the generator structure according to an embodiment of the present invention;

[0037] Figure 6 This is the spatial discriminator structure according to an embodiment of the present invention;

[0038] Figure 7 This is the structure of the time discriminator in an embodiment of the present invention;

[0039] Figure 8 This is the JoCrCAPNet architecture, a joint collaboration and cross-attention network model according to an embodiment of the present invention.

[0040] Figure 9 This is a diagram of a dual-path cross-attention structure according to an embodiment of the present invention. Detailed Implementation

[0041] To enable those skilled in the art to better understand and implement the technical solutions of the present invention, the present invention will be described in detail below with reference to the accompanying drawings and specific embodiments. The following embodiments are only used to more clearly illustrate the technical solutions of the present invention and should not be construed as limiting the scope of protection of the present invention.

[0042] This invention provides a method for predicting joint angles during infant crawling, specifically as follows: Figure 1 As shown, it includes the following steps.

[0043] S1. Obtain the three-dimensional coordinate data of the joints of the limbs of infants and toddlers when they crawl, and convert the three-dimensional coordinate data into real joint angle data V_real; generate three-dimensional enhanced simulated crawling coordinate data using randomly sampled noise vectors, and convert the three-dimensional enhanced simulated crawling coordinate data into simulated joint angle data V_gen; mix V_real and V_gen to obtain the joint angle extended dataset V_mixed.

[0044] Specifically, the process involves acquiring the three-dimensional coordinate data of the joints of an infant's limbs during crawling, converting the three-dimensional coordinate data into real joint angle data V_real, generating three-dimensional enhanced simulated crawling coordinate data through a spatiotemporal dual discriminator GAN network, converting the three-dimensional enhanced simulated crawling coordinate data into simulated joint angle data V_gen, and mixing V_real and V_gen to obtain the joint angle extended dataset V_mixed.

[0045] This invention proposes to collect three-dimensional coordinate data of the main joints of the limbs of infants and young children during crawling by using multiple optical motion capture cameras. Based on this, data augmentation is used to solve the problem of data scarcity, and crawling motion coordination patterns are extracted to predict joint angles. Since the motion coordination relationship between joints can explain the crawling motion state at a deeper level, this can make the crawling motion more accurate in terms of data generation and joint angle prediction, and further provide a reference for the compliant control of crawling rehabilitation training robots.

[0046] This embodiment uses the Raptor-E 3D motion capture system from Motion Analysis, Inc. (USA) for kinematic data acquisition (sampling rate 100Hz). Subjects wore only diapers, and reflective marker balls were precisely attached to key anatomical landmarks (lateral acromion, lateral epicondyle of the humerus, ulnar styloid process, posterior superior iliac spine, lateral knee line, lateral malleolus, and scapula). Before the formal experiment, subjects acclimatized to the environment for 10 minutes on a dedicated 360cm × 120cm crawling mat. Valid data must meet the following criteria: a linear crawling sequence containing three or more consecutive complete crawling cycles with a trajectory deviation angle of less than 5 degrees, while excluding transitional gaits at the beginning and end phases.

[0047] The collected data, consisting of the three-dimensional coordinates of the major joints of the limbs during infant crawling, first underwent a cleanliness process by detecting non-physiological mutation points and missing segments to remove invalid cycles. Then, to address the issue of inconsistent cycle lengths, the motion sequences were segmented using the right wrist joint contact point as a reference, and the cycle length was standardized to 0%–100% through cubic spline interpolation. Next, a local coordinate system was constructed with the contact point as the origin, transforming the global coordinates to the local space. Finally, Min-Max normalization was used to eliminate individual body size differences, resulting in a standardized crawling dataset. This dataset was used to train the generative model for data augmentation.

[0048] In the angle calculation section, the crawling motion data generated by the model and the original collected crawling motion standardized dataset are first subjected to inverse Min-Max normalization to restore the joint 3D coordinates back to the original physical scale. Body joint markers are as follows: Figure 2 As shown, the joint angles of the eight joints (elbow, shoulder, hip, and knee) on both sides are calculated using the spatial vector angle calculation formula, and then normalized by dividing by 180. The three-dimensional coordinate system is established as follows. Figure 3 As shown, a three-dimensional coordinate system is established using the shoulder, elbow, and wrist. For example, the formula for calculating the elbow joint angle is as follows:

[0049] ;

[0050] In the formula, ( S x , S y ,S z )and( E x , E y , E z ) represent the positions of the shoulder joint and elbow joint in three-dimensional space, respectively. W x , W y , W z () represents the position of the wrist joint. The difference in position in the same dimension is obtained by subtracting the coordinates of adjacent joints.

[0051] A crawling motion data generation module is used to expand training data. Based on the traditional Generative Adversarial Network (GAN) model, and based on the principle of human motion coordination, it considers the spatiotemporal correlation between joint activities during infant movement from both temporal and spatial perspectives. This integrates the laws of human joint coordination with time series characteristics, and proposes a method such as... Figure 4 The method for generating crawling motion data based on a spatiotemporal dual discriminator GAN is shown. This model extracts joint spatial correlation features through a spatial discriminator and captures temporal variation patterns through a temporal discriminator. The two work together to supervise the generator to learn spatiotemporal characteristics, ultimately generating infant crawling data that conforms to real motion characteristics.

[0052] The model is based on the game-theoretic framework of GANs and incorporates the loss function design of WGANs. This loss function, consisting of four parts—reconstruction loss, adversarial loss, supervision loss, and embedding loss—works together to optimize the generator and discriminator. Training is then completed through dynamic adversarial interaction between the generator and the dual discriminator. During feature reconstruction, the generator decodes high-dimensional spatial features compressed along the time axis into temporal motion trajectories. During training, the temporal discriminator and spatial discriminator simultaneously receive real-world crawling motion data and their generated results, respectively judging their authenticity. The temporal discriminator focuses on analyzing the continuous temporal patterns of joint movements, while the spatial discriminator focuses on evaluating the spatial constraints between joints. Through a collaborative supervision mechanism, the dual discriminator drives the generator's generated data to simultaneously satisfy the temporal rationality of the crawling motion sequence and the spatial coordination of the joint structure.

[0053] like Figure 5As shown, the generator adopts a three-stage architecture of "spatial compression - channel expansion - temporal reconstruction," aiming to map high-dimensional spatial features into temporal motion trajectories that conform to biomechanical laws. The random input vector z corresponds to the high-dimensional spatial features of the main joint groups in infant crawling. It is progressively compressed in spatial dimension through two levels of one-dimensional convolutional layers, while simultaneously expanding the channel dimension to explicitly model the collaborative features of the joint kinematic chains (such as shoulder-elbow-wrist). The compressed spatial features are fused into a 500-dimensional feature vector through a fully connected (Dense Layer), and finally decoded into... (Corresponding to the three-dimensional coordinates X / Y / Z of the 12 joints in a single cycle of 101 time steps of crawling motion).

[0054] like Figure 6 As shown, the spatial discriminator is designed based on a hierarchical one-dimensional convolutional architecture, whose structure is adapted to the core constraints of the WGAN loss function. The first convolutional block 1 (kernel size 3, stride 3) extracts motion patterns along the three-dimensional coordinate axes of the joints to verify the biomechanical rationality of single-joint motion; the second convolutional block 2 (kernel size 3, stride 3) captures the spatial correlation of the limb kinematic chain (e.g., wrist-elbow-shoulder) through cross-joint feature fusion. The features are compressed into linear discriminant values ​​(without the sigmoid activation function) by a fully connected layer to meet the WGAN requirement for the discriminator's Lipschitz continuity. Through a stepped stride design (stride 3 corresponds to the cardinality of the joint coordinate dimension and the number of joints in the limb kinematic chain), the network achieves a progressive evaluation from local joint coordination to the rationality of the overall kinematic chain, ensuring that the generated trajectory conforms to both the single-joint three-dimensional motion constraints and the spatial dynamics of multi-joint linkage.

[0055] like Figure 7 As shown, the temporal discriminator evaluates the continuity of motion dynamics through a two-stage convolutional structure. The first-stage convolutional layer captures local dynamic features of adjacent time steps, achieving fine-grained supervision of motion coherence through a dense sliding convolution strategy. The second-stage convolutional layer extracts long-span time step correlation information and cross-time step beat coordination features (e.g., the temporal correlation between the end of wrist joint movement and the beginning of elbow joint movement). Features are filtered through a max-pooling layer to identify key motion event nodes (e.g., the transition point between the support phase and the swing phase in the crawling cycle), and finally, the discriminant value is directly output through a fully connected network without activation layers. This design also abandons the Sigmoid probability constraint of traditional GAN ​​activation functions, and satisfies the WGAN's requirement for the discriminator's Lipschitz continuity by outputting discriminant values ​​linearly, thereby achieving accurate supervision of the temporal rationality of the generated data.

[0056] S2. Decompose V_real to obtain the time-based fundraising matrix. H and pose coordination matrix W ;use WProjecting V_real yields the true low-dimensional cooperability coefficient sequence C_real; using W Projecting V_gen yields the simulated low-dimensional cooperative coefficient sequence C_gen. H Enhance C_gen to generate a time-enhanced simulation covariance coefficient sequence C_gen_enhanced; mix C_real with C_gen_enhanced to form an extended covariance coefficient sequence C_mixed.

[0057] Decompose the real joint angle data V_rea to obtain the time recruitment matrix. H and pose coordination matrix W ;use W Projecting the real joint angle data V_real, we obtain the real low-dimensional cooperability coefficient sequence C_real; using... W Projecting V_gen yields the simulated low-dimensional cooperative coefficient sequence C_gen. H C_gen is enhanced to generate a temporal enhanced simulation covariance coefficient sequence C_gen_enhanced; C_real and C_gen_enhanced are mixed to form an extended covariance coefficient sequence C_mixed; V_mixed is input into the joint angle encoder to capture local joint angle features.

[0058] This embodiment employs Principal Component Analysis (PCA) based on SVD, which can represent a complex original data matrix as the product of several smaller and simpler submatrices. These submatrices describe the important properties of the matrix. X Indicates to n Variables conduct m The dataset obtained from the observations, among which Indicates the first n The first variable was used for the... m Second observation.

[0059] ;

[0060] right X The PCA implementation is as follows, for the dataset X After centralization, singular value decomposition yields:

[0061] ;

[0062] In the formula X It is a temporal pose matrix describing the training set of the prediction model. Yes X Center the mean matrix. yes A result obtained after scaling A left singular vector matrix of order 1. yes The result is obtained after scaling by a certain ratio. yes A result obtained after scaling A right singular vector matrix of order 1. It is data The result is obtained after scaling by a certain ratio. It is A diagonal singular value matrix of order X, whose elements on the main diagonal are... These are called singular values, and they are arranged in descending order. The magnitude of a singular value reflects the importance of the singular vector carried by the corresponding singular value vector. Represents a left singular value vector. Represents a singular value, Represents a right singular value vector. Together they constitute a collaborative pattern of a certain principal component. This represents the number of data points in the training set (totaling the number of periods × 101 data points). This refers to the number of joints.

[0063] First, collaborative patterns were extracted from the joint angle data in the training set, including joint angle data from eight joints: left elbow, left shoulder, left hip, left knee, right elbow, right shoulder, right hip, and right knee. Therefore, it is a... m A data matrix of ×8 is called a pose matrix. V To address this, we used principal component analysis based on SVD to extract the coordinated patterns achieved by the limbs, from the pose matrix. V Extracting the pose coordination matrix W and time-based fundraising matrix H As shown in the formula:

[0064] ;

[0065] in, This embodiment indicates the number of major joints in the limbs during crawling. n =8, This represents the number of principal components obtained through matrix decomposition. This represents time, which is standardized to a fixed crawling cycle in this embodiment, i.e., 0~100%. This embodiment prioritizes retaining the first two main collaborative features with more obvious characteristics. This is because infants and young children have not yet developed mature neural control mechanisms, and their central nervous system has not yet formed highly optimized motor coordination patterns. The time recruitment matrix, i.e., the motor input coordination pattern, is the manifestation of the behavior of movement in time and space; therefore, this embodiment uses it as the main coordination pattern for subsequent joint angle prediction. Each motor input coordination pattern is... m× The matrix form, which contains Motion input curve of joint angle data.

[0066] The model improves the multi-head attention mechanism by incorporating joint coordination information, such as... Figure 8 As shown, this embodiment employs a "dual encoder-decoder" and "cross-attention" architecture, proposing a Joint Coordination and Cross-Attention Network (JoCrCAPNet). JoCrCAPNet is used to train V_mixed and C_mixed to obtain joint angle data for future time steps during crawling. The Joint Coordination and Cross-Attention Network model JoCrCAPNet consists of a coordinating coefficient encoder and a joint angle encoder, an LSTM layer, a dual-path cross-attention module, an LTMCell layer, and a fully connected layer connected sequentially. The joint angle encoder captures local joint angle features from V_mixed, and the coordinating coefficient encoder captures global motion features from C_mixed. The prediction process of the model consists of two LSTM modules, one above and one below, which are encoder modules. The encoder is a joint angle encoder, whose input is the original crawling joint angle sequence of the infant, and the input shape is (batch_size, sequence_length, feature_dim). The encoder above is a co-coefficient encoder, whose input is the PCA co-coefficient sequence pre-calculated from the original joint angle sequence, and the input shape is also (batch_size, sequence_length, feature_dim).

[0067] Where `batch_size` is the batch size, `sequence_length` is the sequence length, and `feature_dim` is the feature dimension of the sequence. The original joint angle sequence and the PCA co-coefficient sequence have the same `batch_size` and `sequence_length`. The `feature_dim` of the original joint angle sequence is 8-dimensional (eight joint angles: left elbow, left shoulder, left hip, left knee, right elbow, right shoulder, right hip, and right knee), while the `feature_dim` of the PCA co-coefficient sequence is 2-dimensional (including the co-coefficients of the first two main co-coefficient features extracted from the original joint angle sequence). The output of the joint angle encoder is the encoded feature `enc_raw` and the state (`h_raw`, `c_raw`), and the output of the co-coefficient encoder is the encoded feature `enc_pca` and the state (`h_pca`, `c_pca`).

[0068] Cooperative joint attention allows the decoder to extract local details related to the cooperative motion pattern from the output of the raw joint angle encoder based on the current state. The decoder's current hidden state serves as the query vector, actively retrieving fine-grained motion features matching the current cooperative pattern in the joint feature space. The attention is applied to the high-dimensional temporal feature key vector and value vector of the joint angle encoder output enc_raw, containing details of angle changes in the eight joints. Through this mechanism, the model can capture local dynamics associated with low-dimensional cooperative patterns (such as the periodicity of crawling motion) from the raw joint motion data.

[0069] S3. Construct a joint coordination and cross-attention network model JoCrCAPNet. Input V_mixed and C_mixed into the JoCrCAPNet model for training to obtain a model that predicts the joint angle data of the future time step of the infant crawling. By using the JoCrCAPNet model to predict the infant crawling data, the joint angle data of the future time step of the infant crawling can be obtained.

[0070] A dual-path cross-attention network is used to cross-fuse local joint angle features and global motion features to obtain fused features; based on the fused features, the joint angle data of infants and toddlers in future time steps during crawling are predicted.

[0071] like Figure 9As shown, the dual-path cross-attention network consists of a co-coefficient encoder and a joint angle encoder, an LSTM network, a cross-attention layer, an LTM Cell layer, and a fully connected layer connected sequentially. The dual-path cross-attention mechanism achieves dynamic fusion of global co-coefficient patterns and local joint features in the following way: h_decoder and c_decoder are the parameters of the decoder in the dual-path cross-attention network. Through initialization and continuous updating, h_decoder and c_decoder are responsible for transforming latent variables back to the distribution of the original data. First, the final hidden states of the dual encoders (joint angle encoder and co-coefficient encoder) are element-wise added to generate the initial state of the decoder, where h_decoder serves as the query vector Query for multi-head attention. During temporal prediction, the joint-to-co-attention module retrieves key contextual information from the output of the co-coefficient encoder (representing the global motion pattern after PCA dimensionality reduction) based on the current decoder state: using the decoder hidden state as the Query and the temporal step features (Key / Value) in the co-coefficient feature space as the retrieval object, the attention weights are used to locate the motion phase most relevant to the current prediction (such as the crawling propulsion phase). Simultaneously, the joint attention module works in reverse to capture local details from the raw joint data that match the global pattern. This bidirectional mechanism achieves deep coupling between global motion patterns and local biomechanical characteristics by dynamically adjusting attention weights (such as strengthening the co-activation intensity of the shoulder-knee joint pair), ultimately improving the accuracy and biorhythmic plausibility of joint angle prediction.

[0072] The context vectors output from the two cross-attention modules are concatenated with the decoder input to form a comprehensive input feature. The autoregressive decoder output initializes the decoder input with the last frame data of the joint sequence; it then performs predictions for 20 time steps (controlled by the parameter `output_steps=20`): this includes calculating the dual-path attention context, updating the decoder state via LTMCell, mapping the hidden state to the prediction output using a fully connected layer, and finally using the current prediction value as the input for the next step. The output integrates and stacks the prediction results from each time step, ultimately outputting joint angle predictions with dimensions (batch_size, 20, 8).

[0073] This embodiment compares and validates the above scheme with LSTM, BiLSTM, and AttLSTM networks. First, the joint angle sample set M is divided into training and test sets in an 8:2 ratio. Then, the training set is expanded 1:1 by the number of cycles using generated data. Next, to avoid information leakage, a PCA model retaining the first two principal components is constructed based on the training set. Then, its projection matrix is ​​applied to transform the joint angle data of the training and test sets respectively, generating a sequence of co-variable coefficients. Then, 5-fold cross-validation is used on the training set to select the optimal model parameters. Finally, the test set data is input into the optimal prediction model to obtain the corresponding joint angle prediction MAE, RMSE, and R² values. The results are shown in Table 1. This invention has better prediction results for the left elbow, right elbow, left knee, and right knee joints.

[0074] Table 1. Prediction results of different models for bilateral elbow and knee joints

[0075]

[0076] The experiment augmented the data of infants' crawling movements, extracted the multi-joint motion coordination patterns during the crawling process, and used the multi-joint motion input coordination patterns to predict the joint angles of crawling movements. This deeper level of multi-joint motion coordination made the prediction of joint angles in crawling movements more effective.

[0077] Based on the same concept, the present invention also provides a joint angle prediction system for infant crawling, comprising:

[0078] The dataset construction module is used to obtain the three-dimensional coordinate data of the limb joints of infants and toddlers during crawling, and convert the three-dimensional coordinate data into real joint angle data V_real; it uses randomly sampled noise vectors to generate three-dimensional enhanced simulated crawling coordinate data, and converts the three-dimensional enhanced simulated crawling coordinate data into simulated joint angle data V_gen; it mixes V_real and V_gen to obtain the joint angle extended dataset V_mixed.

[0079] The data processing module is used to decompose V_real to obtain the time-based fundraising matrix. H and pose coordination matrix W ;use W Projecting V_real yields the true low-dimensional cooperability coefficient sequence C_real; using W Projecting V_gen yields the simulated low-dimensional cooperative coefficient sequence C_gen. H Enhance C_gen to generate a time-enhanced simulation covariance coefficient sequence C_gen_enhanced; mix C_real with C_gen_enhanced to form an extended covariance coefficient sequence C_mixed.

[0080] The model training and application module is used to construct the joint coordination and cross-attention network model JoCrCAPNet. V_mixed and C_mixed are input into the JoCrCAPNet model for training to obtain a model that predicts the joint angle data of the future time step of an infant's crawling. By using the JoCrCAPNet model to predict the crawling data of an infant, the joint angle data of the future time step of the infant's crawling can be obtained.

[0081] This invention also provides a computer device. At the hardware level, the computer device includes a processor, an internal bus, a network interface, memory, and non-volatile memory, and may also include other hardware required for business operations. The processor reads the corresponding computer program from the non-volatile memory into the memory and then runs it to implement the aforementioned method for predicting joint angles during infant crawling.

[0082] The present invention also provides a computer-readable storage medium storing a computer program that can be used to execute the above-described method for predicting joint angles during infant crawling.

[0083] Specific limitations regarding the calculation system for predicting joint angles in infant crawling can be found in the limitations of the joint angle prediction method for infant crawling described above, and will not be repeated here. Each module in the aforementioned joint angle prediction system for infant crawling can be implemented entirely or partially through software, hardware, or a combination thereof. These modules can be embedded in or independent of the processor in a computer device, or stored in the memory of a computer device as software, so that the processor can call and execute the corresponding operations of each module.

[0084] It should be noted that the specific embodiments described above enable those skilled in the art to more fully understand the present invention, but do not limit the present invention in any way. Therefore, although the present invention has been described in detail in this specification and embodiments, those skilled in the art should understand that modifications or equivalent substitutions can still be made to the present invention; and all technical solutions and improvements that do not depart from the spirit and scope of the present invention are covered within the protection scope of the patent of the present invention. No reference numerals in the claims should be construed as limiting the scope of the claims.

Claims

1. A method for predicting joint angles during infant crawling, characterized in that, include: Obtain the three-dimensional coordinate data of the limb joints of infants crawling, and convert the three-dimensional coordinate data into real joint angle data V_real; generate three-dimensional enhanced simulated crawling coordinate data using randomly sampled noise vectors, and convert the three-dimensional enhanced simulated crawling coordinate data into simulated joint angle data V_gen; mix V_real and V_gen to obtain the joint angle extended dataset V_mixed; Decompose V_real to obtain the time-based fundraising matrix. H and pose coordination matrix W ; use W Projecting V_real yields the true low-dimensional cooperability coefficient sequence C_real; using W Projecting V_gen yields the simulated low-dimensional cooperative coefficient sequence C_gen. H Enhance C_gen to generate a time-enhanced simulation covariance coefficient sequence C_gen_enhanced; mix C_real with C_gen_enhanced to form an extended covariance coefficient sequence C_mixed. A joint coordination and cross-attention network model, JoCrCAPNet, is constructed. V_mixed and C_mixed are input into the JoCrCAPNet model for training, resulting in a model that predicts the joint angle data of infants and toddlers in future time steps during crawling. By using the JoCrCAPNet model to predict infants and toddlers' crawling data, the joint angle data of infants and toddlers in future time steps during crawling can be obtained. The process of acquiring the three-dimensional coordinate data of the limb joints of an infant crawling and converting the three-dimensional coordinate data into real joint angle data V_real specifically includes: Noise in the 3D coordinate data is eliminated, and missing data in the 3D coordinate data is interpolated and filled to obtain complete crawling data; Based on anatomical calibration of joint centers, and with standard crawling posture as a reference, a local coordinate system is established at each joint center of the complete crawling data; The local coordinate system is normalized according to the proportion of infant limb length to obtain a normalized joint coordinate system, and the relative rotation matrix of the normalized joint coordinate system is constructed. The true joint angle data V_real is calculated by the relative rotation matrix of adjacent normalized joint coordinate systems. The randomly sampled noise vector is processed by a spatiotemporal dual-discriminator GAN network to generate three-dimensional enhanced simulated crawling coordinate data; the spatiotemporal dual-discriminator GAN network is composed of a generator, a spatial discriminator, and a temporal discriminator connected in sequence; including: The generator generates 3D simulated crawling coordinate data based on randomly sampled noise vectors; The real 3D coordinate data and the 3D simulated crawling coordinate data are input into the spatial discriminator and the temporal discriminator respectively. The spatial discriminator extracts the spatial features of the real 3D coordinate data and the 3D simulated crawling coordinate data, and the temporal discriminator extracts the temporal features of the real 3D coordinate data and the 3D simulated crawling coordinate data. The parameters of the spatial discriminator and the temporal discriminator are optimized alternately by utilizing the spatial and temporal features of real 3D coordinate data and 3D simulated crawling coordinate data. The generator parameters are then adjusted in reverse by using the spatial and temporal discriminators, ultimately enabling the spatiotemporal dual discriminator GAN to generate 3D enhanced simulated crawling coordinate data that conforms to biomechanical laws.

2. The method for predicting joint angles during infant crawling according to claim 1, characterized in that, The spatial discriminator consists of an input layer, a first convolutional layer, a ReLU layer, a fully connected layer, and an output layer connected in sequence. The first convolutional layer includes a Conv1d layer and a ReLU layer. The temporal discriminator consists of an input layer, a second convolutional layer, a ReLU layer, a fully connected layer, and an output layer connected in sequence. The second convolutional layer includes a Conv1d layer, a ReLU layer, and a MaxPooling layer.

3. The method for predicting joint angles during infant crawling according to claim 1, characterized in that, The decomposition of V_real yields the time-based fundraising matrix. H and pose coordination matrix W This includes: decoupling the V_real into a time-based recruitment matrix using principal component analysis (PCA). H and pose coordination matrix W, Among them, the time-based fundraising matrix H Describes the temporal dynamics of crawling motion, recording when other joints are activated for coordinated crawling; posture coordination matrix. W Describe the spatial coordination structure of crawling motion and record the spatial association patterns of joint movements.

4. The method for predicting joint angles during infant crawling according to claim 1, characterized in that, The joint angle data of infants crawling in future time steps are obtained by training the joint coordination and cross-attention network model JoCrCAPNet on V_mixed and C_mixed. The joint coordination and cross-attention network model JoCrCAPNet is composed of a coordination coefficient encoder and a joint angle encoder, an LSTM layer, a dual-path cross-attention module, an LTMCell layer, and a fully connected layer connected in sequence. The joint angle encoder is used to capture local joint angle features from V_mixed, and the coordination coefficient encoder is used to capture global motion features from C_mixed.

5. The method for predicting joint angles during infant crawling according to claim 4, characterized in that, The local joint angle features and global motion features are cross-fused by the dual-path cross-attention module to obtain fused features. The fused features are then processed by the LTMCell layer and the fully connected layer to predict the joint angle data of the infant's future time step during crawling. The dual-path cross-attention module consists of an input module, an attention calculation module, and a fusion and output module connected in sequence.

6. A joint angle prediction system for infant crawling, characterized in that, include: The dataset construction module is used to obtain the three-dimensional coordinate data of the joints of the limbs of infants and toddlers when they crawl, and convert the three-dimensional coordinate data into real joint angle data V_real; The process of acquiring the three-dimensional coordinate data of the limb joints of an infant crawling and converting the three-dimensional coordinate data into real joint angle data V_real specifically includes: Noise in the 3D coordinate data is eliminated, and missing data is interpolated to obtain complete crawling data. Based on anatomical calibration of joint centers and using the standard crawling posture as a reference, local coordinate systems are established at each joint center of the complete crawling data. These local coordinate systems are normalized according to the proportion of infant limb length to obtain normalized joint coordinate systems, and a relative rotation matrix of the normalized joint coordinate systems is constructed. The real joint angle data V_real is calculated using the relative rotation matrices of adjacent normalized joint coordinate systems. Three-dimensional enhanced simulated crawling coordinate data is generated using randomly sampled noise vectors. The randomly sampled noise vectors are processed by a spatiotemporal dual-discriminator GAN network to generate the three-dimensional enhanced simulated crawling coordinate data. The spatiotemporal dual-discriminator GAN network consists of a generator, a spatial discriminator, and a temporal discriminator connected sequentially. Specifically, the generator is based on randomly sampled noise vectors... The noise vector is used to generate 3D simulated crawling coordinate data. Real 3D coordinate data and 3D simulated crawling coordinate data are sequentially input into a spatial discriminator and a temporal discriminator, respectively. The spatial discriminator extracts the spatial features of the real 3D coordinate data and the 3D simulated crawling coordinate data, while the temporal discriminator extracts the temporal features. The parameters of the spatial and temporal discriminators are alternately optimized using the spatial and temporal features of the real 3D coordinate data and the 3D simulated crawling coordinate data. The generator parameters are then adjusted inversely using the spatial and temporal discriminators, ultimately enabling the spatiotemporal dual discriminator GAN to generate 3D enhanced simulated crawling coordinate data that conforms to biomechanical principles. The 3D enhanced simulated crawling coordinate data is converted into simulated joint angle data V_gen. V_real and V_gen are mixed to obtain the joint angle extended dataset V_mixed. The data processing module is used to decompose V_real to obtain the time-based fundraising matrix. H and pose coordination matrix W ;use W Projecting V_real yields the true low-dimensional cooperability coefficient sequence C_real; using W Projecting V_gen yields the simulated low-dimensional cooperative coefficient sequence C_gen. H Enhance C_gen to generate a time-enhanced simulation covariance coefficient sequence C_gen_enhanced; mix C_real with C_gen_enhanced to form an extended covariance coefficient sequence C_mixed. The model training and application module is used to construct the joint coordination and cross-attention network model JoCrCAPNet. V_mixed and C_mixed are input into the JoCrCAPNet model for training to obtain a model that predicts the joint angle data of the future time step of an infant's crawling. By using the JoCrCAPNet model to predict the crawling data of an infant, the joint angle data of the future time step of the infant's crawling can be obtained.

7. A computer device, comprising a memory, a processor, and a computer program stored in the memory, characterized in that, When the processor executes the computer program, it implements the method steps of any one of claims 1 to 5.

8. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by a processor, it implements the method steps as described in any one of claims 1 to 5.

Citation Information

Patent Citations

  • Three-dimensional human body posture prediction method and device, medium and product

    CN120299082A

  • Interactive behavior understanding method for posture reconstruction based on features of skeleton and image

    US20250022165A1