Three-dimensional human body posture estimation method and system

By combining a frequency-division Transformer spatiotemporal module and a multilayer perceptron decoder with reinforcement learning and geometric kinematic constraints, the problem of insufficient accuracy and robustness in 3D human pose estimation in existing technologies is solved, achieving high-precision and stable 3D human pose estimation, which is suitable for pedestrian collision warning in blind spots of engineering vehicles.

CN120977018AActive Publication Date: 2025-11-18TONGJI UNIV

Patent Information

Application Number
CN202511509160.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-10-22
Publication Date
2025-11-18
Estimated Expiration
2045-10-22

AI Technical Summary

Technical Problem

Existing technologies suffer from insufficient accuracy, poor robustness, and limited real-time performance when estimating 3D human pose from 2D images, especially in dynamic scenes where they struggle to meet low-latency requirements.

Method used

By employing a frequency-division Transformer spatiotemporal module and a multilayer perceptron decoder, spatial and temporal features of key points in 2D human pose are extracted through nonlinear high-dimensional mapping and geometric kinematic constraints. Combined with a reinforcement learning-driven adaptive unsupervised segmentation algorithm, the loss function is optimized to improve the accuracy and stability of 3D human pose estimation.

Benefits of technology

It significantly improves the robustness and detection accuracy of 3D human pose estimation, and can be quickly applied to pedestrian motion posture trend detection, such as the blind spot pedestrian collision warning function in engineering vehicles, and the training process is more stable.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120977018A_ABST
    Figure CN120977018A_ABST
Patent Text Reader

Abstract

The invention discloses a three-dimensional human body posture estimation method and system, and relates to the technical field of computer vision, and the method comprises the steps: extracting two-dimensional human body posture key points from a monocular video picture sequence, and generating a two-dimensional human body posture key point sequence; projecting the two-dimensional key point sequence to a feature space through nonlinear high-dimensional mapping to generate a high-dimensional feature space matrix; and inputting the high-dimensional feature matrix into a three-dimensional human body posture recognition model fusing motion constraints and frequency division spatial-temporal features to obtain a three-dimensional human body posture key point sequence, and realizing three-dimensional human body posture estimation through three-dimensional coordinates. According to the method, the robustness and the detection precision of the monocular three-dimensional human body posture estimation method are improved. Error values of relative movement speed, skeleton length and skeleton direction of the key points are calculated, so that training is easier to converge, and the training process is more stable.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of computer vision, in particular to a three-dimensional human pose estimation method and system. BACKGROUND

[0002] With the continuous evolution of video monitoring technology and artificial intelligence technology, the practical application demand of human three-dimensional pose automatic analysis and recognition technology based on video images is growing. Such demand shows important value in many fields such as road traffic safety, motion recognition, virtual reality, human-computer interaction, robot operation, specific place monitoring and dangerous place worker pose recognition. The currently widely deployed monitoring camera can only capture a single view of two-dimensional plane image. Researching how to realize human three-dimensional pose estimation method from the two-dimensional image obtained in the actual application scene has become a key problem to be solved. Compared with two-dimensional pose estimation, three-dimensional pose estimation of a single frame image is more challenging, because it needs to accurately restore the complete three-dimensional position of each joint from the blurred and noisy two-dimensional image, and different three-dimensional poses may project the same or similar two-dimensional form. The task of inferring three-dimensional pose only relying on single two-dimensional image is extremely difficult in technology.

[0003] Currently, the mainstream method in the field of human three-dimensional pose estimation can be mainly divided into single-step method and two-step method. Among them, the single-step method directly regresses the three-dimensional coordinates of each joint from the two-dimensional image through an end-to-end model without introducing intermediate representation or additional processing steps. The advantage of this kind of method is that the model architecture is relatively simple, but it is limited by the lack of intermediate constraint mechanism and the scarcity of high-quality three-dimensional pose labeling data, and it is difficult to directly infer three-dimensional space information from two-dimensional projection. In addition, this kind of method has high demand for computing resources, and needs to be optimized through fine hyperparameters to realize stable convergence.

[0004] On the other hand, the two-step method adopts a phased processing strategy: first, the two-dimensional key point detection model is used to extract the two-dimensional coordinates of the joints in the image, and then the two-dimensional to three-dimensional mapping relationship is constructed to realize three-dimensional pose reconstruction. Some methods rely on pre-constructed two-dimensional-three-dimensional pose corresponding library, and output three-dimensional results through nearest neighbor search or parameterized matching. However, this kind of method based on large-scale dictionary learning has obvious defects, which needs to consume a lot of time to construct high-dimensional mapping space in the training stage, and the inference stage needs to traverse the candidate pose library, which limits the real-time performance and is difficult to meet the low delay demand in dynamic scene. SUMMARY

[0005] In view of the above existing problems, the present application is proposed.

[0006] Therefore, the application provides a three-dimensional human posture estimation method to solve the problems of the prior art, i.e., lacking of extracting time sequence feature information, still needing to further improve the accuracy of three-dimensional human posture estimation, lacking of a method for inferring three-dimensional human posture estimation from continuous two-dimensional images, and the related model designed has weak robustness.

[0007] To solve the above technical problems, the application provides the following technical solutions. In a first aspect, the application provides a three-dimensional human posture estimation method, which comprises: extracting two-dimensional human posture key points from a monocular video picture sequence to generate a two-dimensional human posture key point sequence; projecting the two-dimensional human posture key point sequence to a feature space through nonlinear high-dimensional mapping to generate a high-dimensional feature space matrix; inputting the high-dimensional feature space matrix into a three-dimensional human posture recognition model fused with motion constraints and frequency division space-time features to obtain a three-dimensional human posture key point sequence, and realizing three-dimensional human posture estimation through three-dimensional coordinates; The obtained three-dimensional human posture key point sequence comprises: extracting spatial feature information of key points for the high-dimensional feature matrix through a frequency division space Transformer encoder to obtain a spatial feature matrix; performing reshape operation on the spatial feature matrix to convert to a time sequence dimension tensor, extracting time sequence features through a frequency division time Transformer encoder, processing through a multilayer perception decoder, and outputting to obtain the final three-dimensional human posture key point sequence; wherein the three-dimensional human posture estimation model comprises a frequency division Transformer space-time module and a multilayer perception decoder.

[0008] As a preferred scheme of the three-dimensional human posture estimation method, the nonlinear high-dimensional mapping comprises: adopting a multilayer perception MLP as a mapping function, taking a low-dimensional feature vector as input, performing linear transformation through a fully connected layer, outputting a dimension of D, and realizing through adjusting the number of neurons of the last fully connected layer. The two-dimensional human posture key point sequence is divided to obtain a matrix sequence The matrix sequence is mapped to a c-dimensional high-latitude feature space and coordinate encoded to obtain a c-dimensional high-latitude feature space matrix.

[0009] ​As a preferred scheme of the three-dimensional human pose estimation method, the spatial Transformer encoder comprises a frequency division Transformer space-time module, the frequency division Transformer space-time module comprises a frequency division spatial Transformer encoder and a frequency division time Transformer encoder, the frequency division time Transformer encoder has the same structure as the frequency division spatial Transformer encoder; the spatial features are extracted by using the frequency division spatial Transformer encoder module, the spatial features are divided into four parts according to a frequency spectrum diagram by using a reinforcement learning driven adaptive unsupervised segmentation module, self-attention Transformer feature extraction is respectively performed on the four parts, feature information in different spatial frequency domains is obtained, and then the feature information is spliced to obtain the spatial features. The frequency division time Transformer encoder extracts frequency division time features, comprising: a high-frequency time sequence branch for extracting fast motion joint point features and a low-frequency time sequence branch for modeling slow motion joint point features. The time sequence features are input into a geometric kinematics constraint optimization decoder to generate physically reliable three-dimensional human pose key points.

[0010] As a preferred scheme of the three-dimensional human pose estimation method, the three-dimensional human pose recognition model comprises a loss function, the training process is determined according to a calculation result of the loss function, and the parameters of the loss function are dynamically adjusted during the training process. During the training process, the average error is evaluated once every training cycle on a verification set, if the average error does not decrease for consecutive periods, the weights of the subtasks in the loss function are adjusted. The adjacent joint spacing and the angle penalty term are calculated by using a dynamic adjacent bone length and direction generation network, the relative speed change of the joint points is calculated to construct a relative speed loss, and the loss function further comprises a physical-geometric collaborative constraint loss, which is used to fuse the kinematics geometric constraint and the dynamics physical constraint to improve the rationality and robustness of the pose estimation. As a preferred scheme of the three-dimensional human pose estimation method, the loss function comprises a position loss function, a bone length loss function, a bone angle loss function and a relative speed loss function, the loss function is obtained by weighted summation of the position loss function, the bone length loss function, the bone angle loss function and the relative speed loss function, and the difference in the speed of the body parts in adjacent predicted frames is considered to promote the smoothness of the 3D pose sequence.

[0011] ​Stop iterating training the monocular three-dimensional human pose estimation model when the loss function tends to be stable;The frequency division time Transformer encoder includes a frequency division attention operation layer, a batch normalization layer, a Relu layer and a multi-layer perception decoder.

[0012] As a preferred scheme of the three-dimensional human pose estimation method, the processing by the multi-layer perception decoder includes extracting spatial frequency domain feature information of the key points by the improved wavelet transform, inputting the frequency domain feature information into an adaptive unsupervised segmentation algorithm driven by reinforcement learning to obtain four different frequency division sub-band matrices, inputting the sub-band matrix features into the frequency division space Transformer encoder respectively, extracting spatial feature information of different frequency domain sub-bands respectively, and obtaining the spatial feature information after splicing. Specifically, a multi-scale decomposition strategy is designed for the spatial frequency domain characteristics of the human pose key points by the improved wavelet transform algorithm, and a frequency division sub-band matrix with spatial-frequency joint representation capability is generated;The improved wavelet transform algorithm uses a direction-sensitive filter bank to enhance the frequency domain feature capture capability of dynamic poses;The adaptive unsupervised segmentation algorithm driven by reinforcement learning models the segmentation process as a Markov decision process, and the segmentation algorithm selects appropriate segmentation actions according to the frequency domain and spatial feature distribution of the data, including dividing boundaries and merging regions;A multi-objective reward function is designed to comprehensively consider the frequency domain consistency and spatial continuity of the segmentation region.

[0013] As a preferred scheme of the three-dimensional human pose estimation method, the three-dimensional human pose estimation by the three-dimensional coordinates includes reshaping the spatial feature information to convert it into a time sequence dimension tensor, and inputting it into a frequency division time Transformer encoder that fuses geometric kinematics constraints;Through a high-frequency time sequence branch, fast motion joint feature is extracted;Through a low-frequency time sequence branch, slow motion joint feature is extracted. The finally extracted spatio-temporal feature information is input into a geometric dynamics optimization decoder to output the final three-dimensional human pose key point sequence, each key point is represented by a three-dimensional coordinate, and the three-dimensional human pose estimation is realized by the three-dimensional coordinate.

[0014] In a second aspect, the application provides a three-dimensional human pose estimation system, which includes a two-dimensional human pose estimation module responsible for extracting two-dimensional human pose data from a video picture sequence, estimating 17 key point coordinates of each frame by a multi-scale feature fusion deep learning network, and generating a two-dimensional pose key point sequence; The high-dimensional feature mapping module maps the two-dimensional posture key point coordinates to a high-dimensional space to generate a high-dimensional feature matrix; this step realizes the promotion and position coding of the two-dimensional coordinates through coordinate coding and feature space mapping; The three-dimensional posture estimation module adopts a frequency division Transformer space-time module, which includes a frequency division space Transformer encoder and a frequency division time Transformer encoder; an adaptive unsupervised segmentation algorithm driven by reinforcement learning is used to perform frequency domain decoupling segmentation on the high-dimensional feature matrix and extract spatial and time sequence features; after multiple processing steps, a multi-layer perceptron decoder is used to obtain the final three-dimensional human posture key point sequence; The training optimization module optimizes the monocular three-dimensional human posture estimation model through training and validation datasets; by calculating position loss, geometric loss and relative velocity loss functions, the model parameters are adjusted to reduce the loss and improve the accuracy; the human body limb centroid velocity calculation method is used to optimize the velocity estimation when the key points are blocked, and the accuracy of motion analysis is improved.

[0015] In a third aspect, the present application provides a computer device comprising a memory and a processor, the memory storing a computer program, wherein the computer program is executed by the processor to implement any step of the three-dimensional human posture estimation method according to the first aspect of the present application.

[0016] In a fourth aspect, the present application provides a computer readable storage medium having a computer program stored thereon, wherein the computer program is executed by the processor to implement any step of the three-dimensional human posture estimation method according to the first aspect of the present application.

[0017] The beneficial effects of this invention are as follows: This invention is mainly used to estimate the 3D human pose in monocular videos, and it integrates geometric kinematic constraints and frequency-division spatiotemporal features, greatly improving the robustness and detection accuracy of monocular 3D human pose estimation methods. Therefore, it can be quickly applied to fields such as pedestrian motion posture trend detection, for example, the blind spot pedestrian collision warning function in engineering vehicles. A frequency-division spatiotemporal Transformer encoder module integrating geometric kinematic constraints is proposed. The frequency-division spatial Transformer encoder extracts key point spatial features at different motion frequencies, and then the stitched features are fed into the frequency-division temporal Transformer encoder to extract temporal information, further improving detection accuracy and, to some extent, temporal jitter resistance. This invention proposes a new network training loss function that not only focuses on the error value between the final 3D keypoints and the real keypoints, but also calculates the error values ​​for the relative motion velocity, bone length, and bone direction of the keypoints. Therefore, training is easier to converge and the training process is more stable. In existing technologies, keyframes, i.e. the 3D pose of intermediate frames, are obtained through a single-layer perceptron. However, this method can obtain multi-frame results simultaneously by setting a Transformer encoder and a fully connected convolutional layer. Attached Figure Description

[0018] To more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings used in the following description of the embodiments will be briefly introduced. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0019] Figure 1 This is a flowchart of a three-dimensional human pose estimation method.

[0020] Figure 2 This is a schematic diagram of the pose recognition process for a 3D human pose estimation method.

[0021] Figure 3 This is a schematic diagram of key points on the human body, illustrating a 3D human pose estimation method.

[0022] Figure 4 This is a schematic diagram of the training process for a 3D human pose estimation method. Detailed Implementation

[0023] To make the above-mentioned objects, features and advantages of the present invention more apparent and understandable, the specific embodiments of the present invention will be described in detail below with reference to the accompanying drawings.

[0024] In the following description, numerous specific details are set forth in order to provide a thorough understanding of the present application. However, it will be apparent to one skilled in the art that the present application can be practiced without the specific details set forth in this description. In other instances, well-known methods, procedures, components, and circuits have not been described in detail so as not to unnecessarily obscure aspects of the present application.

[0025] Secondly, the "one embodiment" or "an embodiment" referred to herein means a specific feature, structure, or characteristic under discussion. Thus, "one embodiment" does not mean a single embodiment nor is it to be taken individually or selectively from other embodiments.

[0026] Reference is made to Figures 1-4 For one embodiment of the present application, the embodiment provides a three-dimensional human pose estimation method, comprising the following steps: S1: Extracting two-dimensional human pose key points from a monocular video image sequence to generate a two-dimensional human pose key point sequence.

[0027] Further, the two-dimensional human pose estimation is performed on the dataset in the video image sequence to obtain key point information and map it to a high-dimensional space to obtain a high-dimensional feature space matrix.

[0028] Specifically, the two-dimensional human pose estimation on the dataset in the video image sequence in S1 includes the following steps: Step 1.1, preprocessing the three-dimensional human pose estimation dataset to obtain a picture sequence including two-dimensional human poses and corresponding three-dimensional human pose information.

[0029] The preprocessing process includes data cleaning, noise removal, image correction, or other necessary data preparation work, and the goal is to ensure the quality and consistency of the dataset.

[0030] Step 1.2, inputting the above obtained two-dimensional human pose picture sequence into a two-dimensional human pose detector to obtain a two-dimensional human pose key point skeleton, including a sequence of 17 two-dimensional human pose key points.

[0031] Two-dimensional human pose detection usually uses an end-to-end model architecture based on deep learning. This embodiment selects a multi-scale feature fusion network with high precision and real-time balance advantages as the backbone model. The network is based on an improved HRNet (High-Resolution Network) architecture, which maintains spatial detail information through parallel multi-resolution subnetwork interaction. HRNet usually uses a resolution of 384*288 ResNet-50 backbone network, and the main task is to accurately identify and locate human key points from images.

[0032] As shown in Figure 3 , the detection process outputs the standardized human joint skeleton topology, including the head (nose tip, eyes, ears), torso (shoulders, elbows, wrists, spine midpoint), and lower limbs (hips, knees, ankles) a total of 17 anatomical landmark points. This process achieves sub-pixel level positioning accuracy through the joint optimization of Heatmap Regression and coordinate offset prediction, and finally generates a key point sequence: , where f is the frame number of the corresponding input image, for example, the commonly used frame numbers f are 300, 243, 128, 81, etc. J is the 17 key point coordinates. The subscript 2 is the corresponding two-dimensional coordinate value .

[0033] Step 1.3, divide the sequence of two-dimensional human posture key points obtained above, the commonly used division step is 243, 128, 81, etc., and the present application adopts a step of 243 to divide the above sequence.

[0034] The two-dimensional space point coordinates are lifted to a c-dimensional feature space and are encoded to obtain a c-dimensional feature matrix , that is:

[0035] , where is a conversion matrix that converts the dimension from 2 to c, , and the coordinate encoding matrix is a learnable matrix that can encode features and increase the position information of the features.

[0036] It should be noted that in the above S1, lifting the obtained two-dimensional human posture to a high-dimensional space specifically includes dividing the obtained matrix sequence , where J is generally 17 and f is generally 243. Map to a c-dimensional feature space, which is generally a high-dimensional c, mainly used to extract more feature information in the subsequent bidirectional space-time module. Common c dimensions are 128, 256, 512, etc.

[0037] S2: Project the two-dimensional human posture key point sequence into a feature space through nonlinear high-dimensional mapping to generate a high-dimensional feature space matrix.

[0038] Further, as shown in Figure 2 , input the high-dimensional feature space matrix into the monocular three-dimensional human posture estimation model for processing. Specifically, the monocular three-dimensional human posture estimation model includes a frequency division Transformer space-time module and a multi-layer perception decoder, and the processing process is: The spatial feature information of the key points is extracted by a frequency division spatial Transformer encoder for a high-dimensional feature space matrix, to obtain a spatial feature matrix; the spatial feature matrix is reshaped to a time sequence dimension tensor, and then a frequency division time Transformer encoder is used to extract the time sequence features, which are processed by a multilayer perception decoder to output a final three-dimensional human pose key point sequence, realizing three-dimensional human pose estimation.

[0039] Specifically, the frequency division Transformer space-time module in S2 includes a frequency division spatial Transformer encoder and a frequency division time Transformer encoder, wherein the frequency division time Transformer encoder has the same structure as the frequency division spatial Transformer encoder, except that the input data is different.

[0040] In the frequency division spatial Transformer encoder, there is a reinforcement learning driven adaptive unsupervised segmentation algorithm module, a spatial attention operation layer, a batch normalization layer, a Relu layer and a multilayer perception decoder. In the spatial attention operation layer, a multi-head attention mechanism is used to calculate the attention of each head and fuse the features of each head to obtain the final multi-head attention. The batch normalization (BN) method is an effective layer-by-layer normalization method that can normalize the intermediate layer. The main advantage is that normalization makes the optimization data smoother.

[0041] In particular, in the frequency division spatial Transformer encoder, the first step needs to implement frequency domain decoupling segmentation on the input spatial feature tensor by a reinforcement learning driven adaptive unsupervised segmentation algorithm module. Specifically, by improving the wavelet transform , the input high-dimensional feature data is mapped to the frequency domain space, and a reinforcement learning driven adaptive unsupervised segmentation algorithm module is used to generate four sub-band coefficient matrices . Among them, the low-frequency component is concentrated in the upper left corner of the frequency spectrum, encodes the overall contour and joint topology of the human body, and corresponds to the overall structure information of the image; the horizontal high-frequency component corresponds to the vertical edge, reflecting the longitudinal motion trend of the limbs; the vertical high-frequency component corresponds to the horizontal edge, representing the transverse motion characteristics of the limbs; and the diagonal high-frequency is concentrated in the diagonal position of the frequency spectrum, encoding details and local deformation.

[0042] The improved wavelet transform algorithm , a multi-scale decomposition strategy is designed for the spatial frequency characteristics of human pose key points to generate sub-band matrices with spatial-frequency joint representation ability. The improved wavelet transform uses a direction-sensitive filter bank to enhance the frequency domain feature capture ability of dynamic gestures. By designing multiple direction filters (such as 0°, 45°, 90°, 135°), the feature information of a specific direction is extracted. The formula is expressed as:

[0043] where is the direction basis function, usually using Gabor wavelet, D is the action direction in human pose estimation. represents the wavelet coefficient of the signal after wavelet transform at the scale, the position.

[0044] Here, a reinforcement learning driven adaptive unsupervised segmentation algorithm module is used to divide the four sub-band coefficient matrices. The commonly used method is to directly divide the frequency spectrum into four equal size modules, with a length and width of 1 / 2 of the frequency spectrum. This segmentation method has certain defects.

[0045] From the perspective of data characteristics, the information distribution in the frequency spectrum is often uneven, and different regions may carry information of different importance and characteristics. Equal size segmentation cannot fully consider the differences in information distribution, which may lead to important information being improperly segmented into different modules, or regions with similar characteristics being artificially split, thereby destroying the internal structure and relevance of the data. Therefore, a reinforcement learning driven adaptive unsupervised segmentation algorithm is proposed. The main steps are: The algorithm models the segmentation process as a Markov decision process, and the segmentation algorithm interacts with the environment (different frequency domain sub-band data) to select appropriate segmentation actions (such as dividing boundaries, merging regions, etc.) according to the current state (data frequency and spatial feature distribution). After each action, the system will feedback a reward signal, which is evaluated according to the accuracy of the segmentation result, region consistency, etc.

[0046] The segmentation algorithm continuously tries and learns, adjusts its strategy according to the reward signal, and maximizes the long-term cumulative reward, so as to adaptively find the optimal segmentation method and extract spatial feature information of different frequency domain sub-bands. Multi-objective reward function, comprehensive consideration of frequency domain consistency, spatial continuity and semantic integrity of segmentation region.

[0047] The multi-objective reward function mainly evaluates the reward from two dimensions of frequency domain consistency and spatial consistency. Specifically: Calculate the Fourier transform of the segmented region to extract the amplitude spectrum. Calculate the frequency spectrum entropy:

[0048] wherein is the normalized spectral energy.

[0049] The frequency domain reward function is:

[0050] The boundary gradient of the segmented region is calculated: the Sobel operator is used to extract the edge. The edge continuity is calculated. The spatial reward function is:

[0051] Multi-objective reward function (comprehensive):

[0052] wherein , are the weight coefficients, which are initially set to 0.5, 0.5. In the model training stage, adaptive learning can be achieved through gradient descent. represents the continuous edge length, represents the total edge length.

[0053] The adaptive unsupervised segmentation algorithm module adjusted by the multi-objective reward function is obtained , the formula is expressed as:

[0054] In the above formula, represents the adaptive unsupervised segmentation algorithm module driven by reinforcement learning, which contains a multi-objective reward function for segmentation effect evaluation. represents the improved two-dimensional wavelet transform function. Next, a multi-head attention mechanism and a multi-layer perceptron decoder are used to process the four sub-band coefficient matrix features, and the specific processing process is as follows:

[0055]

[0056]

[0057]

[0058]

[0059]

[0060]

[0061]

[0062] where LN is Layer Normalization, MSA is Multi-Head Self-Attention, and MLP is Multi-Layer Perceptron decoder, represents the result obtained by the first linear layer of the Multi-Layer Perceptron decoder, represents the result obtained by the second linear layer of the l-th layer of the Multi-Layer Perceptron.

[0063] After processing the characteristics of the four sub-band matrices, the , , , are spliced to obtain the final result: , , ) In the above formula, is the final feature matrix obtained after splicing the sub-band matrix characteristics, and the number of channels is the sum of the channel numbers of all sub-band characteristics. Concate() is a splicing function that stacks the input N sub-band feature matrices in order along the channel dimension to generate a fused multi-scale feature representation.

[0064] In the temporal Transformer module, the feature layer is divided into two modules, i.e., high-speed and low-speed modules, by motion speed, and the speed threshold for segmentation is which is a trainable threshold parameter. The subsequent operations are the same as the spatial Transformer, which will not be described in detail here.

[0065] In the multi-head attention mechanism, the correlation features of each head are obtained by extracting the correlation features of Q, K, and V through the Attention multi-head attention mechanism Generally, 8 heads are more commonly used, wherein .

[0066]

[0067]

[0068] wherein, the function is to divide the input feature Y into three parts to obtain Q, K, and V. In addition, the attention formula is represented as:

[0069] The features obtained by each head attention mechanism are connected, and finally the feature information output by the attention mechanism is obtained.

[0070] In the embodiment, the above S2 specifically includes the following sub-steps: ​Step 2.1, the features obtained by the spatial Transformer encoder are reshaped into a spatial dimension tensor After reshaping, the tensor is converted into a time dimension tensor , which is input into the frequency division time Transformer encoder to extract time sequence features.

[0071] Step 2.2, in the frequency division time sequence features, the high-speed motion key point time sequence features and the low-speed key point time sequence features are divided, which correspond to the key point trajectory time sequence features of different motion speeds. In the forward time sequence feature extraction, the time dimension tensor is input into the time Transformer encoder. The structure of the time Transformer encoder is the same as that of the spatial Transformer encoder, except that the input data is different.

[0072] It should be noted that in the multi-head attention mechanism calculation in the Transformer, the attention of each head is calculated first, and the features of each head are fused to obtain the final multi-head attention, and then the features are sent to the regularization layer for data regularization. Then, through the multi-layer perception decoder, the obtained two parts of the space-time self-attention state are updated separately to obtain the space-time feature matrix, and are regularized.

[0073] S3: input the high-dimensional feature space matrix into the monocular three-dimensional human pose estimation model with geometric kinematic constraints to obtain a three-dimensional human pose key point sequence, and realize three-dimensional human pose estimation through three-dimensional coordinates.

[0074] Further, the process of extracting spatial correlation features through the Transformer bidirectional space-time module mainly includes reshaping the high-dimensional feature space matrix obtained by S1 into a spatial dimension tensor , i.e., mapping from 243*17*512 to 17*243*512, and then performing a Split operation on the mapped matrix, i.e., Q, K, V = Split(Y), so the feature dimensions of Q, K, and V are 17*81*512.

[0075] into the spatial Transformer encoder to extract key point global features. For example Figure 2As shown in the Transformer encoder, a spatial attention operation layer, a batch normalization layer and a Relu layer are included; in the spatial attention operation layer, a multi-head attention mechanism is adopted, the attention of each head is calculated first, and the features of each head are fused to obtain the final multi-head attention; the batch normalization (BN) method is an effective layer-by-layer normalization method, which can normalize the intermediate layer, and the main advantage is that normalization will make the optimization data smoother.

[0076] The process of extracting bidirectional time sequence correlation features by the Transformer mainly includes: reshaping the spatial information, i.e., mapping from 17*243*512 to 243*17*512, and the subsequent process is the same as that of the spatial Transformer, i.e., obtaining a feature matrix containing time sequence information .

[0077] A multi-layer perceptron decoder is adopted to input the time sequence feature information into a fully connected convolutional layer, reduce the c-dimensional features in the feature space to 3 dimensions, and obtain the final three-dimensional human pose key point sequence.

[0078] Specifically, the information on the time sequence is input into a fully connected convolutional layer to obtain the final three-dimensional human pose key point sequence, including the following specific steps: A multi-layer perceptron decoder is adopted to reduce the c-dimensional features in the feature space to 3 dimensions:

[0079] The final output three-dimensional human pose sequence result is:

[0080] This is different from most current methods and is more convenient. The present application can obtain multiple frame results at the same time, for example, when f=243, 243 frames of three-dimensional pose output results can be obtained.

[0081] The matrix feature output by the space-time module is 243*17*512, and the fully connected convolutional layer here reduces the dimension 512 to 3, i.e., the matrix output is 243*17*3, which is the final output result.

[0082] Before applying the monocular three-dimensional human pose estimation model to pose estimation, the model is trained and verified using the Human3.6M human pose estimation dataset.

[0083] The training step flow is as shown in Figure 4 . (1) Obtain a two-dimensional pose sequence.

[0084] (2) Perform data normalization operation.

[0085] (3) Initialize the monocular three-dimensional human pose estimation model of the frequency division space-time feature.

[0086] (4) Train the monocular three-dimensional human pose estimation model using the data set, and iterate the monocular three-dimensional human pose estimation model.

[0087] (5) Calculate the loss function between the result output by the model and the true value.

[0088] (6) Determine whether the loss is stable.

[0089] (7) When the loss is still decreasing and not stable, repeat step (4).

[0090] (8) When the loss has stabilized, the training is complete, and the trained monocular three-dimensional human pose estimation model is obtained.

[0091] Set the position loss function and the skeletal geometry constraint function and the relative velocity loss function to predetermined weights, and the loss function is expressed by the following formula:

[0092] wherein, and are weight coefficients, and their sum is 1; is the position loss function; is the geometry loss function; is the physical loss function; is the relative velocity loss function, which considers the difference in body part velocity in adjacent predicted frames, promoting the smoothness of the 3D pose sequence. When the loss function tends to be stable, stop iterating the training of the monocular three-dimensional human pose estimation model.

[0093] Where it needs to be introduced The relative velocity loss function. The setting of the velocity loss function will be invalid when the joint is blocked, for example, the joint of the athlete's foot is missing due to blocking when running, so the velocity loss function cannot be accurately calculated when the value is empty, which seriously affects the accuracy of the motion analysis.

[0094] To overcome this defect, the application provides a human body limb joint centroid velocity calculation method. According to the correlation of the human body limb joints, the 17 joints in the human body can be divided into 5 parts: left hand region, right hand region, left leg region, right leg region, and middle torso region. The 17 joints of the human body are divided into 5 regions (left hand, right hand, left leg, right leg, and torso), and the centroid coordinates of each region are the average values of the joint coordinates in the region, which are expressed by the formula:

[0095] wherein, respectively represent the key points in the two-dimensional coordinates and axis coordinates.

[0096] Meanwhile, by introducing the flocking flight characteristics of birds, a human body motion speed calculation model based on group behavior simulation is constructed. The scheme uses the obstacle avoidance-following mechanism in bird flocking motion to optimize the speed estimation algorithm when the human body joint is blocked, and the specific improvements are as follows: the core dynamics equation of bird flocking motion (such as the separation-alignment-aggregation model) is combined with the kinematic constraint conditions of human body, wherein the separation behavior corresponds to the minimum safety distance constraint between joint nodes, the alignment behavior corresponds to the consistency of adjacent frame motion trend, and the aggregation behavior corresponds to the maintenance of spatial connectivity of limb segments, and then the improved speed calculation formula suitable for the occlusion scene is derived.

[0097] On the basis of the traditional Cucker-Smale group motion model, the kinematic constraint of human body is fused, and the speed update formula in the occlusion scene is derived:

[0098] wherein, is a dynamic weight based on joint distance, is a visible neighborhood set of the i th joint is a limb segment kinetic potential function; , , is a behavior weight adjustment parameter (obtained by training the CMU motion database).

[0099] The traditional loss function focuses on the difference between the predicted three-dimensional key point coordinates and the real three-dimensional key point coordinates, and the relative speed loss function focuses on the difference between the relative motion speed of the real three-dimensional key point and the corresponding three-dimensional key point in the previous frame image. The relative motion information of the key point adopted by the present application can reduce the jitter of the three-dimensional human body posture between adjacent frames, so as to better converge the training process.

[0100] It should be noted that the present application improves the detection accuracy of the three-dimensional human body posture by extracting the frequency division space-time feature information; and reduces the jitter of the three-dimensional human body posture between adjacent frames by using the method of obtaining attention between adjacent frames by fusing time and space features; and the relative speed information is integrated into the loss function, so that the training process is more robust and stable.

[0101] Embodiment 2, refer toFigure 3 For an embodiment of the present application, a multi-discipline collaborative construction deployment optimization method is provided, and economic benefit calculation and simulation experiments are used to scientifically demonstrate the beneficial effects of the present application.

[0102] The method provided by the present application mainly completes training and testing on a double Nvidia RTX2080Ti GPU display card, the running environment is Pytorch1.8.1 and Python3.8, and the data set is a Human3.6M data set collected by a visual sensor. Human3.6M is one of the largest motion capture data sets, including 3.6 million 3D human poses and corresponding images. The video data contained in the data set is shot by four high-resolution progressive scan cameras at a speed of 50 Hz. The data set involves the activities of 11 professional actors in 17 scenes: discussion, smoking, taking pictures, making phone calls, etc., and provides accurate 3D joint positions and high-resolution videos.

[0103] As shown in Figure 3 , the two-dimensional key points are 17 key nodes of a two-dimensional human body extracted by using a CPN network and real GroundTruth. In the specific parameters of the experiment, N1 and N2 are the number of time sequences and spatial submodules of the Transformer in the single channel, both of which are set to 4. The continuous frame number is set to 243, which is a commonly used frame number in related research. The batch_size is set to 1024 during training, the initial learning rate is set to 0.0004, the training times epoch are 120, and the learning rate is attenuated by 99% for each training epoch.

[0104] 1) The Human3.6M data set is divided into S1, S5, S7 and S8 as the training set, and S9 and S11 as the test set.

[0105] 2) Input the video image into the trained two-dimensional human pose detector CPN to obtain a two-dimensional human skeleton key point sequence, and save it.

[0106] 3) Divide the input skeleton key point sequence into a matrix with a batch size of 1024, f=243, J=17, and coordinates of 2.

[0107] 4) The 243*17*2 matrix obtained in 3) is input into a neural network linear layer with an input dimension of 2 and an output dimension of 512 to obtain a 243*17*512 matrix, and then a position encoding matrix 17*512 is added to each frame feature.

[0108] 5) The 243*17*512 obtained in 4) is reshaped to map to a high spatial dimension, i.e., 17*243*512.

[0109] 6) The matrix obtained in 5) is input into a Transformer encoding layer with a depth of 8. Each Transformer encoding layer first divides the input matrix into the same 3 matrix features, i.e., Q, K, and V with dimensions of 17*81*512, by the Split function. Then, the multi-head attention mechanism is calculated, and in this embodiment, 8 heads are used, i.e., 8 different self-attention mechanisms are used to capture features of different scales, and each attention mechanism is scaled to prevent the activation layer from being biased due to excessive dimensions. Then, the attention mechanisms are fused to obtain a feature matrix with dimensions of 17*512. Finally, the obtained result is sent to a 2-layer perceptron layer, the first layer of which increases the 512-dimensional feature to a 4*512 layer, and then passes through a sigmoid activation layer and a dropout layer. Then, the 4*512 feature is reduced to a 512 layer, and then passes through a sigmoid activation layer and a dropout layer. In this embodiment, the features before entering the Transformer, the features obtained by the multi-head attention, and the features obtained by the multi-layer perceptron are connected in residual to prevent gradient disappearance and gradient explosion.

[0110] 7) The bidirectional spatiotemporal information obtained in step 6, with a matrix dimension of 243*17*512, is input into a fully connected layer with an input of 512 and an output of 3, i.e., an output matrix with a dimension of 243*17*3 is obtained.

[0111] To verify the effectiveness and practicability of the method of the present application, a verification example on the Human3.6M dataset is given below. Table 1 shows the detection results of the example on the test set, and the measurement standard is MPJPE (Mean Per Joint Position Error), which is the average error value of all key points. The full name and abbreviation of the 15 action categories in the dataset are defined as: Directions (Dir), Discussion (Disc), Eating (Eat), Greeting (Greet), Talking on the phone (Phone), Taking photo (Photo), Posing (Pose), Making purchases (Purch), Sitting (Sit), Sitting Down (SitD), Smoking (Smoke), Waiting (Wait), Walking dog (WalkD), Walking (Walk), and Walking together (WalkT).

[0112] Table 1 Comparison table of action effect of the present application

[0113] The present application selects the current mainstream method in recent years for comparison, as shown in Table 1, which is the result of the present application method on the Human3.6M data set using MPJPE error for evaluation.

[0114] The two-dimensional joint points compared are detected by the CPN network, the comparison result is the MPJPE error, the comparison data set is Human3.6M, and the bold part is the best estimation effect. According to the results in Table 1, 15 common actions are compared, and it can be found that the MPJPE error of the method provided by the present application is 41.6mm, which is better than other methods.

[0115] The embodiment also provides a three-dimensional human pose estimation system, comprising: The two-dimensional human pose estimation module is responsible for extracting two-dimensional human pose data from a video picture sequence, estimating 17 key point coordinates of each frame through a multi-scale feature fusion deep learning network, and generating a two-dimensional pose key point sequence.

[0116] The high-dimensional feature mapping module maps the two-dimensional pose key point coordinates to a high-dimensional space to generate a high-dimensional feature space matrix; this step realizes the promotion and position coding of the two-dimensional coordinates through coordinate coding and feature space mapping.

[0117] The three-dimensional pose estimation module adopts a frequency division Transformer space-time module, which includes a frequency division space Transformer encoder and a frequency division time Transformer encoder; an adaptive unsupervised segmentation algorithm driven by reinforcement learning is used to perform frequency domain decoupling segmentation on the high-dimensional feature matrix and extract spatial and time sequence features; after multiple processing steps, a multi-layer perceptron decoder is used to obtain the final three-dimensional human pose key point sequence.

[0118] The training optimization module optimizes the monocular three-dimensional human pose estimation model through training and verification data sets; by calculating the position loss, geometric loss and relative speed loss function, the model parameters are adjusted to reduce the loss and improve the accuracy; the human body limb centroid speed calculation method is used to optimize the speed estimation when the joint point is blocked, and the accuracy of motion analysis is improved.

[0119] The embodiment also provides a computer device suitable for the case of the three-dimensional human pose estimation method, comprising: a memory and a processor; the memory is used to store computer executable instructions, and the processor is used to execute the computer executable instructions to realize the three-dimensional human pose estimation method proposed in the above embodiment.

[0120] The computer device can be a terminal, which includes a processor, a memory, a communication interface, a display screen and an input device connected by a system bus. The processor of the computer device is configured to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for running the operating system and the computer program in the non-volatile storage medium. The communication interface of the computer device is configured to perform wired or wireless communication with an external terminal. The wireless communication can be achieved by WIFI, a carrier network, NFC (Near Field Communication) or other technologies. The display screen of the computer device can be a liquid crystal display screen or an electronic ink display screen. The input device of the computer device can be a touch layer overlaid on the display screen, or a key, a trackball or a touchpad arranged on the shell of the computer device, or an external keyboard, a touchpad or a mouse, etc.

[0121] The embodiment also provides a storage medium having a computer program stored thereon, the program being executed by a processor to implement the three-dimensional human pose estimation method proposed in the above embodiment. The storage medium can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as a static random access memory (SRAM), an electrically erasable programmable read-only memory (EEPROM), an erasable programmable read-only memory (EPROM), a programmable read-only memory (PROM), a read-only memory (ROM), a magnetic memory, a flash memory, a magnetic disk or an optical disk.

[0122] In summary, the present application extracts two-dimensional human posture key points from a monocular video image sequence to generate a two-dimensional human posture key point sequence, projects the two-dimensional human posture key point sequence to a feature space through nonlinear high-dimensional mapping to generate a high-dimensional feature space matrix, inputs the high-dimensional feature space matrix into a monocular three-dimensional human posture estimation model with geometric kinematics constraints to obtain a three-dimensional human posture key point sequence, and realizes three-dimensional human posture estimation through three-dimensional coordinates. The three-dimensional human posture key point sequence is obtained by improving wavelet transform to extract the frequency domain features of the image, and then generating four sub-band coefficient matrices through a reinforcement learning driven adaptive unsupervised segmentation algorithm module. The frequency band sub-matrix is input into a spatial Transformer encoder to extract spatial features in the frequency domain and combine them. The extracted spatial feature information is dimensionally reshaped into a time series dimension tensor. The high-frequency and low-frequency time series branches are used to extract the key point features of fast and slow motion, respectively. The extracted spatiotemporal features are input into a decoder, and the three-dimensional human posture key point sequence is output after optimization.

[0123] It should be noted that the above embodiments are only used to illustrate the technical solutions of the present application and not to limit it. Although the present application has been described in detail with reference to the preferred embodiments, those skilled in the art should understand that the technical solutions of the present application can be modified or replaced equivalently without departing from the spirit and scope of the present application, and they should be covered in the scope of the claims of the present application.

Claims

1. A three-dimensional human pose estimation method, characterized in that: This includes extracting two-dimensional human pose key points from monocular video image sequences and generating two-dimensional human pose key point sequences. The two-dimensional human posture key point sequence is projected onto the feature space through a nonlinear high-dimensional mapping to generate a high-dimensional feature space matrix. The high-dimensional feature space matrix is ​​input into a three-dimensional human posture recognition model that integrates motion constraints and frequency-division spatiotemporal features to obtain a three-dimensional human posture key point sequence, and three-dimensional human posture estimation is achieved through three-dimensional coordinates. The process of obtaining the 3D human pose keypoint sequence includes: extracting spatial feature information of keypoints from a high-dimensional feature matrix using a frequency-division spatial Transformer encoder to obtain a spatial feature matrix; reshaping the spatial feature matrix to convert it into a temporal tensor; extracting temporal features using a frequency-division temporal Transformer encoder; processing the temporal features using a multilayer perceptron decoder; and outputting the final 3D human pose keypoint sequence. The 3D human pose estimation model includes a frequency-division Transformer spatiotemporal module and a multilayer perceptron decoder.

2. The three-dimensional human pose estimation method as described in claim 1, characterized in that: The nonlinear high-dimensional mapping includes using a multilayer perceptron (MLP) as the mapping function, taking a low-dimensional feature vector as input, performing a linear transformation through a fully connected layer, and outputting a dimension of D, which is achieved by adjusting the number of neurons in the last fully connected layer. The two-dimensional human pose key point sequence is divided into a matrix sequence. ,Will Mapping to a c-dimensional high-dimensional feature space and performing coordinate encoding yields a c-dimensional high-dimensional feature space matrix.

3. The three-dimensional human pose estimation method as described in claim 2, characterized in that: The spatial Transformer encoder includes a frequency-division Transformer spatiotemporal module comprising a frequency-division spatial Transformer encoder and a frequency-division time Transformer encoder, wherein the frequency-division time Transformer encoder has the same structure as the frequency-division spatial Transformer encoder; the frequency-division spatial Transformer encoder module is used to extract spatial features, and a reinforcement learning-driven adaptive unsupervised segmentation module is applied to divide the spatial features into four parts according to the spectrogram, and self-attention Transformer feature extraction is performed on each part to obtain feature information in different spatial frequency domains, which are then concatenated to obtain the spatial features; The frequency division time Transformer encoder extracts frequency division time features, including: The high-frequency temporal branch is used to extract features of fast-moving joints; the low-frequency temporal branch is used to model features of slow-moving joints. Temporal features are input into a decoder optimized by geometric kinematic constraints to generate physically reliable 3D human pose key points.

4. The three-dimensional human pose estimation method as described in claim 3, characterized in that: The three-dimensional human pose recognition model includes constructing a loss function, judging the training progress based on the calculation result of the loss function, and dynamically adjusting the parameters of the loss function during the training process. During the training process, each The average error is evaluated once on the validation set during each training epoch. If the error is continuous... If the loss does not decrease within a certain number of cycles, the weights of the sub-tasks in the loss function will be adjusted. The distance between adjacent joints and the angle penalty term are calculated by generating a network that dynamically considers the length and direction of adjacent bones; the relative velocity loss is constructed by calculating the relative velocity change of joints. The constructed loss function also includes: physical-geometric co-constraint loss, which improves the rationality and robustness of attitude estimation by integrating kinematic geometric constraints and dynamic physical constraints.

5. The three-dimensional human pose estimation method as described in claim 4, characterized in that: The loss function is obtained by weighted summation of position loss function, bone length loss function, bone angle loss function, and relative velocity loss function; it considers the difference in velocity of body parts in adjacent prediction frames to promote the smoothness of 3D pose sequences; When the loss function tends to stabilize, the iterative training of the monocular 3D human pose estimation model is stopped; the frequency division time Transformer encoder includes a frequency division attention operation layer, a batch normalization layer, a ReLU layer and a multilayer perceptron decoder.

6. The three-dimensional human pose estimation method as described in claim 5, characterized in that: The processing via the multilayer perceptron decoder includes: extracting the spatial frequency domain feature information of key points through an improved wavelet transform; inputting the frequency domain feature information into an adaptive unsupervised segmentation algorithm driven by reinforcement learning to obtain four different frequency sub-band matrices; inputting the sub-band matrix features into the frequency-division spatial Transformer encoder to extract the spatial feature information of different frequency sub-bands; and then concatenating the sub-bands to obtain the spatial feature information. Specifically, by using an improved wavelet transform algorithm, a multi-scale decomposition strategy is designed for the spatial-frequency domain characteristics of key points in human posture, generating a frequency-divided sub-band matrix with joint spatial-frequency domain representation capabilities; wherein, the improved wavelet transform algorithm employs a direction-sensitive filter bank to enhance the ability to capture the frequency domain features of dynamic posture. The reinforcement learning-driven adaptive unsupervised segmentation algorithm models the segmentation process as a Markov decision process. The segmentation algorithm selects appropriate segmentation actions based on the frequency domain and spatial feature distribution of the data, including dividing boundaries and merging regions. It designs a multi-objective reward function that comprehensively considers the frequency domain consistency and spatial continuity of the segmented regions.

7. The three-dimensional human pose estimation method as described in claim 6, characterized in that: The method of achieving 3D human pose estimation through 3D coordinates includes: reshaping the spatial feature information into a temporal dimensional tensor, inputting it into a frequency-division time Transformer encoder that integrates geometric kinematic constraints; and extracting fast motion joint features through high-frequency temporal branches. Extract features of slow-moving joints by using low-frequency temporal branches; The extracted spatiotemporal feature information is input into the geometric dynamics optimization decoder, and the final 3D human pose key point sequence is output. The 3D coordinates of each key point are represented by the 3D human pose key point sequence, and 3D human pose estimation is achieved through the 3D coordinates.

8. A three-dimensional human pose estimation system, based on the three-dimensional human pose estimation method according to any one of claims 1 to 7, characterized in that: It includes a 2D human pose estimation module, which is responsible for extracting 2D human pose data from video image sequences, estimating the coordinates of 17 key points in each frame through a deep learning network that fuses multi-scale features, and generating a 2D pose key point sequence. The high-dimensional feature mapping module maps the coordinates of two-dimensional pose key points to a high-dimensional space, generating a high-dimensional feature matrix; through coordinate encoding and feature space mapping, it realizes the enhancement and position encoding of two-dimensional coordinates; The 3D pose estimation module adopts a frequency-division Transformer spatiotemporal module, which includes a frequency-division spatial Transformer encoder and a frequency-division temporal Transformer encoder. A reinforcement learning-driven adaptive unsupervised segmentation algorithm is used to perform frequency domain decoupling segmentation on a high-dimensional feature matrix and extract spatial and temporal features. After multiple processing steps, the final 3D human pose keypoint sequence is obtained using a multilayer perceptron decoder. The training optimization module optimizes the monocular 3D human pose estimation model by training and validating datasets; it adjusts model parameters to reduce losses and improve accuracy by calculating position loss, geometric loss, and relative velocity loss functions; it optimizes velocity estimation when joint occlusion by using a human limb center of mass velocity calculation method; and it improves the accuracy of motion analysis by simulating bird flock cooperative behavior to optimize velocity estimation when joint occlusion.

9. A computer device comprising a memory and a processor, wherein the memory stores a computer program, characterized in that: When the processor executes the computer program, it implements the steps of the three-dimensional human pose estimation method according to any one of claims 1 to 7.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by the processor, it implements the steps of the three-dimensional human pose estimation method according to any one of claims 1 to 7.

Citation Information

Patent Citations

  • End-to-end multi-view three-dimensional human body posture estimation method and system and storage medium

    CN112560757A

  • Frequency and space mixed domain multi-modal fusion three-dimensional target detection framework building method

    CN118506352A

  • Three-dimensional human body posture estimation method based on bidirectional spatial-temporal characteristics, program product and electronic equipment

    CN118522071A

  • Office building energy consumption prediction method and system

    CN120217609A

Cited By

  • Robot motion training method and system based on human motion video

    CN121267938A

  • Intelligent warehouse abnormal behavior identification method and system based on reinforcement learning

    CN121686564A

  • Human shape related data visualization method and system based on artificial intelligence

    CN121746593A

  • Self-adaptive time modeling driven motion scene human body posture estimation method and medium

    CN122067319A