A lightweight three-dimensional human pose estimation method for uncalibrated multi-view cameras

Through the lightweight three-dimensional human posture estimation method of uncalibrated multi-view cameras, the dual-stream space-time converter and procrustes method are used to solve the inaccurate estimation problem of camera calibration and occlusion environment in multi-view attitude estimation, and high-precision and low-cost three-dimensional human posture estimation is achieved, which is suitable for various application scenarios.

CN119832644BActive Publication Date: 2025-06-10HANGZHOU DIANZI UNIV

Patent Information

Application Number
CN202510300662.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-03-14
Publication Date
2025-06-10
Estimated Expiration
2045-03-14

AI Technical Summary

Technical Problem

The existing multi-view pose estimation method requires calibration of the camera, which is costly and difficult to apply to various fields, and the estimation is inaccurate in an occluded environment, high training cost and poor generalization ability.

Method used

A lightweight three-dimensional human posture estimation method for uncalibrated multi-view cameras is proposed. The two-dimensional pose sequence of multi-view angles is converted into three-dimensional poses through a dual-stream space-time transformer, and the procrustes method is used to align poses at different perspectives to avoid camera calibration, and the generalization ability of the model is improved through a dual-stage training framework.

Benefits of technology

It realizes the reliable three-dimensional human posture estimation results without camera calibration in multi-view pose estimation. It is suitable for videos taken by mobile cameras, reducing training costs and hardware requirements, and improving the generalization ability and estimation accuracy of the model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119832644B_ABST
    Figure CN119832644B_ABST
Patent Text Reader

Abstract

The present invention discloses a lightweight three-dimensional human body posture estimation method for an uncalibrated multi-view camera. The method first collects a three-dimensional human body posture estimation data set of any number of viewpoints, and defines a two-dimensional posture sequence and a corresponding three-dimensional posture sequence from a dynamic viewpoint. Secondly, a bidirectional Mamba model is used to extract the features of the multi-view two-dimensional posture sequence, obtain a prediction result, and pre-train and fine-tune the bidirectional Mamba module. Finally, a multi-view data set shot by an uncalibrated movable camera is used to train the bidirectional Mamba model end-to-end, and the three-dimensional posture of each viewpoint is output. The postures of different viewpoints are aligned to one of the viewpoints through procrustes to complete the human body posture estimation. The present invention adopts multi-view posture estimation, so that the human body posture can still be estimated with high precision in an environment with more occlusions, and the camera calibration process is avoided.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of computer vision, and particularly relates to a lightweight three-dimensional human pose estimation method for uncalibrated multi-view cameras. It mainly involves using a two-stream spatio-temporal transformer to obtain three-dimensional poses from multiple views of two-dimensional pose sequences, and then aligning the poses from different views to the same coordinate system through the Procrustes method to calculate errors and optimize the neural network, so as to obtain a method for higher-precision three-dimensional human poses. Background Art

[0002] Pose estimation refers to accurately inferring the pose information of the human body from images, and estimating the position coordinates of human key points, including joint angles, positions, and movement trajectories, etc. It is applied to a wide range of fields, such as in the field of sports medicine, for sports that are prone to injury like alpine skiing and gymnastics. Through pose estimation, these sports can be digitized, the movement techniques of athletes can be improved, and the probability of injury can be reduced; at the same time, in the field of autonomous driving, the downstream tasks of pose estimation can be used to predict the possible behaviors of pedestrians, improving the obstacle avoidance ability of autonomous driving.

[0003] Three-dimensional pose estimation can be divided into single-view pose estimation and multi-view pose estimation according to the number of views used during model training. Some single-view pose estimation methods directly obtain the three-dimensional human pose from images through neural networks, and some other methods use the two-dimensional pose sequence obtained from two-dimensional pose estimation as the model input to obtain the three-dimensional pose. Traditional multi-view pose estimation often needs to obtain the two-dimensional pose sequence from the upstream task, or obtain the two-dimensional joint heat map from the image through the backbone convolutional neural network, and then jointly infer the three-dimensional pose through various types of neural networks.

[0004] However, in actual applications, existing methods often encounter various challenges:

[0005] (1) Single-view pose estimation often has large errors in the estimated pose when the person is largely occluded.

[0006] (2) Traditional multi-view pose estimation requires camera calibration, and the cost of camera calibration is very high, making it difficult to apply these methods to various fields. Also, because the datasets containing camera calibration data are few, only a small amount of data can be used for training during model training, making it difficult for the model to achieve ideal effects in specific fields.

[0007] (3) The production of pose estimation datasets is difficult, and professional personnel and equipment are required to collect the three-dimensional position data of human joints;

[0008] (4) The requirements of the training network for time and hardware are too high. The number of parameters to be adjusted often reaches the order of millions. If there is a dataset in a new target application scenario, the entire network has to be retrained, resulting in too high costs.

[0009] (5) The neural network has poor transfer and generalization ability. It often performs well on the dataset but not well in an environment with a certain gap from the dataset. Summary of the Invention

[0010] The technical problem to be solved by the present invention is to propose a lightweight three-dimensional human pose estimation method for uncalibrated multi-view cameras to solve the problem of camera calibration required in multi-view pose estimation. The present invention can obtain reliable estimation results in multi-view pose estimation, avoid camera calibration, and can also reliably estimate human poses for videos shot by moving cameras, effectively expanding the scope of use of this method and better serving downstream tasks of pose estimation.

[0011] In order to better improve the generalization ability of the model in different application environments, the training process of the model in the present invention is divided into two stages. The first stage is the pre-training stage, aiming to learn useful motion representations. The second stage can select a similar dataset for training according to different application environments, thereby improving the generalization ability of the model in various situations.

[0012] The object of the present invention is achieved through the following technical solutions: A lightweight three-dimensional human pose estimation method for uncalibrated multi-view cameras, the method comprising the following steps:

[0013] In step S1, a three-dimensional human pose estimation dataset with any number of viewpoints is collected, and there are no special requirements for the cameras used, which is applicable to uncalibrated movable cameras. Define the two-dimensional pose sequences from dynamic viewpoints and the corresponding three-dimensional pose sequences . Where is the number of viewpoints, is the length of the pose sequence, is the number of human skeleton joints defined by the model, is the number of input channels, is the number of output channels.

[0014] In step S2, a two-way Mamba model is used to extract the features of the multi-view two-dimensional pose sequences. The two-way Mamba model mainly includes a preprocessing module and multiple groups of cascaded two-way Mamba modules. The preprocessing module can map the input two-dimensional pose sequences into a high-dimensional latent space , where is the dimension number of the latent space. Then In multiple cascaded bidirectional Mamba modules, each bidirectional Mamba module contains two branches. For one branch, the input data first passes through the temporal Mamba module and then through the spatial Mamba module. For the other branch, it first passes through the spatial Mamba module and then through the temporal Mamba module. The two branches are fused by directly adding vectors. After passing through multiple bidirectional Mamba modules with the same model parameter structure, the obtained prediction results are compared with the in the dataset to calculate the loss, and the parameters of the bidirectional Mamba model are optimized through backpropagation.

[0015] In step S3, during the pre-training phase of the bidirectional Mamba model, the orthogonal projection method is used to map the three-dimensional poses in the dataset to a complete two-dimensional skeleton pose sequence, and then partial missing two-dimensional poses are obtained through random masking to achieve the purpose of data augmentation. In this stage, the three-dimensional key point loss, velocity loss, and two-dimensional reprojection loss need to be calculated, as follows:

[0016] ;

[0017] where is the three-dimensional key point loss, is the velocity loss, is the two-dimensional reprojection loss; and represent the predicted value and the actual value of the three-dimensional true pose; and represent the predicted value and the actual value of the velocity of the key points between video frames, respectively, that is: ; and represent the predicted value and the actual value of the two-dimensional reprojection pose, respectively, is the confidence of the two-dimensional pose provided by the dataset. The subscript represents each video frame, represents each key point.

[0018] In step S4, during the fine-tuning phase of the bidirectional Mamba model, a simple task head is added to the pre-trained model to estimate the three-dimensional pose. The two-dimensional pose sequences from multiple perspectives in the same scene are stacked and input into the bidirectional Mamba model, and then a part of the key points and perspectives are masked through random masking to simulate the challenges of occlusion in the real environment and achieve the effect of data augmentation. The random masking is implemented by adding a control gate in the input channels of each perspective to mask the joint points or perspectives, as follows:

[0019]

[0020]

[0021] where is the motion feature viewpoint output.

[0022] When the viewpoint is valid, otherwise it is invalid. is the probability of having valid views:

[0023]

[0024] In step S5, the entire two-way Mamba model is trained end-to-end using a multi-view dataset captured by an uncalibrated movable camera. The model outputs the 3D poses of each view. The poses of different views are aligned to one view by translation, rotation, and scaling through the procrustes method, achieving the results of pose alignment and unified coordinate system. In addition to calculating the 3D keypoint loss, velocity loss, and 2D reprojection loss in this stage, the multi-view loss also needs to be calculated, specifically as follows:

[0025] First, predict the 3D pose output for each view through the model, denoted as: , where: ;

[0026] Then use the first view as the reference view. If there are other views, i.e., , the 3D pose estimated from other views can be aligned with the pose estimated from the first view through the procrustes method, denoted as the alignment of the 3D joint points of the two views:

[0027]

[0028]

[0029] The procrustes alignment calculates the scale , rotation and translation between two sets of corresponding 3D point relationships. The pose after procrustes alignment is denoted as:

[0030]

[0031] Finally, use the procrustes alignment to calculate the multi-view consistency loss:

[0032]

[0033] Among them is a binary indicator function used to determine whether to include the multi-view consistency loss represents the weighted multi-view consistency loss 、 represents the weight of the corresponding loss

[0034] The advantages of the present invention compared with the prior art are as follows

[0035] Aiming at the problem that single-view pose estimation is incorrect in an occluded environment, the present invention adopts multi-view pose estimation, enabling high-precision human pose estimation even in an environment with a lot of occlusion

[0036] Aiming at the problem that traditional multi-view pose estimation requires camera calibration, the present invention uses the Procrustes method to unify the coordinate systems of poses from different views through translation, rotation, and scaling operations, avoiding the process of camera calibration

[0037] Aiming at the problems of high training cost and poor generalization ability of traditional pose estimation, the present invention proposes a two-stage training framework. It learns the motion representation of the human body through pre-training and then adapts to various tasks through fine-tuning. When there is a new usage environment, it only needs to be fine-tuned with the dataset of this usage environment on the basis of the original pre-trained model, greatly reducing the training time and hardware requirements and improving the generalization ability of the model BRIEF DESCRIPTION OF THE DRAWINGS

[0038] Figure 1 is the overall flowchart of the present invention

[0039] Figure 2 is the architecture diagram of the pose improvement model of the present invention

[0040] Figure 3 is the architecture diagram of the bidirectional time / space Mamba module adopted by the present invention

[0041] Figure 4 is the visualization result diagram of the model trained according to the present invention

[0042] Figure 5 is the index comparison diagram of the model trained according to the present invention and other models DETAILED DESCRIPTION OF THE INVENTION

[0043] To enable those skilled in the art to better understand the technical solutions of the present disclosure, the present disclosure will be further described in detail below with reference to the drawings and specific embodiments

[0044] Such as Figure 1 and Figure 2As shown in the figure, a multi-view 3D human pose estimation method based on an uncalibrated moving camera includes the following steps:

[0045] S1. Collect 3D human pose estimation datasets with any number of viewpoints. There are no special requirements for the cameras used, and it is applicable to uncalibrated movable cameras. Define the 2D pose sequences from dynamic viewpoints and the corresponding 3D pose sequences . Where is the number of viewpoints, is the length of the pose sequence, is the number of human skeleton joints defined by the model,

[0046] is the number of input channels and the number of output channels. Figure 3 S2. Use the bidirectional Mamba model to extract the features of the multi-view 2D pose sequences. This model mainly consists of a preprocessing module and multiple groups of cascaded bidirectional Mamba modules, as shown. In the preprocessing module, the 2D pose sequence is mapped to a high-dimensional latent space through a fully connected layer , where is the dimension number of the latent space. Then, the time step is converted into an embedding vector , whose function is to encode the time information into a vector with the same dimension as the latent features, which can help the network generate a feature representation that conforms to the time dynamics at each time point. The embedding vector will be input into the MLP layer together with other features to further enhance the features and finally generate the motion feature

[0047] . is then input into multiple groups of cascaded bidirectional Mamba modules. Each Mamba module contains two branches. For one branch, the input data first passes through the temporal Mamba module and then through the spatial Mamba module. For the other branch, it first passes through the spatial Mamba module and then through the temporal Mamba module. The two branches are fused by directly adding the vectors. After passing through multiple groups of bidirectional Mamba modules with the same model parameter structure, the obtained prediction result is calculated with in the dataset to obtain the loss, and the model parameters are optimized through backpropagation.

[0048] The trained model adopts a 3D human pose estimation method based on the bidirectional Mamba module, and the main content includes:

[0049] Temporal Mamba Module: It is used to capture the temporal dependencies in the motion sequence and adopts a bidirectional strategy to enhance the temporal modeling ability. The input feature sequence is guided by temporal step embedding for denoising and then processed through normalization, linear projection, one-dimensional convolution, and a state space model to generate the final output sequence in the temporal dimension. Specifically as follows:

[0050] (1) Apply a normalization operation to the input feature sequence so that the model can avoid training instability caused by differences in feature ranges and accelerate the training process at the same time.

[0051] (2) Map the normalized feature sequence to two new feature spaces respectively through two fully connected layers, denoted as and respectively, where is the higher feature dimension after mapping, which can help the model capture more feature information and thus better learn the patterns of the time series.

[0052] (3) Perform one-dimensional convolution on the mapped sequence and apply the SiLU activation function to obtain . One-dimensional convolution can capture the local temporal dependencies in the sequence, and the SiLU activation function helps to avoid the vanishing gradient problem and provides non-linear modeling ability at the same time.

[0053] (4) Perform two linear transformations on the convolved features respectively to obtain new features and . Through further linear transformation, the model can map the convolved features to a new representation space and enhance the model's understanding of features in different dimensions.

[0054] (5) Then pass through a fully connected layer and use the log-SoftPlus function to activate the output of the linear transformation to ensure that the intermediate parameter of the output time scale is a positive number. The calculation formula is as follows:

[0055]

[0056] where is the fully connected layer,[[]]ID=48]] are learnable parameters, and the log-SoftPlus function is a smoother ReLU-like activation function.

[0057] (6) For and , and two groups of parameters respectively for the feature dimension ​Perform summation to obtain the extended features and , where are learnable parameters. This transformation process enhances the representational ability of the features, especially for the cross-time-step dependencies in time series.

[0058] (7) Use the transformed features , , and as the Mamba module parameters, and the feature as the Mamba module input to obtain the new feature . The role of the Mamba module is to further enhance the model's understanding of motion data by capturing the dependencies between space and time. Add the feature element-wise with , apply a fully connected layer and the SiLU activation function to obtain the output of the temporal Mamba module.

[0059] Spatial Mamba module: Used to learn the spatial dependencies of human postures within a single frame, adopting the same bidirectional strategy and structure as the temporal Mamba module. Its processed features are unfolded along the latent space dimension, and the output in the spatial dimension is generated through linear projection, convolution, and state space modeling, while ensuring the consistency of the output and input in shape. The specific steps are as follows:

[0060] Given the output of the temporal Mamba module, first transpose it along the last two dimensions to obtain . Use as the input of the spatial Mamba module. The subsequent calculation steps are similar to those of the temporal Mamba module, and finally obtain .

[0061] In a bidirectional Mamba module, the final output comes from the outputs of the spatial Mamba module at the ends of the two branches and the output of the temporal Mamba module. The output

[0062] (1) First, transform into , then concatenate it with to obtain .

[0063] (2) Then, perform attention calculation on using a linear transformation and perform a softmax operation to obtain the normalized attention weights , and then is split into and according to the last dimension.

[0064] (3) Finally, use the attention weights to perform weighted fusion on the input and to obtain :

[0065]

[0066] where represents element-wise multiplication.

[0067] The outputs of N bidirectional Mamba modules pass through a fully connected layer to obtain the prediction result , which is used for the calculation of the subsequent loss.

[0068] S3. Use the method of orthogonal projection to map the prediction result to the complete two-dimensional skeleton pose sequence, and then obtain the partially missing two-dimensional pose through random occlusion , achieving the purpose of data augmentation. At this stage, it is necessary to calculate the 3D key point loss, velocity loss, and 2D reprojection loss, as follows:

[0069] where is the 3D key point loss, is the velocity loss, is the 2D reprojection loss; and represent the predicted value and the actual value of the 3D true pose; and represent the predicted value and the actual value of the velocity of the key points between video frames, respectively, that is: ; and represent the predicted value and the actual value of the 2D reprojection pose, respectively. The in the subscript represents each video frame, represents each key point; is the confidence of the 2D pose provided by the dataset.

[0070] S4. Add a simple task head to the pre-trained model obtained in S2 for estimating the 3D pose. Stack the 2D pose sequences of multiple views in the same scene and input them into the bidirectional Mamba model. Occlude a part of the key points and views through random occlusion to simulate the challenges of occlusion in the real environment and achieve the effect of data augmentation. The random occlusion is implemented as: Add a control gate , to occlude the joint points or viewpoints, specifically as follows

[0071]

[0072]

[0073] where is the motion feature, is the viewpoint output.

[0074] When , the viewpoint is valid, otherwise it is invalid. is the probability of having

[0075]

[0076] S5. Train the entire network end-to-end using a multi-viewpoint dataset captured by an uncalibrated movable camera. The model outputs the 3D poses of each viewpoint. Align the poses of different viewpoints to one of the viewpoints by translation, rotation, and scaling through the procrustes method to achieve pose alignment and unified coordinate system. In addition to calculating the 3D key point loss, velocity loss, and 2D reprojection loss in this stage, it is also necessary to calculate the multi-viewpoint loss, specifically as follows:

[0077] First, predict the 3D pose output by the model for each viewpoint, denoted as: , where:

[0078] Then, use the first viewpoint as the reference view. If there are other viewpoints, i.e., , align the 3D pose estimated from other viewpoints with the pose estimated from the first viewpoint through the procrustes method, denoted as the alignment of the 3D joint points of the two viewpoints.

[0079]

[0080]

[0081] The procrustes alignment calculates the scale , rotation , and translation between two sets of 3D point correspondence relationships. The pose after procrustes alignment is denoted as:

[0082]

[0083] Finally, the Procrustes alignment is used to calculate the multi-view consistency loss:

[0084]

[0085] The total loss at this stage is:

[0086] where is a binary indicator function used to determine whether to include the multi-view consistency loss, represents the weighted multi-view consistency loss, 、 represent the weights of the corresponding losses.

[0087] The above content is a further detailed description of the present invention in combination with specific / preferred embodiments. It cannot be determined that the specific implementation of the present invention is only limited to these descriptions. For those of ordinary skill in the technical field to which the present invention belongs, without departing from the concept of the present invention, they can also make several substitutions or modifications to these described embodiments, and these substitution or modification methods should all be regarded as belonging to the protection scope of the present invention.

[0088] The parts not detailed in the present invention belong to the well-known technology in the art.

[0089] Example:

[0090] During the experiment, the Human3.6M and AMASS human pose datasets were selected as the pre-training datasets, and the Human3.6M and SkiPose datasets were used as the fine-tuning datasets. The following are the specific experimental steps and parameter settings:

[0091] Before model training, the dataset needs to be standardized. The specific operation is to set the center point of the human pelvis as the coordinate origin to generate the standardized pose data that can be directly used by the model.

[0092] In the pre-training stage, the following main parameters were set:

[0093] Number of batches: 128; Learning rate: 0.001; Learning rate decay: 0.99; Number of video frames: 243 frames; Number of key points: 17; Two 3090 graphics cards were used in the experiment, equipped with the proposed bidirectional Mamba network architecture. Compared with similar Transformer models, Mamba significantly reduces the training cost. Taking the training time of each epoch as an example, the Mamba network only needs 40 minutes, while similar Transformer models require 30 hours.

[0094] The pre-training results of the Mamba model are as Figure 4As shown, through visualization, it can be seen that the model has high accuracy in capturing pose sequences.

[0095] In the fine-tuning stage, the initial learning rate was adjusted to 0.0001, and the model was fine-tuned on the Human3.6M and SkiPose datasets respectively. Through comparative experiments with the MotionBERT model, as Figure 5 shown, it can be observed that on the SkiPose dataset, the coincidence degree between the prediction results of the Mamba model and the ground truth is significantly better than that of MotionBERT. The final metric comparisons are shown in Table 1, further verifying the superiority of the Mamba model.

[0096] Explanation of metrics: To quantify the model performance, two metrics were adopted: the mean per-joint position error (MPJPE) and the mean per-joint position error after Procrustes alignment (P-MPJPE):

[0097] MPJPE: It represents the average Euclidean distance between the joint positions predicted by the model and the ground truth positions. The smaller the value, the higher the prediction accuracy of the model for poses.

[0098] P-MPJPE: Based on MPJPE, further Procrustes alignment is performed to remove the scale, rotation, and translation errors of the prediction results, so as to better measure the performance of the model in shape matching.

[0099] Table 1 Metric Comparisons

[0100]

[0101] The experimental results show that the Mamba model has achieved the best performance in terms of both MPJPE and P-MPJPE metrics on the SkiPose dataset.

Claims

1. A lightweight 3D human pose estimation method for uncalibrated multi-view cameras, characterized in that: The following steps are involved: Step S1, collecting a 3D human body pose estimation dataset of any number of viewing angles, and defining a 2D pose sequence and a corresponding 3D pose sequence from N dynamic viewing angles; Step S2, using the bidirectional Mamba model, extracting features of the multi-view two-dimensional posture sequence to obtain a prediction result; Step S3, pre-training and fine-tuning the bidirectional Mamba module; Step S4, using the multi-view dataset shot by an uncalibrated mobile camera to train the bidirectional Mamba model end-to-end, output the three-dimensional posture of each view, align the postures of different views to one of the views through procrustes, and complete the human body posture estimation; the specific implementation process is as follows: The bidirectional Mamba model is trained end-to-end using a multi-view dataset shot by an uncalibrated mobile camera. The model outputs the 3D posture of each view. The postures of different views are aligned to one of them by translation, rotation and scaling through procrustes, so that the postures are aligned and the coordinate system is unified. In addition to calculating the 3D key point loss, speed loss and 2D reprojection loss, this stage also calculates the multi-view loss, as follows: First, the model predicts the 3D posture of each view output, which is recorded as: C out = 3, where N is the number of viewpoints, T is the length of the pose sequence, J is the number of human skeleton joints defined by the model, and C in is the number of input channels, C out is the number of output channels; Then use the first view as the reference view. If there are other views, that is, N ≥ 2, the 3D pose estimated from other views is aligned with the pose estimated from the first view through the procrustes method; The posture after procrustes alignment is recorded as: Finally, procrustes alignment is used to calculate the multi-view consistency loss: The total loss at this stage is: Where II is a binary indicator function that determines whether to include multi-view consistency loss, and λ MV L MV represents the weighted multi-view consistency loss, λ 3D , O Represents the weight of the corresponding loss.

2. The lightweight 3D human pose estimation method for an uncalibrated multi-view camera according to claim 1, characterized in that: The step S1 specifically includes: collecting a 3D human posture estimation dataset with any number of viewing angles, using an uncalibrated movable camera; defining a 2D posture sequence from N dynamic viewing angles and the corresponding 3D pose sequence Where N is the number of viewpoints, T is the length of the pose sequence, J is the number of human skeleton joints defined by the model, and C in is the number of input channels, C out is the number of output channels.

3. The lightweight 3D human pose estimation method for an uncalibrated multi-view camera according to claim 2, characterized in that: The bidirectional Mamba model comprises a preprocessing module and a plurality of groups of bidirectional Mamba modules connected in series; The preprocessing module takes the input 2D pose sequence Mapping into latent space Where D is the dimension of the latent space; then x f Input multiple sets of bidirectional Mamba modules in series; each bidirectional Mamba module contains two branches, one of which first passes through the time Mamba module and then the space Mamba module, and the other first passes through the space Mamba module and then the time Mamba module. The two branches are fused by direct vector addition; after passing through multiple sets of bidirectional Mamba modules with the same model parameter structure, the prediction results are obtained. With the data set Calculate the loss and optimize the bidirectional Mamba model parameters by back-propagation.

4. The lightweight 3D human pose estimation method for an uncalibrated multi-view camera according to claim 3, characterized in that: The specific implementation process of the time Mamba module is as follows: For the input feature sequence Apply the normalization operation to normalize the feature sequence x f Mapped to two new feature spaces through two fully connected layers, respectively, and Perform a 1D convolution on the mapped sequence y and apply the SiLU activation function to obtain For the convolution feature y o ′ Perform two linear transformations respectively to obtain new features and Then y o ′After passing through the fully connected layer, the output of the linear transformation is activated using the log-SoftPlus function to ensure that the output time scale intermediate parameter Δ o is a positive number; A o With Δ o , B o With Δ o The two sets of parameters sum the feature dimension E respectively to obtain the extended features and in is a learnable parameter; The transformed features and C o As a Mamba module parameter, feature y o ' is used as the input of the Mamba module to obtain new features Add features f and z element by element, apply a fully connected layer and SiLU activation function to get the output of the temporal Mamba module 5. The lightweight 3D human pose estimation method for an uncalibrated multi-view camera according to claim 4, characterized in that: The specific implementation process of the spatial Mamba module is as follows: Output of the Mamba module at a given time First, transpose it along the last two dimensions, and we get F' TMM As the input of the spatial Mamba module, the subsequent calculation steps are the same as those of the temporal Mamba module, and we get 6. The lightweight 3D human pose estimation method for an uncalibrated multi-view camera according to claim 5, characterized in that: In the bidirectional Mamba module, the output comes from the spatial Mamba module and time Mamba module The output of the bidirectional Mamba module is obtained through adaptive fusion The specific process of adaptive fusion is as follows: Convert to Again with Splicing Use linear transformation to calculate attention on α, and perform softmax operation to get attention weight Then α norm Split by the last dimension and Finally, the attention weights are used to TMM and f′ SMM Perform weighted fusion to obtain f Mamba : f Mamba =f TMM ⊙α TMM +f′ SMM ⊙α SMM ; Where ⊙ represents element-wise multiplication; The output of N bidirectional Mamba modules After a fully connected layer, the prediction result is obtained Used for calculation of subsequent losses.

7. The lightweight 3D human pose estimation method for an uncalibrated multi-view camera according to claim 6, characterized in that: The specific implementation process of step S3 is as follows: Step S3.1: In the pre-training stage, orthogonal projection is used to project the prediction results into a complete 2D skeleton pose sequence, and then the partially defective 2D pose is obtained by random masking. To achieve the purpose of data enhancement; at this stage, the 3D key point loss, speed loss and 2D reprojection loss are calculated as follows: Where L 3D is the 3D keypoint loss, L O is the speed loss, L 2D is the 2D reprojection loss; With X t,j Represents the predicted and actual values ​​of the three-dimensional true pose; With O t,j Respectively represent the predicted value and actual value of the speed of the key point between video frames, namely: With x t,j Represent the predicted value and actual value of the two-dimensional reprojection posture, respectively. The t in the subscript represents each video frame, j represents each key point, and δ t,j is the confidence of the 2D pose provided by the dataset; Step S3.2, in the fine-tuning phase, a task head is added to the pre-trained model to estimate the 3D pose; The 2D posture sequences of multiple perspectives in the same scene are superimposed and input into the bidirectional Mamba model. Then, some key points and perspectives are masked by random masking to simulate the occlusion challenges in the real environment. The random masking is implemented as follows: adding a control gate G in the input channel of each perspective i ∈0,1, occludes joint points or viewpoints.

Citation Information

Patent Citations

  • Robot action prediction method based on MAMBA and selective memory three-dimensional space

    CN118769250A

  • Three-dimensional motion capture and intelligent analysis system and method based on monocular camera

    CN119169701A

Cited By

  • Collaborative fusion three-dimensional human body posture estimation method based on Mama model

    CN121505693A