Three-dimensional attitude estimation method based on data expansion and multi-view positioning
By introducing data expansion and multi-view positioning technology in three-dimensional human posture estimation, combining two-dimensional human posture estimation calculation method and dual space processing, the problems of low pose estimation accuracy and high computational complexity in the existing technology are solved, and efficient and accurate 3-dimensional human posture estimation is achieved.
Patent Information
- Application Number
- CN202510512324.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-23
- Publication Date
- 2025-05-27
- Estimated Expiration
- 2045-04-23
AI Technical Summary
The existing three-dimensional human posture estimation technology has problems such as low accuracy, high computational complexity, robustness and insufficient generalization ability in single-view and complex scenarios.
Using a three-dimensional pose estimation method based on data expansion and multi-view positioning, a multi-view 3D human pose estimation model is combined with a two-dimensional human pose estimation algorithm and dual spatial processing to efficiently and accurately estimate the three-dimensional human pose from image frames of different perspectives.
It effectively solves the occlusion and perspective limitation problems in single-view pose estimation, improves the accuracy of multi-person matching, significantly improves the robustness and generalization ability of the model in complex scenarios, and realizes high-precision three-dimensional human pose estimation.
Smart Images

Figure CN120048003A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of computer vision, and in particular to a three-dimensional posture estimation method based on data expansion and multi-view positioning. Background Art
[0002] Existing 3D human pose estimation technologies face many challenges. Single-view pose estimation methods rely only on single-view images and are easily affected by occlusion and perspective changes, resulting in inaccurate pose estimation. For example, in complex scenes, when a human body is partially occluded by an object, it is difficult for a single-view method to accurately determine the pose of the occluded part, thus affecting the accuracy of the overall pose estimation. Although multi-view pose estimation can alleviate the occlusion problem, it is difficult to match human bodies between different views. Due to differences in camera perspectives, the appearance and size of the same human body in different images vary greatly. Traditional methods find it difficult to accurately associate the same human body in different views, which reduces the reliability of pose estimation.
[0003] From the perspective of model performance, it is difficult for existing posture estimation models to strike a balance between computational efficiency and accuracy. Some high-precision models have complex structures, numerous parameters, high computational complexity, and require powerful hardware support, making them difficult to apply to resource-constrained devices such as mobile terminals and embedded systems. Lightweight models have high computational efficiency but often sacrifice accuracy and cannot meet application scenarios that require high accuracy in posture estimation. In addition, the current model is not adaptable enough to complex scenes and diverse human motions. In complex environments such as light changes and background interference, as well as in the face of some rare or special human motions, the robustness and generalization capabilities of the model need to be improved. Summary of the invention
[0004] In order to solve the above problems, the present invention proposes a three-dimensional posture estimation method and device based on data expansion and multi-view positioning. By introducing data expansion and multi-view positioning technology, combined with two-dimensional human posture estimation and dual space processing methods, efficient and accurate three-dimensional human posture estimation is achieved from image frames of different perspectives.
[0005] The specific plan is as follows:
[0006] On the one hand, the 3D pose estimation method based on data expansion and multi-view positioning includes:
[0007] S1, obtain image frames representing scenes from different perspectives and input them into the multi-view Figure 3 D Human body posture estimation model; the multi-view Figure 3 D The human posture estimation model includes input module, positioning module, detection module and fusion module.
[0008] S2, the input module identifies the two-dimensional bounding box of the human body in the image frame through the human body detector;
[0009] S3, the positioning module approximates the two-dimensional bounding box by fitting an ellipse, obtains the geometric characteristic equation of the two-dimensional bounding box, maps the geometric characteristic equation to the dual space, models the two-dimensional ellipse and the three-dimensional human body in the geometric characteristic equation by the dual conic curve equation and the dual quadratic surface equation, obtains the linear relationship between the two-dimensional ellipse and the three-dimensional human body, solves the linear relationship, and obtains the correlation information of the position of the same human body in different image frames;
[0010] S4, the detection module analyzes the associated information through the distilled two-dimensional human posture estimation algorithm to obtain the two-dimensional human posture;
[0011] S5, the fusion module preprocesses and encodes the two-dimensional human posture, obtains the corresponding key point tokens, integrates the key point tokens under different views, and predicts the three-dimensional posture of the human body in the world coordinate system.
[0012] Furthermore, before S1, the method further includes: performing data expansion on the two-dimensional human body posture by a grid-based human body posture data set generator to obtain the expanded two-dimensional human body posture, and performing data expansion on the multi-view human body posture based on the expanded two-dimensional human body posture. Figure 3 D. Train the human body posture estimation model;
[0013] The grid-based human posture dataset generator takes three-dimensional mesh vertices and a camera calibration matrix as input, renders a three-dimensional mesh in the camera space of each view through the camera calibration matrix, generates a corresponding two-dimensional image, and obtains a noisy two-dimensional posture based on the two-dimensional image; associates the noisy two-dimensional posture with the true value of the three-dimensional posture output by the grid-based human posture dataset generator, and constructs an expanded two-dimensional human posture pair.
[0014] Furthermore, the linear relationship between the two-dimensional ellipse and the three-dimensional human body is as follows:
[0015] ;
[0016] in, represents the overall scale factor; represents the dual conic section; represents the dual quadratic matrix; express and The relationship between them is determined by the projection matrix.
[0017] Furthermore, the linear relationship between the two-dimensional ellipse and the three-dimensional human body is transformed into a linear system, specifically:
[0018] ;
[0019] Among them, the matrix , is the Kronecker product, the matrix ,matrix ; For the matrix The vectorized form of ; For the matrix The vectorized form of ; It is a function that serializes the lower triangular elements of a symmetric matrix.
[0020] Furthermore, in S3, the linear relationship between the two-dimensional ellipse and the three-dimensional human body is solved to obtain the correlation information of the position of the same human body in different image frames. The calculation formula is as follows:
[0021] ;
[0022] Among them, the matrix 2D information representing multi-view geometric relationships and object detection, including matrices and matrix ; Represents the relevant parameters of the quadratic surface corresponding to the object in 3D space; represents a column vector of length 6F, where all elements are 0; F represents the total number of image frames;
[0023] Matrix-based Obtain the correlation information of the position of the same human body in different image frames.
[0024] Further, in S4, the detection module further includes a multi-cycle Transformer module provided with a self-distillation loss function, and the multi-cycle Transformer module is used to perform Figure 3 D. Self-distillation of human pose estimation model;
[0025] Among them, the self-distillation loss function includes key point labeling distillation loss and visual labeling distillation loss;
[0026] The key point labeling distillation loss is used to measure the difference in key point labels between different cycles, and the calculation formula is as follows:
[0027] ;
[0028] in, represents the key point marking distillation loss; represents mean square error; Indicates the key point marking output of the i-th cycle;
[0029] The visual mark distillation loss is used to evaluate the change of visual marks between cycles, and the calculation formula is as follows:
[0030] ;
[0031] in, Indicates visual marker distillation loss; Indicates visual markers; Indicates the number of cycles.
[0032] The present invention adopts the above technical solution and has the following beneficial effects:
[0033] (1) The present invention effectively solves the problems of occlusion and viewing angle limitation in single-view pose estimation by introducing multi-view constraints and distilled 2D human pose estimation technology, while improving the accuracy of multi-person matching.
[0034] (2) The present invention uses a grid-based human posture dataset generator for data expansion, which enhances the diversity and authenticity of the training data, thereby significantly improving the robustness and generalization ability of the model in complex scenarios;
[0035] (3) The present invention realizes the effective integration of key point information from different views through a fusion module, thereby achieving high-precision 3D human posture estimation at a low computational cost. BRIEF DESCRIPTION OF THE DRAWINGS
[0036] Figure 1 A flow chart of a three-dimensional posture estimation method based on data expansion and multi-view positioning according to an embodiment of the present invention;
[0037] Figure 2 A diagram of the steps of three-dimensional posture estimation based on data expansion and multi-view positioning according to an embodiment of the present invention;
[0038] Figure 3 This is a flow chart of the detection module execution of an embodiment of the present invention;
[0039] Figure 4 The present invention is a flowchart of a grid-based human posture dataset generator according to an embodiment of the present invention. DETAILED DESCRIPTION
[0040] The present invention is further described in detail below in conjunction with the embodiments and drawings, but the embodiments of the present invention are not limited thereto. Figure 1 As shown, the flow chart of the three-dimensional posture estimation method based on data expansion and multi-view positioning of the present invention includes:
[0041] S1, obtain image frames representing scenes from different perspectives and input them into the multi-view Figure 3 D Human body posture estimation model; the multi-view Figure 3 D The human posture estimation model includes input module, positioning module, detection module and fusion module.
[0042] Specifically, before S1, the 2D human body posture is expanded by a grid-based human body posture dataset generator to obtain the expanded 2D human body posture. Figure 3 D. Train the human body posture estimation model;
[0043] The grid-based human posture dataset generator takes three-dimensional mesh vertices and a camera calibration matrix as input, renders a three-dimensional mesh in the camera space of each view through the camera calibration matrix, generates a corresponding two-dimensional image, and obtains a noisy two-dimensional posture based on the two-dimensional image; the noisy two-dimensional posture is associated with the true value of the three-dimensional posture output by the grid-based human posture dataset generator to construct an expanded two-dimensional human posture.
[0044] The 3D mesh vertices are derived from the AMASS human shape dataset, which contains a rich set of human shapes and is presented in a standard 3D mesh model. They can be converted to SMPL vertices, and then the required key points are obtained through regression. The camera calibration matrix clarifies the position, angle and other parameters of the camera in the scene, providing a basis for subsequent rendering and projection. Next, the 3D mesh is randomly positioned in the scene to increase the diversity of the dataset. Subsequently, the open source software Body Visualizer and Trimesh are used to render the 3D mesh in the camera space of each view according to the camera calibration matrix to generate the corresponding 2D image. After that, the rendered 2D image is input into the same 2D pose estimator as that used in inference to obtain a noisy 2D pose. This process simulates the noise of 2D pose estimation in actual applications, making the generated dataset more realistic.
[0045] S2, the input module identifies the two-dimensional bounding box of the human body in the image frame through the human detector.
[0046] Specifically, the image frames of the three-dimensional scene received from different perspectives , assuming there is a set Individuals are in a three-dimensional scene, and everyone can Each object is detected in any of the images. In each image frame A two-dimensional bounding box The identification and bounding box are given by the human detector. The bounding box consists of three parameters Definition, where and are the height and width of the bounding box, respectively. is a 2D vector defining the center of the bounding box.
[0047] S3, the positioning module approximates the two-dimensional bounding box by fitting an ellipse, obtains the geometric characteristic equation of the two-dimensional bounding box, maps the geometric characteristic equation to the dual space, models the two-dimensional ellipse and three-dimensional human body in the geometric characteristic equation through the dual conic section equation and the dual quadratic surface equation, obtains the linear relationship between the two-dimensional ellipse and the three-dimensional human body, solves the linear relationship, and obtains the correlation information of the position of the same human body in different image frames.
[0048] Specifically, the linear relationship between the two-dimensional ellipse and the three-dimensional human body is:
[0049] ;
[0050] in, represents the overall scale factor; represents the dual conic section; represents the dual quadratic matrix; express and The relationship between them is determined by the projection matrix.
[0051] Specifically, the linear relationship between the two-dimensional ellipse and the three-dimensional human body is transformed into a linear system, specifically:
[0052] ;
[0053] Among them, the matrix , is the Kronecker product, the matrix ,matrix ; For the matrix The vectorized form of ; For the matrix The vectorized form of ; It is a function that serializes the lower triangular elements of a symmetric matrix.
[0054] Specifically, the linear relationship between the two-dimensional ellipse and the three-dimensional human body is solved to obtain the correlation information of the position of the same human body in different image frames. The calculation formula is as follows:
[0055] ;
[0056] Among them, the matrix 2D information representing multi-view geometric relationships and object detection, including matrices and matrix ; Represents the relevant parameters of the quadratic surface corresponding to the object in 3D space; represents a column vector of length 6F, where all elements are 0; F represents the total number of image frames;
[0057] Matrix-based Obtain the correlation information of the position of the same human body in different image frames.
[0058] Specifically, the positioning module converts the rectangular bounding box into an elliptical box that is easier to express with mathematical formulas. Associate an ellipse inscribed in the bounding box Specifically, each ellipse is is centered and aligned with the image axis, and its axis lengths are equal to and , we express each ellipse using the homogeneous quadratic form of the conic section equation:
[0059] .
[0060] in, It belongs to the symmetric matrix Homogeneous vector of the general 2D point that defines the conic section.
[0061] In order to associate the same person in different image frames, we use the homogeneous quadratic form of the quadratic surface equation to represent the ellipsoid in three-dimensional space, that is, the person in three-dimensional space:
[0062] ;
[0063] in, It belongs to the symmetric matrix A homogeneous 3D point on a quadratic surface defined by Projecting onto the image plane results in a conic section, denoted as . and The relationship between Definition, where is the camera’s intrinsic parameter matrix, is the rotation matrix, is the translation vector of the camera associated with each image frame.
[0064] Specifically, vector For the linear system solver, It can be solved ,in and Specifically:
[0065] , ;
[0066] get After passing The estimated dual quadratic surface is obtained by operation ,Right now Then, apply (Inverse operation of the adjoint operator) Get the estimated quadratic surface of the original space ,Right now By The reconstruction of Two-dimensional best-fit ellipse , and finally associate the positions of the same human body in different image frames.
[0067] Specifically, Figure 2 As shown, in this embodiment, the input multi-view image frames are first received, and the person in each image frame is identified by the human detector, and the two-dimensional bounding box of the human body is marked at the same time, providing position information for subsequent processing. Then, using the multi-view constraint method, the two-dimensional rectangular bounding box is first converted into an elliptical bounding box, and the linear relationship between the two-dimensional ellipse and the ellipsoid corresponding to the three-dimensional human body is solved, and the ellipsoid corresponding to the human body in the three-dimensional space is found to fit the projected ellipse with the two-dimensional elliptical bounding box, thereby estimating the position of the human body in the three-dimensional scene, and finally the same human body in different views is associated with each other. Subsequently, the position information of the same human body on different image frames after association is passed to the two-dimensional human body posture estimator after distillation, and the two-dimensional posture estimation of multiple views of different human bodies is performed respectively, and the two-dimensional skeleton of the human body in different views is independently extracted. On this basis, the system uses a fusion module trained by a data set obtained by a grid-based human body posture data set generator. Among them, the spatial posture transformer independently processes each two-dimensional posture skeleton to add joint type information to it, and outputs a vector containing the feature information of a single key point and the relationship information between different key points in the view. Finally, the received key point information is connected to the key points of all views through the fusion posture transformer and the three-dimensional posture in the world coordinate system is predicted.
[0068] S4, the detection module analyzes the associated information of the position of the same human body in different image frames through the distilled two-dimensional human body posture estimation algorithm to obtain the two-dimensional human body posture.
[0069] Specifically, the detection module also includes a multi-cycle Transformer module, which detects multiple Figure 3 D. Self-distillation of human pose estimation model;
[0070] The self-distillation loss function included in the multi-cycle Transformer module includes a key point labeling distillation loss and a visual labeling distillation loss;
[0071] The key point labeling distillation loss is used to measure the difference in key point labels between different cycles, and the calculation formula is as follows:
[0072] ;
[0073] in, represents the key point marking distillation loss; represents mean square error; Indicates the key point marking output of the i-th cycle;
[0074] The visual mark distillation loss is used to evaluate the change of visual marks between cycles, and the calculation formula is as follows:
[0075] ;
[0076] in, Indicates visual marker distillation loss; Indicates visual markers; Indicates the number of cycles.
[0077] Specifically, Figure 3 As shown in Figure 2, the multi-loop Transformer module allows the tokenized features to be circulated through the small Transformer model multiple times, ultimately achieving results comparable to those of a deep Transformer network. , extract features through the backbone network , then Divide into patches, flatten them and use a linear projection function to convert them into visual markers . Use additional Learnable tags To express Then, the keypoint tags are concatenated with the visual tags and fed into the Transformer encoder layer. For the multi-loop Transformer module, we denote the output keypoint tags and visual tags of each loop as In the In this cycle, we will and As input, the output is and . Self-distillation based on multi-loop Transformer modules is to avoid the extra computation caused by multi-loops. This method regards the complete reasoning process of the multi-loop Transformer module as the "teacher" and the single forward multi-loop Transformer module as the "student". Since the input and output tags in the Transformer layer are in the same vector space, distillation operations can be performed between the output tags of different loops, and information loss can be minimized. During the training phase, all loops of the multi-loop Transformer module are distilled at each inference. Take the multi-loop Transformer module as an example. Output of the next cycle and For Output of the next cycle and Distillation is performed to gradually transfer knowledge from subsequent cycles to the first cycle, so that the output of the first cycle can contain more complete reasoning information. To ensure that the tags learned through self-distillation can be accurately used for pose estimation, the output tags of all cycles of the multi-cycle Transformer module are , respectively, get the prediction results through the same prediction head , and compare these predictions with the true values to calculate the loss. In this way, the correctness of the labeling is constrained, allowing the model to be optimized towards accurately estimating the human body posture. Figure 4 The process of extracting 2D poses from multi-view images, generating a 3D model through 3D rendering, regressing key points to obtain accurate 3D poses, and finally reprojecting these 3D poses back to the 2D plane to generate noisy 2D poses is demonstrated.
[0078] S5, the fusion module preprocesses and encodes the two-dimensional human posture, obtains the corresponding key point tokens, and integrates the key point tokens under different views to predict the three-dimensional posture of the human body in the world coordinate system.
[0079] The fusion module is divided into two steps. The first is to use the spatial posture transformer to obtain the key point token, which is a vector containing various information of the key points, from the human body posture obtained by the detection module. Then these vectors are used to predict the three-dimensional human body through the fusion posture transformer. In fact, the two-dimensional posture estimation obtains the various key points of the human body, and finally the key points are connected to obtain the human body posture.
[0080] Specifically, in this embodiment, the received two-dimensional posture skeleton of the same human body in different image frames Pass it through the spatial attitude transformer, where , Represents the number of anatomical key points. Each 2D coordinate in is mapped to dimensional vector space, we get the formula:
[0081] ;
[0082] in Representative Key Points .
[0083] In order to make the model aware of the joint type information, for each linearly projected vector Add embeddings learned specifically for each keypoint type ,get:
[0084] ;
[0085] Anatomical embedding Contains feature information of different joint types, through In addition, the joint type information is incorporated into the feature representation of each keypoint, so that the subsequent model can better distinguish different types of joints.
[0086] Will get Key point tokens It is obtained through multi-head self-attention mechanism processing and multi-layer perceptron processing. indivual Dimensional vector , these output vectors not only contain the feature information of a single key point, but also integrate the relationship information between different key points in the view.
[0087] For all view keypoint tokens received Through the fusion posture transformer, in order to reduce the complexity of the attention module, the Keypoint tokens are concatenated into a vector , the 3D position encoding learned for each view , get the view token:
[0088]
[0089] The 3D position encoding contains the position information of the camera in 3D space. By adding it to the key point information, the model can take into account the differences in camera positions of different views during the fusion process, thereby more accurately fusing the information of different views and improving the accuracy of 3D pose prediction.
[0090] Set the view token for all views Fusion embedding is obtained through multi-head self-attention mechanism processing and multi-layer perceptron processing , the embedding combines the information of all views.
[0091] Embed the fused Passed to the regression head. The regression head consists of a weighted sum and a 1-layer The weighted sum will be The view corresponds to The vectors are weighted and combined, and the weights are learned through training. These weights reflect the importance of different views in the final prediction. By learning these weights, the model can reasonably integrate the information of each view according to different input conditions. Then, after 1 layer Processing, and finally predicting the 3D posture skeleton in the world coordinate system , complete from multi-view Figure 2 The conversion from 3D pose to 3D pose.
[0092] Although the present invention has been specifically shown and described in conjunction with the preferred embodiments, it should be understood by those skilled in the art that various changes may be made to the present invention in form and details without departing from the spirit and scope of the present invention as defined by the appended claims, all of which are within the scope of protection of the present invention.
Claims
1. A three-dimensional pose estimation method based on data expansion and multi-view positioning, characterized in that: include: S1, acquiring image frames representing scenes from different viewing angles and inputting them into a multi-view 3D human posture estimation model; the multi-view 3D human posture estimation model includes an input module, a positioning module, a detection module and a fusion module; S2, the input module identifies the two-dimensional bounding box of the human body in the image frame through the human body detector; S3, the positioning module approximates the two-dimensional bounding box by fitting an ellipse, obtains the geometric characteristic equation of the two-dimensional bounding box, maps the geometric characteristic equation to the dual space, models the two-dimensional ellipse and the three-dimensional human body in the geometric characteristic equation by the dual conic curve equation and the dual quadratic surface equation, obtains the linear relationship between the two-dimensional ellipse and the three-dimensional human body, solves the linear relationship, and obtains the correlation information of the position of the same human body in different image frames; S4, the detection module analyzes the associated information through the distilled two-dimensional human posture estimation algorithm to obtain the two-dimensional human posture; S5, the fusion module preprocesses and encodes the two-dimensional human posture, obtains the corresponding key point tokens, integrates the key point tokens under different views, and predicts the three-dimensional posture of the human body in the world coordinate system.
2. The three-dimensional posture estimation method based on data expansion and multi-view positioning according to claim 1, characterized in that: Before S1, the method further includes: performing data expansion on the two-dimensional human body posture by a grid-based human body posture data set generator to obtain the expanded two-dimensional human body posture, and training a multi-view 3D human body posture estimation model based on the expanded two-dimensional human body posture; The grid-based human posture dataset generator takes three-dimensional mesh vertices and a camera calibration matrix as input, renders a three-dimensional mesh in the camera space of each view through the camera calibration matrix, generates a corresponding two-dimensional image, and obtains a noisy two-dimensional posture based on the two-dimensional image; associates the noisy two-dimensional posture with the true value of the three-dimensional posture output by the grid-based human posture dataset generator, and constructs an expanded two-dimensional human posture pair.
3. The three-dimensional posture estimation method based on data expansion and multi-view positioning according to claim 1, characterized in that: In S3, the linear relationship between the two-dimensional ellipse and the three-dimensional human body is as follows: ; in, represents the overall scale factor; represents the dual conic section; represents the dual quadratic matrix; express and The relationship between them is determined by the projection matrix.
4. The three-dimensional posture estimation method based on data expansion and multi-view positioning according to claim 3 is characterized in that: The linear relationship between the two-dimensional ellipse and the three-dimensional human body is converted into a linear system, specifically: ; Among them, the matrix , is the Kronecker product, the matrix ,matrix ; For the matrix The vectorized form of ; For the matrix The vectorized form of ; It is a function that serializes the lower triangular elements of a symmetric matrix.
5. The three-dimensional posture estimation method based on data expansion and multi-view positioning according to claim 4 is characterized in that: In S3, the linear relationship between the two-dimensional ellipse and the three-dimensional human body is solved to obtain the correlation information of the position of the same human body in different image frames. The calculation formula is as follows: ; Among them, the matrix 2D information representing multi-view geometric relationships and object detection, including matrices and matrix ; Represents the relevant parameters of the quadratic surface corresponding to the object in 3D space; represents a column vector of length 6F, where all elements are 0; F represents the total number of image frames; Matrix-based Obtain the correlation information of the position of the same human body in different image frames.
6. The three-dimensional posture estimation method based on data expansion and multi-view positioning according to claim 1, characterized in that: In S4, the detection module further includes a multi-cycle Transformer module provided with a self-distillation loss function, and the multi-view 3D human posture estimation model is self-distilled through the multi-cycle Transformer module; Among them, the self-distillation loss function includes key point labeling distillation loss and visual labeling distillation loss; The key point labeling distillation loss is used to measure the difference in key point labels between different cycles, and the calculation formula is as follows: ; in, represents the key point marking distillation loss; represents mean square error; Indicates the key point marking output of the i-th cycle; The visual mark distillation loss is used to evaluate the change of visual marks between cycles, and the calculation formula is as follows: ; in, Indicates visual marker distillation loss; Indicates visual markers; Indicates the number of cycles.
Citation Information
Patent Citations
3D posture estimation method based on multi-view deep sensor frame
CN108389227A
Three-dimensional human body posture estimation method based on feature fusion and sample enhancement
CN111428586A
Three-dimensional human pose estimation method and related apparatus
US20220415076A1
2-d and 3-d pose estimation of articles from 2-d images
WO2003030738A1