3D pose estimation method based on data augmentation and multi-view localization

Through a three-dimensional pose estimation method based on data expansion and multi-view positioning, combined with dual space technology and multi-cycle Transformer module, the single-view occlusion and viewing angle limitations are solved, the multi-view matching accuracy and model robustness are improved, and efficient and high-precision three-dimensional human pose estimation is achieved.

CN120048003BActive Publication Date: 2025-08-08HUAQIAO UNIVERSITY
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510512324.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-04-23
Publication Date
2025-08-08
Estimated Expiration
2045-04-23

AI Technical Summary

Technical Problem

The existing three-dimensional human posture estimation technology is susceptible to occlusion and viewing angle changes in single-view methods, and the human body matching is difficult in multi-view methods, and the model is difficult to balance between computing efficiency and accuracy, especially on resource-constrained devices, and the robustness and generalization capabilities are insufficient.

Method used

Using a method based on data expansion and multi-view positioning, a multi-view 3D human posture estimation model is used, combined with input modules, positioning modules, detection modules and fusion modules, dual space technology and multi-cycle Transformer modules are used to self-distillate to achieve efficient estimation of two-dimensional human posture to three-dimensional posture.

Benefits of technology

It effectively solves the problems of single-view occlusion and perspective limitation, improves the accuracy of multi-person matching, enhances the robustness and generalization ability of the model in complex scenarios, and realizes high-precision three-dimensional human posture estimation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120048003B_ABST
    Figure CN120048003B_ABST
Patent Text Reader

Abstract

This invention discloses a three-dimensional pose estimation method based on data augmentation and multi-view localization, which relates to the field of computer vision. The method includes: using a human body detector in an input module to identify the two-dimensional bounding boxes of human bodies in image frames with different viewpoints; a localization module using dual space technology to process these bounding boxes, obtaining a linear relationship between two-dimensional and three-dimensional human bodies by fitting ellipses, thereby solving the position association information of the same human body in different image frames; a detection module using an optimized two-dimensional human pose estimation algorithm to analyze this association information to obtain accurate two-dimensional human pose; and a fusion module preprocessing and encoding these two-dimensional poses to generate key point tokens, and integrating the key point tokens from different views to predict the three-dimensional pose of the human body in the world coordinate system. By processing image frames of scenes with different viewpoints and using dual space technology, the present invention achieves efficient estimation of human pose from two-dimensional to accurate three-dimensional human pose.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of computer vision, and in particular to a three-dimensional posture estimation method based on data expansion and multi-view positioning. Background Art

[0002] Existing 3D human pose estimation technologies face numerous challenges. Single-view pose estimation methods rely solely on single-view images and are susceptible to occlusion and perspective changes, resulting in inaccurate pose estimation. For example, in complex scenes, when a human body is partially occluded by an object, single-view methods struggle to accurately determine the pose of the occluded part, affecting the accuracy of the overall pose estimation. While multi-view pose estimation can alleviate the occlusion problem, it makes it difficult to match human bodies between different views. Due to differences in camera perspective, the appearance and size of the same human body in different images vary significantly. Traditional methods struggle to accurately associate the same human body in different views, which in turn reduces the reliability of pose estimation.

[0003] From a performance perspective, existing pose estimation models struggle to strike a balance between computational efficiency and accuracy. Some high-precision models have complex structures, numerous parameters, and high computational complexity. These models require powerful hardware and are difficult to apply to resource-constrained devices, such as mobile terminals and embedded systems. Lightweight models, while computationally efficient, often sacrifice accuracy, failing to meet the demands of applications requiring high pose estimation accuracy. Furthermore, current models lack adaptability to complex scenes and diverse human motions. Their robustness and generalization capabilities need to be improved in complex environments such as changing lighting and background interference, as well as for rare or unique human motions. Summary of the Invention

[0004] In order to solve the above problems, the present invention proposes a three-dimensional pose estimation method and device based on data expansion and multi-view positioning. By introducing data expansion and multi-view positioning technologies, combined with two-dimensional human pose estimation and dual space processing methods, it achieves efficient and accurate estimation of three-dimensional human pose from image frames with different perspectives.

[0005] The specific plan is as follows:

[0006] On the one hand, the 3D pose estimation method based on data expansion and multi-view positioning includes:

[0007] S1, obtain image frames representing scenes from different perspectives and input them into the multi-view Figure 3 D Human body posture estimation model; the multi-view Figure 3 D The human posture estimation model includes input module, positioning module, detection module and fusion module.

[0008] S2, the input module identifies the two-dimensional bounding box of the human body in the image frame through the human detector;

[0009] S3: The positioning module approximates the 2D bounding box by fitting an ellipse to obtain the geometric characteristic equation of the 2D bounding box. The geometric characteristic equation is mapped to the dual space. The 2D ellipse and the 3D human body in the geometric characteristic equation are modeled using the dual conic section equation and the dual quadratic surface equation. The linear relationship between the 2D ellipse and the 3D human body is obtained. The linear relationship is solved to obtain the correlation information of the position of the same human body in different image frames.

[0010] S4, the detection module analyzes the associated information through the distilled 2D human pose estimation algorithm to obtain the 2D human pose;

[0011] In S5, the fusion module preprocesses and encodes the two-dimensional human posture to obtain the corresponding key point tokens, integrates the key point tokens under different views, and predicts the three-dimensional posture of the human body in the world coordinate system.

[0012] Furthermore, before S1, the method further includes: performing data expansion on the two-dimensional human body posture by a grid-based human body posture dataset generator to obtain the expanded two-dimensional human body posture, and performing data expansion on the multi-view human body posture based on the expanded two-dimensional human body posture. Figure 3 D. Train the human body pose estimation model;

[0013] The grid-based human pose dataset generator takes three-dimensional mesh vertices and a camera calibration matrix as input, renders a three-dimensional mesh in the camera space of each view using the camera calibration matrix, generates a corresponding two-dimensional image, and obtains a noisy two-dimensional pose based on the two-dimensional image; the noisy two-dimensional pose is associated with the true value of the three-dimensional pose output by the grid-based human pose dataset generator to construct an expanded two-dimensional human pose pair.

[0014] Furthermore, the linear relationship between the two-dimensional ellipse and the three-dimensional human body is as follows:

[0015] ;

[0016] in, represents the overall scale factor; represents the dual conic section; represents the dual quadratic matrix; express and The relationship between them is determined by the projection matrix.

[0017] Furthermore, the linear relationship between the two-dimensional ellipse and the three-dimensional human body is transformed into a linear system, specifically:

[0018] ;

[0019] Among them, the matrix , is the Kronecker product, the matrix ,matrix ; is a matrix The vectorized form of ; is a matrix The vectorized form of ; It is a function that serializes the lower triangular elements of a symmetric matrix.

[0020] Furthermore, in S3, the linear relationship between the two-dimensional ellipse and the three-dimensional human body is solved to obtain the correlation information of the position of the same human body in different image frames. The calculation formula is as follows:

[0021] ;

[0022] Among them, the matrix 2D information representing multi-view geometric relationships and object detection, including matrices and matrix ; Represents the relevant parameters of the quadratic surface corresponding to the object in 3D space; represents a column vector of length 6F, where all elements are 0; F represents the total number of image frames;

[0023] Matrix-based Obtain the correlation information of the position of the same human body in different image frames.

[0024] Furthermore, in S4, the detection module further includes a multi-cycle Transformer module provided with a self-distillation loss function, and the multi-cycle Transformer module is used to detect the multi-view Figure 3 D. Self-distillation of human pose estimation model;

[0025] Among them, the self-distillation loss function includes key point label distillation loss and visual label distillation loss;

[0026] The key point labeling distillation loss is used to measure the difference in key point labels between different cycles, and the calculation formula is as follows:

[0027] ;

[0028] in, represents the key point labeling distillation loss; represents the mean square error; Indicates the key point mark i-th cycle output;

[0029] The visual marker distillation loss is used to evaluate the change of visual markers between cycles and is calculated as follows:

[0030] ;

[0031] in, Indicates visual marker distillation loss; Indicates visual markers; Indicates the number of cycles.

[0032] The present invention adopts the above technical solution and has the following beneficial effects:

[0033] (1) This paper introduces multi-view constraints and distilled 2D human pose estimation technology, effectively solving the problems of occlusion and view limitation in single-view pose estimation, while improving the accuracy of multi-person matching.

[0034] (2) This paper uses a grid-based human posture dataset generator for data augmentation, which enhances the diversity and authenticity of the training data, thereby significantly improving the robustness and generalization ability of the model in complex scenarios;

[0035] (3) The present invention realizes the effective integration of key point information from different views through a fusion module, and achieves high-precision 3D human pose estimation at a low computational cost. BRIEF DESCRIPTION OF THE DRAWINGS

[0036] Figure 1 Flowchart of a method for three-dimensional pose estimation based on data expansion and multi-view positioning according to an embodiment of the present invention;

[0037] Figure 2 A diagram showing the steps of three-dimensional pose estimation based on data expansion and multi-view positioning according to an embodiment of the present invention;

[0038] Figure 3 This is a flow chart of the detection module execution according to an embodiment of the present invention;

[0039] Figure 4 This is a flowchart of a grid-based human posture dataset generator according to an embodiment of the present invention. DETAILED DESCRIPTION

[0040] The present invention will be described in further detail below with reference to the examples and accompanying drawings, but the embodiments of the present invention are not limited thereto. Figure 1 As shown in FIG, the flow chart of the 3D pose estimation method based on data expansion and multi-view positioning of the present invention includes:

[0041] S1, obtain image frames representing scenes from different perspectives and input them into the multi-view Figure 3 D Human body posture estimation model; the multi-view Figure 3 D The human posture estimation model includes input module, positioning module, detection module and fusion module.

[0042] Specifically, before S1, the 2D human posture data is expanded by the grid-based human posture dataset generator to obtain the expanded 2D human posture, and the multi-view Figure 3 D. Train the human body pose estimation model;

[0043] The grid-based human pose dataset generator takes three-dimensional mesh vertices and a camera calibration matrix as input, renders a three-dimensional mesh in the camera space of each view using the camera calibration matrix, generates a corresponding two-dimensional image, and obtains a noisy two-dimensional pose based on the two-dimensional image; the noisy two-dimensional pose is associated with the true value of the three-dimensional pose output by the grid-based human pose dataset generator to construct an expanded two-dimensional human pose.

[0044] The 3D mesh vertices are derived from the AMASS human shape dataset, which contains a rich collection of human shapes presented as standard 3D mesh models that can be converted to SMPL vertices, and then the required key points are obtained through regression. The camera calibration matrix specifies the position, angle, and other parameters of the camera in the scene, providing a basis for subsequent rendering and projection. Next, the 3D mesh is randomly positioned in the scene to increase the diversity of the dataset. Subsequently, using the open source software Body Visualizer and Trimesh, the 3D mesh is rendered in the camera space of each view based on the camera calibration matrix to generate the corresponding 2D image. The rendered 2D image is then input into the same 2D pose estimator as used during inference to obtain a noisy 2D pose. This process simulates the noise conditions of 2D pose estimation in real applications, making the generated dataset more realistic.

[0045] S2, the input module identifies the two-dimensional bounding box of the human body in the image frame through the human detector.

[0046] Specifically, the image frames of the three-dimensional scene received from different perspectives , assuming there is a set Individuals are in a three-dimensional scene, and everyone can Each object is detected in any of the images. In each image frame A two-dimensional bounding box The identification, bounding box is given by the human detector. The bounding box consists of three parameters Definition, where and are the height and width of the bounding box, respectively, is a two-dimensional vector defining the center of the bounding box.

[0047] S3, the positioning module approximates the two-dimensional bounding box by fitting an ellipse, obtains the geometric characteristic equation of the two-dimensional bounding box, maps the geometric characteristic equation to the dual space, models the two-dimensional ellipse and three-dimensional human body in the geometric characteristic equation through the dual conic section equation and the dual quadratic surface equation, obtains the linear relationship between the two-dimensional ellipse and the three-dimensional human body, solves the linear relationship, and obtains the correlation information of the position of the same human body in different image frames.

[0048] Specifically, the linear relationship between the two-dimensional ellipse and the three-dimensional human body is:

[0049] ;

[0050] in, represents the overall scale factor; represents the dual conic section; represents the dual quadratic matrix; express and The relationship between them is determined by the projection matrix.

[0051] Specifically, the linear relationship between the two-dimensional ellipse and the three-dimensional human body is transformed into a linear system, specifically:

[0052] ;

[0053] Among them, the matrix , is the Kronecker product, the matrix ,matrix ; is a matrix The vectorized form of ; is a matrix The vectorized form of ; It is a function that serializes the lower triangular elements of a symmetric matrix.

[0054] Specifically, the linear relationship between the two-dimensional ellipse and the three-dimensional human body is solved to obtain the correlation information of the position of the same human body in different image frames. The calculation formula is as follows:

[0055] ;

[0056] Among them, the matrix 2D information representing multi-view geometric relationships and object detection, including matrices and matrix ; Represents the relevant parameters of the quadratic surface corresponding to the object in 3D space; represents a column vector of length 6F, where all elements are 0; F represents the total number of image frames;

[0057] Matrix-based Obtain the correlation information of the position of the same human body in different image frames.

[0058] Specifically, the positioning module converts the rectangular bounding box into an elliptical box that is easier to express with mathematical formulas, by Associate an ellipse inscribed in the bounding box Specifically, each ellipse is is centered and aligned with the image axis, with axis lengths equal to and , we express each ellipse using the homogeneous quadratic form of the conic section equation:

[0059] .

[0060] in, It belongs to the symmetric matrix A homogeneous vector defining a general two-dimensional point of the conic section.

[0061] In order to associate the same person in different image frames, we use the homogeneous quadratic form of the quadratic surface equation to represent the ellipsoid in three-dimensional space, that is, the person in three-dimensional space:

[0062] ;

[0063] in, It belongs to the symmetric matrix The quadratic surface is defined by a homogeneous 3D point. Each quadratic surface Projecting it onto the image plane will result in a conic section, denoted as . and The relationship between the projection matrix Definition, where is the camera’s intrinsic parameter matrix, is the rotation matrix, is the camera's translation vector associated with each image frame.

[0064] Specifically, vector For linear system solvers, Can be solved ,in and Specifically:

[0065] , ;

[0066] get After passing The estimated dual quadratic surface is obtained by operation ,Right now Then, apply (Inverse operation of the adjoint operator) to obtain the estimated quadratic surface of the original space ,Right now By the quadratic surface The reconstruction of the quadratic curve Two-dimensional best-fit ellipse , and finally associate the positions of the same human body in different image frames.

[0067] Specifically, such as Figure 2 As shown, this embodiment first receives input multi-view image frames and uses a human detector to identify people in each frame, annotating the two-dimensional bounding boxes of the people to provide position information for subsequent processing. Next, using a multi-view constraint method, the two-dimensional rectangular bounding box is converted into an elliptical bounding box. By solving the linear relationship between the two-dimensional ellipse and the ellipsoid corresponding to the three-dimensional person, the ellipsoid corresponding to the person in three-dimensional space is found so that its projected ellipse fits the two-dimensional elliptical bounding box. This allows the person's position in the three-dimensional scene to be estimated, ultimately correlating the same person in different views. Subsequently, the correlated position information of the same person in different image frames is passed to a distilled two-dimensional human pose estimator, which performs two-dimensional pose estimation on multiple views of the different people and independently extracts two-dimensional skeletons of the people in different views. Based on this, the system utilizes a fusion module trained with a dataset obtained by a grid-based human pose dataset generator. The spatial pose transformer independently processes each two-dimensional pose skeleton, adding joint type information to it and outputting a vector containing the feature information of each key point and the relationship information between different key points in that view. Finally, the received key point information is connected by the fusion pose transformer to predict the three-dimensional pose in the world coordinate system.

[0068] S4, the detection module analyzes the correlation information of the position of the same human body in different image frames through the distilled two-dimensional human body posture estimation algorithm to obtain the two-dimensional human body posture.

[0069] Specifically, the detection module also includes a multi-cycle Transformer module, which is used to detect multiple Figure 3 D. Self-distillation of human pose estimation model;

[0070] The self-distillation loss function included in the multi-cycle Transformer module includes key point labeling distillation loss and visual labeling distillation loss;

[0071] The key point labeling distillation loss is used to measure the difference in key point labels between different cycles, and the calculation formula is as follows:

[0072] ;

[0073] in, represents the key point labeling distillation loss; represents the mean square error; Indicates the key point mark i-th cycle output;

[0074] The visual marker distillation loss is used to evaluate the change of visual markers between cycles and is calculated as follows:

[0075] ;

[0076] in, Indicates visual marker distillation loss; Indicates visual markers; Indicates the number of cycles.

[0077] Specifically, such as Figure 3 As shown in the figure, the multi-loop Transformer module can make the tokenized features circulate through the small Transformer model multiple times, and finally achieve the results comparable to the deep Transformer network. , extract features through the backbone network , then Divide into patches, flatten them and use a linear projection function to convert them into visual markers . Use additional Learnable tags To express Then, the keypoint tags are concatenated with the visual tags and fed into the Transformer encoder layer. For the multi-loop Transformer module, we denote the output keypoint labels and visual labels of each cycle as In the In this cycle, we will and As input, the output is and . Self-distillation based on multi-loop Transformer modules is to avoid the extra computational effort brought by multi-loops. This method regards the complete reasoning process of the multi-loop Transformer module as a "teacher" and the single forward multi-loop Transformer module as a "student". Since the input and output tags in the Transformer layer are in the same vector space, distillation operations can be performed between the output tags of different loops, and information loss can be minimized. During the training phase, all loops of the multi-loop Transformer module are distilled each time they are inferred. As an example, the multi-loop Transformer module Output of the secondary loop and For the first Output of the secondary loop and Distillation is performed to gradually transfer knowledge from subsequent cycles to the first cycle, so that the output of the first cycle can contain more complete reasoning information. To ensure that the tags learned through self-distillation can be accurately used for pose estimation, the output tags of all cycles of the multi-cycle Transformer module are marked during training. , respectively, obtain the prediction results through the same prediction head , and compare these predictions with the true values to calculate the loss. In this way, the correctness of the labeling is constrained, allowing the model to be optimized towards accurately estimating the human body posture. Figure 4 The process of extracting 2D poses from multi-view images, generating a 3D model through 3D rendering, regressing key points to obtain accurate 3D poses, and finally reprojecting these 3D poses back to the 2D plane to generate noisy 2D poses is demonstrated.

[0078] In S5, the fusion module preprocesses and encodes the two-dimensional human posture to obtain the corresponding key point tokens, and integrates the key point tokens under different views to predict the three-dimensional posture of the human body in the world coordinate system.

[0079] The fusion module is divided into two steps. The first is to use the spatial posture transformer to obtain the key point tokens, which are vectors containing various information of the key points, from the human body posture obtained by the detection module. These vectors are then used to predict the three-dimensional human body through the fusion posture transformer. In fact, the two-dimensional posture estimation obtains the various key points of the human body, and finally the key points are connected to obtain the human body posture.

[0080] Specifically, in this embodiment, the received two-dimensional posture skeleton of the same human body in different image frames Pass it through the spatial attitude transformer, where , Represents the number of anatomical key points. Each 2D coordinate in is mapped to dimensional vector space, we get the formula:

[0081] ;

[0082] in Representative Key points .

[0083] In order to make the model aware of joint type information, for each linearly projected vector Add embeddings learned specifically for each keypoint type ,get:

[0084] ;

[0085] Anatomical embedding Contains feature information of different joint types, through By adding, the joint type information is incorporated into the feature representation of each keypoint, so that the subsequent model can better distinguish different types of joints.

[0086] Will get Key point tokens Obtained through multi-head self-attention mechanism processing and multi-layer perceptron processing indivual dimensional vector ,These output vectors not only contain the feature information of a single key point, but also integrate the relationship information between different key points in the view.

[0087] For all view keypoint tokens received By fusion pose transformer, in order to reduce the complexity of attention module, each view Keypoint tokens are concatenated into a vector , the 3D position encoding learned for each view , get the view token:

[0088]

[0089] The 3D position encoding contains the camera's position information in 3D space. By adding it to the key point information, the model can take into account the differences in camera positions between different views during the fusion process, thereby more accurately fusing information from different views and improving the accuracy of 3D pose prediction.

[0090] Set the view token for all views Fusion embedding is obtained through multi-head self-attention mechanism processing and multi-layer perceptron processing , this embedding integrates information from all views.

[0091] Embed the fused Passed to the regression head. The regression head consists of a weighted sum and a 1-layer The weighted sum will be The corresponding view The vectors are weighted and combined, and the weights are learned through training. These weights reflect the importance of different views in the final prediction. By learning these weights, the model can reasonably integrate the information of each view according to different input conditions. Then, after 1 layer Processing, and finally predicting the 3D posture skeleton in the world coordinate system , complete from multi-view Figure 2 3D pose conversion.

[0092] Although the present invention has been particularly shown and described in conjunction with preferred embodiments, it will be understood by those skilled in the art that various changes in form and details may be made to the present invention without departing from the spirit and scope of the invention as defined in the appended claims, and all such changes are within the scope of protection of the present invention.

Claims

1. A three-dimensional pose estimation method based on data augmentation and multi-view positioning, characterized in that: include: S1, acquiring image frames representing scenes from different perspectives and inputting them into a multi-view 3D human pose estimation model; the multi-view 3D human pose estimation model includes an input module, a positioning module, a detection module and a fusion module; S2, the input module identifies the two-dimensional bounding box of the human body in the image frame through the human detector; S3: The positioning module approximates the 2D bounding box by fitting an ellipse to obtain the geometric characteristic equation of the 2D bounding box. The geometric characteristic equation is mapped to the dual space. The 2D ellipse and the 3D human body in the geometric characteristic equation are modeled using the dual conic section equation and the dual quadratic surface equation. The linear relationship between the 2D ellipse and the 3D human body is obtained. The linear relationship is solved to obtain the correlation information of the position of the same human body in different image frames. S4, the detection module analyzes the associated information through the distilled 2D human pose estimation algorithm to obtain the 2D human pose; S5, the fusion module preprocesses and encodes the 2D human posture to obtain the corresponding key point tokens, integrates the key point tokens under different views, and predicts the 3D posture of the human body in the world coordinate system; The linear relationship between the two-dimensional ellipse and the three-dimensional human body is specifically: Among them, β if represents the overall scale factor; represents the dual conic section; represents the dual quadratic matrix; P f Represents Q i and C if The relationship between them is given by the projection matrix; The linear relationship between the two-dimensional ellipse and the three-dimensional human body is converted into a linear system, specifically: Among them, the matrix is the Kronecker product, matrix D∈R 6×9 , matrix E∈R 16×10 ; is a matrix The vectorized form of is a matrix The vectorized form of vech() is a function for serializing the lower triangular elements of a symmetric matrix; Solve the linear relationship between the two-dimensional ellipse and the three-dimensional human body to obtain the correlation information of the position of the same human body in different image frames. The calculation formula is as follows: M i oh i =0 6F ; Among them, the matrix M i ∈R 6F×(10+F) Represents 2D information of multi-view geometric relationships and object detection, including matrix G f and matrix ω i Represents the relevant parameters of the quadratic surface corresponding to the object in 3D space; 0 6F represents a column vector of length 6F, where all elements are 0; F represents the total number of image frames; Based on the matrix M i Obtain the correlation information of the position of the same human body in different image frames.

2. The three-dimensional pose estimation method based on data expansion and multi-view positioning according to claim 1, characterized in that: Before S1, the method further includes: performing data expansion on the two-dimensional human body posture using a grid-based human body posture dataset generator to obtain the expanded two-dimensional human body posture, and training a multi-view 3D human body posture estimation model based on the expanded two-dimensional human body posture; The grid-based human pose dataset generator takes three-dimensional mesh vertices and a camera calibration matrix as input, renders a three-dimensional mesh in the camera space of each view using the camera calibration matrix, generates a corresponding two-dimensional image, and obtains a noisy two-dimensional pose based on the two-dimensional image; the noisy two-dimensional pose is associated with the true value of the three-dimensional pose output by the grid-based human pose dataset generator to construct an expanded two-dimensional human pose pair.

3. The three-dimensional pose estimation method based on data expansion and multi-view positioning according to claim 1, characterized in that: In S4, the detection module further includes a multi-cycle Transformer module provided with a self-distillation loss function, and self-distills the multi-view 3D human pose estimation model through the multi-cycle Transformer module; Among them, the self-distillation loss function includes key point label distillation loss and visual label distillation loss; The key point labeling distillation loss is used to measure the difference in key point labels between different cycles, and the calculation formula is as follows: Among them, L kt represents the key point labeling distillation loss; MSE() represents the mean square error; KT i Indicates the key point mark i-th cycle output; The visual marker distillation loss is used to evaluate the change of visual markers between cycles and is calculated as follows: Among them, L vt Indicates visual mark distillation loss; VT i Indicates a visual mark; N indicates the number of cycles.

Citation Information

Patent Citations

  • 3D posture estimation method based on multi-view deep sensor frame

    CN108389227A

  • Three-dimensional human body posture estimation method based on feature fusion and sample enhancement

    CN111428586A