Training method and estimation method of single-view multi-person three-dimensional attitude estimation model

By introducing masking strategies and pose mask training in a single-view multi-person three-dimensional pose estimation model, combined with the Transformer attention mechanism, the occlusion and lack of depth information of multi-person pose estimation in a single-view perspective is solved, and the robustness and estimation accuracy of the model are improved.

CN120356037AActive Publication Date: 2025-07-22CENT SOUTH UNIV
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202510846988.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-24
Publication Date
2025-07-22
Estimated Expiration
2045-06-24

AI Technical Summary

Technical Problem

In a single-view multi-person scenario, human posture estimation is estimated to face the problem of lack of depth information, which makes it impossible to accurately restore the relative position of multiple people in three-dimensional space, the occlusion phenomenon affects joint position detection, and the timing inconsistency of multiple people's movements in dynamic scenarios, making it difficult to effectively separate individual motion characteristics.

Method used

By introducing a variety of masking strategies to simulate occlusion phenomena, structured pose representation and pose mask training methods are adopted, and the Transformer attention and timing attention mechanism are combined to improve the robustness and estimation accuracy of the model in occlusion scenarios.

Benefits of technology

It improves the robustness and estimation accuracy of the model in complex motion environments, can effectively handle joint position detection in occlusion scenarios, and enhances the understanding of depth information and learning ability on time series.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120356037A_ABST
    Figure CN120356037A_ABST
Patent Text Reader

Abstract

The invention discloses a training method and an estimation method of a single-view multi-person three-dimensional attitude estimation model, and the training method of the single-view multi-person three-dimensional attitude estimation model comprises the steps: collecting multi-person attitude pictures; processing the multi-person posture picture by adopting structured posture representation to obtain a joint point set corresponding to the multi-person posture picture; generating a mask value according to a preset mask strategy, and performing mask processing on the joint point set by using the mask value to obtain a mask attitude tensor; and pre-training the backbone network for three-dimensional attitude estimation by using the mask attitude tensor to obtain a trained backbone network. According to the method, attitude mask training is introduced, so that the model can learn and adapt to various possible shielding conditions, and the robustness and estimation precision of the model in a complex motion environment are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of computer vision technology, and particularly relates to a training method and an estimation method for a single-view multi-person three-dimensional pose estimation model. Background Art

[0002] Multi-person three-dimensional pose estimation is an important research direction in the field of computer vision, which aims to infer the positions of key joint points of multiple human bodies by analyzing image or video data. This technology can be applied to many industries, such as group sports analysis, intelligent monitoring, and virtual reality, etc.

[0003] However, in the single-view multi-person scenario, human pose estimation faces great challenges. First, in a single view, due to the lack of depth information, the model cannot accurately restore the relative positions of multiple people in three-dimensional space. Especially when limbs overlap, it is difficult for the model to infer the real spatial distribution through a single view. Second, occlusion problems are common in multi-person interaction scenarios, and these occlusion phenomena affect the accurate detection of joint positions. In addition, the inconsistent time series of multi-person actions in dynamic scenarios also makes it difficult for the model to effectively separate the motion features of individuals in the time series, exacerbating the difficulty of pose estimation.

[0004] In view of the above challenges, how to provide a training method for a single-view multi-person three-dimensional pose estimation model is an urgent technical problem to be solved in this field. Summary of the Invention

[0005] To solve the above technical problems, the purpose of this application is to provide a training method and an estimation method for a single-view multi-person three-dimensional pose estimation model. By adopting various masking strategies, such as masking a small part of joints or an entire limb segment, etc., to simulate common occlusion phenomena in actual scenarios, and prompting it to extract and learn useful information from the data, thereby improving the robustness and inference ability of the model in occlusion scenarios.

[0006] To achieve the above purpose, this application provides a training method for a single-view multi-person three-dimensional pose estimation model: The above object of this application is achieved through the following technical solutions: A training method for a single-view multi-person three-dimensional pose estimation model includes: Collect multi-person pose pictures; Using a structured pose representation to process the multi-person pose pictures to obtain a set of joint points corresponding to the multi-person pose pictures; Generating a mask value corresponding to each joint point in the set of joint points according to a preset masking strategy; Using the mask value to perform masking processing on the set of joint points to obtain a masked pose tensor; Pre-train the backbone network for three-dimensional pose estimation using the masked pose tensor to obtain a trained backbone network.

[0007] Preferably, define the set of joint points , and after processing the set of joint points through a fully connected layer, obtain a new set of joint points P, ; where R is the entity space, N is the number of people in the multi-person pose image, T is the number of time frames, K is the number of joint points included in each person in the multi-person pose image, D is the feature dimension, and each element in the set of joint points P represents the d-th coordinate value of the j-th joint point of the i-th person in the t-th frame, i ∈ {1, 2,..., N}, t ∈ {1, 2,..., T}, j ∈ {1, 2,..., K}; Define the set of mask values M, , where R is the entity space, N is the number of people in the multi-person pose image, T is the number of time frames, K is the number of joint points included in each person in the multi-person pose image, and each element in the set of mask values represents the mask value of the j-th joint point of the i-th person in the t-th frame, and the calculation formula of is: where represents visible, represents occluded. If is visible, then takes the value of 1. If is occluded, then takes the value of 0; The masking the set of joint points with the mask value to obtain a masked pose tensor includes: Fusing each mask value with the corresponding coordinate value to obtain an occluded masked pose tensor , and the calculation formula is: .

[0008] Preferably, the backbone network includes: a single-person spatial attention module, a spatial attention module between multiple people, a single-person joint temporal attention module, a multi-person joint temporal attention module, and a single-person spatio-temporal attention module connected in sequence.

[0009] Preferably, the single-person spatial attention module is used to: Take the masked pose tensor as the first input tensor X of this module, , where R is the entity space, N is the number of people in the multi-person pose image, T is the number of time frames, K is the number of joint points included in each person in the multi-person pose image, and D is the feature dimension; According to the first input tensor X, for each human body instance i and time t, extract the corresponding first feature unit token , token , , representing the input feature high-dimensional vector of the i-th person at the t-th frame, i ∈ {1, 2,..., N}, t ∈ {1, 2,..., T}; Through the spatial attention mechanism, according to the first feature unit token , obtain each feature unit token of the first output tensor Y , token .

[0010] Preferably, the spatial attention module between multiple people is used for: Take the first output tensor Y in the single-person spatial attention module as the second input tensor of this module ; According to the second input tensor , extract and obtain second feature units token , token , where represents the input feature high-dimensional vector of the t-th frame; Through the spatial attention mechanism, according to the second feature unit token , obtain each feature unit token of the second output tensor , token , token .

[0011] Preferably, the single-person joint temporal attention module includes: Take the second output tensor in the spatial attention module between multiple people as the third input tensor of this module ; According to the third input tensor , extract and obtain the third feature unit token , token , representing the input feature high-dimensional vector of the j-th joint of the i-th person; Through the temporal attention mechanism, according to the third feature unit token , obtain each feature unit token of the third output tensor ​ , token 。

[0012] Preferably, the multi-person joint temporal attention module includes: Using the third output tensor in the single-person joint temporal attention module as the fourth input tensor of this module ; According to the fourth input tensor , extracting the fourth feature unit token , token , where represents the high-dimensional input feature vector of the j-th joint of all individuals; Through the temporal attention mechanism, according to the fourth feature unit token , obtaining each feature unit token of the fourth output tensor , token , token 。

[0013] Preferably, the single-person spatio-temporal attention module includes: Using the fourth output tensor in the multi-person joint temporal attention module as the fifth input tensor of this module ; According to the fifth input tensor , extracting the fifth feature unit token , token , where represents the high-dimensional input feature vector of the i-th person; Through the spatio-temporal attention mechanism, according to the fifth feature unit token , obtaining each feature unit token of the fifth output tensor , token , token 。

[0014] Preferably, the backbone network includes at least two attention units connected in sequence, where each of the attention units includes the single-person spatial attention module, the spatial attention module between multiple persons, the single-person joint temporal attention module, the multi-person joint temporal attention module, and the single-person spatio-temporal attention module connected in sequence.

[0015] The second object of this application is to provide a single-view multi-person three-dimensional pose estimation method.

[0016] The above-mentioned second application object of this application is achieved through the following technical solutions: A single-view multi-person three-dimensional pose estimation method, comprising: Collect multi-person pose pictures; Using a structured pose representation, process the multi-person pose pictures to obtain a set of joint points corresponding to the multi-person pose pictures; According to the multi-person pose pictures, construct a symmetric affinity matrix in the local space , and a symmetric affinity matrix in the global space , , where R represents the real number space and K is the total number of joint points; According to the symmetric affinity matrix in the local space and the symmetric affinity matrix in the global space , obtain the overall topology matrix A; The set of joint points passes through a fully connected layer to obtain a high-dimensional feature vector P. Multiply the overall topology matrix A and the high-dimensional feature vector P to perform matrix multiplication to obtain a high-dimensional feature vector ; Input the high-dimensional feature vector into the trained backbone network obtained by the training method of the single-view multi-person three-dimensional pose estimation model as described above, and perform three-dimensional mapping through a multi-layer perceptron to obtain a pose estimation result.

[0017] The present application provides a training method for a single-view multi-person three-dimensional pose estimation model. By introducing pose mask training, by adding a mask to the input data during the training phase to simulate the situation of information loss, the model can learn and adapt to various possible occlusion situations, improving the robustness and estimation accuracy of the model in a complex motion environment.

[0018] The present application also introduces the Transformer attention and temporal attention mechanisms, thereby efficiently capturing the context information of the human body in the time series, comprehensively understanding the complex dependencies between joints, realizing the joint learning and reasoning of the root joint position and body joint displacement by the model, and improving the learning ability of the model in the time series.

[0019] The present application also introduces additional pose information on the structured pose, combines the local and global space topologies of individual joints, fuses the feature vectors of joint space relationships in the pose feature tensor, provides more relationship constraints and expressions for pose estimation, thereby enhancing the model's understanding of depth information, and finally using the trained backbone network to perform three-dimensional pose estimation, thus achieving accurate estimation of the pose after occlusion. Description of the Drawings

[0020] To more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the accompanying drawings required for the description of the embodiments or the prior art. Obviously, the accompanying drawings in the following description are only some embodiments recorded in the present application. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.

[0021] Figure 1 Schematic diagram of a single-view multi-person three-dimensional pose estimation model training method in an embodiment of the present application; Figure 2 Schematic diagram of pose mask training in an embodiment of the present application; Figure 3 Schematic diagram of the backbone network architecture in an embodiment of the present application; Figure 4 Schematic diagram of a single-view multi-person three-dimensional pose model estimation method in an embodiment of the present application. Detailed implementation manners

[0022] To make the objectives, technical solutions, and advantages of the embodiments of the present application clearer, the following will clearly and completely describe the technical solutions in the embodiments of the present application with reference to the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are some, but not all, of the embodiments of the present application. Based on the embodiments of the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts fall within the protection scope of the present application.

[0023] In addition, the technical features in the various embodiments or individual embodiments provided by the present application can be combined with each other arbitrarily to form a feasible technical solution. Such combination is not restricted by the order of steps and / or the structural composition mode, but must be based on what can be achieved by those of ordinary skill in the art. When the combination of technical solutions results in contradictions or cannot be implemented, it should be considered that such combination of technical solutions does not exist and is not within the protection scope required by the present application.

[0024] In the embodiments provided by the present application, it should be understood that the disclosed methods and systems can be implemented in other ways. The system embodiments described below are only illustrative. For example, the division of units and modules is only a logical function division. In actual implementation, there can be other division methods. For example, multiple units or modules can be combined, or integrated into another system, or some features can be ignored, or not executed. In addition, the coupling or direct coupling communication connection between the components shown or discussed with each other can be through some interfaces. The indirect coupling or communication connection of devices or modules can be electrical, mechanical, or other forms.

[0025] In addition, in each embodiment of the present application, each functional unit can be entirely integrated in a processor, or each unit can be separately used as a device alone, or two or more units can be integrated in a device; each functional unit in each embodiment of the present application can be implemented in the form of hardware, or in the form of a combination of hardware and software functional units.

[0026] Those of ordinary skill in the art can understand that all or part of the steps of implementing the following method embodiments can be completed through program instructions and related hardware. The foregoing program instructions can be stored in a computer-readable storage medium. When the program instructions are executed, the steps of the following method embodiments are executed; and the foregoing storage medium includes: various media that can store program codes, such as removable storage devices, read-only memory (ROM), magnetic disks, or optical discs.

[0027] It should be understood that in the present application, if the terms "system", "device", "unit", and / or "module" are used, they are only a method for distinguishing different components, elements, parts, portions, or assemblies at different levels. However, if other words can achieve the same purpose, the term can be replaced by other expressions.

[0028] In addition, the terms "first" and "second" are only used for descriptive purposes, and cannot be understood as indicating or implying relative importance or implicitly specifying the quantity of the indicated technical features. Thus, the features defined with "first" and "second" can explicitly or implicitly include one or more of such features. In the description of the present application, the meaning of "a plurality of" and "several" is two or more, unless otherwise specifically defined.

[0029] It should be noted that the structures, ratios, sizes, etc. shown in the drawings of this specification are only used to cooperate with the content disclosed in the specification for those who are familiar with this technology to understand and read, and are not used to limit the conditions under which the present application can be implemented. Therefore, they do not have a substantial technical meaning. Any modification of the structure, change of the proportional relationship, or adjustment of the size, without affecting the effects that the present application can produce and the purposes that can be achieved, should still fall within the scope covered by the technical content disclosed in the present application.

[0030] If a flowchart is used in the present application, the flowchart is used to illustrate the operations performed by the system according to the embodiments of the present application. It should be understood that the operations before or after do not necessarily need to be executed precisely in sequence. On the contrary, they can be executed in reverse order or simultaneously. At the same time, other operations can also be added to these processes, or one or several operations can be removed from these processes.

[0031] It should also be noted that in this text, terms such as "including", "comprising" or any other variants thereof are intended to cover non-exclusive inclusion, such that an article or device including a series of elements not only includes those elements, but also includes other elements not expressly listed, or elements inherent to such article or device. Without further limitation, an element defined by the statement "including one..." does not exclude the existence of additional identical elements in the article or device including the above elements.

[0032] The embodiments of this application are written in a progressive manner.

[0033] As Figure 1 shown, the embodiments of this application provide a method for training a single-view multi-person three-dimensional pose estimation model, including: S101. Collect multi-person pose pictures; Specifically, the multi-person pose pictures include various pose datasets, specifically: MuCo-3DHP, MuPoTS-3D, Human3.6M, MPI-INF-3DHP, and CMU Panoptic datasets.

[0034] S102. Use structured pose representation to process the multi-person pose pictures to obtain a set of joint points corresponding to the multi-person pose pictures; Specifically, the common multi-person pose estimation method in the prior art is to estimate the human pose by inferring the body joint point coordinates of all person instances in the image. Let N be the number of people in the image and K be the number of joint points included in each human body. represents the coordinates of the j-th body joint point of the i-th person, represents the set of joint points of all person instances in the image, denoted as: ; Therefore, the joint point coordinates in the 2D case are , and the joint point coordinates in the 3D case are . This application uses SPR (Structured pose representation), and SPR aims to unify the position information of human instances and human joint points, providing a single-stage solution for multi-person pose estimation. This application introduces a root joint in SPR to represent the position of the person instance.

[0035] Replace the 2D coordinates with 3D coordinates, and use to represent the coordinate position of the r-th root joint of the i-th person, then represents the coordinate position of the j-th joint of the i-th person, represents the coordinate position of the j-th body joint of the i-th person Displacement relative to the introduced r root joint coordinate positions The structured relationship between the r root joints and the position of the j-th body joint is: ; Therefore, the set of joint points of the human body posture can be expressed as: ;

[0036] Among them, is the set of joint points of the SPR of all person instances in the image data, N is the number of people in the image data, and K is the number of joint points included in each human body.

[0037] Based on the structured posture, additional posture information is introduced to model the joints. Combining the local and global spatial topological relationships of individual joints, the eigenvectors of the joint spatial relationships are fused in the posture feature tensor, providing more relationship constraints and expressions for posture estimation, thereby enhancing the model's understanding of depth information.

[0038] S103. Generate a mask value corresponding to each joint point in the set of joint points according to a preset mask strategy; Specifically, according to the preset mask strategy, simulate the situation where part of the posture is occluded, extract key features from the data, and generate mask values; Among them, the mask strategy is specifically divided into: Full-body joint mask: Randomly select some joints of the human body in the data for masking, such as wrists, knees, etc. This strategy can help the model use the full-body joint context information to infer the positions of the occluded key points.

[0039] Full-limb segment mask: Randomly select an entire limb segment, such as an arm or a leg, for masking. This strategy simulates the situation of severe occlusion, enabling the model to infer the posture of the occluded limb through the remaining visible joints.

[0040] In-limb joint mask: Randomly select some joints in a certain limb segment for masking and remove the parts other than the limb. This strategy enables the model to infer the positions of the invisible joints within the limb according to the context information of the limb.

[0041] In other embodiments of the present application, an implementation manner of this step is: Randomly or according to a set ratio, select some joints or limbs of the human body to be occluded. Then, set the set of multi-person 2D posture tensors as , and the set of mask values is M, ; Wherein, R represents the real number space, N is the number of people in the image (i ∈ {1, 2,..., N}), T is the number of time frames (t ∈ {1, 2,..., T}), K is the number of joint points included in each human body (j ∈ {1, 2,..., K}), D is the feature dimension, and the set After being processed by a fully connected layer, the set P is obtained. ; In the set P, represents the d-th coordinate value of the j-th joint point of the i-th person in the t-th frame, represents the mask value of the j-th joint point of the i-th person in the t-th frame, which is expressed as: ; Among them, visible represents visible, occluded represents occluded, that is, if is visible, then is 1, if is occluded, then is 0.

[0042] S104. Use the mask value to perform mask processing on the joint point set to obtain a masked pose tensor; Specifically, according to the mask value, fuse the pose tensor to obtain a masked pose tensor , and the calculation formula is: .

[0043] S105. Use the masked pose tensor to pre-train the backbone network for three-dimensional pose estimation to obtain a trained backbone network, and the training process is as Figure 2 shown.

[0044] Through the above steps, by introducing additional pose information, the joint relationship is structured. Through pose mask training, the backbone network can learn and adapt to various possible occlusion situations, prompting it to extract and learn how to recover the occluded human joints from the data, and improving the robustness and estimation accuracy of the model in complex motion environments.

[0045] Currently, in the single-view multi-person scenario, human pose estimation faces great challenges. First, in the single-view scenario, due to the lack of depth information, the model cannot accurately restore the relative positions of multiple people in the three-dimensional space. Especially when the limbs overlap, it is difficult for the model to infer the true spatial distribution through a single view. Second, occlusion problems are common in multi-person interaction scenarios, and these occlusion phenomena affect the accurate detection of joint positions. In addition, the inconsistent temporal sequences of multi-person actions in dynamic scenarios also make it difficult for the model to effectively separate the motion features of individuals in the time series, exacerbating the difficulty of pose estimation.

[0046] In the above embodiments, by introducing pose mask training, the learning ability of the backbone network for context information and structural priors is strengthened, enabling the backbone network to learn and adapt to various possible occlusion situations. By adopting various mask strategies, such as masking a small number of joints or an entire limb segment, etc., to simulate common occlusion phenomena in actual scenarios, the backbone network is prompted to extract and learn useful information from the data, thereby improving the robustness and inference ability of the model in occlusion scenarios.

[0047] Preferably, the backbone network specifically includes, as Figure 3 shown, a single-person spatial attention module, a spatial attention module between multiple persons, a single-person joint temporal attention module, a multi-person joint temporal attention module, and a single-person spatio-temporal attention module connected in sequence.

[0048] By introducing a temporal attention mechanism, the backbone network efficiently captures the context information of the human body in the time series, comprehensively understands the complex dependencies between joints, realizes the joint learning and inference of the root joint position and body joint displacement of the model, and improves the learning ability of the model in the time series.

[0049] In some embodiments, the single-person spatial attention module is used to: take the masked pose tensor as the first input tensor X of this module, , where R is the entity space, N is the number of people in the multi-person pose picture, T is the number of time frames, K is the number of joint points included in each person in the multi-person pose picture, and D is the feature dimension; according to the first input tensor X, for each human body instance i and time t, extract the corresponding first feature unit token , token (where token represents the smallest feature unit mapped from the key point position coordinates), represents the high-dimensional input feature vector of the i-th person in the t-th frame, i ∈ {1, 2,..., N}, t ∈ {1, 2,..., T}; through the spatial attention mechanism, according to the first feature unit token , obtain each feature unit token , token .

[0050] In some embodiments, the spatial attention module between multiple persons is used to: take the first output tensor Y in the single-person spatial attention module as the second input tensor of this module ; according to the second input tensor , extract and obtain second feature units token , token ; through the spatial attention mechanism, according to the second feature unit token , obtain the second output tensor for each feature unit token of token .

[0051] In some embodiments, the single-person joint temporal attention module is configured to: use the second output tensor in the spatial attention module among multiple persons as the third input tensor of this module ; according to the third input tensor , extract the third feature unit token of token ; through the temporal attention mechanism, according to the third feature unit token , obtain the third output tensor for each feature unit token of token .

[0052] In some embodiments, the multi-person joint temporal attention module is configured to: use the third output tensor in the single-person joint temporal attention module as the fourth input tensor of this module ; according to the fourth input tensor , extract the fourth feature unit token of token ; through the temporal attention mechanism, according to the fourth feature unit token , obtain the fourth output tensor for each feature unit token of token .

[0053] In some embodiments, the single-person spatio-temporal attention module is configured to: use the fourth output tensor in the multi-person joint temporal attention module as the fifth input tensor of this module ; according to the fifth input tensor , extract the fifth feature unit token of token ; through the spatio-temporal attention mechanism, according to the fifth feature unit token , obtain the fifth output tensor for each feature unit token of token .

[0054] In some embodiments, the backbone network is optionally at least 2 attention units connected in sequence, where each attention unit includes a single-person spatial attention module, a spatial attention module between multiple persons, a single-person joint temporal attention module, a multi-person joint temporal attention module, and a single-person spatio-temporal attention module connected in sequence.

[0055] By stacking multiple layers of the self-attention mechanism, higher-level and more abstract feature representations can be extracted layer by layer, and the context dependencies can be refined iteratively in each layer, thereby capturing complex semantic associations from local to global.

[0056] As Figure 4 shown, an embodiment of the present application provides a single-view multi-person three-dimensional pose estimation method, including: S201. Collect multi-person pose pictures; Specifically, collect single-view multi-person three-dimensional pose pictures that require pose estimation.

[0057] S202. Use a structured pose representation to process the multi-person pose pictures to obtain a set of joint points corresponding to the multi-person pose pictures; S203. Construct a symmetric affinity matrix , in the local space and a symmetric affinity matrix in the global space, where R represents the real number space and K is the total number of joint points; Specifically, based on the local topological relationship of body bone connections, construct a symmetric affinity matrix in the local space, , if two joint points in the human body structure are physically connected, the corresponding element in the affinity matrix is a non-zero value, otherwise it is zero. Based on the mutual connection between joint points in any pose during movement, construct a symmetric affinity matrix in the global space, .

[0058] S204. According to the symmetric affinity matrix in the local space and the symmetric affinity matrix in the global space, obtain the overall topological matrix A; Specifically, by fusing the symmetric affinity matrix in the local space and the symmetric affinity matrix in the global space, obtain the overall topological matrix A, and the calculation formula is: ; S205. The set of joint points passes through a fully connected layer to obtain a high-dimensional feature vector P. Perform matrix multiplication on the overall topological matrix A and the high-dimensional feature vector P to obtain a high-dimensional feature vector ; Specifically, the set of joint points passes through a fully connected layer to obtain a high-dimensional feature vector P. The overall topological matrix A and the high-dimensional feature vector P are subjected to matrix multiplication calculation to obtain a high-dimensional feature vector . The calculation formula is as follows: .

[0059] S206. Input the fused high-dimensional feature vector into the trained backbone network obtained by the training method of the above single-view multi-person three-dimensional pose estimation model, and perform three-dimensional mapping through a multi-layer perceptron to obtain a pose estimation result.

[0060] Specifically, the fused high-dimensional feature vector is used as the first input tensor and input into the trained backbone network. The high-dimensional feature dimension is subjected to three-dimensional mapping through a multi-layer perceptron, and finally a three-dimensional pose estimation result is obtained.

[0061] By providing a single-view multi-person three-dimensional pose estimation method, the present application combines the local and global spatial topological relationships of individual joints, fuses the feature vectors of joint spatial relationships in the pose feature tensor, and inputs them into the trained backbone network, so that the joint positions of the occluded parts can be estimated more accurately for the occluded three-dimensional poses.

[0062] The above serial numbers of the embodiments of the present application are only for description and do not represent the advantages and disadvantages of the embodiments.

[0063] In the above embodiments of the present application, the descriptions of each embodiment have their own emphases. For the parts not detailed in a certain embodiment, reference may be made to the relevant descriptions of other embodiments.

[0064] The above description of the disclosed embodiments enables those skilled in the art to implement or use the present application. Various modifications to these embodiments will be obvious to those skilled in the art, and the general principles defined herein can be implemented in other embodiments without departing from the spirit or scope of the present application. Therefore, the present application will not be limited to these embodiments shown herein, but will be accorded the widest scope consistent with the principles and novel features disclosed herein.

Claims

1. A training method for a single-view multi-person three-dimensional pose estimation model, characterized in that, Including: Collecting multi-person pose pictures; Using a structured pose representation to process the multi-person pose pictures to obtain a set of joint points corresponding to the multi-person pose pictures; Generating a mask value corresponding to each joint point in the set of joint points according to a preset mask strategy; Using the mask value to perform mask processing on the set of joint points to obtain a masked pose tensor; Using the masked pose tensor to pre-train a backbone network for 3D pose estimation to obtain a trained backbone network.

2. The training method of a single-view multi-person three-dimensional pose estimation model according to claim 1, characterized in that Define the set of joint points , and process the set of joint points through a fully connected layer to obtain a new set of joint points P, , where R is the entity space, N is the number of people in the multi-person pose picture, T is the number of time frames, K is the number of joint points included in each person in the multi-person pose picture, D is the feature dimension, and each element in the set of joint points P represents the d-th coordinate value of the j-th joint point of the i-th person in the t-th frame, where i ∈ {1, 2,..., N}, t ∈ {1, 2,..., T}, and j ∈ {1, 2,..., K}; Define the set of mask values M, , where R is the entity space, N is the number of people in the multi-person pose image, T is the number of time frames, K is the number of joint points included in each person in the multi-person pose image, and each element in the set of mask values represents the mask value of the j-th joint point of the i-th person in the t-th frame, and its calculation formula is: ; Among them, represents visible, represents occluded. If is visible, then the value is 1. If is occluded, then the value is 0; The using the mask value to perform mask processing on the set of joint points to obtain a masked pose tensor includes: Fuse each mask value with the corresponding coordinate value to obtain the masked pose tensor The calculation formula is as follows: 。 3. The training method of a single-view multi-person three-dimensional pose estimation model according to claim 2, characterized in that, The backbone network includes: a single-person spatial attention module, a spatial attention module between multiple persons, a single-person joint temporal attention module, a multi-person joint temporal attention module, and a single-person spatio-temporal attention module connected in sequence.

4. The training method of a single-view multi-person three-dimensional pose estimation model according to claim 3, characterized in that, The single-person spatial attention module is used for: Take the masked pose tensor as the first input tensor X of this module, , where R is the entity space, N is the number of people in the multi-person pose image, T is the number of time frames, K is the number of joint points included in each person in the multi-person pose image, and D is the feature dimension; According to the first input tensor X, for each human instance i and time t respectively, extract the corresponding first feature unit token , token , represents the high-dimensional input feature vector of the i-th person in the t-th frame, where i ∈ {1, 2,..., N} and t ∈ {1, 2,..., T}; Through the spatial attention mechanism, each feature unit token of the first output tensor Y is obtained according to the first feature unit token , and each feature unit token of the first output tensor Y is obtained , token .

5. The training method of a single-view multi-person three-dimensional pose estimation model according to claim 4, characterized in that, The spatial attention module between multiple persons is used for: Use the first output tensor Y in the single-person space attention module as the second input tensor of this module ; According to the second input tensor , extract to obtain second feature unit tokens , the token , where represents the high-dimensional input feature vector of the t-th frame; Through the spatial attention mechanism, based on the second feature unit token , obtain each feature unit token of the second output tensor , token , token .

6. The training method of a single-view multi-person three-dimensional pose estimation model according to claim 5, characterized in that The single-person joint temporal attention module includes: Use the second output tensor in the spatial attention module among multiple people as the third input tensor of this module ; According to the third input tensor , a third feature unit token is extracted , the token , represents the input feature high-dimensional vector of the j-th joint of the i-th person; Through the temporal attention mechanism, based on the third feature unit token , obtain the third output tensor for each feature unit token , token .

7. The training method of a single-view multi-person three-dimensional pose estimation model according to claim 6, characterized in that, The multi-person joint temporal attention module includes: Take the third output tensor in the single-person joint temporal attention module as the fourth input tensor of this module ; According to the fourth input tensor , extract the fourth feature unit token , the token , where represents the high-dimensional input feature vector of the j-th joint of all individuals; Through the sequential attention mechanism, according to the token of the fourth feature unit , obtain each feature unit token of the fourth output tensor , token , token .

8. The training method of a single-view multi-person three-dimensional pose estimation model according to claim 7, characterized in that The single-person spatio-temporal attention module includes: Take the fourth output tensor in the multi-person joint temporal attention module as the fifth input tensor of this module ; According to the fifth input tensor , extract the fifth feature unit token , the token , where represents the high-dimensional input feature vector of the i-th person; Through the spatio-temporal attention mechanism, according to the token of the fifth feature unit , obtain the fifth output tensor for each feature unit token , token .

9. The training method of a single-view multi-person three-dimensional pose estimation model according to claim 3, characterized in that, The backbone network includes: at least two attention units connected in sequence, where each attention unit includes the single-person spatial attention module, the spatial attention module between multiple persons, the single-person joint temporal attention module, the multi-person joint temporal attention module, and the single-person spatio-temporal attention module connected in sequence.

10. A single-view multi-person three-dimensional pose estimation method, characterized in that, Including: Collecting multi-person pose pictures; Using a structured pose representation to process the multi-person pose pictures to obtain a set of joint points corresponding to the multi-person pose pictures; Construct a symmetric affinity matrix in the local space based on the multi-person pose image , and a symmetric affinity matrix in the global space , , where R represents the real number space and K is the total number of joint points; Based on the symmetric affinity matrix in the local space and the symmetric affinity matrix in the global space , the overall topological matrix A is obtained; The set of joint points passes through a fully connected layer to obtain a high-dimensional feature vector P. The overall topological matrix A and the high-dimensional feature vector P are subjected to matrix multiplication calculation to obtain a high-dimensional feature vector ; Input the high-dimensional feature vector into the trained backbone network obtained by the training method of the single-view multi-person three-dimensional pose estimation model according to any one of claims 1-9, and perform three-dimensional mapping through a multi-layer perceptron to obtain a pose estimation result.

Citation Information

Patent Citations

  • Multi-person human body posture estimation method

    CN111339903A

  • Three-dimensional human body posture estimation method based on adaptive shielding

    CN117854112A

  • Shielding scene two-dimensional attitude estimation method and system based on spatio-temporal information

    CN119625846A

  • 3D human body posture estimation method and system, electronic equipment and storage medium

    CN120014713A

  • 3D Human Pose Estimation System

    US20220051437A1