An end-to-end multi-person pose estimation method and device
Through the end-to-end multi-person pose estimation method, the visual feature encoder and decoder are used to decode and fine-tune the pose information directly from the original image, solving the problems of low efficiency and high computational volume of existing methods, and achieving efficient pose estimation and end-to-end optimization.
Patent Information
- Application Number
- CN202210658293.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-06-10
- Publication Date
- 2025-06-27
- Estimated Expiration
- 2042-06-10
AI Technical Summary
The existing multi-person pose estimation method is inefficient, has a large amount of calculation, and is difficult to achieve end-to-end optimization.
The end-to-end multi-person pose estimation method is adopted to obtain multi-scale fusion features through visual feature encoder, and combine the pose decoder and the joint node decoder to directly decode the pose information from the original image and fine-tune it, avoiding the step of converting the multi-person pose estimation into a single-person pose estimation.
The calculation amount of multi-person pose estimation is reduced, the estimation efficiency is improved, and the end-to-end optimization is achieved, avoiding the prior information required for fine-tuning the results.
Smart Images

Figure CN115131820B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of pose estimation, and particularly to an end-to-end multi-person pose estimation method and device. Background Art
[0002] Currently, the following scheme can be adopted for multi-person pose estimation: First, a human detection network is used to obtain single-person images from the original image; then, a single-person pose estimation network is used to obtain the pose estimation results of each single-person image; according to the overlap rate between the single-person detection box and other human detection boxes in the image, it is judged whether the single-person image is occluded; if there is no occlusion, the pose estimation result is directly output, otherwise, the graph representation is used to finely adjust the single-person pose result.
[0003] This scheme converts the multi-person pose estimation problem into a single-person pose estimation problem. In the scheme, the network calculation amount of the single-person pose estimation network is large, and the computational complexity increases linearly with the increase of the number of targets. For targets with severe occlusion, the results also need to be finely adjusted through a separate graph neural network, which takes a long time, resulting in low efficiency of the scheme for multi-person pose estimation. And if this scheme is optimized, since the implementation of this scheme involves prior information, end-to-end optimization cannot be carried out. Summary of the Invention
[0004] This application discloses an end-to-end multi-person pose estimation method and device to reduce the calculation amount of multi-person pose estimation, improve the efficiency of multi-person pose estimation, and enable end-to-end optimization.
[0005] According to the first aspect of the embodiments of this application, an end-to-end multi-person pose estimation method is provided, and the method includes:
[0006] Input the downsampled feature map corresponding to the obtained original image into a visual feature encoder, so that the visual feature encoder performs multi-scale information fusion on the input downsampled feature map to obtain a multi-scale fusion feature;
[0007] Input the multi-scale fusion feature into a pose decoder, so that the pose decoder decodes pose information from the multi-scale fusion feature, and the pose information includes: at least one candidate pose and the pose confidence corresponding to the candidate pose;
[0008] Input the pose information and the multi-scale fusion feature into a joint point decoder, so that for each candidate pose whose pose confidence meets the condition, the joint point decoder finely adjusts the position information of each joint point in the candidate pose according to the multi-scale fusion feature and outputs the position information of each finely adjusted joint point; the position information of each finely adjusted joint point in the candidate pose constitutes the target pose.
[0009] According to a second aspect of the embodiments of the present application, an end-to-end multi-person pose estimation device is provided. The device includes:
[0010] A multi-scale fusion feature acquisition module, configured to input the downsampled feature map corresponding to the obtained original image into a visual feature encoder, so that the visual feature encoder performs multi-scale information fusion on the input downsampled feature map to obtain multi-scale fusion features;
[0011] A pose information acquisition module, configured to input the multi-scale fusion features into a pose decoder, so that the pose decoder decodes pose information from the multi-scale fusion features, where the pose information includes: at least one candidate pose and a pose confidence corresponding to the candidate pose;
[0012] A target pose acquisition module, configured to input the pose information and the multi-scale fusion features into a joint point decoder, so that for each candidate pose whose pose confidence meets the condition, the joint point decoder fine-tunes the position information of each joint point in the candidate pose according to the multi-scale fusion features and outputs the position information of each fine-tuned joint point; the position information of each fine-tuned joint point in the candidate pose constitutes the target pose.
[0013] The technical solutions provided by the embodiments of the present application may include the following beneficial effects:
[0014] As can be seen from the above technical solutions, the solution provided by the present application can input the downsampled feature map corresponding to the obtained original image into a visual feature encoder to obtain multi-scale fusion features, then input the multi-scale fusion features into a pose decoder to decode pose information including candidate poses from the multi-scale fusion features, and then input the pose information and the multi-scale fusion features into a joint point decoder to realize fine-tuning of the joint points in the candidate pose to obtain the target pose. This solution does not require converting multi-person pose estimation into single-person pose estimation first, reducing the computational complexity of multi-person pose estimation, and the solution provided by the present application does not require a separately set graph neural network that requires prior information to finely adjust the result of pose estimation, improving the efficiency of pose estimation, and since the pose decoder and joint point decoder in this solution can be optimized end-to-end through self-learning.
[0015] It should be understood that the above general description and the following detailed description are only exemplary and explanatory, and cannot limit the present application. BRIEF DESCRIPTION OF THE DRAWINGS
[0016] The accompanying drawings herein are incorporated into the specification and constitute a part of this specification, showing embodiments consistent with this specification, and are used together with the specification to explain the principles of this specification.
[0017] Figure 1Flowchart of a multi-person pose estimation method provided by an embodiment of the present application;
[0018] Figure 2 Block diagram of a pose decoder provided by an embodiment of the present application;
[0019] Figure 3 Block diagram of a joint decoder provided by an embodiment of the present application;
[0020] Figure 4 Schematic block diagram of a multi-person pose estimation device provided by an embodiment of the present application;
[0021] Figure 5 Schematic diagram of an electronic device provided by an embodiment of the present application. Detailed implementation manners
[0022] Here, exemplary embodiments will be described in detail, and examples are shown in the drawings. When the following description refers to the drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The implementation manners described in the following exemplary embodiments do not represent all implementation manners consistent with the present application. On the contrary, they are merely examples of devices and methods consistent with some aspects of the present application as detailed in the appended claims.
[0023] The terms used in the present application are only for the purpose of describing specific embodiments and are not intended to limit the present application. The singular forms "a", "the", and "said" used in the present application and the appended claims are also intended to include the plural forms unless the context clearly indicates otherwise. It should also be understood that the term "and / or" as used herein refers to and includes any or all possible combinations of one or more of the associated listed items.
[0024] It should be understood that although the terms first, second, third, etc. may be used in the present application to describe various information, such information should not be limited to these terms. These terms are only used to distinguish the same type of information from each other. For example, without departing from the scope of the present application, the first information may also be referred to as the second information, and similarly, the second information may also be referred to as the first information. Depending on the context, the word "if" as used herein may be interpreted as "when" or "while" or "in response to determining".
[0025] To enable those skilled in the art to better understand the technical solutions provided by the embodiments of the present application and to make the above-mentioned objects, features, and advantages of the embodiments of the present application more obvious and understandable, the method embodiments provided by the embodiments of the present application will be described below with reference to the drawings.
[0026] Please refer to Figure 1 , Figure 1A method flowchart of an end-to-end multi-person pose estimation method is provided. As an example, this method can be used in an electronic device such as a PC.
[0027] As Figure 1 shown, the method includes the following steps:
[0028] Step 101: Input the downsampled feature map corresponding to the obtained original image into a Visual Feature Encoder, so that the Visual Feature Encoder performs multi-scale information fusion on the input downsampled feature map to obtain a multi-scale fusion feature.
[0029] As an example, the downsampled feature map corresponding to the original image can be obtained in the following way: Input the original image into a backbone network, which is formed by connecting K downsampling processing layers with different multiples; Determine the downsampled feature map output by each downsampling processing layer in the backbone network as the downsampled feature map input to the Visual Feature Encoder (the obtained downsampled feature maps can be collectively referred to as K-level feature maps). For example, if the backbone network is formed by connecting downsampling processing layers with multiples of 8 times, 16 times, and 32 times, after inputting the original image into the backbone network, the original image will be downsampled by 8 times, 16 times, and 32 times respectively through the downsampling processing layers with multiples of 8 times, 16 times, and 32 times, and then 3 feature maps with the dimensions of the original image reduced by 8 times, 16 times, and 32 times are obtained. These 3 feature maps are the downsampled feature maps input to the Visual Feature Encoder in this step 101. In this example, downsampling the original image is to obtain the feature information of each person object in the feature maps of each dimension in the original image, and to improve the efficiency of feature extraction.
[0030] Further, when inputting the downsampled feature map corresponding to the original image into the Visual Feature Encoder, the downsampled feature map can be first flattened into a feature vector sequence, and the feature vector sequence is input into the Visual Feature Encoder. Among them, the feature vector sequence includes at least one vector.
[0031] In this example, the downsampled feature map can be flattened into a feature vector in the following way: Flatten the feature matrix corresponding to each feature map in the downsampled feature map into a vector to be spliced, and then splice the vectors to be spliced corresponding to each feature map together to obtain the above-mentioned feature vector.
[0032] In the embodiment of the present application, after inputting the above-mentioned feature vector into the Visual Feature Encoder, the Visual Feature Encoder can extract the joint point information of the people included in the original image, etc. as the multi-scale fusion feature through multi-scale fusion feature fusion of the feature vector. Here, the joint point information at least includes the position coordinates of the joint points of the people in each feature map.
[0033] Step 102: Input the multi-scale fusion features into a Pose Decoder to decode pose information from the multi-scale fusion features by the Pose Decoder.
[0034] Among them, the pose information decoded by the Pose Decoder in Step 102 includes: at least one candidate pose and the pose confidence corresponding to the candidate pose. The pose confidence corresponding to the candidate pose is used to indicate the probability that the candidate pose is the pose of a certain person existing in the original image.
[0035] As an embodiment, the above Pose Decoder may at least include a pose self-attention layer, a pose cross-attention layer, and a pose information calculation layer. The method for decoding pose information from multi-scale fusion features by the Pose Decoder will be described in detail after introducing the Figure 1 method embodiment shown and in combination with Figure 2 which will not be elaborated here for the time being.
[0036] Optionally, the Pose Decoder in the embodiments of the present application may be implemented by various self-learning algorithms, such as being implemented through a Transformer algorithm architecture, and the present application is not limited thereto.
[0037] Step 103: Input the above pose information and the multi-scale fusion features into a Joint Decoder, so that for each candidate pose whose pose confidence meets the condition, the Joint Decoder fine-tunes the position information of each joint point in the candidate pose according to the multi-scale fusion features and outputs the position information of each fine-tuned joint point; the position information of each fine-tuned joint point in the candidate pose constitutes the target pose.
[0038] In this embodiment, a threshold may be preset. If the pose confidence of the candidate pose obtained by the Pose Decoder is greater than the threshold, it is determined that the pose confidence of the candidate pose meets the condition, and it is determined that the pose of the candidate pose is the pose existing in the original image, and then the candidate pose is input into the Joint Decoder.
[0039] As an implementation, the Joint Decoder in the embodiments of the present application may be implemented by various self-learning algorithms, or may also be implemented through a Transformer algorithm architecture like the Pose Decoder, and the present application is not limited thereto. The specific process of fine-tuning the position information of each joint point in the candidate pose by the Joint Decoder according to the multi-scale fusion features and outputting the position information of each fine-tuned joint point in Step 103 will be described in detail after introducing the Figure 1 method embodiment shown and in combination with Figure 3 which will not be elaborated here for the time being.
[0040] Thus far, it is completed. Figure 1The process shown
[0041] Through Figure 1 It can be seen from the method embodiments shown that in the embodiments of the present application, the downsampled feature map corresponding to the obtained original image can be directly input into the visual feature encoder to obtain multi-scale fusion features, and then the multi-scale fusion features are input into the pose decoder to decode the pose information including candidate poses from the multi-scale fusion features, and then the pose information and the multi-scale fusion features are input into the joint decoder to fine-tune the joints in the candidate pose to obtain the target pose. This solution does not require converting multi-person pose estimation into single-person pose estimation first, reducing the computational complexity of multi-person pose estimation, and the solution provided by the present application does not require a separately set graph neural network with prior information to fine-tune the result of pose estimation, improving the efficiency of pose estimation, and since the pose decoder and joint decoder in this solution can be optimized end-to-end through self-learning.
[0042] The following combines Figure 2 For the above Figure 1 The process of obtaining the pose information will be described in detail. As an embodiment, as Figure 2 shown, the pose decoder in this embodiment at least includes a pose self-attention layer (Pose-to-Pose Attention), a pose cross-attention layer (Deformable Feature-to-Pose Attention), and a pose information calculation layer.
[0043] Among them, the pose self-attention layer can mine the relationship between different Poses in the original image based on the Self-Attention method and a preset pose queue (Pose Querise) representing the positions of N pose points of the whole body. The relationship between different Poses mined by the pose self-attention layer at least includes: the overlapping occlusion situation between different Poses. In specific implementation, for each Pose, the relationship between this Pose and all other Poses will be mined, and then by calculating the fusion weight coefficient between this Pose and other Poses, the fusion result after fusing this Pose and other Poses is output, and this fusion result is the result obtained by updating the coordinates of each joint point in this Pose according to the coordinates and other information of each joint point in other Poses.
[0044] Optionally, when specifically implementing the mining of the relationship between one Pose and another Pose, the coordinates of the joint points included in the two Poses can be used to determine whether there is an overlapping area between the two Poses to avoid the influence of the overlapping area on pose judgment.
[0045] In this embodiment, there is a Pose Query representing the position of pose points of the whole body. The initial value of the Pose Query is a vector composed of the initial coordinates of a preset number of human joint points (such as the relationships between joint points like the nose, eyes, hands, elbow joints of a person, etc.). The initial coordinates of the human joint points can be set randomly. The initial coordinates of the preset number of human joint points can indicate the pose of a human body, and the pose of the human body refers to poses such as standing, running, squatting, etc.
[0046] The pose cross-attention layer performs cross-attention operations on the fusion results corresponding to each Pose and the multi-scale fusion features to obtain the pose embedding representations (Pose Embedding) corresponding to N Poses. Here, the pose cross-attention layer can obtain the Pose Embedding corresponding to N Poses through the following steps: The pose cross-attention layer is used to fuse the fusion results corresponding to N Poses output by the pose self-attention layer and the multi-scale fusion feature F. For example, according to the joint point information in the multi-scale fusion feature, the coordinates of the joint points in the fusion result corresponding to each Pose are updated to map the joint points of the human body existing in the original image to the joint point coordinates included in the fusion result of this Pose, and this Pose containing the updated coordinates of each joint point is determined as the pose embedding representation (Pose Embedding).
[0047] By performing cross-attention operations on the relationships between different Poses mined and the multi-scale fusion features, the above-mentioned pose cross-attention layer can achieve feature fusion between the multi-scale fusion feature F and N Poses, and obtain the fusion feature corresponding to each Pose (i.e., Pose Embedding).
[0048] The pose information calculation layer infers the pose information of the entire body in the original image through N Pose Embeddings. The pose information in this embodiment includes: N candidate poses and the pose confidence degrees corresponding to the candidate poses.
[0049] Optionally, the pose information calculation layer can determine the pose information of the entire body in the original image corresponding to each Pose Embedding and the pose confidence degree of the Pose corresponding to each Pose Embedding according to a preset human pose estimation network (such as a linear mapping network) based on N Pose Embeddings.
[0050] Furthermore, in this embodiment, a Pose with a pose confidence degree greater than a preset threshold can be determined as a candidate pose.
[0051] Preferably, when obtaining the pose information through the above-mentioned pose decoder, multiple iterations can be performed between the pose self-attention layer and the pose cross-attention layer, so as to obtain a Pose Embedding with more information, making the pose information of the whole body in the original image inferred by the pose information calculation layer more accurate. Among them, multiple iterations between the pose self-attention layer and the pose cross-attention layer can be achieved by setting multiple layers of pose self-attention layers and pose cross-attention layers. For example, setting two iterations between the pose self-attention layer and the pose cross-attention layer can be achieved by setting "pose self-attention layer - pose cross-attention layer - pose self-attention layer - pose cross-attention layer".
[0052] Optionally, based on the fact that when the pose self-attention layer in this embodiment first mines the relationship between poses, it mines the relationship between preset Pose Querise, and there is no association between the preset Pose Querise and the original image. Therefore, this embodiment can set at least two iterations between the pose self-attention layer and the pose cross-attention layer to ensure that the relationship between each pose in the original image can be mined.
[0053] The following combines Figure 3 with the above Figure 1 to detail the adjustment process of the joint point coordinates.
[0054] As an embodiment, as Figure 3 shown, the joint point decoder in this embodiment at least includes a joint point self-attention layer (Joint-to-Joint Attention), a joint point cross-attention layer (Deformable Feature-to-JointAttention), and a joint point coordinate adjustment layer:
[0055] The joint point self-attention layer models the correlation between each joint point according to the position information of each joint point in each candidate pose for each candidate pose, and obtains the modeling result corresponding to the candidate pose.
[0056] In this embodiment, the joint point self-attention layer is used to mine the correlation between each joint point in each candidate pose.
[0057] Exemplarily, for a candidate pose, based on the position information of each joint point in the pose, it can be determined which joint point of the human body each joint point is, and according to the relevant relationships between the human body joint points, the joint points are modeled. Modeling the joint points means associating the joint points with relevant relationships together as a visual model, that is, connecting the joint points in the candidate pose together. For example, connecting the arm joint points, the elbow joint points, and the shoulder joint points together, and connecting the head joint points and the neck joint points together, etc., to form a model of the joint points of the human body.
[0058] The joint point cross-attention layer calculates cross-attention between the modeling result corresponding to the candidate pose and the multi-scale fusion features to obtain the fusion features of each joint point in the candidate pose.
[0059] Optionally, when the joint point cross-attention layer calculates cross-attention between the modeling result corresponding to the candidate pose and the multi-scale fusion features, it means mapping the position information of each joint point in the multi-scale fusion features to the modeling result to correct the position information of the joint points in the modeling result that is inconsistent with the position information of each joint point in the original image.
[0060] Furthermore, based on the fact that the candidate pose is composed of the feature information of at least one joint point, the fusion features of each joint point in the candidate pose can be obtained through the following method:
[0061] For the position information of each joint point in the candidate pose, taking this joint point as the origin, according to the CrossAttention method, P feature points near the origin position are adaptively selected as representative feature points from the joint points of the multi-scale fusion features; the features of the representative feature points are aggregated to obtain the fusion feature of this joint point.
[0062] Optionally, in this embodiment, when selecting P feature points as representative feature points, P is preset, and the Cross Attention method will adaptively select feature points from around the feature point serving as the origin according to the required number P of representative feature points.
[0063] The joint point coordinate adjustment layer is used to finely adjust the position information of each joint point based on the fusion features of each joint point and output the finely adjusted joint point position information.
[0064] Optionally, the position information of each joint point can be finely adjusted based on the fusion features of each joint point. According to the fusion features of each joint point and the preset correlation between human joint points, the position information of the joint points in the input modeling result that does not conform to the correlation between human joint points can be adjusted. For example, when the elbow joint point of the human body is too far from the shoulder joint point and the hand joint point, the coordinates of the elbow joint point are adjusted according to the position information of the shoulder joint point and the hand joint point.
[0065] Preferably, when fine-tuning the coordinates in the candidate pose through the above joint point decoder, multiple iterations can be performed between the joint point self-attention layer and the joint point cross-attention layer, so as to obtain the fusion features of the joint points with more information, making the position information of the joint points after fine-tuning by the joint point coordinate adjustment layer more accurate. Among them, multiple iterations between the joint point self-attention layer and the joint point cross-attention layer can be achieved by setting multiple joint point self-attention layers and joint point cross-attention layers. For example, setting two iterations between the joint point self-attention layer and the joint point cross-attention layer can be achieved by setting "joint point self-attention layer - joint point cross-attention layer - joint point self-attention layer - joint point cross-attention layer".
[0066] The method provided by the embodiments of the present application has been described above. Next, the device provided by the embodiments of the present application will be described:
[0067] See Figure 4 , Figure 4 which is an end-to-end multi-person pose estimation device provided by the embodiments of the present application. As an embodiment, the device can be used in an electronic device such as a PC. The device includes:
[0068] A multi-scale fusion feature acquisition module 401, configured to input the downsampled feature map corresponding to the obtained original image into a visual feature encoder, so that the visual feature encoder performs multi-scale information fusion on the input downsampled feature map to obtain multi-scale fusion features.
[0069] A pose information acquisition module 402, configured to input the multi-scale fusion features into a pose decoder, so that the pose decoder decodes pose information from the multi-scale fusion features, where the pose information includes: at least one candidate pose and the pose confidence corresponding to the candidate pose.
[0070] The target pose acquisition module 403 is configured to input the pose information and the multi-scale fusion features into a joint point decoder, so that for each candidate pose whose pose confidence meets the condition, the joint point decoder fine-tunes the position information of each joint point in the candidate pose according to the multi-scale fusion features and outputs the position information of each fine-tuned joint point; the position information of each fine-tuned joint point in the candidate pose constitutes the target pose.
[0071] Optionally, the device further includes:
[0072] The downsampled feature map acquisition module is configured to input the original image into a backbone network, which is formed by connecting downsampling processing layers with different multiples; determine the downsampled feature map output by each downsampling processing layer in the backbone network as the downsampled feature map;
[0073] Inputting the obtained downsampled feature map corresponding to the original image into the visual feature encoder includes: flattening the downsampled feature map into a sequence of feature vectors and inputting the sequence of feature vectors into the visual feature encoder.
[0074] Optionally, the pose decoder at least includes a pose self-attention layer, a pose cross-attention layer, and a pose information calculation layer;
[0075] The pose self-attention layer mines the relationships between different poses in the original image based on the self-attention method and a preset pose queue representing the positions of N body-wide pose points, and obtains the fusion result after fusing each pose with other poses;
[0076] The pose cross-attention layer performs cross-attention operations on the fusion results corresponding to each pose and the multi-scale fusion features to obtain pose embedding representations corresponding to N poses;
[0077] The pose information calculation layer infers the whole-body pose information in the original image through the N pose embedding representations, and the pose information includes: N candidate poses and the pose confidence corresponding to the candidate poses.
[0078] Optionally, the joint point decoder at least includes a joint point self-attention layer, a joint point cross-attention layer, and a joint point coordinate adjustment layer;
[0079] The joint point self-attention layer models the correlation between each joint point according to the position information of each joint point in each candidate pose for each candidate pose, and obtains the modeling result corresponding to the candidate pose;
[0080] The joint point cross-attention layer calculates cross-attention between the modeling results corresponding to the candidate poses and the multi-scale fusion features to obtain the fusion features of each joint point in the candidate poses;
[0081] The joint point coordinate adjustment layer is used to finely adjust the position information of each joint point based on the fusion features of each joint point and output the finely adjusted joint point position information.
[0082] Optionally, the candidate pose consists of the feature information of at least one joint point; the fusion features of each joint point in the candidate pose are obtained through the following method:
[0083] For the position information of each joint point in the candidate pose, taking this joint point as the origin, according to the cross-attention method, adaptively select the feature points near the origin position in the multi-scale fusion features as representative feature points;
[0084] Aggregate the features of the representative feature points to obtain the fusion features of this joint point.
[0085] Correspondingly, the embodiment of the present application also provides a hardware structure diagram of an electronic device, specifically as Figure 5 shown. This electronic device can be the device for implementing the above end-to-end multi-person pose estimation method. As Figure 5 shown, this hardware structure includes: a processor and a memory.
[0086] Among them, the memory is used to store machine-executable instructions;
[0087] The processor is used to read and execute the machine-executable instructions stored in the memory to implement the method embodiment of the corresponding end-to-end multi-person pose estimation method as shown above.
[0088] As an example, the memory can be any electronic, magnetic, optical or other physical storage device that can contain or store information, such as executable instructions, data, etc. For example, the memory can be: volatile memory, non-volatile memory or similar storage media. Specifically, the memory can be RAM (Random Access Memory), flash memory, storage drives (such as hard disk drives), solid state drives, any type of storage disk (such as optical discs, DVDs, etc.), or similar storage media, or a combination thereof.
[0089] So far, the description of the Figure 5 shown electronic device is completed. The above are only the preferred embodiments of the present application and are not intended to limit the present application. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of the present application shall be included within the scope of protection of the present application.
Claims
1. An end-to-end multi-person pose estimation method, characterized in that, The method includes: Inputting the downsampled feature map corresponding to the obtained original image into a visual feature encoder, so that the visual feature encoder performs multi-scale information fusion on the input downsampled feature map to obtain a multi-scale fusion feature; Inputting the multi-scale fusion feature into a pose decoder, so that the pose decoder decodes pose information from the multi-scale fusion feature, where the pose information includes: at least one candidate pose and a pose confidence corresponding to the candidate pose; Inputting the pose information and the multi-scale fusion feature into a joint point decoder, so that for each candidate pose whose pose confidence meets the condition, the joint point decoder fine-tunes the position information of each joint point in the candidate pose according to the multi-scale fusion feature and outputs the position information of each fine-tuned joint point; the position information of each fine-tuned joint point in the candidate pose constitutes the target pose.
2. The method according to claim 1, wherein The downsampled feature map corresponding to the original image is obtained through the following steps: inputting the original image into a backbone network, where the backbone network is formed by connecting downsampling processing layers with different multiples; determining the downsampled feature map output by each downsampling processing layer in the backbone network as the downsampled feature map; Inputting the downsampled feature map corresponding to the obtained original image into the visual feature encoder includes: flattening the downsampled feature map into a sequence of feature vectors and inputting the sequence of feature vectors into the visual feature encoder.
3. The method according to claim 1, wherein The pose decoder at least includes a pose self-attention layer, a pose cross-attention layer, and a pose information calculation layer; The pose self-attention layer mines the relationships between different poses in the original image based on the self-attention method and a preset pose queue representing the positions of N body-wide pose points, and obtains a fusion result after fusing each pose with other poses; The pose cross-attention layer performs cross-attention operations on the fusion results corresponding to each pose and the multi-scale fusion feature to obtain pose embedding representations corresponding to N poses; The pose information calculation layer infers the whole-body pose information in the original image through the N pose embedding representations, where the pose information includes: N candidate poses and pose confidences corresponding to the candidate poses.
4. The method according to claim 1, characterized in that, The joint point decoder at least includes a joint point self-attention layer, a joint point cross-attention layer, and a joint point coordinate adjustment layer; The joint point self-attention layer models the correlation between each joint point for each candidate pose based on the position information of each joint point in the candidate pose, and obtains a modeling result corresponding to the candidate pose; The joint point cross-attention layer performs cross-attention calculations on the modeling result corresponding to the candidate pose and the multi-scale fusion feature to obtain the fusion features of each joint point in the candidate pose; The joint point coordinate adjustment layer is used to fine-tune the position information of each joint point based on the fusion features of each joint point and output the position information of the fine-tuned joint point.
5. The method according to claim 4, wherein The candidate pose is composed of the feature information of at least one joint point; the fusion features of each joint point in the candidate pose are obtained through the following method: For the position information of each joint point in the candidate pose, taking this joint point as the origin, according to the cross-attention method, P feature points located near the position of the origin are adaptively selected from the multi-scale fusion features as representative feature points; Aggregate the features of the representative feature points to obtain the fusion feature of this joint point.
6. An end-to-end multi-person pose estimation device, characterized in that, The device includes: A multi-scale fusion feature acquisition module, configured to input the downsampled feature map corresponding to the obtained original image into a visual feature encoder, so that the visual feature encoder performs multi-scale information fusion on the input downsampled feature map to obtain multi-scale fusion features; A pose information acquisition module, configured to input the multi-scale fusion features into a pose decoder, so that the pose decoder decodes pose information from the multi-scale fusion features, and the pose information includes: at least one candidate pose and the pose confidence corresponding to the candidate pose; A target pose acquisition module, configured to input the pose information and the multi-scale fusion features into a joint point decoder, so that for each candidate pose whose pose confidence meets the condition, the joint point decoder fine-tunes the position information of each joint point in the candidate pose according to the multi-scale fusion features and outputs the position information of each fine-tuned joint point; the position information of each fine-tuned joint point in the candidate pose constitutes the target pose.
7. The device according to claim 6, characterized in that, The device further includes: A downsampled feature map acquisition module, configured to input the original image into a backbone network, and the backbone network is formed by connecting downsampling processing layers with different multiples; determining the downsampled feature map output by each downsampling processing layer in the backbone network as the downsampled feature map; Inputting the downsampled feature map corresponding to the obtained original image into the visual feature encoder includes: flattening the downsampled feature map into a sequence of feature vectors, and inputting the sequence of feature vectors into the visual feature encoder.
8. The device according to claim 6, characterized in that, The pose decoder at least includes a pose self-attention layer, a pose cross-attention layer, and a pose information calculation layer; The pose self-attention layer, based on the self-attention method and a preset pose queue representing the positions of N pose points of the whole body, mines the relationships between different poses in the original image to obtain the fusion result after fusing each pose with other poses; The pose cross-attention layer, by performing cross-attention operations on the fusion results corresponding to each pose and the multi-scale fusion features, obtains the pose embedding representations corresponding to N poses; The pose information calculation layer, through the N pose embedding representations, infers the pose information of the whole body in the original image, and the pose information includes: N candidate poses and the pose confidence corresponding to the candidate poses.
9. The device according to claim 6, wherein The joint point decoder at least includes a joint point self-attention layer, a joint point cross-attention layer, and a joint point coordinate adjustment layer; The joint point self-attention layer, for each candidate pose, models the correlation between each joint point according to the position information of each joint point in the candidate pose to obtain the modeling result corresponding to the candidate pose; The keypoint cross-attention layer calculates cross-attention between the modeling results corresponding to the candidate poses and the multi-scale fusion features to obtain the fusion features of each keypoint in the candidate poses; The keypoint coordinate adjustment layer is used to finely adjust the position information of each keypoint based on the fusion features of each keypoint and output the finely adjusted keypoint position information.
10. The device according to claim 9, characterized in that, The candidate pose consists of the feature information of at least one keypoint; the fusion features of each keypoint in the candidate pose are obtained through the following method: For the position information of each keypoint in the candidate pose, taking this keypoint as the origin, according to the cross-attention method, adaptively select the feature points near the origin position in the multi-scale fusion features as representative feature points; Aggregate the features of the representative feature points to obtain the fusion features of this keypoint.
Citation Information
Patent Citations
Human body posture estimation method and device
CN113095106A
System for estimating a pose of one or more persons in a scene
US11074711B1