A Multi-Person Pose Estimation Method and Device Based on Context Instance Decoupling

Through a multi-person pose estimation model based on context instance decoupling, the backbone network, instance information abstract module and heat map estimation module are used to solve the bounding box cropping and key point assembly error problems in multi-person pose estimation, achieving higher robustness and efficiency, and are suitable for smart cities and digital retinal technologies.

CN114926895BActive Publication Date: 2025-07-18PEKING UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210339901.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-04-01
Publication Date
2025-07-18
Estimated Expiration
2042-04-01

AI Technical Summary

Technical Problem

The existing multi-person pose estimation methods have bounding box cutting errors, key point assembly errors and long-distance regression problems, which lack robustness and real-timeness, making it difficult to meet the needs of smart cities and digital retinal technologies.

Method used

A multi-person pose estimation model based on context instance decoupling, including a backbone network, an instance information abstract module, a global feature decoupling module and a heat map estimation module, is used to generate the probability distribution of each key point of each person through training sample input and preset loss function optimization.

Benefits of technology

It improves the robustness and efficiency of multi-person pose estimation, reduces the challenge of key point grouping, avoids the difficulty of long-distance regression of the single-stage regression method, and achieves higher accuracy and real-time.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114926895B_ABST
    Figure CN114926895B_ABST
Patent Text Reader

Abstract

This application relates to the technical fields of deep learning and pose estimation. More specifically, this application relates to a multi-person pose estimation method and apparatus based on context instance decoupling. The method includes: obtaining a preset number of images containing multiple persons; using the images containing multiple persons as training samples and inputting them into a multi-person pose estimation model based on context instance decoupling for training; using the trained multi-person pose estimation model based on context instance decoupling to perform pose estimation on a target image; wherein, the multi-person pose estimation model based on context instance decoupling is provided with an instance information abstraction module, a global feature decoupling module, and a heatmap estimation module. The method and apparatus of this application can explore context clues in a larger range, thereby being robust to spatial detection errors, and are superior in both accuracy and efficiency.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the technical fields of deep learning and pose estimation. More specifically, this application relates to a multi-person pose estimation method and apparatus based on context instance decoupling. Background Art

[0002] The multi-person pose estimation (MPPE) technology is a technology for detecting all people in an image and locating key points for each person. As an important step in human activity understanding, human-computer interaction, human syntactic analysis, etc., MPPE has attracted more and more attention.

[0003] Currently, the commonly used multi-person pose estimation methods include the top-down estimation method, the bottom-up estimation method, and the single-stage regression method. However, these methods have problems such as bounding box cropping errors, key point assembly errors, and long-distance regression, and cannot achieve good robustness. Summary of the Invention

[0004] Based on the above technical problems, the present invention aims to decouple the poses of multiple people in a target image based on context instance decoupling (CID), that is, use a trained multi-person pose estimation model based on context instance decoupling to perform pose estimation on the target image, wherein the multi-person pose estimation model based on context instance decoupling is provided with an instance information abstraction module, a global feature decoupling module, and a heat map estimation module.

[0005] The first aspect of the present invention provides a multi-person pose estimation method based on context instance decoupling, the method comprising:

[0006] Obtaining a preset number of images containing multiple people;

[0007] Using the images containing multiple people as training samples and inputting them into a multi-person pose estimation model based on context instance decoupling for training;

[0008] Using the trained multi-person pose estimation model based on context instance decoupling to perform pose estimation on the target image;

[0009] wherein the multi-person pose estimation model based on context instance decoupling is provided with an instance information abstraction module, a global feature decoupling module, and a heat map estimation module.

[0010] In some embodiments of the present invention, the multi-person pose estimation model based on context instance decoupling further includes a backbone network. Using the trained multi-person pose estimation model based on context instance decoupling to perform pose estimation on the target image includes:

[0011] Inputting the target image into the backbone network to obtain a global feature map, wherein the target image includes multiple people, and the global feature map contains three-dimensional features of all people;

[0012] Input the global feature map into the instance information abstraction module and the global feature decoupling module respectively;

[0013] Obtain the instance features of each person in the target image through the instance information abstraction module;

[0014] Input the instance features of each person into the global feature decoupling module, and the global feature decoupling module decouples an instance feature perception map based on the global feature map and the instance features of each person;

[0015] Input the instance feature perception map into the heat map estimation module to obtain the probability distribution of each key point of each person in the target image.

[0016] In some embodiments of the present invention, obtaining the instance features of each person in the target image through the instance information abstraction module includes:

[0017] Input the global feature map into the heat map module;

[0018] Extract the center point coordinates of each person;

[0019] Sample at the corresponding positions in the global feature map according to the center point coordinates of each person to obtain the instance features of each person in the target image.

[0020] In some embodiments of the present invention, before obtaining the instance features of each person in the target image, it further includes: recalibrating the center point features of each person based on spatial attention or channel attention.

[0021] In some embodiments of the present invention, the global feature decoupling module decouples an instance feature perception map based on the global feature map and the instance features of each person, including:

[0022] Based on the mapping relationship between the instance features of each person and the global feature map, recalibrate the instance features of each person from the spatial dimension to obtain a first instance feature perception map;

[0023] Based on the mapping relationship between the instance features of each person and the global feature map, recalibrate the instance features of each person from the channel dimension to obtain a second instance feature perception map;

[0024] Fuse the first instance perception map and the second instance feature perception map to obtain an instance feature perception map.

[0025] In some embodiments of the present invention, based on the mapping relationship between the instance features of each person and the global feature map, recalibrating the instance features of each person from the spatial dimension to obtain a first instance feature perception map includes:

[0026] Generate a spatial mask for each person in the global feature map to represent the weight of the foreground feature of each person;

[0027] Increase the weight of the foreground feature, recalibrate the spatial position in the instance feature of each person, and obtain the first instance feature perception map.

[0028] In some embodiments of the present invention, based on the mapping relationship between the instance feature of each person and the global feature map, recalibrate the instance feature of each person from the channel dimension to obtain the second instance feature perception map, including: reweighting the global feature map based on the person feature in the channel dimension to generate the second instance feature perception map.

[0029] In some embodiments of the present invention, fusing the first instance perception map and the second instance feature perception map to obtain the instance feature perception map, including: performing a weighted sum of the first instance feature perception map and the second instance feature perception map to obtain the instance feature perception map.

[0030] In some embodiments of the present invention, inputting the instance feature perception map into the heatmap estimation module to obtain the probability distribution of each key point of each person in the target image, including:

[0031] Input the instance feature perception map into the heatmap estimation module to obtain the heatmap corresponding to each person in the target image;

[0032] Among them, the heatmap corresponding to each person in the target image contains the probability distribution of each key point.

[0033] In some other embodiments of the present invention, training in the multi-person pose estimation model based on context instance decoupling includes training the multi-person pose estimation model based on context instance decoupling through a preset loss function, and the preset loss function is:

[0034]

[0035] Among them, represents the loss of the instance information abstraction module, represents the loss of the global feature decoupling module, and λ represents the weight.

[0036] The second aspect of the present invention provides a multi-person pose estimation device based on context instance decoupling, and the device includes:

[0037] An acquisition module for acquiring a preset number of images containing multiple people;

[0038] A training module for inputting the image containing multiple people as a training sample into the multi-person pose estimation model based on context instance decoupling for training;

[0039] A pose estimation module, configured to perform pose estimation on a target image by using a trained multi-person pose estimation model based on context instance decoupling;

[0040] Wherein, the multi-person pose estimation model based on context instance decoupling is provided with an instance information abstraction module, a global feature decoupling module, and a heatmap estimation module.

[0041] A third aspect of the present invention provides a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the following steps are implemented:

[0042] Obtain a preset number of images containing multiple persons;

[0043] Use the images containing multiple persons as training samples and input them into the multi-person pose estimation model based on context instance decoupling for training;

[0044] Perform pose estimation on a target image by using the trained multi-person pose estimation model based on context instance decoupling;

[0045] Wherein, the multi-person pose estimation model based on context instance decoupling is provided with an instance information abstraction module, a global feature decoupling module, and a heatmap estimation module.

[0046] A fourth aspect of the present invention provides a computer program product, including a computer program, and when the computer program is executed by a processor, the following steps are implemented:

[0047] Obtain a preset number of images containing multiple persons;

[0048] Use the images containing multiple persons as training samples and input them into the multi-person pose estimation model based on context instance decoupling for training;

[0049] Perform pose estimation on a target image by using the trained multi-person pose estimation model based on context instance decoupling;

[0050] Wherein, the multi-person pose estimation model based on context instance decoupling is provided with an instance information abstraction module, a global feature decoupling module, and a heatmap estimation module.

[0051] The technical solutions provided in the embodiments of the present application at least have the following technical effects or advantages:

[0052] In the technical solution provided in the embodiment of the present application, a trained multi-person pose estimation model based on context instance decoupling is used to perform pose estimation on a target image. A backbone network, an instance information abstraction module, a global feature decoupling module, and a heatmap estimation module are provided in the multi-person pose estimation model based on context instance decoupling. The target image including multiple persons is input into the backbone network to obtain a global feature map, where the global feature map contains the three-dimensional features of all persons. The global feature map is respectively input into the instance information abstraction module and the global feature decoupling module. The instance features of each person in the target image are obtained through the instance information abstraction module, and the instance features of each person are input into the global feature decoupling module. The global feature decoupling module decouples an instance feature perception map based on the global feature map and the instance features of each person, and the instance feature perception map is input into the heatmap estimation module to obtain the probability distribution of each key point of each person in the target image, which can explore context clues in a larger range, so as to be robust to spatial detection errors, reduce the challenge of key point grouping, and avoid the difficulty of long-distance regression faced by single-stage regression methods. Experiments show that the technical solution provided in the embodiment of the present application is superior to other estimation methods in terms of efficiency and accuracy.

[0053] It should be understood that the above general description and the following detailed description are only exemplary and explanatory, and cannot limit the present invention. BRIEF DESCRIPTION OF THE DRAWINGS

[0054] By reading the following detailed description of the preferred embodiments, various other advantages and benefits will become clear to those of ordinary skill in the art. The drawings are only for the purpose of showing the preferred embodiments and are not considered to be a limitation of the present application. Moreover, throughout the drawings, the same reference numerals are used to represent the same components. In the drawings:

[0055] Figure 1 A schematic diagram of the steps of a multi-person pose estimation method based on context instance decoupling in an exemplary embodiment of the present application is shown;

[0056] Figure 2 A schematic diagram of the working process of a multi-person pose estimation method based on context instance decoupling in an exemplary embodiment of the present application is shown;

[0057] Figure 3 A schematic diagram of extracting the center point of each person in the target image in an exemplary embodiment of the present application is shown;

[0058] Figure 4 A schematic diagram of the effect after training using the loss function of formula (9) in an exemplary embodiment of the present application is shown;

[0059] Figure 5 A schematic diagram of comparing the present application with other estimation methods is shown;

[0060] Figure 6 Shows a schematic comparison diagram of the multi-person pose estimation method based on context instance decoupling and the RoIAlign method in the experiment of this application;

[0061] Figure 7 Shows a schematic comparison diagram of three pose estimation methods, namely COCO, CrowdPose, and OCHuman, in the experiment of this application;

[0062] Figure 8 Shows a schematic structural diagram of the multi-person pose estimation device based on context instance decoupling in an exemplary embodiment of this application;

[0063] Figure 9 Shows a schematic structural diagram of a computer device provided in an exemplary embodiment of this application;

[0064] Figure 10 Shows a schematic diagram of a storage medium provided in an exemplary embodiment of this application. Detailed implementation manners

[0065] Hereinafter, embodiments of the present application will be described with reference to the accompanying drawings. However, it should be understood that these descriptions are merely exemplary and are not intended to limit the scope of the present application. In addition, in the following description, descriptions of well-known structures and technologies are omitted to avoid unnecessarily confusing the concepts of the present application. It is obvious to those skilled in the art that the present application can be implemented without one or more of these details. In other examples, some well-known technical features in the art are not described to avoid confusing the present application.

[0066] It should be noted that the terms used herein are only for describing specific embodiments and are not intended to limit the exemplary embodiments according to the present application. As used herein, unless the context clearly indicates otherwise, the singular form is also intended to include the plural form. In addition, it should also be understood that when the terms "comprise" and / or "include" are used in this specification, they indicate the presence of the described features, wholes, steps, operations, elements, and / or components, but do not exclude the presence or addition of one or more other features, wholes, steps, operations, elements, components, and / or combinations thereof.

[0067] Now, exemplary embodiments according to the present application will be described in more detail with reference to the accompanying drawings. However, these exemplary embodiments can be implemented in many different forms and should not be construed as being limited only to the embodiments set forth herein. The drawings are not drawn to scale, and some details may be enlarged for the purpose of clear illustration, and some details may be omitted. The shapes of various regions and layers shown in the drawings, as well as their relative sizes and positional relationships, are merely exemplary, and may deviate in practice due to manufacturing tolerances or technical limitations. Those skilled in the art can design regions / layers with different shapes, sizes, and relative positions according to actual needs.

[0068] The following will describe several embodiments in conjunction with the Figure 1-10 accompanying drawings of the specification to describe the exemplary embodiments according to the present application. It should be noted that the following application scenarios are only shown for the convenience of understanding the spirit and principle of the present application, and the embodiments of the present application are not limited in this regard. On the contrary, the embodiments of the present application can be applied to any applicable scenario.

[0069] Currently, common multi-person pose estimation (MPPE) methods include top-down estimation methods, bottom-up estimation methods, and single-stage regression methods. However, the top-down estimation method has problems such as incorrect bounding box cropping, the bottom-up estimation method has problems such as incorrect key point localization, and the single-stage regression method has problems such as long-distance regression. These methods cannot achieve good robustness and real-time performance. Also, considering that accurate and efficient MPPE is an important technology for realizing intelligent acquisition and perception of human information in a large number of videos, MPPE is also an important technical issue in the digital retina architecture.

[0070] Therefore, in some exemplary embodiments of the present application, focusing on the digital retina architecture, a multi-person pose estimation method based on context instance decoupling is provided, as Figure 1 shown. The method includes: S1, obtaining a preset number of images containing multiple persons; S2, using the images containing multiple persons as training samples and inputting them into a multi-person pose estimation model based on context instance decoupling for training; S3, using the trained multi-person pose estimation model based on context instance decoupling to perform pose estimation on a target image; wherein, the multi-person pose estimation model based on context instance decoupling is provided with an instance information abstraction module, a global feature decoupling module, and a heat map estimation module. The multi-person pose estimation method based on context instance decoupling can be applied to the digital retina architecture for intelligent acquisition and perception of human information, and is superior to the top-down estimation method, the bottom-up estimation method, and the single-stage regression method, because these methods cannot achieve good robustness and real-time performance and do not meet the requirements of current smart city and digital retina technologies.

[0071] Figure 2The working process of the present method is illustrated as follows. Figure 2 As shown, the multi-person pose estimation model based on context instance decoupling includes a backbone network, an instance information abstraction module, a global feature decoupling module, and a heatmap estimation module. Both the instance information abstraction module and the heatmap estimation module encapsulate a heatmap module. For example, when we input an image containing multiple people into the multi-person pose estimation model based on context instance decoupling, the image passes through the backbone network, the instance information abstraction module, the global feature decoupling module, and the heatmap estimation module in sequence, and finally outputs a heatmap, that is, the probability distribution of each key point of each person in the target image, which is also the estimation result of the input image.

[0072] In a specific implementation, the multi-person pose estimation model based on context instance decoupling further includes a backbone network. Using the trained multi-person pose estimation model based on context instance decoupling to perform pose estimation on a target image includes: inputting the target image into the backbone network to obtain a global feature map, where the target image includes multiple people and the global feature map contains the three-dimensional features of all people; inputting the global feature map into the instance information abstraction module and the global feature decoupling module respectively; obtaining the instance features of each person in the target image through the instance information abstraction module; inputting the instance features of each person into the global feature decoupling module, and the global feature decoupling module decouples an instance feature perception map based on the global feature map and the instance features of each person; inputting the instance feature perception map into the heatmap estimation module to obtain the probability distribution of each key point of each person in the target image.

[0073] It can be seen that the goal of the estimation method of the present application is to estimate the positions of the pose key points of each person in the target image, which can be expressed by formula (1):

[0074]

[0075] where MPPE represents multi-person pose estimation, I represents the target image, represents the j-th pose key point of the i-th person in the target image, and m and n respectively represent the total number of people in the target image and the number of key points each person has. For example, in COCO pose estimation, n is selected to be 17, while in the CrowdPose method, n is selected to be 14.

[0076] The present application sets a heatmap module in both the instance information abstraction module and the heatmap estimation module to locate the key points with it, and finally converts the decoupled global feature map into a heatmap, indicating the probability distribution map of each key point, to obtain our pose estimation result. The working method of the heatmap module can be expressed by formula (2):

[0077]

[0078] Among them, HM represents the heat map module, I represents the target image, and F represents the global feature map obtained by processing the target image through the backbone network. represents an n-channel heat map. If we want to perform inverse encoding using the result obtained in formula (2), it can be expressed by formula (3):

[0079]

[0080] The context instance decoupling proposed in this application decouples the multi-person feature map into a set of instance-aware feature maps, where each map represents the clues of a specific person and retains the context clues to infer his / her key points. In a specific implementation manner, the instance features of each person in the target image obtained through the instance information abstraction module include: inputting the global feature map into the heat map module; extracting the center point coordinates of each person; sampling at the corresponding positions in the global feature map according to the center point coordinates of each person to obtain the instance features of each person in the target image. Of course, it can be understood that before obtaining the instance-aware feature map of each person in the target image, it also includes: recalibrating the center point features of each person based on spatial attention or channel attention. Extracting the center point features of each person, as Figure 3 shown. The working process of this implementation manner or the working process of the instance information abstraction module can be expressed by formula (4):

[0081]

[0082] where, f (i) represents the instance-aware feature of the i-th person in the target image, which is a one-dimensional feature. The global feature map contains the three-dimensional features of all people and can be expressed as H*W*C. Then the size of this one-dimensional feature is C. The purpose here is to effectively separate the instances while retaining rich context clues for subsequent estimation of the key point positions. IIA represents the instance information abstraction module, and F represents the global feature map.

[0083] In a specific implementation manner, the global feature decoupling module decouples the instance feature perception map based on the global feature map and the instance feature of each person, including: recalibrating the instance feature of each person from the spatial dimension based on the mapping relationship between the instance feature of each person and the global feature map to obtain the first instance feature perception map; recalibrating the instance feature of each person from the channel dimension based on the mapping relationship between the instance feature of each person and the global feature map to obtain the second instance feature perception map; fusing the first instance perception map and the second instance feature perception map to obtain the instance feature perception map. Recalibrating the instance feature of each person from the spatial dimension based on the mapping relationship between the instance feature of each person and the global feature map to obtain the first instance feature perception map includes: generating a spatial mask for each person in the global feature map to represent the weight of the foreground feature of each person; increasing the weight of the foreground feature and recalibrating the spatial position in the instance feature of each person to obtain the first instance feature perception map. Recalibrating the instance feature of each person from the channel dimension based on the mapping relationship between the instance feature of each person and the global feature map to obtain the second instance feature perception map includes: reweighting the global feature map based on the person feature in the channel dimension to generate the second instance feature perception map. Fusing the first instance perception map and the second instance feature perception map to obtain the instance feature perception map includes: performing a weighted sum on the first instance feature perception map and the second instance feature perception map to obtain the instance feature perception map. The working process of this implementation manner or the working process of the global feature decoupling module can be represented by formula (5):

[0084]

[0085] In addition, it should be noted that the training in the multi-person pose estimation model based on context instance decoupling includes training the multi-person pose estimation model based on context instance decoupling through a preset loss function, and the preset loss function is represented by formula (6):

[0086]

[0087] Wherein, represents the loss of the instance information abstraction module, represents the loss of the global feature decoupling module, and λ represents the weight.

[0088] It should be noted that the regression method generates the key point coordinates based on the feature of the person center point. In this application, the center point is also selected as the key point, that is, the person center point is located through the heat map module, and finally the model outputs the heat map, that is, the probability distribution of each key point of each person in the target image is obtained. Locating the person center point through the heat map module can be represented by formula (7):

[0089]

[0090] Among them, C represents the heatmap corresponding to the center point. As Figure 2 shown, the heatmap, also known as the thermogram, indicates that each pixel is the center of a person. C can be input into formula (2) for inverse coding positioning to determine the center point position. Taking the center point feature of each person as the representative feature and re-labeling it on the global feature can be expressed by formula (8):

[0091]

[0092] It should be noted that in the actual application of the pose estimation method, we expect it to have strong discrimination ability to effectively distinguish visually similar people. In other words, if two adjacent or overlapping people have similar appearances, their features may be similar, resulting in cases of person decoupling failure. To enhance the discrimination of person features, in a preferred training method, the present application performs contrastive loss training on IIA to ensure the discrimination of each f (i) . Given a set of person features {f(i)}, we constrain the i-th person feature by minimizing the similarity between the i-th person feature and other features, which can be expressed by formula (9):

[0093]

[0094] where, represents the normalized feature of the i-th person, and τ represents the temperature coefficient, preferably set to 0.05.

[0095] In some embodiments of the present application, a spatial mask is generated for each person in the global feature map to represent the weight of the foreground feature of each person; the weight of the foreground feature is increased to recalibrate the spatial position in the global feature map, and the first instance feature perception map can be expressed by formula (10):

[0096]

[0097] where, M represents the spatial mask, also known as the foreground mask, represents the instance perception feature map. When generating the spatial mask, the spatial position I (i) (x i , y i ) of the i-th person in the image is considered, a relative covariance map is generated, and then the inner product of the instance feature and the feature at each spatial position is calculated. The generation of the spatial mask produces a mapping indicating pixel-level feature similarity, which can be expressed by formula (11):

[0098]

[0099] where, and Used to indicate the similarity of pixel-level features, and Sigmoid represents the activation function. Experiments have proved that increasing the weight of foreground features enhances the discriminability of human features, making it more robust to occlusion and interference from adjacent humans with similar appearances. The instance-aware feature map can better focus on the foreground of each person and ensure the generation of reliable keypoint heatmaps.

[0100] Channels play an important role in the encoding context, and each channel can be recoded as a feature detector. Therefore, in a specific implementation, different channels are decoupled, including: reweighting the global feature map based on human features in the channel dimension and generating a second instance feature-aware map, which can be represented by formula (12):

[0101]

[0102] Among them, represents the recalibration result, represents the product, and f (i) can be regarded as the human features of the clues reserved for the i-th person. The formula does not decouple a person into a specific channel of the feature map, but ensures that different people show different channel distributions. Of course, in order to achieve better results, the ability to decouple channels can be further trained according to formula (9). While retaining the context instances, the performance of channel recalibration can be further enhanced.

[0103] Fusing the different spatial positions and channels to obtain the global feature map can be represented by formula (13):

[0104]

[0105] Among them, ReLU represents the activation function, and Conv represents the convolution.

[0106] Figure 4 It shows that training a good model using the loss function of formula (9) and then recognizing the target image results in better performance. Compared with the bottom-up method that relies on keypoint grouping, the end-to-end training feature and robustness to detection errors of this application show better performance and efficiency. Figure 5 It also shows a comparison of several different pose estimation methods. Such as Figure 5As shown, (a) is a top-down estimation method for detecting human bounding boxes and performing pose estimation on each bounding box. (b) is a bottom-up estimation method that first detects body key points and then groups them into corresponding people. (c) is a single-stage regression method that regresses the coordinates of pose key points based on human features. Obviously, the estimation method described in this application is highly robust to occlusion. Compared with the top-down and bottom-up estimation methods, this method has end-to-end trainability, is more robust to detection errors, and alleviates the challenge of key point grouping. It also avoids the difficulty of long-distance regression faced by the single-stage regression method. Experiments show that the multi-person pose estimation method based on context instance decoupling proposed in this application is superior to other estimation methods in terms of both efficiency and accuracy.

[0107] In some embodiments of the present application, in order to achieve better training results, a ground truth heat map is also used when training the multi-person pose estimation model based on context instance decoupling. When specifically training with the preset loss function described in formula (6), a Gaussian heat map generated from real coordinates can be used to accurately locate key points. The specific process can be represented by formulas (14) to (17):

[0108]

[0109]

[0110]

[0111]

[0112] Among them, represents the ground truth heat map, x, y j represents the spatial coordinates of the th key point, and α and β represent hyperparameters. α is preferably 2, and β is preferably 4. Through these formulas, the difference between the heat map corresponding to each person and the ground truth is measured to continuously adjust individual parameters.

[0113] A brief description of the experiments conducted for this application is as follows. The purpose of the experiment is to evaluate the method proposed in this application. We conducted the evaluation on the widely used multi-person pose estimation benchmark. All experiments were carried out on Pytorh. We used HRNet-W32 as the backbone network for all experiments and followed most of the configurations of the above-mentioned multi-person pose estimation model based on context instance decoupling. Set m equal to 30 in IIA, and set λ in formula (6) to 4. During the training process, the size of each image was adjusted to 512*512, and the learning rate of all layers was set to 0.001. We trained the model for 35 epochs on COCO. For CrowdPose, we trained the model for 300 stages and divided the learning rate by 10 at the 200th and 260th stages. The batch size was set to 20, and 40 for Ohuman, CrowdPose, and COCO. Table 1 shows the comparison of spatial and channel recalibration. As shown in Table 1, Spatial represents space, and Chennel represents channels. This loss helps channel recalibration and improves the performance from 64.9% to 65.3%. The spatial recalibration contrast loss reaches 64.6%, which indicates that the recalibration method can effectively separate people. Fusing channel and spatial recalibration always obtains the best performance with and without contrast loss. We conclude that learning discriminative people and jointly considering spatial decoupling and channel decoupling is very important for decoupling "people" in the target image.

[0114] Table 1 Comparison of Spatial and Channel Recalibration

[0115]

[0116] Moreover, considering that it is difficult to encode clues of a large number of people when the channel dimension is too small, and a too large channel dimension will increase the storage and computational costs, this implementation tested different embedding dimensions from 8 to 64 and reported their performance, as shown in Table 2. This shows that the smaller the dimension, the lower the performance. Setting too large a dimension (e.g., 64) will no longer significantly improve the performance. We set the embedding dimension to 32 as a reasonable trade-off between accuracy and computational cost.

[0117] Table 2 Comparison of Different Embedding Dimensions of Channels

[0118]

[0119] Figure 6 Shows a comparison of the multi-person pose estimation method based on context instance decoupling with RoIAlign in terms of the time cost of generating instance features for different numbers of people from the global feature map. As Figure 6 shown, RoIAlign +Output a feature map with a size of 14×14, and then upsample it to 56×56. RoIAlign* directly outputs a feature map with a size of 56×56. Obviously, the estimation efficiency of the method of the present application is much higher than that of RoIAlign + and RoIAlign*.

[0120] Moreover, the memory consumption of several methods was compared, as shown in Table 3. In these comparison works, HrHRNeT follows a bottom-up process, and DEKR and FCPose belong to single-stage regression methods. Therefore, among the three of them, HrHRNet is more efficient. However, the context instance decoupling method (CID) of the present application generally performs better than HrHRNet. Compared with the two single-stage regression methods, our method consumes more memory but achieves faster inference speed and better accuracy.

[0121] Table 3 Comparison of memory consumption of several methods

[0122]

[0123] In addition, Tables 4, 5, and 6 also list the comparisons of other estimation methods in various parameters.

[0124] Table 4 Comparison between the present application and COCO

[0125]

[0126] Table 5 Comparison between the present application and CrowdPose

[0127]

[0128] Table 6 Comparison between the present application and OCHuman

[0129]

[0130] The top-down method, bottom-up method, and single-stage regression method in Tables 4 to 6. It can be seen from Table 4 that compared with the top-down method, CID obtains better performance, 5.8% higher than Mask R-CNN. This shows that our decoupling strategy is better than the cropping of the bounding box. CID is also better than many bottom-up methods. It can also be seen from Tables 5 and 6 that CID has greater advantages. In addition, Figure 7 Figure 3 shows a comparison diagram of three pose estimation methods of COCO, CrowdPose, and OCHuman. It can be observed that the method of the present application can obtain reliable and accurate pose estimation even in challenging situations such as severe occlusion and human overlap, and can be applied to the processing of actual scenarios in smart cities, such as pedestrian perception and analysis in digital retina technology.

[0131] It should be understood that the above general description and the following detailed description are only exemplary and explanatory, and do not limit the present invention.

[0132] In some embodiments of the present application, a multi-person pose estimation device based on context instance decoupling for a digital retina architecture is also provided, as Figure 8 shown, the device includes:

[0133] An acquisition module 801, configured to acquire a preset number of images containing multiple persons;

[0134] A training module 802, configured to use the images containing multiple persons as training samples and input them into a multi-person pose estimation model based on context instance decoupling for training;

[0135] A pose estimation module 803, configured to perform pose estimation on a target image by using the trained multi-person pose estimation model based on context instance decoupling;

[0136] Wherein, the multi-person pose estimation model based on context instance decoupling is provided with an instance information abstraction module, a global feature decoupling module, and a heat map estimation module. It can be understood that the settings of the various modules of the device are integrated with the existing digital retina architecture, so the device can be applied to the digital retina architecture for intelligent acquisition and perception of human information, and the perception effect is accurate.

[0137] It should also be emphasized that the system provided in the embodiments of the present application can acquire and process relevant data based on artificial intelligence technology. Among them, artificial intelligence (AI) is to use a digital computer or a machine controlled by a digital computer to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use knowledge to obtain the best results of theory, methods, technologies, and application systems. Artificial intelligence basic technologies generally include technologies such as sensors, dedicated artificial intelligence chips, cloud computing, distributed storage, big data processing technology, operation / interaction systems, and mechatronics. Artificial intelligence software technologies mainly include several major directions such as computer vision technology, robotics, biometric technology, speech processing technology, natural language processing technology, and machine learning / deep learning.

[0138] Next, please refer to Figure 9 , which shows a schematic diagram of a computer device provided by some embodiments of the present application. As Figure 9As shown, the computer device 2 includes: a processor 200, a memory 201, a bus 202, and a communication interface 203. The processor 200, the communication interface 203, and the memory 201 are connected through the bus 202. A computer program that can run on the processor 200 is stored in the memory 201. When the processor 200 runs the computer program, it executes the multi-person pose estimation method based on context instance decoupling provided in any of the foregoing embodiments of the present application.

[0139] Among them, the memory 201 may include a high-speed random access memory (RAM), and may also include a non-volatile memory, such as at least one disk memory. Through at least one communication interface 203 (which can be wired or wireless), a communication connection is realized between this system network element and at least one other network element, and the Internet, a wide area network, a local area network, a metropolitan area network, etc. can be used.

[0140] The bus 202 may be an ISA bus, a PCI bus, an EISA bus, etc. The bus can be divided into an address bus, a data bus, a control bus, etc. Among them, the memory 201 is used to store a program. After receiving an execution instruction, the processor 200 executes the program. The multi-person pose estimation method based on context instance decoupling disclosed in any of the foregoing embodiments of the present application can be applied to the processor 200 or implemented by the processor 200.

[0141] The processor 200 may be an integrated circuit chip with the ability to process signals. In the implementation process, each step of the above method can be completed by the integrated logic circuit of the hardware in the processor 200 or the instructions in the form of software. The above-mentioned processor 200 may be a general-purpose processor, including a central processing unit (CPU for short), a network processor (NP for short), etc.; it may also be a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components. It can implement or execute the various methods, steps and logic block diagrams disclosed in the embodiments of the present application. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc. The steps of the method disclosed in combination with the embodiments of the present application can be directly embodied as being executed and completed by the hardware decoding processor, or executed and completed by a combination of the hardware and software modules in the decoding processor. The software module may be located in a mature storage medium in the art such as a random access memory, a flash memory, a read-only memory, a programmable read-only memory or an electrically erasable programmable memory, a register, etc. This storage medium is located in the memory 201, and the processor 200 reads the information in the memory 201 and combines its hardware to complete the steps of the above method.

[0142] The embodiment of the present application also provides a computer-readable storage medium corresponding to the multi-person pose estimation method based on context instance decoupling provided in the foregoing embodiment. Please refer to Figure 10 , Figure 10 The computer-readable storage medium shown is an optical disc 30, on which a computer program (i.e., a program product) is stored. When the computer program is run by a processor, it will execute the multi-person pose estimation method based on context instance decoupling provided in any of the foregoing embodiments.

[0143] In addition, examples of the computer-readable storage medium may also include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other optical and magnetic storage media, which will not be elaborated here one by one.

[0144] The computer-readable storage medium provided in the above embodiment of the present application and the quantum key distribution channel allocation method in the space-division multiplexing optical network provided in the embodiment of the present application are based on the same inventive concept and have the same beneficial effects as the methods adopted, run or implemented by the application programs stored therein.

[0145] The embodiments of the present application also provide a computer program product, including a computer program which, when executed by a processor, implements the steps of the multi-person pose estimation method based on context instance decoupling provided by any of the foregoing embodiments, including: obtaining a preset number of images containing multiple persons; using the images containing multiple persons as training samples and inputting them into a multi-person pose estimation model based on context instance decoupling for training; using the trained multi-person pose estimation model based on context instance decoupling to perform pose estimation on a target image; wherein, the multi-person pose estimation model based on context instance decoupling is provided with an instance information abstraction module, a global feature decoupling module, and a heatmap estimation module.

[0146] It should be noted that: The algorithms and displays provided herein are not inherently related to any particular computer, virtual device, or other equipment. Various general-purpose devices can also be used in conjunction with the teachings herein. The structure required to construct such devices is obvious from the above description. In addition, the present application is not directed to any specific programming language. It should be understood that the content of the present application described herein can be implemented using various programming languages, and the description of a specific language above is to disclose the best implementation mode of the present application. In the specification provided herein, a large number of specific details are set forth. However, it can be understood that the embodiments of the present application can be practiced without these specific details. In some instances, well-known methods, structures, and technologies are not shown in detail so as not to obscure the understanding of this specification.

[0147] Similarly, it should be understood that, in order to streamline the present application and assist in understanding one or more of the various inventive aspects, in the above description of the exemplary embodiments of the present application, the various features of the present application are sometimes grouped together into a single embodiment, figure, or description thereof. However, the disclosed method should not be construed as reflecting the intention that the claimed present application requires more features than are expressly recited in each claim. Rather, as reflected in the following claims, the inventive aspects lie in less than all the features of the single foregoing disclosed embodiment. Thus, the claims following the detailed description are hereby expressly incorporated into the detailed description, where each claim itself serves as a separate embodiment of the present application.

[0148] Those skilled in the art can understand that the modules in the devices in the embodiments can be adaptively changed and arranged in one or more devices different from the embodiments. The modules or units or components in the embodiments can be combined into one module or unit or component, and in addition, they can be divided into multiple sub-modules or sub-units or sub-components. Except that at least some of such features and / or processes or units are mutually exclusive, any combination can be used to combine all the features disclosed in this specification and all the processes or units of any method or device so disclosed. Unless otherwise explicitly stated, each feature disclosed in this specification can be replaced by an alternative feature that provides the same, equivalent or similar purpose.

[0149] Each component embodiment of the present application can be implemented in hardware, or in software modules running on one or more processors, or in a combination thereof. Those skilled in the art should understand that a microprocessor or a digital signal processor (DSP) can be used in practice to implement some or all of the functions of some or all of the components in the virtual machine creation device according to the embodiments of the present application. The present application can also be implemented as a device or device program for executing part or all of the methods described herein. The program for implementing the present application can be stored on a computer-readable medium, or can be in the form of one or more signals. Such signals can be downloaded from an Internet website, or provided on a carrier signal, or provided in any other form.

[0150] As described above, only the preferred specific embodiments of the present application are provided, but the protection scope of the present application is not limited thereto. Any changes or substitutions that can be easily thought of by those skilled in the technical field of the present application within the technical scope disclosed by the present application should be covered by the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims.

Claims

1. A multi-person pose estimation method based on context instance decoupling, characterized in that, The method includes: Obtaining a preset number of images containing multiple people; Using the images containing multiple people as training samples and inputting them into a multi-person pose estimation model based on context instance decoupling for training; Performing pose estimation on a target image using the trained multi-person pose estimation model based on context instance decoupling; Wherein, the multi-person pose estimation model based on context instance decoupling is provided with an instance information abstraction module, a global feature decoupling module, and a heatmap estimation module; The multi-person pose estimation model based on context instance decoupling further includes a backbone network. Performing pose estimation on a target image using the trained multi-person pose estimation model based on context instance decoupling includes: Inputting the target image into the backbone network to obtain a global feature map, wherein the target image includes multiple people, and the global feature map contains the three-dimensional features of all people; Respectively inputting the global feature map into the instance information abstraction module and the global feature decoupling module; Obtaining the instance features of each person in the target image through the instance information abstraction module; Inputting the instance features of each person into the global feature decoupling module, and the global feature decoupling module decouples an instance feature perception map based on the global feature map and the instance features of each person; Inputting the instance feature perception map into the heatmap estimation module to obtain the probability distribution of each key point of each person in the target image.

2. The method for multi-person pose estimation based on context instance decoupling according to claim 1, wherein, The obtaining the instance features of each person in the target image through the instance information abstraction module includes: Inputting the global feature map into the heatmap module; Extracting the center point coordinates of each person; Sampling at the corresponding positions in the global feature map according to the center point coordinates of each person to obtain the instance features of each person in the target image.

3. The method for multi-person pose estimation based on context instance decoupling according to claim 2, wherein Before obtaining the instance features of each person in the target image, it further includes: recalibrating the center point features of each person based on spatial attention or channel attention.

4. The multi-person pose estimation method based on context instance decoupling according to claim 1, wherein The global feature decoupling module decouples an instance feature perception map based on the global feature map and the instance features of each person, including: Based on the mapping relationship between the instance features of each person and the global feature map, recalibrating the instance features of each person in the spatial dimension to obtain a first instance feature perception map; Based on the mapping relationship between the instance features of each person and the global feature map, recalibrating the instance features of each person in the channel dimension to obtain a second instance feature perception map; Fusing the first instance feature perception map and the second instance feature perception map to obtain an instance feature perception map.

5. The multi-person pose estimation method based on context instance decoupling according to claim 4, wherein Based on the mapping relationship between the instance features of each person and the global feature map, recalibrating the instance features of each person in the spatial dimension to obtain a first instance feature perception map, including: Generating a spatial mask for each person in the global feature map to represent the weight of the foreground feature of each person; Increasing the weight of the foreground feature and recalibrating the spatial position in the instance features of each person to obtain a first instance feature perception map.

6. The multi-person pose estimation method based on context instance decoupling according to claim 5, wherein Decoupling different channels includes: reweighting the global feature map based on person features in the channel dimension to generate a second instance feature perception map.

7. The method for multi-person pose estimation based on context instance decoupling according to claim 6, wherein Fusing the first instance feature perception map and the second instance feature perception map to obtain an instance feature perception map includes: obtaining the instance feature perception map by performing a weighted sum of the first instance feature perception map and the second instance feature perception map.

8. The multi-person pose estimation method based on context instance decoupling according to claim 1, characterized in that, Inputting the instance feature perception map into a heatmap estimation module to obtain the probability distribution of each key point of each person in the target image, including: Inputting the instance feature perception map into a heatmap estimation module to obtain the heatmap corresponding to each person in the target image; Wherein, the heatmap corresponding to each person in the target image contains the probability distribution of each key point.

9. The multi-person pose estimation method based on context instance decoupling according to claim 1, wherein Training in the multi-person pose estimation model based on context instance decoupling includes training the multi-person pose estimation model based on context instance decoupling through a preset loss function, and the preset loss function is: Among them, represents the loss of the instance information abstraction module, represents the loss of the global feature decoupling module, and λ represents the weight.

10. A multi-person pose estimation device based on context instance decoupling, characterized in that, The device includes: An acquisition module, configured to acquire a preset number of images containing multiple people; A training module, configured to use the images containing multiple people as training samples and input them into the multi-person pose estimation model based on context instance decoupling for training; A pose estimation module, configured to use the trained multi-person pose estimation model based on context instance decoupling to perform pose estimation on a target image; Wherein, the multi-person pose estimation model based on context instance decoupling is provided with an instance information abstraction module, a global feature decoupling module, and a heatmap estimation module; The multi-person pose estimation model based on context instance decoupling further includes a backbone network. Using the trained multi-person pose estimation model based on context instance decoupling to perform pose estimation on a target image includes: Inputting the target image into the backbone network to obtain a global feature map, wherein the target image includes multiple people, and the global feature map contains the three-dimensional features of all people; Inputting the global feature map into the instance information abstraction module and the global feature decoupling module respectively; Obtaining the instance features of each person in the target image through the instance information abstraction module; Inputting the instance features of each person into the global feature decoupling module, and the global feature decoupling module decouples an instance feature perception map based on the global feature map and the instance features of each person; Inputting the instance feature perception map into the heatmap estimation module to obtain the probability distribution of each key point of each person in the target image.

11. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1-9.

12. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1-9.

Citation Information

Patent Citations

  • Multi-person posture estimation method based on global information integration

    CN110135375A

  • Human body action posture intelligent estimation method and device based on convolutional neural network

    CN112052886A