Multi-person Pose Estimation Method, Apparatus, Electronic Device, and Machine-readable Storage Medium
By generating dynamic convolution kernel weights for each human body in the input image for convolution processing, the problem of excessive calculation volume of traditional multi-person pose estimation is solved, and the calculation efficiency and accuracy are improved.
Patent Information
- Application Number
- CN202111521832.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-12-13
- Publication Date
- 2025-07-29
- Estimated Expiration
- 2041-12-13
AI Technical Summary
The calculation volume of traditional multi-person pose estimation scheme is too large, and the calculation complexity increases linearly with the increase in the number of targets.
By determining the positions of each human body in the input image, generating corresponding convolution kernel weights, and using these weights for convolution processing, the convolution kernel is dynamically adjusted to estimate the position of the pose point, avoiding the use of the same fixed convolution kernel for each target.
The calculation amount of multi-person pose estimation is reduced, and the calculation efficiency is improved. Especially when the number of targets increases, the calculation amount does not increase significantly, which improves the accuracy and efficiency of the estimation.
Smart Images

Figure CN114333050B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of computer vision technology, and particularly to a multi-person pose estimation method, apparatus, electronic device, and machine-readable storage medium. Background Art
[0002] Human pose estimation, abbreviated as pose estimation, is an important task in computer vision and an essential step for a computer to understand human actions and behaviors.
[0003] Human pose estimation mainly predicts the key points of the human body, that is, predicts the position coordinates of each key point of the human body, and then determines the spatial position relationship between the key points based on prior knowledge to obtain the predicted human pose.
[0004] However, it is found in practice that in traditional human pose estimation schemes, for the problem of multi-person pose estimation, the detection box of each target is first obtained through an independent detection network, then single-person pose estimation is performed on each target in turn, and finally the pose estimation result is finely adjusted according to the occlusion situation of the target. Therefore, the entire network has a large amount of computation, and the computational complexity increases linearly with the increase in the number of targets. Summary of the Invention
[0005] In view of this, this application provides a multi-person pose estimation method, apparatus, electronic device, and machine-readable storage medium to at least solve the problem of excessive computation in traditional multi-person pose estimation schemes.
[0006] Specifically, this application is implemented through the following technical solutions:
[0007] According to the first aspect of the embodiments of this application, a multi-person pose estimation method is provided, including:
[0008] Determine the human positions of each human body in the input image;
[0009] According to the human positions of each human body, generate corresponding convolution kernel weights for each human body in the input image;
[0010] According to the convolution kernel weights corresponding to each human body in the input image, through convolution processing, obtain the pose point positions of each human body in the input image.
[0011] According to the second aspect of the embodiments of this application, a multi-person pose estimation apparatus is provided, including:
[0012] A determination unit, configured to determine the human positions of each human body in the input image;
[0013] A generation unit, configured to generate corresponding convolution kernel weights for each human body in the input image according to the human positions of each human body;
[0014] A pose estimation unit, configured to obtain the position of the pose points of each person in the input image through convolution processing according to the convolution kernel weights corresponding to each person in the input image.
[0015] According to a third aspect of the embodiments of the present application, there is provided an electronic device, including a processor and a memory, where the memory stores machine-executable instructions that can be executed by the processor, and the processor is configured to execute the machine-executable instructions to implement the multi-person pose estimation method provided in the first aspect.
[0016] According to a fourth aspect of the embodiments of the present application, there is provided a machine-readable storage medium, where machine-executable instructions are stored in the machine-readable storage medium, and when the machine-executable instructions are executed by a processor, the multi-person pose estimation method provided in the first aspect is implemented.
[0017] The technical solution provided by the present application can at least bring the following beneficial effects:
[0018] By determining the human body positions of each person in the input image, and respectively generating corresponding convolution kernel weights for each person in the input image according to the human body positions of each person, furthermore, the position of the pose points of each person in the input image can be obtained through convolution processing according to the convolution kernel weights corresponding to each person in the input image. By generating different convolution kernel weights for different people, when the number of targets increases, the amount of calculation will not increase significantly, reducing the amount of calculation for multi-person pose estimation. BRIEF DESCRIPTION OF THE DRAWINGS
[0019] Figure 1 is a schematic flowchart of a multi-person pose estimation method shown in an exemplary embodiment of the present application;
[0020] Figure 2 is a schematic flowchart of obtaining the position of the pose points of each person in the input image shown in an exemplary embodiment of the present application;
[0021] Figure 3 is a schematic diagram of the offset of the feature points in the target area to the true position of the pose points shown in an exemplary embodiment of the present application;
[0022] Figure 4 is a schematic diagram of a multi-person pose estimation algorithm framework shown in an exemplary embodiment of the present application;
[0023] Figure 5 is a schematic structural diagram of a multi-person pose estimation device shown in an exemplary embodiment of the present application;
[0024] Figure 6 is a schematic hardware structure diagram of an electronic device shown in an exemplary embodiment of the present application. Detailed Implementation Manner
[0025] Here, exemplary embodiments will be described in detail, and examples thereof are shown in the accompanying drawings. When the following description refers to the accompanying drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with the present application. On the contrary, they are merely examples of devices and methods consistent with some aspects of the present application as detailed in the appended claims.
[0026] The terms used in the present application are only for the purpose of describing specific embodiments and are not intended to limit the present application. The singular forms "a", "the", and "said" used in the present application and the appended claims are also intended to include the plural forms unless the context clearly indicates otherwise.
[0027] To enable those skilled in the art to better understand the technical solutions provided by the embodiments of the present application and to make the above-mentioned objects, features, and advantages of the embodiments of the present application more apparent and understandable, the technical solutions in the embodiments of the present application will be further described in detail below with reference to the accompanying drawings.
[0028] Please refer to Figure 1 , which is a schematic flowchart of a multi-person pose estimation method provided by an embodiment of the present application. As Figure 1 shown, the multi-person pose estimation method may include the following steps:
[0029] Step S100: Determine the human body positions of each human body in the input image.
[0030] Step S110: Generate corresponding convolution kernel weights for each human body in the input image according to the human body positions of each human body.
[0031] In the embodiments of the present application, in order to reduce the computational complexity of determining the positions of human body pose points in a multi-person pose estimation scenario, when determining the positions of human body pose points of different human bodies in the same input image, it is no longer necessary to use the same fixed convolution kernel for each human body to determine the positions of human body pose points. Instead, the convolution kernel used for determining the positions of human body pose points can be dynamically adjusted according to the human body positions of different human bodies, so that when the number of targets increases, the computational complexity will not increase significantly.
[0032] Exemplarily, the human body positions of each human body in the input image can be determined by means of object detection, and corresponding convolution kernel weights can be generated for each human body in the input image according to the human body positions of each human body, so that in the subsequent process of determining the positions of human body pose points of each human body, the convolution kernel used for determining the positions of human body pose points can be dynamically adjusted according to the convolution kernel weights corresponding to each human body.
[0033] Step S120: Based on the convolution kernel weights corresponding to each human body in the input image, through convolution processing, obtain the position of the pose points of each human body in the input image.
[0034] In the embodiments of the present application, when the convolution kernel weights corresponding to each human body in the input image are determined in the above manner, based on the convolution kernel weights corresponding to each human body in the input image, through convolution processing, obtain the position of the pose points of each human body in the input image.
[0035] Exemplarily, for any human body, when determining the position of the pose points of the human body through convolution processing, the convolution kernel can be dynamically adjusted according to the convolution kernel weights corresponding to the human body, and convolution processing is performed according to the adjusted convolution kernel.
[0036] It can be seen that in Figure 1 the method flow shown, by determining the human body positions of each human body in the input image, and based on the human body positions of each human body, generating corresponding convolution kernel weights for each human body in the input image respectively. Furthermore, based on the convolution kernel weights corresponding to each human body in the input image, through convolution processing, obtain the position of the pose points of each human body in the input image. By generating different convolution kernel weights for different human bodies, when the number of targets increases, the computational amount will not increase significantly, reducing the computational amount of multi-person pose estimation.
[0037] In some embodiments, as Figure 2 shown, in step S120, based on the convolution kernel weights corresponding to each human body in the input image, through convolution processing, obtain the position of the pose points of each human body in the input image, which can be achieved through the following steps:
[0038] Step S121: Use the pose estimation model to perform convolution processing on the target feature map based on the convolution kernel weights corresponding to each human body in the input image, to obtain the initial position of the pose points of each human body in the target feature map of the input image; the target feature map is obtained by the feature extraction layer of the pose estimation model performing feature extraction on the input image;
[0039] Step S122: Fine-tune the initial position of the pose points in the target feature map, and output the final position of the pose points of each human body in the target feature map of the input image.
[0040] Exemplarily, the feature extraction layer of the pose estimation model can be used to perform feature extraction on the input features to obtain the feature map corresponding to the input image (referred to as the target feature map in this article), and use the pose estimation model to perform convolution processing on the target feature map based on the convolution kernel weights corresponding to each human body in the input image determined in the above manner, to obtain the position of the pose points of each human body in the target feature map of the input image (which can be called the initial position of the pose points).
[0041] Exemplarily, considering that there are usually certain downsampling errors in the pose point positions obtained on the feature map (i.e., the initial pose point positions), in order to compensate for the trade-off errors caused by downsampling and improve the accuracy of human pose estimation, the initial pose point positions in the obtained target feature map can be fine-tuned to obtain the final pose point positions of each human body in the input image in the target feature map.
[0042] In one example, in step S121, according to the convolutional kernel weights corresponding to each human body in the input image, performing convolutional processing on the target feature map to obtain the initial pose point positions of each human body in the input image in the target feature map may include:
[0043] Performing at least one convolutional processing on the target feature map to obtain a key point feature map;
[0044] For any human body in the input image, according to the convolutional kernel weight corresponding to this human body, performing convolutional processing on the key point feature map to obtain the initial pose point position of this human body in the target feature map.
[0045] Exemplarily, when the convolutional kernel weights corresponding to different human bodies in the input image are obtained in the above manner, the initial pose point positions of each human body in the target feature map can be obtained respectively according to the convolutional kernel weights corresponding to each human body and the target feature map through convolutional processing.
[0046] Exemplarily, at least one convolutional processing can be performed on the target feature map to obtain a key point feature map.
[0047] Exemplarily, the key points may include some or all of the key points for human pose estimation such as the nose, eyes, shoulders, elbows, and knees.
[0048] For any human body in the input image, according to the convolutional kernel weight corresponding to this human body, performing convolutional processing on the key point feature map to obtain the initial pose point position of this human body in the target feature map.
[0049] In one example, in step S122, fine-tuning the initial pose point positions in the target feature map and outputting the final pose point positions of each human body in the input image in the target feature map may include:
[0050] Using the short-range offset branch of the pose estimation model to determine the offsets corresponding to each feature point in the target feature map;
[0051] For any initial pose point of any human body in any input image, according to the offset corresponding to the initial position of this pose point in the target feature map, fine-tuning the initial position of this pose point to obtain the final position of this pose point.
[0052] Exemplarily, to improve the efficiency and accuracy of pose point position adjustment, a short-range offset branch can be added to the pose estimation model to fine-tune the initial position of the pose points determined by the pose estimation model.
[0053] During the training process of the pose estimation model, the short-range offset branch can be trained, and the short-range offset branch can output the offsets of some or all of the feature points in the target feature map relative to the human pose points.
[0054] For example, as Figure 3 shown, during the training process of the pose estimation model, for any pre-annotated human pose point (such as the hollow point in Figure 3 , and the hollow point is the ground truth point), it can be determined that in the target feature map, for each feature point within the area with a preset radius centered on this pose point (which can be called the target area) (such as the solid points within the virtual circle shown in Figure 3 ), the offset relative to this pose point is learned by the short-range offset branch in the pose estimation model. Furthermore, after the pose estimation model is trained, based on the short-range offset branch, the offset of the predicted initial position of the pose point (for the ground truth of the pose point shown in Figure 3 , the predicted initial position of the pose point is usually one of the solid points within the virtual circle) relative to the true position of the pose point can be determined.
[0055] To improve the model training and prediction efficiency, a short-range offset branch can be added to the pose estimation model to fine-tune the initial position of the pose points determined by the pose estimation model. Thus, during the model training and the process of predicting the human pose point positions, it can be carried out in an end-to-end manner, and the model directly outputs the adjusted positions of the human pose points of each person, without the need to fine-tune the pose point positions through other specialized networks after obtaining the pose point positions.
[0056] In some embodiments, in step S100, determining the human positions of each person in the input image may include:
[0057] Using the feature extraction layer of the pose estimation model to extract features from the input image to obtain the feature map corresponding to the input image;
[0058] Using the classification branch of the pose estimation model to classify each feature point in the feature map to obtain the human positions of each person in the input image.
[0059] Exemplarily, the pose estimation model may further include a classification branch for distinguishing feature points corresponding to different humans, and further, for distinguishing the positions of different humans in a multi-person pose estimation scenario.
[0060] Correspondingly, for the input image, the feature extraction layer of the pose estimation model can be used to extract features from it to obtain the feature map corresponding to the input image, and the classification branch of the pose estimation model can be used to classify each feature point in the feature map to obtain the human positions of each human in the input image, that is, the positions of different humans can be directly distinguished through the pose estimation model, without the need to distinguish the pose points through post-processing methods such as clustering after feature extraction, improving the efficiency of human position determination.
[0061] In one example, the above-mentioned extraction of features from the input image to obtain the feature map corresponding to the input image may include:
[0062] Extract features from the input image to obtain the multi-scale features of the input image; the multi-scale features include N1 feature maps of different scales; N1 is an integer greater than 1;
[0063] According to the multi-scale features of the input image, through feature fusion processing, N2 fused feature maps of different scales are obtained; N2 is an integer greater than 1.
[0064] Exemplarily, in order to improve the accuracy of human pose estimation, the multi-scale features of the input image can be obtained for multi-scale feature fusion in subsequent processes to obtain the fused feature map, and human pose estimation is performed based on the fused feature map.
[0065] Exemplarily, feature maps of different scales can be obtained by means of downsampling.
[0066] Exemplarily, the downsampling factor can be 2n, where n is a positive integer.
[0067] Exemplarily, the multi-scale features may include, but are not limited to, feature maps of multiple (denoted as N1 in this article) scales among scales such as 4 times, 8 times, 16 times, 32 times, 64 times, etc.
[0068] For the obtained multi-scale features of the input image, multiple (denoted as N2 in this article) fused feature maps of different scales can be obtained through feature fusion processing.
[0069] Exemplarily, when obtaining fused feature maps of different scales through feature fusion processing, new fused feature maps can also be obtained by downsampling the fused feature maps.
[0070] For example, assuming that the multi-scale features in step S100 include an 8x feature map, a 16x feature map, and a 32x feature map, then through feature fusion processing, such as using FPN (Feature Pyramid Network), an 8x fused feature map, a 16x fused feature map, and a 32x fused feature can be obtained, and by downsampling the 32x fused feature map, a 64x fused feature map and a 128x fused feature map can be obtained.
[0071] Exemplarily, N1 can be greater than N2, less than N2, or equal to N2.
[0072] Exemplarily, for the N2 different-scale fused feature maps obtained in the above manner, one feature map (i.e., the above-mentioned target feature map) can be selected for human pose estimation.
[0073] Exemplarily, in order to balance the semantic information representation ability and the geometric detail information representation ability, the target feature map can be an 8x fused feature map.
[0074] As an example, the above-mentioned classification of each feature point in the feature map to obtain the human positions of each human body in the input image may include:
[0075] Classifying each feature point in the N2 different-scale fused feature maps respectively to obtain the human positions in the N2 different-scale fused feature maps;
[0076] According to the NMS algorithm, filtering the human positions in the N2 different-scale fused feature maps to obtain the filtered human positions.
[0077] Exemplarily, considering that the sizes of the human bodies in the input image may vary. For example, a human body farther from the image acquisition device will be relatively smaller, while a human body closer to the image acquisition device will be relatively larger.
[0078] In addition, due to the differences in the semantic information representation ability and the geometric detail information representation ability of feature maps at different scales, the semantic information representation ability of high-scale feature maps is relatively strong, but the geometric detail information representation ability is relatively weak, and the semantic information representation ability of low-scale feature maps is relatively weak, but the geometric detail information representation ability is relatively strong.
[0079] Therefore, in order to improve the accuracy of human detection, each human body in the input image can be detected from the fused feature maps at different scales respectively.
[0080] Exemplarily, the human position of a larger human body can be determined based on the fused feature map of a higher level (smaller scale), and the human position of a smaller human body can be determined based on the fused feature map of a lower level (larger scale).
[0081] In addition, considering that the same human body may be detected in the fused feature maps at different scales, in order to avoid performing human pose estimation multiple times for the same human body, for the human body positions obtained by performing human body detection on the fused feature maps of N2 different scales, the human body positions detected in the fused feature maps of the N2 different scales can be filtered according to the NMS (Non-Maximum Suppression) algorithm to filter out duplicate human body positions and obtain the filtered human body positions.
[0082] As an example, generating corresponding convolution kernel weights for each human body according to the human body positions of each human body may include:
[0083] For any one of the fused feature maps of the N2 different scales, generating corresponding convolution kernel weights according to the filtered human body positions in the fused feature map.
[0084] Exemplarily, after filtering the human body positions in each fused feature map in the above manner, for any one of the fused feature maps of the N2 different scales, corresponding convolution kernel weights can be generated according to the filtered human body positions in the fused feature map.
[0085] As an example, generating corresponding convolution kernel weights according to the filtered human body positions in the fused feature map may include:
[0086] Extracting the eigenvalue of the feature point of the filtered human body position in the fused feature map, and using the eigenvalue as the convolution kernel weight of the human body corresponding to the human body position.
[0087] Exemplarily, for any one of the fused feature maps of the N2 different scales of feature, corresponding convolution kernel weights can be extracted from the fused feature map according to the filtered human body positions in the fused feature map, the eigenvalue of the feature point of the filtered human body position in the fused feature map is extracted, and the eigenvalue is used as the convolution kernel weight of the human body corresponding to the human body position. Furthermore, according to the convolution kernel weight of the human body, through convolution processing, the position of the pose point corresponding to the human body can be obtained, and dynamic convolution processing for estimating pose points using different convolution kernels for different human bodies can be realized.
[0088] To enable those skilled in the art to better understand the technical solutions provided in the embodiments of the present application, the technical solutions provided in the embodiments of the present application will be described below in conjunction with specific embodiments.
[0089] Please refer to Figure 4 , which is a schematic diagram of a multi-person pose estimation algorithm framework provided in the embodiments of the present application. As shown in Figure 4As shown, after the input image enters the algorithm framework, the backbone network can output feature maps with 8x, 16x, and 32x downsampling (corresponding to C3 - C5 in Figure 4 respectively), that is, the above N1 = 3.
[0090] Exemplarily, the backbone network can be a Residual Network (abbreviated as ResNet) or a High-Resolution Network (abbreviated as HRNet).
[0091] For Figure 4 the C3 - C5 output by the backbone network in, it can be input into a scale fusion structure, such as an FPN structure. The FPN structure outputs fused features P3 - P7, which are respectively 8x fused feature maps, 16x fused feature maps,..., 128x fused feature maps, that is, the above N2 = 5; among them, P6 can be obtained by performing 2x downsampling on P5, and P7 can be obtained by performing 2x downsampling on P6.
[0092] In Figure 4 the shown algorithm framework, in addition to the branches for classification and box regression in the head, a branch for dynamic convolution kernel prediction (Dynamic Filters) can also be added.
[0093] For any scale of the fused feature maps in P3 - P7, on the one hand, the human body detection can be performed on the fused feature map respectively through the classification branch and the box regression branch in the head, and the human body position can be determined. Based on the human body position in the fused feature map, the dynamic convolution kernel prediction is performed through the branch for dynamic convolution kernel prediction to obtain the convolution kernel weights corresponding to each human body.
[0094] In this embodiment, considering that the same human body may be detected in the fused feature maps of different scales, therefore, in order to improve the accuracy and efficiency of human pose estimation, the human body positions detected in P3 - P7 can be filtered according to the NMS algorithm to obtain the filtered human body positions.
[0095] On the other hand, by performing convolution processing on P3 (taking the above target feature map as P3 as an example), a key point feature map (Keyoint Feature) can be obtained.
[0096] Exemplarily, convolution processing can be performed on P3 to obtain a key point feature map, that is, input P3 into KP-Net to obtain the corresponding key point feature map.
[0097] In this embodiment, when the convolutional kernel weights corresponding to each human body and the above-mentioned keypoint feature map are obtained, the convolutional kernel weights corresponding to each human body (one human body can be called an instance) can be used to perform convolutional processing (which can be called dynamic convolutional processing) on the keypoint feature map to obtain the pose point positions of each human body on the 8-fold downsampled feature map (corresponding to Figure 4 the Keypoint Map in it, which can be called the original pose point positions).
[0098] Exemplarily, one human body (i.e., one instance) can correspond to one Keypoint Map.
[0099] Since the pose estimation task is very sensitive to position, there is a certain downsampling error in the pose point positions predicted by the network on the 8-fold downsampled map. To compensate for the trade-off error caused by downsampling, a short-range offset prediction branch is added to the P3 feature map. Through this short-range offset prediction branch, the original pose points are fine-tuned in position.
[0100] Exemplarily, during the training process, for any original pose point P i , the original pose point P i =(x i , y i ) can be mapped to the 8-fold downsampled map to obtain the mapped point P i ′=(x i / 8, y i / 8). Then, with P i ′ as the center point, a circular region (i.e., the above-mentioned target region) with a radius of R (i.e., the above-mentioned preset radius) is constructed. The feature points falling within this circular region need to determine their relative displacements (i.e., offsets) to P i ′. Thus, the short-range offset prediction branch can learn the offsets of each feature point within the above-mentioned target region to the true position of the pose point.
[0101] Correspondingly, when performing multi-person pose estimation, for the initial pose point positions obtained in the above manner, the short-range offset prediction branch can be used to determine the offset of the initial pose point position relative to the true position of the pose point, and the initial pose point position is fine-tuned according to this offset to obtain the final pose point position.
[0102] In this embodiment, during testing, for the human body positions in each fused feature map processed according to the NMS algorithm, the weights of the convolution kernels can be extracted, and the convolution operation output can be obtained on the 8-fold downsampled feature map (i.e., the above-mentioned Keypoint Feature). After passing through softmax (regression model), the initial positions of the pose points can be obtained, and then the obtained initial pose point positions can be fine-tuned through the short-range offset branch to obtain the final pose estimation result.
[0103] It can be seen that the embodiment of the present application provides an efficient pose estimation scheme based on dynamic convolution, which belongs to the one-stage algorithm. The model structure is simple, the hardware execution efficiency is high, and compared with the top-down pose estimation scheme in the traditional scheme, the calculation amount is significantly reduced; compared with the bottom-up pose estimation scheme, there is no need to perform post-processing of key point grouping.
[0104] In addition, the multi-person pose estimation scheme provided by the embodiment of the present application can achieve end-to-end training optimization and has high performance.
[0105] The method provided by the present application has been described above. Next, the device provided by the present application will be described:
[0106] Please refer to Figure 5 , which is a schematic structural diagram of a multi-person pose estimation device provided by an embodiment of the present application. As Figure 5 shown, the multi-person pose estimation device may include:
[0107] A determination unit 510, configured to determine the human body positions of each human body in the input image;
[0108] A generation unit 520, configured to generate corresponding convolution kernel weights for each human body in the input image according to the human body positions of each human body;
[0109] A pose estimation unit 530, configured to obtain the pose point positions of each human body in the input image through convolution processing according to the convolution kernel weights corresponding to each human body in the input image.
[0110] In some embodiments, the pose estimation unit 530 obtains the pose point positions of each human body in the input image through convolution processing according to the convolution kernel weights corresponding to each human body in the input image, including:
[0111] Using a pose estimation model, performing convolution processing on a target feature map according to the convolution kernel weights corresponding to each human body in the input image to obtain the initial pose point positions of each human body in the input image in the target feature map; the target feature map is obtained by the feature extraction layer of the pose estimation model performing feature extraction on the input image;
[0112] Fine-tune the initial positions of the pose points in the target feature map, and output the final positions of the pose points of each human body in the input image in the target feature map.
[0113] In some embodiments, the pose estimation unit 530 performs convolution processing on the target feature map according to the convolution kernel weights corresponding to each human body in the input image to obtain the initial positions of the pose points of each human body in the input image in the target feature map, including:
[0114] Perform at least one convolution processing on the target feature map to obtain a key point feature map;
[0115] For any human body in the input image, perform convolution processing on the key point feature map according to the convolution kernel weight corresponding to the human body to obtain the initial position of the pose point of the human body in the target feature map.
[0116] In some embodiments, the pose estimation unit 530 fine-tunes the initial pose point positions in the target feature map and outputs the final positions of the pose points of each human body in the input image in the target feature map, including:
[0117] Use the short-range offset branch of the pose estimation model to determine the offset corresponding to each feature point in the target feature map;
[0118] For any initial pose point of any human body in any input image, fine-tune the initial position of the pose point according to the offset corresponding to the initial position of the pose point in the target feature map to obtain the final position of the pose point.
[0119] In some embodiments, the determination unit 510 determines the human body positions of each human body in the input image, including:
[0120] Use the feature extraction layer of the pose estimation model to extract features from the input image to obtain the feature map corresponding to the input image;
[0121] Use the classification branch of the pose estimation model to classify each feature point in the feature map to obtain the human body positions of each human body in the input image.
[0122] In some embodiments, the determination unit 510 extracts features from the input image to obtain the feature map corresponding to the input image, including:
[0123] Extract features from the input image to obtain the multi-scale features of the input image; the multi-scale features include N1 feature maps of different scales; N1 is an integer greater than 1;
[0124] Based on the multi-scale features of the input image, through feature fusion processing, N2 fused feature maps of different scales are obtained; N2 is an integer greater than 1.
[0125] In some embodiments, the determining unit 510 classifies each feature point in the feature map to obtain the human body positions of each human body in the input image, including:
[0126] Classify each feature point in the N2 fused feature maps of different scales respectively to obtain the human body positions in the N2 fused feature maps of different scales;
[0127] According to the non-maximum suppression (NMS) algorithm, filter the human body positions in the N2 fused feature maps of different scales to obtain the filtered human body positions.
[0128] In some embodiments, the generating unit 520 generates corresponding convolution kernel weights for each human body in the input image according to the human body positions of each human body, including:
[0129] For any one of the N2 fused feature maps of different scales, generate corresponding convolution kernel weights according to the filtered human body positions in this fused feature map.
[0130] In some embodiments, the generating unit 520 generates corresponding convolution kernel weights according to the filtered human body positions in this fused feature map, including:
[0131] Extract the feature values of the feature points at the filtered human body positions in the fused feature map, and use this feature value as the convolution kernel weight of the human body corresponding to this human body position.
[0132] An embodiment of the present application provides an electronic device, including a processor and a memory. Among them, the memory stores machine-executable instructions that can be executed by the processor, and the processor is used to execute the machine-executable instructions to implement the multi-person pose estimation method described above.
[0133] Please refer to Figure 6 , which is a schematic hardware structure diagram of an electronic device provided by an embodiment of the present application. The electronic device may include a processor 601 and a memory 602 storing machine-executable instructions. The processor 601 and the memory 602 can communicate via a system bus 603. And by reading and executing the machine-executable instructions corresponding to the multi-person pose estimation logic in the memory 602, the processor 601 can execute the multi-person pose estimation method described above.
[0134] The memory 602 mentioned in this article can be any electronic, magnetic, optical or other physical storage device that can contain or store information such as executable instructions, data, and so on. For example, the machine-readable storage medium can be: RAM (Random Access Memory), volatile memory, non-volatile memory, flash memory, storage drives (such as hard disk drives), solid state drives, any type of storage disk (such as optical discs, DVDs, etc.), or similar storage media, or a combination thereof.
[0135] In some embodiments, a machine-readable storage medium is also provided, such as Figure 6 the memory 602 in [description not provided], and the machine-readable storage medium stores machine-executable instructions that, when executed by a processor, implement the multi-person pose estimation method described above. For example, the storage medium can be ROM, RAM, CD-ROM, magnetic tapes, floppy disks, and optical data storage devices, etc.
[0136] It should be noted that in this article, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article or device comprising a series of elements not only includes those elements, but also includes other elements not expressly listed, or elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "comprising an..." does not exclude the existence of additional identical elements in the process, method, article or device comprising the element.
[0137] The above are only the preferred embodiments of the present application and are not intended to limit the present application. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of the present application shall be included within the scope of protection of the present application.
Claims
1. A multi-person pose estimation method, characterized in that, Including: Using the feature extraction layer of the pose estimation model to extract features from the input image to obtain a feature map corresponding to the input image; wherein, the extracting features from the input image to obtain the feature map corresponding to the input image includes: extracting features from the input image to obtain multi-scale features of the input image; the multi-scale features include N1 feature maps of different scales; N1 is an integer greater than 1; according to the multi-scale features of the input image, through feature fusion processing, N2 fused feature maps of different scales are obtained; N2 is an integer greater than 1; Using the classification branch of the pose estimation model to classify each feature point in the feature map to obtain the human body positions of each human body in the input image; Generating corresponding convolution kernel weights for each human body in the input image according to the human body positions of each human body; According to the convolution kernel weights corresponding to each human body in the input image, through convolution processing, obtaining the pose point positions of each human body in the input image.
2. The method according to claim 1, characterized in that The obtaining the pose point positions of each human body in the input image through convolution processing according to the convolution kernel weights corresponding to each human body in the input image includes: Using the pose estimation model, according to the convolution kernel weights corresponding to each human body in the input image, performing convolution processing on the target feature map to obtain the initial pose point positions of each human body in the input image in the target feature map; the target feature map is obtained by the feature extraction layer of the pose estimation model extracting features from the input image; Fine-tuning the initial pose point positions in the target feature map and outputting the final pose point positions of each human body in the input image in the target feature map.
3. The method according to claim 2, wherein The performing convolution processing on the target feature map according to the convolution kernel weights corresponding to each human body in the input image to obtain the initial pose point positions of each human body in the input image in the target feature map includes: Performing at least one convolution processing on the target feature map to obtain a key point feature map; For any human body in the input image, performing convolution processing on the key point feature map according to the convolution kernel weight corresponding to the human body to obtain the initial pose point position of the human body in the target feature map.
4. The method according to claim 2, wherein The fine-tuning the initial pose point positions in the target feature map and outputting the final pose point positions of each human body in the input image in the target feature map includes: Using the short-range offset branch of the pose estimation model to determine the offset amount corresponding to each feature point in the target feature map; For any initial pose point of any human body in any input image, fine-tuning the initial pose point position according to the offset amount corresponding to the initial pose point position in the target feature map to obtain the final pose point position of the pose point.
5. The method according to claim 1, characterized in that, The classifying each feature point in the feature map to obtain the human body positions of each human body in the input image includes: Classifying each feature point in the N2 fused feature maps of different scales respectively to obtain the human body positions in the N2 fused feature maps of different scales; According to the non-maximum suppression (NMS) algorithm, filter the human body positions in the N2 fused feature maps of different scales to obtain the filtered human body positions.
6. The method according to claim 5, characterized in that, Generating corresponding convolution kernel weights for each human body in the input image according to the human body positions of each human body respectively includes: For any one of the N2 fused feature maps of different scales, generate corresponding convolution kernel weights according to the filtered human body positions in this fused feature map.
7. The method according to claim 6, wherein Generating corresponding convolution kernel weights according to the filtered human body positions in this fused feature map includes: Extract the feature values of the feature points of the filtered human body positions in the fused feature map, and use this feature value as the convolution kernel weight of the human body corresponding to this human body position.
8. A multi-person pose estimation device, characterized in that, Including: A determination unit, configured to use the feature extraction layer of the pose estimation model to perform feature extraction on the input image to obtain the feature map corresponding to the input image; wherein, performing feature extraction on the input image to obtain the feature map corresponding to the input image includes: performing feature extraction on the input image to obtain the multi-scale features of the input image; the multi-scale features include N1 feature maps of different scales; N1 is an integer greater than 1; according to the multi-scale features of the input image, through feature fusion processing, obtain N2 fused feature maps of different scales; N2 is an integer greater than 1; Use the classification branch of the pose estimation model to classify each feature point in the feature map to obtain the human body positions of each human body in the input image; A generation unit, configured to generate corresponding convolution kernel weights for each human body in the input image according to the human body positions of each human body; A pose estimation unit, configured to obtain the pose point positions of each human body in the input image through convolution processing according to the convolution kernel weights corresponding to each human body in the input image.
9. The device according to claim 8, characterized in that, The pose estimation unit obtaining the pose point positions of each human body in the input image through convolution processing according to the convolution kernel weights corresponding to each human body in the input image includes: Using the pose estimation model, perform convolution processing on the target feature map according to the convolution kernel weights corresponding to each human body in the input image to obtain the initial pose point positions of each human body in the input image in the target feature map; the target feature map is obtained by performing feature extraction on the input image by the feature extraction layer of the pose estimation model; Fine-tune the initial pose point positions in the target feature map and output the final pose point positions of each human body in the input image in the target feature map; Among them, the pose estimation unit performing convolution processing on the target feature map according to the convolution kernel weights corresponding to each human body in the input image to obtain the initial pose point positions of each human body in the input image in the target feature map includes: Perform at least one convolution processing on the target feature map to obtain a key point feature map; For any one human body in the input image, perform convolution processing on the key point feature map according to the convolution kernel weight corresponding to this human body to obtain the initial pose point position of this human body in the target feature map. Among them, the pose estimation unit fine-tunes the positions of the initial pose points in the target feature map and outputs the final positions of the pose points of each human body in the input image in the target feature map, including: Determining the offset corresponding to each feature point in the target feature map by using the short-range offset branch of the pose estimation model; For any initial pose point of any human body in any of the input images, fine-tuning the initial position of the pose point according to the offset corresponding to the initial position of the pose point in the target feature map to obtain the final position of the pose point; and / or The determination unit classifies each feature point in the feature map to obtain the human body positions of each human body in the input image, including: Classifying each feature point in the N2 fusion feature maps of different scales respectively to obtain the human body positions in the N2 fusion feature maps of different scales; Filtering the human body positions in the N2 fusion feature maps of different scales according to the non-maximum suppression (NMS) algorithm to obtain the filtered human body positions; Among them, the generation unit generates corresponding convolution kernel weights for each human body in the input image according to the human body positions of each human body, including: For any one of the N2 fusion feature maps of different scales, generating corresponding convolution kernel weights according to the filtered human body positions in the fusion feature map; Among them, the generation unit generates corresponding convolution kernel weights according to the filtered human body positions in the fusion feature map, including: Extracting the eigenvalues of the feature points at the filtered human body positions in the fusion feature map, and using the eigenvalue as the convolution kernel weight of the human body corresponding to the human body position.
10. An electronic device, characterized in that, It includes a processor and a memory. The memory stores machine-executable instructions that can be executed by the processor, and the processor is used to execute the machine-executable instructions to implement the method according to any one of claims 1-7.
11. A machine-readable storage medium, characterized in that, The machine-readable storage medium stores machine-executable instructions, and when the machine-executable instructions are executed by the processor, the method according to any one of claims 1-7 is implemented.