Multi-person 3D posture estimation method, device and electronic equipment
The feature map is extracted through the backbone network and combined with the two-dimensional pose regression model and the depth regression model for multi-stage regression calculation, the problem of unclear depth information of three-dimensional human body pose estimation in a single image is solved, and more accurate three-dimensional pose estimation results are achieved.
Patent Information
- Application Number
- CN202111593360.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-12-23
- Publication Date
- 2025-05-09
- Estimated Expiration
- 2041-12-23
AI Technical Summary
In the prior art, when using a single image to estimate the three-dimensional human body posture of multiple people, the depth information is not clear enough, resulting in inaccurate estimation results and inability to effectively perceive advanced features such as human body scale and background position.
A multi-person three-dimensional pose estimation method is proposed. The feature map is extracted through the backbone network, combined with the two-dimensional pose regression model and the depth regression model, and the multi-stage regression calculation is performed, and the three-dimensional pose estimation result is finally calculated.
By combining the two-dimensional pose guidance features with the features in depth regression, the accuracy of the depth regression results is improved, making the results of multi-person three-dimensional pose estimation more accurate.
Smart Images

Figure CN114550282B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the technical field of posture estimation, and in particular to a method, device and electronic equipment for multi-person three-dimensional posture estimation. Background Art
[0002] Human body pose is generally represented in the form of a skeleton. The task of estimating the 3D human pose of a single person from a single image has made significant progress. A more realistic and challenging task has attracted more and more attention, namely: estimating the 3D human pose of multiple people from a single image.
[0003] At present, the application of single-stage regression models to achieve multi-person 3D human pose estimation has the following problems in real-world scenarios: Since the depth information of a single image is not clear enough, the estimation result obtained by using a single image to achieve 3D human pose estimation is not accurate. Using absolute depth for supervision cannot perceive high-level features such as human scale and background position. Summary of the invention
[0004] In view of this, the purpose of the present application is to propose a multi-person three-dimensional posture estimation method, device and electronic device to solve or partially solve the above-mentioned technical problems.
[0005] Based on the above objectives, the present application provides a method for multi-person 3D posture estimation, comprising:
[0006] Get multi-person images taken by the camera;
[0007] Extracting features of the multi-person image based on a backbone network to obtain a feature map, wherein the backbone network is a model trained based on a deep learning method;
[0008] Performing regression calculation on the feature map based on a two-dimensional posture regression model to obtain a two-dimensional posture result, wherein the two-dimensional posture regression model is a model trained based on a deep learning method;
[0009] Performing regression calculation on the feature map based on the two-dimensional posture result and the deep regression model to obtain a deep regression result, wherein the deep regression model is a model trained based on a deep learning method;
[0010] A three-dimensional posture estimation result is calculated based on the two-dimensional posture result and the depth regression result.
[0011] Furthermore, the two-dimensional posture result includes: human body key points, human body center key point coordinates and human body key point offset mapping set.
[0012] Furthermore, the performing regression calculation on the feature map based on the two-dimensional posture regression model includes:
[0013] Based on the feature map, each human key point offset map in the human key point offset map set corresponding to each human key point is calculated according to the following formula:
[0014] O1=F1(X1),O2=F2(X2),...,O n =F n (X n )
[0015] Among them, F m is the mth regressor corresponding to the mth human key point, O m is the offset mapping of the mth human key point, X m It is the mth sub-feature map corresponding to the mth human key point obtained by decomposing the feature map, 1≤m≤n, n is the number of offset mappings in the human key point offset mapping set, wherein the regressor is a model trained based on the adaptive convolutional learning model.
[0016] Further, the performing regression calculation on the feature map based on the two-dimensional posture result and the depth regression model to obtain the depth regression result includes:
[0017] Outputting a root depth regression result and a relative depth regression result through the depth regression model based on the feature map;
[0018] Obtaining a two-dimensional posture guidance feature through convolution calculation based on the feature map;
[0019] Calculate a root key point depth regression result based on the two-dimensional posture guidance feature and the root depth regression result;
[0020] The root key point depth regression result and the relative depth regression result are combined to obtain the depth regression result.
[0021] Furthermore, the two-dimensional posture guidance feature is obtained by convolution calculation based on the feature map, including:
[0022] Outputting a predetermined number of decomposed feature maps by convolution calculation based on the feature map;
[0023] Based on each of the decomposed feature graphs, the human body key point offset mapping set is output through its corresponding regressor, wherein the regressor is a model trained based on an adaptive convolutional learning model;
[0024] A two-dimensional posture guidance feature is obtained by adopting a position query operation based on the human body key point offset mapping set.
[0025] Further, the root key point depth regression result is calculated based on the two-dimensional posture guidance feature and the root depth regression result, including:
[0026] The root keypoint depth regression result is calculated according to the following formula:
[0027] Z=F Z (cat{X,D(O1'),D(O2')...D(O n-1 ')})
[0028] Among them, F Z is the root keypoint deep regressor in the deep regression model, cat{} is the connection operation on the channel dimension, X is the feature map, D(O k ') is the position query operation on the kth human key point, O k ' is the kth offset mapping in the human key point offset mapping set, 1≤k≤n-1, n is the number of offset mappings in the human key point offset mapping set, wherein the root key point deep regressor is a model trained based on the adaptive convolutional learning model.
[0029] Furthermore, the adaptive convolutional learning model is obtained based on the following formula:
[0030]
[0031] Among them, y(q) is the output result of the adaptive convolutional learning model, is the offset of the key point of the human body, q is the coordinate of the key point of the center of the human body, W i is the weight of the convolution kernel in the adaptive convolution learning model, and j is the size of the convolution kernel in the adaptive convolution learning model.
[0032] Further, the calculating of the three-dimensional posture estimation result based on the two-dimensional posture result and the depth regression result includes:
[0033] Get the camera intrinsic matrix;
[0034] Calculate the coordinates of the key points of the human body in a two-dimensional coordinate system based on the two-dimensional posture result and the depth regression result;
[0035] The three-dimensional posture estimation result is calculated according to the following formula:
[0036] [X,Y,Z] T =ZK -1 [x,y,1] T
[0037] Among them, [X, Y, Z] is the three-dimensional posture estimation result, [x, y] in [x, y, 1] is the coordinates of the key points of the human body in the two-dimensional coordinate system, K is the camera intrinsic matrix, and Z is the depth regression result.
[0038] Based on the same inventive concept, the present application also provides a multi-person 3D posture estimation device, comprising:
[0039] An image acquisition module is configured to acquire images of multiple people taken by a camera;
[0040] A feature map generation module is configured to extract features from the multi-person image based on a backbone network to obtain a feature map, wherein the backbone network is a model trained based on a deep learning method;
[0041] A two-dimensional posture generation module is configured to perform regression calculation on the feature map based on a two-dimensional posture regression model to obtain a two-dimensional posture result, wherein the two-dimensional posture regression model is a model trained based on a deep learning method;
[0042] A depth regression module is configured to perform regression calculation on the feature map based on the two-dimensional posture result and a depth regression model to obtain a depth regression result, wherein the depth regression model is a model trained based on a deep learning method;
[0043] The three-dimensional estimation module is configured to calculate a three-dimensional posture estimation result based on the two-dimensional posture result and the depth regression result.
[0044] Based on the same inventive concept, the present application also provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the method described above when executing the program.
[0045] From the above, it can be seen that the method, device and electronic device for multi-person 3D posture estimation provided by the present application associate 2D posture regression with depth regression, and no longer need to group key points of the human body, thus simplifying the process of the multi-person 3D posture estimation method. By merging the 2D posture guidance features with the features in the root depth regression, the accuracy of the estimation result of the root absolute depth in the depth regression result is improved, making the result of multi-person 3D posture estimation more accurate. BRIEF DESCRIPTION OF THE DRAWINGS
[0046] In order to more clearly illustrate the technical solutions in the present application or related technologies, the drawings required for use in the embodiments or related technical descriptions are briefly introduced below. Obviously, the drawings described below are only embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.
[0047] Figure 1 A schematic diagram of a flow chart of a method for estimating three-dimensional postures of multiple persons according to an embodiment of the present application;
[0048] Figure 2 A schematic diagram of the process of obtaining the deep regression results of an embodiment of the present application;
[0049] Figure 3 This is a schematic diagram of the structure of a multi-person 3D posture estimation device according to an embodiment of the present application;
[0050] Figure 4 A schematic diagram of the hardware structure of an electronic device according to an embodiment of the present application;
[0051] Figure 5 Schematic diagram of the overall framework of multi-person 3D posture estimation according to an embodiment of the present application. DETAILED DESCRIPTION
[0052] In order to make the objectives, technical solutions and advantages of the present application more clearly understood, the present application is further described in detail below in combination with specific embodiments and with reference to the accompanying drawings.
[0053] It should be noted that, unless otherwise defined, the technical terms or scientific terms used in the embodiments of the present application should be understood by people with ordinary skills in the field to which the present application belongs. The "first", "second" and similar words used in the embodiments of the present application do not represent any order, quantity or importance, but are only used to distinguish different components. "Including" or "comprising" and similar words mean that the elements or objects appearing in front of the word cover the elements or objects listed after the word and their equivalents, without excluding other elements or objects. "Connect" or "connected" and similar words are not limited to physical or mechanical connections, but may include electrical connections, whether direct or indirect.
[0054] As described in the background technology, 3D human posture estimation refers to:
[0055] The purpose of this task is to generate a corresponding image based on a given image of a person, so that it has a three-dimensional spatial posture. The application of three-dimensional human posture estimation is very wide, including human-computer interaction, motion analysis, rehabilitation training, etc. It can also provide information such as bone structure for other computer vision tasks (such as behavior recognition). There are generally two ways to represent the human body: the first is to represent the human body posture in the form of a skeleton, which is composed of a series of human body key points and the lines between the key points; the other is a parameterized human body model, which represents the human body posture and body shape in the form of a network.
[0056] The one-stage method and the two-stage method refer to:
[0057] The single-stage method based on statistical learning methods directly regresses the information of 3D key points from the original image, and can obtain more effective information from image features, such as occlusion. The two-stage method first estimates the 2D human posture, and then derives the corresponding relative 3D human posture based on the obtained 2D human posture results combined with anthropological constraints or statistical data. The two-stage method can be further divided into template-based methods based on statistical learning methods and optimization methods.
[0058] The top-down approach and bottom-up approach refer to:
[0059] The top-down method is to first detect the human body in the image, obtain the detection frame of each person, and then crop the area within the detection frame and send it to the next posture estimation network. The advantage of the top-down method is that it can ensure that the number of posture points to be detected by the posture estimation network is fixed and all belong to the same person; the bottom-up method is to directly estimate the posture of the input image, and then allocate the detected posture key points through a certain strategy. The advantage of the bottom-up method is that the posture estimation network can better obtain global information.
[0060] The embodiments of the present application are described in detail below with reference to the accompanying drawings.
[0061] This application provides a method for multi-person 3D posture estimation. Figure 1 , including the following steps:
[0062] Step S101: Acquire a multi-person image taken by a camera.
[0063] Step S102, extract features from the multi-person image based on the backbone network to obtain a feature map, wherein the backbone network is a model trained based on a deep learning method. HRNet (Deep High-ResolutionRepresentation Learning for Human Pose Estimation) is used as the backbone network. HRNet was proposed by Microsoft Research. It changes the traditional serial connection of high and low resolution convolutions to parallel connection of high and low resolution convolutions. It learns rich high-resolution representations by maintaining high resolution throughout the process and exchanging information between high and low resolution representations multiple times, and achieves the best performance in the human pose estimation task of multiple data sets. During the entire feature map extraction process, HRNet gradually adds low-resolution feature map sub-networks to the high-resolution feature map main network in parallel, and different networks realize multi-scale fusion and feature extraction to maintain the high resolution of the feature map, so that the human feature map obtained by the final network output is more accurate. The fixed-size image I is input into the backbone network HRNet, and the human feature map X is predicted by the output of HRNet. The human feature map X output by the backbone network is the human key point features extracted by HRNet, which is stored in the form of a tensor to provide information for further accurate prediction of the position and depth of human key points. Its scale is B*C*H*W, where B is the batch size, that is, the number of input images at a time, which can be set to 1 during the test process, C is the number of channels of the feature map, and H and W are the length and width of the feature map, respectively. The feature map can be used as the data basis for subsequent two-dimensional posture regression and depth regression.
[0064] Step S103, based on the two-dimensional posture regression model, regression calculation is performed on the multi-person image to obtain a two-dimensional posture result, wherein the two-dimensional posture regression model is a model trained based on a deep learning method. The network structure refers to DEKR (DEKR: Bottom-Up Human Pose Estimation Via Disentangled KeypointRegression) proposed by Microsoft, and uses a central key point and n human key point offset mappings to locate the position of the person in a given image, that is, the possibility measure of a certain joint point at a certain position in a two-dimensional image. The center map is modeled as a heat map based on a Gaussian distribution, and the heat map value represents the confidence of the center position. DEKR uses multi-branch parallel convolution to predict the representation of n human key point regressions, so that each representation is concentrated on the corresponding key point area. The two-dimensional posture regression model using multi-branch parallel convolution can reduce the parameters of the network and reduce the amount of calculation. The two-dimensional posture result can be used as the data basis for the subsequent calculation of the deep regression result.
[0065] Step S104, based on the two-dimensional posture result and the deep regression model, the feature map is regressed to obtain a deep regression result, wherein the deep regression model is a model trained based on a deep learning method. The deep regression result includes the absolute depth value of the root key point and the relative depth value of the human body key point, rather than predicting the absolute depth value of all human body key points. Such a representation enables the deep regression result to retain the relevant information of the body key points and improves the stability of the overall training. The pelvis was selected as the root key point of the human body during the design. By using parallel branches, the two-dimensional posture regression and the depth regression are separated from each other to prevent them from affecting each other. At the same time, there is no need to group the human body key points, which simplifies the process of the multi-person three-dimensional posture estimation method. In order to obtain the final result, the deep regression branch improves the accuracy of the root key point prediction by sharing the output features of the two-dimensional posture regression branch.
[0066] Step S105 , obtaining a three-dimensional posture estimation result based on the merging of the two-dimensional posture result and the depth regression result.
[0067] In some embodiments, the two-dimensional posture result includes: human body key points, human body center key point coordinates and human body key point offset mapping set.
[0068] Specifically, the human body key points include 17 human body joints. The coordinates of the human body center key points refer to the average coordinates of all visible human body key points of each character. The human body key point offset mapping refers to the offset distance between the human body key points other than the human body center key points and the human body center key points. The human body key point offset mappings corresponding to all human body key points are combined to form a human body key point offset mapping set. The two-dimensional coordinates of the human body key points in the feature map can be obtained through the two-dimensional posture results, providing a data basis for the subsequent calculation of the two-dimensional posture guidance features.
[0069] In some embodiments, the performing regression calculation on the feature map based on the two-dimensional posture regression model includes:
[0070] Based on the feature map, each human key point offset map in the human key point offset map set corresponding to each human key point is calculated according to the following formula:
[0071] O1=F1(X1),O2=F2(X2),...,O n =F n (X n )
[0072] Among them, F m is the mth regressor corresponding to the mth human key point, O m is the offset mapping of the mth human key point, X mIt is the mth sub-feature map corresponding to the mth human key point obtained by decomposing the feature map, 1≤m≤n, n is the number of offset mappings in the human key point offset mapping set, wherein the regressor is a model trained based on the adaptive convolutional learning model.
[0073] Specifically, the regressors use the same structure and independently predict the offset mapping of the corresponding human key points. By using the regressors with the same structure, the parameters and computational complexity in the deep learning method can be reduced. The obtained human key point offset mapping can provide a data basis for the subsequent calculation of the two-dimensional posture guidance feature.
[0074] In some embodiments, the feature map is regressed based on the two-dimensional posture result and the depth regression model to obtain a depth regression result, referring to Figure 2 ,include:
[0075] Step S201, outputting the root depth regression result and the relative depth regression result through the deep regression model based on the feature map. The deep regression model and the two-dimensional posture regression model use the same structure to output the absolute depth value of the root key point and the relative depth values of the remaining human body key points through parallel branching. The pelvis is used as the root key point here. Since the absolute depth values of all human body key points are not output through the deep regression model, but the absolute depth values of other human body key points are calculated by adding the absolute depth value of the root key point to the relative depth values of other human body key points, not only the relevant information of other human body key points is retained, but also the stability of the deep regression model training is improved by utilizing the characteristic that the regressed relative depth value has better stability than the regressed absolute depth value during training.
[0076] Step S202: Obtain a two-dimensional posture guidance feature through convolution calculation based on the feature map. The absolute depth of the key points of the human body can be partially expressed by the human body scale, so the human body scale is obtained through the two-dimensional posture guidance feature to improve the accuracy of the absolute depth regression of the key points of the human body.
[0077] Step S203: Calculate the root key point depth regression result based on the two-dimensional posture guidance feature and the root depth regression result. The features of the root key point are enriched by the features around the human body key point provided by the two-dimensional posture guidance feature, so as to improve the accuracy of the absolute depth regression of the root key point.
[0078] Step S204: merge the root key point depth regression result and the relative depth regression result to obtain the depth regression result.
[0079] In some embodiments, the step of obtaining the two-dimensional posture guidance feature through convolution calculation based on the feature map includes:
[0080] Outputting a predetermined number of decomposed feature maps by convolution calculation based on the feature map;
[0081] Based on each of the decomposed feature graphs, the human body key point offset mapping set is output through its corresponding regressor, wherein the regressor is a model trained based on an adaptive convolutional learning model;
[0082] A two-dimensional posture guidance feature is obtained by adopting a position query operation based on the human body key point offset mapping set.
[0083] Specifically, an adaptive convolution learning model is used to obtain a human key point offset mapping set through multi-branch parallel convolution, so that the confidence of each human key point at a certain position in the feature map can be obtained, and the regression of each human key point is concentrated in the corresponding human key point area. Then, a two-dimensional posture guidance feature is obtained based on the position query operation, providing a data basis for the absolute depth regression of the subsequent root key point.
[0084] In some embodiments, the calculating the root key point depth regression result based on the two-dimensional posture guidance feature and the root depth regression result includes:
[0085] The root keypoint depth regression result is calculated according to the following formula:
[0086] Z=F Z (cat{X,D(O1'),D(O2')...D(O n-1 ')})
[0087] Among them, F Z is the root keypoint deep regressor in the deep regression model, cat{} is the connection operation on the channel dimension, X is the feature map, D(O k ') is the position query operation on the kth human key point, O k ' is the kth offset mapping in the human key point offset mapping set, 1≤k≤n-1, n is the number of offset mappings in the human key point offset mapping set, wherein the root key point deep regressor is a model trained based on the adaptive convolutional learning model.
[0088] Specifically, through the connection operation of the two-dimensional posture guided feature and the feature map, the feature map of the root key point regression is enriched, making the result of the root key point regression more accurate.
[0089] In some embodiments, the adaptive convolutional learning model is obtained based on the following formula:
[0090]
[0091] Among them, y(q) is the output result of the adaptive convolutional learning model, is the offset of the key point of the human body, q is the coordinate of the key point of the center of the human body, W i is the weight of the convolution kernel in the adaptive convolution learning model, and j is the size of the convolution kernel in the adaptive convolution learning model. By adopting an adaptive convolution learning model with the same structure, the parameters of the deep learning network can be reduced during the training process, and the amount of calculation can be reduced.
[0092] In some embodiments, the calculating and obtaining the three-dimensional posture estimation result based on the two-dimensional posture result and the depth regression result includes:
[0093] Get the camera intrinsic matrix;
[0094] Calculate the coordinates of the key points of the human body in a two-dimensional coordinate system based on the two-dimensional posture result and the depth regression result;
[0095] The three-dimensional posture estimation result is calculated according to the following formula:
[0096] [X,Y,Z] T =ZK -1 [x,y,1] T
[0097] Wherein, [X,Y,Z] is the three-dimensional pose estimation result, [x,y] in [x,y,1] is the coordinates of the key points of the human body in the two-dimensional coordinate system, K is the camera intrinsic matrix, and Z is the depth regression result. The accuracy of the three-dimensional pose estimation results of multiple people is improved by improving the accuracy of the depth regression results and the calculation of the camera intrinsic matrix.
[0098] It should be noted that the method of the embodiment of the present application can be performed by a single device, such as a computer or server. The method of this embodiment can also be applied to a distributed scenario and completed by multiple devices cooperating with each other. In the case of such a distributed scenario, one of the multiple devices can only perform one or more steps in the method of the embodiment of the present application, and the multiple devices will interact with each other to complete the described method.
[0099] It should be noted that the above describes some embodiments of the present application. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recorded in the claims can be performed in an order different from that in the above embodiments and still achieve the desired results. In addition, the processes depicted in the accompanying drawings do not necessarily require the specific order or continuous order shown to achieve the desired results. In some embodiments, multitasking and parallel processing are also possible or may be advantageous.
[0100] Based on the same inventive concept, corresponding to any of the above-mentioned embodiment methods, the present application also provides a multi-person three-dimensional posture estimation device.
[0101] refer to Figure 3 , the multi-person three-dimensional posture estimation device comprises:
[0102] The image acquisition module 301 is configured to acquire images of multiple people taken by a camera;
[0103] A feature map generation module 302 is configured to extract features from the multi-person image based on a backbone network to obtain a feature map, wherein the backbone network is a model trained based on a deep learning method;
[0104] A two-dimensional posture generation module 303 is configured to perform regression calculation on the feature map based on a two-dimensional posture regression model to obtain a two-dimensional posture result, wherein the two-dimensional posture regression model is a model trained based on a deep learning method;
[0105] A depth regression module 304 is configured to perform regression calculation on the feature map based on the two-dimensional posture result and a depth regression model to obtain a depth regression result, wherein the depth regression model is a model trained based on a deep learning method;
[0106] The three-dimensional estimation module 305 is configured to calculate the three-dimensional posture estimation result based on the two-dimensional posture result and the depth regression result. For the convenience of description, the above device is described by functions divided into various modules. Of course, when implementing this application, the functions of each module can be implemented in the same or multiple software and / or hardware.
[0107] The device of the above embodiment is used to implement the corresponding multi-person three-dimensional posture estimation method, device and electronic device in any of the above embodiments, and has the beneficial effects of the corresponding method embodiment, which will not be repeated here.
[0108] Based on the same inventive concept, corresponding to any of the above-mentioned embodiments and methods, the present application also provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the program, the method, device, and electronic device for estimating three-dimensional postures of multiple people described in any of the above embodiments are implemented.
[0109] Figure 4A more specific schematic diagram of the hardware structure of an electronic device provided in this embodiment is shown, and the device may include: a processor 1010, a memory 1020, an input / output interface 1030, a communication interface 1040, and a bus 1050. The processor 1010, the memory 1020, the input / output interface 1030, and the communication interface 1040 are connected to each other through the bus 1050 in the device.
[0110] The processor 1010 can be implemented by a general-purpose CPU (Central Processing Unit), a microprocessor, an application-specific integrated circuit (ASIC), or one or more integrated circuits, and is used to execute relevant programs to implement the technical solutions provided in the embodiments of this specification.
[0111] The memory 1020 may be implemented in the form of ROM (Read Only Memory), RAM (Random Access Memory), static storage device, dynamic storage device, etc. The memory 1020 may store an operating system and other application programs. When the technical solutions provided in the embodiments of this specification are implemented by software or firmware, the relevant program codes are stored in the memory 1020 and are called and executed by the processor 1010.
[0112] The input / output interface 1030 is used to connect the input / output module to realize information input and output. The input / output module can be configured in the device as a component (not shown in the figure), or it can be externally connected to the device to provide corresponding functions. The input device may include a keyboard, a mouse, a touch screen, a microphone, various sensors, etc., and the output device may include a display, a speaker, a vibrator, an indicator light, etc.
[0113] The communication interface 1040 is used to connect a communication module (not shown) to realize communication interaction between the device and other devices. The communication module can realize communication through a wired mode (such as USB, network cable, etc.) or a wireless mode (such as mobile network, WIFI, Bluetooth, etc.).
[0114] The bus 1050 includes a path that transmits information between the various components of the device (eg, the processor 1010, the memory 1020, the input / output interface 1030, and the communication interface 1040).
[0115] It should be noted that, although the above device only shows the processor 1010, the memory 1020, the input / output interface 1030, the communication interface 1040 and the bus 1050, in the specific implementation process, the device may also include other components necessary for normal operation. In addition, it can be understood by those skilled in the art that the above device may also only include the components necessary for implementing the embodiments of the present specification, and does not necessarily include all the components shown in the figure.
[0116] The electronic device of the above-mentioned embodiment is used to implement the corresponding multi-person three-dimensional posture estimation method, device and electronic device in any of the above-mentioned embodiments, and has the beneficial effects of the corresponding method embodiment, which will not be repeated here.
[0117] Based on the same inventive concept, corresponding to any of the above-mentioned embodiment methods, the present application also provides a non-transitory computer-readable storage medium, wherein the non-transitory computer-readable storage medium stores computer instructions, and the computer instructions are used to enable the computer to execute the multi-person three-dimensional posture estimation method, device and electronic device as described in any of the above embodiments.
[0118] The computer-readable medium of this embodiment includes permanent and non-permanent, removable and non-removable media, and information storage can be implemented by any method or technology. Information can be computer-readable instructions, data structures, modules of programs, or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technology, read-only compact disk read-only memory (CD-ROM), digital versatile disk (DVD) or other optical storage, magnetic cassettes, magnetic tape magnetic disk storage or other magnetic storage devices or any other non-transmission media that can be used to store information that can be accessed by a computing device.
[0119] The computer instructions stored in the storage medium of the above embodiment are used to enable the computer to execute the multi-person three-dimensional posture estimation method, device and electronic device as described in any of the above embodiments, and have the beneficial effects of the corresponding method embodiments, which will not be repeated here.
[0120] Based on the same inventive concept, on the basis of the implementation plans corresponding to the above-mentioned various example methods, the following specific implementations may be implemented.
[0121] refer to Figure 5, the image (corresponding to the multi-person image taken by the camera) is first extracted through the skeleton network-HRNet to obtain the feature map X, and then a central key point heat map A (corresponding to the coordinates of the central key points of the human body) is generated through multiple parallel branches. The heat value in A represents the probability of the joint being located at this position; J other key point offset maps (corresponding to the human body key point offset mapping) represent the offset distance between the other key points of each person and their central key point; first, we obtain the candidate center point position coordinates according to the heat map A, and the position coordinates minus the values of the corresponding center point positions in the J offset maps can obtain the position of each key point coordinate of each person, so as to obtain the two-dimensional human posture of the human body. In the depth regression branch, at the root node position of the absolute root depth map, that is, the pelvic point coordinates of the 2D key point, the root node depth of each person is obtained, and at the relative root depth map position, that is, the other key points except the 2D key points, the root relative depth of each other key point is obtained, and added to the root absolute depth of the corresponding person, the absolute depth of each key point of each person can be obtained (corresponding to the depth regression result). At this point, we have obtained the position of each person's key points in two-dimensional coordinates and the depth of the key points in camera coordinates. Using these results and the camera's intrinsic matrix, we can reconstruct the three-dimensional posture through the perspective camera model.
[0122] Those skilled in the art should understand that the discussion of any of the above embodiments is merely illustrative and is not intended to imply that the scope of the present application (including the claims) is limited to these examples. In line with the concept of the present application, the technical features in the above embodiments or different embodiments may be combined, the steps may be implemented in any order, and there are many other variations of the different aspects of the embodiments of the present application as described above, which are not provided in detail for the sake of simplicity.
[0123] In addition, to simplify the description and discussion, and in order not to make the embodiments of the present application difficult to understand, the known power supply / ground connection with the integrated circuit (IC) chip and other components may or may not be shown in the provided drawings. In addition, the device can be shown in the form of a block diagram to avoid making the embodiments of the present application difficult to understand, and this also takes into account the fact that the details of the implementation of these block diagram devices are highly dependent on the platform to be implemented in the embodiments of the present application (that is, these details should be fully within the scope of understanding of those skilled in the art). In the case of elaborating specific details (e.g., circuits) to describe exemplary embodiments of the present application, it is obvious to those skilled in the art that the embodiments of the present application can be implemented without these specific details or when these specific details are changed. Therefore, these descriptions should be considered to be illustrative rather than restrictive.
[0124] Although the present application has been described in conjunction with specific embodiments of the present application, many replacements, modifications and variations of these embodiments will be apparent to those skilled in the art from the foregoing description. For example, other memory architectures (e.g., dynamic RAM (DRAM)) may use the embodiments discussed.
[0125] The embodiments of the present application are intended to cover all such substitutions, modifications and variations that fall within the broad scope of the appended claims. Therefore, any omissions, modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the embodiments of the present application should be included in the scope of protection of the present application.
Claims
1. A method for multi-person three-dimensional posture estimation, characterized in that: include: Get multi-person images taken by the camera; Extracting features of the multi-person image based on a backbone network to obtain a feature map, wherein the backbone network is a model trained based on a deep learning method; Performing regression calculation on the feature map based on a two-dimensional posture regression model to obtain a two-dimensional posture result, wherein the two-dimensional posture regression model is a model trained based on a deep learning method; Performing regression calculation on the feature map based on the two-dimensional posture result and the deep regression model to obtain a deep regression result, wherein the deep regression model is a model trained based on a deep learning method; Calculate a three-dimensional posture estimation result based on the two-dimensional posture result and the depth regression result; The step of performing regression calculation on the feature map based on the two-dimensional posture result and the depth regression model to obtain a depth regression result includes: Outputting a root depth regression result and a relative depth regression result through the depth regression model based on the feature map; Obtaining a two-dimensional posture guidance feature through convolution calculation based on the feature map; Calculate a root key point depth regression result based on the two-dimensional posture guidance feature and the root depth regression result; Combining the root key point depth regression result and the relative depth regression result to obtain the depth regression result; The step of calculating the root key point depth regression result based on the two-dimensional posture guidance feature and the root depth regression result includes: The root keypoint depth regression result is calculated according to the following formula: Z=F Z (cat{X,D(O1'),D(O2')...D(On-1')}) Among them, F Z is the root keypoint deep regressor in the deep regression model, cat{} is the connection operation on the channel dimension, X is the feature map, D(O k ') is the position query operation on the kth human key point, O k ' is the kth offset mapping in the human body key point offset mapping set, 1≤k≤n-1, n is the number of offset mappings in the human body key point offset mapping set, wherein the root key point deep regressor is a model trained based on an adaptive convolutional learning model; The two-dimensional posture result includes: human body key points, human body center key point coordinates and human body key point offset mapping set; The step of obtaining a two-dimensional posture guidance feature through convolution calculation based on the feature map includes: Outputting a predetermined number of decomposed feature maps by convolution calculation based on the feature map; Based on each of the decomposed feature maps, the human body key point offset mapping set is output through its corresponding regressor, wherein the regressor is a model trained based on an adaptive convolutional learning model; based on the human body key point offset mapping set, a position query operation is used to obtain a two-dimensional posture guidance feature.
2. The method according to claim 1, characterized in that The performing regression calculation on the feature map based on the two-dimensional posture regression model includes: Based on the feature map, each human key point offset map in the human key point offset map set corresponding to each human key point is calculated according to the following formula: O1=F1(X1),O2=F2(X2),...,O n =F n (X n ) Among them, F m is the mth regressor corresponding to the mth human key point, O m is the offset mapping of the mth human key point, X m It is the mth sub-feature map corresponding to the mth human key point obtained by decomposing the feature map, 1≤m≤n, n is the number of offset mappings in the human key point offset mapping set, wherein the regressor is a model trained based on the adaptive convolutional learning model.
3. The method according to claim 1, characterized in that The adaptive convolutional learning model is obtained based on the following formula: Among them, y(q) is the output result of the adaptive convolutional learning model, is the offset of the key point of the human body, q is the coordinate of the key point of the center of the human body, W i is the weight of the convolution kernel in the adaptive convolution learning model, and j is the size of the convolution kernel in the adaptive convolution learning model.
4. The method according to claim 1, characterized in that: The calculating and obtaining the three-dimensional posture estimation result based on the two-dimensional posture result and the depth regression result includes: Get the camera intrinsic matrix; Calculate the coordinates of key points of the human body in a three-dimensional coordinate system based on the two-dimensional posture result and the depth regression result; The three-dimensional posture estimation result is calculated according to the following formula: [X,Y,Z] T =ZK -1 [x,y,1] T Among them, [X, Y, Z] is the three-dimensional posture estimation result, [x, y] in [x, y, 1] is the coordinates of the key points of the human body in the two-dimensional coordinate system, K is the camera intrinsic matrix, and Z is the depth regression result.
5. A multi-person three-dimensional posture estimation device, characterized in that: include: An image acquisition module is configured to acquire images of multiple people taken by a camera; A feature map generation module is configured to extract features from the multi-person image based on a backbone network to obtain a feature map, wherein the backbone network is a model trained based on a deep learning method; A two-dimensional posture generation module is configured to perform regression calculation on the feature map based on a two-dimensional posture regression model to obtain a two-dimensional posture result, wherein the two-dimensional posture regression model is a model trained based on a deep learning method; A depth regression module is configured to perform regression calculation on the feature map based on the two-dimensional posture result and a depth regression model to obtain a depth regression result, wherein the depth regression model is a model trained based on a deep learning method; A three-dimensional estimation module is configured to calculate a three-dimensional posture estimation result based on the two-dimensional posture result and the depth regression result; The deep regression module is specifically configured as follows: Outputting a root depth regression result and a relative depth regression result through the depth regression model based on the feature map; Obtaining a two-dimensional posture guidance feature through convolution calculation based on the feature map; Calculate a root key point depth regression result based on the two-dimensional posture guidance feature and the root depth regression result; Combining the root key point depth regression result and the relative depth regression result to obtain the depth regression result; The deep regression module is further configured as follows: The root keypoint depth regression result is calculated according to the following formula: Z=F Z (cat{X,D(O1'),D(O2')...D(On-1')}) Among them, F Z is the root keypoint deep regressor in the deep regression model, cat{} is the connection operation on the channel dimension, X is the feature map, D(O k ') is the position query operation on the kth human key point, O k ' is the kth offset mapping in the human body key point offset mapping set, 1≤k≤n-1, n is the number of offset mappings in the human body key point offset mapping set, wherein the root key point deep regressor is a model trained based on an adaptive convolutional learning model; The two-dimensional posture result includes: human body key points, human body center key point coordinates and human body key point offset mapping set; The deep regression module is further configured as follows: Outputting a predetermined number of decomposed feature maps by convolution calculation based on the feature map; Based on each of the decomposed feature graphs, the human body key point offset mapping set is output through its corresponding regressor, wherein the regressor is a model trained based on an adaptive convolutional learning model; A two-dimensional posture guidance feature is obtained by adopting a position query operation based on the human body key point offset mapping set.
6. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the program, the method according to any one of claims 1 to 4 is implemented.