Human body 3D posture estimation method, device, computer equipment and storage medium
By using the object detection network, fully connected neural network and encoder-decoder structure in human 3D pose estimation, the problem of inaccurate estimation of human 3D pose in the prior art is solved, and higher estimation accuracy and reduced annotation dependence are achieved.
Patent Information
- Application Number
- CN202111014652.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-08-31
- Publication Date
- 2025-05-06
- Estimated Expiration
- 2041-08-31
AI Technical Summary
In the prior art, the human body 3D posture node labeling cost is high and when there are not enough labeling images, the human body 3D posture estimates are not accurate enough.
By obtaining the target picture, input it to the target detection network to obtain the 2D key point coordinates and semantic feature maps, the embedding feature vector is extracted using a fully connected neural network, and encoded and decoded through the encoder-decoder structure to finally determine the 3D key point coordinates of the target to be identified.
It improves the accuracy of human 3D pose estimation and reduces the dependence on human 3D pose node annotation.
Smart Images

Figure CN113920529B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of computer vision, and in particular to a method, device, computer equipment and storage medium for estimating 3D posture of a human body. Background Art
[0002] 3D human pose estimation is a technology that recognizes the 3D movements of a human body from 2D images.
[0003] In the prior art, an end-to-end convolutional neural network is usually used to directly predict the 3D joint positions of the human body from the input image. However, the use of traditional convolutional neural networks requires labeling according to the 3D posture nodes of the human body, training the convolutional neural network through the labeled information, and using the trained convolutional neural network to estimate the input image. However, labeling the 3D posture nodes of the human body is a labor-intensive task, and when there are not enough labeled images, the prior art may not estimate the 3D posture of the human body accurately. Summary of the invention
[0004] Based on this, the embodiments of the present application provide a method, apparatus, computer device and storage medium for estimating 3D posture of a human body, which can solve the problem of inaccurate 3D posture estimation of a human body in the prior art.
[0005] In a first aspect, a method for human 3D posture recognition is provided, the method comprising:
[0006] Acquire a target image, wherein the target image includes a target to be identified;
[0007] Input the target image into the target detection network to obtain multiple 2D key point coordinates and semantic feature maps of the target to be identified;
[0008] Each 2D key point coordinate of the target to be identified is passed through a fully connected neural network to obtain an embedding feature vector corresponding to each 2D key point coordinate;
[0009] Decomposing the semantic feature graph of the target to be identified to obtain a decomposed feature vector, and performing a dimensionality reduction operation on the decomposed feature vector to obtain a reduced dimensionality feature vector of a preset dimension;
[0010] Inputting the reduced-dimensional feature vector into an encoder for encoding to obtain an encoded vector, wherein the number of the reduced-dimensional feature vector is the same as the number of nodes of the encoder;
[0011] Decoding the embedding feature vector and the encoding vector by a decoder to obtain a decoding vector;
[0012] The decoded vector is input into the fully connected neural network to determine the 3D key point coordinates of the target to be identified.
[0013] In one embodiment, the decoder includes three layers, and the decoding is performed by the decoder according to the embedding feature vector and the encoding vector to obtain a decoding vector, including:
[0014] The embedding feature vector and the encoding vector are input into a first decoder for decoding to obtain a first decoding vector; the first decoding vector and the encoding vector are input into a second decoder for decoding to obtain a second decoding vector; the second decoding vector and the encoding vector are input into a third decoder for decoding to obtain a decoding vector.
[0015] In one embodiment, the first decoder, the second decoder and the third decoder have the same structure.
[0016] In one embodiment, the step of inputting the reduced dimension feature vector into an encoder for encoding to obtain an encoded vector comprises:
[0017] Transforming each eigenvector in the dimension-reduced eigenvector into three first transformed eigenvectors through three transformation matrices;
[0018] Inputting the first transformed feature vector into a Multi-head attention network for calculation to obtain a first feedback vector having the same number and dimension as the reduced-dimensional feature vector;
[0019] After adding the first feedback vector to the reduced-dimensional feature vector, a normalization algorithm is used for processing, and each vector in the processed normalized vector is input into a 2-layer fully connected feedforward network and then added to the normalized vector, and then the added vector is normalized to obtain a coding vector.
[0020] In one embodiment, the semantic feature map includes the number of feature channels, image height and image width.
[0021] In one embodiment, the object detection network includes MaskRCNN.
[0022] In one of the embodiments, the fully connected neural network is a two-layer fully connected neural network.
[0023] In a second aspect, a human 3D posture recognition device is provided, the device comprising:
[0024] A target detection module is used to input the target image into a target detection network to obtain multiple 2D key point coordinates and a semantic feature map of the target to be identified;
[0025] A fully connected network module, used for obtaining an embedding feature vector corresponding to each 2D key point coordinate of the target to be identified through a fully connected neural network;
[0026] A processing module, used for decomposing the semantic feature graph of the target to be identified to obtain a decomposed feature vector, and performing a dimensionality reduction operation on the decomposed feature vector to obtain a reduced dimensionality feature vector of a preset dimension;
[0027] An encoding module, used for inputting the reduced-dimensionality feature vector into an encoder for encoding to obtain an encoded vector, wherein the number of the reduced-dimensionality feature vector is the same as the number of nodes of the encoder;
[0028] A decoding module, used for decoding the embedding feature vector and the encoding vector through a decoder to obtain a decoding vector;
[0029] A determination module is used to input the decoding vector into the fully connected neural network to determine the 3D key point coordinates of the target to be identified.
[0030] In a third aspect, a computer device is provided, including a memory and a processor, wherein the memory stores a computer program, and when the processor executes the computer program, the human body 3D posture recognition method described in any one of the first aspects is implemented.
[0031] In a fourth aspect, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the human body 3D posture recognition method described in any one of the first aspects is implemented.
[0032] The technical solution provided by the embodiment of the present application obtains a target image with a target to be identified, and inputs the target image into a target detection network to obtain multiple 2D key point coordinates and a semantic feature map of the target to be identified; obtains a corresponding embedding feature vector for each obtained 2D key point coordinate through a fully connected neural network; and decomposes and reduces the dimension of the semantic feature map of the target to be identified to obtain a reduced dimension feature vector of a preset dimension; inputs the reduced dimension feature vector into an encoder for encoding to obtain an encoded vector, and then decodes it through a decoder according to the embedding feature vector and the encoded vector to obtain a decoded vector; and inputs the decoded vector into a fully connected neural network to determine the 3D key point coordinates of the target to be identified. It can be seen that compared with the prior art, the present application improves the accuracy of 3D posture estimation of the human body through the encoder-decoder method. BRIEF DESCRIPTION OF THE DRAWINGS
[0033] Figure 1 A flow chart of a human 3D posture recognition method provided in an embodiment of the present application;
[0034] Figure 2 A block diagram of a human 3D posture recognition device provided in an embodiment of the present application;
[0035] Figure 3 A schematic diagram of a computer device provided in an embodiment of the present application. DETAILED DESCRIPTION
[0036] In order to make the objectives, technical solutions and advantages of the present application clearer, the implementation methods of the present application will be further described in detail below with reference to the accompanying drawings.
[0037] Please refer to Figure 1 , which shows a flow chart of a method for estimating a 3D human body posture provided by an embodiment of the present application. The method for estimating a 3D human body posture may include the following steps:
[0038] Step 101, obtaining a target image.
[0039] Among them, the target image includes a target to be identified. In the embodiment of the present application, the target to be identified is a person, and there is only one person in the acquired target image.
[0040] In one embodiment of the present application, if the acquired picture includes multiple persons, the picture can be identified and cropped using existing image recognition technology to crop the picture into a target picture with only one person.
[0041] Step 102: Input the target image into the target detection network to obtain multiple 2D key point coordinates and semantic feature maps of the target to be identified.
[0042] Among them, the target detection network can be a MaskRCNN algorithm. By inputting the target image into MaskRCNN, multiple 2D key point coordinates of the target to be identified in the target image can be obtained. In one embodiment of the present application, it can be expressed as p1=(x1,y1), p2=(x2,y2),...pk=(xk,yk), where k is the predefined number of human key points. The semantic feature map of the target to be identified after RoI Align can also be obtained. The semantic feature map is denoted as sem_feat_map, and the dimension is (C,H,W), C is the number of feature channels, H is the height of the image after convolution, and W is the width of the image after convolution.
[0043] Step 103: Obtain an embedding feature vector corresponding to each 2D key point coordinate of the target to be identified through a fully connected neural network.
[0044] Among them, the fully connected neural network is a two-layer fully connected neural network. Each 2D key point coordinate of the target to be identified is extracted through the two-layer fully connected neural network to extract the embedding feature vector corresponding to each 2D key point coordinate. In one embodiment of the present application, a 256-dimensional embedding feature vector is extracted.
[0045] Step 104 , decomposing the semantic feature graph of the target to be identified to obtain a decomposed feature vector, and performing a dimensionality reduction operation on the decomposed feature vector to obtain a reduced dimensionality feature vector of a preset dimension.
[0046] Among them, the semantic feature map of the target to be identified is decomposed to obtain a decomposed feature vector.
[0047] In one embodiment of the present application, a feature map sem_feat_map with dimensions (C, H, W) is decomposed into H×W C-dimensional feature vectors, and H×W is denoted as m. Since C is greater than 256, it is necessary to perform a dimensionality reduction operation on the decomposed feature vectors, and transform the m C-dimensional feature vectors into m 256-dimensional feature vectors through the dimensionality reduction operation.
[0048] Step 105: input the dimension-reduced feature vector into an encoder for encoding to obtain an encoded vector.
[0049] Among them, the number of reduced-dimensional feature vectors is the same as the number of nodes of the encoder. In the embodiment of the present application, each feature vector in the reduced-dimensional feature vector is transformed into three first transformed feature vectors through three transformation matrices; the first transformed feature vector is input into the Multi-head attention network for calculation to obtain a first feedback vector with the same number and dimension as the reduced-dimensional feature vector; after adding the first feedback vector to the reduced-dimensional feature vector, a normalization algorithm is used for processing, and each vector in the processed normalized vector is input into a 2-layer fully connected feedforward network, and then added to the normalized vector, and then the added vector is normalized to obtain a coding vector.
[0050] In one embodiment of the present application, the m 256-dimensional feature vectors obtained in step 104 are input into a single-layer encoder, each feature vector is input into a node of the encoder, and the encoder has a total of m input nodes. The specific encoding process is as follows:
[0051] S1, each eigenvector is transformed into three eigenvectors Q, K, V through three transformation matrices, and a total of 3×m eigenvectors Qi, Ki, Vi (i ranges from 1 to m) are obtained;
[0052] S2, input the above 3×m feature vectors into the multi-head attention layer to obtain m 256-dimensional vectors;
[0053] S3, adding the m vectors obtained in step S2 to the m vectors input by the encoder to obtain m 256-dimensional vectors;
[0054] S4, performs LayerNorm operation on the m 256-dimensional vectors obtained in S3 to obtain normalized m 256-dimensional vectors;
[0055] S5, for the m vectors obtained in S4, input each vector into a 2-layer fully connected feedforward network to obtain m 256-dimensional vectors. Among them, the m vectors share one feedforward network;
[0056] S6, adds the m vectors obtained by S5 and S4, and performs LayerNorm operation on the added m vectors to obtain m normalized 256-dimensional vectors, recorded as encoder_feat_i (i ranges from 1 to m).
[0057] Step 106: Decode the embedding feature vector and the encoding vector through a decoder to obtain a decoding vector.
[0058] The embedding feature vector and the encoding vector are input into the first decoder for decoding to obtain a first decoding vector; the first decoding vector and the encoding vector are input into the second decoder for decoding to obtain a second decoding vector; the second decoding vector and the encoding vector are input into the third decoder for decoding to obtain a decoding vector. The first decoder, the second decoder and the third decoder have the same structure.
[0059] In one embodiment of the present application, the k 256-dimensional embedding feature vectors obtained in step 103, or the k 256-dimensional vectors output by each decoding layer, are recorded as p_feat_j (j ranges from 1 to k), together with the m vectors encoder_feat_i (i ranges from 1 to m) obtained in step 105, and input into the decoding layer. In this embodiment, a total of three decoding layers (decoders) are set. The specific process of one decoding includes:
[0060] S1, transform each eigenvector in the k vectors p_feat_j into three eigenvectors Q, K, V through three transformation matrices, and obtain a total of 3×k eigenvectors Qj, Kj, Vj (j ranges from 1 to k);
[0061] S2, input the above 3×k feature vectors into the Multi-head attention layer to obtain k 256-dimensional vectors;
[0062] S3, add the k vectors obtained in S2 to the k vectors of the decoded input of this layer to obtain k 256-dimensional vectors;
[0063] S4, perform LayerNorm operation on the k 256-dimensional vectors obtained in S3 to obtain normalized k 256-dimensional vectors;
[0064] S5, transform each of the k vectors obtained in S4 into vector Q through the transformation matrix, and transform each of the m vectors encoder_feat_i into K, V through two different transformation matrices. A total of 2×m+k vectors are obtained, denoted as Qj (j from 1 to k), Ki,Vi (i from 1 to m);
[0065] S6, input Q, K, V obtained in S5 into the Multi-head attention layer to obtain k 256-dimensional vectors;
[0066] S7, adding the k vectors obtained in S6 and the k vectors obtained in 4) to obtain k 256-dimensional vectors;
[0067] S8, performing LayerNorm operation on the k 256-dimensional vectors obtained in S7 to obtain normalized k 256-dimensional vectors;
[0068] S9, for the k vectors obtained in S8, input each vector into a 2-layer fully connected feedforward network to obtain k 256-dimensional vectors. Among them, the k vectors share one feedforward network;
[0069] S10 adds the k vectors obtained by S8 and S9, and performs LayerNorm operation on the added k vectors to obtain k normalized 256-dimensional vectors as the k 256-dimensional vectors output by each decoding layer and as the input of the next decoding layer.
[0070] In this embodiment, when the number of decoding layers is 3, the above iteration is performed 3 times to obtain a decoding vector.
[0071] Step 107, input the decoded vector into a fully connected neural network to determine the 3D key point coordinates of the target to be identified.
[0072] In one embodiment of the present application, for the k 256-dimensional vectors output by the last decoding layer, each vector is input into a two-layer fully connected neural network, the output of which is the 3D coordinate of each key point, wherein each vector shares a fully connected network.
[0073] Please refer to Figure 2 , which shows a block diagram of a human 3D posture recognition device 200 provided in an embodiment of the present application. Figure 2 As shown, the device 200 may include: an acquisition module 201, a target detection module 202, a fully connected network module 203, a processing module 204, an encoding module 205, a decoding module 206, and a determination module 207.
[0074] An acquisition module 201 is used to acquire a target image, wherein the target image includes a target to be identified;
[0075] The target detection module 202 is used to input the target image into the target detection network to obtain multiple 2D key point coordinates and semantic feature maps of the target to be identified;
[0076] A fully connected network module 203 is used to obtain an embedding feature vector corresponding to each 2D key point coordinate of the target to be identified through a fully connected neural network;
[0077] The processing module 204 is used to decompose the semantic feature graph of the target to be identified to obtain a decomposed feature vector, and perform a dimension reduction operation on the decomposed feature vector to obtain a reduced dimension feature vector of a preset dimension;
[0078] The encoding module 205 is used to input the reduced-dimensionality feature vector into the encoder for encoding to obtain an encoded vector, wherein the number of the reduced-dimensionality feature vector is the same as the number of nodes of the encoder;
[0079] A decoding module 206, configured to decode the embedding feature vector and the encoding vector through a decoder to obtain a decoding vector;
[0080] The determination module 207 is used to input the decoded vector into a fully connected neural network to determine the 3D key point coordinates of the target to be identified.
[0081] For the specific definition of the human body 3D posture recognition device, please refer to the definition of the human body 3D posture recognition method above, which will not be repeated here. Each module in the above-mentioned human body 3D posture recognition device can be implemented in whole or in part by software, hardware and a combination thereof. The above-mentioned modules can be embedded in or independent of the processor in the computer device in the form of hardware, or can be stored in the memory in the computer device in the form of software, so that the processor can call and execute the operations corresponding to the above modules.
[0082] In one embodiment, a computer device is provided. The computer device may be a server, and its internal structure diagram may be as follows: Figure 3As shown. The computer device includes a processor, a memory and a network interface connected through a system bus. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The database of the computer device is used for human 3D posture recognition data. The network interface of the computer device is used to communicate with an external terminal through a network connection, and the display is used to display the posture recognition result. When the computer program is executed by the processor, a human 3D posture recognition method is implemented.
[0083] Those skilled in the art will understand that Figure 3 The structure shown in the figure is only a block diagram of a part of the structure related to the solution of the present application, and does not constitute a limitation on the computer device to which the solution of the present application is applied. The specific computer device may include more or fewer components than those shown in the figure, or combine certain components, or have a different arrangement of components.
[0084] In one embodiment of the present application, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the steps of the above-mentioned human body 3D posture recognition method are implemented.
[0085] The computer-readable storage medium provided in this embodiment has similar implementation principles and technical effects to those of the above method embodiments, and will not be described in detail here.
[0086] Those of ordinary skill in the art can understand that all or part of the processes in the above-mentioned embodiment methods can be completed by instructing related hardware through a computer program, and the computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above-mentioned methods. Among them, any reference to memory, storage, database or other media used in the embodiments provided in this application may include non-volatile and / or volatile memory. Non-volatile memory may include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM) or flash memory. Volatile memory may include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in M forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link (SyMchliMk) DRAM (SLDRAM), memory bus (RaMbus) direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM), etc.
[0087] The technical features of the above-described embodiments may be arbitrarily combined. To make the description concise, not all possible combinations of the technical features in the above-described embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.
[0088] The above-described embodiments only express several implementation methods of the present application, and the descriptions thereof are relatively specific and detailed, but they cannot be construed as limiting the scope of the patent application. It should be pointed out that, for a person of ordinary skill in the art, several variations and improvements can be made without departing from the concept of the present application, and these all belong to the protection scope of the present application. Therefore, the protection scope of the patent application shall be subject to the attached claims.
Claims
1. A method for human 3D posture recognition, characterized in that: The method comprises: Acquire a target image, wherein the target image includes a target to be identified; Input the target image into the target detection network to obtain multiple 2D key point coordinates and semantic feature maps of the target to be identified; Each 2D key point coordinate of the target to be identified is passed through a fully connected neural network to obtain an embedding feature vector corresponding to each 2D key point coordinate; Decomposing the semantic feature graph of the target to be identified to obtain a decomposed feature vector, and performing a dimensionality reduction operation on the decomposed feature vector to obtain a reduced dimensionality feature vector of a preset dimension; Inputting the reduced-dimensional feature vector into an encoder for encoding to obtain an encoded vector, wherein the number of the reduced-dimensional feature vector is the same as the number of nodes of the encoder; Decoding the embedding feature vector and the encoding vector by a decoder to obtain a decoding vector; The decoded vector is input into the fully connected neural network to determine the 3D key point coordinates of the target to be identified.
2. The method according to claim 1, characterized in that The decoder includes three layers, and the decoding is performed by the decoder according to the embedding feature vector and the encoding vector to obtain a decoding vector, including: The embedding feature vector and the encoding vector are input into a first decoder for decoding to obtain a first decoding vector; the first decoding vector and the encoding vector are input into a second decoder for decoding to obtain a second decoding vector; the second decoding vector and the encoding vector are input into a third decoder for decoding to obtain a decoding vector.
3. The method according to claim 2, characterized in that The first decoder, the second decoder, and the third decoder have the same structure.
4. The method according to claim 1, characterized in that: The step of inputting the dimension-reduced feature vector into an encoder for encoding to obtain an encoded vector comprises: Transforming each eigenvector in the dimension-reduced eigenvector into three first transformed eigenvectors through three transformation matrices; Inputting the first transformed feature vector into a Multi-head attention network for calculation to obtain a first feedback vector having the same number and dimension as the reduced-dimensional feature vector; After adding the first feedback vector to the reduced-dimensional feature vector, a normalization algorithm is used for processing, and each vector in the processed normalized vector is input into a 2-layer fully connected feedforward network and then added to the normalized vector, and then the added vector is normalized to obtain a coding vector.
5. The method according to claim 1, characterized in that The semantic feature map includes the number of feature channels, image height and image width.
6. The method according to claim 1, characterized in that The target detection network includes MaskRCNN.
7. The method according to claim 1, characterized in that The fully connected neural network is a two-layer fully connected neural network.
8. A human body 3D posture recognition device, characterized in that: The device comprises: An acquisition module, used to acquire a target image, wherein the target image includes a target to be identified; A target detection module is used to input the target image into a target detection network to obtain multiple 2D key point coordinates and a semantic feature map of the target to be identified; A fully connected network module, used for obtaining an embedding feature vector corresponding to each 2D key point coordinate of the target to be identified through a fully connected neural network; A processing module, used for decomposing the semantic feature graph of the target to be identified to obtain a decomposed feature vector, and performing a dimensionality reduction operation on the decomposed feature vector to obtain a reduced dimensionality feature vector of a preset dimension; An encoding module, used for inputting the reduced-dimensionality feature vector into an encoder for encoding to obtain an encoded vector, wherein the number of the reduced-dimensionality feature vector is the same as the number of nodes of the encoder; A decoding module, used for decoding the embedding feature vector and the encoding vector through a decoder to obtain a decoding vector; A determination module is used to input the decoding vector into the fully connected neural network to determine the 3D key point coordinates of the target to be identified.
9. A computer device, characterized in that: The method comprises a memory and a processor, wherein the memory stores a computer program, and when the computer program is executed by the processor, the method for recognizing a 3D posture of a human body as claimed in any one of claims 1 to 7 is implemented.
10. A computer-readable storage medium, characterized in that: A computer program is stored thereon, and when the computer program is executed by a processor, the human body 3D posture recognition method as described in any one of claims 1 to 7 is implemented.
Citation Information
Patent Citations
A human skeleton key point detection method and system
CN109784149A
Method and device for human body action recognition
CN112587129A