A 3D reconstruction method, device, equipment and computer storage medium
Through the combination of 3D pose regression network and vertex regression network, the problems of expensive equipment and inaccurate prediction in the existing three-dimensional model construction methods are solved, and the accuracy of the three-dimensional model structure is improved.
Patent Information
- Application Number
- CN202111336174.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-11-10
- Publication Date
- 2025-07-04
- Estimated Expiration
- 2041-11-10
AI Technical Summary
The existing three-dimensional model construction methods rely on expensive equipment and are inefficient. The prediction results of the reconstruction method based on model parameters are inaccurate, resulting in inaccurate three-dimensional model structure.
The 3D pose regression network and the vertex regression network are used to obtain the 2D pose information of the target object, and the 3D pose regression network is used to extract the 3D pose information, and then splice it with the 2D pose information and input it into the vertex regression network to build a three-dimensional model.
The accuracy of the three-dimensional model structure is improved, and the accuracy of the vertex coordinate information corresponding to the three-dimensional model is enhanced by fully utilizing the posture information of different dimensions.
Smart Images

Figure CN114219890B_ABST
Abstract
Description
Technical Field
[0001] This application belongs to the field of three-dimensional reconstruction, and particularly relates to a three-dimensional reconstruction method, device, equipment, and computer storage medium. Background Art
[0002] The construction of early three-dimensional models generally relied on devices such as three-dimensional scanners or multi-view cameras to scan the surface information of the target to establish the corresponding three-dimensional model. This method requires expensive and complex equipment to establish a three-dimensional model, with high modeling costs and low efficiency.
[0003] With the continuous development of deep learning technology, the three-dimensional reconstruction method based on model parameters has become a research hotspot. The three-dimensional reconstruction method based on model parameters obtains the corresponding three-dimensional model by training and optimizing the pose parameters and shape parameters. This method depends on the optimization degree of the pose parameters and shape parameters, often resulting in inaccurate prediction results and inaccurate three-dimensional model structures. Summary of the Invention
[0004] The purpose of the embodiments of this application is to provide a three-dimensional reconstruction method, device, equipment, and computer storage medium, which can improve the accuracy of the three-dimensional model structure in the three-dimensional model construction task.
[0005] The embodiments of this application are implemented as follows. In the first aspect, the embodiments of this application provide a three-dimensional reconstruction method. The method includes: obtaining the 2D pose information of the target object, where the 2D pose information includes the position coordinates of multiple key points on the target object; inputting the 2D pose information into the trained three-dimensional reconstruction model for processing to obtain the three-dimensional model of the target object; the three-dimensional reconstruction model includes a 3D pose regression network and a vertex regression network; the 3D pose regression network is used to extract 3D pose information from the 2D pose information; the vertex regression network is used to process the concatenated information of the 2D pose information and the 3D pose information to obtain the three-dimensional model of the target object.
[0006] In one embodiment, the 3D pose regression network includes a first deformation layer, a first fully connected layer, multiple first residual blocks, a second fully connected layer, and a second deformation layer connected in sequence; the first deformation layer is used to convert the position coordinates of multiple key points into a feature vector in a preset format; the second deformation layer is used to convert the feature information output by the second fully connected layer into 3D pose information.
[0007] In one embodiment, the multiple first residual blocks include a normalization layer, an activation function layer, and a third fully connected layer connected in sequence.
[0008] In one embodiment, the vertex regression network includes a plurality of first graph convolutional layers, a third deformation layer, a fourth fully connected layer, a fourth deformation layer, an intermediate layer, and a plurality of second graph convolutional layers connected in sequence. The intermediate layer includes a plurality of second residual blocks and upsampling layers arranged alternately; the third deformation layer, the fourth fully connected layer, and the fourth deformation layer are used to map the feature map output by the plurality of first graph convolutional layers and input it into the intermediate layer.
[0009] In one embodiment, the second residual block includes a plurality of graph convolutional layers connected in sequence.
[0010] In one embodiment, the process of obtaining the three-dimensional reconstruction model includes: training the 3D pose regression initial network using the first loss function and the first training set to obtain an updated 3D pose regression network; training the updated 3D pose regression network and the vertex regression initial network using the second loss function and the second training set to obtain the three-dimensional reconstruction model.
[0011] In one embodiment, the second loss function includes a mesh loss, a 3D loss, a mesh surface normal loss, and a mesh surface edge loss.
[0012] In a second aspect, an embodiment of the present application provides a three-dimensional reconstruction device, which includes: an acquisition unit for acquiring 2D pose information of a target object, where the 2D pose information includes the position coordinates of a plurality of key points on the target object;
[0013] a processing unit for inputting the 2D pose information into a trained three-dimensional reconstruction model for processing to obtain a three-dimensional model of the target object;
[0014] The three-dimensional reconstruction model includes a 3D pose regression network and a vertex regression network; the 3D pose regression network is used to extract 3D pose information from the 2D pose information; the vertex regression network is used to process the concatenated information of the 2D pose information and the 3D pose information to obtain a three-dimensional model of the target object.
[0015] In a third aspect, an embodiment of the present application provides a terminal device, where the device includes: a processor and a memory, the memory is used to store a computer program, and the processor is used to call and run the computer program from the memory, so that the device executes the method described in any item of the first aspect.
[0016] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium, in which a computer program is stored. When the computer program is executed by a processor, the processor is caused to execute the method described in any item of the first aspect.
[0017] Fifth aspect, an embodiment of the present application provides a computer program product, which includes computer program code. When the computer program code is run on a computer, it causes the computer to execute the method described in any one of the first aspect.
[0018] It can be understood that the beneficial effects of the above second aspect to fifth aspect can be referred to the relevant descriptions in the first aspect above, and will not be elaborated here.
[0019] The 3D reconstruction method, device, equipment and computer storage medium provided by the present application extract 3D pose information from the 2D pose information of the target object by using a 3D pose regression network, splice the extracted 3D pose information and 2D pose information to form 5D pose information, and then input the 5D pose information into the vertex regression network for processing to obtain vertex coordinate information corresponding to the 3D model of the target object, so as to construct a 3D model corresponding to the target object. This method can process the spliced 5D position coordinates based on the vertex regression network to make full use of the pose information of different dimensions, make the vertex coordinate information corresponding to the obtained 3D model more accurate, and thus improve the accuracy of the 3D model structure. BRIEF DESCRIPTION OF THE DRAWINGS
[0020] Figure 1 is a network framework diagram of a 3D reconstruction model provided by an embodiment of the present application;
[0021] Figure 2 is a comparison schematic diagram of a human joint tree provided by an embodiment of the present application;
[0022] Figure 3 is a human body mesh effect diagram of a 3D reconstruction method provided by an embodiment of the present application;
[0023] Figure 4 is a reconstruction effect diagram based on a human body 3D reconstruction model provided by an embodiment of the present application;
[0024] Figure 5 is a structural schematic diagram of a 3D reconstruction device provided by an embodiment of the present application;
[0025] Figure 6 is a structural schematic diagram of a terminal device provided by an embodiment of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0026] To make the objectives, technical solutions and advantages of the present application clearer, the embodiments of the present application will be further described in detail below in conjunction with the accompanying drawings.
[0027] For a 3D model construction task, this application provides a 3D reconstruction method. After obtaining the 2D pose information of a target object, a 3D pose regression network is used to extract 3D pose information from the 2D pose information of the target object. The extracted 3D pose information and 2D pose information are concatenated to form 5D pose information, and then the 5D pose information is input into a vertex regression network for processing to obtain vertex coordinate information corresponding to the 3D model of the target object, thereby constructing a 3D model corresponding to the target object. This method can process the concatenated 5D position coordinates based on the vertex regression network to make full use of pose information in different dimensions, making the vertex coordinate information corresponding to the obtained 3D model more accurate, thereby improving the accuracy of the 3D model structure.
[0028] The following uses specific embodiments to elaborate on the technical solutions of this application in detail. The embodiments described below with reference to the accompanying drawings are exemplary and are intended to explain this application, and should not be construed as a limitation to this application.
[0029] First, in combination with Figure 1 An exemplary introduction is given to a 3D reconstruction model provided by this application. This 3D reconstruction model can be deployed in a 3D reconstruction processing device, which can be a mobile terminal such as a smartphone, a tablet computer, a camera, etc., or a device capable of 3D model reconstruction such as a desktop computer, a robot, a server, etc.
[0030] The 3D reconstruction model provided by this application includes a 3D pose regression network and a vertex regression network. Among them, the 3D pose regression network includes a first deformation layer, a first fully connected layer, a plurality of first residual blocks, a second fully connected layer, and a second deformation layer connected in sequence. The first deformation layer is used to convert the input 2D pose information into a feature vector in a preset format. For example, for a human body 3D reconstruction model, the conversion process can be R m×n →R m·n conversion, where m can represent the number of key points in the human joint tree, and n can represent the dimension of the human 2D pose information; the first fully connected layer is used to map the feature vector converted by the first deformation layer into a feature vector in a higher dimension; the plurality of first residual blocks are used to extract feature vectors in a higher dimension from different 2D pose information; the second fully connected layer is used to inverse map the feature vector in a higher dimension into a feature vector corresponding to the dimension of the 3D pose information, and the second deformation layer is used to convert the feature vector output by the second fully connected layer into 3D pose information correspondingly. For example, for a human body 3D reconstruction model, the conversion process can be R m·n →R m×n conversion, where m can represent the number of key points in the human joint tree, and n can represent the dimension of the human 3D pose information.
[0031] In one example, as Figure 1Each of the multiple first residual blocks includes a normalization layer, an activation function layer, and a third fully connected layer connected in sequence. The multiple first residual blocks are used to extract higher-dimensional features from different 2D pose information. See Figure 1 , the multiple first residual blocks may include two first residual blocks, and the number of first residual blocks can be designed according to the accuracy of the 3D pose information corresponding to the 2D pose information output by the second deformation layer.
[0032] When performing the three-dimensional model reconstruction task, the 3D pose regression network first maps the 2D pose information into a higher-dimensional feature vector, then uses multiple first residual blocks to extract higher-dimensional features from different 2D pose information, and then performs an inverse mapping on the higher-dimensional features to obtain the 3D pose information corresponding to the above 2D pose information. It is not difficult to understand that when the above 2D pose information includes the two-dimensional position coordinates of multiple key points corresponding to the target object, the 3D pose information correspondingly includes the three-dimensional position coordinates of multiple key points corresponding to the target object.
[0033] In the embodiment of the present application, the vertex regression network includes Figure 1 multiple first graph convolutional layers, a third deformation layer, a fourth fully connected layer, a fourth deformation layer, an intermediate layer, and multiple second graph convolutional layers connected in sequence as shown. The multiple first graph convolutional layers are used to extract a feature map corresponding to the graph structure from the information after splicing the 2D pose information and the 3D pose information; the third deformation layer, the fourth fully connected layer, and the fourth deformation layer are used to implement the mapping between the feature map of the graph structure and the feature map corresponding to the three-dimensional model, where the three-dimensional model includes multiple graph structures.
[0034] For ease of representation, Figure 1 a thick line is used in
[0035] to divide the intermediate layer part. The intermediate layer includes multiple second residual blocks and upsampling layers arranged alternately, where each second residual block includes multiple third graph convolutional layers connected in sequence. The intermediate layer is used to extract feature maps of different resolutions corresponding to the three-dimensional model. When performing the three-dimensional model reconstruction task, the second residual blocks are used to extract feature maps of corresponding resolutions, and the upsampling layers are used to connect feature maps of adjacent resolutions. The feature maps output by the intermediate layer are feature maps with higher dimensions.
[0036] It should be noted that the multiple first graph convolutional layers, the multiple second graph convolutional layers, and the multiple third graph convolutional layers connected in sequence in the second residual block all include a graph convolutional layer, a normalization layer, and an activation function layer.
[0037] To reduce the computational complexity, in the embodiments of the present application, each graph convolutional layer uses Chebyshev graph convolution, that is, the Chebyshev polynomial is used to construct the graph convolution algorithm. Among them, the Chebyshev polynomial includes the following formula:
[0038] T0(x) = 1; T1(x) = x;
[0039] T n+1 (x) = 2xT n (x) - T n-1 (x) (1)
[0040] Exemplarily, a graph topology structure G = (V, A, F) is constructed, where V represents vertices, A represents edges, and F represents the features of vertices. The normalized Laplacian operator of the graph topology structure G can be expressed as:
[0041] L = I - D -1 / 2 AD -1 / 2 (2)
[0042] Among them, D represents the degrees of each vertex in the graph topology structure G in the diagonal matrix, I is the identity matrix, A represents the edges in the graph topology structure G, and the scaled Laplacian operator corresponding to the graph G is expressed as:
[0043]
[0044] The operation of performing graph convolution on the graph topology structure G is defined as:
[0045]
[0046] In the above formula represents the input feature, represents the output feature, N represents the number of vertices in the graph topology structure G, K represents the use of the K-order Chebyshev polynomial, represents the feature change matrix, and R represents a real number. Therefore, substituting the above formulas (1), (2), and (3) into formula (4), the feature map corresponding to the graph topology structure G can be extracted.
[0047] Exemplarily, the embodiments of the present application provide a method for building an upsampling layer in the intermediate layer. The process of building the upsampling layer in the intermediate layer includes constructing a basic three-dimensional model grid and an upsampling process for extracting features from the above three-dimensional model grid.
[0048] First, construct a basic 3D model grid. The process of constructing the basic 3D model grid includes: Since the 3D model grid is a set of vertices, edges, and vertex features, that is, a set of graph structures. Therefore, define the graph topology as G = (V, A, F), where V ∈ R n×3 represents the vertex set in the 3D model grid, n represents the number of vertices in the 3D model grid, and R is the set of real numbers; F ∈ R n×f represents the features corresponding to the vertices of the 3D model grid, and f represents the dimension of the features; A ∈ {0, 1} n×n is an adjacency matrix. For example, assume that vertex i and vertex j are two vertices in the adjacency matrix. If vertex i and vertex j are connected, then A i,j = 1, and conversely, if vertex i and vertex j are not connected, then A i,j = 0.
[0049] Then, construct a sampling process corresponding to the above basic 3D model grid. Exemplarily, this sampling process can be obtained by improving the existing surface simplification algorithm, that is, on the basis of the surface simplification algorithm, set the following conditions: ① (v i , v j ) is a valid vertex pair if and only if (v i , v j ) is an edge; ② The vertex pair contraction condition is that the value of is either v i or v j . At the same time, each time select the vertex pair with the smallest contraction cost for contraction, and perform vertex pair contraction iteratively in a loop to obtain the sampled target grid. Among them, the contraction cost of the vertex pair is shown in formula (5):
[0050]
[0051] In the above formula (5), represents the coordinate vector corresponding to the contraction result of the vertex pair (v i , v j ), and Q i and Q j are the 4×4 symmetric matrices corresponding to vertices v i and v j respectively.
[0052] During the sampling process, the contraction of vertex pairs will lose and / or reorganize some points, edges, and mesh faces on the original grid (i.e., the basic 3D model grid), which is achieved by changing the adjacency matrix A in the graph topology G. Therefore, a corresponding downsampling matrix D ∈ {0, 1} can be established based on the above sampling process n×m and m > n.
[0053] Exemplarily, if vertex vp is retained during downsampling, then the vertex v p corresponding downsampling matrix D(q,p) = 1; if the vertex v p is discarded during downsampling, then the vertex v p corresponding downsampling matrix D(q,p) = 0, where p and q respectively represent the indices corresponding to the vertex.
[0054] Among them, establishing the corresponding downsampling process includes: defining the graph topology structure corresponding to the 3D model as: defining the corresponding downsampling matrix as: where represents rounding down, represents the number of vertices, and C represents the sampling frequency.
[0055] According to the above steps, define the characteristics of the vertices during downsampling as: or where According to V c = D c V c+1 it can be obtained that
[0056] Finally, after constructing the downsampling process corresponding to the basic 3D model grid, by performing an inverse mapping on the above downsampling process, an upsampling process corresponding to the basic 3D model grid can be obtained. The construction of the upsampling process includes:
[0057] Set the conditions corresponding to the upsampling process: ① For the vertices retained during downsampling, they are also retained during upsampling, that is, if D(q,p) = 1, then the corresponding U(p,q) = 1; ② For the vertices v p discarded during downsampling, that is, D(q,p) = 0, then in the corresponding upsampling process, project the vertex v p onto the barycentric coordinates of the grid (i,j,k) closest to the 3D model grid M c+1 during downsampling, where the barycentric coordinates can be obtained through and w + w i + w j + w k = 1, and v i , v j , v k ∈V c+1 , so U(p,i) = w i , U(p,j) = w j , U(p,k) = w k . In this way, the upsampling matrix can be established The vertex features of the 3D model grid corresponding to the upsampling process can be defined as
[0058] In summary, based on the above basic 3D model grid and the corresponding upsampling process, an upsampling layer in the middle layer as shown in Figure 1 can be constructed. By combining with the second residual block, a feature map of one resolution can be extracted. In order to extract the features of 3D model grids with different resolutions, the second residual blocks and upsampling layers corresponding to the types of resolutions can be set.
[0059] Exemplarily, referring to Figure 1 , in the embodiment of the present application, five second residual blocks and four upsampling layers are alternately arranged in the middle layer, that is, a middle layer that can extract five feature maps with different resolutions is established. This building mode can effectively relieve the learning pressure of the network model.
[0060] It should be noted that the number of the second residual blocks and upsampling layers alternately arranged in the middle layer can be designed according to actual application requirements, and the present application does not make any limitation thereto.
[0061] It should be understood that according to actual experiments, by combining the third graph convolutional layer and the upsampling layer to build the middle layer, the computational complexity of the 3D reconstruction model can be effectively reduced, and the memory consumption of the hardware device corresponding to the 3D reconstruction model can be reduced.
[0062] It should be noted that the network model provided by the present application has universality. It can be applied to any 3D model reconstruction task or a task with the 3D model reconstruction effect as an evaluation index. For example, various model reconstruction tasks such as human 3D model reconstruction and 3D cultural relic model reconstruction.
[0063] It can be understood that for different 3D model reconstruction tasks, the initial 3D reconstruction model can be trained by designing corresponding training sets and loss functions, so as to obtain a 3D reconstruction model applicable to different 3D model reconstruction tasks.
[0064] According to actual application requirements, the execution entity for training the 3D reconstruction model and the execution entity for performing the 3D model reconstruction task using the 3D reconstruction model can be the same or different.
[0065] Next, taking the human 3D model reconstruction task as an example, the training process and effect of the 3D model provided by the present application will be exemplarily described.
[0066] First step, obtain a corresponding training set for the human body three-dimensional model reconstruction task. The training set can directly utilize the human body image samples in existing human body pose datasets. For example, the human body image samples in the Human 3.6M dataset, the Common Objects in Context (COCO) dataset, and / or the Multiperson Composited 3D Human Pose (MuCo-3DHP) dataset, etc. The training set can also be human body image samples collected by mobile phones, cameras, etc. The training set can also be human body image samples obtained from public video websites.
[0067] After obtaining the image samples for the human body three-dimensional model reconstruction task, detect the 2D key points of the human body in the obtained image samples to obtain the 2D pose information samples of the human body. Among them, the 2D pose information samples of the human body include the position coordinates of multiple key points corresponding to the human body. It should be understood that the method for obtaining the 2D pose information samples of the human body can be to implement the detection of 2D key points of the human body through an open-source human body pose recognition project, so as to obtain the 2D pose information samples of the human body. Exemplarily, open-source human body pose recognition projects such as OpenPose, AlphPose, and High-Resolution Network (HRNet), etc.
[0068] According to the actual application, the method for obtaining the 2D pose information samples can also be to directly obtain the 2D pose information samples from the dataset. Exemplarily, use the data in the Archive of Motion Capture as Surface Shapes (AMASS) dataset to generate 2D pose information samples. The process of obtaining the 2D pose information samples includes: first, use the data in the AMASS dataset to obtain the corresponding human body mesh, and extract the 3D pose information samples of the human body from the obtained human body mesh; then project the above-extracted 3D pose information samples of the human body according to the predefined camera parameters to obtain the 2D pose information samples.
[0069] In one example, the present application uses the open-source human body pose recognition project HRNet to predict the position coordinates of each key point in the image sample in a preset human body joint tree, so as to obtain the 2D pose information of the human body. It should be understood that since the open-source human body pose recognition project HRNet is used to obtain the 2D pose information of the human body, it can avoid selecting the key points applicable to the subsequent three-dimensional reconstruction model from the relatively large number of predicted human body key points when using other open-source projects, and further shorten the time for building the human body three-dimensional reconstruction model.
[0070] In a possible implementation, the human joint tree can be composed of human key points in different formats (such as COCO format, humam3.6M format, or MPII format, etc.), or the human joint tree can also be custom-designed by the user according to actual application requirements. This application does not make any restrictions on this.
[0071] To improve the universality of the application, expand the reusability of data, and meet the actual application requirements, refer to Figure 2 (a) in this document shows the human joint tree in COCO format used in this application. Two new key points are added to it. The two new key points are shown as the 17th key point and the 18th key point in Figure 2 (b). The 17th key point is determined by taking the average of the 11th and 12th key points. Similarly, the 18th key point is determined by taking the average of the 5th and 6th key points. Therefore, this application obtains the 2D pose information of the human body based on the human joint tree with a total of 19 key points (i.e., 0, 1, ……, 17, 18) and the human pose recognition project HRNet.
[0072] It should be understood that according to different actual application requirements, different human joint trees and / or different human pose recognition projects (for example, custom human pose recognition methods) can be used to obtain the 2D pose information of the human body. When the tasks of 3D model construction are different, different methods can be used to obtain the 2D pose information of different target objects. This application does not make any restrictions on this.
[0073] In the second step, the initial 3D reconstruction model is iteratively trained using the human 2D pose information samples obtained from the training set and a preset loss function to obtain a 3D reconstruction model. The preset loss function is used to describe the loss between the predicted human 3D model and the real human 3D model samples.
[0074] After the initial 3D reconstruction model is built, the obtained human 2D pose information samples are input into the initial 3D reconstruction model. The initial 3D reconstruction model processes the human 2D pose information samples to obtain a predicted human 3D model.
[0075] For the human 3D model reconstruction task, exemplarily, the following loss functions are used to train the 3D pose initial regression network and the vertex initial regression network in the initial 3D reconstruction model respectively.
[0076] First of all, according to actual applications, it can be obtained that the 2D pose information of the human body obtained by this application using the open-source human pose recognition project HRNet usually contains errors. Therefore, in order to improve the robustness of the 3D pose regression network in processing the 2D pose information of the human body, when training the 3D pose initial regression network, the input human 2D pose information includes human 2D pose information samples and errors.
[0077] Then, based on the Human 3.6M dataset and the COCO dataset, this application first trains the 3D pose initial regression network, and a trained 3D pose updated regression network can be obtained. The loss function used is as follows:
[0078] L pose = ||P 3D - P 3D* ||1 (6)
[0079] In the above formula (6), P 3D represents the human 3D pose information predicted by the 3D pose initial regression network, and P 3D* represents the human 3D pose information in the human 2D pose information sample. Based on the above dataset and the loss function in formula (6), the 3D pose initial regression network is iteratively trained until the network converges, and then the trained 3D pose updated regression network can be obtained.
[0080] According to the actual experimental data, this application iteratively trains the 3D pose initial regression network 60 times and then obtains the trained 3D pose updated regression network.
[0081] Finally, based on the Human 3.6M dataset, the COCO dataset, and the AMASS dataset, the trained 3D pose updated regression network and the vertex initial regression network are trained to obtain the trained 3D pose regression network and the trained vertex regression network, that is, the human three-dimensional reconstruction model. The loss functions used in the above training process are shown in formula (7) and include mesh loss, 3D loss, mesh surface normal loss, and mesh surface edge loss.
[0082] loss = λ a L v + λ b L j + λ c L n + λ d L e (7)
[0083] Among them, λ a , λ b , λ c and λ d are constants; L v = ||M - M * ||1 represents the human mesh loss value, M * represents the true value of the human mesh, and M represents the predicted value of the human mesh; L j = ||JM - J 3D* ||1 represents the human 3D joint loss value, J 3D*Represents the true value of the 3D joints of the human body, JM represents the predicted value of the 3D joints of the human body, and J ∈ R v×N is a matrix of the 3D joints of the human body extracted from the human mesh; Represents the surface normal loss of the human mesh, f represents the triangular faces of the human mesh, represents the unit normal vector of f, m i and m j respectively represent the coordinates of two vertices in f; Represents the surface edge loss of the human mesh.
[0084] Based on the above training set and the loss function in formula (7), the trained 3D pose update regression network and vertex initial regression network can be iteratively trained. When the network converges, a trained 3D reconstruction model can be obtained.
[0085] Similarly, according to the actual experimental data, in this application, after iteratively training the trained 3D pose update regression network and vertex initial regression network 15 times, the network converges, and finally a trained 3D reconstruction model is obtained.
[0086] In an example, for the human 3D model reconstruction task, after obtaining the 2D pose information of the human body in this application, the obtained 2D pose information of the human body is input into the trained 3D reconstruction model for processing to obtain the human 3D reconstruction model. Among them, the process of the trained 3D reconstruction model processing the 2D pose information of the human body further includes:
[0087] ① Perform standard normalization processing on the obtained 2D pose information of the human body, that is, subtract the average value from the obtained 2D pose information of the human body and then divide by the standard deviation.
[0088] ② Input the 2D pose information of the human body after standard normalization processing into the 3D pose regression network for processing to obtain 3D pose information. The 3D pose regression network includes, as Figure 1 shown, a first deformation layer, a first fully connected layer, two first residual blocks, a second fully connected layer, and a second deformation layer connected in sequence. Among them, both of the two first residual blocks include a normalization layer, an activation function layer, and a third fully connected layer connected in sequence.
[0089] In this application, the 2D pose information of the human body is defined as P 2D ∈ R J×2 , R is a real number, J represents the number of key points in the human joint tree. Since the human joint tree used in this application includes 19 key points, therefore, J = 19.
[0090] The first deformation layer expands the corresponding 2D pose information of the human body into a feature vector, that is, R 19×2 → R 38The first fully-connected layer is used to map the 38-dimensional feature vector into a 4096-dimensional feature vector, i.e., R 38 →R 4096 Two first residual blocks are used to learn 4096-dimensional features according to different 2D pose information. The second fully-connected layer is used to map the 4096-dimensional feature vector into a 57-dimensional feature vector, i.e., R 57 The second deformation layer corresponds to the first deformation layer and is used to convert the 57-dimensional feature vector into the corresponding 3D pose information, i.e., through R 57 →R 19×3 to obtain P 3D ∈R J×3 .
[0091] ③ Concatenate the 3D pose information and 2D pose information obtained in the second step, and then input them into the trained vertex regression network for processing to obtain a three-dimensional model corresponding to the human body. The trained vertex regression network includes three first graph convolutional layers, a third deformation layer, a fourth fully-connected layer, a fourth deformation layer, an intermediate layer, and two second graph convolutional layers connected in sequence as shown in Figure 1 . Among them, the intermediate layer includes five second residual blocks and four upsampling layers, and the above five second residual blocks and four upsampling layers are alternately arranged.
[0092] Exemplarily, for the human body three-dimensional model reconstruction task, the embodiment of the present application can generate a human body mesh topology structure corresponding to the human body three-dimensional model according to the Skinned Multi-Person Linear (SMPL) template, and the human body mesh topology structure is defined as: The corresponding downsampling matrix is defined as: where represents rounding down, represents the number of vertices, and the sampling frequency C is set to 4 during the downsampling process. Since the number of vertices of the human body mesh of the SMPL template is 6890, when the sampling frequency of the downsampling is set to 4 (i.e., D0, D1, D2, and D3), the human body mesh effect as shown in Figure 3 can be generated. represents the human body mesh topology structure with 6890 vertices, represents the human body mesh topology structure with 1723 vertices obtained after the first downsampling, represents the human body mesh topology structure with 431 vertices obtained after the second downsampling, represents the human body mesh topology structure with 108 vertices obtained after the third downsampling, represents the human body mesh topology structure with 27 vertices obtained after the fourth downsampling.
[0093] In this application, after splicing the 3D pose information and the 2D pose information, the spliced 5D pose information is obtained, that is, P ∈ R J×5 ; Three first graph convolutional layers are used to obtain the feature map F of the corresponding graph structure G=(V, A, F) from the spliced 5D pose information P , where F P is initialized to P, that is, F P = P ∈ R J×5 , and after three first graph convolutional processes, F P ∈ R J×64 is output. That is to say, the feature of each vertex obtained after three first graph convolutional processes is 64-dimensional.
[0094] Input F P ∈ R J×64 into the third deformation layer, the fourth fully connected layer, and the fourth deformation layer for processing, and the mapping of the vertex feature map in the graph structure can be completed, that is, the G in the graph structure G=(V, A, F) P corresponding feature map F P ∈ R J×64 is mapped to the corresponding feature map on the human body mesh topology , where |V4| represents the number of vertices in the human body mesh topology .
[0095] Based on the inverse mapping of the downsampling process, an upsampling layer applied to the intermediate layer in this application is constructed, and the obtained upsampling matrix group is The adjacent two second residual blocks are connected by an upsampling layer. In this way, according to the predefined topology successively pass through for convolution operations, so as to extract the corresponding feature map F based on the human body mesh topology with different numbers of vertices c , that is, F c = U c F c+1 , c = [3, 2, 1, 0]. After being processed by the five second residual blocks in the intermediate layer, the feature map F0 ∈ R corresponding to the human body mesh topology is finally obtained 6890×128 , where 6890 is the number of vertices corresponding to the human body mesh topology , and the feature dimension of each vertex is 128-dimensional.
[0096] Two second graph convolutional layers in the vertex regression network are used to reduce the dimension of the feature of each vertex in the feature map output by the intermediate layer, that is, the feature map F0 ∈ R corresponding to the human body mesh topology 6890×128The feature dimension of each vertex in [model name] is reduced from 128 dimensions to 3 dimensions. The feature map output after being processed by two second graph convolutional layers is the human body mesh topology The corresponding V0 ∈ R in [model name] N×3 , that is, the position information of the mesh vertices in the human body mesh topology. After obtaining V0 ∈ R N×3 , the three-dimensional human body model can be reconstructed.
[0097] Next, take Figure 1 the three-dimensional reconstruction model shown as an example to illustrate the performance of the three-dimensional reconstruction model provided by this application.
[0098] Table 1
[0099]
[0100] As shown in Table 1, it is the comparison result of the three-dimensional reconstruction model provided by this application and other methods trained with the same training set. The three-dimensional reconstruction model provided by this application and other methods are all trained using the same data set. The training data sets are the COCO data set and the Human3.6M data set, and the test sets use the Human3.6M data set and the 3DPW data set. In Table 1, other methods include: Human Mesh Recovery (HMR), the SPIN method that learns to reconstruct the 3D human pose and shape by fitting a cyclic model, Graph Convolutional Mesh Regression (GraphCMR), and the Pose2Mesh method of the graph convolutional network for 3D human pose and mesh recovery from 2D human poses. The Mean Per Joint Position Error (MPJPE) represents the accuracy of human pose estimation, that is, the error between the predicted 3D joint position and the true value. The lower the value of MPJPE, the more accurate the prediction; PA MPJPE represents the average error of the joints after doing the alignment transformation; MPVPE represents the average error of the predicted three-dimensional human body mesh and the positions of each vertex of the true three-dimensional human body mesh.
[0101] It can be seen from Table 1 that under the same test set, the three-dimensional reconstruction model provided by this application has very good experimental results.
[0102] As shown in Table 2, it is the comparison result of the three-dimensional reconstruction model provided by this application and other methods after being trained with their respective corresponding training sets on the Human 3.6M data set and the 3D Poses in the Wild (3DPW) data set. This application uses the Human 3.6M data set, the COCO data set, and the AMASS data set to train the three-dimensional reconstruction model.
[0103] Table 2
[0104]
[0105] Since 3DPW is a dataset collected in an outdoor environment and Human3.6M is a dataset collected in a laboratory environment, as can be seen from Table 2, after different methods are trained using their respective corresponding training sets, it shows that the 3D reconstruction model provided by this application has good experimental results in a dataset under a non-laboratory environment.
[0106] As shown in Table 3, it is the comparison result of the 3D reconstruction model provided by this application and other methods in other performance aspects. Among them, GPU mem reflects the usage of video memory when the number of samples in each batch of the training set during the training process is 64; No.param represents the number of parameters of the network model; Avg.time represents the average time of model inference.
[0107] Table 3
[0108] methods GPU mem No.Param Avg.time MPJPE PA MPJPE Pose2mesh 6G 8.8M 132ms 64.9mm 48.7mm 3D reconstruction model 0.9G 2.5M 34ms 64.6mm 47.4mm
[0109] As can be seen from Table 3, the 3D reconstruction model provided by this application has good experimental results.
[0110] Such as Figure 4 is an example of a human 3D model obtained after processing different pictures using the human 3D reconstruction model provided by this application. From Figure 4 the four examples listed, it can be seen that the method for processing the 2D pose information of the target object in the to-be-processed image by the 3D reconstruction model provided by this application can make full use of the pose information in different dimensions, making the vertex coordinate information corresponding to the obtained 3D model more accurate. Compared with the prior art, it can significantly improve the accuracy of the 3D model.
[0111] Based on the same inventive concept, the embodiment of this application also provides a 3D reconstruction device. Such as Figure 5 shown, the embodiment of this application also provides a 3D reconstruction device. The 3D reconstruction device 100 includes:
[0112] An acquisition unit 101, configured to acquire 2D pose information of a target object, where the 2D pose information includes the position coordinates of multiple key points corresponding to the target object;
[0113] A processing unit 102, configured to input the 2D pose information of the target object into a trained 3D reconstruction model for processing to obtain a 3D model of the target object; the 3D reconstruction model includes a 3D pose regression network and a vertex regression network; the 3D pose regression network is configured to extract 3D pose information from the 2D pose information; after splicing the 2D pose information and the 3D pose information, input them into the vertex regression network for processing to obtain the 3D model of the target object.
[0114] In one embodiment, the 3D pose regression network includes a first deformation layer, a first fully connected layer, a plurality of first residual blocks, a second fully connected layer, and a second deformation layer connected in sequence; the first deformation layer is used to convert the position coordinates of multiple key points into a feature vector in a preset format; the second deformation layer is used to convert the data output by the second fully connected layer into 3D pose information.
[0115] In one embodiment, the plurality of first residual blocks include a normalization layer, an activation function layer, and a third fully connected layer connected in sequence.
[0116] In one embodiment, the vertex regression network includes a plurality of first graph convolutional layers, a third deformation layer, a fourth fully connected layer, a fourth deformation layer, an intermediate layer, and a plurality of second graph convolutional layers connected in sequence. The intermediate layer includes a plurality of second residual blocks and upsampling layers alternately arranged; the third deformation layer, the fourth fully connected layer, and the fourth deformation layer are used to map the feature map output by the plurality of first graph convolutional layers into the intermediate layer.
[0117] In one embodiment, the second residual block includes a plurality of graph convolutional layers connected in sequence.
[0118] In one embodiment, the process of obtaining the three-dimensional reconstruction model includes: training the initial 3D pose regression network using a first loss function and a first training set to obtain an updated 3D pose regression network; training the updated 3D pose regression network and the initial vertex regression network using a second loss function and a second training set to obtain the three-dimensional reconstruction model.
[0119] In one embodiment, the second loss function includes a mesh loss, a 3D loss, a mesh surface normal loss, and a mesh surface edge loss.
[0120] Those skilled in the art can clearly understand that, for the convenience and simplicity of description, only the above-mentioned division of each functional unit and module is used as an example. In practical applications, the above functions can be allocated to different functional units and modules according to needs, that is, the internal structure of the device is divided into different functional units or modules to complete all or part of the functions described above. Each functional unit and module in the embodiment can be integrated into a processing unit, or each unit can exist physically alone, or two or more units can be integrated into one unit. The above integrated unit can be implemented in the form of hardware or in the form of a software functional unit. In addition, the specific names of each functional unit and module are only for the convenience of mutual distinction and do not limit the protection scope of the present application. The specific working processes of the units and modules in the above system can refer to the corresponding processes in the foregoing method embodiments and will not be repeated here.
[0121] Based on the same inventive concept, an embodiment of the present application further provides a terminal device, and the terminal device 200 is as follows Figure 6 shown.
[0122] As Figure 6 shown, the terminal device 200 of this embodiment includes: a processor 201, a memory 202, and a computer program 203 stored in the memory 202 and executable on the processor 201. The computer program 203 can be run by the processor 201 to generate instructions, and the processor 201 can implement the steps in the above-mentioned embodiments of each privilege authentication method according to the instructions. Alternatively, when the processor 201 executes the computer program 203, it implements the functions of each module / unit in the above-mentioned device embodiments.
[0123] Exemplarily, the computer program 203 can be divided into one or more modules / units. One or more modules / units are stored in the memory 202 and executed by the processor 201 to complete the present application. One or more modules / units can be a series of computer program instruction segments capable of performing specific functions, and these instruction segments are used to describe the execution process of the computer program 203 in the terminal device 200.
[0124] Those skilled in the art can understand that Figure 6 merely examples of the terminal device 200 do not constitute a limitation on the terminal device 200, and it may include more or fewer components than shown in the figure, or combine some components, or different components. For example, the terminal device 200 may further include input / output devices, network access devices, buses, etc.
[0125] The processor 201 may be a central processing unit (CPU), or may also be other general-purpose processors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc.
[0126] The memory 202 may be an internal storage unit of the terminal device 200, such as the hard disk or memory of the terminal device 200. The memory 202 may also be an external storage device of the terminal device 200, such as a plug-in hard disk equipped on the terminal device 200, a Smart Media Card (SMC), a Secure Digital (SD) card, a Flash Card, etc. Further, the memory 202 may also include both the internal storage unit and the external storage device of the terminal device 200. The memory 202 is used to store computer programs and other programs and data required by the terminal device 200. The memory 202 may also be used to temporarily store data that has been output or will be output.
[0127] The terminal device provided in this embodiment may execute the above method embodiment, and its implementation principle and technical effect are similar, which will not be elaborated here.
[0128] This application embodiment also provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the method of the above method embodiment is implemented.
[0129] This application embodiment also provides a computer program product. When the computer program product runs on the terminal device, the terminal device is enabled to execute the method of the above method embodiment.
[0130] If the above integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it may be stored in a computer-readable storage medium. Based on such an understanding, to implement all or part of the processes in the above method embodiments of this application, a computer program may be used to instruct relevant hardware to complete. The computer program may be stored in a computer-readable storage medium. When the computer program is executed by a processor, the steps of the above method embodiments may be implemented. Among them, the computer program includes computer program code, and the computer program code may be in the form of source code, object code, executable file or some intermediate form, etc. The computer-readable storage medium may at least include: any entity or device that can carry the computer program code to the photographing device / terminal device, a recording medium, a computer memory, a Read-Only Memory (ROM), a Random Access Memory (RAM), an electrical carrier signal, a telecommunication signal, and a software distribution medium. For example, a USB flash drive, a mobile hard disk, a magnetic disk, or an optical disc, etc.
[0131] References to "an embodiment" or "some embodiments" etc. described in this application mean that a particular feature, structure, or characteristic described in connection with the embodiment is included in one or more embodiments of this application. Thus, statements such as "in one embodiment", "in some embodiments", "in other some embodiments", "in still other embodiments", etc. that appear in different places in this specification are not necessarily all referring to the same embodiment, but mean "one or more but not all embodiments", unless otherwise specifically emphasized. The terms "comprising", "including", "having" and their variants all mean "including but not limited to", unless otherwise specifically emphasized.
[0132] In the description of this application, it should be understood that the terms "first" and "second" are only used for descriptive purposes and cannot be construed as indicating or implying relative importance or implicitly specifying the quantity of the indicated technical features. Thus, features defined with "first" and "second" may explicitly or implicitly include at least one of such features.
[0133] In addition, in this application, unless otherwise clearly defined and limited, terms such as "connected", "coupled" etc. should be understood in a broad sense. For example, it may be a mechanical connection or an electrical connection; it may be a direct connection or an indirect connection through an intermediate medium, and it may be the internal connection of two components or the interaction relationship between two components, unless otherwise clearly limited. For those of ordinary skill in the art, the specific meanings of the above terms in this application can be understood according to specific circumstances.
[0134] The above embodiments are only used to illustrate the technical solutions of this application, rather than to limit them; although this application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements on some or all of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the scope of the technical solutions of the embodiments of this application.
Claims
1. A three-dimensional reconstruction method, characterized in that, The method includes: Obtaining 2D pose information of a target object, where the 2D pose information includes the position coordinates of multiple key points on the target object; Inputting the 2D pose information into a trained three-dimensional reconstruction model for processing to obtain a three-dimensional model of the target object; The three-dimensional reconstruction model includes a 3D pose regression network and a vertex regression network; the 3D pose regression network is used to extract 3D pose information from the 2D pose information; the vertex regression network is used to process the concatenated information of the 2D pose information and the 3D pose information to obtain the three-dimensional model of the target object; The 3D pose regression network includes a first deformation layer, a first fully connected layer, multiple first residual blocks, a second fully connected layer, and a second deformation layer connected in sequence; the first deformation layer is used to convert the position coordinates of the multiple key points into a feature vector in a preset format; the second deformation layer is used to convert the feature information output by the second fully connected layer into the 3D pose information; The multiple first residual blocks include a normalization layer, an activation function layer, and a third fully connected layer connected in sequence; The vertex regression network includes multiple first graph convolutional layers, a third deformation layer, a fourth fully connected layer, a fourth deformation layer, an intermediate layer, and multiple second graph convolutional layers connected in sequence, and the intermediate layer includes multiple second residual blocks and upsampling layers arranged alternately; The third deformation layer, the fourth fully connected layer, and the fourth deformation layer are used to map the feature map input output by the multiple first graph convolutional layers into the intermediate layer; The second residual block includes multiple graph convolutional layers connected in sequence.
2. The method according to claim 1, wherein The process of obtaining the three-dimensional reconstruction model includes: Training an initial 3D pose regression network using a first loss function and a first training set to obtain an updated 3D pose regression network; Training the updated 3D pose regression network and an initial vertex regression network using a second loss function and a second training set to obtain the three-dimensional reconstruction model.
3. The method according to claim 2, wherein The second loss function includes a mesh loss, a 3D loss, a mesh surface normal loss, and a mesh surface edge loss.
4. A three-dimensional reconstruction device, characterized in that, The device includes: An acquisition unit for acquiring 2D pose information of a target object, where the 2D pose information includes the position coordinates of multiple key points on the target object; A processing unit for inputting the 2D pose information into a trained three-dimensional reconstruction model for processing to obtain a three-dimensional model of the target object; The three-dimensional reconstruction model includes a 3D pose regression network and a vertex regression network; the 3D pose regression network is used to extract 3D pose information from the 2D pose information; the vertex regression network is used to process the concatenated information of the 2D pose information and the 3D pose information to obtain the three-dimensional model of the target object; The 3D pose regression network includes a first deformation layer, a first fully connected layer, multiple first residual blocks, a second fully connected layer, and a second deformation layer connected in sequence; the first deformation layer is used to convert the position coordinates of the multiple key points into a feature vector in a preset format; the second deformation layer is used to convert the feature information output by the second fully connected layer into the 3D pose information; The multiple first residual blocks include a normalization layer, an activation function layer, and a third fully-connected layer that are connected in sequence; The vertex regression network includes a plurality of first graph convolutional layers, a third deformation layer, a fourth fully-connected layer, a fourth deformation layer, an intermediate layer, and a plurality of second graph convolutional layers that are connected in sequence. The intermediate layer includes a plurality of second residual blocks and upsampling layers that are alternately arranged; The third deformation layer, the fourth fully-connected layer, and the fourth deformation layer are used to map the feature map output by the plurality of first graph convolutional layers into the intermediate layer; The second residual block includes a plurality of graph convolutional layers that are connected in sequence.
5. A terminal device, characterized in that, The device includes: a processor and a memory. The memory is used to store a computer program, and the processor is used to call and run the computer program from the memory, so that the device executes the method according to any one of claims 1 to 3.
6. A computer-readable storage medium, characterized in that, A computer program is stored in the computer-readable storage medium. When the computer program is executed by a processor, the processor is caused to execute the method according to any one of claims 1 to 3.