Method, apparatus, device and storage medium for reconstructing a three-dimensional model from three-view drawings
By obtaining back projection and constructing candidate edges from three views, combining deep neural networks and parameterized model fitting, the efficiency and complexity problems when reconstructing complex three-dimensional models in the existing technology are solved, and efficient three-dimensional model reconstruction is achieved.
Patent Information
- Application Number
- CN202210328333.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-03-30
- Publication Date
- 2025-05-27
- Estimated Expiration
- 2042-03-30
AI Technical Summary
When reconstructing a three-dimensional model from three views, it is difficult to effectively deal with complex objects, especially objects formed by higher order curves or B-spline curves, resulting in high search costs, complex processes and low efficiency.
By obtaining the reverse projection of the two-dimensional edges of each view in three-dimensional space, a three-dimensional candidate edge with the reverse projection of three different views intersecting, and a deep neural network is used to construct the candidate edges to form a loop of the actual face, and finally obtain the three-dimensional model by fitting the parameterized model.
The simple and efficient reconstruction of three-dimensional models in CAD is achieved, avoiding limitations on geometric shape types, expanding the scope of application, and improving the efficiency of complex object reconstruction.
Smart Images

Figure CN114693874B_ABST
Abstract
Description
Technical Field
[0001] The present disclosure generally relates to the technical field of model construction, and particularly relates to a method, apparatus, device and storage medium for reconstructing a three-dimensional model from three views. Background Art
[0002] Reconstructing a three-dimensional object from three orthographic projection views has been a long-standing problem in the field of computer-aided design. A successful solution to this problem would enable users to conveniently create a three-dimensional object model by drawing sketches on a two-dimensional interactive interface.
[0003] In the prior art, orthographic projection is one of the most commonly used methods in three-dimensional model construction. Its steps are as follows: First, all possible three-dimensional vertices generated from the two-dimensional vertices in the orthographic views are obtained. These methods gradually enumerate all possible three-dimensional edges, three-dimensional faces, and solid bodies, and then re-project these candidates onto each view to check for matches. The disadvantages of this method are as follows: The types of objects are limited to some common types of faces, including planes and limited quadratic surfaces, such as cylinders, spheres, and tori. When complex objects appear, such as higher-order curves or B-spline curves formed by two intersecting quadratic surfaces, the search cost becomes extremely high. Additionally, the process of searching for all valid faces, especially curved surfaces, is very complex, and each step often requires multiple re-projections to check the necessary conditions. Due to the ambiguity of the intermediate steps, backtracking and heuristic methods may also be adopted, resulting in a complex process and low efficiency. Summary of the Invention
[0004] In view of the above-mentioned defects or deficiencies in the prior art, it is desirable to provide a method, apparatus, device and storage medium for reconstructing a three-dimensional model from three views, which can meet the specific requirements of currently reconstructing a three-dimensional model from three views.
[0005] Based on one aspect of the embodiments of the present invention, embodiments of the present application provide a method for reconstructing a three-dimensional model from three views, the method comprising:
[0006] Obtaining the back-projection of the two-dimensional edges of each view in three-dimensional space;
[0007] According to the back-projection of the two-dimensional edges of each view in three-dimensional space, obtaining three-dimensional candidate edges constructed by all points where the back-projections of three different views intersect, the three different views including a front view, a top view, and a side view;
[0008] According to the three-dimensional candidate edges constructed by all points where the back-projections of the three different views intersect, obtaining a loop of actual faces formed by the candidate edges constructed through a deep neural network, and each loop of the actual faces corresponds to an actual face of a three-dimensional space object;
[0009] Obtain a three-dimensional model obtained by fitting a parametric model based on the loop of the actual surface formed by the candidate edges constructed by the deep neural network.
[0010] In another embodiment, the three-dimensional candidate edges constructed from all the points where the back-projections of the two-dimensional edges of each view in three-dimensional space intersect include:
[0011] Obtain a plurality of three-dimensional voxels for voxelizing the three-dimensional space;
[0012] Based on the plurality of three-dimensional voxels for voxelizing the three-dimensional space, obtain the positions of the two-dimensional views onto which each of the three-dimensional voxels is projected in three different views;
[0013] Based on the positions of the two-dimensional views onto which each of the three-dimensional voxels is projected in three different views, obtain the matching line information of each of the three-dimensional voxels in the two-dimensional views of each of the different views;
[0014] Based on the matching line information of each of the three-dimensional voxels in the two-dimensional views of each of the different views, obtain the three-dimensional candidate voxels in the three-dimensional voxels;
[0015] Based on the three-dimensional candidate voxels, obtain the three-dimensional candidate voxels having the same matching line in the two-dimensional views of the three different views and construct them into three-dimensional candidate edges.
[0016] In another embodiment, the three-dimensional candidate edges are associated with the matching lines of three different views of the two-dimensional view.
[0017] In another embodiment, obtaining the loop of the actual surface formed by the candidate edges constructed by the deep neural network based on the three-dimensional candidate edges constructed from all the points where the back-projections of the three different views intersect includes:
[0018] Based on the three-dimensional candidate edges constructed from all the points where the back-projections of the three different views intersect, obtain the common edges for constructing the surfaces of the three-dimensional space, and the common edge is a candidate edge shared by two adjacent three-dimensional space surfaces;
[0019] Based on the common edges for constructing the surfaces of the three-dimensional space and the Pointer Net deep neural network, obtain the loop of the actual surface constructed by the three-dimensional candidate edges.
[0020] In another embodiment, the Pointer Net deep neural network includes:
[0021] Use an encoder to obtain a context embedding vector for each input vector, and the input vector is composed of the common edges;
[0022] At each decoding step, the decoder outputs a pointer vector;
[0023] The pointer vector is compared with the context embedding vector through dot product, and the comparison result scores are normalized using the softmax algorithm to obtain a valid probability distribution over the input set;
[0024] The parameters of the model are learned by maximizing the conditional probability of the training set.
[0025] In another embodiment, when the encoder is used to obtain context embedding vectors for each input vector, each input vector is composed of common edges, and each common edge embedding vector includes:
[0026] Value embedding, representing the coordinate values of the common edge;
[0027] Position embedding, representing the position of the input vector.
[0028] Based on another aspect of the embodiments of the present invention, a device for reconstructing a three-dimensional model from three views is disclosed. The device includes:
[0029] A candidate line generation module, configured to obtain the back-projection of the two-dimensional edges of each view in three-dimensional space; based on the back-projection of the two-dimensional edges of each view in three-dimensional space, obtain three-dimensional candidate edges constructed by all the intersection points of the back-projections of three different views, where the three different views include a front view, a top view, and a side view;
[0030] A face detection module based on a deep neural network, configured to obtain a loop of the actual faces formed by the candidate edges constructed by the deep neural network based on the three-dimensional candidate edges constructed by all the intersection points of the back-projections of the three different views, and each loop of the actual faces corresponds to an actual face of the three-dimensional space object;
[0031] A model reconstruction module, configured to obtain a three-dimensional model obtained by fitting a parameterized model based on the loop of the actual faces formed by the candidate edges constructed by the deep neural network.
[0032] Based on yet another aspect of the embodiments of the present invention, an electronic device is disclosed. The electronic device includes one or more processors and a memory. The memory is used to store one or more programs; when the one or more programs are executed by the processor, the processor implements the method for reconstructing a three-dimensional model from three views provided in the embodiments of the present invention.
[0033] Based on yet another aspect of the embodiments of the present invention, a computer-readable storage medium storing a computer program is disclosed. When the computer program is executed, it implements the method for reconstructing a three-dimensional model from three views provided in the embodiments of the present invention.
[0034] In the embodiments of the present application, by obtaining the reverse projection of the two-dimensional edges of each view in three-dimensional space; according to the reverse projection of the two-dimensional edges of each view in three-dimensional space, obtaining three-dimensional candidate edges constructed by all the intersection points of the reverse projections of three different views; according to the three-dimensional candidate edges constructed by all the intersection points of the reverse projections of the three different views, obtaining the loop of the actual surface formed by the candidate edges constructed by a deep neural network; according to the loop of the actual surface formed by the candidate edges constructed by the deep neural network, obtaining a three-dimensional model obtained by fitting a parametric model. The present application can solve the problem of three-dimensional model reconstruction in CAD in a simple and efficient way, avoiding the limitation on the type of geometric shape when restoring the topological information of an object, and having a wider application range. BRIEF DESCRIPTION OF THE DRAWINGS
[0035] By reading the detailed description of the non-limiting embodiments with reference to the following drawings, other features, objects, and advantages of the present application will become more apparent:
[0036] Figure 1 is an application scenario diagram of a method for reconstructing a three-dimensional model from three views provided by an embodiment of the present application;
[0037] Figure 2 is a flowchart of a method for reconstructing a three-dimensional model from three views provided by an embodiment of the present application;
[0038] Figure 3 is a schematic structural diagram of a device for reconstructing a three-dimensional model from three views provided by an embodiment of the present application;
[0039] Figure 4 is an internal structure diagram of an electronic device provided by an embodiment of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0040] The following further describes the present application in detail with reference to the drawings and embodiments. It can be understood that the specific embodiments described herein are only used to explain the related invention, rather than limiting the invention. Additionally, it should be noted that for the sake of description, only the parts related to the invention are shown in the drawings.
[0041] It should be noted that, without conflict, the embodiments in the present application and the features in the embodiments can be combined with each other. The present application will be described in detail below with reference to the drawings and embodiments.
[0042] The method for reconstructing a three-dimensional model from three views provided by the present application can be applied to, for example Figure 1In the application environment shown. The method for reconstructing a three-dimensional model from three views is applied to an apparatus for reconstructing a three-dimensional model from three views. The apparatus for reconstructing a three-dimensional model from three views can be configured in the terminal 102 or the server 104, or partially configured in the terminal 102 and partially configured in the server 104, and the method for reconstructing a three-dimensional model from three views is completed through the interaction between the terminal 102 and the server 104.
[0043] Among them, the terminal 102 and the server 104 can communicate through a network.
[0044] Among them, the terminal 102 can be but is not limited to various personal computers, laptop computers, smart phones, tablet computers, and portable wearable devices. The terminal 102 needs to have the functions of receiving, viewing, editing, and sharing a shared three-dimensional model scene. The server 104 can be implemented by an independent server or a server cluster composed of multiple servers.
[0045] In one embodiment, as Figure 2 shown, a method for reconstructing a three-dimensional model from three views is provided. This embodiment mainly takes the application of this method to Figure 1 the terminal 102 in it as an example for illustration.
[0046] Please refer to Figure 2 , which shows an exemplary process of the method for reconstructing a three-dimensional model from three views that can apply the embodiments of the present application.
[0047] As Figure 2 shown, in step 210, the reverse projection of the two-dimensional edges of each view in three-dimensional space is obtained.
[0048] Specifically, the three-dimensional model of an object can be represented by surfaces. An object is constructed by a finite number of bounded surfaces. Each surface is represented by a set of edges that form its boundary, and each edge is represented by its two endpoints. A view consists of three orthogonal projections, namely the front view, the top view, and the side view. Each view can be regarded as a graph composed of two-dimensional edges and vertices. Each three-dimensional edge of the object must be projected onto one or more two-dimensional elements in each view.
[0049] In step 220, according to the reverse projection of the two-dimensional edges of each view in three-dimensional space, three-dimensional candidate edges constructed by all the points where the reverse projections of three different views intersect are obtained. The three different views include the front view, the top view, and the side view.
[0050] Specifically, the two-dimensional lines of the front view, the top view, and the side view are respectively represented as: Then three-dimensional candidate edges that may be located on the boundary of the object are generated. It is achieved by matching the parameterized lines or curves in different views and calculating the parameter representation in three-dimensional space.
[0051] Specifically, in an embodiment of the present application, the three-dimensional candidate edges constructed by obtaining all the intersection points of the reverse projections of the two-dimensional edges of each view in three-dimensional space include:
[0052] Obtaining a plurality of three-dimensional voxels for voxelizing the three-dimensional space;
[0053] Based on the plurality of three-dimensional voxels for voxelizing the three-dimensional space, obtaining the positions of each of the three-dimensional voxels projected onto the two-dimensional views of three different views;
[0054] Based on the positions of each of the three-dimensional voxels projected onto the two-dimensional views of three different views, obtaining the matching line information of each of the three-dimensional voxels in the two-dimensional views of each of the different views;
[0055] Based on the matching line information of each of the three-dimensional voxels in the two-dimensional views of each of the different views, obtaining the three-dimensional candidate voxels in the three-dimensional voxels;
[0056] Based on the three-dimensional candidate voxels, obtaining the three-dimensional candidate voxels having the same matching lines in the two-dimensional views of the three different views and constructing them into three-dimensional candidate edges.
[0057] Specifically, use a grid V with d 3 three-dimensional voxels to voxelize the three-dimensional space, project each three-dimensional voxel v ∈ V onto the two-dimensional views, and check whether it lies on any line. If a three-dimensional voxel has at least one matching line in each two-dimensional view, then the three-dimensional voxel is regarded as a three-dimensional candidate voxel, and all the three-dimensional candidate voxels having the same matching lines in the two-dimensional views of three different views are grouped into one candidate edge.
[0058] For example, a generated candidate edge can be expressed as Each edge e n is associated with a unique set of two-dimensional views (I n , J n , K n ) of three different views, where I n , J n and K n are respectively the index sets of the matching lines in the front view, top view, and side view. If the three-dimensional voxel projects to the intersection point of the two-dimensional lines in that view, each set may contain multiple indices. The set may contain multiple indices of lines, such as when these three-dimensional voxels project to the intersection points of the two-dimensional lines in the corresponding views. The three-dimensional candidate edge is associated with the matching lines of the three different views of the two-dimensional views.
[0059] In step 230, for the three-dimensional candidate edges constructed based on all the points where the back-projections of the three different views intersect, obtain the loops of the actual faces formed by the candidate edges constructed through a deep neural network, and each loop of the actual face corresponds to an actual face of the three-dimensional space object.
[0060] Specifically, in an embodiment of the present application, the obtaining the loops of the actual faces formed by the candidate edges constructed through a deep neural network for the three-dimensional candidate edges constructed based on all the points where the back-projections of the three different views intersect includes:
[0061] For the three-dimensional candidate edges constructed based on all the points where the back-projections of the three different views intersect, obtain the common edges for constructing the three-dimensional space faces, and the common edge is a candidate edge shared by two adjacent three-dimensional space faces;
[0062] Based on the common edges for constructing the three-dimensional space faces and the Pointer Net deep neural network, obtain the loops of the actual faces constructed by the three-dimensional candidate edges.
[0063] Specifically, in an embodiment of the present application, the Pointer Net deep neural network includes:
[0064] Use an encoder to obtain a context embedding vector for each input vector, where the input vector is composed of the common edges;
[0065] At each decoding step, the decoder outputs a pointer vector;
[0066] Compare the pointer vector with the context embedding vector through dot product, and normalize the comparison result scores using the softmax algorithm to obtain an effective probability distribution on the input set;
[0067] Learn the parameters of the model by maximizing the conditional probability of the training set.
[0068] Specifically, in an embodiment of the present application, when using the encoder to obtain a context embedding vector for each input vector, each input vector is composed of common edges, and each common edge embedding vector includes:
[0069] Value embedding, representing the coordinate value of the common edge;
[0070] Position embedding, representing the position of the input vector.
[0071] Specifically, the common edge is divided along the length of the edge into edges with different directions. Each edge just has two common edges pointing in opposite directions. A face can be conveniently represented as one or more loops formed by common edges. In the embodiment of the present application, a loop is a closed path, and each edge is exactly shared by two faces, and one face corresponds to one common edge.
[0072] The order of the common edges of a face is defined such that, when viewed from the direction of the common edge, the face is always on the left side of the common edge. Transforming the recognition of a face into a sequence generation problem, starting from an arbitrary candidate common edge C n1 to generate a sequence of common edges to represent a face f m ={C n1 , …, C nT}, where n is an integer between 1 and N_E. To detect all faces F, we can use each candidate common edge in the candidate common edges C as the starting common edge and repeat this process N_E times.
[0073] Specifically, in the embodiments of the present application, a Pointer Net deep neural network is used to recognize the face model. The Pointer Net deep neural network aims to generate an output sequence, and the sequence composition comes from the input sequence. Specifically, given a sequence of input vectors P = {p 1 , p 2 ,...}, the Pointer Net deep neural network learns the conditional probability: where N=(n 1 , n 2 ,..., n T ) is a sequence composed of T indices, and each index is between 1 and |P|.
[0074] The Pointer Net deep neural network adopts an encoder-decoder architecture. It uses the encoder to obtain the context embedding vector w i for each input vector. At each decoding step t, the decoder outputs a pointer vector u t , and compares it with the input embedding vector through the dot product. The softmax is used to normalize the resulting scores to obtain a valid probability distribution over the input set:
[0075]
[0076] u t = Decoder(N<t, p; θ);
[0077]
[0078] The parameters of the model are learned by maximizing the conditional probability of the training set.
[0079] For the input vectors and embedding vectors. The input vectors consist of all candidate co-edges C. A stop token [EOS] is added to indicate the end of the sequence. Two different embedding vectors are used for each co-edge: one is the value embedding, representing the coordinate values of the co-edge; the other is the position embedding, representing the position of the token in the sequence. Since the lengths of the co-edges are different, a fixed number of points are sampled to represent each co-edge. To avoid sampling errors caused by directly sampling in the grid of 3D voxels, the 3D voxels are projected onto each view and sampled on the 2D lines, and these points are sorted in the direction of the co-edge. The sampled points are flattened and two linear layers are applied to obtain an embedding vector of 512 dimensions.
[0080] For the output vectors and embedding vectors. Since each face is defined as a sequence of candidate co-edges, the output sequence can be represented as f m ={c n1 ,c n2 ...,c nT}. When a face consists of multiple loops, except for the loop to which the starting co-edge belongs, the co-edges in other loops are sorted in ascending order of their indices.
[0081] Since each face is predicted from each candidate co-edge, there are many duplicate face predictions in the results. If two output sequences are composed of the same set of co-edges, they must correspond to the same face and the two can be merged.
[0082] In addition, duplicate predictions can filter out incorrect output sequences. If c′ appears in the true output sequence containing the starting co-edge c, then c should also appear in the true output sequence with the starting co-edge c′. A consistency metric is proposed for the predictions of all faces. Specifically, let f(c) denote the predicted sequence of the face starting with the co-edge c, and the consistency score of the face prediction is defined as f(c n1 )={c n1 ,c n2 ,…,c nT} as follows:
[0083]
[0084] Discard the predictions of faces with a consistency score lower than the threshold δ.
[0085] In step 240, according to the loop of the actual face formed by the candidate edges constructed by the deep neural network, obtain the 3D model obtained by fitting the parametric model for the loop of the actual face.
[0086] Specifically, in the embodiments of the present application, a reconstruction algorithm is implemented based on the BRepBuilderAPIs of Open CASCADE Technology. By identifying candidate edges belonging to the same pair of opposite faces, vertices are recovered from the intersection points of the edges. To restore the geometric information of the object, the points on the edges are fitted into a parametric model. The parametric model is restricted to line segments and circular arcs. Based on the prediction of the faces, the BRepBuilderAPI_MakeWire is used to connect the edges into wires, and a parametric model of the face is obtained therefrom through BRepBuilderAPI_MakeFace.
[0087] The method for reconstructing a three-dimensional model from three views in the present application obtains the reverse projection of the two-dimensional edges of each view in three-dimensional space; based on the reverse projection of the two-dimensional edges of each view in three-dimensional space, it obtains all the points where the reverse projections of the three different views intersect to construct three-dimensional candidate edges; based on the three-dimensional candidate edges constructed from all the points where the reverse projections of the three different views intersect, it obtains the loop of the actual faces formed by the candidate edges constructed through a deep neural network; based on the loop of the actual faces formed by the candidate edges constructed through the deep neural network, it obtains the three-dimensional model obtained by fitting the parametric model of the loop of the actual faces. The present application can solve the problem of three-dimensional model reconstruction in CAD in a simple and efficient manner, avoiding restrictions on the geometric shape type when restoring the topological information of the object, and having a wider application range.
[0088] It should be understood that although Figure 2 the steps in the flowchart of Figure 2 are shown in sequence according to the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless there is a clear indication in this article, the execution of these steps has no strict order restriction, and these steps can be executed in other orders. Moreover,
[0089] Figure 3 is a schematic structural diagram of a device for reconstructing a three-dimensional model from three views provided by an embodiment of the present application. As Figure 3 shown, the device for reconstructing a three-dimensional model from three views includes:
[0090] a candidate line generation module, a face detection module based on a deep neural network, and a model reconstruction module;
[0091] A candidate line generation module, configured to obtain the back-projection of the two-dimensional edges of each view in three-dimensional space; and based on the back-projections of the two-dimensional edges of each view in three-dimensional space, obtain three-dimensional candidate edges constructed by all the points where the back-projections of three different views intersect, where the three different views include a front view, a top view, and a side view.
[0092] A surface detection module based on a deep neural network, configured to obtain a loop of the actual surface formed by the candidate edges constructed by the deep neural network based on the three-dimensional candidate edges constructed by all the points where the back-projections of the three different views intersect, and each loop of the actual surface corresponds to an actual surface of the three-dimensional space object.
[0093] A model reconstruction module, configured to obtain a three-dimensional model obtained by fitting a parametric model based on the loop of the actual surface formed by the candidate edges constructed by the deep neural network.
[0094] Specifically, in another embodiment of the present application, the candidate line generation module is configured to obtain a plurality of three-dimensional voxels obtained by voxelizing the three-dimensional space; based on the plurality of three-dimensional voxels obtained by voxelizing the three-dimensional space, obtain the positions of each three-dimensional voxel projected onto the two-dimensional views of three different views; based on the positions of each three-dimensional voxel projected onto the two-dimensional views of three different views, obtain the matching line information of each three-dimensional voxel in the two-dimensional views of each different view; based on the matching line information of each three-dimensional voxel in the two-dimensional views of each different view, obtain three-dimensional candidate voxels in the three-dimensional voxels; and based on the three-dimensional candidate voxels, obtain three-dimensional candidate edges constructed by the three-dimensional candidate voxels having the same matching line in the two-dimensional views of the three different views.
[0095] Specifically, in another embodiment of the present application, the surface detection module based on a deep neural network is configured to obtain a common edge for constructing a three-dimensional space surface based on the three-dimensional candidate edges constructed by all the points where the back-projections of the three different views intersect, where the common edge is a candidate edge shared by two adjacent three-dimensional space surfaces; and based on the common edge for constructing the three-dimensional space surface and the Pointer Net deep neural network, obtain a loop of the actual surface constructed by the three-dimensional candidate edges.
[0096] The device for reconstructing a three-dimensional model from three orthographic views in this application obtains the back-projections in three-dimensional space of the two-dimensional edges of each view through a candidate line generation module; based on the back-projections in three-dimensional space of the two-dimensional edges of each view, it obtains the three-dimensional candidate edges constructed by all the intersection points of the back-projections of three different views, where the three different views include the front view, the top view, and the side view; a surface detection module based on a deep neural network obtains, according to the three-dimensional candidate edges constructed by all the intersection points of the back-projections of the three different views, the loops of the actual surfaces formed by the candidate edges constructed through the deep neural network, and each loop of the actual surface corresponds to an actual surface of the three-dimensional space object; a model reconstruction module obtains, according to the loops of the actual surfaces formed by the candidate edges constructed through the deep neural network, the three-dimensional model obtained by fitting a parametric model to the loops of the actual surfaces. This application can solve the problem of three-dimensional model reconstruction in CAD in a simple and efficient way, avoiding the limitation on the geometric shape type when restoring the topological information of an object, and has a wider application range.
[0097] For the specific limitations on the device for reconstructing a three-dimensional model from three orthographic views, reference can be made to the limitations on the method for reconstructing a three-dimensional model from three orthographic views in the above text, which will not be elaborated here. Each module in the above device for reconstructing a three-dimensional model from three orthographic views can be implemented in whole or in part by software, hardware, and their combination. The above-mentioned modules can be embedded in the processor of a computer device in hardware form or be independent of it, or can be stored in the memory of the computer device in software form, so that the processor can call and execute the operations corresponding to the above-mentioned modules.
[0098] In particular, according to an embodiment of the present disclosure, as Figure 4 shown, the present invention discloses an electronic device, which includes one or more processors and a memory. The memory is used to store one or more programs; when the one or more programs are executed by the processor, the processor implements the method for reconstructing a three-dimensional model from three orthographic views according to the embodiments of the present invention.
[0099] In particular, according to an embodiment of the present disclosure, the method for reconstructing a three-dimensional model from three orthographic views described in any of the above embodiments can be implemented as a computer software program. For example, an embodiment of the present disclosure includes a computer program product, which includes a computer program tangibly embodied on a machine-readable medium, and the computer program includes program code for executing the method for reconstructing a three-dimensional model from three orthographic views. In such an embodiment, the computer program can be downloaded and installed from a network through a communication part, and / or installed from a removable medium.
[0100] The one or more programs are executed to perform various appropriate actions and processes by a program stored in a read-only memory (ROM) or a program stored in a random access memory (RAM). In the random access memory (RAM), there are software programs for the server to complete corresponding services, as well as various programs and data required for vehicle driving operations. The server, the hardware devices it controls, the read-only memory (ROM), and the random access memory (RAM) are connected to each other via a bus, and various input / output interfaces are also connected to the bus.
[0101] The following components are connected to the input / output interface: an input part including a keyboard, a mouse, etc.; an output part including a cathode ray tube (CRT), a liquid crystal display (LCD), etc. and a speaker, etc.; and a communication part including a network interface card such as a LAN card, a modem, etc. The communication part performs communication processing via a network such as the Internet. A drive is also connected to the input / output interface as needed. Removable media, such as magnetic disks, optical disks, magneto-optical disks, semiconductor memories, etc., are installed on the drive as needed so that computer programs read from them can be installed into the memory as needed.
[0102] In particular, according to an embodiment of the present disclosure, the method for reconstructing a three-dimensional model from three views described in any of the above embodiments can be implemented as a computer software program. For example, an embodiment of the present disclosure includes a computer program product that includes a computer program tangibly embodied on a machine-readable medium, and the computer program includes program code for executing the method for reconstructing a three-dimensional model from three views. In such an embodiment, the computer program can be downloaded and installed from a network through the communication part, and / or installed from a removable medium.
[0103] The units or modules involved in the embodiments described in this application can be implemented in software or in hardware. The described units or modules can also be provided in a processor. The names of these units or modules do not, in some cases, constitute a limitation on the units or modules themselves.
[0104] The above description is only a preferred embodiment of this application and an explanation of the applied technical principles. Those skilled in the art should understand that the scope of the invention involved in this application is not limited to the technical solution formed by the specific combination of the above technical features, and should also cover other technical solutions formed by any combination of the above technical features or their equivalent features without departing from the inventive concept. For example, a technical solution formed by mutually replacing the above features with (but not limited to) technical features having similar functions disclosed in this application.
Claims
1. A method for reconstructing a three-dimensional model from three views, characterized in that, the method includes: Obtaining the back-projection of each two-dimensional edge in each view into three-dimensional space; According to the back-projection of each two-dimensional edge in each view into three-dimensional space, obtaining three-dimensional candidate edges constructed by all points where the back-projections of three different views intersect, and the three different views include a front view, a top view, and a side view; According to the three-dimensional candidate edges constructed by all points where the back-projections of the three different views intersect, obtaining a loop of the actual surface formed by constructing candidate edges through a deep neural network, and each loop of the actual surface corresponds to the actual surface of the three-dimensional space object; According to the loop of the actual surface formed by constructing candidate edges through the deep neural network, obtaining a three-dimensional model obtained by fitting a parametric model for the loop of the actual surface; The obtaining of three-dimensional candidate edges constructed by all points where the back-projections of three different views intersect according to the back-projection of each two-dimensional edge in each view into three-dimensional space includes: Obtaining a plurality of three-dimensional voxels for voxelizing three-dimensional space; According to the plurality of three-dimensional voxels for voxelizing three-dimensional space, obtaining the positions of each three-dimensional voxel projected onto the two-dimensional views of three different views; According to the positions of each three-dimensional voxel projected onto the two-dimensional views of three different views, obtaining the matching line information of each three-dimensional voxel in the two-dimensional views of each different view; According to the matching line information of each three-dimensional voxel in the two-dimensional views of each different view, obtaining three-dimensional candidate voxels in the three-dimensional voxels; According to the three-dimensional candidate voxels, obtaining three-dimensional candidate voxels having the same matching line in the two-dimensional views of the three different views and constructing them into three-dimensional candidate edges.
2. The method according to claim 1, characterized in that, the three-dimensional candidate edges are associated with the matching lines of the two-dimensional views of the three different views.
3. The method according to claim 1, characterized in that, The obtaining of a loop of the actual surface formed by candidate edges constructed through a deep neural network according to the three-dimensional candidate edges constructed by all points where the back-projections of the three different views intersect includes: According to the three-dimensional candidate edges constructed by all points where the back-projections of the three different views intersect, obtaining the common edges for constructing the three-dimensional space surface, and the common edge is a candidate edge shared by two adjacent surfaces; According to the common edges for constructing the three-dimensional space surface and the Pointer Net deep neural network, obtaining a loop of the actual surface constructed by the three-dimensional candidate edges.
4. The method according to claim 3, characterized in that, the Pointer Net deep neural network includes: Using an encoder to obtain a context embedding vector for each input vector, and the input vector is composed of the common edges; In each decoding step, the decoder outputs a pointer vector; Comparing the pointer vector with the context embedding vector through dot product, and normalizing the comparison result scores using the softmax algorithm to obtain an effective probability distribution on the input set; Learning the parameters of the model by maximizing the conditional probability of the training set.
5. The method according to claim 4, wherein, when using an encoder to obtain context embedding vectors for each input vector, each input vector is composed of common edges, and each common edge embedding vector includes: a value embedding, representing the coordinate value of the common edge; a position embedding, representing the position of the input vector.
6. An apparatus for reconstructing a three-dimensional model from three views, wherein, the apparatus includes: a candidate line generation module, configured to obtain the reverse projection of the two-dimensional edges of each view in three-dimensional space; and based on the reverse projection of the two-dimensional edges of each view in three-dimensional space, obtain three-dimensional candidate edges constructed by all points where the reverse projections of three different views intersect, the three different views including a front view, a top view, and a side view; the obtaining three-dimensional candidate edges constructed by all points where the reverse projections of three different views intersect based on the reverse projection of the two-dimensional edges of each view in three-dimensional space includes: obtaining a plurality of three-dimensional voxels obtained by voxelizing three-dimensional space; based on the plurality of three-dimensional voxels obtained by voxelizing three-dimensional space, obtaining the positions of each of the three-dimensional voxels projected onto the two-dimensional views of three different views; based on the positions of each of the three-dimensional voxels projected onto the two-dimensional views of three different views, obtaining the matching line information of each of the three-dimensional voxels in each of the two-dimensional views of the different views; based on the matching line information of each of the three-dimensional voxels in each of the two-dimensional views of the different views, obtaining three-dimensional candidate voxels in the three-dimensional voxels; and based on the three-dimensional candidate voxels, obtaining three-dimensional candidate voxels having the same matching line in the two-dimensional views of the three different views and constructing them into three-dimensional candidate edges; a face detection module based on a deep neural network, configured to obtain, based on the three-dimensional candidate edges constructed by all points where the reverse projections of the three different views intersect, a loop of an actual face formed by the candidate edges constructed by the deep neural network, and each loop of the actual face corresponds to an actual face of a three-dimensional space object; a model reconstruction module, configured to obtain a three-dimensional model obtained by fitting a parameterized model based on the loop of the actual face formed by the candidate edges constructed by the deep neural network.
7. An electronic device, wherein, the device includes one or more processors and a memory, and the memory is used to store one or more programs; when the one or more programs are executed by the processor, the processor is caused to implement the method according to any one of claims 1 to 5.
8. A computer-readable storage medium storing a computer program, wherein, when the computer program is executed, it implements the method according to any one of claims 1 to 5.
Citation Information
Patent Citations
Multi-view-based 3D reconstruction method
CN104851129A
Three-dimensional shape restoration method, its device, program, and recording medium
JP2009294956A