Three-dimensional prediction model training method, three-dimensional reconstruction method and device

By training the target 3D prediction model, using the 2D view of the sample object to generate a predicted coordinate sequence and combining it with the labeled coordinate sequence for training, the problem of high dependence on 2D views in the existing technology is solved, and the accuracy and robustness of 3D model reconstruction are improved.

CN116721217BActive Publication Date: 2025-09-09HANGZHOU QUNHE INFORMATION TECHNOLOGIES CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310767466.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-06-26
Publication Date
2025-09-09
Estimated Expiration
2043-06-26

AI Technical Summary

Technical Problem

Existing technologies are highly dependent on two-dimensional views in reconstructing three-dimensional objects, resulting in poor reconstruction accuracy and robustness. They are particularly sensitive to drawing errors, which easily lead to reconstruction failure.

Method used

By training the target three-dimensional prediction model, the two-dimensional sample views of the sample object are used to generate a prediction coordinate sequence, and the preset model is trained in combination with the labeled coordinate sequence to generate a target coordinate sequence for reconstructing the three-dimensional model, reducing the dependence on the two-dimensional views and improving the robustness and accuracy of the model.

Benefits of technology

It achieves 3D model reconstruction that is insensitive to drawing errors, improves reconstruction accuracy and robustness, and reduces dependence on 2D views.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116721217B_ABST
    Figure CN116721217B_ABST
Patent Text Reader

Abstract

The present disclosure provides a training method, a 3D prediction model, a 3D reconstruction method, and a device. The method comprises inputting a first sample sequence of a sample object into a preset 3D prediction model to obtain a predicted coordinate sequence, wherein the first sample sequence is obtained based on a two-dimensional sample view of the sample object. The preset 3D prediction model is trained based on the labeled coordinate sequence and the predicted coordinate sequence of the sample object to obtain a target 3D prediction model. The target 3D prediction model is used to generate a target coordinate sequence of the object to be modeled based on the two-dimensional view of the object to be modeled. The target coordinate sequence is used to generate a 3D model of the object to be modeled. The training method, the 3D reconstruction method, and the device provided by the present disclosure can generate a 3D model of the object to be modeled through the target 3D prediction model. The method has low dependence on two-dimensional views and is insensitive to errors in drawings, thereby improving the accuracy and robustness of 3D model reconstruction.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of three-dimensional reconstruction technology, and in particular to a three-dimensional prediction model training method, a three-dimensional reconstruction method and a device. Background Art

[0002] Reconstructing a 3D object from three views is the process of reconstructing a 3D model of an object based on its two-dimensional views. It is a long-standing research topic in computer-aided design. Improving the accuracy of the reconstructed 3D model is a key issue in 3D object reconstruction. On the other hand, robust reconstruction of 3D models from 2D drawings (such as drawing errors, missing line segments, and other noise issues) has not received much attention in the industry. Summary of the Invention

[0003] The present disclosure provides a three-dimensional prediction model training method, a three-dimensional reconstruction method and a device.

[0004] According to one aspect of the present disclosure, a method for training a three-dimensional prediction model is provided, comprising: inputting a first sample sequence of a sample object into a preset three-dimensional prediction model to obtain a prediction coordinate sequence, wherein the first sample sequence is obtained based on a two-dimensional sample view of the sample object; training the preset three-dimensional prediction model based on the labeled coordinate sequence and the prediction coordinate sequence of the sample object to obtain a target three-dimensional prediction model, wherein the target three-dimensional prediction model is used to generate a target coordinate sequence of the object to be modeled based on the two-dimensional view of the object to be modeled, and the target coordinate sequence is used to generate a three-dimensional model of the object to be modeled.

[0005] According to another aspect of the present disclosure, a three-dimensional reconstruction method is provided, comprising: generating a first sequence based on a two-dimensional view of an object to be modeled; inputting the first sequence into a target three-dimensional prediction model to obtain a target coordinate sequence, wherein the target three-dimensional prediction model is trained according to the method of any of the above embodiments; and generating a three-dimensional model of the object to be modeled based on the target coordinate sequence.

[0006] According to another aspect of the present disclosure, a training device for a three-dimensional prediction model is provided, comprising: an input unit for inputting a first sample sequence of a sample object into a preset three-dimensional prediction model to obtain a prediction coordinate sequence, wherein the first sample sequence is obtained based on a two-dimensional sample view of the sample object; a training unit for training the preset three-dimensional prediction model based on the labeled coordinate sequence and the prediction coordinate sequence of the sample object to obtain a target three-dimensional prediction model, wherein the target three-dimensional prediction model is used to generate a target coordinate sequence of the object to be modeled based on the two-dimensional view of the object to be modeled, and the target coordinate sequence is used to generate a three-dimensional model of the object to be modeled.

[0007] According to another aspect of the present disclosure, a three-dimensional reconstruction apparatus is provided, comprising: a first generation unit for generating a first sequence based on a two-dimensional view of an object to be modeled; a prediction unit for inputting the first sequence into a target three-dimensional prediction model to obtain a target coordinate sequence, wherein the target three-dimensional prediction model is trained according to the method of any of the above-mentioned embodiments; and a second generation unit for generating a three-dimensional model of the object to be modeled based on the target coordinate sequence.

[0008] According to another aspect of the present disclosure, an electronic device is provided, comprising: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute any method in the embodiments of the present disclosure.

[0009] According to another aspect of the present disclosure, a non-transitory computer-readable storage medium storing computer instructions is provided, wherein the computer instructions are used to enable the computer to execute any method according to the embodiments of the present disclosure.

[0010] According to another aspect of the present disclosure, a computer program product is provided, including a computer program. When the computer program is executed by a processor, the computer program implements any one of the methods according to the embodiments of the present disclosure.

[0011] The training method, 3D reconstruction method, and apparatus of the 3D prediction model provided by the embodiments of the present disclosure obtain a predicted coordinate sequence by inputting a first sample sequence of a sample object into a preset 3D prediction model, wherein the first sample sequence is obtained based on a 2D sample view of the sample object; the preset 3D prediction model is trained based on the labeled coordinate sequence and the predicted coordinate sequence of the sample object to obtain a target 3D prediction model, the target 3D prediction model is used to generate a target coordinate sequence of the object to be modeled based on the 2D view of the object to be modeled, and the target coordinate sequence is used to generate a 3D model of the object to be modeled. The training method, 3D reconstruction method, and apparatus of the 3D prediction model provided by the embodiments of the present disclosure enable the target 3D prediction model to generate a 3D model of the object to be modeled. Compared with the method of modeling based on the correspondence relationship between the 2D view and the 3D model in the related art, the 3D reconstruction performed using the target 3D prediction model has low dependence on the 2D view, is insensitive to errors in the drawing, and improves the accuracy and robustness of the 3D model reconstruction.

[0012] It should be understood that the contents described in this section are not intended to identify the key or important features of the embodiments of the present disclosure, nor are they intended to limit the scope of the present disclosure. Other features of the present disclosure will become readily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS

[0013] The accompanying drawings are used to better understand the present invention and do not constitute a limitation of the present invention.

[0014] Figure 1 3D prediction model training method and 3D reconstruction method of the embodiment of the present disclosure are applied;

[0015] Figure 2 is a flowchart of a training method for a three-dimensional prediction model provided in accordance with an embodiment of the present disclosure;

[0016] Figure 3 is a schematic structural diagram of a three-dimensional prediction model provided according to an embodiment of the present disclosure;

[0017] Figure 4 is a flowchart of a three-dimensional reconstruction method provided according to an embodiment of the present disclosure;

[0018] Figure 5 is a process diagram of generating a three-dimensional model according to the three-dimensional reconstruction method provided by an embodiment of the present disclosure;

[0019] Figure 6A is a schematic diagram of global editing of a 3D model according to a 3D reconstruction method provided by an embodiment of the present disclosure;

[0020] Figure 6B is a schematic diagram of local editing of a 3D model according to a 3D reconstruction method provided by an embodiment of the present disclosure;

[0021] Figure 7 3D prediction model training device according to an embodiment of the present disclosure;

[0022] Figure 8 is a schematic structural diagram of a three-dimensional reconstruction device provided according to an embodiment of the present disclosure;

[0023] Figure 9 It is a block diagram of an electronic device used to implement the three-dimensional prediction model training method and the three-dimensional reconstruction method of the embodiment of the present disclosure. DETAILED DESCRIPTION

[0024] The following description of exemplary embodiments of the present disclosure is made in conjunction with the accompanying drawings, including various details of the embodiments of the present disclosure to facilitate understanding, which should be considered as merely exemplary. Therefore, it should be appreciated by those skilled in the art that various changes and modifications may be made to the embodiments described herein without departing from the scope of the present disclosure. Similarly, for the sake of clarity and conciseness, descriptions of well-known functions and structures are omitted in the following description.

[0025] The embodiments of the present disclosure provide a training method for a three-dimensional prediction model, a three-dimensional reconstruction method and an apparatus method, an apparatus, an electronic device and a storage medium. Specifically, the training method for a three-dimensional prediction model, the three-dimensional reconstruction method and the apparatus of the embodiments of the present disclosure can be executed by an electronic device, wherein the electronic device can be a terminal or a server and other devices. The terminal can be a smart phone, a tablet computer, a laptop computer, an intelligent voice interaction device, a smart home appliance, a wearable smart device, an aircraft, an intelligent vehicle-mounted terminal and other devices. The terminal can also include a client, which can be an audio client, a video client, a browser client, an instant messaging client or a mini-program, etc. The server can be an independent physical server, or a server cluster or distributed system composed of multiple physical servers, or a cloud server that provides cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, content delivery networks (CDNs), and basic cloud computing services such as big data and artificial intelligence platforms.

[0026] Reconstructing a 3D object from 2D three-view images is a long-standing research topic in computer-aided design. 3D reconstruction methods in related technologies reconstruct 3D models by explicitly modeling the correspondence between the 3D model and the 2D views. The reconstruction process is as follows: (1) Generate 3D vertices from 2D vertices; (2) Generate 3D line segments from 3D vertices; (3) Generate 3D faces from 3D line segments; (4) Generate 3D volumes from 3D faces; and (5) Construct a 3D model from the 3D volumes.

[0027] However, this method is heavily dependent on the accuracy of the drawings and is very sensitive to errors in the drawings. Once there are errors or missing elements in the drawings, reconstruction will fail, and the reconstruction accuracy and robustness will be poor.

[0028] In order to solve at least one of the above problems, the embodiments of the present disclosure provide a training method, a three-dimensional reconstruction method and an apparatus for a three-dimensional prediction model, wherein a first sample sequence of a sample object is input into a preset three-dimensional prediction model to obtain a predicted coordinate sequence, wherein the first sample sequence is obtained based on a two-dimensional sample view of the sample object; the preset three-dimensional prediction model is trained according to the labeled coordinate sequence and the predicted coordinate sequence of the sample object to obtain a target three-dimensional prediction model, and the target three-dimensional prediction model is used to generate a target coordinate sequence of the object to be modeled based on the two-dimensional view of the object to be modeled, and the target coordinate sequence is used to generate a three-dimensional model of the object to be modeled. The training method, the three-dimensional reconstruction method and the apparatus for a three-dimensional prediction model provided by the embodiments of the present disclosure enable the target three-dimensional prediction model to generate a three-dimensional model of the object to be modeled. Compared with the method of modeling based on the correspondence relationship between the two-dimensional view and the three-dimensional model in the related art, the three-dimensional reconstruction performed using the target three-dimensional prediction model has low dependence on the two-dimensional view, is insensitive to errors in the drawing, and improves the accuracy and robustness of the three-dimensional model reconstruction.

[0029] The embodiments of the present disclosure will be described in detail below with reference to the accompanying drawings.

[0030] Figure 1 This is a schematic diagram of the structure of a system for applying the training method of the 3D prediction model and the 3D reconstruction method of the embodiment of the present disclosure. Figure 1 The system includes a terminal 110 and a server 120, etc.; the terminal 110 and the server 120 are connected via a network, for example, via a wired or wireless network connection.

[0031] In which, the server 120 can be used to input the first sample sequence of the sample object into a preset three-dimensional prediction model to obtain a predicted coordinate sequence, wherein the first sample sequence is obtained based on the two-dimensional sample view of the sample object; according to the labeled coordinate sequence and the predicted coordinate sequence of the sample object, the preset three-dimensional prediction model is trained to obtain a target three-dimensional prediction model.

[0032] The terminal 110 can be used to display a graphical user interface. The terminal is used to interact with the user through the graphical user interface, for example, by downloading and installing the corresponding client and running it through the terminal, for example, by calling and running the corresponding applet, for example, by logging into a website to present the corresponding graphical user interface, etc. In an embodiment of the present disclosure, the terminal 110 may have a two-dimensional view input interface to obtain a two-dimensional view of the object to be modeled. The server 120 can generate a first sequence based on the two-dimensional view of the object to be modeled; input the first sequence into the target three-dimensional prediction model to obtain a target coordinate sequence, and generate a three-dimensional model of the object to be modeled based on the target coordinate sequence. The terminal 120 can also be used to display the three-dimensional model.

[0033] In this embodiment, the 3D model is displayed and the 2D view is input via terminal 110, while the 3D reconstruction method is executed by the server. Of course, in other embodiments, the 3D model can be displayed and the 2D view can be input via the server, or the 3D reconstruction method can be executed on the terminal. Furthermore, the model training method and the 3D reconstruction method can be executed by different servers.

[0034] It should be noted that the above application scenarios are only shown to facilitate understanding of the spirit and principles of the present disclosure, and the embodiments of the present disclosure are not limited in this respect. On the contrary, the embodiments of the present disclosure can be applied to any applicable scenario.

[0035] It should be noted that the order in which the following embodiments are described does not limit the priority order of the embodiments.

[0036] Figure 2 This is a flow chart of a method for training a three-dimensional prediction model according to an embodiment of the present disclosure; please refer to Figure 2 The embodiment of the present disclosure provides a three-dimensional prediction model training method 200, including the following steps S201 and S202.

[0037] Step S201 : Input a first sample sequence of a sample object into a preset three-dimensional prediction model to obtain a prediction coordinate sequence, wherein the first sample sequence is obtained based on a two-dimensional sample view of the sample object.

[0038] In step S202, a preset three-dimensional prediction model is trained based on the labeled coordinate sequence and the predicted coordinate sequence of the sample object to obtain a target three-dimensional prediction model. The target three-dimensional prediction model is used to generate a target coordinate sequence of the object to be modeled based on the two-dimensional view of the object to be modeled. The target coordinate sequence is used to generate a three-dimensional model of the object to be modeled.

[0039] The preset 3D prediction model may be a model built based on a neural network, which refers to a 3D prediction model before or during training. The target 3D prediction model is a trained 3D prediction model obtained by training the preset 3D prediction model.

[0040] The sample object can be a three-dimensional object used as a sample, for example, a cabinet composed of multiple panels. 2D sample view of the sample object Includes three 2D views of the sample object, i.e., orthogonal projection views They are front view, top view and side view respectively.

[0041] The first sample sequence is obtained according to the two-dimensional sample view of the sample object and includes the coordinates of all line segments in the two-dimensional sample view.

[0042] The annotated coordinate sequence includes the annotated coordinates (i.e., real coordinates) of each face in the three-dimensional object. Taking the cabinet structure as an example, the cabinet structure includes a plurality of planar panels (i.e., components of the sample object). It can be understood that the planar panels are usually placed along the coordinate axis, that is, the planar panels are parallel to one of the three mutually perpendicular coordinate axes of the cabinet. Therefore, the planar panels can be represented by rectangular blocks, each of which has 6 degrees of freedom (DOF), corresponding to the starting and ending coordinates of the three coordinate axes. For example, the coordinates of the two diagonally opposite corners of the rectangular block are (x min ,y min ,z min ) and (x max ,y max ,z max ), then the cuboid can be represented as Cuboid(x min ,y min ,z min ,x max ,y max ,z max ). It can be understood that a cuboid has 6 faces, and the coordinates of these six faces are x min 、y min 、z min 、x max 、y max 、z max .

[0043] Expanding the coordinates of all components (all cuboids) in the sample object into a one-dimensional sequence generates a labeled coordinate sequence. In the cabinet scenario, every six labeled coordinates in the labeled coordinate sequence can be used to represent a flat panel.

[0044] It can be understood that if the sample object has A components, each component includes B faces, and the sample object has a total of M faces, then M can be determined based on A and B. When the sample object is a cabinet, the cabinet includes A planar panels, each planar panel includes 6 faces (B=6), then the number of faces of the sample object, M, can be A multiplied by B. Alternatively, M can also be equal to A+1 multiplied by B. In this case, the bounding box of the entire cabinet can also be regarded as a large rectangular parallelepiped. In addition to the faces of the A components, the M faces can also include the coordinates of the 6 faces of the bounding box (not necessarily the actual faces of the cabinet). The annotated coordinate sequence can also include M annotated coordinates, and these M annotated coordinates correspond one-to-one to the M faces, that is, each annotated coordinate is the coordinate of a face.

[0045] In addition, the coordinates can be coordinate values ​​or face identifiers determined by the attachment relationship, that is, the coordinate values ​​in the cuboid can be defined by specifying numerical values ​​or attachment operations.min ,y min ,z min ,x max ,y max ,z max ) can be either a coordinate value or a pointer that points to a coordinate of the attached cuboid (the identifier of the attached face).

[0046] The attachment operation can also be referred to as an attachment relationship. It is understood that in the cabinet scenario, an attachment relationship can refer to the contact and assembly of two surfaces. Specifically, taking two planar panels as an example, each of planar panel 1 and planar panel 2 can include two end surfaces (two surfaces at both ends in the thickness direction) and four side surfaces (four surfaces surrounding the thickness direction).

[0047] In some embodiments, a side surface of the planar panel 1 is assembled on an end surface of the planar panel 2 , which means that there is an attachment relationship between the two surfaces.

[0048] Conversely, if one end face of planar panel 1 is assembled to one side face of planar panel 2, it can also be considered that the two have a dependent relationship. However, if one end face of planar panel 1 is assembled to one end face of planar panel 2, or one side face of planar panel 1 is assembled to one side face of planar panel 2, it is not a dependent relationship.

[0049] For example, if surface 1 (labeled as C1) of planar plate 1 is attached to surface 2 of planar plate 2, then in the labeled coordinate sequence, the labeled coordinates of surface 2 can be represented by the label C1 of surface 1. Of course, the coordinate values ​​of surface 2 can also be used.

[0050] It is understood that the above embodiments of the dependency relationship are not intended to limit the present disclosure. In other embodiments, the dependency relationship may have other assembly forms and still be applicable to the methods of the embodiments of the present disclosure.

[0051] The predicted coordinate sequence is the output of the model obtained by taking the first sample sequence as the input of the preset three-dimensional prediction model. The number of predicted coordinates included in the predicted coordinate sequence is consistent with the number and arrangement order of the annotated coordinates in the annotated coordinate sequence.

[0052] The loss function value of the preset three-dimensional prediction model can be calculated by combining the predicted coordinate sequence and the labeled coordinate sequence. If the preset convergence condition is not met, the parameters of the preset three-dimensional prediction model are adjusted until the loss function value meets the preset convergence condition, thereby obtaining the target three-dimensional prediction model.

[0053] When in use, the user can generate a first sequence of two-dimensional views of the object to be modeled, and then input the first sequence into the target three-dimensional prediction model to obtain the target coordinate sequence of the object to be modeled, and then generate a three-dimensional model of the object to be modeled according to the target coordinate sequence.

[0054] In this embodiment, a target 3D prediction model can be obtained by training a preset 3D prediction model. This target 3D prediction model can be used to generate a 3D model of the object to be modeled. This allows the process of obtaining a 3D model from a 2D view to be converted into the input and output of the 3D prediction model. This establishes a correspondence between the 2D view and the 3D model, effectively utilizing the 3D prediction model to convert the 2D view into a 3D model.

[0055] Furthermore, the related art modeling approach relies on the correspondence between 2D views and 3D models. If the 2D drawings contain incorrect coordinates, the generation of 3D vertices from 2D vertices will inevitably result in errors, and the 3D model construction will inevitably fail. Because the 3D model reconstruction of this embodiment utilizes a 3D prediction model and does not rely on vertex correspondence, it has a low dependency on the 2D views, is insensitive to errors in the drawings, and has a high tolerance for errors in the 2D views, thereby improving the accuracy and robustness of 3D model reconstruction.

[0056] The construction of the first sample sequence is explained below. In some embodiments, method 200 may further include: acquiring multiple line segments contained in the two-dimensional sample view; converting the coordinates of the multiple line segments into a line segment coordinate sequence; and converting the line segment coordinate sequence into the first sample sequence.

[0057] It is understood that all line segments in the two-dimensional sample view can be extracted. Taking the two-dimensional sample view in CAD format as an example, each view (front view, top view, or side view) can include multiple line segments, and the line segments can include visible line segments represented by solid lines and invisible line segments represented by dashed lines. Of course, in some cases, all line segments in the view can be visible line segments.

[0058] Acquiring multiple line segments contained in a two-dimensional sample view may involve acquiring all line segments in a front view, a top view, and a side view.

[0059] After obtaining all the line segments, the line segments can be converted into a line segment coordinate sequence, which can include the coordinates of all line segments. It can be understood that the line segments in the two-dimensional view are two-dimensional line segments. A two-dimensional line segment has four degrees of freedom, corresponding to the coordinate values ​​of the two endpoints of the line segment. For example, a line segment can include the starting point coordinates (x1, y1) and the end point coordinates (x2, y2). The coordinates of the line segment can be flattened to form a sequence of four line segment coordinates (coordinate values) (x1, y1, x2, y2).

[0060] Flatten the coordinates of all line segments into a one-dimensional sequence That is, the line segment coordinate sequence. Each line segment coordinate v in represents a coordinate value of the line segment (such as x1, y1, x2 or y2), and The length N v It is 4 times the number of all line segments in the three views. For example, Every four line segment coordinates in can be used to represent a line segment, for example, v1 to v4 represent a line segment, v5 to v8 represent a line segment, and so on.

[0061] After obtaining the line segment coordinate sequence, the line segment coordinate sequence can be processed, for example, using a word embedding method, to obtain a first sample sequence.

[0062] In the above manner, the two-dimensional view can be converted into a first sample sequence that can be processed by the three-dimensional prediction model.

[0063] Furthermore, in some embodiments, converting the coordinates of the plurality of line segments into a line segment coordinate sequence may include: sorting the plurality of line segments according to a first preset rule; and converting the coordinates of the sorted plurality of line segments into a line segment coordinate sequence.

[0064] It is understood that before generating a line segment coordinate sequence, the line segments can be sorted according to a first preset rule. For example, they can be sorted by view first. The view sorting can rely on a custom view sorting order. For example, the line segments can be classified according to the view (front view, top view, and side view) in which they are located, and the line segments of the front view are arranged before the line segments of the top view, and the line segments of the side view are arranged at the end. Then, the line segments in each view category can be sorted according to the coordinate values. For example, the line segments in each view category can be sorted in the order of the coordinate values ​​of the horizontal coordinates of the starting points of the line segments.

[0065] In this embodiment, the first preset rule may also be other rules, such as directly sorting according to the size of the coordinate values ​​of the line segments.

[0066] Then, the four coordinate values ​​of the start points and end points of the sorted multiple line segments may be expanded to obtain a one-dimensional line segment coordinate sequence.

[0067] By sorting the line segments, the order of the line segment coordinates in the line segment coordinate sequence can be defined, so that the ordered line segment coordinate sequence can be used to train the preset three-dimensional prediction model, and the correspondence between the line segment coordinate sequence and the input two-dimensional sample view can be established, which can improve the accuracy of the prediction results.

[0068] In some embodiments, the line segment coordinate sequence includes N line segment coordinates, the first sample sequence includes N first features, the rth first feature among the N first features is obtained based on the rth line segment coordinate among the N line segment coordinates, and N is a positive integer; the rth first feature includes at least one of the following: a numerical feature, a view feature, a view position feature, a line segment position feature, and a visual feature.

[0069] Among them, the numerical feature is used to represent the coordinate value of the rth line segment coordinate; the view feature is used to represent the two-dimensional sample view where the rth line segment coordinate is located; the line segment position feature is used to represent the relative position of the first line segment where the rth line segment coordinate is located in the two-dimensional sample view, and the first line segment is one of multiple line segments; the coordinate position feature is used to represent the relative position of the rth line segment coordinate in the first line segment; and the visual feature is used to represent the visual state of the first line segment.

[0070] It can be understood that N in the line segment coordinate sequence is The length N v Each first feature in the first sample sequence is obtained based on a line segment coordinate in the line segment coordinate sequence. For example, using the rth line segment coordinate v in the first sample sequence r The rth first feature E(v r ). The value of r is a positive integer greater than or equal to 1 and less than or equal to N.

[0071] The line segment coordinate sequence Each element v in r Get its feature embedding through word embedding:

[0072] E(v r )=E value (v r )+E view (v r )+E edge (v r )+E coord (v r )+E type (v r )

[0073] Among them, E(v r ) represents the rth first feature, E value (v r ) represents the numerical feature of the rth first feature, that is, the coordinate value used to represent the coordinate of the rth line segment, that is, E value Represents the discretized coordinate value.

[0074] E view (v r) represents the view feature of the r-th first feature, that is, it is used to represent the two-dimensional sample view where the r-th line segment coordinate is located, that is, E view Indicates which view the line segment comes from (front, top, or side).

[0075] E edge (v r ) represents the line segment position feature of the rth first feature, that is, it is used to represent the first line segment where the rth line segment coordinate is located (v in multiple line segments) r The relative position of the line segment where E edge Indicates the relative position of the line segment in the corresponding view.

[0076] E coord (v r ) represents the coordinate position feature of the rth first feature, that is, it is used to represent the relative position of the rth line segment coordinate in the first line segment, that is, E coord Indicates the relative position of the coordinate on the line segment.

[0077] E type (v r ) represents the visual feature of the rth first feature, which is used to represent the visual state of the first line segment, that is, E type Indicates whether the line segment is visible.

[0078] It can be understood that the first feature may include at least one of a numerical feature, a view feature, a view position feature, a line segment position feature, and a visual feature.

[0079] In this embodiment, a first sample sequence can be obtained based on a line segment coordinate sequence. By setting the first feature in the first sample sequence to include at least one of a numerical feature, a view feature, a view position feature, a line segment position feature, and a visual feature, the first sample sequence can be made to correspond to each feature in a two-dimensional sample view, thereby converting the two-dimensional view into a sequence that can be recognized by a three-dimensional prediction model, which helps to solve the three-dimensional reconstruction problem using the three-dimensional prediction model.

[0080] The above is an explanation of the method for generating the first sample sequence. The following will explain the generation of the annotation coordinate sequence.

[0081] In some embodiments, method 200 may further include the following steps 1 to 3.

[0082] Step 1: Determine a directed acyclic graph of the sample object. The directed acyclic graph includes a dependency relationship set and a face set. The dependency relationship in the dependency relationship set is used to represent two faces with a dependency relationship among the M faces of the sample object. The face set includes the coordinates of the M faces.

[0083] Step 2: sort the directed acyclic graph according to the second preset rule to obtain an ordered graph of sample objects.

[0084] Step three: convert the ordered graph into a labeled coordinate sequence, where the labeled coordinate sequence includes M labeled coordinates arranged according to a preset rule, and the labeled coordinates in the M labeled coordinates are the coordinates of the faces in the M faces.

[0085] It can be understood that when generating the labeled coordinate sequence, the shape program describing the sample object can be converted into a directed acyclic graph. As mentioned above, each plane plate contains 6 faces, and each face corresponds to a coordinate value of Cuboid. The shape program of the sample object can be expressed as a directed acyclic graph A directed acyclic graph includes a set of dependency relationships and a set of faces. The face set corresponds to the graph The vertex set and dependency relationship set of the graph The edge set of .

[0086] It is understandable that the face set (picture The vertex set of the cabinet body contains the faces of all the flat panels. That is, all M faces of the cabinet. Taking f1 as an example, f1 represents the coordinates of the first face in the face set.

[0087] The dependency relationship set ε(Fig. The edge set) represents all dependency relations ε={e1,…,e |ε|}. Among them, each edge e i→j =(f i ,f j ) is a directed edge, starting from f i , the end point is f j , represents the i-th face f in the face set i The jth face f attached to the face set j The dependency relationship set ε can also be expressed as an adjacency matrix Specifically, if f i Attached to f j , then A ij =1; otherwise, A ij =0.

[0088] Then, the directed acyclic graph can be sorted according to the second preset rule to obtain an ordered graph. The order of the vertices (faces) in π.

[0089] For example, first Perform topological sorting, which makes the graph The starting point of any directed edge (dependency relationship) is placed before the end point. Then, for the remaining vertices (faces) without connection relationships, for example, several parallel plane plates, if there is no order, they can be sorted by coordinate values. After the above steps, an ordered graph can be obtained. Its apex The order is based on π. Of course, in addition to the above-mentioned sorting method, the second preset rule can also have other methods, such as directly sorting according to the size of the coordinate value.

[0090] In step 3, similar to the flattening of the first sample sequence, the ordered graph can be Flattened into a one-dimensional sequence of labeled coordinates

[0091] It is understood that in some embodiments, when the first face corresponding to the marked coordinates is located in the dependency set, the marked coordinates are the identifier of the second face that has a dependency relationship with the first face. Alternatively, when the first face corresponding to the marked coordinates is not located in the dependency set, the marked coordinates are the coordinate values ​​of the first face. The first face and the second face are different faces among the M faces. In the following, the first face is taken as f i , the second side is f j Let’s take this as an example to explain.

[0092] Specifically, the coordinate sequence is marked The i-th annotation coordinate (element) g in i yes:

[0093]

[0094] Among them, g i is the i-th labeled coordinate, which means the i-th face (f i )’s labeled coordinates, f i π represents the i-th face (f i ), represents the i-th face (f i ) and the j-th face (f j ) have a dependent relationship, Represents the dependency relationship in the dependency relationship set, which is represented by the adjacency matrix. represents the i-th face (f i ) and the j-th face (f j ) has a dependency relationship, that is, the dependency relationship set has the i-th face (f i ) and the j-th face (f j ) between them, represents the i-th face (f i) has no dependency relationship with the rest of the faces, that is, the i-th face is not in the dependency relationship set. i ) and the j-th face (f j ) are all faces in the face set. Wherein, the values ​​of i and j are both positive integers from 1 to M, and i is not equal to j.

[0095] It can be understood that the dependency relationship set may only record two faces that have a dependency relationship. If a face does not have a dependency relationship with other faces, it may not be recorded in the dependency relationship set.

[0096] When determining whether the labeled coordinates are coordinate values ​​or face identifiers, the face identifier is preferred. For example, if the first face corresponding to the labeled coordinates is in the dependency relationship set, that is, there is a second face that has a dependency relationship with the first face, then the identifier of the second face is used as g. i If the first face is not in the dependency set, that is, there is no face with a dependency relationship with the first face among the remaining faces, then the coordinate value of the first face is taken as g i The value of .

[0097] By using the identifier of the face to which the dependency relationship points as the value of the labeled coordinate when there is a dependency relationship, the output of the three-dimensional prediction model can prioritize the relationship between faces rather than the coordinate value, reducing the dependence on the numerical value and improving the error tolerance of the prediction of the target coordinate sequence, thereby improving the accuracy and robustness of the three-dimensional reconstruction.

[0098] Finally, after obtaining the labeled coordinate sequence, two special tokens, [SOS] and [EOS], can be added to represent the start and end symbols of the output sequence, respectively. [SOS] can be placed before the labeled coordinate sequence, and [EOS] can be placed after the labeled coordinate sequence. Since the length of the labeled coordinate sequence varies for different sample objects, the start and end symbols can be used to represent the complete labeled coordinate sequence.

[0099] It can be understood that the directed acyclic graph represents the coordinates of the faces and the relationship between the faces, but it does not specify the arrangement order of the faces. We can customize the arrangement order of the faces through the second preset rule.

[0100] Since there's a one-to-one correspondence between annotated coordinates and faces, defining the face ordering determines the order of the annotated coordinates, and thus the order of the predicted coordinates in the predicted coordinate sequence output by the 3D prediction model. In other words, by defining the order of the annotated coordinate sequence, the order of the target coordinate sequence output by the trained 3D prediction model can be made consistent with the ordering of the annotated coordinate sequence.

[0101] In this embodiment, the labeled coordinates in the labeled coordinate sequence can be matched with the faces of the three-dimensional model in the sample object, that is, the labeled coordinates can be matched with the three-dimensional model, so that the target coordinate sequence obtained by the target three-dimensional prediction model can correspond to the three-dimensional model of the object to be modeled, that is, the correspondence between the output of the three-dimensional prediction model and the three-dimensional model is realized, so that the three-dimensional model can be obtained using the target coordinate sequence, which helps to solve the three-dimensional reconstruction problem using the three-dimensional prediction model.

[0102] In some embodiments, the sample object includes A parts; method 200 also includes: obtaining a coordinate set of the parts in the A parts, the coordinate set of the parts including the coordinates of B faces constituting the parts; obtaining a face set based at least on the coordinate set of the parts, wherein A and B are positive integers, and the value of M is related to the values ​​of A and B.

[0103] Taking the cabinet scenario as an example, the cabinet includes A flat panels, each of which includes B faces. M can be equal to A multiplied by B.

[0104] For each plane plate, its coordinate set can be obtained as follows:

[0105] Cuboid(x min ,y min ,z min ,x max ,y max ,z max ) represents the coordinates of the six faces that make up the plane plate. Then the coordinates of all the components (all the plane plates) in the sample object are put into a set to obtain a face set.

[0106] This embodiment can represent all parts of the sample object in the form of a coordinate set, so that the labeled coordinate sequence obtained based on the face set can correspond to the various parts of the sample object, thereby converting the three-dimensional reconstruction problem into a sample construction problem of the three-dimensional prediction model, so that the target three-dimensional prediction model can be directly used to solve the difficulty of three-dimensional reconstruction.

[0107] Figure 3 This is a schematic diagram of the structure of the three-dimensional prediction model provided according to the embodiment of the present disclosure; please refer to Figure 3 ,The structure and processing process of the three-dimensional preset model are described below.

[0108] In some embodiments, the sample object includes M faces, the predicted coordinate sequence includes M predicted coordinates, the predicted coordinates among the M predicted coordinates are predicted coordinates of faces among the M faces, and M is a positive integer greater than or equal to 2.

[0109] Inputting the first sample sequence of the sample object into a preset three-dimensional prediction model to obtain a prediction coordinate sequence in step S201 may include the following steps 4 to 6.

[0110] Step 4: Input the first sample sequence into the encoder in the preset three-dimensional prediction model to obtain a first embedded feature.

[0111] Step 5: Input the first embedded feature and the t-1th historical sample sequence into the decoder of the preset three-dimensional prediction model to obtain the tth prediction coordinate, where the t-1th historical sample sequence is obtained based on the t-1th prediction coordinate before the tth prediction coordinate, and t is a positive integer greater than or equal to 2 and less than or equal to M;

[0112] Step six: obtain a predicted coordinate sequence based on at least the t-th predicted coordinate.

[0113] like Figure 3 As shown, the preset three-dimensional prediction model may include an encoder (Transformer Encoder) and a decoder (Transformer Decoder).

[0114] The first sample sequence E(v1) to E(v N ) Input the encoder to obtain the first embedding feature.

[0115] For the t-th predicted coordinate in the predicted coordinate sequence, when t=1, the first embedded feature is used as the input of the decoder to obtain the first predicted coordinate.

[0116] When t is a positive integer greater than 1 and less than or equal to M, the t-1th historical sample sequence can be obtained based on the t-1th predicted coordinate before the t-th predicted coordinate, and then the t-1th historical sample sequence and the first embedded feature are used as the input of the decoder to obtain the t-th predicted coordinate.

[0117] Taking t=3 as an example, after obtaining the second predicted coordinate, the second historical sample sequence can be obtained based on the second predicted coordinate and the first predicted coordinate. Then, the second historical sample sequence and the first embedded feature are used as the input of the decoder to predict the third predicted coordinate.

[0118] Finally, a predicted coordinate sequence can be obtained based on the obtained M predicted coordinates. It will be appreciated that the preset three-dimensional prediction model generates the predicted coordinate sequence based on an autoregressive neural network, processing the output of the previous historical moment and using it as input to predict the predicted coordinates for the current moment. Of course, in other embodiments, other neural networks, such as recurrent neural networks, can also be used to implement the prediction of the predicted coordinate sequence.

[0119] It can be understood that M is a positive integer greater than or equal to 2. In this embodiment, when the sample object is a cabinet, M can be an integer multiple of 6.

[0120] In this embodiment, the output of the previous historical moment of the preset three-dimensional prediction model is processed and used as input to predict the predicted coordinates of the current moment, so that the current coordinates can be predicted in combination with the previous predicted coordinates. The potential correlation between the various predicted coordinates is utilized to simplify complex problems, and the prediction accuracy is high, thereby improving the accuracy of the target three-dimensional prediction model in predicting the target coordinates.

[0121] In some embodiments, in step five, inputting the first embedded feature and the t-1th historical sample sequence into the decoder of the preset three-dimensional prediction model to obtain the tth prediction coordinate may include: inputting the first embedded feature and the t-1th historical sample sequence into the decoder to obtain the tth hidden feature; determining the tth conditional probability distribution result based on the tth hidden feature; and determining the tth prediction coordinate based on the tth conditional probability distribution result.

[0122] like Figure 3 , the first embedded feature and the t-1th historical sample sequence are used as the input of the decoder, and the output of the decoder is the tth hidden feature h t Then we can use h t The t-th conditional probability distribution result is obtained. It can be understood that the t-th conditional probability distribution result is the conditional probability distribution result of the t-th predicted coordinate, that is, the candidate values ​​of the t-th predicted coordinate and the probability of each candidate value.

[0123] Then, the tth predicted coordinate can be determined according to the tth conditional probability distribution result.

[0124] In this embodiment, the prediction of the predicted coordinates can be converted into a series of conditional probability distribution problem solving problems, so that the model structure of the conditional probability distribution solution can be used to obtain the predicted coordinates, thereby converting the geometric problem from a two-dimensional view to a three-dimensional model into a conditional probability distribution problem that can be solved using a three-dimensional prediction model, overcoming the difficulty of using a model to solve.

[0125] In some embodiments, determining the tth conditional probability distribution result based on the tth hidden feature may include the following sub-steps 1 to 3.

[0126] Sub-step 1: Determine a first conditional probability distribution and an attachment probability based on the t-th hidden feature, where the first conditional probability distribution is the probability distribution of candidate values ​​of the t-th predicted coordinate.

[0127] Sub-step two: determine a second conditional probability distribution based on the t-th hidden feature and the t-1th historical hidden feature sequence, where the t-1th historical hidden feature sequence includes the t-1th hidden features output by the encoder that precede the t-th hidden feature; the second conditional probability distribution is the probability distribution of the candidate attachment relationships of the t-th predicted coordinate, where the candidate attachment relationships are used to point to the faces among the M faces that have an attachment relationship with the current face, and the current face is the face among the M faces corresponding to the t-th predicted coordinate.

[0128] Sub-step three: determining the tth conditional probability distribution result based on the first conditional probability distribution, the second conditional probability distribution, and the dependent probability.

[0129] It can be understood that in order to solve the conditional probability distribution problem, we decompose the joint distribution into a series of conditional probability distribution solutions:

[0130]

[0131] in, Indicates that given the first sample sequence Solving the predicted coordinate sequence The probability distribution of . Predicted coordinate sequence is a sequence of predicted coordinates arranged according to a preset rule, Sorting method and labeling coordinate sequence For details, please refer to the above embodiment of marking coordinate sequence.

[0132] g′ t Represents the predicted coordinate sequence The t-th predicted coordinate in the t-th predicted coordinate can be a coordinate value (value ) or face identification (dependency Indicates that the coordinate value is The surface and coordinate values ​​are The surface has an attachment relationship, q is a positive integer greater than or equal to 1 and less than t). This value is the same as the annotation coordinate g i The values ​​are the same, please refer to the marked coordinate g for details i The value of .

[0133] Indicates that given the first sample sequence Based on the t-1th historical sample sequence Solve for the predicted coordinate g′ t The probability distribution of is the t-th conditional probability distribution result. By solving the t-th conditional probability distribution result from t = 1 to t = M, we can get

[0134] It can be understood that the t-th conditional probability distribution result can include a fixed-length dictionary (the set of all possible discrete coordinate values) and a variable-length output sequence That is g′ t The candidate values ​​can be coordinate values ​​from a fixed-length dictionary or from a variable-length output sequence. The ID of the face to which the dependency in .

[0135] It can be understood that the fixed-length dictionary can correspond to the first conditional probability distribution ( Abbreviation p vocab ). Variable-length output sequence It can correspond to the second conditional probability distribution ( Abbreviation p attach ).

[0136] g′ t Candidate values ​​for p vocab The predicted coordinate values ​​can also come from p attach The ID of the face to which the predicted dependency refers.

[0137] For the first conditional probability distribution, that is, the distribution of the previous fixed length, the discrete probability distribution in the classification problem can be used to solve it. Let h t Representing the hidden feature output by the model at time t (the tth hidden feature), it can be projected to the dimension of the discrete value set size through a linear layer (linear), and normalized to a reasonable probability distribution through the normalized exponential function (softmax):

[0138]

[0139] For the second conditional probability distribution, in order to generate a The probability distribution on can be obtained using a Pointer Network. First, a linear layer is used to predict a pointer vector. Then, the inner product between the pointer vector and the feature predicted at the previous time t-1 (the t-1th historical hidden feature sequence) is calculated. Finally, the softmax is used to normalize it to a reasonable probability distribution:

[0140]

[0141] Among them, h <t It refers to the t-1 hidden features before time t, that is, the t-1th historical hidden feature sequence. The value of k is a positive integer from 1 to t-1.

[0142] In addition, according to h t Determine an attachment probability w tTo weigh the first conditional probability distribution and the second conditional probability distribution, we can avoid directly comparing the two probability distributions. It can be understood that the probability of the first conditional probability distribution is in the range of 0-1, and the probability value of the second conditional probability distribution is also in the range of 0-1. Through the attachment probability, these two conditional probabilities can be combined in the same range of 0-1. The attachment probability can be predicted by a linear layer linear and Sigmoid (S-type function) activation function σ(·): w t =σ(linear(h t )).

[0143] The t-th conditional probability distribution result is obtained based on the first conditional probability distribution, the second conditional probability distribution, and the dependent probability, for example, by weightedly combining the first conditional probability distribution and the second conditional probability distribution using the dependent probability:

[0144]

[0145] Among them, concat means finding the weighted distribution.

[0146] like Figure 3 , taking t=3 as an example, the decoder can obtain the third hidden feature h3 based on the first embedded feature and the second historical sample sequence, and the first conditional probability distribution p can be obtained based on the third hidden feature h3 vocab And the attachment probability w3. Then, the second historical hidden feature sequence (including h1 and h2) can be determined to continue to obtain the second conditional probability distribution p according to the second historical hidden feature sequence and the third hidden feature h3. attach Finally, we rely on the first conditional probability distribution p vocab , the second conditional probability distribution p attach The third prediction coordinate can be determined based on the third conditional probability distribution result.

[0147] In this embodiment, the preset three-dimensional prediction model uses a standard encoder and decoder architecture, and the parameters of the preset three-dimensional prediction model can be optimized by a standard cross-entropy loss function. The encoder and decoder can both be composed of 6 Transformer layers. Given a first sample sequence {E(v1), E(v2), ...} consisting of word embeddings, the encoder encodes it to obtain a context embedding feature (first embedding feature). At time t, the decoder receives the context embedding feature and the output sequence of the previous t-1 time (the t-1th historical sample sequence) {E(g1), E(g2), ...}, and the decoder predicts the hidden feature h at time t t .

[0148] Through this embodiment, the prediction of the t-th predicted coordinate can be realized, and the value of the predicted coordinate can be realized as a numerical value or an identifier of the surface pointed to by the dependency relationship.

[0149] In some embodiments, determining the tth predicted coordinate according to the tth conditional probability distribution result may include: determining a candidate result corresponding to the maximum probability from the tth conditional probability distribution result; and obtaining the tth predicted coordinate based on the candidate result.

[0150] It can be understood that the t-th conditional probability distribution result includes the candidate values ​​of the t-th predicted coordinate and the probability of each candidate value.

[0151] The candidate value (ie, candidate result) corresponding to the maximum probability can be determined from the candidate values, and then the t-th predicted coordinate is obtained according to the candidate result.

[0152] By determining the predicted coordinates with maximum probability, the accuracy of the prediction can be improved. At the same time, it has low dependence on two-dimensional views and is insensitive to errors in drawings, thereby improving the accuracy and robustness of three-dimensional model reconstruction.

[0153] In some embodiments, when the candidate result is a candidate numerical value in the first conditional probability distribution, the tth predicted coordinate is the candidate numerical value; or, when the candidate result is a candidate dependency in the second conditional probability distribution, the tth predicted coordinate is the identifier of the face pointed to by the candidate dependency.

[0154] It can be understood that since the t-th conditional probability distribution result may come from the first conditional probability distribution or the second conditional probability distribution, when the candidate result is a candidate value from the first conditional probability distribution, the candidate value can be used as the value of the t-th predicted coordinate.

[0155] When the candidate result is a candidate attachment relationship from the second conditional probability distribution, the identifier of the face that has an attachment relationship with the current face and to which the candidate attachment relationship points may be used as the value of the t-th predicted coordinate.

[0156] By determining the value of the t-th predicted coordinate through this embodiment, the output of the three-dimensional prediction model can be a coordinate value or a face identifier.

[0157] In some embodiments, the t-1th historical sample sequence includes t-1 transformation features, and the sth transformation feature among the t-1 transformation features is obtained based on the sth prediction coordinate; the sth transformation feature includes at least one of the following: a surface numerical feature, a component position feature, and a surface position feature; wherein the surface numerical feature is used to represent the coordinate value of the sth prediction coordinate; the component position feature is used to represent the relative position of the first component where the sth prediction coordinate is located in the sample object, and the first component is one of the multiple components included in the sample object; the surface position feature is used to represent the relative position of the surface corresponding to the sth prediction coordinate in the first component.

[0158] When generating the t-1th historical sample sequence based on the first t-1 predicted coordinates, the corresponding transformation features can be obtained by embedding each of the first t-1 predicted coordinates, and the t-1th historical sample sequence can be generated based on the obtained t-1 transformation features. The s-th predicted coordinate g′ in s yes:

[0159]

[0160] The values ​​of s and q are both positive integers from 1 to M, and s is not equal to q.

[0161] For example, the corresponding g′ s Get its feature embedding through word embedding:

[0162]

[0163] Among them, E(g′ s ) is the sth transformation feature.

[0164] The face value feature representing the s-th transformed feature, that is, the coordinate value used to represent the s-th predicted coordinate, that is, E value Represents the discretized coordinate value.

[0165] The component position feature representing the s-th transformed feature is used to represent the relative position of the first component at the s-th predicted coordinate in the sample object, where the first component is one of the multiple components included in the sample object. plank Indicates the relative position of the plane panel corresponding to the predicted coordinates in the cabinet model.

[0166] The surface position feature representing the s-th transformed feature is used to represent the relative position of the surface corresponding to the s-th predicted coordinate in the first component. faceIndicates the relative position of the surface corresponding to the predicted coordinate in the plane plate.

[0167] It can be understood that although the model can predict the predicted coordinates, the value of the predicted coordinates can be the coordinate value or the face identifier. When determining the sth transformation feature, when the value of the sth predicted coordinate is the face identifier, that is, when g′ s Corresponding to the edge When the corresponding E value Use the coordinates of the pointed face As input to calculate the s-th conversion feature.

[0168] It can be understood that the conversion feature may include at least one of a surface value feature, a component position feature, and a surface position feature.

[0169] This embodiment can obtain the sth transformation feature based on the sth predicted coordinate, and then obtain the t-1th historical sample sequence. By setting the transformation feature in the historical sample sequence to include at least one of the surface numerical feature, the component position feature and the surface position feature, when the target three-dimensional prediction model predicts the object to be modeled, the target coordinates in the target coordinate sequence output by the target three-dimensional prediction model can correspond to the position features of each component and surface in the three-dimensional model, so that the three-dimensional model of the object to be modeled can be obtained using the target coordinate sequence.

[0170] Figure 4 This is a flow chart of the 3D reconstruction method provided by the embodiment of the present disclosure. Figure 4 The embodiment of the present disclosure further provides a three-dimensional reconstruction method 400, comprising the following steps S401 to S403.

[0171] Step S401: Generate a first sequence based on a two-dimensional view of an object to be modeled.

[0172] Step S402: input the first sequence into a target three-dimensional prediction model to obtain a target coordinate sequence, wherein the target three-dimensional prediction model is trained according to the three-dimensional prediction model training method of each of the above embodiments.

[0173] Step S403: generating a three-dimensional model of the object to be modeled according to the target coordinate sequence.

[0174] The target three-dimensional prediction model refers to a preset three-dimensional prediction model trained according to the above three-dimensional prediction model training method.

[0175] A 2D view is similar to a 2D sample view, which is an orthogonal projection of the object to be modeled. They are front view, top view and side view respectively.

[0176] The first sequence is a sequence obtained based on a two-dimensional view. The first sequence is determined in the same manner as the first sample sequence obtained based on a two-dimensional sample view. For details, reference may be made to the above embodiments of determining the first sample sequence.

[0177] The first sequence is input into the target three-dimensional prediction model to obtain a target coordinate sequence, where the target coordinate sequence includes a plurality of target coordinates, and these target coordinates respectively correspond to the coordinates of all faces in the object to be modeled.

[0178] The method of inputting the first sequence into the target three-dimensional prediction model to obtain the target coordinate sequence is the same as the method of inputting the first sample sequence into the preset three-dimensional prediction model to obtain the predicted coordinate sequence. For details, please refer to the various embodiments of determining the predicted coordinate sequence.

[0179] Since the target coordinates in the target coordinate sequence correspond to the respective faces of the object to be modeled, a three-dimensional model of the object to be modeled can be obtained from the target coordinate sequence.

[0180] Figure 5 This is a process diagram of generating a three-dimensional model according to the three-dimensional reconstruction method provided by the embodiment of the present disclosure; please refer to Figure 5 , take the cabinet shown in Figure (a) as an example. Figure 5 (b) to (i) show the timing diagram of the three-dimensional modeling of the cabinet.

[0181] First, the two-dimensional drawing of the cabinet can be processed to obtain a first sequence, and then the first sequence is input into the target three-dimensional prediction model.

[0182] The target coordinates in the target coordinate sequence output by the target three-dimensional prediction model (i.e., the values ​​in Cuboid()) are as follows:

[0183] bbox=Cuboid(-0.35,-0.23,-0.76,0.35,0.23,0.76)

[0184] plank1=Cuboid(bbox1,bbox2,bbox3,-0.34,bbox5,bbox6)

[0185] plank2=Cuboid(0.34,bbox2,bbox3,bbox4,bbox5,bbox6)

[0186] plank3=Cuboid(plank14,bbox2,-0.70,plank21,bbox5,-0.69)

[0187] plank4=Cuboid(plank14,bbox2,0.75,plank21,bbox5,bbox6)

[0188] plank5=Cuboid(plank14,0.21,plank36,plank21,0.22,plank43)

[0189] plank6=Cuboid(plank14,bbox2,bbox3,plank21,-0.21,plank33)

[0190] plank7=Cuboid(plank14,0.21,bbox3,plank21,bbox5,plank33)

[0191] In the process of 3D modeling, we can first get the bbox ( Figure 5 (b) The six target coordinates of the cabinet frame (-0.35, -0.23, -0.76, 0.35, 0.23, 0.76), which are the maximum dimensions of the cabinet. Then plank1 ( Figure 5 The six target coordinates (bbox1, bbox2, bbox3, -0.34, bbox5, bbox6) of the plane plate 1 (shown in dark in (c)) can be modeled based on these six coordinates. Figure 5 (c) The planar plate is shown in dark color.

[0192] By analogy, we can obtain the 3D model of the seven planar panels (plank1 to plank7) of the entire cabinet. The order of the bbox and plank1 to plank7 is determined by the second preset arrangement rule mentioned above. This means that each component and its corresponding coordinates are sorted when generating the annotated coordinate sequence. This ordering rule is also used for prediction when obtaining the target coordinate sequence.

[0193] It can be understood that the prediction order of the target coordinate sequence is bbox, plank1, plank2, ..., plank7. That is, first get the first target coordinate -0.35 in the bbox, then get -0.23 based on -0.35 and the first sequence, and predict all the target coordinates in sequence.

[0194] Among them, taking plank1=Cuboid(bbox1,bbox2,bbox3,-0.34,bbox5,bbox6) as an example, after predicting the 6 target coordinates in bbox, the first target coordinate of plank1 can be predicted. At this time, bbox1 refers to the identifier of the first target coordinate in bbox (that is, the identifier of the first face in bbox, the corresponding value is -0.35). Bbox1 indicates that the first face of plank1 has an attachment relationship with the first face in bbox. It can be understood that here the identifier bbox1 of the face pointed to by the attachment relationship is used to represent the first target coordinate of plank1. Similarly, plank21 refers to the identifier of the first face (target coordinate) in plank2.

[0195] Furthermore, it is understood that the number of target coordinates in this embodiment is the total number of the six faces of the cuboid corresponding to all planar panels and the six faces of the cuboid at the maximum dimension of the entire cabinet, which is 48. In other embodiments, only the total number of the six faces of the cuboid corresponding to all planar panels, which is 42, may be used. This depends on whether the annotated coordinate sequence used in the training of the preset 3D prediction model includes the coordinates of the six faces of the cuboid at the maximum dimension of the entire cabinet.

[0196] In this embodiment, a three-dimensional model of the object to be modeled can be generated through the target three-dimensional prediction model, so that the process of obtaining a three-dimensional model from a two-dimensional view can be converted into the input and output of a three-dimensional prediction model, and a correspondence between the two-dimensional view and the three-dimensional model is established, that is, the process of using a three-dimensional prediction model to solve the process from a two-dimensional view to a three-dimensional model is realized.

[0197] Furthermore, the related art modeling approach relies on the correspondence between 2D views and 3D models. If the 2D drawings contain incorrect coordinates, the generation of 3D vertices from 2D vertices will inevitably result in errors, and the 3D model construction will inevitably fail. Because the 3D model reconstruction of this embodiment utilizes a 3D prediction model and does not rely on vertex correspondence, it has a low dependency on the 2D views, is insensitive to errors in the drawings, and has a high tolerance for errors in the 2D views, thereby improving the accuracy and robustness of 3D model reconstruction.

[0198] In some embodiments, after generating the three-dimensional model of the object to be modeled according to the target coordinate sequence in step S403 , the method 400 further includes: modifying the target coordinate sequence to edit the three-dimensional model.

[0199] It can be understood that since the target coordinate sequence preferably uses the face identifier to identify the target coordinate, when modifying the three-dimensional model, there is no need to modify the coordinates of each component. It is only necessary to modify the value in the target coordinate sequence to realize the editing of the three-dimensional model, thereby improving the modification efficiency of the three-dimensional model.

[0200] The editing of a 3D model may include global editing and local editing. Global editing may include scaling the entire 3D model. Figure 6A This is a schematic diagram of a global editing of a 3D model according to a 3D reconstruction method provided by an embodiment of the present disclosure; please refer to Figure 6A The first row shows the 3D model before editing. The second row shows the new 3D model after the 3D model is enlarged as a whole by modifying the target coordinate sequence.

[0201] Local editing can include local modifications to one or several components in a 3D model. Figure 6B This is a schematic diagram of local editing of a 3D model according to the 3D reconstruction method provided by the embodiment of the present disclosure; please refer to Figure 6B The first row shows the 3D model before editing. The second row shows how to modify the height and width of one or more plane panels in the 3D model by modifying the target coordinate sequence.

[0202] The 3D prediction model training method and 3D reconstruction method provided by this disclosure utilize a three-view reconstruction algorithm based on sequential modeling. This method implicitly establishes the correspondence between 2D views and 3D models using a self-attention mechanism, making it highly robust to errors in the drawings. Furthermore, the reconstructed model is also suitable for downstream applications, such as model editing.

[0203] In a specific embodiment, a method for reconstructing a 3D model from three views of a cabinet is proposed. Thus, the reconstruction of the three views is modeled as a sequence generation problem. The input of the 3D prediction model is three orthogonal projection views. It includes front, top, and side views. Each view contains the coordinates of the 2D line segments and their intersection points. Visible lines are represented by solid lines, and invisible lines by dashed lines. The output of the 3D prediction model is a 3D model.

[0204] 1. Shape Program

[0205] First, we define a domain-specific language (DSL) for describing the 3D model of a cabinet. Typically, a cabinet consists of a set of planar panels, and these panels are usually placed along coordinate axes. Therefore, we use cubes (cuboids) along the coordinate axes to represent the planar panels.

[0206] Each cube has 6 degrees of freedom, corresponding to the starting and ending coordinates of the three coordinate axes: Cuboid (x min ,y min ,z min ,x max ,y max ,z max ).

[0207] The modeling language can define the coordinates of a cube by specifying numerical values ​​(coordinate values) or attachment operations (face identifiers). In other words, each coordinate in the above formula can be a coordinate value or a pointer to a coordinate (face identifier) ​​of the attached cube.

[0208] Then the cabinet shape program is converted into a directed acyclic graph. Each plane panel contains 6 faces, and each face corresponds to a coordinate value of Cuboid. The shape program can be expressed as a graph The set of its vertices contains the faces of all the flat panels that make up the cabinet The set of its edges represents all dependency relations ε={e1,…,e |ε|}. Each edge e i→j =(f i ,f j ) is a directed edge, starting from f i , the end point is f j , represents the i-th face f i Attached to the jth face f j ε can also be expressed as an adjacency matrix Specifically, if f i Attached to f j , then A ij =1; otherwise, A ij =0.

[0209] 2. Input sequence (first sequence or first sample sequence)

[0210] Convert all line segments in the three-view image into input sequences.

[0211] First, sort the two-dimensional segments by view, then by coordinate value. Then, flatten the segments into a one-dimensional sequence. Because each two-dimensional line segment has 4 degrees of freedom, corresponding to the x and y coordinate values ​​of the two endpoints, The length N v It is 4 times the number of all line segments in the three views.

[0212] will sequence Each element v in iGet its feature embedding (first feature) through word embedding:

[0213] E(v i )=E value (v i )+E view (v i )+E edge (v i )+E coord (v i )+E type (v i )

[0214] Among them, E value Represents the discretized coordinate value, E view Indicates which view the line segment comes from (front view, top view or side view), E edge Indicates the relative position of the line segment in the corresponding view, E coord Indicates the relative position of the coordinates on the line segment, E type Indicates whether the line segment is visible.

[0215] 3. Output sequence (predicted coordinate sequence, labeled coordinate sequence or target coordinate sequence)

[0216] The diagram corresponding to the board program will be described Convert to sequence

[0217] Specifically, we need to define the graph The order of vertices in the graph is π: First, we Perform topological sorting, which makes the starting point of any directed edge in the graph come before the end point. Then, for the remaining unconnected vertices, we sort them by coordinate values. After the above steps, we get an ordered graph. Its apex The order is according to π.

[0218] Similar to the input sequence, the ordered graph Flattened into a one-dimensional sequence Its i-th element is:

[0219]

[0220] Finally, two special tokens [SOS] and [EOS] are added to represent the start and end symbols of the output sequence respectively.

[0221] The corresponding f in the output sequence i π The element g i Get its feature embedding (transformation feature) through word embedding:

[0222] E(g i )=E(f i π )=E value (f i π )+E plank (f i π )+E face (f i π )

[0223] Among them, E value Represents the discretized coordinate value, E plank Indicates the relative position of the corresponding plane panel in the cabinet model, E face Indicates the relative position of the corresponding surface in the plane plate. i Corresponding to the edge When the corresponding E value Use the pointed face f i π as input.

[0224] 4. 3D prediction model design

[0225] In order to solve the above Seq2Seq problem, the joint distribution is decomposed into a series of conditional probability distribution solutions:

[0226]

[0227] Among them, g t It can be a numerical value f i π Or dependent operation This probability distribution consists of a fixed-length dictionary (the set of all possible discrete coordinate values) and a variable-length output sequence constitute.

[0228] like Figure 3 , the previous fixed length distribution is the discrete probability distribution in the classification problem. Let h t Representing the hidden features of the model output at time t, it is projected to the dimension of the discrete value set size through a linear layer and normalized to a reasonable probability distribution through softmax:

[0229]

[0230] In order to generate a The probability distribution of the above can be referenced by the Pointer Network. First, a linear layer is used to predict a pointer vector. Then, the inner product between the pointer vector and the feature predicted at the previous time t-1 (the t-1th historical hidden feature sequence) is calculated. Finally, the softmax is used to normalize it to a reasonable probability distribution:

[0231]

[0232] By predicting an attachment probability w t To weigh these two distributions, we avoid directly comparing the two probability distributions mentioned above. The attachment probability is predicted by a linear layer and the Sigmoid activation function σ(·): w t =σ(linear(h t )). The final probability distribution is a combination of the two weighted distributions above:

[0233]

[0234] The parameters of the predicted 3D model are optimized using a standard cross-entropy loss function. The predicted 3D model uses a standard encoder-decoder architecture, both consisting of 6 Transformer layers. Given an input sequence of embedded features {E(v1), E(v2), …}, the encoder encodes it to obtain the context embedded feature (the first embedded feature). At time t, the decoder receives the context embedded feature and the output sequence of the previous t-1 time (the t-1th historical sample sequence) {E(g1), E(g2), …}, and the decoder predicts the hidden feature h at time t. t .

[0235] The three-dimensional prediction model provided in this embodiment can implement a three-view cabinet reconstruction algorithm based on Transformer. At the same time, for the cabinet model composed of flat panels, a set of domain-specific languages ​​is proposed with reference to the modeling habits of professionals.

[0236] Figure 7 This is a schematic diagram of the structure of the training device of the three-dimensional prediction model provided by the embodiment of the present disclosure; please refer to Figure 7 The embodiment of the present disclosure also provides a three-dimensional prediction model training device 700, which includes the following units.

[0237] The input unit 701 is configured to input a first sample sequence of the sample object into a preset three-dimensional prediction model to obtain a predicted coordinate sequence, wherein the first sample sequence is obtained based on a two-dimensional sample view of the sample object.

[0238] The training unit 702 is used to train a preset three-dimensional prediction model based on the labeled coordinate sequence and the predicted coordinate sequence of the sample object to obtain a target three-dimensional prediction model. The target three-dimensional prediction model is used to generate a target coordinate sequence of the object to be modeled based on the two-dimensional view of the object to be modeled, and the target coordinate sequence is used to generate a three-dimensional model of the object to be modeled.

[0239] In some embodiments, the sample object includes M faces, the predicted coordinate sequence includes M predicted coordinates, the predicted coordinates among the M predicted coordinates are the predicted coordinates of the faces among the M faces, and M is a positive integer greater than or equal to 2; the input unit 701 is also used to: input the first sample sequence into the encoder in the preset three-dimensional prediction model to obtain a first embedded feature; input the first embedded feature and the t-1th historical sample sequence into the decoder of the preset three-dimensional prediction model to obtain the tth predicted coordinate, wherein the t-1th historical sample sequence is based on the t-1th predicted coordinate before the tth predicted coordinate, and t is a positive integer greater than or equal to 2 and less than or equal to M; and obtain the predicted coordinate sequence based at least on the tth predicted coordinate.

[0240] In some embodiments, the input unit 701 is further used to: input the first embedded feature and the t-1th historical sample sequence into the decoder to obtain the tth hidden feature; determine the tth conditional probability distribution result based on the tth hidden feature; and determine the tth predicted coordinate based on the tth conditional probability distribution result.

[0241] In some embodiments, the input unit 701 is also used to: determine a first conditional probability distribution and an attachment probability based on the tth hidden feature, the first conditional probability distribution being the probability distribution of candidate values ​​of the tth predicted coordinate; determine a second conditional probability distribution based on the tth hidden feature and the t-1th historical hidden feature sequence, the t-1th historical hidden feature sequence including the t-1 hidden features output by the encoder that are located before the tth hidden feature; the second conditional probability distribution is the probability distribution of candidate attachment relationships of the tth predicted coordinate, the candidate attachment relationships are used to point to a face among the M faces that has an attachment relationship with the current face, and the current face is the face among the M faces corresponding to the tth predicted coordinate; determine the tth conditional probability distribution result based on the first conditional probability distribution, the second conditional probability distribution and the attachment probability.

[0242] In some embodiments, the input unit 701 is further used to: determine the candidate result corresponding to the maximum probability from the t-th conditional probability distribution result; and obtain the t-th predicted coordinate based on the candidate result.

[0243] In some embodiments, when the candidate result is a candidate numerical value in the first conditional probability distribution, the tth predicted coordinate is the candidate numerical value; or, when the candidate result is a candidate dependency in the second conditional probability distribution, the tth predicted coordinate is the identifier of the face pointed to by the candidate dependency.

[0244] In some embodiments, the t-1th historical sample sequence includes t-1 transformation features, and the sth transformation feature among the t-1 transformation features is obtained based on the sth prediction coordinate; the sth transformation feature includes at least one of the following: a surface numerical feature, a component position feature, and a surface position feature; wherein the surface numerical feature is used to represent the coordinate value of the sth prediction coordinate; the component position feature is used to represent the relative position of the first component where the sth prediction coordinate is located in the sample object, and the first component is one of the multiple components included in the sample object; the surface position feature is used to represent the relative position of the surface corresponding to the sth prediction coordinate in the first component.

[0245] In some embodiments, the apparatus 700 further includes: a conversion unit configured to obtain multiple line segments contained in the two-dimensional sample view; convert the coordinates of the multiple line segments into a line segment coordinate sequence; and convert the line segment coordinate sequence into a first sample sequence.

[0246] In some embodiments, the conversion unit is further configured to: sort the plurality of line segments according to a first preset rule; and convert the coordinates of the sorted plurality of line segments into a line segment coordinate sequence.

[0247] In some embodiments, a line segment coordinate sequence includes N line segment coordinates, a first sample sequence includes N first features, the rth first feature among the N first features is obtained based on the rth line segment coordinate among the N line segment coordinates, and N is a positive integer; the rth first feature includes at least one of the following: a numerical feature, a view feature, a view position feature, a line segment position feature, and a visual feature; wherein the numerical feature is used to represent the coordinate value of the rth line segment coordinate; the view feature is used to represent the two-dimensional sample view where the rth line segment coordinate is located; the line segment position feature is used to represent the relative position of the first line segment where the rth line segment coordinate is located in the two-dimensional sample view, where the first line segment is one of multiple line segments; the coordinate position feature is used to represent the relative position of the rth line segment coordinate in the first line segment; and the visual feature is used to represent the visual state of the first line segment.

[0248] In some embodiments, the device 700 also includes: a sorting unit, used to determine a directed acyclic graph of the sample object, the directed acyclic graph includes a dependency relationship set and a face set, the dependency relationship in the dependency relationship set is used to represent two faces with dependency relationships among the M faces of the sample object, and the face set includes the coordinates of the M faces; according to the second preset rule, the directed acyclic graph is sorted to obtain an ordered graph of the sample object; the ordered graph is converted into a labeled coordinate sequence, the labeled coordinate sequence includes M labeled coordinates arranged according to preset rules, and the labeled coordinates among the M labeled coordinates are the coordinates of the faces among the M faces.

[0249] In some embodiments, when the first face corresponding to the marked coordinate is located in the dependency set, the marked coordinate is the identifier of the second face that has a dependency relationship with the first face; or, when the first face corresponding to the marked coordinate is not located in the dependency set, the marked coordinate is the coordinate value of the first face; wherein the first face and the second face are different faces among the M faces.

[0250] In some embodiments, the sample object includes A parts; the device also includes: an acquisition unit, used to obtain a coordinate set of parts in the A parts, the coordinate set of the parts includes the coordinates of B faces constituting the parts; at least based on the coordinate set of the parts, a face set is obtained, where A and B are positive integers, and the value of M is related to the values ​​of A and B.

[0251] Figure 8 is a schematic diagram of the structure of a three-dimensional reconstruction device according to an embodiment of the present disclosure; please refer to Figure 8 , an embodiment of the present disclosure provides a three-dimensional reconstruction device 800, including the following units.

[0252] The first generating unit 801 is configured to generate a first sequence according to a two-dimensional view of the object to be modeled.

[0253] The prediction unit 802 is configured to input the first sequence into a target three-dimensional prediction model to obtain a target coordinate sequence, wherein the target three-dimensional prediction model is trained according to the three-dimensional prediction model training method of any of the above embodiments.

[0254] The second generating unit 803 is configured to generate a three-dimensional model of the object to be modeled according to the target coordinate sequence.

[0255] In some embodiments, after generating a three-dimensional model of the object to be modeled according to the target coordinate sequence, the apparatus further includes: an editing unit configured to modify the target coordinate sequence to edit the three-dimensional model.

[0256] According to an embodiment of the present disclosure, the present disclosure also provides an electronic device, a readable storage medium, and a computer program product.

[0257] An embodiment of the present disclosure provides an electronic device, comprising: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the method of any one of the above embodiments.

[0258] An embodiment of the present disclosure provides a non-transitory computer-readable storage medium storing computer instructions, wherein the computer instructions are used to enable a computer to execute the method of any one of the above embodiments.

[0259] An embodiment of the present disclosure provides a computer program product, including a computer program, which implements the method described in any one of the above embodiments when executed by a processor.

[0260] For the description of specific functions and examples of each module and submodule of the device in the embodiment of the present disclosure, please refer to the relevant description of the corresponding steps in the above method embodiment, which will not be repeated here.

[0261] Figure 9 A schematic block diagram of an example electronic device 900 that can be used to implement embodiments of the present disclosure is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital assistants, cellular phones, smartphones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the present disclosure described and / or claimed herein.

[0262] like Figure 9 As shown, the device 900 includes a computing unit 901, which can perform various appropriate actions and processes according to a computer program stored in a read-only memory (ROM) 902 or a computer program loaded from a storage unit 908 into a random access memory (RAM) 903. Various programs and data required for the operation of the device 900 can also be stored in the RAM 903. The computing unit 901, the ROM 902, and the RAM 903 are connected to each other via a bus 904. An input / output (I / O) interface 905 is also connected to the bus 904.

[0263] Various components in the device 900 are connected to the I / O interface 905, including an input unit 906, such as a keyboard, a mouse, etc.; an output unit 907, such as various types of displays, speakers, etc.; a storage unit 908, such as a magnetic disk, an optical disk, etc.; and a communication unit 909, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 909 allows the device 900 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks.

[0264] The computing unit 901 can be a variety of general-purpose and / or specialized processing components with processing and computing capabilities. Some examples of the computing unit 901 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various computing units that run machine learning model algorithms, a digital signal processor (DSP), and any appropriate processor, controller, microcontroller, etc. The computing unit 901 performs the various methods and processes described above, such as the training method and the 3D prediction model 3D reconstruction method. For example, in some embodiments, the training method and the 3D reconstruction method of the 3D prediction model can be implemented as a computer software program, which is tangibly contained in a machine-readable medium, such as the storage unit 908. In some embodiments, part or all of the computer program can be loaded and / or installed on the device 900 via the ROM 902 and / or the communication unit 909. When the computer program is loaded into the RAM 903 and executed by the computing unit 901, one or more steps of the training method and the 3D reconstruction method described above can be performed. Alternatively, in other embodiments, the computing unit 901 may be configured to execute the three-dimensional prediction model training method and the three-dimensional reconstruction method in any other appropriate manner (for example, by means of firmware).

[0265] Various embodiments of the systems and techniques described herein can be implemented in digital electronic circuit systems, integrated circuit systems, field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), system-on-chip systems (SOCs), programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include being implemented in one or more computer programs that are executable and / or interpreted on a programmable system comprising at least one programmable processor, which can be a special purpose or general purpose programmable processor that can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit data and instructions to the storage system, the at least one input device, and the at least one output device.

[0266] The program code for implementing the method of the present disclosure can be written in any combination of one or more programming languages. These program codes can be provided to a processor or controller of a general-purpose computer, a special-purpose computer, or other programmable data processing device so that when the program code is executed by the processor or controller, the functions / operations specified in the flow chart and / or block diagram are implemented. The program code can be executed entirely on the machine, partially on the machine, as a stand-alone software package, partially on the machine and partially on a remote machine, or entirely on a remote machine or server.

[0267] In the context of the present disclosure, a machine-readable medium can be a tangible medium that can contain or store a program for use by or in conjunction with an instruction execution system, device or equipment. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, device or equipment, or any suitable combination of the foregoing. A more specific example of a machine-readable storage medium can include an electrical connection based on one or more lines, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0268] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user can provide input to the computer. Other types of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, voice input, or tactile input).

[0269] The systems and techniques described herein can be implemented in a computing system that includes back-end components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes front-end components (e.g., a user computer having a graphical user interface or a web browser through which a user can interact with implementations of the systems and techniques described herein), or a computing system that includes any combination of such back-end components, middleware components, or front-end components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include a local area network (LAN), a wide area network (WAN), and the Internet.

[0270] A computer system may include a client and a server. The client and server are generally remote from each other and typically interact through a communication network. The client-server relationship arises through computer programs running on the respective computers and having a client-server relationship with each other. The server may be a cloud server, a server in a distributed system, or a server integrated with a blockchain.

[0271] It should be understood that the various forms of the processes shown above can be used to reorder, add, or delete steps. For example, the steps described in this disclosure can be performed in parallel, sequentially, or in a different order, as long as the desired results of the technical solutions disclosed in this disclosure can be achieved. This is not limited herein.

[0272] The above specific embodiments do not constitute a limitation on the scope of protection of this disclosure. Those skilled in the art will appreciate that various modifications, combinations, sub-combinations, and substitutions may be made based on design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the principles of this disclosure shall be included within the scope of protection of this disclosure.

Claims

1. A training method for a three-dimensional prediction model, comprising: Inputting a first sample sequence of a sample object into a preset three-dimensional prediction model to obtain a predicted coordinate sequence, wherein the sample object is a three-dimensional object; the first sample sequence is obtained based on a two-dimensional sample view of the sample object; the two-dimensional sample view of the sample object includes a two-dimensional three-view image of the sample object, and the first sample sequence includes coordinates of all line segments in the two-dimensional sample view; Training the preset three-dimensional prediction model based on the labeled coordinate sequence of the sample object and the predicted coordinate sequence to obtain a target three-dimensional prediction model, wherein the target three-dimensional prediction model is used to generate a target coordinate sequence of the object to be modeled based on the two-dimensional view of the object to be modeled, and the target coordinate sequence is used to generate a three-dimensional model of the object to be modeled; The sample object includes M faces, the predicted coordinate sequence includes M predicted coordinates, a predicted coordinate in the M predicted coordinates is a predicted coordinate of a face in the M faces, and M is a positive integer greater than or equal to 2; The step of inputting the first sample sequence of the sample object into a preset three-dimensional prediction model to obtain a predicted coordinate sequence includes: Inputting the first sample sequence into an encoder in the preset three-dimensional prediction model to obtain a first embedding feature; the first embedding feature is a context embedding feature; Inputting the first embedded feature and the t-1th historical sample sequence into a decoder of the preset three-dimensional prediction model to obtain a t-th prediction coordinate, wherein the t-1th historical sample sequence is obtained based on the t-1th prediction coordinate before the t-th prediction coordinate, and t is a positive integer greater than or equal to 2 and less than or equal to M; A predicted coordinate sequence is obtained based on at least the t-th predicted coordinate.

2. The method according to claim 1, wherein Inputting the first embedded feature and the t-1th historical sample sequence into the decoder of the preset three-dimensional prediction model to obtain the tth predicted coordinate includes: Inputting the first embedded feature and the t-1th historical sample sequence into the decoder to obtain the tth hidden feature; Determine the tth conditional probability distribution result according to the tth hidden feature; Determine the tth predicted coordinate according to the tth conditional probability distribution result.

3. The method according to claim 2, wherein: Determining the tth conditional probability distribution result according to the tth hidden feature includes: Determine a first conditional probability distribution and an attachment probability based on the t-th hidden feature, wherein the first conditional probability distribution is a probability distribution of candidate values ​​of the t-th predicted coordinate; Determining a second conditional probability distribution based on the t-th hidden feature and a t-1-th historical hidden feature sequence, where the t-1-th historical hidden feature sequence includes the t-1 hidden features output by the encoder and located before the t-th hidden feature; the second conditional probability distribution is a probability distribution of candidate attachment relationships for the t-th predicted coordinate, where the candidate attachment relationships are used to point to a face among the M faces that has a dependency relationship with a current face, where the current face is the face among the M faces corresponding to the t-th predicted coordinate; The tth conditional probability distribution result is determined according to the first conditional probability distribution, the second conditional probability distribution, and the attachment probability.

4. The method according to claim 3, wherein: Determining the t-th predicted coordinate according to the t-th conditional probability distribution result includes: Determine the candidate result corresponding to the maximum probability from the t-th conditional probability distribution result; The t-th predicted coordinate is obtained based on the candidate results.

5. The method according to claim 4, wherein In a case where the candidate result is a candidate value in the first conditional probability distribution, the t-th predicted coordinate is the candidate value; Alternatively, when the candidate result is a candidate dependency relationship in the second conditional probability distribution, the t-th predicted coordinate is the identifier of the face pointed to by the candidate dependency relationship.

6. The method according to any one of claims 1 to 5, wherein: The t-1th historical sample sequence includes t-1 transformation features, the sth transformation feature among the t-1 transformation features is obtained based on the sth prediction coordinate; the sth transformation feature includes at least one of the following: a surface numerical feature, a component position feature, and a surface position feature; Wherein, the face numerical feature is used to represent the coordinate value of the s-th predicted coordinate; The component position feature is used to represent the relative position of the first component where the s-th predicted coordinate is located in the sample object, where the first component is one of the multiple components included in the sample object; The surface position feature is used to represent the relative position of the surface corresponding to the s-th predicted coordinate in the first component.

7. The method according to any one of claims 1 to 5, further comprising: Acquire a plurality of line segments contained in the two-dimensional sample view; Converting the coordinates of the plurality of line segments into a line segment coordinate sequence; The line segment coordinate sequence is converted into a first sample sequence.

8. The method according to claim 7, wherein: Converting the coordinates of the plurality of line segments into a line segment coordinate sequence includes: sorting the plurality of line segments according to a first preset rule; The sorted coordinates of the plurality of line segments are converted into a line segment coordinate sequence.

9. The method according to claim 7, wherein: The line segment coordinate sequence includes N line segment coordinates, the first sample sequence includes N first features, an r-th first feature among the N first features is obtained based on the r-th line segment coordinate among the N line segment coordinates, where N is a positive integer; the r-th first feature includes at least one of the following: a numerical feature, a view feature, a view position feature, a line segment position feature, and a visual feature; Wherein, the numerical feature is used to represent the coordinate value of the r-th line segment coordinate; The view feature is used to represent the two-dimensional sample view where the r-th line segment coordinate is located; The line segment position feature is used to represent the relative position of the first line segment where the r-th line segment coordinate is located in the two-dimensional sample view, where the first line segment is one of the multiple line segments; The view position feature is used to represent the relative position of the r-th line segment coordinate in the first line segment; The visual feature is used to indicate the visual state of the first line segment.

10. The method according to any one of claims 1 to 5, further comprising: Determining a directed acyclic graph of the sample object, the directed acyclic graph comprising a dependency relationship set and a face set, wherein the dependency relationship in the dependency relationship set is used to represent two faces having a dependency relationship among M faces of the sample object, and the face set comprises coordinates of the M faces; Sorting the directed acyclic graph according to a second preset rule to obtain an ordered graph of the sample objects; The ordered graph is converted into the labeled coordinate sequence, where the labeled coordinate sequence includes M labeled coordinates arranged according to a preset rule, and a labeled coordinate among the M labeled coordinates is a coordinate of a surface among the M surfaces.

11. The method according to claim 10, wherein: In a case where the first surface corresponding to the marked coordinates is located in the dependency relationship set, the marked coordinates are identifiers of the second surface having a dependency relationship with the first surface; Alternatively, when the first surface corresponding to the marked coordinates is not located in the dependency relationship set, the marked coordinates are the coordinate values ​​of the first surface; The first surface and the second surface are different surfaces among the M surfaces.

12. The method according to claim 10, wherein: The sample object includes A components; The method further comprises: Obtaining a coordinate set of a component in the A components, the coordinate set of the component including coordinates of B faces constituting the component; The face set is obtained based at least on the coordinate set of the component, wherein A and B are positive integers, and the value of M is related to the values ​​of A and B.

13. A three-dimensional reconstruction method, comprising: generating a first sequence according to a two-dimensional view of the object to be modeled; inputting the first sequence into a target three-dimensional prediction model to obtain a target coordinate sequence, wherein the target three-dimensional prediction model is trained by the method according to any one of claims 1 to 12; A three-dimensional model of the object to be modeled is generated according to the target coordinate sequence.

14. The method according to claim 13, after generating the three-dimensional model of the object to be modeled according to the target coordinate sequence, the method further comprises: The target coordinate sequence is modified to edit the three-dimensional model.

15. A training device for a three-dimensional prediction model, comprising: an input unit, configured to input a first sample sequence of a sample object into a preset three-dimensional prediction model to obtain a predicted coordinate sequence, wherein the sample object is a three-dimensional object; the first sample sequence is obtained based on a two-dimensional sample view of the sample object; the two-dimensional sample view of the sample object includes a two-dimensional three-view image of the sample object, and the first sample sequence includes coordinates of all line segments in the two-dimensional sample view; a training unit, configured to train the preset three-dimensional prediction model based on the labeled coordinate sequence of the sample object and the predicted coordinate sequence to obtain a target three-dimensional prediction model, wherein the target three-dimensional prediction model is used to generate a target coordinate sequence of the object to be modeled based on a two-dimensional view of the object to be modeled, and the target coordinate sequence is used to generate a three-dimensional model of the object to be modeled; The sample object includes M faces, the predicted coordinate sequence includes M predicted coordinates, a predicted coordinate in the M predicted coordinates is a predicted coordinate of a face in the M faces, and M is a positive integer greater than or equal to 2; Wherein, the input unit is further used for: Inputting the first sample sequence into an encoder in the preset three-dimensional prediction model to obtain a first embedding feature; the first embedding feature is a context embedding feature; Inputting the first embedded feature and the t-1th historical sample sequence into a decoder of the preset three-dimensional prediction model to obtain a t-th prediction coordinate, wherein the t-1th historical sample sequence is obtained based on the t-1th prediction coordinate before the t-th prediction coordinate, and t is a positive integer greater than or equal to 2 and less than or equal to M; A predicted coordinate sequence is obtained based on at least the t-th predicted coordinate.

16. The device according to claim 15, wherein The input unit is further configured to: Inputting the first embedded feature and the t-1th historical sample sequence into the decoder to obtain the tth hidden feature; Determine the tth conditional probability distribution result according to the tth hidden feature; Determine the tth predicted coordinate according to the tth conditional probability distribution result.

17. The device according to claim 16, wherein The input unit is further configured to: Determine a first conditional probability distribution and an attachment probability based on the t-th hidden feature, wherein the first conditional probability distribution is a probability distribution of candidate values ​​of the t-th predicted coordinate; Determining a second conditional probability distribution based on the t-th hidden feature and a t-1-th historical hidden feature sequence, where the t-1-th historical hidden feature sequence includes the t-1 hidden features output by the encoder that precede the t-th hidden feature; The second conditional probability distribution is a probability distribution of candidate attachment relationships of the t-th predicted coordinate, where the candidate attachment relationship is used to point to a face having an attachment relationship with a current face among the M faces, where the current face is the face corresponding to the t-th predicted coordinate among the M faces; The tth conditional probability distribution result is determined according to the first conditional probability distribution, the second conditional probability distribution, and the attachment probability.

18. The device according to claim 17, wherein The input unit is further configured to: Determine the candidate result corresponding to the maximum probability from the t-th conditional probability distribution result; The t-th predicted coordinate is obtained based on the candidate results.

19. The device according to claim 18, wherein In a case where the candidate result is a candidate value in the first conditional probability distribution, the t-th predicted coordinate is the candidate value; Alternatively, when the candidate result is a candidate dependency relationship in the second conditional probability distribution, the t-th predicted coordinate is the identifier of the face pointed to by the candidate dependency relationship.

20. The device according to any one of claims 15 to 19, wherein: The t-1th historical sample sequence includes t-1 transformation features, the sth transformation feature among the t-1 transformation features is obtained based on the sth prediction coordinate; the sth transformation feature includes at least one of the following: a surface numerical feature, a component position feature, and a surface position feature; Wherein, the face numerical feature is used to represent the coordinate value of the s-th predicted coordinate; The component position feature is used to represent the relative position of the first component where the s-th predicted coordinate is located in the sample object, where the first component is one of the multiple components included in the sample object; The surface position feature is used to represent the relative position of the surface corresponding to the s-th predicted coordinate in the first component.

21. The apparatus according to any one of claims 15 to 19, further comprising: a conversion unit, configured to obtain a plurality of line segments contained in the two-dimensional sample view; Converting the coordinates of the plurality of line segments into a line segment coordinate sequence; The line segment coordinate sequence is converted into a first sample sequence.

22. The device according to claim 21, wherein The conversion unit is also used for: sorting the plurality of line segments according to a first preset rule; The sorted coordinates of the plurality of line segments are converted into a line segment coordinate sequence.

23. The device according to claim 21, wherein The line segment coordinate sequence includes N line segment coordinates, the first sample sequence includes N first features, an r-th first feature among the N first features is obtained based on the r-th line segment coordinate among the N line segment coordinates, where N is a positive integer; the r-th first feature includes at least one of the following: a numerical feature, a view feature, a view position feature, a line segment position feature, and a visual feature; Wherein, the numerical feature is used to represent the coordinate value of the r-th line segment coordinate; The view feature is used to represent the two-dimensional sample view where the r-th line segment coordinate is located; The line segment position feature is used to represent the relative position of the first line segment where the r-th line segment coordinate is located in the two-dimensional sample view, where the first line segment is one of the multiple line segments; The view position feature is used to represent the relative position of the r-th line segment coordinate in the first line segment; The visual feature is used to indicate the visual state of the first line segment.

24. The apparatus according to any one of claims 15 to 19, further comprising: A sorting unit is used to determine a directed acyclic graph of the sample object, the directed acyclic graph including a dependency relationship set and a face set, the dependency relationship in the dependency relationship set is used to represent two faces with dependency relationships among M faces of the sample object, and the face set includes the coordinates of the M faces; according to a second preset rule, the directed acyclic graph is sorted to obtain an ordered graph of the sample object; and the ordered graph is converted into the annotated coordinate sequence, the annotated coordinate sequence including M annotated coordinates arranged according to a preset rule, and the annotated coordinates among the M annotated coordinates are the coordinates of a face among the M faces.

25. The apparatus according to claim 24, wherein In a case where the first surface corresponding to the marked coordinates is located in the dependency relationship set, the marked coordinates are identifiers of the second surface having a dependency relationship with the first surface; Alternatively, when the first surface corresponding to the marked coordinates is not located in the dependency relationship set, the marked coordinates are the coordinate values ​​of the first surface; The first surface and the second surface are different surfaces among the M surfaces.

26. The apparatus according to claim 24, wherein The sample object includes A components; The device further comprises: An acquisition unit is used to obtain a coordinate set of a component among the A components, wherein the coordinate set of the component includes the coordinates of B faces constituting the component; and obtain the face set based at least on the coordinate set of the component, wherein A and B are positive integers and the value of M is related to the values ​​of A and B.

27. A three-dimensional reconstruction device comprising: A first generating unit, configured to generate a first sequence according to a two-dimensional view of the object to be modeled; a prediction unit, configured to input the first sequence into a target three-dimensional prediction model to obtain a target coordinate sequence, wherein the target three-dimensional prediction model is trained according to the method according to any one of claims 1 to 12; The second generating unit is configured to generate a three-dimensional model of the object to be modeled according to the target coordinate sequence.

28. The apparatus according to claim 27, after generating the three-dimensional model of the object to be modeled according to the target coordinate sequence, the apparatus further comprises: An editing unit is used to modify the target coordinate sequence to edit the three-dimensional model.

29. An electronic device comprising: at least one processor; as well as a memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the method according to any one of claims 1 to 14.

30. A non-transitory computer-readable storage medium storing computer instructions, wherein: The computer instructions are used to cause the computer to execute the method according to any one of claims 1-14.

31. A computer program product comprising a computer program, which, when executed by a processor, implements the method according to any one of claims 1 to 14.

Citation Information

Patent Citations

  • Three-dimensional reconstruction method, device, computer equipment and storage medium

    CN112581597A

  • Three-dimensional model generation method and device based on single freehand sketch and electronic equipment

    CN113129447A