3D Model Reconstruction Method, Device, Equipment, and Storage Medium

By extracting visual features from two-dimensional images and using visual language models to identify and retrieve component models, the problem of inaccuracy of the three-dimensional model caused by ignoring the annotation layer in the prior art is solved, and a more accurate and complete three-dimensional model reconstruction is achieved.

CN119068126BActive Publication Date: 2025-08-01HANGZHOU QUNHE INFORMATION TECHNOLOGIES CO LTD
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202411568087.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-11-04
Publication Date
2025-08-01
Estimated Expiration
2044-11-04

AI Technical Summary

Technical Problem

In the prior art, when reconstructing a three-dimensional model, the labeled layer information in the CAD drawing is usually ignored, resulting in inaccurate and complete model reconstruction.

Method used

By extracting visual features from two-dimensional images, using visual language models to identify component information combined with prompt information, and retrieving and adjusting component models from the database, the three-dimensional model is reconstructed.

Benefits of technology

It improves the accuracy and flexibility of model reconstruction results, can effectively utilize labeled layer information, and improves the integrity and accuracy of the three-dimensional model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119068126B_ABST
    Figure CN119068126B_ABST
Patent Text Reader

Abstract

The present disclosure provides a three-dimensional model reconstruction method, apparatus, device, and storage medium, relating to the field of computer technologies, and particularly to the fields of three-dimensional model reconstruction and three-dimensional modeling technologies. The method includes: extracting visual features of an object to be reconstructed from a two-dimensional image; using a vision-language model to identify component information of the object based on the visual features and prompt information, where the prompt information is used to represent a three-dimensional reconstruction task that the vision-language model needs to perform; retrieving a component model of the object based on the component information of the object; and adjusting the component model of the object based on the component information of the object to reconstruct a three-dimensional model of the object.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure relates to the field of computer technology, and in particular to the fields of three-dimensional model reconstruction and three-dimensional modeling technology. Background Art

[0002] Drawings designed using computer-aided design software, such as CAD drawings, usually include a geometric layer and a dimensioning layer. The geometric layer contains the orthographic projections that describe the geometric information of the three-dimensional model, and the dimensioning layer may contain dimension markings, functional symbols, etc. During the process of reconstructing a three-dimensional model using the design drawings, usually only the geometric layer is mainly used, while the dimensioning information in the dimensioning layer is ignored. Summary of the Invention

[0003] This disclosure provides a three-dimensional model reconstruction method, apparatus, device, and storage medium to solve or alleviate one or more technical problems in the prior art.

[0004] In a first aspect, this disclosure provides a three-dimensional model reconstruction method, including:

[0005] Extracting visual features of an object to be reconstructed from a two-dimensional image; using a vision-language model to identify part information of the object based on the visual features and prompt information, where the prompt information is used to represent the three-dimensional reconstruction task that the vision-language model needs to perform; retrieving part models of the object based on the part information of the object; and adjusting the part models of the object based on the part information of the object to reconstruct the three-dimensional model of the object.

[0006] In a second aspect, this disclosure provides a three-dimensional model reconstruction apparatus, including:

[0007] An extraction module for extracting visual features of an object to be reconstructed from a two-dimensional image;

[0008] An identification module for using a vision-language model to identify part information of the object based on the visual features and prompt information, where the prompt information is used to represent the three-dimensional reconstruction task that the vision-language model needs to perform;

[0009] A retrieval module for retrieving part models of the object based on the part information of the object;

[0010] An adjustment module for adjusting part models of the object based on the part information of the object to reconstruct the three-dimensional model of the object.

[0011] In a third aspect, an electronic device is provided, including:

[0012] At least one processor; and

[0013] A memory communicatively connected to the at least one processor; where

[0014] The memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute any method in the embodiments of the present disclosure.

[0015] In a fourth aspect, a non-transitory computer-readable storage medium storing computer instructions is provided, wherein the computer instructions are used to cause the computer to execute any method in the embodiments of the present disclosure.

[0016] In a fifth aspect, a computer program product is provided, including a computer program which, when executed by a processor, implements any method in the embodiments of the present disclosure.

[0017] It should be understood that the content described in this part is not intended to identify the key or important features of the embodiments of the present disclosure, nor is it used to limit the scope of the present disclosure. Other features of the present disclosure will become easily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS

[0018] In the drawings, unless otherwise specified, the same reference numerals throughout the several views denote the same or similar components or elements. These drawings are not necessarily drawn to scale. It should be understood that these drawings depict only some embodiments provided in accordance with the present disclosure and should not be regarded as limiting the scope of the present disclosure.

[0019] Figure 1 is a schematic flowchart of a three-dimensional model reconstruction method according to an embodiment of the present disclosure;

[0020] Figure 2 is a schematic flowchart of a three-dimensional model reconstruction method according to another embodiment of the present disclosure;

[0021] Figure 3 is a schematic flowchart of a three-dimensional model reconstruction method according to another embodiment of the present disclosure;

[0022] Figure 4 is an example diagram of default values of component-specific parameters;

[0023] Figure 5 and Figure 6 is an example diagram after changing the number of compartments in component-specific parameters;

[0024] Figure 7 and Figure 8 is an example diagram after changing the grid width in component-specific parameters;

[0025] Figure 9 and Figure 10 is an example diagram after changing the position of the shutter baffle in component-specific parameters;

[0026] Figure 11 is a schematic flowchart of a three-dimensional model reconstruction method according to another embodiment of the present disclosure;

[0027] Figure 12 is a schematic diagram of an example of a model network structure;

[0028] Figure 13 is an example diagram of an input drawing;

[0029] Figure 14 is Figure 13 an example diagram of the reconstruction result of;

[0030] Figure 15 is an example diagram of an input drawing;

[0031] Figure 16 is Figure 15 an example diagram of the reconstruction result of;

[0032] Figure 17 is an example diagram expressing a shape program using a programming language;

[0033] Figure 18 is an example diagram of a dialogue of the model;

[0034] Figure 19 is a schematic structural diagram of a three-dimensional model reconstruction apparatus according to an embodiment of the present disclosure;

[0035] Figure 20 is a schematic structural diagram of a three-dimensional model reconstruction apparatus according to another embodiment of the present disclosure;

[0036] Figure 21 is a block diagram of an electronic device for implementing the embodiments of the present disclosure. Detailed implementation manners

[0037] The present disclosure will be further described in detail below with reference to the accompanying drawings. The same reference numerals in the drawings denote elements having the same or similar functions. Although various aspects of the embodiments are shown in the drawings, the drawings do not have to be drawn to scale unless otherwise specified.

[0038] In addition, for better explaining the present disclosure, numerous specific details are given in the following detailed implementation manners. Those skilled in the art should understand that the present disclosure can be implemented without some of these specific details. In some instances, methods, means, elements, and circuits well known to those skilled in the art are not described in detail so as to highlight the gist of the present disclosure.

[0039] Figure 1 is a schematic flowchart of a three-dimensional model reconstruction method according to an embodiment of the present disclosure, and the method may include:

[0040] S110. Extract the visual features of the object to be reconstructed from the two-dimensional image;

[0041] S120. Use a vision-language model to identify the component information of the object based on the visual features and the prompt information, so as to reconstruct the three-dimensional model of the object; wherein, the prompt information is used to represent the three-dimensional reconstruction task that the vision-language model needs to perform;

[0042] S130. Retrieve the component model of the object based on the component information of the object;

[0043] S140. Adjust the component model of the object based on the component information of the object, so as to reconstruct the three-dimensional model of the object.

[0044] In the embodiments of the present disclosure, the two-dimensional image may be a planar image composed of two-dimensional elements and not containing depth information, and may also be referred to as a two-dimensional picture, a planar picture, etc. The two-dimensional image may include, for example, bitmaps such as photos and pictures, and may also include vector graphics such as CAD planar drawings. The visual features may be extracted based on the geometric layer and / or the annotation layer of the two-dimensional image, rather than only based on the geometric layer. There are various ways to extract image features. In one example, through image extraction techniques such as OCR technology, the visual features of the object to be modeled can be extracted from the geometric layer and / or the annotation layer of the two-dimensional image. In another example, through an artificial intelligence model such as a vision-language model, the visual features can be extracted from the image. For example, a two-dimensional CAD image contains the geometric layer and the annotation layer of a cabinet. From this two-dimensional CAD drawing, not only the visual features corresponding to the geometric layer of the cabinet can be extracted, but also the visual features corresponding to the annotation layer of the cabinet can be extracted.

[0045] In the embodiments of the present disclosure, the vision-language model can understand and interpret the association between images and texts, and generate accurate and vivid natural language descriptions according to the images. The vision-language model may include a visual encoder, a multilayer perceptron (MLP), and a language model part, etc. For example, a vision-language model is a multimodal large model. The multimodal large model may include a visual encoder using a vision transformer (ViT), an MLP, and a large language model. For example, the multimodal large model may be a Mini-InternVL-1.5-2B model, which may include a visual encoder of InternViT-300M, an MLP, and a language model of InternLM2-1.8B. The encoder, MLP, language model, etc. in the above vision-language model may also be of other types, which are not limited in the embodiments of the present disclosure. The prompt information may include the content input to the model for describing the model task. The prompt information may also be referred to as a prompt word, a prompt word template, a prompt, etc.

[0046] In the embodiments of the present disclosure, the visual features extracted from the two-dimensional image and the hint information used to describe the three-dimensional reconstruction task can be input into the vision-language model, and the part information of each object in the two-dimensional image can be obtained through the processing of the vision-language model. An object may include one or more parts. For example, a cabinet may include part information such as a cabinet body and a door. The three-dimensional model of the object can be reconstructed using the part information of a certain object. For example, from a two-dimensional image of a bicycle, the information of multiple parts of the target object "bicycle" can be extracted, which are "body, handlebar, seat, pedal and two wheels" respectively. Based on the information of multiple parts of the "bicycle", the three-dimensional models of each part can be constructed respectively, and then the three-dimensional model of the whole object "bicycle" can be formed.

[0047] In the embodiments of the present disclosure, the part models of the object to be reconstructed can be retrieved from the database according to the part information obtained by the vision-language model. The parameters of the part model retrieved from the database are adjusted according to the identified part information, and the three-dimensional model of the object is established based on the adjusted part model.

[0048] According to the embodiments of the present disclosure, by combining the vision-language model and the hint information to reconstruct the model of the two-dimensional image, the flexibility of the input format is improved, and the accuracy of the model reconstruction result is enhanced.

[0049] Figure 2 FIG. is a schematic flowchart of a three-dimensional model reconstruction method according to another embodiment of the present disclosure, and this method may include one or more features of the above three-dimensional model reconstruction method. In one implementation, step S110 extracts the visual features of the object to be reconstructed from the two-dimensional image, including:

[0050] S210, using the vision-language model to divide the two-dimensional image into multiple image patches;

[0051] S220, inputting the multiple image patches into the visual encoder and the multi-layer perceptron MLP of the vision-language model to obtain the visual features of the object; wherein, the visual features include the vector representation mapped based on the multiple image patches.

[0052] In the embodiments of the present disclosure, a two-dimensional image is input into a vision-language model, and the vision-language model divides the two-dimensional image into multiple image patches. For example, when image A is input into the vision-language model, image A can be divided into multiple image patches such as A1, A2, A3, etc. Each image patch can represent a local area of the two-dimensional image. There are various ways to extract image patches, such as the sliding window method, the method based on convolutional operations, and the grid method. Among them, the sliding window method can slide a window of a fixed size on the original image, with a certain step size for each slide, so as to extract the local area of the image as an image patch. The method based on convolutional operations can use convolutional operations to extract image patches, and the convolutional kernel can slide on the image and perform convolutional operations. Convolutional operations can not only extract local features of the image, but also control the size and position of the extracted image patches by adjusting the size and step size of the convolutional kernel. The grid method can evenly divide the image into multiple image patches according to the number of rows and columns of the grid.

[0053] In the embodiments of the present disclosure, the multiple segmented image patches are input into the vision encoder part of the vision-language model, image features can be extracted, a multi-layer perceptron (MLP) is used to align the image features with text features, and the image features are processed to obtain the visual features of each image patch, and then the visual features of the object are obtained. The visual features can include visual tokens. The visual tokens can be vector representations obtained by mapping each of the image patches into which the image is segmented, and these vector representations can be regarded as basic processing units in the model.

[0054] According to the embodiments of the present disclosure, using a vision-language model to segment a two-dimensional image and then extracting visual features through a vision encoder and an MLP can improve the accuracy of the visual feature extraction results.

[0055] Figure 3 It is a schematic flowchart of a three-dimensional model reconstruction method according to another embodiment of the present disclosure, and this method can include one or more features of the above three-dimensional model reconstruction method. In one implementation, in step S120, the vision-language model is used to identify the component information of the object based on the visual features and the prompt information, including:

[0056] S310. Input the visual features and the prompt information into the language model part of the vision-language model to identify the identifiers and parameters of one or more components of the object;

[0057] In step S130, the component model of the object is retrieved based on the component information of the object, including:

[0058] S320. Retrieve one or more component models from the database according to the identities of the one or more components;

[0059] Step S140 adjusts the component model of the object based on the component information of the object to reconstruct the three-dimensional model of the object, including:

[0060] S330. Adjust the parameters of the corresponding component model according to the parameters of the one or more components to reconstruct the three-dimensional model of the object.

[0061] In the embodiments of the present disclosure, the image features and prompt information extracted by the visual encoder part of the vision-language model are input into the language model part of the vision-language model. After being processed by the language model part, the identities (Identity document, ID) and parameters of one or more components included in the object in the two-dimensional image can be obtained. The identity of a component can be a number used to uniquely identify the component, and this number can be composed of numbers, characters, etc. The identity of the component can be used as an index in the database, and based on the ID of the component, the pre-stored component model in the database that matches this ID can be found. The pre-stored component model can have some default parameters. According to the identified parameters of the component, the default parameters of the component model corresponding to this component can be adjusted.

[0062] In the embodiments of the present disclosure, the default parameters of different component models are usually not exactly the same. For example, component models can be constructed according to the type of item. For example, component models of cabinets, doors, drawers, fixed plate components, movable plate components, etc. can be constructed respectively. Another example is that component models can be constructed according to the composition characteristics of the components. For example, a component model of a cabinet including one grid is constructed. Another example is that a component model of a cabinet including two grids on the top and one grid on the bottom is constructed.

[0063] In the embodiments of the present disclosure, according to the parameters of the components identified by the language model part, the posture, size, position, etc. of the component model can also be adjusted, so that the multiple component models of a certain object are combined into the three-dimensional model of the object.

[0064] According to the embodiments of the present disclosure, by obtaining the identities and parameters of each component of the object through the language model part of the vision-language model, the corresponding component models can be retrieved from the database and the parameters of the component models can be adjusted, thereby completing the reconstruction of the three-dimensional model, quickly and accurately finding the component models of the object, and further improving the accuracy of the model reconstruction result.

[0065] In one implementation manner, step S330 adjusts the parameters of the corresponding component model according to the parameters of the one or more components to reconstruct the three-dimensional model of the object, including:

[0066] Adjust the pose parameters and specific parameters of the corresponding component model according to the parameters of the one or more components to obtain the three-dimensional model of the object; wherein, the pose parameters include component position, component size, and component rotation angle.

[0067] In the embodiments of the present disclosure, the parameters of the components identified by the language model part may include the pose parameters of the components. For example, the pose parameters of the components include component position, component size, component rotation angle, etc. The component size can also be a specific parameter of the component. The component models in the database may not have pose parameters or may be set with default pose parameters. The pose parameters of a certain component identified can be used to adjust the pose parameters of the component model corresponding to the component. For example, based on the position of a component of an object in the original two-dimensional image, the three-dimensional coordinates (x, y, z) of the component can be obtained. For another example, the component size of a certain object identified from the original two-dimensional image may include one or more of the length, width, height, etc. of the component of the object. The size of the component model of the component can be adjusted according to one or more of these sizes. For another example, the component rotation angle of a certain object identified from the original two-dimensional image can be the angle between a certain line on the component of the object and a certain reference line or reference plane. For example, the angle between the axis of symmetry of the component and the line representing the ground. Based on the identified component rotation angle of the component, the angle between the component model corresponding to the component and the plane representing the ground can be adjusted.

[0068] For example, if the length of a certain component M1 identified from the original two-dimensional image is 30 cm and the length of the component model m1 corresponding to the component M1 retrieved from the database is 5 cm, then the length of the component model m1 can be adjusted to 30 cm, and other sizes can be adjusted according to the recognition result. For another example, if the rotation angle of a certain component M2 identified from the original two-dimensional image is 90°, and the rotation angle of the component model m2 corresponding to the component M2 retrieved from the database is 180°, then the rotation angle of the component model m2 can be adjusted to 90°.

[0069] In the embodiments of the present disclosure, the default parameters of the component models in the database may include the specific parameters of the component models. The specific parameters of the component can be the parameters unique to the component, and the specific parameters of different components can be completely different or partially different. For example, an example of the specific parameters of a cabinet component model is as Figure 4 shown. The specific parameters of the cabinet component model include N = 2, NKA = 873, NKB = 873, DBXX = 1, where N represents the number of cabinet compartments, NKA and NKB represent the width of each compartment, and DBXX represents the position of the door baffle. Figures 5 to 10 is an example after changing the default specific parameters of the above component model according to the recognition result. For another example, in the component model of a combined drawer, the specific parameters may include the number of drawers, the width and length of the drawers, etc.

[0070] The default parameters of the component model of a component can be adjusted according to the specific parameters of the component recognized by the language model part. For example, if the component model of a component in the database includes two grids, and the default length, width, and height of each grid are 40 cm, 20 cm, and 60 cm respectively. If it is recognized from the two-dimensional image that the cabinet body includes 3 grids, and the length, width, and height of grid A are 50 cm, 24 cm, and 80 cm respectively, then the length of grid A in the component model can be adjusted to 50 cm, the width can be adjusted to 24 cm, and the height can be adjusted to 80 cm. If the parameters of the other two grids B and C recognized from the two-dimensional image only include a length of 100 cm and do not include the width or height, the lengths of grids B and C can be adjusted to 100 cm, the width can be adjusted to 24 cm according to the first grid A, and the height can be adjusted to 80 cm according to the first grid A.

[0071] If an object includes multiple components, after adjusting the pose parameters and specific parameters of the component models corresponding to the respective components of the object according to the recognition result, a three-dimensional model composed of the respective components can be obtained.

[0072] According to an embodiment of the present disclosure, by adjusting the parameters of the component models retrieved from the database according to the parameters of the respective components of the object in the original two-dimensional image, a three-dimensional model that conforms to the shape, pose, position, etc. of the object in the two-dimensional image can be obtained, improving the accuracy of the model reconstruction result.

[0073] Figure 11 It is a schematic flowchart of a three-dimensional model reconstruction method according to another embodiment of the present disclosure, and this method may include one or more features of the above three-dimensional model reconstruction method. In one implementation, this method further includes:

[0074] S510. Establish a parameterized component model in the database in advance; wherein, the component model includes the identifier of the component and its corresponding specific parameters.

[0075] In an embodiment of the present disclosure, each component model can be parameterized to represent the component model in a mathematical language description. The component model may include the identifier of the parameterized component model in the database and the specific parameters corresponding to the parameterized component model. Further, the component model can be represented in the form of code or data serialization. For example, the form of code may include code formats such as Python code format and Java code format, and the data serialization format may include data serialization forms such as Json format and YAML format.

[0076] For example, a parameterized component model can be represented as where is the unique ID of the parametric component model in the database, is the pose parameter of the component model, which can specifically include the center point position , size and rotation angle etc. is a set of parameters specific to the component model. The formula can represent a model (such as a 3D model) assembled from N parametric component models.

[0077] Through the embodiments of the present disclosure, by pre-modeling the components, the components can be represented parametrically, which can improve the efficiency of obtaining 3D models.

[0078] In one implementation, the identifier of the component has an association relationship with the token of the component model; wherein, the feature of the token of the component model is obtained by splicing the text feature of the component name and the image feature of the component preview image. In one way, the vision-language model can look up the corresponding token based on the identifier of the component, and then look up the component model corresponding to the token. In another way, the vision-language model can identify the token from the 2D model, and then look up the identifier or component model corresponding to the token, etc.

[0079] In the embodiments of the present disclosure, the token of the component model can be used in the vision-language model to represent the component, and the token of the component model and the identifier of the component can have an association relationship. For example, the component name and component preview image of each component model can be set in advance, and the token feature of the component model can be generated based on the text feature of the component name and the image feature of the component preview image. For example, the text feature can be extracted from the component name through a multimodal model such as the CLIP (Contrastive Language-Image Pre-Training) model, the image feature can be extracted from the component preview image, and the text feature and the image feature are spliced in the feature dimension to obtain the feature of the token of the component model.

[0080] According to the embodiments of the present disclosure, the token of the component model can help the vision-language model better understand the component information, improve the accuracy of the component model extracted from the database, and also improve the efficiency of extracting the component model.

[0081] In one implementation, the vision-language model is obtained by fine-tuning a benchmark large model based on training data; wherein, the training data includes sample images and their corresponding parametric component models.

[0082] In the embodiments of the present disclosure, the sample image may include the feature vector of the image, and the parameterized component model may be in the format of code. The training data may further include the prompt information for the model to perform the 3D reconstruction task. After inputting the training data into the baseline large model, the parameters of the baseline large model can be supervised and fine-tuned according to the output result of the baseline large model. After the output result of the baseline large model meets the training objective, the training can be stopped to obtain the trained vision language model.

[0083] According to the embodiments of the present disclosure, using the supervised fine-tuning method to fine-tune the large model to obtain the vision language model can improve the accuracy of the output result of the vision language model, and further improve the accuracy of the model reconstruction result.

[0084] In one implementation, the two-dimensional image includes a geometry layer and / or an annotation layer. The geometry layer may include the orthographic projection for describing the geometric information of the 3D model, and the annotation layer may include dimension annotations, functional symbols, etc. Examples of functional symbols may include surface types, manufacturing instructions, etc.

[0085] According to the embodiments of the present disclosure, by obtaining the information of the geometry layer and the annotation layer in the two-dimensional image, the accuracy of generating the component model can be improved, and further the accuracy of the model reconstruction result can be improved.

[0086] In one implementation, the two-dimensional image is a bitmap or a vector graph.

[0087] In the embodiments of the present disclosure, the two-dimensional image can have various formats, such as a bitmap composed of pixels, a vector graph composed of vector graphic elements. The bitmap can also be called a dot matrix image or a raster image, etc., and can be composed of multiple pixel points. The vector graph can also be called an object-oriented image or a drawing image, etc. The graphic elements in the vector graph can have attributes such as color, shape, contour, size, and screen position. The vector graph can be drawn by some software, such as AutoCAD, etc.

[0088] According to the embodiments of the present disclosure, two-dimensional images in different formats can be used as the input images of the vision language model, which can expand the application scope of the model reconstruction scenario.

[0089] The embodiments of the present disclosure propose a 3D model reconstruction method, and a schematic diagram of an example of a model network structure is as Figure 12 shown.

[0090] The embodiments of the present disclosure can model the reconstruction problem of the 3D model as a visual question answering problem and implement it by fine-tuning the Vision Language Model.

[0091] In an application scenario, the input of the model can be a 2D image such as a complete CAD drawing (including geometric layers and annotation layers) and a prompt question, or a bitmap in the format of GIF or JPEG and a prompt question. The output of the model can include natural language (such as program code) for describing the 3D model.

[0092] An example of a method for reconstructing a 3D customized cabinet model from a CAD drawing is as follows:

[0093] 1. Algorithm input

[0094] The input of the algorithm can be a complete CAD drawing, which includes geometric layers and annotation layers, as shown in Figure 13 and Figure 15 .

[0095] 2. Algorithm output

[0096] In actual modeling, using a set of pre-modeled parametric component models can speed up the design process, reduce design errors, and ensure that the product can be correctly manufactured. For example, model the components of a customized cabinet. For example, these component models can include the basic cabinet body, doors, drawers, fixed plates, movable plates, etc. Each component model can be pre-modeled as a parametric model.

[0097] The customized cabinet can be assembled from a set of pre-modeled component models, for example, represented by the formula . Among them, each component model can be characterized as , where is the unique ID of the parametric component model in the database, is the pose parameter, including the center point position , dimensions and the rotation angle , is a set of specific parameters of the component models. If the component model has no parameters, .

[0098] An example of the specific parameters of a cabinet component model can be seen in Figures 4 to 10 . Figure 4 Among them, the default values of the specific parameters of the cabinet component model can include: the number of cabinet grids N = 2, the width of each grid NKA = 873 and NKB = 873, and the position of the flush door baffle DBXX = 1. Figure 5 is an example where the number of cabinet grids is changed to N = 1. Figure 6 is an example where the number of cabinet grids is changed to N = 3. Figure 7 is an example where the width of the cabinet grids is changed to NKA = 1100 and NKB = 646. Figure 8It is an example where the width of the cabinet grid is changed to NKA = 646 and NKB = 1100. Figure 9 It is an example where the position of the cabinet door baffle is changed to DBXX = 2. Figure 10 It is an example where the position of the cabinet door baffle is changed to DBXX = 3.

[0099] The part ID is a randomly generated number without actual semantics. To help the neural network model better understand part information, a set of additional special tokens can be used to represent parts. For example, tokens are obtained through two features of each part: (1) part name and (2) part preview image. Among them, the part preview image can use default parameters and is rendered through a fixed camera view. The CLIP model can extract the text features from the part name and the image features from the part preview image, and splice these two sets of features in the feature dimension to serve as the token of each part.

[0100] Since the Large Language Model (LLM) has good code generation ability, the 3D model can be represented in the format of code. For example, using programming languages such as Python as the proxy language, Figure 17 Shows an example of a shape program expressed in the Python programming language. Among them, bbox_0 to bbox_5 represent the pose parameters of multiple parts included in an object, and model_0 to model_5 represent the IDs of these parts and their corresponding specific parameters. According to the example of the shape program, the preview image of the reconstruction result can be run and displayed in the corresponding software tool. See Figure 14 and Figure 16 .

[0101] 3. Algorithm Model

[0102] The algorithm model used can be a vision-language model, such as the Mini-InternVL-1.5-2B model. The vision-language model can include a vision encoder such as InternViT-300M and a language model such as InternLM2-1.8B. Among them, the InternViT-300M model can be used to extract image features. Figure 18 Provides an example diagram of a complete conversation. Among them, the example of the prompt information input to the language model can include "Reconstruct cabinet from image. The bounding box reference code is as follows: ……", and the output can include a shape program describing the cabinet expressed in Python. Among them, Represents image features, which can be part of the prompt information and input into the model together.

[0103] Use supervised fine-tuning to fine-tune the Mini-InternVL-1.5-2B model. For the training samples of the model, refer to Figure 18 the examples in. During inference, the drawing and prompt input can be provided, and the algorithm will automatically generate the program code describing an object such as the cabinet structure.

[0104] Figure 19 is a schematic flowchart of a three-dimensional model reconstruction device 1000 according to an embodiment of the present disclosure. The device may include:

[0105] An extraction module 1010, configured to extract the visual features of the object to be reconstructed from the two-dimensional image;

[0106] An identification module 1020, configured to use a vision-language model to identify the component information of the object based on the visual features and prompt information, so as to reconstruct the three-dimensional model of the object; wherein, the prompt information is used to represent the three-dimensional reconstruction task that the vision-language model needs to execute;

[0107] An access module 1030, configured to access the component model of the object based on the component information of the object;

[0108] An adjustment module 1040, configured to adjust the component model of the object based on the component information of the object, so as to reconstruct the three-dimensional model of the object.

[0109] Figure 20 is a schematic flowchart of a three-dimensional model reconstruction device 1100 according to another embodiment of the present disclosure. The device may include one or more features of the above three-dimensional model reconstruction device. In one implementation, the extraction module 1010 includes:

[0110] A segmentation sub-module 1111, configured to use the vision-language model to segment the two-dimensional image into multiple image blocks;

[0111] A visual feature sub-module 1112, configured to input the multiple image blocks into the visual encoder and multi-layer perceptron MLP of the vision-language model to obtain the visual features of the object; wherein, the visual features include the vector representation mapped based on the multiple image blocks.

[0112] In one implementation, the identification module 1020 is configured to input the visual features and the prompt information into the language model part of the vision-language model to obtain the identifiers and parameters of one or more components of the object;

[0113] Retrieval module 1030, configured to retrieve one or more component models from a database according to the identifiers of the one or more components;

[0114] Adjustment module 1040, configured to adjust the parameters of the corresponding component model according to the parameters of the one or more components, so as to reconstruct the three-dimensional model of the object.

[0115] In one embodiment, the adjustment module 1040 is configured to adjust the pose parameters and specific parameters of the corresponding component model according to the parameters of the one or more components, so as to obtain the three-dimensional model of the object; wherein, the pose parameters include component position, component size, and component rotation angle.

[0116] In one embodiment, the device further includes:

[0117] Modeling module 1130, configured to pre-establish a parameterized component model in a database; wherein, the component model includes the identifier of the component and its corresponding specific parameters.

[0118] In one embodiment, the identifier of the component has an associated relationship with the token of the component model; wherein, the features of the token of the component model are obtained by splicing the text features of the component name and the image features of the component preview image.

[0119] In one embodiment, the vision-language model is obtained by fine-tuning a benchmark large model based on training data; wherein, the training data includes sample images and their corresponding parameterized component models.

[0120] In one embodiment, the two-dimensional image includes a geometric layer and / or an annotation layer.

[0121] In one embodiment, the two-dimensional image is a bitmap or a vector graph.

[0122] Figure 21 It is a structural block diagram of an electronic device according to an embodiment of the present disclosure. As Figure 21 shown, the electronic device includes: a memory 1210 and a processor 1220. The memory 1210 stores a computer program that can run on the processor 1220. The number of the memory 1210 and the processor 1220 can be one or more. The memory 1210 can store one or more computer programs. When the one or more computer programs are executed by the electronic device, the electronic device executes the method provided in the above method embodiment. The electronic device may further include: a communication interface 1230, configured to communicate with external devices and perform data interaction and transmission.

[0123] If the memory 1210, the processor 1220, and the communication interface 1230 are implemented independently, the memory 1210, the processor 1220, and the communication interface 1230 can be interconnected through a bus and communicate with each other. The bus can be an Industry Standard Architecture (ISA) bus, a Peripheral Component Interconnect (PCI) bus, an Extended Industry Standard Architecture (EISA) bus, or the like. The bus can be divided into an address bus, a data bus, a control bus, etc. For the sake of simplicity of representation, Figure 21 it is represented by only one thick line in Figure 21 , but it does not mean that there is only one bus or one type of bus.

[0124] Optionally, in specific implementation, if the memory 1210, the processor 1220, and the communication interface 1230 are integrated on a single chip, the memory 1210, the processor 1220, and the communication interface 1230 can communicate with each other through an internal interface.

[0125] It should be understood that the above-mentioned processor can be a Central Processing Unit (CPU), or can also be other general-purpose processors, Digital Signal Processors (DSPs), Application Specific Integrated Circuits (ASICs), Field Programmable Gate Arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor can be a microprocessor or any conventional processor, etc. It is worth noting that the processor can be a processor that supports the Advanced RISC Machines (ARM) architecture.

[0126] Further, optionally, the above-mentioned memory may include a read-only memory and a random access memory, and may further include a non-volatile random access memory. The memory may be a volatile memory or a non-volatile memory, or may include both volatile and non-volatile memories. Among them, the non-volatile memory may include a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), or a flash memory. The volatile memory may include a random access memory (RAM), which is used as an external cache. By way of example but not limitation, many forms of RAM are available. For example, static random access memory (SRAM), dynamic random access memory (DRAM), synchronous dynamic random access memory (SDRAM), double data rate synchronous dynamic random access memory (DDR SDRAM), enhanced synchronous dynamic random access memory (ESDRAM), synchlink dynamic random access memory (SLDRAM), and direct rambus random access memory (DR RAM).

[0127] In the above embodiments, it can be implemented in whole or in part by software, hardware, firmware, or any combination thereof. When implemented using software, it can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer instructions are loaded and executed on a computer, the processes or functions according to the embodiments of the present disclosure are generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions can be stored in a computer-readable storage medium, or transmitted from one computer-readable storage medium to another computer-readable storage medium. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center by wire (e.g., coaxial cable, optical fiber, Digital Subscriber Line (DSL)) or wirelessly (e.g., infrared, Bluetooth, microwave, etc.). The computer-readable storage medium can be any available medium that can be accessed by a computer, or a data storage device such as a server or data center that includes one or more integrated available media. The available medium can be a magnetic medium (e.g., floppy disk, hard disk, magnetic tape), an optical medium (e.g., Digital Versatile Disc (DVD)), or a semiconductor medium (e.g., Solid State Disk (SSD)), etc. It should be noted that the computer-readable storage medium mentioned in the present disclosure can be a non-volatile storage medium, in other words, it can be a non-transitory storage medium.

[0128] Those of ordinary skill in the art can understand that all or part of the steps to implement the above embodiments can be completed by hardware, or can be completed by a program instructing relevant hardware. The program can be stored in a computer-readable storage medium, and the storage medium mentioned above can be a read-only memory, a magnetic disk, an optical disc, etc.

[0129] In the description of the embodiments of the present disclosure, the descriptions referring to terms such as "one embodiment", "some embodiments", "example", "specific example", or "some examples" mean that the specific features, structures, materials, or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present disclosure. Moreover, the specific features, structures, materials, or characteristics described can be combined in a suitable manner in any one or more embodiments or examples. In addition, without contradiction, those skilled in the art can combine and combine the different embodiments or examples described in this specification and the features of the different embodiments or examples.

[0130] In the description of the embodiments of the present disclosure, unless otherwise specified, " / " means "or". For example, A / B may mean A or B. "And / or" herein is merely a description of the association relationship of associated objects, indicating that there can be three relationships. For example, A and / or B may mean: A exists alone, A and B exist simultaneously, and B exists alone.

[0131] In the description of the embodiments of the present disclosure, the terms "first" and "second" are only used for descriptive purposes and cannot be construed as indicating or implying relative importance or implicitly specifying the quantity of the indicated technical features. Thus, the features defined with "first" and "second" may explicitly or implicitly include one or more of such features. In the description of the embodiments of the present disclosure, unless otherwise specified, "a plurality of" means two or more.

[0132] The foregoing are only exemplary embodiments of the present disclosure and are not intended to limit the present disclosure. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present disclosure shall be included within the protection scope of the present disclosure.

Claims

1. A three-dimensional model reconstruction method, comprising: Pre-establishing a parameterized component model in a database; wherein, the component model includes an identifier of the component and its corresponding specific parameters; Extracting visual features of an object to be reconstructed from a two-dimensional image; Using a vision-language model to identify component information of the object based on the visual features and prompt information; wherein, the prompt information is used to represent a three-dimensional reconstruction task that the vision-language model needs to perform; Searching in the database based on the identifier of the component in the component information of the object and retrieving the component model of the object that matches the identifier of the component; Adjusting the component model of the object based on the parameters of the component in the component information of the object to reconstruct the three-dimensional model of the object; Wherein, using a vision-language model to identify component information of the object based on the visual features and prompt information includes: inputting the visual features and the prompt information into the language model part of the vision-language model to identify the identifiers and parameters of one or more components of the object; Retrieving the component model of the object based on the identifier of the component in the component information of the object includes: retrieving one or more component models from the database according to the identifier of the one or more components; Adjusting the component model of the object based on the parameters of the component in the component information of the object to reconstruct the three-dimensional model of the object includes: adjusting the parameters of the corresponding component model according to the parameters of the one or more components to reconstruct the three-dimensional model of the object; wherein, adjusting the parameters of the corresponding component model according to the parameters of the one or more components to reconstruct the three-dimensional model of the object includes: adjusting the pose parameters and specific parameters of the corresponding component model according to the parameters of the one or more components to obtain the three-dimensional model of the object.

2. The method according to claim 1, wherein, Extracting visual features of an object to be reconstructed from a two-dimensional image includes: Using the vision-language model to segment the two-dimensional image into multiple image patches; Inputting the multiple image patches into the vision encoder and multi-layer perceptron MLP of the vision-language model to obtain the visual features of the object; wherein, the visual features include vector representations mapped based on the multiple image patches.

3. The method according to claim 1, wherein The pose parameters include component position, component size, and component rotation angle.

4. The method according to claim 1, wherein, The identifier of the component has an associated relationship with the token of the component model; wherein, the features of the token of the component model are obtained by splicing the text features of the component name and the image features of the component preview image.

5. The method according to any one of claims 1 to 4, wherein The vision-language model is obtained by fine-tuning a baseline large model based on training data; wherein, the training data includes sample images and their corresponding parameterized component models.

6. The method according to any one of claims 1 to 4, wherein The two-dimensional image includes a geometry layer and / or an annotation layer.

7. The method according to any one of claims 1 to 4, wherein The two-dimensional image is a bitmap or a vector map.

8. A three-dimensional model reconstruction apparatus, comprising: A modeling module for pre-establishing a parameterized component model in a database; wherein, the component model includes an identifier of the component and its corresponding specific parameters; An extraction module for extracting visual features of an object to be reconstructed from a two-dimensional image; An identification module, configured to use a vision-language model to identify component information of the object based on the visual features and prompt information; wherein, the prompt information is used to represent a 3D reconstruction task that the vision-language model needs to perform; An extraction module, configured to search for and extract a component model of the object that matches the identifier of the component in the database based on the identifier of the component in the component information of the object; An adjustment module, configured to adjust the component model of the object based on the parameters of the component in the component information of the object to reconstruct a 3D model of the object; Wherein, the identification module is configured to input the visual features and the prompt information into a language model part of the vision-language model to obtain identifiers and parameters of one or more components of the object; the extraction module is configured to extract one or more component models from the database according to the identifiers of the one or more components; the adjustment module is configured to adjust parameters of the corresponding component model according to the parameters of the one or more components to reconstruct a 3D model of the object; Wherein, the adjustment module is configured to adjust pose parameters and specific parameters of the corresponding component model according to the parameters of the one or more components to obtain a 3D model of the object.

9. The device according to claim 8, wherein, The extraction module includes: A segmentation sub-module, configured to use the vision-language model to segment the 2D image into multiple image patches; A visual feature sub-module, configured to input the multiple image patches into a visual encoder and a multi-layer perceptron (MLP) of the vision-language model to obtain visual features of the object; wherein, the visual features include vector representations mapped based on the multiple image patches.

10. The device according to claim 8, wherein, The pose parameters include component position, component size, and component rotation angle.

11. The apparatus according to claim 8, wherein, There is an association relationship between the identifier of the component and the token of the component model; wherein, the features of the token of the component model are obtained by splicing text features of the component name and image features of the component preview image.

12. The device according to any one of claims 10 to 11, wherein, The vision-language model is obtained by fine-tuning a benchmark large model based on training data; wherein, the training data includes sample images and their corresponding parameterized component models.

13. The apparatus according to any one of claims 10 to 11, wherein, The 2D image includes a geometry layer and / or an annotation layer.

14. The device according to any one of claims 10 to 11, wherein The 2D image is a bitmap or a vector map.

15. An electronic device, comprising: At least one processor; And A memory communicatively connected to the at least one processor; wherein, The memory stores instructions executable by the at least one processor, and when the instructions are executed by the at least one processor, the at least one processor is enabled to execute the method according to any one of claims 1-7.

16. A non-transitory computer-readable storage medium storing computer instructions, wherein, The computer instructions are used to cause the computer to execute the method according to any one of claims 1-7.

17. A computer program product, comprising a computer program, where the computer program, when executed by a processor, implements the method according to any one of claims 1-7.

Citation Information

Patent Citations

  • Three-dimensional reconstruction method and device and computer storage medium

    CN115587928A

  • Three-dimensional reconstruction model training method, three-dimensional reconstruction method and device

    CN117237538A

  • Model training method and device, storage medium and electronic equipment

    CN117635822A

  • Drawing processing method and device, equipment and storage medium

    CN118115692A

  • Combined three-dimensional representation and scene reconstruction method and device for large language model

    CN118365796A