A 3D model reconstruction method and device
By using grid construction models and ARCore technology in three-dimensional scanning technology to uniformly process the three-dimensional model structure and color, the problem of low structure and color matching in the existing technology is solved, and the overall quality of the three-dimensional model is improved.
Patent Information
- Application Number
- CN202211721385.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-12-30
- Publication Date
- 2025-06-17
- Estimated Expiration
- 2042-12-30
AI Technical Summary
In the existing three-dimensional scanning technology, the structural reconstruction and color reconstruction of the three-dimensional model are usually two independent steps, resulting in a low degree of matching between the final reconstruction of the three-dimensional model structure and the color, and a large difference in reality.
By obtaining image data under multiple angles and multiple time series, building a three-dimensional model structure based on the grid construction model, and then converting image data into 3D point cloud information with colors through ARCore, matching point clouds and grid data, realizing unified and coordinated processing of structure and color.
The unified and coordinated processing of structure and color during the construction of three-dimensional model is realized, reducing the problem of mismatch between structure and color and low authenticity, and improving the overall quality of the three-dimensional model.
Smart Images

Figure CN115830244B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of three-dimensional scanning technology, and provides a three-dimensional model reconstruction method and device. Background Art
[0002] Three-dimensional scanning refers to a high-tech that integrates light, machinery, electricity, and computer technologies. It is mainly used to collect and analyze the geometric structure and appearance data of an object or environment, and perform three-dimensional reconstruction on the collected data to obtain a three-dimensional digital model of the scanned object. When using a three-dimensional scanner to scan an impression, impression data is obtained from multiple angles respectively. Using the high-precision impression data obtained by the three-dimensional scanner, three-dimensional reconstruction is performed to finally obtain a three-dimensional model.
[0003] Regarding the final effect of three-dimensional scanning, including the overall structure of the three-dimensional model and the color attached to the three-dimensional surface, in the prior art, the establishment of the overall structural characteristics of three-dimensional scanning and the color information on the three-dimensional surface are two independent steps, that is, they are independently performed through a surface reconstruction algorithm and a color reconstruction algorithm respectively. Through these two independent processing processes, although the reconstruction of the final three-dimensional model can be achieved, due to the independent steps, the independent algorithms will lead to a low matching degree between the structure and color of the finally reconstructed three-dimensional model, and a large difference from the authenticity of the environment. Summary of the Invention
[0004] To solve the above technical problems, this application provides a method and device that can link and unify structure reconstruction and color reconstruction, realize the fusion of the two, and improve the unity and coordination of the three-dimensional model construction process.
[0005] To achieve the above object, the technical solutions adopted in the embodiments of this application are as follows:
[0006] In a first aspect, a three-dimensional model reconstruction method is provided. The method includes: obtaining a plurality of consecutive frame images at different angles and different time series, performing feature extraction and feature fusion on the plurality of consecutive frame images based on a grid construction model to obtain grid data; processing the plurality of consecutive frame images based on ARcore to obtain a point cloud feature estimate, where the point cloud feature estimate includes color features; matching a plurality of points in the point cloud feature estimate and the grid data to obtain a matching relationship between the point cloud feature estimate and the grid data, and performing color assignment on the grid data based on the matching relationship to obtain the final grid data as the three-dimensional model.
[0007] Further, the grid construction model includes a reverse mapping layer, a gate sequence layer, and a multi-layer perception layer connected layer by layer.
[0008] Further, the grid construction model further includes an image encoder, which encodes multiple consecutive frame images and sends them to the inverse mapping layer.
[0009] Further, the inverse mapping layer performs inverse mapping on multiple consecutive frame images to obtain multiple initial 3D voxel features, and performs averaging processing on the multiple 3D voxel features to obtain 3D voxel features.
[0010] Further, the gate sequence layer fuses the 3D acceleration features corresponding to the consecutive frames under multiple time series to obtain the fused 3D voxel features.
[0011] Further, the multi-layer perceptron layer processes the fused 3D voxel features to obtain the TSDF transparency prediction value and the SDF prediction value of the 3D voxel features.
[0012] Further, matching multiple points in the point cloud feature and the grid data includes: obtaining a transformation matrix between the grid data and the point cloud feature based on the ICP algorithm, and obtaining the relationship between points in the point cloud feature and the grid data based on the transformation matrix.
[0013] Further, obtaining the matching relationship between the point cloud feature and the grid data includes: determining multiple points closest to each surface of the grid data in the point cloud feature, and assigning colors to each surface of the grid data according to the color features of the points.
[0014] Further, assigning colors to the grid data based on the matching relationship to obtain the final grid data as a three-dimensional model includes: comparing colors based on the colored points. When the colors of multiple points are the same, the color of the point is determined as the current color; if they are inconsistent, interpolation calculation can be performed according to the distance between each point and the surface edge to obtain the progressive color of the surface.
[0015] In a second aspect, a three-dimensional model reconstruction device is provided. The device includes: a grid data construction module for obtaining grid data based on multiple consecutive frame images at different angles and different time series through a grid construction model; a point cloud feature construction module for processing multiple consecutive frame images based on ARcore to obtain point cloud feature estimation; a three-dimensional model construction module for matching and constructing the grid data and the point cloud feature to obtain the final grid data as a three-dimensional model
[0016] In a third aspect, a terminal device is provided, including: at least one processor; and a memory, where the memory stores computer instructions that can be run on the processor, and when the instructions are executed by the processor, the steps of the method described in any one of the above are implemented.
[0017] Fourthly, a computer-readable storage medium is provided. The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps of the method described in any one of the above are implemented.
[0018] In the technical solution provided by the embodiments of the present application, by acquiring image data of a static environment under multiple angles and multiple time series, a three-dimensional model structure is built based on a grid construction model, and then the acquired image data is converted into colored 3D point cloud information through ARCore. By matching the previously obtained three-dimensional model with the point cloud, a colored three-dimensional model can be obtained. It realizes the unified and coordinated processing of the structure and color during the construction process of the three-dimensional model, and reduces the problems of mismatch between the structure and color and low authenticity caused by multiple independent steps. Description of the Drawings
[0019] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.
[0020] The methods, systems, and / or programs in the drawings will be further described according to exemplary embodiments. These exemplary embodiments will be described in detail with reference to the drawings. These exemplary embodiments are non-limiting exemplary embodiments, where the example numbers represent similar mechanisms in the various views of the drawings.
[0021] Figure 1 It is a schematic structural diagram of a terminal device provided by the embodiments of the present application.
[0022] Figure 2 It is a schematic diagram of a method shown in some embodiments of the present application.
[0023] Figure 3 It is a schematic block diagram of a device shown in some embodiments of the present application. Detailed Embodiments
[0024] In order to better understand the above technical solutions, the following will make a detailed description of the technical solutions of the present application through the drawings and specific embodiments. It should be understood that the embodiments of the present application and the specific features in the embodiments are detailed descriptions of the technical solutions of the present application, rather than limitations on the technical solutions of the present application. Without conflict, the technical features in the embodiments of the present application and the embodiments can be combined with each other.
[0025] In the following detailed description, many specific details are set forth by way of example in order to provide a thorough understanding of the relevant teachings. However, it will be apparent to those skilled in the art that the present application may be practiced without these details. In other instances, well-known methods, procedures, systems, components, and / or circuits have been described at a relatively high level without detail in order to avoid unnecessarily obscuring aspects of the present application.
[0026] Flowcharts are used in the present application to illustrate the execution processes performed by the systems according to the embodiments of the present application. It should be clearly understood that the execution processes of the flowcharts may not be executed in sequence. On the contrary, these execution processes may be executed in reverse order or simultaneously. Additionally, at least one other execution process may be added to the flowchart. One or more execution processes may be deleted from the flowchart.
[0027] Before further elaborating on the embodiments of the present invention, the nouns and terms involved in the embodiments of the present invention are described. The nouns and terms involved in the embodiments of the present invention are applicable to the following explanations.
[0028] (1) Responsive to, which is used to indicate the conditions or states upon which the performed operations depend. When the dependent conditions or states are satisfied, one or more of the performed operations may be real-time or may have a set delay; without special instructions, there is no limitation on the execution order of the multiple operations performed.
[0029] (2) Based on, which is used to indicate the conditions or states upon which the performed operations depend. When the dependent conditions or states are satisfied, one or more of the performed operations may be real-time or may have a set delay; without special instructions, there is no limitation on the execution order of the multiple operations performed.
[0030] Embodiments of the present application provide a terminal device. Figure 1 Shown is a schematic diagram of an embodiment of the terminal device provided by the present invention. As Figure 1 shown, embodiments of the present invention include the following devices: at least one processor 120; and a memory 110, where the memory 110 stores computer instructions, i.e., computer programs, that can run on the processor.
[0031] In this embodiment, the memory, the processor, and the communication unit are directly or
[0032] indirectly electrically connected to each other to achieve data transmission or interaction. For example, these elements may be electrically connected to each other through one or more communication buses or signal lines. The memory is used to store specific information and programs, and the communication unit is used to send the processed information to the corresponding user terminal.
[0033] In this embodiment, the storage module is divided into two storage areas. One storage area is a program storage unit, and the other storage area is a data storage unit. The program storage unit is equivalent to the firmware area.
[0034] The read and write permissions of this area are set to the read-only mode, and the data stored therein cannot be erased or changed. However, the data in the data storage unit can be erased, read, or written. When the capacity of the data storage area is full
[0035] the newly written data will overwrite the earliest historical data.
[0036] Among them, the memory can be, but is not limited to, Random Access Memory (RAM), Read Only Memory (ROM), Programmable Read-Only Memory (PROM), Erasable Programmable Read-Only Memory (EPROM), Electrically Erasable Programmable Read-Only Memory (EEPROM), etc.
[0037] The processor may be an integrated circuit chip with signal processing capabilities. The above-mentioned processor can be a general-purpose processor, including a Central Processing Unit (CPU), a Network Processor (NP), etc.; it can also be a Digital Signal Processor (DSP), an Application Specific Integrated Circuit (ASIC), a Field Programmable Gate Array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components. It can implement or execute the various methods, steps, and logic block diagrams disclosed in the embodiments of the present invention. The general-purpose processor can be a microprocessor or the processor can also be any conventional processor, etc.
[0038] When the computer instructions in this embodiment are executed by the processor, the following method is implemented:
[0039] Step S210. Obtain a plurality of consecutive frame images with different angles and different time series, and perform feature extraction and feature fusion on the plurality of consecutive frame images based on a grid construction model to obtain grid data.
[0040] Refer to Figure 2 For the above, when the computer instructions in this embodiment are executed by the processor, the following method is implemented:
[0041] Step S210. Obtain a plurality of consecutive frame images with different angles and different time series, and perform feature extraction and feature fusion on the plurality of consecutive frame images based on a grid construction model to obtain grid data.
[0042] In this embodiment, the construction of the 3D model is for the construction of the 3D model of a static environment or a static object in a static environment. The basic processing method for the construction of the 3D model is to collect images of the construction object and process the collected images to obtain a 3D model that meets the target requirements, where the 3D model is a computerized graphic of a static environment or a static object. Among them, the acquisition of continuous frame images is based on an image acquisition device that is set outside or inside the terminal device and communicates with the memory or processor in the terminal device. In this embodiment, the image acquisition device used can be an RGB camera device with color imaging ability. Among them, the definitions of different angles and different time sequences are respectively that for a static environment or a static object, full-angle image acquisition is performed. Because there is a time sequence before and after for the image frames during the image acquisition process, the continuous frame images are formed by presenting the acquired images in the order of time sorting, that is, the time sequence.
[0043] In this embodiment, the multiple continuous frame images obtained by the image acquisition device are initial image data. In order to construct a 3D model subsequently, the planar image data needs to be converted into mesh data with 3D expression characteristics, that is, mesh data. Since the mesh data, that is, the grid data, is the basic data of the 3D model, the core of this step is to extract features from the obtained multiple planar continuous frame images through a network construction model and fuse the multiple features to obtain the grid data.
[0044] The methods for feature extraction and feature fusion are implemented based on the grid construction model in this embodiment. The grid construction model proposed for this embodiment includes an image encoder, an inverse mapping layer, a gate sequence layer, and a multi-layer perceptron layer that are connected layer by layer.
[0045] Among them, the image encoder is mainly used to encode multiple continuous frame images and send them to the inverse mapping layer, and an existing encoder structure can be used for implementation. The inverse mapping layer is to perform inverse mapping on multiple continuous frame images to obtain multiple initial 3D voxel features. In this embodiment, in order to obtain complete and real 3D voxel features, according to different camera angles, the visible features presented by the voxel features at multiple angles are also different. Therefore, it is necessary to perform average processing on the voxel features at multiple angles to form 3D voxel features with higher credibility. Then, average processing is performed on multiple 3D voxel features to obtain 3D voxel features. The gate sequence layer fuses the 3D voxel features corresponding to multiple continuous frames under multiple time sequences to obtain the fused 3D voxel features. The multi-layer perceptron layer mainly processes the fused 3D voxel features to obtain the TSDF transparency prediction value and the SDF prediction value of the 3D voxel features, and completes the generation of the final grid data.
[0046] Step S220. Process the multiple consecutive frame images based on ARcore to obtain a point cloud feature estimation, where the point cloud feature estimation includes color features.
[0047] In this embodiment, through the depth estimation function and the camera pose estimation function carried by the ARCore framework, the captured RGB image is directly converted into a colored point cloud.
[0048] Step S230. Match multiple points in the point cloud feature estimation and the mesh data to obtain a matching relationship between the point cloud feature estimation and the mesh data, and assign colors to the mesh data based on the matching relationship to obtain the final mesh data as a three-dimensional model.
[0049] The mesh data and the corresponding point cloud estimation are obtained through Step S210 and Step S220 respectively. For Step S230, it is mainly to fuse the information obtained in the above two steps to realize the corresponding color assignment in the point cloud estimation of the mesh data.
[0050] Among them, the process of matching the point cloud features with multiple points in the mesh data includes: obtaining a transformation matrix between the mesh data and the point cloud features based on the ICP algorithm, and obtaining the relationship between points in the point cloud features and the mesh data based on the transformation matrix. Among them, the ICP algorithm is an existing algorithm and will not be elaborated in this embodiment.
[0051] After obtaining the relationship between points in the point cloud features and the mesh data, it is necessary to determine the matching relationship between the point cloud features and the mesh data, that is, to perform color assignment processing on the point cloud features in the mesh data. That is, determine multiple points in the point cloud features that are closest to each surface of the mesh data, and assign colors to each surface of the mesh data according to the color features of the points. In this embodiment, for the points after color assignment, color comparison is performed. When the colors of multiple points are the same, the color of this point is determined as the current color; if they are inconsistent, interpolation calculation can be performed according to the distance between each point and the surface edge to obtain the progressive color of this surface.
[0052] And, referring to Figure 3 , this embodiment provides a three-dimensional model reconstruction device 300, and the device includes: a mesh data construction module 310, configured to construct a mesh model based on multiple consecutive frame images at different angles and different time series to obtain mesh data. A point cloud feature construction module 320, configured to process multiple consecutive frame images based on ARcore to obtain a point cloud feature estimation. A three-dimensional model construction module 330, configured to match and construct the mesh data and the point cloud features to obtain the final mesh data as a three-dimensional model.
[0053] In the technical solution provided by the embodiments of the present application, by acquiring image data of a static environment under multiple angles and multiple time series, a three-dimensional model structure is built based on a grid construction model, and then the acquired image data is converted into colored 3D point cloud information through ARCore. By matching the previously obtained three-dimensional model with the point cloud, a colored three-dimensional model can be obtained. This realizes the unified collaborative processing of the structure and color during the construction of the three-dimensional model, and reduces the problems of mismatch between the structure and color and low authenticity caused by multiple independent steps.
[0054] It should be understood that for technical terms that have not been explained above, those skilled in the art can definitely determine their meanings through deduction before and after according to the content disclosed above, and no limitations are made here.
[0055] Those skilled in the art can definitely determine some preset, benchmark, predetermined, set, and preference-labeled technical features / technical terms according to the content disclosed above, such as thresholds, threshold intervals, threshold ranges, etc. For some technical feature terms that have not been explained, those skilled in the art can definitely deduce them reasonably and without doubt based on the logical relationship of the context, so as to clearly and completely implement the above technical solution. Prefixes of technical feature terms that have not been explained, such as "first", "second", "example", "target", etc., can be deduced and determined without doubt according to the context.
[0056] Suffixes of technical feature terms that have not been explained, such as "set", "list", etc., can also be deduced and determined without doubt according to the context.
[0057] The above content disclosed by the embodiments of the present application is clear and complete for those skilled in the art. It should be understood that the process of those skilled in the art deducing and analyzing technical terms that have not been explained is based on the content recorded in the present application. Therefore, the above content is not a judgment on the creativity of the overall solution.
[0058] The basic concepts have been described above. Obviously, for those skilled in the art, the above
[0059] detailed disclosure is only for example and does not constitute a limitation to the present application. Although not explicitly stated here, those skilled in the art can make various modifications, improvements, and corrections to the present application. Such modifications, improvements, and corrections are proposed in the present application, so such modifications, improvements, and corrections still belong to the spirit and scope of the exemplary embodiments of the present application.
[0060] 5 At the same time, this application uses specific terms to describe the embodiments of this application. For example, "an embodiment", "one embodiment", and / or "some embodiments" mean a certain feature, structure, or characteristic related to at least one embodiment of this application. Therefore, it should be emphasized and noted that the "one embodiment" or "an embodiment" or "an alternative embodiment" mentioned twice or more in different parts of this specification does not necessarily refer to the same embodiment. In addition, certain features, structures, or characteristics in at least one embodiment of this application can be appropriately combined.
[0061] Furthermore, those of ordinary skill in the art can understand that various aspects of this application can be illustrated and described by several patentable types or situations, including any new and useful process, machine, product, or combination of substances, or any new and useful improvement thereof.
[0062] Accordingly, various aspects of this application can be executed entirely by hardware, entirely by software (including firmware, resident software, microcode, etc.), or by a combination of hardware and software. The above-mentioned hardware or software can all be referred to as "units", "components", or "systems". In addition, various aspects of this application can be embodied as a computer product located in at least one computer-readable medium, and the product includes computer-readable program code.
[0063] A computer-readable signal medium may contain a propagated data signal containing computer program code, for example, on a baseband or as part of a carrier wave. This propagated signal may have various forms of manifestation, including electromagnetic form, optical form, etc., or a suitable combination of forms. A computer-readable signal medium can be any computer-readable medium other than a computer-readable storage medium, and this medium can be connected to an instruction execution system, device, or equipment to implement communication, propagation, or transmission for use of the program. The program code located on the computer-readable signal medium can be propagated through any suitable medium, including radio, cable, fiber optic cable, RF, or similar media, or any combination of the above media.
[0064]
[0065]
[0066] The computer program code required for the implementation of various aspects of the present application can be written in any combination of one or more programming languages, including object-oriented programming languages such as Java, Scala, Smalltalk, Eiffel, JADE, Emerald, C++, C#, VB.NET, Python, etc., or conventional programming languages similar thereto, such as the "C" programming language, Visual Basic, Fortran 2003, Perl, COBOL 2002, PHP, ABAP, dynamic programming languages such as Python, Ruby, and Groovy, or other programming languages. The programming code can be executed entirely on the user's computer, or executed as a stand-alone software package on the user's computer, or partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the latter case, the remote computer can be connected to the user's computer through any network form, such as a local area network (LAN) or a wide area network (WAN), or connected to an external computer (e.g., through the Internet), or in a cloud computing environment, or used as a service such as software as a service (SaaS).
[0067] In addition, unless specifically stated in the claims, the order of the processing elements and sequences, the use of numerical letters, or the use of other names in the present application are not used to limit the order of the processes and methods of the present application. Although some currently useful embodiments of the invention have been discussed through various examples in the above disclosure, it should be understood that such details are only for illustrative purposes, and the appended claims are not limited to the disclosed embodiments. On the contrary, the claims are intended to cover all modifications and equivalent combinations that conform to the essence and scope of the embodiments of the present application. For example, although the system components described above can be implemented by hardware devices, they can also be implemented only through software solutions, such as installing the described system on an existing server or mobile device.
[0068] It should also be understood that, in order to simplify the presentation of the disclosure of the present application and thus help the understanding of at least one embodiment of the invention, in the foregoing description of the embodiments of the present application, multiple features are sometimes grouped into one embodiment, drawing, or description thereof. However, this disclosure method does not mean that the features required by the subject matter of the present application are more than those mentioned in the claims. In fact, the features of the embodiments are less than all the features of the individual embodiments disclosed above.
Claims
1. A three-dimensional model reconstruction method, characterized in that, The method includes: Obtaining a plurality of consecutive frame images at multiple different angles and different time series, and performing feature extraction and feature fusion on the plurality of consecutive frame images based on a grid construction model to obtain grid data; Processing the plurality of consecutive frame images based on ARcore to obtain a point cloud feature estimation, where the point cloud feature estimation includes color features; Matching a plurality of points in the point cloud feature estimation and the grid data to obtain a matching relationship between the point cloud feature estimation and the grid data, and performing color assignment on the grid data based on the matching relationship to obtain the final grid data as a three-dimensional model.
2. The three-dimensional model reconstruction method according to claim 1, characterized in that, The grid construction model includes an inverse mapping layer, a gate sequence layer, and a multi-layer perceptron layer connected layer by layer.
3. The three-dimensional model reconstruction method according to claim 2, characterized in that, The grid construction model further includes an image encoder, and the image encoder encodes the plurality of consecutive frame images and sends them to the inverse mapping layer.
4. The three-dimensional model reconstruction method according to claim 3, characterized in that, The inverse mapping layer performs inverse mapping on the plurality of consecutive frame images to obtain a plurality of initial 3D voxel features, and performs averaging processing on the plurality of 3D voxel features to obtain 3D voxel features.
5. The three-dimensional model reconstruction method according to claim 3, characterized in that, The gate sequence layer fuses the 3D voxel features corresponding to the plurality of consecutive frames under multiple time series to obtain fused 3D voxel features.
6. The three-dimensional model reconstruction method according to claim 3, characterized in that, The multi-layer perceptron layer processes the fused 3D voxel features to obtain a TSDF transparency prediction value and an SDF prediction value for the 3D voxel features.
7. The three-dimensional model reconstruction method according to claim 1, characterized in that, Matching the plurality of points in the point cloud feature and the grid data includes: Obtaining a transformation matrix between the grid data and the point cloud feature based on the ICP algorithm, and obtaining the relationship between points in the point cloud feature and the grid data based on the transformation matrix.
8. The three-dimensional model reconstruction method according to claim 7, characterized in that, Obtaining the matching relationship between the point cloud feature and the grid data includes: Determining a plurality of points in the point cloud feature that are closest to each surface of the grid data, and performing color assignment on each surface of the grid data according to the color features of the points.
9. The three-dimensional model reconstruction method according to claim 8, characterized in that, Performing color assignment on the grid data based on the matching relationship to obtain the final grid data as a three-dimensional model includes: Performing color comparison based on the points after color assignment. When the colors of a plurality of points are the same, the color of the point is determined as the current color; if they are inconsistent, interpolation calculation can be performed according to the distance between each point and the surface edge to obtain the progressive color of the surface.
10. A three-dimensional model reconstruction device, characterized in that, The device includes: A grid data construction module for obtaining grid data based on a grid construction model from a plurality of consecutive frame images at multiple different angles and different time series; A point cloud feature construction module for processing a plurality of consecutive frame images based on ARcore to obtain a point cloud feature estimation; A three-dimensional model construction module for matching and constructing the grid data and the point cloud feature to obtain the final grid data as a three-dimensional model.
Citation Information
Patent Citations
Three-dimensional portrait real-time reconstruction and rendering method based on multi-view depth camera
CN112348957A
Static portrait three-dimensional reconstruction and dynamic face fusion method based on depth camera
CN112562083A