A three-dimensional reconstruction method and device based on structured three-dimensional Gaussian representation
Patent Information
- Application Number
- CN202411630317.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-11-14
- Publication Date
- 2026-09-11
- Estimated Expiration
- 2044-11-14
AI Technical Summary
高斯基元(Gauss ian Pr imit ives)作为一种有效的概率分布表示方式,已被广泛应用于三维物体表面和特征点的描述中,但传统高斯基元的非结构化表示无法满足结构化场景中精细三维重建的需求
[0018] 1) This application utilizes an optimal transfer algorithm to transform the unstructured Gaussian representation of the target object into a structured Gaussian representation. Then, using the textual description of the target object, a diffusion model is employed to complete the structured Gaussian representation, followed by 3D reconstruction. This application progressively optimizes the Gaussian representation of the target object using an optimal transfer algorithm, a textual description of the target object, and a diffusion model. This effectively addresses the shortcomings of traditional unstructured Gaussian representations in 3D reconstruction and achieves accurate Gaussian representation completion, enabling refined 3D reconstruction of the object.
Smart Images

Figure CN119625169B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of computer vision technology, and in particular to a three-dimensional reconstruction method and apparatus based on structured three-dimensional Gaussian representation. Background Technology
[0002] With the development of computer vision technology, 3D reconstruction technology has been widely used in many fields. Gaussian parameters, as an effective probability distribution representation, have been widely used in the description of 3D object surfaces and feature points. However, the unstructured representation of traditional Gaussian parameters cannot meet the needs of fine 3D reconstruction in structured scenes. Summary of the Invention
[0003] This application provides a three-dimensional reconstruction method and apparatus based on structured three-dimensional Gaussian representation, so as to improve the accuracy of the three-dimensional reconstruction model through structured Gaussian primitive representation.
[0004] This application provides the following solution:
[0005] According to a first aspect, a 3D reconstruction method based on structured 3D Gaussian representation is provided. The method includes: acquiring target images containing multiple viewpoints of a target object, wherein the target images from multiple viewpoints have viewpoint gaps relative to the full-view target image; generating an unstructured Gaussian representation of the target object based on the target images; establishing a transfer matrix describing the distribution of unstructured and structured Gaussian elements, solving the transfer matrix using an optimal transfer algorithm to obtain an optimal transfer matrix, and generating a structured Gaussian representation of the target object from the unstructured Gaussian elements based on the optimal transfer matrix; acquiring a text description of the target object, using a pre-trained diffusion model to complete the structured Gaussian elements based on the text description to obtain complete structured Gaussian elements; and performing 3D reconstruction using the complete structured Gaussian elements to obtain a 3D reconstruction model corresponding to the target object.
[0006] According to one achievable method in an embodiment of this application, generating an unstructured Gaussian representation of the target object based on the target image includes: performing image segmentation on the target image to obtain a target object region; extracting feature points from the target object region to obtain multiple feature points; clustering the feature points using a clustering algorithm; and fitting the clustering results to a Gaussian distribution to obtain unstructured Gaussian units.
[0007] According to one achievable method in an embodiment of this application, the method further includes: the structured Gaussian primitives are represented using a structured mesh, and the resolution of the structured Gaussian primitive representation is adjusted according to a scaling factor, wherein the scaling factor includes the size of the structured mesh; or, the structured Gaussian primitives are represented using a point cloud, and the resolution of the structured Gaussian primitive representation is adjusted according to a scaling factor, wherein the scaling factor includes the number of samples or the sampling density.
[0008] According to one achievable method in an embodiment of this application, the step of using a pre-trained diffusion model to complete the structured Gaussian units based on the text description to obtain complete structured Gaussian units includes: encoding the text description to obtain a high-dimensional vector representation corresponding to the text description; fusing the high-dimensional vector and the structured Gaussian units to obtain a fused vector; and generating the complete structured Gaussian units from the diffusion model based on the fused vector.
[0009] According to one achievable method in an embodiment of this application, encoding the text description to obtain a high-dimensional vector representation corresponding to the text description includes: loading trained low-rank matrix parameters, combining the low-rank matrix parameters with the parameter matrix of the language model to obtain a combined matrix; and using the combined matrix to generate a high-dimensional vector representation corresponding to the text description.
[0010] According to one achievable method in an embodiment of this application, generating the complete structured Gaussian unit from the diffusion model based on the fusion vector includes: acquiring control information and generating guidance information based on the control information using a control network;
[0011] The complete structured Gaussian unit is generated by the diffusion model based on the guiding information and the fusion vector.
[0012] According to one achievable method in an embodiment of this application, the step of using the complete structured Gaussian units to perform 3D reconstruction to obtain a 3D reconstruction model corresponding to the target object includes: performing 3D reconstruction using the complete structured Gaussian units to obtain a pre-reconstructed model of the target object; rasterizing the pre-reconstructed model to obtain 2D images from multiple angles, the 2D images including feature information of their corresponding viewpoints; determining the correspondence between different viewpoints based on the feature information; and generating a 3D reconstruction model corresponding to the target object based on the correspondence.
[0013] According to a second aspect, a 3D reconstruction apparatus based on structured 3D Gaussian representation is provided. The apparatus includes: a target image acquisition unit configured to acquire target images containing multiple perspectives of a target object, wherein the multiple perspective target images have perspective gaps compared to the full-view target image; an unstructured Gaussian primitive generation unit configured to generate an unstructured Gaussian primitive representation corresponding to the target object based on the target images; a structured Gaussian primitive generation unit configured to establish a transfer matrix describing the distribution of unstructured Gaussian primitives and the distribution of structured Gaussian primitives, solve the transfer matrix using an optimal transfer algorithm to obtain an optimal transfer matrix, and generate a structured Gaussian primitive representation corresponding to the target object from the unstructured Gaussian primitives based on the optimal transfer matrix; a Gaussian primitive completion unit configured to acquire a text description of the target object, complete the structured Gaussian primitives based on the text description using a pre-trained diffusion model to obtain complete structured Gaussian primitives; and a 3D reconstruction unit configured to perform 3D reconstruction using the complete structured Gaussian primitives to obtain a 3D reconstruction model corresponding to the target object.
[0014] According to a third aspect, a computer-readable storage medium is provided having a computer program stored thereon, which, when executed by a processor, implements the steps of the method described in any one of the first aspects above.
[0015] According to the fourth aspect, an electronic device is provided, comprising:
[0016] One or more processors; and a memory associated with the one or more processors, the memory being used to store program instructions that, when read and executed by the one or more processors, perform the steps of the method described in any one of the first aspects above.
[0017] According to the specific embodiments provided in this application, the following technical effects are disclosed:
[0018] 1) This application utilizes an optimal transfer algorithm to transform the unstructured Gaussian representation of the target object into a structured Gaussian representation. Then, using the textual description of the target object, a diffusion model is employed to complete the structured Gaussian representation, followed by 3D reconstruction. This application progressively optimizes the Gaussian representation of the target object using an optimal transfer algorithm, a textual description of the target object, and a diffusion model. This effectively addresses the shortcomings of traditional unstructured Gaussian representations in 3D reconstruction and achieves accurate Gaussian representation completion, enabling refined 3D reconstruction of the object.
[0019] 2) This application uses image segmentation, feature point extraction and clustering algorithms to obtain unstructured Gaussian units of the target object, which enhances the accuracy of feature extraction, reduces computational complexity and simplifies the generation process of unstructured Gaussian units.
[0020] 3) This application adjusts the density and resolution of structured Gaussian elements by setting a scaling factor, which can flexibly control the resolution of the 3D model under different computing resources and accuracy requirements, and meet the needs of different application scenarios.
[0021] 4) This application encodes text descriptions into high-dimensional vectors and fuses them with structured Gaussian primitives, which helps the model understand the high-level semantic information of the target object and realizes joint modeling of visual data and text data.
[0022] 5) This application uses a combination of low-rank matrix parameters and language models, which can more effectively extract and express key information in text descriptions, and better integrate text information with Gaussian primitives generated during the 3D reconstruction process, thereby improving the accuracy of the reconstruction model and the effect of detail restoration.
[0023] 6) This application introduces control information when generating structured Gaussian primitives using a diffusion model. A control network is used to generate guiding information, which is then combined with the fusion vector to make the generated Gaussian primitives more accurate and consistent with expectations. This guiding mechanism ensures that the model follows a specific direction during completion, thereby avoiding interference from irrelevant information and improving the stability and reliability of 3D reconstruction.
[0024] 7) This application first generates a pre-reconstructed model, and then generates two-dimensional images from multiple perspectives through rasterization to analyze and determine the feature information and correspondences of different perspectives. This method can enhance the detail representation and visual consistency of the three-dimensional model, ensuring that the reconstructed model can accurately display the true shape and features of the target object from different angles.
[0025] Of course, any product implementing this application does not necessarily need to achieve all of the advantages described above at the same time. Attached Figure Description
[0026] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0027] Figure 1 A flowchart of a 3D reconstruction method based on structured 3D Gaussian representation provided in an embodiment of this application;
[0028] Figure 2A schematic diagram illustrating the process of the 3D reconstruction method based on structured 3D Gaussian representation provided in this application embodiment;
[0029] Figure 3 A structural block diagram of a 3D reconstruction device based on structured 3D Gaussian representation provided in an embodiment of this application;
[0030] Figure 4 A schematic block diagram of an electronic device provided in an embodiment of this application. Detailed Implementation
[0031] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. All other embodiments obtained by those skilled in the art based on the embodiments of this application are within the scope of protection of this application.
[0032] The terminology used in the embodiments of this invention is for the purpose of describing particular embodiments only and is not intended to limit the invention. The singular forms “a,” “the,” and “the” as used in the embodiments of this invention and the appended claims are also intended to include the plural forms unless the context clearly indicates otherwise.
[0033] It should be understood that the term "and / or" used in this article is merely a description of the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A existing alone, A and B existing simultaneously, and B existing alone. Additionally, the character " / " in this article generally indicates that the preceding and following related objects have an "or" relationship.
[0034] Depending on the context, the word "if" as used here can be interpreted as "when," "when," "in response to determination," or "in response to detection." Similarly, depending on the context, the phrase "if determination" or "if detection (of the stated condition or event)" can be interpreted as "when determination," "in response to determination," "when detection (of the stated condition or event)," or "in response to detection (of the stated condition or event)."
[0035] Several 3D reconstruction techniques exist, typically using unstructured Gaussian units for direct reconstruction. However, due to the lack of organization and strict topological constraints in unstructured data, the generated 3D models often struggle to achieve high-precision detail representation, especially in complex shapes where unstructured representations may be inaccurate. Furthermore, when reconstructing a target object, multiple 2D images containing different perspectives are often used as target images. However, these target images often contain some degree of viewpoint loss, leading to incomplete, distorted, discontinuous, or void-like 3D reconstruction models, severely impacting the quality of the reconstruction.
[0036] In view of this, this application provides a new approach. Figure 1 A flowchart of a 3D reconstruction method based on structured 3D Gaussian representation provided in this application embodiment is shown below. Figure 1 As shown, the method may include the following steps:
[0037] Step 101: Obtain target images containing multiple perspectives of the target object, wherein the multiple perspective target images have missing perspectives compared to the full-view target image.
[0038] Step 102: Generate an unstructured Gaussian representation of the target object based on the target image.
[0039] Step 103: Establish a transfer matrix describing the distribution of unstructured Gaussian elements and the distribution of structured Gaussian elements. Solve the transfer matrix using the optimal transfer algorithm to obtain the optimal transfer matrix. Based on the optimal transfer matrix, generate the structured Gaussian representation of the target object from the unstructured Gaussian elements.
[0040] Step 104: Obtain the text description of the target object, and use the pre-trained diffusion model to complete the structured Gaussian units based on the text description to obtain the complete structured Gaussian units.
[0041] Step 105: Use the complete structured Gaussian units to perform 3D reconstruction and obtain the 3D reconstruction model corresponding to the target object.
[0042] As can be seen from the above process, this application transforms the unstructured Gaussian representation of the target object into a structured Gaussian representation using an optimal transfer algorithm. Then, using the textual description of the target object, a diffusion model is employed to complete the structured Gaussian representation before 3D reconstruction is performed. This application utilizes an optimal transfer algorithm, the textual description of the target object, and a diffusion model to progressively optimize the Gaussian representation of the target object, effectively addressing the shortcomings of traditional unstructured Gaussian representations in 3D reconstruction and achieving refined reconstruction and completion of the object.
[0043] The following describes in detail each step of the above process and the effects that can be further produced, with reference to the embodiments. First, step 101, namely "acquiring target images containing multiple perspectives of the target object, wherein the target images with multiple perspectives have missing perspectives compared to the target image with full perspective", will be described in detail with reference to the embodiments.
[0044] This application aims to perform 3D reconstruction on a target image, which is an image containing a target object. This application acquires multiple target images containing multiple viewpoints, providing multi-view information through different imaging angles. Notably, this application can perform 3D reconstruction even when the target image has missing viewpoints. Therefore, the multiple viewpoint target images acquired by this application have missing viewpoints compared to the full-view target image.
[0045] Target images can be acquired through various means, such as public data platforms or by taking pictures from different imaging angles using imaging devices. After obtaining an image containing the target object from a data platform or imaging device, preprocessing operations such as cropping and geometric correction can be performed on the image to obtain a higher quality target image.
[0046] The following describes step 102, namely "generating an unstructured Gaussian representation of the target object based on the target image", in detail with reference to the embodiments.
[0047] Unstructured Gaussian elements are a probabilistic representation of the distribution of feature points on an object. They are typically composed of parameters such as the mean, variance, and weights of the distribution and are used to describe the features of local regions on the object's surface. This representation constructs a point cloud or a continuous distribution model in three-dimensional space, which can reflect the object's shape and structure to a certain extent.
[0048] The unstructured Gaussian representation of the target object can be obtained from the target image in a variety of ways. One feasible way is to extract point clouds from multiple target images and obtain a three-dimensional point cloud representation of the target object through multi-view image matching techniques, such as Structure from Motion (SFM) or Multi-View Stereo (MVS). Then, Gaussian fitting is performed on the point cloud to generate unstructured Gaussian units.
[0049] As another possible approach, the following steps are included:
[0050] First, image segmentation can be performed on the target image to obtain the target object region. Besides the target object, the target image may also contain other objects or background patterns. Image segmentation is needed to extract the target object region from the target image; the segmented region is the target object region. Image segmentation can be achieved using algorithms such as edge detection and semantic segmentation.
[0051] Next, feature points are extracted from the target object region to obtain multiple feature points. Feature point extraction aims to identify key points from the two-dimensional image of the target object that can uniquely describe the object's features. These feature points typically possess the following properties: stability: they remain consistent under different viewpoints, lighting, and noise conditions; discriminability: they have good discriminability and can be used to match the same feature in different images; repeatability: they can appear repeatedly in multiple images, providing necessary data support for 3D reconstruction.
[0052] Feature point extraction can be achieved using various algorithms. For example, the Harris corner detection method can be used to find corners with strong responses by calculating the autocorrelation matrix of the image. Alternatively, scale-invariant feature transformation algorithms can be used to extract feature points that are invariant to scale and rotation by detecting extreme points in multi-scale space.
[0053] Clustering algorithms are used to cluster feature points, and the clustering results are fitted with a Gaussian distribution to obtain unstructured Gaussian units. Specifically, a set of feature points extracted from the image is selected as the input to the clustering algorithm; the selected clustering algorithm is used to cluster the feature points. The clustering results will form multiple clusters, each containing feature points with similar characteristics; the effectiveness of the clustering can be evaluated using metrics such as the silhouette coefficient and the Davies-Bouldin index to assess the quality of the clustering.
[0054] After clustering, a Gaussian distribution is fitted based on the clustering results. Specifically, for each cluster, the set of feature points is extracted; for each cluster, its mean, covariance matrix, and other statistical parameters are calculated to form the basis of the unstructured Gaussian distribution. These parameters describe the spatial distribution of feature points. Using the obtained parameters, an unstructured Gaussian distribution model is constructed. Each unstructured Gaussian unit consists of multiple Gaussian distributions, defined by its mean and variance, describing the distribution of feature points in that region. Furthermore, for multiple clustering results, it may be necessary to merge multiple unstructured Gaussian units to form a complete unstructured Gaussian unit representation. For example, similar Gaussian units can be merged into a larger Gaussian unit, depending on the requirements.
[0055] This application can also standardize unstructured Gaussian elements. Standardization can balance the differences between different unstructured Gaussian elements, avoid the influence of numerical deviations on the solution process of the transfer matrix, and thus improve the accuracy of the optimal transfer matrix solution.
[0056] The following describes in detail step 103, namely, "establishing a transfer matrix describing the distribution of unstructured Gaussian elements and the distribution of structured Gaussian elements, solving the transfer matrix using the optimal transfer algorithm to obtain the optimal transfer matrix, and generating the structured Gaussian representation of the target object from the unstructured Gaussian elements according to the optimal transfer matrix", with reference to the embodiments.
[0057] A transfer matrix is a mathematical tool used to describe the relationship between two distributions. In this application, the transfer matrix is used to represent the correspondence between unstructured Gaussian distributions and structured Gaussian distributions. Specifically, each element of the transfer matrix can be viewed as a "transfer amount" from one primitive distribution (unstructured) to another primitive distribution (structured).
[0058] The steps for constructing a transfer matrix include: First, matching feature points in unstructured Gaussian elements and structured Gaussian elements. A correspondence can be established by calculating the distance or similarity metric between them; based on the feature point matching, the elements of the transfer matrix are assigned values. Generally, each element T of the matrix... ij The "transfer amount" representing the distance from the i-th unstructured Gaussian unit to the j-th structured Gaussian unit can be calculated using a distance metric (e.g., Euclidean distance) or a similarity function. To ensure the validity of the transfer matrix, it may be necessary to normalize it so that the sum of each row or column equals 1. This helps to treat the transfer matrix as a probability distribution in subsequent processing.
[0059] The optimal transmission algorithm is a mathematical tool used to find the optimal solution to transform one distribution into another. Its goal is to minimize transmission cost. The optimal transmission matrix can be obtained by solving for the transmission matrix using the optimal transmission algorithm, which can be achieved through the following steps:
[0060] First, we define a transmission cost function, which is a function based on distance or similarity metrics, representing the "transportation cost" from unstructured primitives to structured primitives.
[0061] Next, the transfer matrix is solved using optimization algorithms (such as Sinkhorn distance, linear programming, etc.) to find the optimal transfer matrix. Each element in the optimal transfer matrix represents the optimal transfer amount from each unstructured Gaussian cell to a structured Gaussian cell. An iterative algorithm is used to solve for the optimal transfer matrix until convergence is achieved. The iterative process adjusts the elements of the transfer matrix to reduce the total transfer cost.
[0062] Preferably, a regularization term can be added to the transmission cost function. The regularization term can introduce prior knowledge or other information (such as marginal distribution, smoothness, sparsity, etc.), thereby giving the optimization problem a clearer direction.
[0063] After obtaining the optimal transfer matrix, the structured Gaussian representation of the target object is obtained based on the optimal transfer matrix. This process can be achieved in several ways. For example, the unstructured source Gaussian units can be directly combined using a weighted method to obtain the structured Gaussian representation of the target object; or the unstructured Gaussian units can be weighted and combined into structured Gaussian units using the weights in the optimal transfer matrix; or the parameters of the structured Gaussian units can be iteratively updated using the expectation-maximization algorithm to make the distribution of the structured Gaussian units match the transfer rules guided by the optimal transfer matrix as closely as possible.
[0064] Structured Gaussian elements can be represented in various ways, such as using a point cloud approach. The Gaussian elements generated by the optimal transfer matrix are treated as a sparse point cloud distributed in space, recording the 3D coordinates, density, and other attributes of each point. The resolution of the structured Gaussian element representation is adjusted according to a scaling factor, which may include the number of samples or the sampling density.
[0065] Structured Gaussian elements can also be represented by structured meshes. A structured mesh divides three-dimensional space into regular mesh cells, each of which can contain one or more Gaussian elements to describe the spatial characteristics of an object within that cell. The denser the mesh (higher resolution), the more detailed information each cell contains. Therefore, the resolution of the structured Gaussian element representation can be adjusted according to a scaling factor, which includes the size of the structured mesh.
[0066] Resolution here typically refers to the fineness or sampling density of Gaussian pixels distributed in 3D space. The resolution of Gaussian pixels can be adjusted using different representation methods to maintain the aspect ratio of an object, and different levels of detail can be achieved through specific scaling factors. Resolution is adjusted by increasing or decreasing the value of the scaling factor. For example, low resolution can use fewer voxel meshes to represent the overall structure, while high resolution uses a denser voxel mesh to capture more detail.
[0067] The following describes in detail step 104, namely, "obtaining a text description of the target object, using a pre-trained diffusion model to complete the structured Gaussian units based on the text description, and obtaining complete structured Gaussian units," with reference to an embodiment.
[0068] This application utilizes a pre-trained diffusion model to complete structured Gaussian primitives. The diffusion model is a generative model that progressively transforms data into noise and then reconstructs the data from the noise. In the scenario of Gaussian primitive structured completion, the diffusion model can complete missing or incomplete parts from existing structured Gaussian primitives by simulating the progressive generation process of data.
[0069] When using the diffusion model for completion, this application guides the reasoning of the diffusion model through the textual description of the target object. Figure 2 This is a schematic diagram illustrating the process of the 3D reconstruction method based on structured 3D Gaussian representation provided in this application embodiment. The text description of the target object can be text that describes the characteristics of the target object, including but not limited to: the text description can be the category information of the target object, such as "cat," "building," "plant," etc., which helps the diffusion model generate Gaussian elements that conform to specific category characteristics; it can be a description of the shape, size, position, or spatial structure of the target object, such as "circle," "rectangle," "distributed in the left area," etc., which can affect the distribution parameters of the Gaussian elements, making the completion result more consistent with the geometric features of the object; it can be a description of the visual data, color, or texture of the target object, such as "blue with stripes," which helps the diffusion model generate more distinctive features when generating the completed Gaussian elements; for target objects with motion characteristics, the text description can be a description of its behavior or posture, such as "moving forward," "tilting," etc., which helps complete the distribution of Gaussian elements that reflects dynamics.
[0070] Text descriptions can be obtained in various ways, including manual annotation and automated extraction. Manual annotation involves humans providing precise object descriptions, suitable for more detailed scene requirements. Automated extraction utilizes image tagging tools, feature extraction models, or natural language processing models to automatically generate descriptions from data. Furthermore, knowledge bases or metadata tags can be used to find detailed descriptions of the target object, or users can provide descriptions based on specific needs to meet personalized completion requirements. These description methods provide rich semantic information to the generative model, enhancing the accuracy and consistency of the completion results.
[0071] After obtaining the text description, the text description and the structured Gaussian units can be input together into the pre-trained diffusion model. The diffusion model will then complete the structured Gaussian units based on the text description, thus obtaining the complete structured Gaussian units.
[0072] As one possible implementation, this application can further encode the text description to convert it into a high-dimensional vector representation; fuse the high-dimensional vector and the structured Gaussian unit to obtain a fused vector; and generate the complete structured Gaussian unit from the diffusion model based on the fused vector.
[0073] Further encoding of the text description can be achieved using traditional methods such as sentence encoding and multimodal encoding; alternatively, a language model can be used. This involves fusing the high-dimensional vector representation of the text description with structured Gaussian primitives. This step can be implemented in several ways, such as weighted summation, convolution, and attention mechanisms. The goal of this fusion is to combine the semantic information of the text with the information from the Gaussian distribution, forming a fused vector that incorporates features from both.
[0074] As another feasible approach, a high-dimensional vector representation of the text description can be generated using a large language model. Specifically, the low-rank adaptation of large language models (LoRA) technique can be utilized. By loading pre-trained low-rank matrix parameters, the large language model can be made capable of adapting to specific tasks. Specifically, the pre-trained low-rank matrix parameters are loaded and combined with the parameter matrix of the language model to obtain a combined matrix; this combined matrix is then used to generate the high-dimensional vector representation of the text description. The combination of the low-rank matrix parameters and the parameter matrix of the language model can involve adding, weighted summing, convolution, etc., to superimpose the low-rank matrix onto the original model's parameters, resulting in higher accuracy of the generated results.
[0075] Language models can be based on Large Language Models (LLMs) or ordinary pre-trained language models. Large Language Models (LLMs) are deep learning models trained on massive amounts of text data, capable of generating natural language text or understanding the meaning of language text. They are characterized by their enormous scale and massive number of parameters (typically exceeding tens of billions), and are usually based on deep learning architectures such as the Transformer architecture. The difference between LLMs and ordinary pre-trained language models lies in the scale of their parameters. When the parameter scale exceeds a certain level, the model achieves significant performance improvements and exhibits capabilities not found in smaller models, such as in-context learning capabilities. It can learn complex patterns in language and perform a wide range of tasks, including text summarization, translation, sentiment analysis, multi-turn dialogue, and more. Therefore, to distinguish them from traditional pre-trained language models, models with a parameter scale exceeding a certain level are called LLMs. In general, language models with a parameter scale exceeding tens of billions implemented based on deep learning architectures can be considered large language models. Common LLMs include: GTP-3 (Generative Pre-trained Transformer 3), T5 (Text-to-Text Transfer Transformer), GTP-4, PaLM (a large language model proposed by Google), LLaMA (Large Language Model Meta AI, a large language model released by Meta AI), and so on.
[0076] As another feasible approach, this application, when generating complete structured Gaussian elements from the fusion vector using a diffusion model, can also guide the Gaussian element completion process through a ControlNet. ControlNet serves as an effective control mechanism, ensuring the model follows specific structural information during generation. Specifically, control information is acquired, and guidance information is generated using the control network based on this control information; the complete structured Gaussian elements are then generated by the diffusion model based on the guidance information and the fusion vector.
[0077] In this system, control information serves as the input to the control network, constraining the generation process of the diffusion model. This control information can be a structural map of the target object, such as an edge map, depth map, or attitude map, or a Gaussian element distribution map describing the density or positional distribution of the target object in space. Based on the control information, the control network outputs guiding information, which constrains the generation process of the diffusion model, ensuring that the generated content conforms to the input conditions.
[0078] The following describes step 105, namely "using complete structured Gaussian units to perform three-dimensional reconstruction and obtain the three-dimensional reconstruction model corresponding to the target object", in detail with reference to the embodiments.
[0079] First, the complete structured Gaussian primitives are preprocessed to extract features containing 3D spatial location and distribution information. Then, based on these structured Gaussian primitives, a 3D reconstruction algorithm is used to convert them into 3D point cloud data of the target object. During this process, a Gaussian distribution-based density estimation algorithm can be employed to complete the point cloud and adjust its density, making the generated point cloud more consistent with the morphological characteristics of the target object. Subsequently, a surface reconstruction algorithm is used to convert the point cloud data into a continuous 3D mesh model. For example, methods such as Poisson surface reconstruction, Marching Cubes, or Delaunay triangulation are used to generate a 3D model with a complete surface structure.
[0080] Preferably, the generated 3D model can be optimized, for example, by removing redundant noise or smoothing, to improve the model's accuracy and visual effect, thereby obtaining a 3D reconstruction model of the target object that conforms to the structured Gaussian elements.
[0081] As an feasible approach, this application utilizes rasterization technology to optimize the 3D reconstruction model. Specifically: 3D reconstruction is performed using complete structured Gaussian units to obtain a pre-reconstructed model of the target object; the pre-reconstructed model is rasterized to obtain 2D images from multiple angles, wherein the 2D images include feature information of their corresponding viewpoints; the correspondence between different viewpoints is determined based on the feature information; and the 3D reconstruction model of the target object is obtained based on the correspondence.
[0082] Rasterization technology is widely used in view synthesis and multi-view geometric reconstruction. This technology projects a 3D object from different viewpoints to generate 2D view images at different angles. These view images can represent the depth information and geometric structure of a scene. By employing rasterization, the target 3D model is converted into multiple 2D projected images according to set camera viewpoints and intrinsic / extrinsic parameters, effectively capturing the features, textures, and details of the object's surface. This application first renders the pre-reconstructed model into 2D images from multiple set angles using rasterization, allowing the feature information from different viewpoints to be fully displayed in the 2D images. Next, by analyzing and extracting features from each of the multi-view images, image registration or stereo vision methods are used to match features between adjacent views to determine the correspondence between different viewpoints. Increasing the number of viewpoints can also improve the accuracy and stability of the reconstruction. Finally, based on the optimized correspondence, a 3D reconstructed model of the target object is obtained. Both the generation of the pre-reconstructed model and the generation of the 3D reconstructed model can be obtained using the 3D reconstruction methods mentioned above, without specific limitations.
[0083] In this application, the pre-trained diffusion model can be trained using the following method:
[0084] Training data comprising multiple training samples is acquired. The training samples are image samples including target objects, and the image samples include full-view image representations of the target object samples. Unstructured Gaussian primitives corresponding to the target objects are obtained using the image samples. The unstructured Gaussian primitives are then structured to obtain structured Gaussian primitive samples corresponding to the target objects. The structured Gaussian primitive samples are represented using a three-dimensional mesh space. The three-dimensional mesh space of the structured Gaussian primitive samples is sampled to obtain three-dimensional local blocks. The regions corresponding to the three-dimensional local blocks are removed from the three-dimensional mesh space of the target object samples to obtain input samples. The input samples are input into the diffusion model to obtain output Gaussian primitives generated by the diffusion model based on the input samples for the input samples. The training objective includes minimizing the difference between the output Gaussian primitives for the input samples and the structured Gaussian primitive samples.
[0085] The foregoing has described specific embodiments of this specification. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recited in the claims may be performed in a different order than that shown in the embodiments and may still achieve the desired result. Furthermore, the processes depicted in the drawings do not necessarily require the specific or sequential order shown to achieve the desired result. In some embodiments, multitasking and parallel processing are possible or may be advantageous.
[0086] According to another embodiment, a three-dimensional reconstruction apparatus based on structured three-dimensional Gaussian representation is provided. Figure 3 A schematic block diagram of a 3D reconstruction apparatus based on a structured 3D Gaussian representation according to one embodiment is shown. The apparatus 300 includes:
[0087] The target image acquisition unit 301 is configured to acquire target images containing multiple perspectives of the target object, where the multiple perspective target images have a missing perspective compared to the full-view target image.
[0088] The unstructured Gaussian generation unit 302 is configured to generate an unstructured Gaussian representation of the target object based on the target image.
[0089] The structured Gaussian generation unit 303 is configured to establish a transfer matrix describing the distribution of unstructured Gaussian elements and the distribution of structured Gaussian elements, solve the transfer matrix using an optimal transfer algorithm to obtain the optimal transfer matrix, and generate the structured Gaussian representation of the target object from the unstructured Gaussian elements based on the optimal transfer matrix.
[0090] The Gaussian completion unit 304 is configured to acquire a textual description of the target object and use a pre-trained diffusion model to complete the structured Gaussian units based on the textual description, thereby obtaining the complete structured Gaussian units.
[0091] The 3D reconstruction unit 305 is configured to perform 3D reconstruction using complete structured Gaussian units to obtain a 3D reconstruction model corresponding to the target object.
[0092] As one possible implementation, the unstructured Gaussian unit 302 can be configured to: perform image segmentation on the target image to obtain the target object region; extract feature points from the target object region to obtain multiple feature points; cluster the feature points using a clustering algorithm; and fit the clustering results to a Gaussian distribution to obtain unstructured Gaussian units.
[0093] As one possible approach, the structured Gaussian primitives (GMPs) are represented using a structured mesh, with the resolution of the GMP representation adjusted according to a scaling factor, which includes the size of the structured mesh; or, the GMPs are represented using a point cloud, with the resolution of the GMP representation adjusted according to a scaling factor, which includes the number of samples or the sampling density.
[0094] As one possible implementation, the Gaussian complement unit 304, when using a pre-trained diffusion model to complete the structured Gaussian units based on the text description to obtain the complete structured Gaussian units, can be configured as follows: encoding the text description to obtain a high-dimensional vector representation corresponding to the text description; fusing the high-dimensional vector and the structured Gaussian units to obtain a fused vector; and generating the complete structured Gaussian units from the diffusion model based on the fused vector.
[0095] As one possible approach, the Gaussian complement unit 304 can be configured to: load pre-trained low-rank matrix parameters, combine the low-rank matrix parameters with the parameter matrix of the language model to obtain a combined matrix, and use the combined matrix to generate a high-dimensional vector representation of the text description.
[0096] As one possible implementation, the Gaussian complement unit 304, when generating the complete structured Gaussian from the diffusion model based on the fusion vector, can be configured to: acquire control information, generate guidance information using a control network based on the control information, and generate the complete structured Gaussian from the diffusion model based on the guidance information and the fusion vector.
[0097] As one possible implementation, the 3D reconstruction unit 305 can be configured to: perform 3D reconstruction using complete structured Gaussian units to obtain a pre-reconstructed model of the target object; rasterize the pre-reconstructed model to obtain 2D images from multiple angles, the 2D images including feature information of their corresponding viewpoints; determine the correspondence between different viewpoints based on the feature information; and generate a 3D reconstruction model corresponding to the target object based on the correspondence.
[0098] The various embodiments in this specification are described in a progressive manner. Similar or identical parts between embodiments can be referred to mutually. Each embodiment focuses on describing the differences from other embodiments. In particular, the device embodiments are basically similar to the method embodiments, so the description is relatively simple; relevant parts can be referred to the descriptions in the method embodiments. The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs. Those skilled in the art can understand and implement this without creative effort.
[0099] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, data stored, data displayed, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties. Furthermore, the collection, use and processing of the relevant data must comply with the relevant laws, regulations and standards of the relevant countries and regions, and corresponding operation portals are provided for users to choose to authorize or refuse.
[0100] In addition, embodiments of this application also provide a computer-readable storage medium storing a computer program thereon, which, when executed by a processor, implements the steps of the method described in any of the foregoing method embodiments.
[0101] And an electronic device comprising: one or more processors; and a memory associated with the one or more processors, the memory being used to store program instructions that, when read and executed by the one or more processors, perform the steps of the method described in any of the foregoing method embodiments.
[0102] This application also provides a computer program product, including a computer program that, when executed by a processor, implements the steps of the method described in any of the foregoing method embodiments.
[0103] in, Figure 4 An exemplary architecture of an electronic device is shown, which may include a processor 410, a video display adapter 411, a disk drive 412, an input / output interface 413, a network interface 414, and a memory 420. The processor 410, video display adapter 411, disk drive 412, input / output interface 413, network interface 414, and memory 420 can communicate with each other via a communication bus 430.
[0104] The processor 410 can be implemented using a general-purpose CPU, microprocessor, application-specific integrated circuit (ASIC), or one or more integrated circuits to execute relevant programs and implement the technical solution provided in this application.
[0105] The memory 420 can be implemented in the form of ROM (Read-Only Memory), RAM (Random Access Memory), static storage device, dynamic storage device, etc. The memory 420 can store the operating system 421 for controlling the operation of the electronic device 400, and the basic input / output system (BIOS) 422 for controlling the low-level operations of the electronic device 400. Additionally, it can store a web browser 423, a data storage management system 424, and a 3D reconstruction device 425 based on structured 3D Gaussian representation, etc. The aforementioned 3D reconstruction device 425 based on structured 3D Gaussian representation can be the application program that specifically implements the aforementioned steps in this embodiment. In summary, when implementing the technical solution provided in this application through software or firmware, the relevant program code is stored in the memory 420 and executed by the processor 410.
[0106] Input / output interface 413 is used to connect input / output modules to realize information input and output. Input / output modules can be configured as components in the device (not shown in the figure) or externally connected to the device to provide corresponding functions. Input devices may include keyboards, mice, touch screens, microphones, various sensors, etc., and output devices may include displays, speakers, vibrators, indicator lights, etc.
[0107] Network interface 414 is used to connect a communication module (not shown in the figure) to enable communication between this device and other devices. The communication module can communicate via wired means (such as USB, Ethernet cable, etc.) or wireless means (such as mobile network, WIFI, Bluetooth, etc.).
[0108] Bus 430 includes a pathway for transmitting information between various components of the device, such as processor 410, video display adapter 411, disk drive 412, input / output interface 413, network interface 414, and memory 420.
[0109] It should be noted that although the above-described device only shows the processor 410, video display adapter 411, disk drive 412, input / output interface 413, network interface 414, memory 420, bus 430, etc., in specific implementations, the device may also include other components necessary for normal operation. Furthermore, those skilled in the art will understand that the above-described device may only include the components necessary for implementing the solution of this application, and does not necessarily include all the components shown in the figures.
[0110] As can be seen from the above description of the embodiments, those skilled in the art can clearly understand that this application can be implemented by means of software plus necessary general-purpose hardware platforms. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, can be embodied in the form of a computer program product. This computer program product can be stored in a storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute the methods described in various embodiments or some parts of the embodiments of this application.
[0111] The technical solutions provided in this application have been described in detail above. Specific examples have been used to illustrate the principles and implementation methods of this application. The descriptions of the above embodiments are only for the purpose of helping to understand the methods and core ideas of this application. Furthermore, those skilled in the art will recognize that, based on the ideas of this application, there will be changes in the specific implementation methods and application scope. Therefore, the content of this specification should not be construed as a limitation of this application.
Claims
1. A method for 3D reconstruction based on structured 3D Gaussian representation, characterized in that, The method includes: Acquire target images containing multiple perspectives of the target object, wherein the multiple perspective images have missing perspectives compared to the full-view target image; Based on the target image, generate an unstructured Gaussian representation of the target object; A transfer matrix describing the distribution of unstructured Gaussian elements and the distribution of structured Gaussian elements is established. The optimal transfer matrix is obtained by solving the transfer matrix using the optimal transfer algorithm. Based on the optimal transfer matrix, the structured Gaussian representation of the target object is generated from the unstructured Gaussian elements. The process involves: acquiring a text description of the target object; loading pre-trained low-rank matrix parameters; combining the low-rank matrix parameters with the parameter matrix of the language model to obtain a combined matrix; using the combined matrix to generate a high-dimensional vector representation corresponding to the text description; fusing the high-dimensional vector with the structured Gaussian elements to obtain a fused vector; acquiring control information; using a control network to generate guidance information based on the control information; and generating complete structured Gaussian elements from a pre-trained diffusion model based on the guidance information and the fused vector. The complete structured Gaussian units are used for 3D reconstruction to obtain a pre-reconstructed model of the target object; the pre-reconstructed model is rasterized to obtain 2D images from multiple angles, the 2D images including feature information of their corresponding viewpoints; the correspondence between different viewpoints is determined based on the feature information; and a 3D reconstruction model corresponding to the target object is generated based on the correspondence.
2. The method of claim 1, wherein, The step of generating the unstructured Gaussian representation of the target object based on the target image includes: The target image is segmented to obtain the target object region; Feature points are extracted from the target object region to obtain multiple feature points; The feature points are clustered using a clustering algorithm, and the clustering results are used to fit a Gaussian distribution to obtain unstructured Gaussian units.
3. The method of claim 1, wherein, The method further includes: The structured Gaussian elements are represented using a structured mesh, and the resolution of the structured Gaussian element representation is adjusted according to a scaling factor, which includes the size of the structured mesh. Alternatively, the structured Gaussian primitives may be represented as point clouds, and the resolution of the structured Gaussian primitive representation may be adjusted according to a scaling factor, which may include the number of samples or the sampling density.
4. A device for 3D reconstruction based on structured 3D Gaussian representation, characterized in that, The apparatus performs the method as described in any one of claims 1-3, the apparatus comprising: The target image acquisition unit is configured to acquire target images containing multiple perspectives of the target object, wherein the multiple perspective target images have missing perspectives compared to the full-view target image; An unstructured Gaussian unit is configured to generate an unstructured Gaussian representation of the target object based on the target image. The structured Gaussian generation unit is configured to establish a transfer matrix describing the distribution of unstructured Gaussian elements and the distribution of structured Gaussian elements, solve the transfer matrix using an optimal transfer algorithm to obtain an optimal transfer matrix, and generate a structured Gaussian representation of the target object from the unstructured Gaussian elements based on the optimal transfer matrix. The Gaussian complement unit is configured to: acquire a text description of the target object; load pre-trained low-rank matrix parameters; combine the low-rank matrix parameters with the parameter matrix of the language model to obtain a combination matrix; use the combination matrix to generate a high-dimensional vector representation corresponding to the text description; fuse the high-dimensional vector and the structured Gaussian complements to obtain a fused vector; acquire control information; use a control network to generate guidance information based on the control information; and generate complete structured Gaussian complements from a pre-trained diffusion model based on the guidance information and the fused vector. The 3D reconstruction unit is configured to perform 3D reconstruction using the complete structured Gaussian units to obtain a pre-reconstructed model of the target object; rasterize the pre-reconstructed model to obtain 2D images from multiple angles, the 2D images including feature information of their corresponding viewpoints; determine the correspondence between different viewpoints based on the feature information; and generate a 3D reconstruction model corresponding to the target object based on the correspondence.
5. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the program is executed by the processor, it implements the steps of the method according to any one of claims 1 to 3.
6. An electronic device, characterized in that, include: One or more processors; And a memory associated with the one or more processors, the memory being used to store program instructions that, when read and executed by the one or more processors, perform the steps of the method according to any one of claims 1 to 3.