A Zero-Shot 3D Generation Method for Enhancing Viewpoint Consistency

By introducing the viewing angle decoupling method VDM and similarity partial order loss PSL in the text-to-3D generation method, combined with the 3D Gaussian splattering technology, the problems of viewing angle inconsistency and geometric collapse in the prior art are solved, and higher viewing angle consistency and generation quality are achieved.

CN119152101BActive Publication Date: 2025-05-27NANJING DITAVI DATA TECH CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202411625074.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-11-14
Publication Date
2025-05-27
Estimated Expiration
2044-11-14

AI Technical Summary

Technical Problem

The existing text-to-3D generation methods have shortcomings in the problem of view angle inconsistency and geometric collapse, resulting in the same 3D object being inconsistent at different perspectives, affecting the realism of the generated results.

Method used

Through the mathematical analysis of multi-faceted problems, the pre-trained text encoder CLIP and the three-dimensional diffusion model PointE are used, and the 3D Gaussian splashing technology and the viewing angle decoupling method VDM are combined to enhance the viewing angle consistency. The specific steps include: obtaining the text prompt words entered by the user, extracting the subject words and constructing a viewpoint description phrase, generating rough 3D point cloud data, fitting an explicit 3D Gaussian model through 3D Gaussian splattering technology, rendering images of multiple viewpoint angles, injecting viewpoint control information, and unconditional denoising results through similarity partial order loss constraints to enhance viewpoint consistency.

Benefits of technology

It significantly improves the geometric and semantic consistency of 3D models at different perspectives, reduces the frequency of "multi-faceted" problems, improves the generation speed and quality, and enhances the user experience and the application value of the model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119152101B_ABST
    Figure CN119152101B_ABST
Patent Text Reader

Abstract

The present invention proposes a zero-shot 3D generation method for enhancing view consistency, aiming to solve the problem of inconsistency presented by the same object from different viewpoints in current 3D generation technologies, namely the "multi-faceted" problem. This problem stems from the typical view preferences of pre-trained generation models. To this end, the present invention uses the view decoupling method VDM to extract view features to eliminate view prior preferences and enhance view control. At the same time, the similarity partial order loss PSL is introduced to optimize the similarity distribution of images between views, ensuring that the generated 3D images maintain a high degree of consistency from different viewpoints. In addition, this technology also combines the 3D Gaussian splash technology to further enhance the rendering effect and detail performance of the model. This solution significantly improves the authenticity and consistency of 3D content, making it more applicable to fields such as virtual reality, game design, and industrial design, and greatly promoting the practical application and development of zero-shot 3D generation technology.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of 3D generation, and particularly relates to a zero-shot 3D generation method for enhancing view consistency. Background Art

[0002] 3D generation technology plays a crucial role in many fields such as innovative industrial design, game design, and virtual reality. In particular, the progress of zero-shot text-to-3D generation technology provides more possibilities for innovation design from scratch, and also provides strong support for user interactive experience and simulation, making the transformation from concept to reality more efficient and intuitive. However, different from the small-span dimensionality increase in the text-to-image generation task, the inherent complexity of the real scene and the scarcity of 3D data make it a huge challenge to generate high-quality 3D content from text. Recently, the DreamFusion technology utilizes the prior knowledge of the mature text-to-image generation model, the Stable Diffusion (SD) model, and through the advanced Score Distillation Sampling (SDS) technology, upgrades the 2D results from the perspective dimension to the 3D world without relying on 3D datasets, achieving unsupervised 3D content generation. In subsequent research, the 3D Gaussian splashing technology is introduced to replace the traditional implicit 3D representation method, and this paradigm is optimized in terms of generation quality, generation speed, and geometric structure.

[0003] However, the existing text-to-3D generation methods are still largely limited by the problems of geometric collapse and view inconsistency. One significant problem is the "multi-faceted" problem (also known as the "multi-face" problem), which usually manifests as the same 3D object presenting different or even contradictory appearances from different perspectives, seriously affecting the realism of the 3D generation results. The emergence of this phenomenon is because a key factor has been overlooked: in order to improve the training efficiency, most existing text generation image models use normalized perspective images (typical perspective images, images in which objects or scenes are presented in common, standard angles) as training data. Such data usually selects common perspectives that can express object features to the greatest extent, such as the front view and the oblique front view. Therefore, in the absence of given specific perspective conditions, pre-trained text generation image models such as the Stable Diffusion (SD) model will generate images in the canonical perspective based on perspective prior knowledge. Recent text-to-3D generation methods optimize the parameterized 3D model by making the rendered images of each perspective of the target 3D object match the generation results of the Stable Diffusion (SD) model. However, when trying to add specific perspective descriptions to the original prompt to guide the model to generate images corresponding to the perspective, the inherent perspective prior knowledge of the Stable Diffusion (SD) model may conflict with the newly added perspective control information, resulting in the ambiguity of perspective semantics, thus causing the view Figure 1 consistency problem. Summary of the Invention

[0004] Objective of the Invention: The technical problem to be solved by the present invention is to provide a zero-shot 3D generation method for enhancing perspective consistency in view of the deficiencies of the prior art, including the following steps:

[0005] Step 1, conduct a mathematical analysis of the multi-faceted problem in the process of generating three-dimensional objects and propose an optimization objective;

[0006] Step 2, obtain the text prompt input by the user and select the topic words in the prompt. Connect the topic words with three different perspective descriptors (connect "front view", "side view", and "back view" to the end of the topic words respectively), construct three phrases containing the topic words with perspective descriptions, input the three phrases containing the topic words into the pre-trained text encoder CLIP respectively, obtain three encoded vectors, extract the encoded parts corresponding to the topic words from the three encoded vectors respectively, and then adopt the method of extracting the vertical component to extract the features of the front view, side view, and back view of the topic word object from the three topic word encodings respectively;

[0007] Step 3, input the user prompt into the pre-trained text-to-3D diffusion model PointE to generate rough three-dimensional point cloud data. At the same time, use the three-dimensional Gaussian splashing technology to fit the rough three-dimensional point cloud to generate an explicitly represented three-dimensional Gaussian model. Then, render the three-dimensional Gaussian model with randomly generated camera parameters to obtain renderings of more than two perspectives;

[0008] Step 4, input the user prompt into the pre-trained text encoder CLIP to obtain the prompt encoded vector, locate the position of the topic word in the encoded vector according to the topic word, and inject the features of the front view, side view, and back view extracted in Step 2 into the topic word encoded part in the user prompt encoded according to the randomly generated camera parameters in Step 3 to obtain the perspective-controlled prompt encoding corresponding to all camera parameters. Input the new prompt encoded vector and the renderings of more than two perspectives obtained in Step 3 Figure 1 into the pre-trained Stable Diffusion model SD;

[0009] Step 5, the Stable Diffusion model SD will generate corresponding two-dimensional images according to the user prompt and the rendered images, extract the unconditional prediction noise generated in the generation process, and obtain the unconditional denoising result. Then, explore the similarity distribution of the images and use the similarity partial order loss to constrain the unconditional denoising result, thereby enhancing the perspective consistency of the three-dimensional Gaussian model.

[0010] In Step 1, the "multi-faceted" problem includes:

[0011] Existing text-to-3D generation methods are still largely limited by the problems of geometric collapse and view inconsistency. One notable problem is the "multi-faceted" problem (also known as the "multi-face" problem), which typically manifests as the same 3D object presenting different or even conflicting appearances from different viewpoints, seriously affecting the realism of the 3D generation results. This phenomenon occurs because a key factor has been overlooked: in order to improve training efficiency, most existing text-to-image generation models use normalized view images (typical view images, where objects or scenes are presented from common, standard angles) as training data. Such data usually selects common viewpoints that can most effectively express object features, such as the front-facing view and the oblique front-facing view. Therefore, in the absence of given specific view conditions, pre-trained text-to-image generation models such as the StableDiffusion model SD will generate images in the canonical view based on view prior knowledge. Recent text-to-3D generation methods optimize parameterized 3D models by matching the rendered images of each view of the target 3D object with the generation results of the StableDiffusion model SD. However, when attempting to add specific view descriptions to the original prompt to guide the model to generate images from the corresponding view, the inherent view prior knowledge of the StableDiffusion model SD may conflict with the newly added view control information, resulting in ambiguity in view semantics and thus triggering multi-view Figure 1 consistency issues.

[0012] In step 1, the mathematical level analysis includes:

[0013] Diffusion models are a class of generative models used to learn and sample from complex distributions. They originate from stochastic processes in statistical physics. The core idea is to gradually "diffuse" data from its original form to random noise and then reverse the process to restore the noise to data through a trained neural network to denoise noise samples. Subsequently, the concept of the denoising diffusion probabilistic model DDPM was proposed, and the denoising diffusion probabilistic model DDPM was used to optimize the following simplified objective:

[0014]

[0015] where, represents the loss function of the denoising diffusion probabilistic model DDPM, which is used to measure the difference between the samples generated by the model and the actual samples. By minimizing the loss function , the denoising diffusion probabilistic model DDPM can gradually learn how to denoise and generate high-quality samples. The expected value represents the calculation of the expected value for the original sample data , noise and time step , where the original sample data is used as the input data of the denoising diffusion probabilistic model DDPM, and the noise denotes random noise that follows a standard normal distribution with a mean of 0 and a variance of 1. During the training process, this noise is added to the samples at time step denotes the diffusion time step, which is used to control the degree of noise addition. In the diffusion model, the time steps range from 1 to and increase gradually. denotes the denoising diffusion probability model DDPM at model parameters Under the condition, according to the noise perturbation time step the sample and time step the predicted noise value; the added noise and the model-predicted noise the square of the two-norm between them measures the difference between the predicted noise and the actual added noise. During the inference process, the denoising diffusion probability model DDPM starts from and samples the previous sample from a normal distribution with a probability density of . The denoising diffusion probability model DDPM highlights the excellent ability of the diffusion model in capturing and simulating real-world image data. This advantage has given rise to a series of innovations and improvements, including the Stable Diffusion model SD. The Stable Diffusion model SD is based on the Latent Diffusion model LDM and highlights two key advantages in design and implementation. First, by performing the diffusion process in a lower-dimensional latent space, the efficiency and speed of image generation are significantly improved, and the computational burden is greatly reduced. Second, by introducing a conditional mechanism, the Latent Diffusion model LDM uses the advanced text encoder CLIP to encode text conditional information into the latent space, enabling the model to efficiently generate highly relevant and realistic images according to specific conditional information. The Latent Diffusion model LDM optimizes the following simplified objective:

[0016]

[0017] where is the loss function of the Latent Diffusion model LDM, and the expected value denotes the expected value calculation for the latent encoding result of the original sample data , the user prompt , the random noise that follows a standard normal distribution, and the time step ; the conditional information encoding denotes the result of encoding the user prompt through the text encoder CLIP, denotes the two-norm of the difference between the added noise and the predicted noise. The method of the present invention is proposed based on the Stable Diffusion model SD and the Latent Diffusion model LDM.

[0018] Under the framework of the denoising diffusion probabilistic model (DDPM) and the latent diffusion model (LDM), the generation of an image can be regarded as a reverse recovery process from the final noise state to the noise-free initial clear state , which is precisely controlled by the conditioning of the user prompt . Specifically, the reverse recovery process involves gradually recovering from a high-noise state to the noise-free original image state, and the probability density of generating the final noise-free initial clear state is expressed by the following formula : :

[0019]

[0020] where represents the maximum number of time steps that can take, and represents the probability distribution of the final state of the diffusion process, which is a standard normal distribution. represents the conditional probability of generating the previous latent representation under the condition of the given user prompt and the current latent representation . The symbol is the differential element in the integral, indicating that in the integration process, a small change is made to each possible state, and through these differential elements, the joint probability distribution of all possible intermediate states is calculated in the process from the initial state to the final state. In the zero-shot text-to-3D generation task, the model will integrate the two-dimensional image representations generated from each perspective in the denoising diffusion probabilistic model (DDPM) process through multiple iterative steps, and can construct the unnormalized probability density function of the three-dimensional parameter

[0021]

[0022] where represents the set of all perspectives, and this expression aggregates each perspective selected from the perspective set The results of generating images through the Denoising Diffusion Probability Model (DDPM) process reveal the contribution of generated images from different perspectives to the overall 3D model parameter estimation. Although the DDPM method performs excellently in generating high-quality 2D images, the randomness in its reverse process and sensitivity to input fluctuations may lead to inconsistent image features generated from different perspectives, thus affecting the quality of the overall 3D model. Specifically, during the training or distillation phase of the model, the pre-trained DDPM may tend to produce low-quality and inconsistent pseudo-ground truths, which is particularly prominent in the multi-view fusion process, resulting in over-smoothed and detail-lacking 3D reconstruction results. To address this challenge, many existing works have introduced the Denoising Diffusion Implicit Model (DDIM) method, which significantly reduces the random variation in each iteration step through a more deterministic recursive process. Although the DDIM effectively improves the quality and consistency of pseudo-ground truths, the "multi-faceted" problem remains significant, so the focus is on the multi-view fitting process outside the diffusion process. The method of the present invention redefines the 3D parameters of the density function as the product of the conditional likelihoods of each iteration step under the perspective control given a series of optimization steps , user prompts and the 3D Gaussian model projection at the corresponding perspective, expressed as:

[0023]

[0024] where represents the maximum number of optimization steps that can take, and represents the 3D model state controlled by the 3D parameters at the iteration step . The user prompt should ideally only contain guidance information related to the generated content and not directly involve perspective information. However, it has been found in the research that the pre-trained diffusion model automatically introduces perspective prior knowledge when parsing the user prompt due to exposure to normalized perspective training data, even though does not explicitly contain perspective information. To precisely control and understand this phenomenon theoretically, from the perspective of how the stable diffusion model understands user prompts, the user prompt is divided into a content part and a perspective prior part , and the content part and the perspective prior part The model-based training experience is a typical perspective preference introduced by the model based on pre-training data when parsing user input. The three-dimensional parameters of the density function The new expression is:

[0025]

[0026] where denotes being constrained by, , this condition indicates that the content part and the perspective prior part do not share any components in the direction in the vector space, ensuring that the generation of content is completely independent of the influence of any specific perspective. Further, taking the logarithm on both sides of the equation gives:

[0027]

[0028] Next, using the chain rule, for the gradient of is expressed as:

[0029]

[0030] where, denotes the partial derivative of the 3D model state with respect to the three-dimensional parameters , denotes the partial derivative of the log probability density with respect to the 3D model state ;

[0031] For it is further expanded using Bayes' theorem:

[0032]

[0033] where, the term is regarded as a constant with respect to the 3D model state , so the result after taking the partial derivative is 0 and this term can be deleted. In the expression the term intuitively reflects the model's instinctive perception of 3D objects in the absence of external perspective information guidance. However, this perception is particularly vulnerable to the influence of view prior knowledge, resulting in the model relying on the most common or prominent perspective features in past experience when not constrained by external perspectives. This reliance can lead to biases in the model when dealing with new perspectives, that is, when facing situations significantly different from the perspectives in the training data, the generated results deviate from the true 3D representation. Additionally, the second term in the expression The point - conditional mutual information (PCMI) can be further extended. PCMI provides a quantitative approach, indicating how the information interaction between different variables in a specific state goes beyond the simple combination of their individual behaviors. It is further decomposed into:

[0034]

[0035] Among them, the view control and the view prior part of the point - conditional mutual information , and can be regarded as constants in the current optimization step. Therefore, the term can be further simplified and expanded using the definition of conditional probability:

[0036]

[0037] This term measures the additional information between and and relative to their independent existence, given . If the view control conflicts with the view prior part in the pre - trained model, that is, << , then for each term, the point - conditional mutual information (PCMI) term approaches 0. Then, the term and the and terms will have an adverse impact on 3D generation simultaneously. Therefore, establishing the view semantic consistency between

[0038] is particularly crucial for alleviating the "multi - face" problem.

[0039] Step 2 includes: .

[0040] Obtain the user - input text prompt and select the topic word in the prompt. For example, in the prompt "a photo of a car", the topic word is "car". Input the topic word directly into the pre - trained CLIP encoder to obtain the encoded vector of the topic word without view description.

[0041] Connect the topic word with the descriptive words of three different views, that is, connect the front - view, side - view, and rear - view descriptions respectively at the end of the topic word to construct three phrases with view descriptions containing the topic word. For example, after connecting the view - descriptive phrases to the topic word "car", it becomes: "car, front - view", "car, side - view", and "car, rear - view".Input three phrases containing the subject word with perspective descriptions into the CLIP encoder respectively and locate the position of the subject word encoding to obtain the subject word encoding containing the corresponding perspective description , different from , the subject word encoding containing the corresponding perspective description contains additional context information, which provides a way to extract the perspective feature encoding unrelated to the content in . This process uses the method of extracting the vertical component and is formulated as:

[0042]

[0043] where , represents the projection part of the subject word encoding containing the corresponding perspective description on the encoding vector of the subject word without perspective description . The perspective feature encoding unrelated to the content obtained by subtracting this projection component represents the refined pure perspective difference. This difference can be applied to the encoding process of other prompt words, thereby realizing perspective control. When generating a 2D image, if there is no perspective-related description in the prompt word, the generated result is often a normalized perspective result. To make the generated view more in line with the diverse needs of users, first decouple the prior perspective features of the subject word encoding in the user prompt, eliminate typical view preferences, and inject perspective control information on this basis. Taking the "back view" as an example, the "back view" feature will be injected after decoupling the "front view" and "side view" features, and the formula is:

[0044]

[0045] where represents the perspective for which prior knowledge needs to be eliminated, including the front and side cases represents the encoding vector of the subject word without perspective description projected in the direction of the perspective feature encoding unrelated to the content , represents the perspective feature encoding of the back view unrelated to the content is a scale coefficient representing the injection intensity of the new perspective

[0046] Step 3 includes:

[0047] Input the user prompt into the pre-trained text-to-3D diffusion model PointE to generate rough 3D point cloud data, and the rough 3D point cloud data contains the basic geometric information of the object surface

[0048] An anisotropic Gaussian volume is a three-dimensional Gaussian distribution that allows for different variances in different directions, i.e., the shape of the distribution can be stretched or scaled along each direction. This property enables it to more flexibly adapt to the local structure of point cloud data. By fitting an anisotropic Gaussian volume, a Gaussian volume that matches the local distribution of the point cloud can be generated in each local region of the point cloud, forming a continuous and smooth representation. This fitting process takes into account the density distribution and local geometric structure of the point cloud, thus more accurately capturing the shape details of the object. Using anisotropic Gaussian volumes to capture the geometric features of the rough three-dimensional point cloud data generated by the three-dimensional diffusion model PointE, and these Gaussian volumes directly define the geometric structure of the rough three-dimensional point cloud data generated by the three-dimensional diffusion model PointE through information such as position, direction, and size, which is an explicit representation. This process fits an explicitly represented three-dimensional Gaussian model;

[0049] The tile-based rasterization technology optimized by the Graphics Processing Unit (GPU) is a technology that efficiently renders a 3D scene into a 2D image by dividing the image into small blocks and utilizing the parallel processing capabilities of the Graphics Processing Unit (GPU). Render the three-dimensional Gaussian model using the tile-based rasterization technology optimized by the Graphics Processing Unit (GPU) with randomly generated camera parameters to obtain rendered images from multiple perspectives.

[0050] Step 4 includes:

[0051] Input the user prompt into the pre-trained text encoder CLIP to obtain the prompt encoding vector, and locate the position of the topic word in the encoding vector according to the topic word.

[0052] Obtain the multiple randomly generated camera parameters in Step 3. The view control process can be adaptively implemented through the camera parameters of each optimization step, specifically including: when the azimuth angle in the camera parameters is within ( ):

[0053]

[0054] where, represents the encoding vector without view description of the topic word Projection in the front view feature encoding unrelated to the content direction;

[0055] When the azimuth angle is within ( ): :

[0056]

[0057] where, the weight adjusted according to the azimuth angle in the camera parameters reflects The injected intensity, when the azimuth angle in the camera parameters is closer to 180° or -180°, that is, when the current viewing angle is closer to the front or back, the weight is larger. Through this perspective control at the encoding level, the Stable Diffusion model SD can more accurately understand the perspective control specified by the rendering camera without being interfered by the common perspective preferences in the Stable Diffusion model SD, making the view prior part more consistent with the perspective control controlled by the camera parameters and obtaining a new prompt encoding after perspective control.

[0058] Next, the new prompt encoding vector and the multiple perspective renderings obtained in step 3 Figure 1 are input into the pre-trained Stable Diffusion model SD.

[0059] Step 5 includes:

[0060] During the training process of the Stable Diffusion model SD, the data used is the pairing of text and images. However, the description of the perspective in the text is usually relatively rough and general, lacking the necessary numericalization and precision. This vague perspective information limits the Stable Diffusion model SD's ability to establish a sound perspective perception during learning, and thus makes it face significant difficulties in converting from 2D images to 3D content. To alleviate this problem, in step 4, by improving the perspective semantic clarity of the user prompt during the image generation process, ensuring the alignment of the user prompt and the generated object in terms of perspective semantics, thus effectively alleviating the dependence of the 3D generated content on the prior perspective knowledge. But the mathematical analysis of the "multi-faceted" problem in step 1 shows that another gradient term as an unconditional term does not interact with the user prompt, and its influence on the perspective of the generated content only depends on the view prior knowledge of the Stable Diffusion model SD. Therefore, this term also potentially injects prior perspective preferences into the generated content. To eliminate this preference, the goal is to find a method that can establish a connection between the unconditional guidance term and the perspective control, or that can make the unconditional term be controlled by the camera parameters, injecting perspective perception ability into the model.

[0061] The contrastive language-image pre-training model is a machine learning model specialized for measuring the similarity between images and texts, obtaining the cosine similarity score between images and texts. Starting from the distribution of cosine similarity scores between the perspective description of the text (such as "back", "side", etc.) and the rendered image, the relationship between perspective control and the rendered image is explored. Uniform sampling is performed on two groups of results with and without the "multi-faceted" problem respectively. The sampling method is to sample 500 evenly spaced camera parameters from the upper hemisphere with a fixed radius. All these camera parameters point to the center of the sphere at the same height, and 500 images are rendered from the 3D scene. The cosine similarity between each image and the perspective description text is calculated through the pre-trained contrastive language-image pre-training model. It is found that the distribution result of the cosine similarity scores of the samples without the "multi-faceted" problem changes approximately periodically with the change of the angle, and this change is approximately continuous.

[0062] Furthermore, according to the characteristics of the unconditional term, the cosine similarity relationship between images and images is explored to obtain the cosine similarity distribution law; first, a reference image is randomly selected from all the rendered images, and the cosine similarity between other rendered images and the reference image is calculated. Considering the random fluctuations of the pitch angle and the field of view size during the optimization process and the flipping operation of the images, more global features need to be obtained. The first-layer vector in the pre-trained encoder ViT is an ideal choice because the first-layer vector in the encoder ViT not only retains the detailed information of each small block of the image but also effectively integrates the overall layout and relationship of the entire image, making the feature representation have a comprehensive global perspective. Through experiments, it is found that the cosine similarity distribution calculated from all the rendered images and the reference image is indeed related to the azimuth angle of the camera parameters. The rendered image corresponding to the azimuth angle closer to the azimuth angle of the reference image has a higher cosine similarity score with the reference image, and the score gradually decreases from the azimuth angle of the reference image to both sides. However, when the "multi-faceted" problem occurs, this cosine similarity distribution is disrupted, which provides a way to explicitly constrain the unconditional denoising result away from the "multi-faceted" problem. Specifically, in each round of optimization steps, a number of perspective controls are randomly generated, and an order is determined for these perspective controls according to the azimuth angle in the , calculate the distance from the points corresponding to other perspective controls to the secant line distance , according to size of camera controls are sorted.

[0063] Next, through these perspective controls obtain the rendered image , rendered image latent representation of , and the noisy latent representation of the rendered image , and obtain the latent representation of the rendered image after passing through the Stable Diffusion model SD, the corresponding unconditional noise prediction result , , represents the Stable Diffusion model SD estimating the noise component after noise addition processing.

[0064] Then, is removed from the noisy latent representation to obtain the unconditional denoising result , and calculate the partial order loss between the unconditional denoising results , and the loss function is :

[0065]

[0066] Among them, represents the number of perspective controls in each iteration, represents the index position among perspective controls, represents the corresponding unconditional denoising result of the first-layer encoding result of the rendered image in the ViT encoder, represents the cosine similarity between This term will penalize the unconditional denoising results that do not satisfy the cosine similarity regular distribution, making them approach the cosine similarity distribution without the "multi-faceted" problem. This partial order loss can repair multi-faceted objects in the initial stage of iterative optimization, thus effectively alleviating the "multi-faceted" problem.

[0067] The present invention also provides a zero-shot 3D generation device for enhancing perspective consistency implemented based on the above method, including:

[0068] A text input module for obtaining the text prompt words input by the user;

[0069] A point cloud generation module for generating a rough 3D point cloud according to text prompts;

[0070] A 3D fitting module for fitting the rough point cloud using 3D Gaussian splashing technology to generate an explicit 3D Gaussian model;

[0071] A multi-view rendering module for rendering the 3D Gaussian model from multiple perspectives to obtain multi-view rendered images;

[0072] A view encoding control module for eliminating typical view preferences and injecting view control to improve the view semantic consistency from text to 2D and then to 3D;

[0073] A text-to-image generation module for adding noise to the multi-view rendered images, predicting the noise according to user prompts and the multi-view rendered images, and updating the 3D Gaussian model parameters according to the prediction results;

[0074] A similarity partial order loss constraint module for establishing view consistency between multi-view images and improving the view perception ability of the 3D model.

[0075] A progress tracking module for monitoring the training progress and recording and displaying various metrics during the training process;

[0076] A model checkpoint management module responsible for saving and loading the model state to facilitate the continuity and recovery of model training;

[0077] A network interface module for handling network communication, receiving external commands, and sending rendered images;

[0078] A video generation module for generating video files based on the rendered images.

[0079] The encoding control module includes:

[0080] A view feature decoupling unit for extracting view feature differences from text prompts;

[0081] A view control unit for using the view feature differences after feature decoupling to eliminate typical view preferences and inject new view control to enhance view semantic clarity;

[0082] The text-to-image generation module includes:

[0083] An image encoding unit for encoding the multi-view images to obtain a latent representation;

[0084] An image noise addition unit for adding noise to the latent representation of the multi-view images;

[0085] An image denoising unit for performing conditional denoising operations on the noisy image according to user prompts, and also performing unconditional denoising operations on the noisy image;

[0086] A noise comparison unit for comparing the predicted noise with the noisy part and calculating the gradient of the noise tensor;

[0087] A parameter update unit for updating the 3D Gaussian model parameters through the backpropagation algorithm according to the noise tensor gradient;

[0088] The similarity partial order loss constraint module includes:

[0089] A camera parameter sorting unit for sorting the randomly generated camera parameters according to the ideal cosine similarity distribution law;

[0090] An unconditional denoising result acquisition unit for selecting the unconditional predicted noise in the image denoising unit according to the sorted camera parameters and obtaining the unconditional denoising result according to the unconditional predicted noise;

[0091] A partial order loss calculation unit for calculating the cosine similarity between the unconditional denoising results and calculating the partial order loss according to the ideal cosine similarity distribution law. This loss will update the 3D Gaussian model parameters together with the parameter update unit during backpropagation.

[0092] The present invention also provides an electronic device, including a processor and a storage medium;

[0093] The storage medium is used for storing instructions;

[0094] The processor is used for executing the steps of the method according to the instructions.

[0095] The present invention also provides a computer-readable storage medium, on which a computer program is stored, and when the program is executed by a processor, the steps of the method are implemented.

[0096] The present invention significantly improves the 3D consistency of the generated target from two aspects. On the one hand, aiming at the conflicts brought by the view prior, a view decoupling method VDM is proposed. The view decoupling method VDM is a view feature decoupling method, which effectively eliminates the view semantic ambiguity caused by typical view preferences and enhances view control when generating 3D content, significantly strengthening the view semantic consistency in the process from text to image to 3D. On the other hand, aiming at the lack of explicit 3D consistency constraints, the present invention explores the distribution law of image similarity and proposes a similarity partial order loss PSL to enhance the view perception ability in the process of promoting from 2D to 3D. Facts have proved that the method of the present invention constructs a strong view semantic association in different modal data and effectively alleviates the "multi-faceted" problem.

[0097] Beneficial effects: A zero-shot 3D generation method for enhancing view consistency provided by the present invention addresses the common multi-view consistency problem in traditional methods, namely the "multi-faceted" problem, and proposes systematic optimization measures, significantly improving the geometric and semantic consistency of 3D models from different views. By introducing the view decoupling method VDM and the similarity partial order loss PSL, the present invention not only enhances the model's view perception ability, but also effectively solves the generation conflicts and view ambiguity problems caused by view priors through fine-grained view control. In addition, the present invention uses 3D Gaussian splashing technology to improve the rendering efficiency and the expressiveness of the model, allowing real-time and efficient processing of large-scale and dynamic scenes, which is particularly important for real-time applications such as virtual reality and game design. The explicit 3D representation combined with high-quality view consistency ensures coherence and authenticity when observed from different views, greatly enhancing the user experience and the application value of the model. Through the comprehensive application of these technologies, the present invention significantly reduces the occurrence frequency of the "multi-faceted" problem and improves the generation speed and quality of 3D content. This method is not only applicable to single scenes, but also can adapt to variable environments and complex scene requirements, providing broader practicality and flexibility for 3D content creation. Therefore, the present invention has broad application prospects in the fields of industrial design, virtual reality, game design, etc., and is expected to promote the development of 3D generation technology and bring innovation and value to related industries. Brief Description of the Drawings

[0098] Figure 1 is the training flowchart of the zero-shot 3D generation method for enhancing view consistency provided in Embodiment 1 of the present invention.

[0099] Figure 2 is the model framework diagram of the zero-shot 3D generation method for enhancing view consistency provided in Embodiment 1 of the present invention. Detailed Embodiments

[0100] The following further detailed description of the present invention is made in conjunction with the accompanying drawings and specific embodiments, and the above and / or other advantages of the present invention will become clearer.

[0101] As Figure 1As shown in the figure, an embodiment of the present invention provides a zero-shot 3D generation method for enhancing perspective consistency. First, the training process is executed: at the start of training, the environment and the 3D Gaussian model are initialized. The 3D Gaussian model is used to represent 3D objects. This step mainly prepares the environment and the model structure for subsequent training. During rendering, the objects in the 3D Gaussian model are projected onto the imaging plane of the camera, and the projected two-dimensional Gaussian distribution is assigned to each tile. Then, the checkpoint is used to restore the previously trained state. If a checkpoint file is provided, the program will load the previous model parameters and state to continue the interrupted training. Otherwise, training will start from scratch. Then, a pre-trained diffusion model (Stable Diffusion model SD) and text encoding are set up to guide the training of the model. The text encoding is a vector representation extracted by using the pre-trained text encoder CLIP with the topic words selected from the received user prompt words. These encodings will be used in the subsequent visual feature decoupling process. Next, the user prompt is input into the pre-trained three-dimensional diffusion model to generate a rough 3D point cloud. Then, it is checked whether all the specified training iterations have been completed. If the number of iterations reaches the preset maximum value, the training will stop. Otherwise, the next iteration is entered. Next, the learning rate is dynamically adjusted, and the learning rate is gradually decreased during training to achieve stable convergence. Then, the camera parameters are configured, and parameters such as position, perspective, and focal length are set. These parameters are used to select the current optimized view. Then, the scene is rendered using the currently set camera parameters to generate a set of images, which will be used to train the 3D Gaussian model. Then, the prompt text is encoded using the pre-trained text encoder CLIP, and perspective feature decoupling is performed through the prepared topic word encoding to eliminate typical perspective preferences and inject perspective control. The rendered images are encoded by the pre-trained image encoder VAE. Next, the text encoding and the rendered image encoding are input into the Stable Diffusion model SD for noise prediction. The obtained unconditional noise prediction and text-guided conditional noise prediction are weighted and aggregated, and the similarity partial order loss PSL, prediction noise loss, scale loss, and total variance loss are calculated using the unconditional noise. Then, the 3D Gaussian model parameters are updated through backpropagation, and processor events are recorded to monitor the training time and resource consumption. It is also checked whether the current iteration meets the conditions for saving the model (such as the specified saving interval). If it meets the conditions, the current model state and parameters are saved. Finally, the current training state, model parameters, and logs are saved to disk for subsequent continued training or evaluation. After all the training steps are completed, the entire process ends.

[0102] As Figure 2 shown, the network framework of this embodiment includes: First, the user input text prompt is obtained, and the topic words in the prompt are selected. For example, in the prompt "A squirrel is eating a hamburger", the topic word is "squirrel". The topic word is directly input into the pre-trained CLIP encoder to obtain the encoding vector of the topic word without perspective description. Connect the subject term with descriptors from three different perspectives, that is, connect "front perspective", "side perspective", and "back perspective" at the end of the subject term respectively, to construct three phrases containing the subject term with perspective descriptions. For example, when the subject term "squirrel" is connected with the perspective description phrases, it becomes: "squirrel, front perspective", "squirrel, side perspective", and "squirrel, back perspective".

[0103] Input the three phrases containing the subject term with perspective descriptions into the CLIP encoder respectively and locate the position of the subject term encoding to obtain the subject term encoding containing the corresponding perspective description , different from , the subject term encoding containing the corresponding perspective description contains additional context information, which provides a way to extract the perspective feature encoding irrelevant to the content in , and this process adopts the method of extracting the vertical component, which can be formulated as:

[0104]

[0105] where , represents the projection part of the subject term encoding containing the corresponding perspective description on the encoding vector of the subject term without perspective description , and the perspective feature encoding irrelevant to the content obtained by subtracting this projection component It represents the pure perspective difference extracted, and this difference value can be applied to the encoding process of other prompt words, thereby realizing perspective control. The user's prompt words are input into the pre-trained text-to-3D diffusion model PointE to generate rough 3D point cloud data, which contains the basic geometric information of the object's surface. The anisotropic Gaussian volume is a three-dimensional Gaussian distribution that allows different variances in different directions, that is, the shape of the distribution can be stretched or scaled along each direction. This property enables it to more flexibly adapt to the local structure of the point cloud data. By fitting the anisotropic Gaussian volume, a Gaussian volume matching the local distribution of the point cloud can be generated in each local area of the point cloud, forming a continuous and smooth representation. This fitting process takes into account the density distribution and local geometric structure of the point cloud, thus more accurately capturing the shape details of the object. The anisotropic Gaussian volume is used to capture the geometric features of the rough 3D point cloud data generated by the 3D diffusion model PointE, and these Gaussian volumes directly define the geometric structure of the rough 3D point cloud data generated by the 3D diffusion model PointE through information such as position, direction, and size, which is an explicit representation. This process fits an explicit 3D Gaussian model. The tile-based rasterization technology optimized by the Graphics Processing Unit (GPU) is a technology that efficiently renders a 3D scene into a 2D image by dividing the image into small blocks and utilizing the parallel processing ability of the Graphics Processing Unit (GPU). The randomly generated camera parameters are used to render the 3D Gaussian model by the tile-based rasterization technology optimized by the Graphics Processing Unit (GPU) to obtain rendered images from multiple perspectives. The user's prompt words are input into the pre-trained text encoder CLIP to obtain the prompt word encoding vector, and the position of the topic word in the encoding vector is located according to the topic word.

[0106] The perspective control process can be adaptively realized through the camera parameters at each optimization step. Specifically, when the azimuth angle in the camera parameters is within ( ):

[0107]

[0108] Among them, represents the encoding vector of the topic word without perspective description projected in the front view feature encoding direction that is irrelevant to the content.

[0109] When the azimuth angle is within ( ), then:

[0110]

[0111] Among them, the weight adjusted according to the azimuth angle in the camera parameters The injected intensity, when the azimuth angle in the camera parameters is closer to 180° or -180°, that is, when the current viewing angle is closer to the front or back, the weight is larger. Through this view control at the coding level, the Stable Diffusion model SD can more accurately understand the view control specified by the rendering camera without being interfered by the common view preferences in the Stable Diffusion model SD, making the view prior part more consistent with the view control controlled by the camera parameters . In this step, a new prompt word encoding after view control will be obtained. Then, the new prompt word encoding vector and the rendering Figure 1 are input into the pre-trained Stable Diffusion model SD.

[0112] Furthermore, explore the cosine similarity relationship between images according to the characteristics of the unconditional term noise. First, randomly select one of all the rendered images as the reference image, and calculate the cosine similarity between other rendered images and the reference image. Considering the random fluctuations of the pitch angle and the field of view size during the optimization process and the flipping operation of the images, more global features need to be obtained. The first-layer vector in the pre-trained encoder ViT is an ideal choice because the first-layer vector in the encoder ViT not only retains the detailed information of each small block of the image but also effectively integrates the overall layout and relationship of the entire image, making the feature representation have a comprehensive global perspective. Through experiments, it is found that the cosine similarity distribution calculated between all the rendered images and the reference image is indeed related to the azimuth angle of the camera parameters. The rendered image corresponding to the azimuth angle closer to the reference image has a higher cosine similarity score with the reference image, and the scores gradually decrease from the azimuth angle of the reference image to both sides. However, when the "multi-faceted" problem occurs, this cosine similarity distribution is disrupted, which provides a way to explicitly constrain the unconditional denoising result away from the "multi-faceted" problem. Specifically, in each round of optimization steps, randomly generate view controls , and determine an order for these view controls according to the azimuth angle in the view control , and this order follows the above cosine similarity distribution rule. Establish a rectangular coordinate system and construct a unit circle with the origin as the center. The angles corresponding to the points on the unit circle start from 0° in the negative direction of the vertical axis and gradually increase to 180° in the positive direction of the horizontal axis and decrease to -180° in the negative direction of the horizontal axis. Place the randomly generated view controls on the unit circle according to the azimuth angle size. Then randomly select one of the view controls as the reference camera, and construct a secant line parallel to the horizontal axis with this point as the starting point , calculate the distance from the points corresponding to other cameras to the secant line , and according to the size of Sort the camera parameters from smallest to largest. Then, control through these viewpoints to obtain the rendered image and the rendered image 's latent representation as well as the noise-added latent representation of the rendered image, and obtain the latent representation of the rendered image After passing through the Stable Diffusion model SD, the corresponding unconditional noise prediction result is obtained. , indicating that the model estimates the noise component after noise addition. Then, is removed from the noise-added latent representation to obtain the unconditional denoising result and calculate the partial order loss between the unconditional denoising results . The loss function is as follows:

[0113]

[0114] where represents the number of viewpoint controls in each iteration , represents the -th index position among the rendered images corresponding unconditional denoising results at the first-layer encoding result of the ViT encoder, , indicating the cosine similarity between The

[0115] term will penalize the unconditional denoising results that do not satisfy the cosine similarity regular distribution, making them converge to the cosine similarity distribution without the "multi-faceted" problem. This partial order loss can repair multi-faceted objects at the initial stage of iterative optimization, thus effectively alleviating the "multi-faceted" problem. Finally, update the model parameters of the 3D Gaussian model through the backpropagation algorithm to improve the quality of the 3D generated objects. Another embodiment of the present invention provides a zero-shot 3D generation device for enhancing viewpoint consistency, including:

[0116] A text input module for obtaining the text prompt words input by the user;

[0117] A point cloud generation module for generating a rough 3D point cloud according to the text prompt words;

[0118] A 3D fitting module, which is used to fit the rough point cloud by using the 3D Gaussian splashing technology to generate an explicit 3D Gaussian model;

[0119] A multi-view rendering module, which is used to render the 3D Gaussian model from multiple perspectives to obtain multi-view rendering images;

[0120] A perspective encoding control module, which is used to eliminate the typical perspective preference and inject perspective control to improve the perspective semantic consistency from text to 2D and then to 3D;

[0121] A text-to-image generation module, which is used to add noise to the multi-view rendering images, predict the noise according to the user prompt and the multi-view rendering images, and update the 3D Gaussian model parameters according to the prediction results;

[0122] A similarity partial order loss constraint module, which is used to establish the perspective consistency between multi-view images and improve the perspective perception ability of the 3D model.

[0123] A progress tracking module, which monitors the training progress and records and displays various metrics during the training process;

[0124] A model checkpoint management module, which is responsible for saving and loading the model state to facilitate the continuity and recovery of model training;

[0125] A network interface module, which processes network communication, receives external commands and sends rendering images;

[0126] A video generation module, which generates video files according to the rendered images.

[0127] The encoding control module includes:

[0128] A perspective feature decoupling unit, which is used to extract the perspective feature difference from the text prompt;

[0129] A perspective control unit, which is used to eliminate the typical perspective preference and inject new perspective control by using the perspective feature difference after feature decoupling to enhance the perspective semantic clarity;

[0130] The text-to-image generation module includes:

[0131] An image encoding unit, which is used to encode the multi-view images to obtain a latent representation;

[0132] An image noise addition unit, which is used to add noise to the latent representation of the multi-view images;

[0133] An image denoising unit, which is used to perform conditional denoising operations on the noise-added images according to the user prompt and also perform unconditional denoising operations on the noise-added images;

[0134] A noise comparison unit, configured to compare the predicted noise with the noise-added part and calculate the gradient of the noise tensor;

[0135] A parameter update unit, configured to update the 3D Gaussian model parameters through the backpropagation algorithm according to the noise tensor gradient;

[0136] The similarity partial order loss constraint module includes:

[0137] A camera parameter sorting unit, configured to sort the randomly generated camera parameters according to the ideal cosine similarity distribution rule;

[0138] An unconditional denoising result obtaining unit, configured to select the unconditional predicted noise in the image denoising unit according to the sorted camera parameters, and obtain the unconditional denoising result according to the unconditional predicted noise;

[0139] A partial order loss calculation unit, configured to calculate the cosine similarity between the unconditional denoising results, and calculate the partial order loss according to the ideal cosine similarity distribution rule. This loss will be used together with the parameter update unit to update the 3D Gaussian model parameters during backpropagation.

[0140] Another embodiment of the present invention provides an electronic device, including a processor and a storage medium;

[0141] The processor is configured to operate according to instructions to execute the steps of the method.

[0142] Another embodiment of the present invention provides a computer-readable storage medium, on which a computer program is stored. When the program is executed by a processor, the steps of the method are implemented.

[0143] The method of the embodiment of the present invention mainly solves the problem of multi-view consistency that commonly exists in the prior art, namely the "multi-faceted" problem. This problem often manifests as inconsistent images generated for the same 3D object from different perspectives, seriously affecting the authenticity and visual effect of the generation result. To solve this problem, the embodiment of the present invention introduces the 3D Gaussian splashing technology, the view decoupling method VDM, and the similarity partial order loss PSL to form a systematic optimization process. First, an input text prompt describing the 3D object to be generated is provided, and the model generates a rough 3D point cloud based on this prompt. Although this initially generated point cloud captures the basic shape of the object, it lacks in details and consistency. To improve the generation accuracy and consistency, the present invention uses the 3D Gaussian splashing technology to fit this rough point cloud and transform it into an explicit 3D representation composed of anisotropic Gaussian bodies. Next, an iterative optimization strategy of 5000 steps is adopted, and each iteration process includes the following key steps: First, select a certain view of the current 3D model and render the projected image of this view; then, input the text prompt, the rendered image, and the camera parameters into the pre-trained diffusion model, and update the explicit parameters of the 3D model through the generated 2D image. Through multiple iterations of optimization, the consistency of the 3D model from different perspectives is continuously enhanced. To further alleviate the multi-faceted problem, the present invention introduces the view decoupling method VDM and the similarity partial order loss PSL. The view decoupling method VDM is a view feature decoupling method, which effectively eliminates the view semantic ambiguity caused by typical view preferences and enhances view control when generating 3D content, significantly strengthening the view semantic consistency in the process from text to image and then to 3D. On the other hand, aiming at the lack of explicit 3D consistency constraints, the present invention explores the distribution law of image similarity and proposes the similarity partial order loss PSL to enhance the view perception ability in the process of promoting from 2D to 3D. Experimental results show that this method significantly reduces the occurrence frequency of the "multi-faceted" problem during the generation process, and the generated 3D model has higher authenticity and consistency from multiple perspectives. Compared with traditional methods, the present invention not only improves the generation speed and quality, but also shows significant advantages in solving the multi-view consistency problem, and can better meet the needs of high-quality 3D content in fields such as industrial design, virtual reality, and game design.

[0144] The present invention provides a zero-shot 3D generation method for enhancing view consistency. There are many methods and ways to specifically implement this technical solution. The above description is only the preferred embodiment of the present invention. It should be noted that for those of ordinary skill in the art, without departing from the principle of the present invention, several improvements and refinements can be made, and these improvements and refinements should also be regarded as the protection scope of the present invention. Each component not clearly defined in this embodiment can be implemented using existing technologies.

Claims

1. A zero-sample 3D generation method for enhancing perspective consistency, characterized in that: The following steps are involved: Step 1: mathematically analyze the multifaceted problems in the process of generating three-dimensional objects and propose optimization goals; Step 2, obtain the text prompt word input by the user and select the theme word in the prompt word, connect the theme word to three different perspective description words: connect the front perspective, side perspective and back perspective to the end of the theme word, construct three phrases containing the theme word with perspective description, input the three phrases containing the theme word into the pre-trained text encoder CLIP respectively, obtain three encoding vectors, extract the encoding parts corresponding to the theme word from the three encoding vectors respectively, and then use the method of extracting vertical components to extract the features of the front perspective, side perspective and back perspective of the theme word object from the three theme word encodings respectively; Step 2 includes: Get the user input text prompt word, select the theme word in the prompt word, and directly input the theme word into the pre-trained CLIP encoder to obtain the encoding vector of the theme word without perspective description Connect the subject word to the three descriptive words of different perspectives, that is, connect the end of the subject word to the front perspective, the side perspective, and the back perspective, respectively, to construct three phrases containing the subject word with perspective descriptions; Input the three phrases containing the subject word with the perspective description into the CLIP encoder respectively and locate the position of the subject word code to obtain the subject word code containing the corresponding perspective description extract Content-independent view feature encoding Using the method of extracting the vertical component, the formula is: in, Indicates the subject word code containing the corresponding perspective description Encoded vector describing the topic word without perspective The projection part on the The prior perspective features are decoupled from the keyword encoding in the user prompt to eliminate the typical view preference, and the perspective control information is injected on this basis. For the back perspective, the back perspective feature is injected after decoupling the front perspective and side perspective features, which is formulated as: Among them, view = {front, side} indicates that the perspective from which prior knowledge needs to be eliminated, including the front and side perspectives. The encoding vector representing the topic word without perspective description Content-independent view feature encoding Direction projection, represents the content-independent back view feature encoding, and η is a scale coefficient; Step 3: Input the user prompt word into the pre-trained text into the 3D diffusion model PointE to generate rough 3D point cloud data. Meanwhile, the rough 3D point cloud is fitted using the 3D Gaussian splashing technique to generate an explicitly represented 3D Gaussian model. Then, the 3D Gaussian model is rendered using randomly generated camera parameters to obtain renderings from more than two perspectives. Step 4: Input the user prompt word into the pre-trained text encoder CLIP to obtain the prompt word encoding vector, and locate the position of the subject word in the encoding vector according to the subject word. According to the camera parameters randomly generated in step 3, the features of the front view, side view and back view extracted in step 2 are respectively injected into the subject word encoding part in the user prompt word encoding, and the prompt word encoding after view control corresponding to all camera parameters is obtained. The new prompt word encoding vector and the rendering images of more than two view angles obtained in step 3 are input into the pre-trained stable diffusion model SD. Step 5: The stable diffusion model SD generates a corresponding two-dimensional image based on the user prompt word and the rendered image, extracts the unconditional prediction noise generated during the generation process, and obtains the unconditional denoising result. It then explores the similarity distribution of the image and uses the similarity partial order loss to constrain the unconditional denoising result, thereby enhancing the perspective consistency of the three-dimensional Gaussian model.

2. The method according to claim 1, characterized in that In step 1, the mathematical level analysis includes: The denoising diffusion probability model DDPM is used to optimize the following simplified objectives: Among them, L DDPM Represents the loss function of the denoising diffusion probability model DDPM; expected value It represents the expected value calculation of the original sample data x, noise ∈ and time step t, where the original sample data x is used as the input data of the denoising diffusion probability model DDPM, and the noise represents random noise that follows a standard normal distribution with a mean of 0 and a variance of 1; the time step t represents the diffusion time step, which is used to control the degree of noise addition. The time step in the diffusion model gradually increases from 1 to t; Denotes the denoising diffusion probability model DDPM in the model parameters Next, according to the noise perturbation of the sample x after time step t t and the noise value predicted at time step t; the added noise ∈ is the same as the model prediction noise The square of the two norm between Measuring the difference between the predicted noise and the actual added noise, during the inference process, the denoising diffusion probability model DDPM is calculated from x t Starting from the probability density Sampling the previous sample x from the normal distribution t-1 ; The following simplified objectives are optimized by the latent diffusion model LDM: Among them, L LDM is the loss function of the latent diffusion model LDM, and the expected value represents the potential encoding result z of the original sample data x, the user prompt c, and the random noise noise∈ and time step t for expected calculation; conditional information encoding represents the result of encoding the user prompt c through the text encoder CLIP, The second norm of the difference between the added noise and the predicted noise; In the framework of the denoising diffusion probability model DDPM and the latent diffusion model LDM, the generation of the image is regarded as the process from the final noise state z T The reverse recovery process to the noise-free initial clear state z0, in the process, the probability density p of generating the final noise-free initial clear state z0 is expressed by the following formula 2D (z0|c): Where T represents the maximum time step that t can take, p(z T ) represents the probability distribution of the final state of the diffusion process, p(z t-1 ∣z t ,c) represents the given user prompt c and the current potential representation z t Under the condition, generate the potential representation z of the previous step t-1 The conditional probability of is the differential element in the integral; Construct a non-normalized probability density function of the three-dimensional parameter θ Among them, Λ represents the set of all viewing angles; Redefine the density function of the three-dimensional parameter θ For a given series of optimization steps τ, the viewing angle control λ τ , user prompt c and 3D Gaussian model projection under the corresponding viewing angle Under the condition, the product of the conditional likelihood of each iteration step τ is expressed as: Where N represents the maximum optimization step that τ can take, and Z θ,τ represents the state of the 3D model controlled by the 3D parameters θ in the iteration step τ; From the perspective of the stable diffusion model’s understanding of user prompts, the user prompt c is divided into content part c c and perspective prior part c v , content part c c and perspective prior part c v is orthogonal; the density function of the three-dimensional parameter θ The new expression is: Where st means subject to constraints, c c ⊥c v , taking the logarithm of both sides of the equation gives: Next, using the chain rule, The gradient of θ It is expressed as: in, Indicates the 3D model state Z θ,τ The partial derivative with respect to the three-dimensional parameter θ is Represents the logarithmic probability density About 3D Model Status Z θ,τ The partial derivative of right Further expansion using Bayes' theorem: Further broken down into: Among them, the viewing angle controls λ τ and perspective prior part c v Point-wise conditional mutual information c c and is considered as a constant in the current optimization step, so PCMI(λ τ ,c v ∣Z θ,τ ) term and expand it using the definition of conditional probability:

3. The method according to claim 2, characterized in that In step 3, the rough three-dimensional point cloud data contains basic geometric information of the object surface.

4. The method according to claim 3, characterized in that Step 4 includes: Input the user prompt word into the pre-trained text encoder CLIP to obtain the prompt word encoding vector, and locate the position of the topic word in the encoding vector according to the topic word; Obtain the randomly generated camera parameters in step 3. The viewing angle control process can be adaptively implemented through the camera parameters of each optimization step, specifically including: when the azimuth angle in the camera parameters is at (-90°, 90°): in, The encoding vector representing the topic word without perspective description Content-independent frontal view feature encoding Directional projection; When the azimuth is (-90°, -180°) ∪ (90°, 180°): Among them, the weight w adjusted according to the azimuth angle in the camera parameters reflects The injected intensity, when the azimuth angle in the camera parameters is closer to 180° or -180°, that is, when the current viewing angle is closer to the front and back sides, the weight w is greater; Next, the new prompt word encoding vector and the perspective rendering image obtained in step 3 are input into the pre-trained stable diffusion model SD.

5. The method according to claim 4, characterized in that Step 5 includes: Uniform sampling is performed on the two groups of results with and without the multi-faceted problem. The sampling method is to sample uniformly spaced camera parameters from the upper hemisphere of a fixed radius. All camera parameters point to the center of the sphere at the same height, and render images from the 3D scene to explore the cosine similarity relationship between images and obtain the cosine similarity distribution law. Randomly select one from all the renderings as the reference image, calculate the cosine similarity between other rendering images and the reference image, randomly generate c_batchsize view control λ in each round of optimization, and determine an order for the view control according to the azimuth in the view control λ, and the order obeys the cosine similarity distribution law; establish a rectangular coordinate system, and construct a unit circle with the origin as the center, the angle corresponding to the point on the unit circle starts from 0° in the negative direction of the vertical axis and gradually increases to 180° in the positive direction of the horizontal axis and decreases to -180° in the negative direction of the horizontal axis, and the randomly generated view control λ is placed on the unit circle according to the azimuth size, then randomly select one of the view controls as the reference camera, and construct a secant s parallel to the horizontal axis on the unit circle with the view control as the starting point, calculate the distance d between the points corresponding to other view controls and the secant s, and sort the c_batchsize view controls according to the size of d; By controlling the viewing angle λ, we obtain the rendered image R, the potential representation ψ of the rendered image R, and the potential representation of the rendered image R after adding noise, and obtain the unconditional noise prediction result ∈ of the potential representation ψ of the rendered image R after the stable diffusion model SD is passed. φ ,∈ φ It represents the noise component of ψ estimated by the stable diffusion model SD after noise processing; Will ∈ φ Remove from the noised latent representation to get the unconditional denoising result And calculate the unconditional denoising result The partial order loss between them, the loss function is Loss partial : Where c_batchsize represents the number of view control λ in each iteration, i represents the index position in c_batchsize view control, Represents the i-th rendered image R i Corresponding unconditional denoising results The first layer encoding result of ViT encoder is express and The cosine similarity between .

6. An electronic device, characterized in that: including processor and storage medium; The storage medium is used to store instructions; The processor is used to execute the steps of the method according to any one of claims 1 to 5 according to the instructions.

7. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the program is executed by a processor, the steps of the method according to any one of claims 1 to 5 are implemented.

Citation Information

Patent Citations

  • Arbitrary track three-dimensional scene construction and roaming video generation method and system guided by plain text

    CN117853686A