Multi-stage text-to-three-dimensional model generation method
Through a multi-stage text-to-3D model generation method, the initial model is generated using a multi-view diffusion model and a 3D perception module. Combined with a multi-step integral optimization and a clone segmentation module, the problems of poor generation quality and structural distortion in existing technologies are solved, and efficient and high-visual-quality 3D model generation is achieved.
Patent Information
- Application Number
- CN202510839634.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-23
- Publication Date
- 2025-10-03
AI Technical Summary
Existing technologies have problems with poor generation quality and structural distortion when generating high-quality three-dimensional models. Especially when dealing with complex geometric structures and high-fidelity textures and lighting effects, it is difficult to strike a balance between generation efficiency and visual quality.
A multi-stage text-to-3D model generation method is adopted, including an initialization module, an optimization module, and a clone segmentation module. A geometrically consistent initial model is generated through a multi-view diffusion model and a 3D perception module. Multi-step integral optimization and trajectory consistency loss function are used for optimization. The LOD perception strategy and geometry estimation module are combined for density control and structure enhancement.
It achieves efficient generation of high-fidelity three-dimensional models that are highly consistent with the semantics of the input text, improves generation efficiency and visual quality, and is suitable for three-dimensional model generation in complex scenes.
Smart Images

Figure CN120747355A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of artificial intelligence technology, and in particular relates to a multi-stage text-to-three-dimensional model generation method. Background Art
[0002] Driven by the wave of digitalization, 3D technology has become a core tool in fields such as architecture, film and television, games, and virtual reality (VR / AR). Its application has gradually penetrated into vertical industries such as education, medical care, and industrial design. Through the intuitive visualization of three-dimensional models, users can accurately simulate complex real-world scenes in virtual space, such as virtual fitting, digital twin factories, and immersive training systems. However, the generation of high-quality 3D content still faces significant challenges: traditional processes rely on professional software (such as Blender and Maya) and manual modeling, which is time-consuming, labor-intensive, and costly; although automated generation technology has lowered the threshold, it is still difficult to strike a balance between generation efficiency and visual quality when dealing with complex geometric structures, high-fidelity textures, and lighting effects.
[0003] In recent years, the rise of text-to-3D generation technologies has provided new solutions to this problem. By directly driving 3D model generation through natural language descriptions, users can quickly create customized 3D content without specialized modeling knowledge. Among them, methods based on fractional distillation sampling (SDS) have shown strong potential by transferring prior knowledge from 2D diffusion models to the 3D optimization process.
[0004] However, existing methods have two major bottlenecks: First, the SDS loss function, which relies on single-step predictions, is random and inconsistent with the gradient direction, resulting in overly smooth surfaces and missing details (such as hair and carved patterns) in the generated 3D models. Second, the initialization geometric deviation (such as the multi-faceted Janus problem) can easily cause the optimization to fall into local optimality, generating structural distortions or incoherent models. To address the above problems, the academic community has proposed two main solutions: one is a feedforward generative model based on 3D data, which directly uses 3D datasets to train diffusion models. Although it can ensure geometric consistency, it is difficult to extend to complex scenes due to the scarcity of 3D data and the cost of annotation. The other is an optimization method based on a 2D diffusion model, which upgrades the 2D generation capability to 3D space through differentiable rendering. However, it is limited by the approximate error of single-step SDS and has difficulty retaining high-frequency details.
[0005] At the same time, the existing adaptive density control mechanism (ADC) mainly dynamically adjusts the point cloud density based on the disparity gradient. Although it has a certain structural adaptability, it faces many limitations in practical applications. Specifically, gradient-based density control is susceptible to interference from the "gradient conflict" phenomenon - in high-frequency detail areas, gradients in different directions may offset each other, resulting in the areas that should be refined not being fully expressed; and the mechanism lacks a comprehensive perception of semantics and geometric characteristics, and is not effective in complex or non-uniform geometric change areas. More importantly, in dynamic scenes or generative tasks, ADC often cannot maintain geometric consistency, which can easily lead to problems such as uneven distribution of point clouds and structural distortion, thereby limiting the application potential of 3DGS in high-quality three-dimensional generation. Therefore, the present invention provides a multi-stage text-to-three-dimensional model generation method. Summary of the Invention
[0006] The purpose of the present invention is to provide a multi-stage text-to-three-dimensional model generation method to solve the technical problems of poor generation quality and structural distortion in the prior art.
[0007] In order to achieve the above object, a multi-stage text-to-3D model generation method is provided, which includes an initialization module, an optimization module and a clone segmentation module;
[0008] The initialization module includes the following steps:
[0009] Randomly generate an initial Gaussian point cloud and render it to obtain an orthographic perspective image;
[0010] Noise the orthogonal view image to obtain a pure noise image;
[0011] Build a 3D perception module based on depth information to obtain residual features;
[0012] Combining text prompt words, residual features and multi-view diffusion model to predict the noise of noisy images;
[0013] Based on the added noise and predicted noise, the SDS loss function is constructed to optimize the initial point cloud to obtain the initial model;
[0014] The optimization module includes the following steps:
[0015] Render color and depth maps of random view angles for the initial model;
[0016] Perform multi-step noise processing on the color image;
[0017] Multi-step prediction of noise in noisy images based on the TCD model;
[0018] Construct a trajectory consistency loss function for multi-step predicted trajectories;
[0019] Aiming at the prediction differences between different perspectives at the same time step, a depth-based multi-view consistency loss function is constructed;
[0020] Construct geometric regularization terms for point cloud distribution;
[0021] The clone segmentation module includes the following steps:
[0022] Based on the gradient information, Gaussian scale and projection area returned by the optimization module loss, an LOD-aware strategy is constructed to screen the Gaussian points that need to be cloned and segmented;
[0023] A geometric estimation module is constructed to extract normal, principal direction and curvature responses to guide the direction of cloning segmentation. Based on Gaussian scale, curvature and saliency score, a saliency-structure coupled perturbation mechanism is constructed. Through the synergistic effect of the above three modules, a three-dimensional Gaussian point cloud model with high semantic consistency with the input text is ultimately generated for high-fidelity text-driven 3D content generation.
[0024] Furthermore, the orthogonal viewing angles include four viewing angles: front, back, left, and right;
[0025] The method for performing noise processing on an orthogonal view image to obtain a pure noise image specifically includes: the mathematical model for performing noise processing on the orthogonal view image is:
[0026]
[0027] Among them, x0 represents the latent variable of the original orthogonal image, ∈ is the Gaussian noise that obeys the standard normal distribution, is the latent variable after adding noise at time step t, represents the forward noise accumulation factor, which is expressed as
[0028] The method of constructing a 3D perception module based on depth information to obtain residual features; combining text prompt words, residual features and a multi-view diffusion model to predict the noise of a noisy image; constructing an SDS loss function based on added noise and predicted noise, and optimizing the initial point cloud to obtain an initial model includes:
[0029] The mathematical model of the 3D perception module is:
[0030]
[0031] Among them, D t It represents the depth map corresponding to the current rendered image, y represents the text prompt word, t∈[0,T] represents the current diffusion step number, f down and f mid Represent the control residuals and intermediate control residuals in the downsampling stage respectively;
[0032] The mathematical model for predicting noise using the multi-view diffusion model is:
[0033]
[0034] Among them, Unet refers to the Unet module of the multi-view diffusion model, and the final initial stage loss function is:
[0035]
[0036] Furthermore, the specific method of rendering the color and depth images of random perspectives for the initialization model and performing multi-step noise processing on the color image includes:
[0037] The rendering image includes using a differentiable renderer to render the current three-dimensional Gaussian point cloud representation G0 at random perspectives to obtain a color image I under the corresponding perspective v and depth map D v , the rendering process is performed by the rendering function Implementation, its mathematical model is:
[0038]
[0039] Among them, c v Represents the camera parameters of the v-th camera view;
[0040] The rendered image is used for the subsequent multi-step noise addition, which includes: constructing a time series segments = {[t0,t1], [t1,t2], ..., [t N-1 ,t N ]}, where each [t i ,t i+1 ] represents a local noise adding stage, given the original prediction samples of each stage and noise prediction According to the forward process of the diffusion model, it is converted into the target time step t i+1 The following noisy sample:
[0041]
[0042] in, Represents the simulated diffusion state at the next time step.
[0043] Furthermore, the specific method for performing multi-step noise prediction on the noise image based on the TCD model includes:
[0044] Step 1: Use the DPM-solver to solve the process from time step t to time step s. The mathematical model is:
[0045]
[0046] Among them, x t→s represents moving from time step t to step size h = λ t -λ s The latent variable after the target step s, λ t Represents the log-SNR parameter (logarithmic signal-to-noise ratio) whose value is equal to σ t ,σ s represents the noise amplitude (variance term) of the diffusion process at time t and s, ε φ (x t ,y,t) represents the noise term predicted by a neural network (such as UNet) and depends on the current latent variable x t , text condition y and current time t, φ represents the network model parameters;
[0047] Step 2: After each step, add a small amount of noise to enhance the anti-interference ability and improve the model's sensitivity to high-frequency details. The mathematical model is:
[0048] x s→g =ODE(x t→s ,y uncond ,s,g)
[0049] where x s→g represents the new latent variable obtained by ODE integration from time step s to g, y uncond Indicates unconditional prompt, t>g>s;
[0050] Repeat steps 1 and 2 several times to get the final Right now:
[0051]
[0052] Among them, k represents k-step operation, so the SDS loss in the final optimization stage is:
[0053]
[0054] Where x0 represents the image obtained by initial rendering.
[0055] Furthermore, the specific method for constructing a trajectory consistency loss function for multi-step predicted trajectories includes:
[0056] The trajectory consistency loss function is constructed for the multi-step prediction trajectory to minimize the multi-step SDS gradient variance, and its mathematical model is:
[0057]
[0058] Where λ = λ max (1-t / T)α , T represents the total number of time steps, α controls the decay rate, λ max is the initial weight, represents the gradient of the SDS loss function at time step t, Var represents the variance, The mean represents the average gradient fluctuation over multiple time steps.
[0059] Furthermore, the specific method of constructing a depth-based multi-view consistency loss function for the prediction differences of different viewpoints at the same time step includes:
[0060] The multi-view consistency loss function includes:
[0061] The pixel p i =(u,v) combined with depth information d i (u,v) is back-projected to the three-dimensional point x in the camera coordinate system i ,Right now
[0062]
[0063] Where K represents the internal parameter matrix;
[0064] Introduce a Gaussian weight function based on angular distance:
[0065]
[0066] Among them, v i Represents the observation direction of the i-th view in the world coordinate system, and finally constructs the multi-view consistency loss function:
[0067]
[0068] in, represents the gradient map reconstructed from view i to view j through 3D projection, Z is the weight normalization factor, G i Represents the gradient map at view angle i.
[0069] Furthermore, the specific method of constructing a geometric regularization term for point cloud distribution includes:
[0070] The geometric regularization terms include:
[0071] Consider the normal consistency between the sampling point and its neighborhood in the point cloud, that is,
[0072]
[0073] Among them, N i represents the k nearest neighbors of its sampling point, n i ,n jare the normal vectors of the point and its neighbors respectively; in order to avoid the points from being concentrated in certain high-density areas, the repulsive regularization term of the distance between points is further introduced
[0074]
[0075] d ij =||x i -x j ||2
[0076] Among them, B represents the total number of samples, i represents the index of the center point currently considered, k represents the number of nearest neighbors considered for each point, ReLU represents the activation function, and d ij Represents x i with x j The distance between represents the sampling point, σ represents the Gaussian bandwidth parameter that controls the distance attenuation speed, and the total loss function of the optimization module is:
[0077]
[0078] in, is the total loss, L SDS represents the SDS loss function in the optimization phase, λ normal ,λ repulsion ,λ mv ,λ TCL Represents the weight parameters of each part of the loss, represents the geometric structure regularization term, represents the repulsion loss between points, represents the multi-view consistency loss, represents the trajectory consistency loss.
[0079] Furthermore, the specific method of constructing an LOD-aware strategy to screen Gaussian points to be cloned and segmented based on the gradient information, Gaussian scale, and projection area returned by the optimization module loss includes:
[0080] The LOD perception strategy includes:
[0081] Construct a significance score:
[0082] s i =normgrad i normscreen i
[0083] Among them, normgrad i and normscreen i are the normalized values of gradient amplitude and projection area respectively;
[0084] Building an LOD-aware strategy:
[0085] M LoD =(M grad ∧M scale )∨(M screen ∧(s i >τ s ))
[0086] Among them, M scale is the scale-aware mask, M grad The mask of points that are sensitive to the optimization loss, M screen represents the screen space perception mask, τ s Represents the significance score filtering threshold.
[0087] Furthermore, the specific method of constructing a geometric estimation module to extract normal, main direction and curvature response to guide the clone segmentation direction includes:
[0088] The geometry estimation module includes:
[0089] For each Gaussian center point p i , construct the centralized neighborhood matrix X from its k nearest neighbors i :
[0090]
[0091] Among them, C i represents the local covariance matrix, Represents X i The transposed matrix, p j Represents the center point p i The neighboring points of i Perform eigendecomposition:
[0092] C i =V i Λ i V i · ,Λ i =diag(λ0,λ1,λ2)
[0093] Among them, V i represents the eigenvector matrix, Λ i represents the eigenvalue diagonal matrix, V i · Indicates V i The transposed matrix of , λ0, λ1, λ2 represent eigenvalues, and diag represents a diagonal matrix;
[0094] By calculating the covariance matrix and performing eigenvalue decomposition, three orthogonal directions can be obtained: the direction corresponding to the minimum eigenvalue is regarded as the normal direction, the direction of the maximum eigenvalue is regarded as the main direction, and the normalized minimum eigenvalue is used as the curvature index:
[0095]
[0096] where κ i represents the normalized curvature, and ε is a positive number close to 0 to prevent numerical stability terms from dividing by zero.
[0097] Furthermore, the specific method for constructing the perturbation mechanism of significance-structure coupling based on Gaussian scale, curvature and significance score includes:
[0098] The disturbance mechanism includes:
[0099] Introducing a structural stability index η i =λ2 / (λ1+ε) reflects the signal-to-noise ratio of the local main direction and constructs the weighted curvature index:
[0100]
[0101] The normalized weighted curvature and significance scores i Combined, define the joint importance score:
[0102]
[0103] Where α represents the weight coefficient, which is used to allocate the number of encryption points N i ∈[1,N max ], and regulate the disturbance intensity:
[0104]
[0105] where p j ′ represents the newly generated sub-Gaussian, μ i represents the position of the i-th original Gaussian point, represents random disturbance (Gaussian noise), σ i Indicates the scale of the original point, d i,j is the disturbance direction;
[0106] In the construction of the perturbation direction, the perturbation direction is adaptively selected based on the local curvature intensity of the Gaussian point. If κ i Above the median offset, the perturbation direction is sampled in a plane orthogonal to the main direction; otherwise, the perturbation direction is along the main direction itself, thereby improving the structural integrity.
[0107] Compared with the prior art, the present invention has the following beneficial effects:
[0108] The present invention establishes an efficient and highly integrated multi-stage text-to-3D model generation method by integrating the initialization module, the optimization module and the clone segmentation module. The method fully utilizes the latest diffusion model technology and 3D reconstruction technology in the three modules. The initialization module realizes the conversion of text into a geometrically consistent rough 3D model through the multi-view diffusion model and the 3D perception module. The optimization module realizes the generation of a highly detailed 3D model consistent with the text semantics through multi-step integral optimization, trajectory consistency loss, multi-view consistency loss and geometric regularization term technology. The clone segmentation realizes precise point cloud density control through the LOD strategy, the geometric estimation module and the perturbation strategy. The modular design of the overall method provides high scalability and flexibility, which can efficiently handle complex text-to-3D model generation tasks and significantly improve user experience.
[0109] Based on the implementation methods provided in the above aspects, this application can also be further combined to provide more implementation methods. BRIEF DESCRIPTION OF THE DRAWINGS
[0110] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.
[0111] Figure 1 is an overall flow chart of the multi-stage text-to-3D model method of the present invention;
[0112] Figure 2 A schematic diagram of the multi-stage text-to-3D model principle of the present invention;
[0113] Figure 3 This is a workflow diagram of the initialization module of the present invention;
[0114] Figure 4 It is the workflow diagram of the optimization module of the present invention;
[0115] Figure 5 This is a workflow diagram of the clone segmentation module of the present invention;
[0116] Figure 6 Schematic diagram of the geometry estimation module of the present invention;
[0117] Figure 7 Generate renderings for the multi-stage text-to-3D model method of the present invention. DETAILED DESCRIPTION
[0118] In order to enable those skilled in the art to better understand the solutions of the present invention, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the embodiments described are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative work should fall within the scope of protection of the present invention. Unless otherwise specified, the technical means used in the embodiments are conventional means well known to those skilled in the art.
[0119] like Figure 1 As shown, the overall process of the multi-stage text-to-3D model generation method of the present invention includes: initialization, optimization and cloning segmentation parts; the system first generates an initial Gaussian point cloud representation through a multi-view diffusion model based on the input text prompt; then enters the optimization part to perform multi-step diffusion and geometric consistency optimization; finally, adaptive control of point cloud density and structural enhancement are achieved through LOD-based strategy and curvature estimation.
[0120] like Figure 2 As shown, the multi-stage text-to-3D model generation method of the present invention includes three main modules, the principle of which is based on the cross-modal fusion between the 2D diffusion model and the 3D Gaussian point cloud; through the input natural language prompts, a geometrically consistent initialization model is generated by combining the multi-view diffusion model (MVDream) with the 3D perception initialization module; a multi-step integral optimization module combined with the TCD model guides point cloud optimization; and finally, density control and high-detail expression are achieved through the cloning segmentation module of the geometric estimation and perturbation mechanism.
[0121] Initialize the module, such as Figure 3 As shown in the figure, the initialization module includes the following key steps: first, the system encodes the input text prompt word and randomly initializes the three-dimensional Gaussian point cloud; second, it renders images of the initial point cloud under four orthogonal perspectives: front, back, left, and right; then these images are subjected to Gaussian noise processing with standard normal distribution to generate pure noise images for diffusion inversion; combined with the depth map generated by rendering, a 3D perception module is constructed to extract geometric residual features across perspectives; finally, these features are combined with the text encoding and input into the Unet module to predict noise, and the initial point cloud representation is optimized by constructing the SDS loss function to obtain a geometrically consistent rough three-dimensional model.
[0122] Optimization modules such as Figure 4As shown in the figure, the optimization module takes the initialized three-dimensional Gaussian point cloud as input, and uses a differentiable renderer to render its color image and depth map from multiple random perspectives; then, the color image is input into the diffusion model for multi-step noise processing, and the noise trajectory of each step is predicted based on the TCD model; in this process, the system constructs a trajectory consistency loss function to maintain the semantic stability of the diffusion path and avoid geometric drift; at the same time, in order to improve the consistency across perspectives, the system constructs a multi-perspective consistency loss function through a depth-guided back-projection mechanism to spatially align the prediction results under different perspectives; in addition, geometric regularization terms are introduced to constrain the point cloud structure, including normal consistency and inter-point repulsion loss, which are used to improve the surface smoothness and structural coherence of the model.
[0123] Clone segmentation module, such as Figure 5 As shown in the figure, the cloning and segmentation module is mainly used for density control and structure enhancement. First, the system calculates the significance score based on factors such as the projection area of the Gaussian point and the image gradient intensity, and constructs the LOD perception strategy based on factors such as the point scale and SDS gradient response to screen out the key areas that need to be encrypted. Then, for these candidate Gaussian points, the geometric estimation module is called to extract the local covariance matrix and calculate the normal vector, main direction and curvature index. In terms of perturbation strategy, the system introduces a structural stability index, combines the significance score with the weighted curvature, calculates the joint importance score, and determines the perturbation direction (main direction or orthogonal plane) and perturbation intensity. Finally, the system performs cloning and splitting operations of sub-Gaussian points along the selected direction to achieve encrypted expression of key structural areas and improve the detail fidelity and structural integrity of the overall model.
[0124] Optionally, the orthogonal viewing angles include four viewing angles: front, back, left, and right.
[0125] Specifically, the present invention first receives a text prompt word input by the user, uses the multi-view diffusion model MVDream to perform conditional generation on it, and outputs multiple image views with geometrically consistent consistency; specifically, the model generates four interrelated view images based on a unified latent space encoding and cross-view feature fusion mechanism, corresponding to the front, back, left, and right directions respectively, ensuring that the images are consistent in shape contours, texture details, and projection relationships, providing high-quality input for subsequent geometric modeling of three-dimensional point clouds or Gaussian representations.
[0126] Optionally, the mathematical model for performing noise addition processing on the orthogonal image is:
[0127]
[0128] Among them, x0 represents the latent variable of the original orthogonal image, ∈ is the Gaussian noise that obeys the standard normal distribution, is the latent variable after adding noise at time step t, represents the forward noise accumulation factor, which is expressed as
[0129]
[0130] Optionally, the mathematical model of the 3D perception module is:
[0131]
[0132] Among them, D t It represents the depth map corresponding to the current rendered image, y represents the text prompt word, t∈[0,T] represents the current diffusion step number, f down and f mid They represent the control residual in the downsampling stage and the intermediate control residual respectively.
[0133] Specifically, the present invention further integrates a 3D perception module to enhance the geometric structure modeling of the multi-view image; specifically, the depth map is used as a guiding condition input to the control module, and the original image latent variables and text prompt vectors are combined to generate intermediate control residual features with geometric consistency, thereby realizing an effective transition from two-dimensional perception to three-dimensional structure.
[0134] Optionally, the noise mathematical model predicted by using the multi-view diffusion model is:
[0135]
[0136] Among them, Unet refers to the Unet module of the multi-view diffusion model, and the final initial stage loss function is
[0137] Specifically, geometrically consistent multi-view images are obtained through the MVDream model in the initialization module, which effectively solves the multi-head phenomenon in existing generation technologies. The geometric inconsistency of the three-dimensional model is further reduced by perceiving geometric information through the 3D perception module, laying a solid foundation for the subsequent optimization stage.
[0138] Optionally, the rendered image includes rendering the current three-dimensional Gaussian point cloud representation G0 at random perspectives using a differentiable renderer to obtain a color image I at a corresponding perspective. v and depth map D v , the rendering process is performed by the rendering function Implementation, its mathematical model is:
[0139]
[0140] Among them, c vRepresents the camera parameters of the v-th camera view. The rendered image is used for subsequent multi-step noise addition.
[0141] Construct time series segments = {[t0,t1],[t1,t2],…,[t N-1 ,t N ]}, where each [t i ,t i+1 ] represents a local noise adding stage, given the original prediction samples of each stage and noise prediction According to the forward process of the diffusion model, it is converted into the target time step t i+1 The following noisy sample:
[0142]
[0143] in, Represents the simulated diffusion state at the next time step.
[0144] Specifically, the present invention adopts a multi-step noise addition strategy to gradually simulate the real diffusion trajectory while maintaining semantic consistency, thereby improving the stability and reconstruction quality of the model.
[0145] Optionally, the performing multi-step prediction of the noise of the noisy image based on the TCD model includes:
[0146] A: Use the DPM-solver to solve the process from time step t to time step s. The mathematical model is:
[0147]
[0148] Among them, x t→s represents moving from time step t to step size h = λ t -λ s The latent variable after the target step s, λ t Represents the log-SNR parameter (logarithmic signal-to-noise ratio) whose value is equal to σ t ,σ s represents the noise amplitude (variance term) of the diffusion process at time t and s, Represents the noise term predicted by a neural network (such as UNet), which depends on the current latent variable x t , text condition y and current time t, Represents the network model parameters.
[0149] B: Add a small amount of noise after each step to enhance the anti-interference ability and improve the model's sensitivity to high-frequency details. The mathematical model is:
[0150] x s→g=ODE(x t→s ,y uncond ,s,g)
[0151] Among them, x s→g represents the new latent variable obtained by ODE integration from time step s to g, y uncond Indicates unconditional prompt, t>g>s;
[0152] Repeat parts A and B several times to get the final Right now
[0153]
[0154] Among them, k represents k-step operation, Refers to the optimization model formed in steps 1 and 2; therefore, the SDS loss in the final optimization stage is:
[0155]
[0156] Specifically, the present invention proposes a multi-step trajectory optimization strategy that uses a multi-step DPM-Solver integral optimization process to replace the traditional single-step SDS prediction mechanism. The traditional SDS method only performs reverse updates based on the noise estimate of the current time step, which cannot accurately restore the true diffusion trajectory and easily leads to problems such as geometric drift and appearance jitter. To address this problem, the present invention introduces a high-order numerical integration method to construct a multi-step prediction path in the continuous time domain, thereby performing a high-precision approximation of the reverse evolution process of the diffusion model.
[0157] Specifically, the present invention effectively stimulates the model's ability to express high-frequency details by introducing a small amount of noise perturbation in the multi-step diffusion process.
[0158] Specifically, in order to address the detail information that is easily weakened by smoothing in diffusion sampling, the present invention designs a detail enhancement mechanism based on random perturbations, which introduces small random perturbations to the samples in each noise addition stage, thereby promoting the model's diversified exploration of local geometry and texture structures; this trace noise perturbation not only avoids the loss of details caused by excessive smoothing in the traditional diffusion process, but also improves the ability to recover texture edges, curvature changes and subtle structures.
[0159] Optionally, a trajectory consistency loss function is constructed for the multi-step predicted trajectory to minimize the multi-step SDS gradient variance, and its mathematical model is:
[0160]
[0161] Where λ = λ max (1-t / T) α , T represents the total number of time steps, α controls the decay rate, λ max is the initial weight, represents the gradient of the SDS loss function at time step t, Var represents the variance, The mean represents the average gradient fluctuation over multiple time steps.
[0162] Specifically, the trajectory consistency loss function is used to constrain the evolutionary consistency between the samples generated at each time step in the multi-step diffusion sampling process; this loss function strengthens the model's global alignment ability for diffusion paths by measuring the difference between the generated state on the trajectories of different time steps and the target diffusion trajectory.
[0163] Specifically, the trajectory consistency loss not only considers the single-step prediction error, but also combines the sample similarity and geometric structure continuity between consecutive time points to avoid the cumulative deviation caused by the backward propagation of local sampling errors; by introducing this loss function, the present invention effectively improves the stability of the multi-step integral optimization strategy and the structural coherence of the generated results, especially enhancing the spatial consistency and detail retention ability of the model in cross-view 3D reconstruction tasks.
[0164] Optionally, the multi-view consistency loss function includes:
[0165] The pixel p i =(u,v) combined with depth information d i (u,v) is back-projected to the three-dimensional point x in the camera coordinate system i ,Right now:
[0166]
[0167] Where K represents the internal parameter matrix.
[0168] Introduce a Gaussian weight function based on angular distance:
[0169]
[0170] Among them, v i Represents the observation direction of the i-th view in the world coordinate system, and finally constructs the multi-view consistency loss function:
[0171]
[0172] in, represents the gradient map reconstructed from view i to view j through 3D projection, Z is the weight normalization factor, G i Represents the gradient map at view angle i.
[0173] Specifically, a multi-view consistency loss function is introduced to constrain the spatial and semantic consistency between the generated results under different camera perspectives. This loss function promotes the model to maintain global consistency of the three-dimensional structure during the multi-view sampling process by comparing the feature matching and geometric alignment of the rendered images from multiple perspectives. Specifically, the loss function not only considers the pixel-level error between perspectives, but also integrates depth information to ensure that the shape and details of the reconstructed model at different observation angles have no significant deviation. By introducing multi-view consistency loss, the present invention significantly reduces rendering artifacts and geometric drift between perspectives, improves the stability and realism of 3D reconstruction, and is suitable for high-precision text-to-3D generation tasks in complex scenes.
[0174] Optionally, the geometric regularization term includes:
[0175] Consider the normal consistency between the sampling point and its neighborhood in the point cloud, that is,
[0176]
[0177] Among them, N i represents the k nearest neighbors of its sampling point, n i ,n j are the normal vectors of the point and its neighbors respectively.
[0178] In order to avoid points concentrating in certain high-density areas, we further introduce an exclusion regularization term for the distance between points:
[0179]
[0180] d ij =||x i -x j ||2
[0181] Among them, B represents the total number of samples, i represents the index of the center point currently considered, k represents the number of nearest neighbors considered for each point, ReLU represents the activation function, and d ij Represents x i with x j The distance between represents the sampling point, σ represents the Gaussian bandwidth parameter that controls the distance attenuation speed, and the total loss function of the optimization module is:
[0182]
[0183] in is the total loss, represents the SDS loss function in the optimization phase, λ normal 、
[0184] λ repulsion ,λ mv ,λ TCLRepresents the weight parameters of each part of the loss, represents the geometric structure regularization term, represents the repulsion loss between points, represents the multi-view consistency loss, represents the trajectory consistency loss.
[0185] Specifically, the present invention effectively improves the surface smoothness and geometric continuity of the three-dimensional reconstructed model by constraining the normal consistency between the sampling point and its neighborhood.
[0186] Specifically, for the local neighborhood structure of point clouds or Gaussian points, the normal vector of each sampling point is calculated, and by designing a normal consistency loss function, the normal direction differences of adjacent points are constrained within a preset range, thereby reducing surface noise and discontinuities; this normal consistency mechanism not only enhances the local surface smoothing effect, but also promotes the retention of geometric features, especially showing better structural stability in edge and detail areas.
[0187] Specifically, this invention effectively prevents excessive clustering of sampling points in a 3D point cloud by introducing a repulsive regularization term for inter-point distance, ensuring uniform distribution and spatial coverage integrity. This regularization term constrains the Euclidean distance between adjacent sampling points, designing a repulsive force function that maintains a certain minimum spacing between adjacent points. This avoids structural redundancy and reconstruction errors caused by local overcrowding in the point cloud.
[0188] Optionally, the LOD perception module includes:
[0189] Construct significance score: s i =normgrad i normscreen i
[0190] Among them, normgrad i and normscreen i are the normalized values of gradient amplitude and projection area, respectively.
[0191] Building LOD-aware strategies: M LoD =(M grad ∧M scale )∨(M screen ∧(s i >τ s ))
[0192] Among them, M scale is the scale-aware mask (indicating it is still in a sparse state), M grad The mask of points that are sensitive to the optimization loss, M screen represents the screen space perception mask, τ sRepresents the significance score filtering threshold.
[0193] Specifically, by introducing the Level of Detail (LOD) perception strategy, 3D point clouds or Gaussian point clouds are adaptively screened based on multi-dimensional features such as gradient, scale, and perspective projection area.
[0194] This strategy selectively retains point cloud data with high geometric complexity and rich details by screening sampling points in key areas, while eliminating or sparsely distributing points in flat or low-detail areas, thereby achieving efficient sparseness and encryption of point clouds. Through this screening mechanism, the LOD-aware strategy effectively ensures the consistency and visual quality of geometric details in the multi-view rendering process, while reducing the consumption of computing resources.
[0195] like Figure 6 As shown in the figure, the geometric estimation module is mainly used to extract core geometric structure information in the local point cloud neighborhood, providing an accurate geometric basis for perturbation strategy and LOD control; the module first constructs a centralized neighborhood matrix from its K nearest neighbors for each Gaussian center point, and calculates its covariance matrix; then the covariance matrix is decomposed into three orthogonal directions; the direction corresponding to the minimum eigenvalue is regarded as the normal vector, the direction of the maximum eigenvalue is defined as the main direction, and the normalized minimum eigenvalue is used as a quantitative indicator of the local curvature response; the curvature estimation process is not only sensitive to complex local geometry, but also has good differentiability and noise resistance, and is suitable for density-aware control and perturbation direction determination.
[0196] Optionally, the geometry estimation module includes:
[0197] For each Gaussian center point p i , construct the centralized neighborhood matrix X from its k nearest neighbors i :
[0198]
[0199] Among them C i represents the local covariance matrix, Represents X i The transposed matrix, p j Represents the center point p i The neighboring points of i Perform eigendecomposition:
[0200] C i =V i Λ i V i ,Λ i =diag(λ0,λ1,λ2)
[0201] Where V irepresents the eigenvector matrix, Λ i represents the eigenvalue diagonal matrix, V i · Indicates V i The transposed matrix of , λ0, λ1, λ2 represent eigenvalues, and diag represents a diagonal matrix. By calculating the covariance matrix and performing eigenvalue decomposition, three orthogonal directions can be obtained: the direction corresponding to the minimum eigenvalue is regarded as the normal direction, and the direction of the maximum eigenvalue is regarded as the main direction. At the same time, the normalized minimum eigenvalue is used as the curvature index, that is:
[0202]
[0203] where κ i represents the normalized curvature, and ε is a positive number close to 0 to prevent numerical stability terms from dividing by zero.
[0204] Specifically, the present invention realizes the accurate capture and description of local geometric features in three-dimensional point clouds or volume rendering scenes through a geometric estimation module; the module provides accurate geometric guidance for subsequent point cloud optimization, detail enhancement and sparse encryption strategies by calculating the normal, curvature and other information of the sampling points; by introducing the geometric estimation module, the present invention effectively improves the geometric consistency and structural integrity of the three-dimensional model, and significantly enhances the quality of detail reconstruction in complex scenes.
[0205] Optionally, the disturbance mechanism includes:
[0206] Introducing a structural stability index η i =λ2 / (λ1+ε) reflects the signal-to-noise ratio of the local main direction and constructs the weighted curvature index,
[0207] The normalized weighted curvature and significance scores i Combined, define the joint importance score:
[0208]
[0209] Where α represents the weight coefficient, which is used to allocate the number of encryption points N i ∈[1,N max ], and regulate the disturbance intensity:
[0210]
[0211] Among them, p j ′ represents the newly generated sub-Gaussian, μ i represents the position of the i-th original Gaussian point, represents random disturbance (Gaussian noise), σ i Indicates the scale of the original point, d i,jis the perturbation direction (main direction or tangential perturbation, automatically decided according to the split mode); in the construction of the perturbation direction, we adopt the “automatic dual-mode” strategy based on the curvature intensity: if κ i If the deviation is above the median, the sample is taken from the plane orthogonal to the main direction; otherwise, the main direction itself is used to improve the structural integrity.
[0212] Example 1
[0213] This embodiment takes the user input text prompt "a red sports car with aerodynamic curves" as an example to demonstrate the specific application process of the multi-stage text-to-3D model generation method of the present invention.
[0214] Initialization phase: After the user enters the text, the system invokes a multi-view diffusion model based on the MVDream architecture to generate geometrically consistent images in the front, back, left, and right directions with a resolution of 512×512. The images are then subjected to normal distribution noise processing to construct a pure noise image as input.
[0215] Combined with the depth map information, the system further constructs a 3D perception module to extract consistent residual features across viewpoints, and inputs them into the multi-view diffusion Unet module together with text prompts to predict the initial noise of each image; by constructing the SDS loss function, the predicted noise is aligned with the noisy latent variable, completing the optimization of the initial Gaussian point cloud and obtaining a rough three-dimensional point cloud representation.
[0216] Optimization phase: The system uses a differentiable renderer to render color and depth images from multiple random viewpoints (sampling 4), and inputs the color image into a multi-step diffusion process. It also uses the DPM-solver and TCD model to perform multi-step noise prediction. It constructs a trajectory consistency loss function through a multi-step sampling path in continuous time to guide the 3D representation to maintain semantic consistency on the diffusion inversion path. The system further calculates the multi-view prediction between different viewpoints. Figure 1 The system eliminates the loss of consistency and projects the images from each perspective back into 3D space based on the camera projection matrix and depth information to ensure that the prediction results from different perspectives are aligned in terms of spatial position and gradient features. To enhance geometric stability, the system imposes regularization terms of normal consistency and distance between points on the point cloud distribution to avoid dense overlap and local noise.
[0217] Clone segmentation stage: The system constructs a significance score based on the projected area and gradient strength of the Gaussian point, and constructs an LOD mask based on the point scale and loss sensitivity. When the significance threshold is met, the points that need to be encrypted are screened out. The covariance matrix is calculated in the local neighborhood of the selected point, and the normal vector, main direction and normalized curvature value are extracted. The system introduces a structural stability index, calculates the weighted curvature response, and determines the perturbation direction (main direction or its orthogonal tangential plane) based on the automatic splitting strategy. For points that meet the perturbation conditions, sub-Gaussian point splitting is performed in the perturbation direction and the initial parameters are optimized, ultimately forming a high-density, structurally coherent detail area expression.
[0218] In summary, this embodiment realizes a fully automatic generation process from text to three-dimensional model through this process, solving the problem that traditional three-dimensional modeling technology has high requirements and is time-consuming.
[0219] Example 2
[0220] This embodiment uses the user input text prompt "a realistic female statue with curly hair and intricate dress details" as an example to illustrate the specific application process of the multi-stage text-to-3D model generation method of the present invention when processing highly complex human structures.
[0221] Initialization phase: After receiving the above-mentioned text prompts, the system calls the MVDream multi-view diffusion model to generate four orthogonal perspective images. The images maintain geometric consistency in texture distribution and spatial configuration, significantly avoiding the "Janus" problem in traditional multi-view synthesis; these images are input into the noise addition module for standard normal noise perturbation, and the 3D perception module constructed in combination with the depth map outputs residual features, which are input into the Unet module together with the text prompts to predict the noise map; the SDS loss function is combined with the above-mentioned noise terms for optimization to obtain the initial three-dimensional Gaussian point cloud model, which has preliminarily captured the figure's form, hairstyle structure outline, and the general outline of the skirt.
[0222] Optimization phase: Four viewpoints are randomly sampled and rendered using a differentiable renderer. After generating color and depth maps, three-stage noise addition is performed on each viewpoint image, and multi-step inversion prediction is performed using the DPM-Solver of the TCD model. To improve the consistency of the multi-step diffusion path, the system constructs a trajectory consistency loss function to guide the outputs of different time steps to maintain progressive coherence in semantics and texture structure. The system simultaneously calculates the three-dimensional back-projection consistency loss under different viewpoints to eliminate the geometric jitter caused by multi-view fusion. Among the geometric regularization terms, the normal consistency constraint enhances the smoothness and continuity of areas such as curly hair and facial details. The inter-point distance exclusion term effectively prevents structural collapse caused by dense aggregation of point clouds in the skirt pleat area.
[0223] During the cloning and segmentation phase, the system automatically identifies geometrically complex regions within the model (such as curly hair and skirt details), which exhibit high gradient amplitudes and large projected areas in the saliency score. After screening using the Level of Dimension (LOD) perception module, the system divides Gaussian points into areas to be infilled and areas not to be infilled, and constructs multiple masks by combining point scale, structural gradient, and SDS gradient response. In the geometry estimation module, the system extracts the local covariance matrix centered on each Gaussian point and performs eigenvalue decomposition to obtain the normal vector, principal direction, and curvature index. The system dynamically allocates the number of infill points by introducing an importance score weighted by curvature and structural stability. The perturbation mechanism employs an "automatic dual-mode" strategy: for regions with curvature above the median (such as curly hair), the perturbation direction is orthogonal to the principal direction to enhance structural coverage density. For smooth skirt areas, the principal direction itself is selected to maintain geometric consistency. Ultimately, by perturbing and cloning new points, the system effectively enhances the point cloud expressiveness of high-frequency detail areas, significantly improving the fidelity and naturalness of the character structure.
[0224] In summary, this embodiment realizes the process from complex text to high-detail three-dimensional model through this process, solving the problems of poor surface texture and distorted geometric structure in the prior art.
[0225] like Figure 7 As shown, it is a display diagram of the effects of Examples 1 and 2. It can be seen from the figure that the results produced by the multi-stage text to three-dimensional model method of the present invention have consistent geometric structures and high surface texture details.
[0226] The above embodiments are only used to illustrate the technical method of the present invention and are not intended to limit the present invention. Although the present invention has been described in detail with reference to the preferred embodiments, those skilled in the art should understand that the technical method of the present invention may be modified or replaced by equivalents without departing from the spirit and scope of the technical method of the present invention.
Claims
1. A multi-stage text-to-3D model generation method, comprising an initialization module, an optimization module, and a clone segmentation module, characterized in that: The initialization module includes the following steps: Randomly generate an initial Gaussian point cloud and render it to obtain an orthographic perspective image; Noise the orthogonal view image to obtain a pure noise image; Build a 3D perception module based on depth information to obtain residual features; Combining text prompt words, residual features and multi-view diffusion model to predict the noise of noisy images; Based on the added noise and predicted noise, the SDS loss function is constructed to optimize the initial point cloud to obtain the initial model; The optimization module includes the following steps: Render color and depth maps of random view angles for the initial model; Perform multi-step noise processing on the color image; Multi-step prediction of noise in noisy images based on the TCD model; Construct a trajectory consistency loss function for multi-step predicted trajectories; Aiming at the prediction differences between different perspectives at the same time step, a depth-based multi-view consistency loss function is constructed; Construct geometric regularization terms for point cloud distribution; The clone segmentation module includes the following steps: Based on the gradient information, Gaussian scale and projection area returned by the optimization module loss, an LOD-aware strategy is constructed to screen the Gaussian points that need to be cloned and segmented; A geometric estimation module is constructed to extract normal, main direction and curvature response to guide the clone segmentation direction; Based on Gaussian scaling, curvature, and saliency scores, a perturbation mechanism for saliency-structure coupling is constructed. Through the synergistic effect of the above three modules, a three-dimensional Gaussian point cloud model with high semantic consistency with the input text is finally generated for high-fidelity text-driven three-dimensional content generation.
2. The multi-stage text to 3D model generation method according to claim 1, characterized in that: In the initialization module, the initial Gaussian point cloud is randomly generated and rendered to obtain orthogonal perspective images including renderings of four perspectives: front, back, left, and right; The method for performing noise processing on the orthogonal view image in the initialization module to obtain a pure noise image specifically includes: the mathematical model for performing noise processing on the orthogonal image is: Among them, x0 represents the latent variable of the original orthogonal image, ∈ is the Gaussian noise that obeys the standard normal distribution, is the latent variable after adding noise at time step t, represents the forward noise accumulation factor, which is expressed as In the initialization module, the 3D perception module is constructed based on the depth information to obtain residual features; the noise of the noisy image is predicted by combining the text prompt words, residual features and multi-view diffusion model; the SDS loss function is constructed based on the added noise and the predicted noise, and the specific method for optimizing the initial point cloud to obtain the initial model includes: The mathematical model of the 3D perception module is: Among them, D t It represents the depth map corresponding to the current rendered image, y represents the text prompt word, t∈[0,T] represents the current diffusion step number, f down and f mid Represent the control residuals and intermediate control residuals in the downsampling stage respectively; The mathematical model for predicting noise using the multi-view diffusion model is: Among them, Unet refers to the Unet module of the multi-view diffusion model, and the final initial stage loss function is:
3. The multi-stage text to 3D model generation method according to claim 1, characterized in that: The specific method of rendering the color and depth images of random perspectives of the initialization model and performing multi-step noise processing on the color image in the optimization module includes: The rendering image includes using a differentiable renderer to render the current three-dimensional Gaussian point cloud representation G0 at random perspectives to obtain a color image I under the corresponding perspective v and depth map D v , the rendering process is performed by the rendering function Implementation, its mathematical model is: Among them, c v Represents the camera parameters of the v-th camera view; The rendered image is used for the subsequent multi-step noise addition, which includes: constructing a time series segments = {[t0,t1], [t1,t2], ..., [t N-1 ,t N ]}, where each [t i ,t i+1 ] represents a local noise adding stage, given the original prediction samples of each stage and noise prediction According to the forward process of the diffusion model, it is converted into the target time step t i+1 The following noisy sample: in, Represents the simulated diffusion state at the next time step.
4. The multi-stage text to 3D model generation method according to claim 1, characterized in that: The specific method of performing multi-step noise prediction on the noise image based on the TCD model in the optimization module includes: Step 1: Use the DPM-solver to solve the process from time step t to time step s. The mathematical model is: Among them, x t→s represents moving from time step t to step size h = λ t -λ s The latent variable after the target step s, λ t Indicates that the log-SNR parameter is equal to σ t ,σ s represents the noise amplitude of the diffusion process at time t and s, ε φ (x t ,y,t) represents the noise term predicted by the neural network, which depends on the current latent variable x t , text condition y and current time t, φ represents the network model parameters, e -h It represents the ODE numerical solution strategy of DPM-Solver; Step 2: After each step, add a small amount of noise to enhance the anti-interference ability and improve the model's sensitivity to high-frequency details. The mathematical model is: x s→g =ODE(x t→s ,y uncond ,s,g) Among them, x s→g represents the new latent variable obtained by ODE integration from time step s to g, y uncond Indicates unconditional prompt, t>g>s; Repeat steps 1 and 2 several times to get the final prediction graph Right now: Among them, k represents k-step operation, Refers to the optimization model formed in steps 1 and 2; therefore, the SDS loss in the final optimization stage is: Where x0 represents the image obtained by initial rendering.
5. The multi-stage text to 3D model generation method according to claim 1, characterized in that: The specific method of constructing a trajectory consistency loss function for the multi-step predicted trajectory in the optimization module includes: The mathematical model of the trajectory consistency loss function constructed for the multi-step prediction trajectory is: Where λ = λ max (1-t / T) α , T represents the total number of time steps, α controls the decay rate, λ max is the initial weight, represents the gradient of the SDS loss function at time step t, Var represents the variance, The mean represents the average gradient fluctuation over multiple time steps.
6. The multi-stage text to 3D model generation method according to claim 1, characterized in that: The specific method of constructing a depth-based multi-view consistency loss function for the prediction differences between different viewpoints at the same time step in the optimization module includes: The multi-view consistency loss function includes: The pixel p i =(u,v) combined with depth information d i (u,v) is back-projected to the three-dimensional point x in the camera coordinate system i ,Right now Where K represents the internal parameter matrix; Introduce a Gaussian weight function based on angular distance: Among them, v i Represents the observation direction of the i-th view in the world coordinate system, and constructs a multi-view consistency loss function: in, represents the gradient map reconstructed from view i to view j through 3D projection, Z is the weight normalization factor, G i Represents the gradient map at view angle i.
7. The multi-stage text to 3D model generation method according to claim 1, characterized in that: The specific method of constructing a geometric regularization term for point cloud distribution in the optimization module includes: The geometric regularization terms include: Consider the normal consistency between the sampling point and its neighborhood in the point cloud, that is, Among them, N i represents the k nearest neighbors of its sampling point, n i ,n j are the normal vectors of the point and its neighbors respectively; in order to avoid the points from being concentrated in certain high-density areas, an exclusion regularization term of the distance between points is further introduced: d ij =||x i -x j ||2 Among them, B represents the total number of samples, i represents the index of the center point currently considered, k represents the number of nearest neighbors considered for each point, ReLU represents the activation function, and d ij Represents x i with x j The distance between represents the sampling point, σ represents the Gaussian bandwidth parameter that controls the distance attenuation speed, and the total loss function of the optimization module is: in, is the total loss, represents the SDS loss function in the optimization phase, λ normal ,λ repulsion ,λ mv ,λ TCL Represents the weight parameters of each part of the loss, represents the geometric structure regularization term, represents the repulsion loss between points, represents the multi-view consistency loss, represents the trajectory consistency loss.
8. The multi-stage text to 3D model generation method according to claim 1, characterized in that: In the clone segmentation module, the specific method of constructing an LOD perception strategy to screen Gaussian points to be cloned and segmented based on the gradient information, Gaussian scale and projection area returned by the optimization module loss includes: The LOD perception strategy includes: Construct a significance score: s i =normgrad i ·normscreen i Among them, normgrad i and normscreen i are the normalized values of gradient amplitude and projection area respectively; Building an LOD-aware strategy: M LoD =(M grad ∧M scale )∨(M screen ∧(s i >τ s )) Among them, M scale is the scale-aware mask, M grad The mask of points that are sensitive to the optimization loss, M screen represents the screen space perception mask, τ s Represents the significance score filtering threshold.
9. The multi-stage text to 3D model generation method according to claim 1, characterized in that: The specific method of constructing the geometric estimation module in the clone segmentation module to extract the normal, main direction and curvature response to guide the clone segmentation direction includes: The geometry estimation module includes: For each Gaussian center point p i , construct the centralized neighborhood matrix X from its k nearest neighbors i : Among them C i represents the local covariance matrix, Represents X i The transposed matrix, p j Represents the center point p i The neighboring points of i Perform eigendecomposition: C i =V i L i V′ i ,L i =diag(λ0,λ1,λ2) Where V i represents the eigenvector matrix, Λ i represents the eigenvalue diagonal matrix, V i · Indicates V i The transposed matrix of , λ0, λ1, λ2 represent eigenvalues, and diag represents a diagonal matrix; By calculating the covariance matrix and performing eigenvalue decomposition, three orthogonal directions can be obtained: the direction corresponding to the minimum eigenvalue is regarded as the normal direction, the direction of the maximum eigenvalue is regarded as the main direction, and the normalized minimum eigenvalue is used as the curvature index: where κ i represents the normalized curvature, and ε is a positive number close to 0 to prevent numerical stability terms from dividing by zero.
10. The multi-stage text to 3D model generation method according to claim 1, characterized in that: The specific method for constructing the saliency-structure coupling perturbation mechanism based on Gaussian scale, curvature and saliency score in the clone segmentation module includes: The disturbance mechanism includes: Introducing a structural stability index η i =λ2 / (λ1+ε) reflects the signal-to-noise ratio of the local main direction and constructs the weighted curvature index: The normalized weighted curvature and significance scores i Combined, define the joint importance score f i : Among them, α represents the weight coefficient, which is used to allocate the number of encryption points N i ∈[1,N max ], and adjust the disturbance intensity: Among them, p j ′ represents the newly generated sub-Gaussian, μ i represents the position of the i-th original Gaussian point, represents random disturbance (Gaussian noise), σ i Indicates the scale of the original point, d i,j is the disturbance direction; In the construction of the perturbation direction, the perturbation direction is adaptively selected based on the local curvature intensity of the Gaussian point. If κ i If the perturbation direction is higher than the median offset, the perturbation direction is sampled in a plane orthogonal to the main direction. Otherwise, the perturbation direction is along the main direction itself to improve the structural integrity.
Citation Information
Cited By
Geometric perception ADMM-based training-text-free three-dimensional digital model generation method and system, terminal and storage medium
CN122089963A
Geometric perception ADMM-based training-free text-based three-dimensional digital model generation method, system, terminal and storage medium
CN122089963B