Three-dimensional data synthesis method, electronic equipment, storage medium and program product
By constructing training sample pairs of images and text and using low-rank adaptation technology to fine-tune the textual graph model, combined with the multi-view consistent texture generation strategy, the problems of scarcity and insufficient diversity of three-dimensional data are solved, and the quality of three-dimensional data generation and the generalization ability of the model are improved.
Patent Information
- Application Number
- CN202510751286.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-06
- Publication Date
- 2025-07-04
- Estimated Expiration
- 2045-06-06
AI Technical Summary
The scarcity of three-dimensional data, insufficient data diversity and limited generalization capabilities of generative models in the prior art have resulted in a limited number of three-dimensional data sets, which are difficult to meet the high-quality and diversified training needs, and the generation results of the model when facing new categories or new scenarios.
By obtaining a textureless three-dimensional mesh model, perform multi-view rendering and combine semantic description text construction training sample pairs, use low-rank adaptation technology to fine-tune the textual graph model, guide the geometric deformation of the three-dimensional mesh model, and introduce multi-view consistency texture generation strategy based on deformation to generate high-quality and diverse three-dimensional data.
It significantly improves the semantic consistency, geometric structure recovery accuracy and texture expression continuity of synthetic three-dimensional data, and enhances the model's reconstruction accuracy, diversity expression ability and cross-category generalization ability.
Smart Images

Figure CN120259590A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of computer technology, and particularly to a three-dimensional data synthesis method, an electronic device, a storage medium, and a program product. Background Art
[0002] With the continuous progress of deep learning and computer graphics technologies, three-dimensional generation models and multi-view diffusion models have shown broad application prospects in fields such as virtual reality, game development, and robot navigation. These models can learn complex geometric structures and texture information from limited training samples, enabling the automated reconstruction and generation of three-dimensional objects or scenes, and becoming important tools for promoting digital content generation and enhancing three-dimensional perception capabilities.
[0003] However, the existing technologies still face the following significant problems in generating high-quality and diverse three-dimensional data: Scarcity of three-dimensional data: Compared with two-dimensional data, the acquisition cost of three-dimensional data is higher and the process is more complex, resulting in a limited number of high-quality and structurally complete three-dimensional data sets, which are difficult to support the training needs of large-scale models.
[0004] Insufficient data diversity: Current mainstream publicly available three-dimensional data sets (such as ShapeNet) usually focus on limited object categories or standard scenes, lacking comprehensive coverage of diverse forms, detailed variations, and complex backgrounds in the real world. This limitation in the sample structure restricts the model's ability to learn sufficient variability expression during training, thereby affecting its performance on complex or novel objects.
[0005] Limited model generalization ability: Existing models generally rely on fixed real data sets for training and are difficult to adapt to new categories or new scenes outside the training data distribution. When faced with unseen object structures or environmental conditions, the generated results often exhibit problems such as deformation distortion and detail loss, indicating insufficient generalization performance and difficulty in meeting the requirements for generality and adaptability in practical applications. Summary of the Invention
[0006] In view of the deficiencies of the existing technologies, this application provides a three-dimensional data synthesis method, an electronic device, a storage medium, and a program product, at least to solve the problems of scarcity of three-dimensional image data, insufficient data diversity, and limited generalization ability of the generation model in the existing technologies.
[0007] To achieve the above objectives and other advantages, some embodiments of this application provide the following aspects: In a first aspect, some embodiments of this application provide a three-dimensional data synthesis method, including: Obtain multiple textureless 3D mesh models, render each of the 3D mesh models based on multiple preset viewpoints to obtain multiple viewpoint images, and construct training sample pairs of images and texts in combination with semantic description texts; Based on the training sample pairs, use the low-rank adaptation technology to fine-tune the pre-trained first text-to-image model to obtain a second text-to-image model with the ability to understand multi-view consistency; Based on the second text-to-image model, model the latent space semantic difference between the rendered image of the deformable 3D mesh model and the target semantic text, construct an optimization constraint and update the parameters of the deformable 3D mesh model to obtain a deformable 3D mesh model that conforms to the target semantics; Based on the deformable 3D mesh model, use a texture generation model to generate a texture image and attach it to the surface of the deformable 3D mesh model to obtain textured synthetic 3D data.
[0008] In a second aspect, some embodiments of the present application further provide an electronic device, where the electronic device includes: One or more processors; and a memory storing computer program instructions, where the computer program instructions, when executed, cause the processors to execute the 3D data synthesis method as described in any one of the above.
[0009] In a third aspect, some embodiments of the present application further provide a computer-readable storage medium, on which computer programs and / or instructions are stored, and when the computer programs and / or instructions are executed by a processor, the 3D data synthesis method as described in any one of the above is implemented.
[0010] In a fourth aspect, some embodiments of the present application further provide a computer program product, including computer programs and / or instructions, and when the computer programs / instructions are executed by a processor, the 3D data synthesis method as described in any one of the above is implemented.
[0011] Compared with the related art, in the solution provided by the embodiments of the present application, by constructing training sample pairs of images and texts and using the low-rank adaptation technology to fine-tune the first text-to-image model, a second text-to-image model with the ability to understand multi-view consistency is obtained, thereby guiding the geometric deformation of the 3D mesh model. On the basis of the deformed 3D mesh, a texture generation strategy with multi-view consistency is introduced, which can ensure that the generated texture images have good continuity and style consistency under multiple viewing angles. Therefore, the 3D data synthesis method provided by the present application effectively improves the comprehensive performance of the synthesized 3D data in terms of semantic consistency, geometric structure recovery accuracy, and texture expression continuity, provides high-quality and high-diversity training data support for downstream 3D generation models, and thus significantly enhances its reconstruction accuracy, diversity expression ability, and cross-category generalization ability. Figure 1 Brief Description of the Drawings
[0012] To more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the accompanying drawings required for the description of the embodiments. Obviously, the accompanying drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other implementation manners can also be obtained based on these drawings.
[0013] Figure 1 is a schematic flowchart of a three-dimensional data synthesis method provided by an embodiment of the present application; Figure 2 is a schematic flowchart of fine-tuning a text-to-image model based on low-rank adaptation provided by an embodiment of the present application; Figure 3 is a schematic flowchart of three-dimensional mesh deformation optimization based on semantic guidance provided by an embodiment of the present application; Figure 4 is based on multiple views Figure 1 is a schematic flowchart of three-dimensional mesh texturing based on multi-view consistency texture generation; Figure 5 is a schematic flowchart of the training process for fine-tuning a three-dimensional generation model based on mixed three-dimensional data provided by an embodiment of the present application Figure 6 is a schematic structural diagram of an electronic device provided by an embodiment of the present application. Detailed implementation manners
[0014] To make the objectives, technical solutions, and advantages of the embodiments of the present application clearer, the following will clearly and completely describe the technical solutions in the embodiments of the present application with reference to the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are some, rather than all, of the embodiments of the present application. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts fall within the scope of protection of the present application.
[0015] As one of the current mainstream open-source three-dimensional data resources, the Objaverse dataset is mainly generated through manual capture, model reconstruction, or graphics synthesis, aiming to approximately restore the appearance of three-dimensional objects in the real world. Although this dataset covers a wide range of object categories and has good annotation consistency, due to the simplification of some three-dimensional model construction processes, there are still certain deficiencies in geometric structure complexity and texture detail fidelity, making it difficult to meet the training requirements of generation models with high requirements for three-dimensional data.
[0016] The DiT (Diffusion Transformer) structure is a text-to-image architecture that incorporates the Transformer network structure into the diffusion model. It combines the step-by-step denoising ability of diffusion probability modeling with the advantages of Transformer in modeling global context for the generation and modeling of high-quality images or feature maps.
[0017] The first embodiment
[0018] The first embodiment of this application relates to a three-dimensional data synthesis method. Referring to Figure 1 as shown, the method may include the following steps: Step S1: Obtain multiple textureless three-dimensional mesh models, render each three-dimensional mesh model based on multiple preset viewpoints to obtain multiple viewpoint images, and construct a training sample pair of images and text in combination with semantic description text.
[0019] Regarding step S1, specifically, in this embodiment, the original three-dimensional mesh models can be from publicly available three-dimensional datasets (such as the Objaverse dataset). To ensure the unity of training data and avoid external texture interference, all texture information in the original three-dimensional mesh models is removed, and only geometric structure data is retained.
[0020] Furthermore, to obtain multiple textureless three-dimensional mesh models, it may include: Select multiple three-dimensional mesh models with texture and material attributes from a publicly available dataset containing multi-category three-dimensional objects, and perform preprocessing on the three-dimensional mesh models. The preprocessing includes deleting texture maps, clearing material parameter bindings, and resetting mesh surface attributes to obtain a white film three-dimensional mesh that only retains the geometric structure.
[0021] Specifically, select multiple three-dimensional mesh models with texture maps and material parameters from publicly available three-dimensional datasets. The three-dimensional datasets may include, but are not limited to, Objaverse, ShapeNet, etc. The selected three-dimensional meshes cover multiple object categories, such as furniture, tools, animals, transportation vehicles, etc. Then perform preprocessing operations on the selected three-dimensional mesh models to obtain white film three-dimensional meshes with a unified style. The preprocessing includes: deleting the texture map files (such as.jpg or.png format textures) bound to the models, clearing the material parameter binding information (such as PBR materials, metallicity, roughness, etc.), resetting the mesh surface attributes (for example, unifying the surface color, removing normal perturbations or specular effects to make the surface present a neutral, textureless state), etc. After the above processing, the obtained three-dimensional mesh models only retain the original geometric structures of vertices, edges, and faces, thus forming textureless white film three-dimensional meshes with standardized structures and consistent visuals.
[0022] Through the above preprocessing method, the unity of the surface style of the 3D model is achieved during the construction of the training data, effectively avoiding the rendering style deviation caused by the difference in material complexity, which helps the subsequent training model to focus more on the learning relationship between structural modeling and semantic expression, thus improving the model's modeling ability for 3D geometric features and cross-category generalization performance.
[0023] Use 3D modeling and rendering software (such as Blender tool) to render each textureless 3D mesh model from multiple preset perspectives. Exemplarily, the preset perspectives may include: front view, back view, left view, right view, front left, front right, back left, back right, and top view, a total of 9 directions, which are used to comprehensively cover the spatial geometric information of the object. The number of the preset perspectives can also be 4 (such as front, back, left, right), or 6 (such as front, back, left, right, top, bottom), and this embodiment does not limit this.
[0024] Stitch the above 9 perspective images in a predetermined order and arrange them in a 3×3 nine-square grid form to generate a combined image containing complete perspective information, which is used as the training input image. Among them, the predetermined order can be the perspective order, the spatial position order, the two-dimensional arrangement order, or any stitching order configured by the user or the program.
[0025] For each stitched combined image, construct its corresponding semantic description text. This description text can be automatically filled and generated based on manually set language template rules, such as according to the perspective order and the object structure features (such as "the front view shows the overall shape of the object", "the top view shows the contour structure"), or, take the above combined image as the input and generate it by combining a multimodal large language model (such as GPT-4o). This guiding text can include the object category, typical structure features, the detailed expression requirements reflected by each perspective, etc. Through the above method, a training sample pair with highly aligned image and text and complete semantics can be constructed to support the low-rank fine-tuning process of the subsequent text-to-image model.
[0026] Step S2: Based on the training sample pair, use the low-rank adaptation technology to fine-tune the pre-trained first text-to-image model to obtain a second text-to-image model with the ability to understand multi-perspective consistency.
[0027] Specifically for step S2, the first text-to-image model in this embodiment adopts a Transformer architecture based on the diffusion mechanism (such as the DiT structure). To achieve efficient parameter fine-tuning, on the premise of keeping the main weight parameters of the first text-to-image model frozen, low-rank adaptation modules (LoRA modules) are respectively inserted into the attention mapping sub-layer and the feed-forward sub-layer in its multi-layer Transformer network.
[0028] During the training process, first, random noise conforming to a Gaussian distribution is added to the combined image input to construct the denoising prediction task of the diffusion model. This model takes the noisy image at the current time step as input and learns to predict the noise residual added to the original combined image. The L2 norm is used as the loss function for model training to measure the difference between the predicted noise and the true noise, and the gradient is calculated based on this loss. Through the backpropagation mechanism, only the trainable parameters in the LoRA module are updated, thereby achieving lightweight fine-tuning of the first text-to-image model.
[0029] After training is completed, the fine-tuned first text-to-image model and the LoRA module with inserted and updated parameters together form the second text-to-image model. This model has the ability to model multi-view consistency and can accurately model the potential associations between combined images from different views and semantic texts. At the same time, this second text-to-image model learns the visual style features of the training images in the latent space and can generate images consistent or similar to the training picture style according to the input text.
[0030] Step S3: Based on the second text-to-image model, model the latent space semantic difference between the rendered image of the three-dimensional mesh model to be deformed and the target semantic text, construct an optimization constraint, and update the parameters of the three-dimensional mesh model to be deformed to obtain a deformed three-dimensional mesh model that conforms to the target semantics.
[0031] Specifically for step S3, the three-dimensional mesh model to be deformed is rendered and stitched from multiple views into a combined image, and Gaussian noise is added. The target semantic text is encoded and jointly input into the fine-tuned second text-to-image model with the image latent vector to predict the noise and backpropagate the loss. Combining the Jacobian field structure and the regularization term, the gradient is passed back to each vertex parameter to achieve geometric structure optimization driven by the target semantics.
[0032] For example, for a three-dimensional mesh model representing a "chair", the system first renders the model from multiple views such as the front, back, left, right, front left, front right, back left, back right, and top to obtain corresponding two-dimensional view images, and stitches these images to form a nine-grid combined image. Subsequently, noise conforming to a Gaussian distribution is added to this combined image to construct the diffusion prediction input.
[0033] The target semantic text can be an automatically generated semantic description, such as: "The nine-square grid shows different views of a chair. The chair consists of a square seat, four vertical legs, and a backrest extending from the rear edge." After this text is input into the text encoder, a latent space semantic vector is obtained, which is input into the fine-tuned second text-to-image model together with the latent vector extracted by the image encoder to predict the noise added to the combined image.
[0034] The loss function is calculated based on the difference between the predicted noise and the real noise, and the gradient of this loss function is backpropagated from the image space to the vertex positions of the 3D mesh through a differentiable rendering mechanism. With the help of the pre-constructed Jacobian field, this optimization signal is used to guide the iterative adjustment of the local geometric structure of the 3D mesh. Finally, on the basis of maintaining the structural rationality, the 3D model of the "chair" deforms into a new model structure that better fits the target semantic description, such as changes like a higher backrest and thicker legs, thus completing the geometric optimization process driven by the target semantics.
[0035] Step S4: Based on the deformed 3D mesh model, use a texture generation model to generate a texture image and attach it to the surface of the deformed 3D mesh model to obtain textured synthetic 3D data.
[0036] Specifically for step S4, after the geometric structure deformation is completed, the obtained deformed 3D mesh model is used as input and fed into a texture generation model (such as the MvPaint model) to generate a texture image suitable for this 3D mesh structure. The texture generation model can generate high-quality and continuous texture maps while maintaining style consistency based on multiple perspective input images, thus avoiding problems such as local jumps, edge misalignments, or style incoherence that may occur in traditional texture mapping methods.
[0037] In this embodiment, the MvPaint network can receive the 2D images rendered from multiple perspectives of the deformed 3D mesh as input, extract local texture features from multiple perspectives through a perspective alignment module, and fuse them into a consistent global texture representation. Finally, one or more texture images are output, and these images can be attached to the 3D mesh surface based on the UV mapping method to obtain textured synthetic 3D data with detail continuity and style unity.
[0038] It is not difficult to find that, compared with the related technologies, in the solution provided by the embodiments of the present application, by constructing a training sample pair of an image and text, and using the low-rank adaptation technology to fine-tune the first text-to-image model, a second text-to-image model with the ability to understand multi-view consistency is obtained, and then the geometric deformation of the 3D mesh model is guided. On the basis of deforming the 3D mesh, a texture generation strategy with multi-view Figure 1 consistency is introduced, which can ensure that the generated texture images have good continuity and style consistency under multiple viewing angles. Therefore, the 3D data synthesis method provided by the present application effectively improves the comprehensive performance of the synthesized 3D data in terms of semantic consistency, geometric structure recovery accuracy, and texture expression continuity, provides high-quality and high-diversity training data support for downstream 3D generation models, and thus significantly enhances their reconstruction accuracy, diversity expression ability, and cross-category generalization ability.
[0039] Second Embodiment
[0040] The second embodiment of the present application relates to a 3D data synthesis method. The second embodiment is an improvement based on the first embodiment. The specific improvement lies in: in the second embodiment of the present application, a specific implementation manner for low-rank adaptation fine-tuning of the first text-to-image model is provided. That is, step S2 may further include the following steps: Step S201: On the basis of freezing the main parameters of the first text-to-image model, low-rank adaptation modules are respectively inserted into the attention mapping sub-layer and the feed-forward sub-layer in the multi-layer Transformer structure of the first text-to-image model; Step S202: Add random noise obeying a Gaussian distribution to the combined image of the training sample pair to construct a diffusion prediction task, and the combined image is formed by splicing multiple perspective images in a predetermined order; Step S203: Input the training sample pair into the first text-to-image model, and predict the noise added to the combined image through the diffusion modeling process; Step S204: Construct a first mean square error loss function based on the difference between the predicted noise and the random noise and perform backpropagation gradient, and only optimize the trainable parameters in the low-rank adaptation module on the premise of keeping the main weights of the model unchanged; Step S205: When the first mean square error loss function reaches the convergence condition, complete the fine-tuning training of the low-rank adaptation module, and jointly constitute a second text-to-image model based on the first text-to-image model and the optimized low-rank adaptation module.
[0041] Specifically, referring to Figure 2As shown, on the basis of freezing the main parameters of the first text-to-image generation model (such as the diffusion Transformer model based on the DiT architecture), LoRA modules are inserted into the attention mapping sub-layer and the feed-forward sub-layer in multiple Transformer layers of the model respectively. The LoRA module realizes efficient parameter reconstruction by introducing two groups of low-rank matrices, such as a pair of low-rank matrices with rank r and constitute, and are inserted into the linear mapping path of the specified layer in the first text-to-image generation model. Since r is much smaller than n, the degree of freedom after multiplying matrices A and B is much smaller than that of a full-parameter matrix. In this way, only a few parameters need to be trained, which can finely tune the entire large model structure and effectively adjust the behavior of the original model under the premise of extremely small parameter overhead.
[0042] Exemplarily, the first text-to-image generation model is a pre-trained text-image diffusion generation network based on DiT, which has good text-image alignment ability. On the basis of this model architecture, a LoRA module is introduced for parameter fine-tuning, thereby constructing a second text-to-image generation model with multi-view consistency modeling ability.
[0043] For each training sample, the de-textured three-dimensional mesh model is differentiably rendered from multiple preset perspectives (9 view directions) to obtain corresponding two-dimensional perspective images. The images under the above multiple perspectives are spliced into a combined image in a predetermined order, and the combination method can be a nine-grid arrangement form. To construct the diffusion modeling task, random noise conforming to the Gaussian distribution is added to the combined image as the image channel perturbation in the training input, so that the model learns how to denoise and restore the original image according to the text-image conditions during the training phase, thereby establishing the latent space mapping ability between semantics and images.
[0044] The combined image added with random noise and the corresponding semantic description text together constitute the conditional input of the diffusion model. The first text-to-image generation model performs diffusion modeling on this basis and predicts the Gaussian noise residual added to the combined image, thereby establishing a semantic alignment mechanism between the text-image inputs. Based on the difference between the predicted noise output by the first text-to-image generation model and the actual random noise added to the combined image, a first mean square error loss function is constructed, and the L2 norm form can also be used to measure the prediction error.
[0045] After calculating the loss function, the gradient calculation and weight update process are performed through the backpropagation mechanism. To ensure the stability of the main structure of the model, the Transformer backbone network of the first text-to-image generation model is kept frozen during the training process, and only the trainable parameters in the LoRA module are updated for gradients. This can significantly reduce the parameter tuning overhead and still achieve effective transfer learning under the condition of a small number of samples (only using about 20 - 60 pairs of image and text samples), thereby improving the modeling ability of the model on the textureless white film images.
[0046] When the first mean squared error loss function converges below a preset error threshold, or when the training process reaches the set maximum number of iterations, the fine-tuning process is considered completed. The first text-to-image generation model after training completion and the optimized LoRA module together constitute the second text-to-image generation model described in this application. Based on the original visual generation ability of the pre-trained model, this text-to-image generation model can perceive and align the structural correlation relationships of the same three-dimensional object in multiple perspective images; at the same time, it can also generate images consistent with the style of the training samples according to the input semantic description text, improving the quality of the latent space mapping from text to image. It improves the structural understanding and visual discrimination ability of white film style images, providing a stable image semantic prior for downstream tasks such as three-dimensional geometry optimization.
[0047] It is not difficult to find that in the solution provided by the embodiments of this application, during the fine-tuning process of the text-to-image generation model, only a small number of image and text sample pairs are used to obtain a lightweight diffusion model with robust performance in the multi-perspective consistency modeling task. Compared with the existing methods that require fine-tuning on a large number of RGB images, this application completes model fine-tuning on de-textured white film data, enabling the model to have stronger generalization ability and semantic expression ability when facing real three-dimensional geometry modeling tasks, providing stable image semantic prior support for subsequent three-dimensional geometry deformation modeling.
[0048] The Third Embodiment
[0049] The third embodiment of this application relates to a three-dimensional data synthesis method. The third embodiment is an improvement based on the first embodiment. The specific improvement lies in: in the third embodiment of this application, for the geometric deformation modeling process, a specific implementation manner combining differentiable rendering, semantic-guided optimization, and Jacobian field regularization constraint is provided. That is, step S3 can further include the following steps: Step S301: Render the three-dimensional mesh model to be deformed through a differentiable rendering module to obtain rendering images at multiple preset perspectives and splice them into a combined image; Step S302: Construct a Jacobian field based on the topological structure of the three-dimensional mesh model to be deformed, which is used to describe the local spatial transformation relationship in the three-dimensional mesh, and introduce a Jacobian regularization term to constrain the change of the vertex positions of the three-dimensional mesh; Step S303: Add Gaussian-distributed random noise to the combined image and input it into the image encoder to obtain an image latent space representation. At the same time, input the target semantic text into the text encoder to obtain a text latent space representation; Step S304: Use the image latent space representation and the text latent space representation as conditional inputs to the second text-to-image generation model, and output the predicted value of the noise added to the above-mentioned combined image; Step S305: Construct a second mean square error loss function based on the difference between the predicted value and the random noise, and relying on the gradient backpropagation mechanism supported by the differentiable rendering module, conduct the gradient of the second mean square error loss function with respect to the image pixels from the image domain to the corresponding vertex positions in the Jacobian field; Step S306: Jointly construct a total optimization objective function by combining the second mean square error loss function and the Jacobian regularization term, and perform multiple rounds of iterative optimization on the vertex positions of the Jacobian field until the loss value of the total optimization objective function converges, obtaining a deformed three-dimensional mesh model consistent with the target semantic text.
[0050] Specifically, as shown in Figure 3 Input the original three-dimensional mesh model to be deformed into the differentiable rendering module (such as the differentiable rendering nvdiffrast.torch library, which provides differentiable rasterization and interpolation operations), and render from multiple preset perspectives respectively. Arrange and splice the images from multiple perspectives in a predetermined order to form a combined image, which is used as the input image for the fine-tuned second text-to-image model.
[0051] It should be explained that differentiable rasterization is to record which triangles and vertices contribute the most to each pixel when projecting three-dimensional vertices onto a two-dimensional plane, and retain the gradient channel. Differentiable interpolation is to interpolate attributes such as image color, normal, and UV from vertices to pixels, and can also backpropagate the error gradient to the vertices.
[0052] It should be noted that in the training sample construction stage and the geometric deformation optimization stage, the same perspective configuration and image arrangement order are adopted. Exemplarily, nine fixed perspective directions can be selected, including the front, back, left, right, front left, front right, back left, back right, and top, and they are spliced into a combined image in a fixed order. The above combined image structure remains consistent during the training of the text-to-image model and the three-dimensional mesh deformation process, which is beneficial for the model to transfer the consistent perception ability and optimization constraints between different stages. This can significantly reduce the risk of semantic drift caused by input changes and improve the generalization ability and accuracy of the fine-tuned second text-to-image model in the geometric optimization stage.
[0053] Based on the topological structure of the original 3D mesh model, a Jacobi field is constructed. The Jacobi field is used to describe the geometric transformation relationship between vertices in the local neighborhood of the 3D mesh. Specifically, the Jacobi field expresses the local gradient information of the mesh during the deformation process by calculating the coordinate differences between adjacent faces or vertex pairs in the mesh. On this basis, a Jacobi regularization term is introduced as an optimization constraint to maintain geometric continuity and smoothness during the deformation process. This regularization term can prevent abnormal structures such as irregular deformations, topological breaks, or local intersections from occurring in the 3D mesh during the optimization process by constraining the norm change of the Jacobi matrix or gradient smoothness, thereby ensuring that the finally generated deformed mesh has good renderability and geometric rationality while maintaining semantic consistency.
[0054] Random noise following a Gaussian distribution is added to the combined image generated in step S301 to construct a damaged image input for simulating the noise prediction task of the diffusion model during training. This noise addition process can be achieved by performing Gaussian perturbations with a standard deviation of 𝜎 on the pixel values of the image. Subsequently, the combined image with added noise is input into an image encoder (such as an image backbone network based on the ViT architecture) to extract the latent space representation vector corresponding to the image, which is used to describe the structure and style information in the multi-view combined image. At the same time, the target semantic text is input into a text encoder (such as the text branch in the CLIP model) to extract the corresponding text latent space representation for characterizing the target semantics specified by the user.
[0055] The image latent space representation and the text latent space representation obtained in step S303 are used as conditional inputs and jointly fed into the fine-tuned second text-to-image model. The random noise added to the image is predicted through the diffusion modeling mechanism, and the predicted value of the noise added to the combined image is output. Subsequently, the difference between this predicted value and the actual random noise added to the combined image is calculated to construct a loss function (such as mean squared error loss or L2 norm loss) for measuring the text-image consistency modeling ability of the second text-to-image model under the current conditions.
[0056] Due to the use of a differentiable rendering module that supports gradient backpropagation, the loss function calculated by the model can not only reflect the difference between the combined image and the predicted image, but also, during the backpropagation process, conduct the gradient of this loss function with respect to the image pixels layer by layer to the structural parameters of the 3D model. Specifically, the pixel error in the image first acts on the rendered image itself, and then backtracks to the vertex attributes on which the image depends through differentiable interpolation operations, and finally conducts to the position parameters of each vertex in the 3D mesh model. Through this mechanism, the model can adjust the mesh structure according to the image error to achieve 3D deformation optimization guided by semantic goals.
[0057] Construct the total optimization objective function by combining the second loss function constructed in step S305 with the Jacobian regularization term introduced in step S302. This optimization objective comprehensively considers the semantic consistency error in the image domain and the spatial smoothness constraint of 3D structure deformation, ensuring that the model maintains geometric stability while meeting semantic guidance. Under this optimization objective, an iterative optimization algorithm (such as the Adam or Adagrad optimizer) is used to update the vertex position parameters in the Jacobian field in multiple rounds. During the optimization process, the total loss function composed of the prediction error of the second text-to-image model and the Jacobian regularization term constraint is gradually minimized until the loss value converges or reaches the preset number of iterations. After the loss function converges, the optimization process terminates. At this time, the 3D mesh model satisfies that the rendered images from multiple perspectives are closer to the images described by the target semantic text; at the same time, the 3D mesh structure is continuous and distortion-free, with good renderability and trainability.
[0058] Since the constructed Jacobian field records the transformation gradient information between vertices in the local area. To convert this local gradient information into a globally consistent vertex coordinate distribution. A Poisson equation solving mechanism is further introduced during the optimization process, that is, the gradient field extracted from the Jacobian field is restored to the deformed 3D mesh coordinates by solving the Poisson equation. The finally output 3D mesh is the deformed mesh model that has been optimized and conforms to the target semantic description. This model will be exported in the standard.ply format for subsequent texture generation or model training.
[0059] It is not difficult to find that in the solution provided by the embodiments of the present application, by introducing a diffusion text-to-image model for joint understanding of images and texts and a differentiable rendering mechanism that supports gradient backpropagation, the optimization of the 3D mesh structure driven by semantic information is realized. With the help of the differentiable rendering module, the loss gradient is effectively transmitted from the image domain to the 3D mesh structure, thus forming an end-to-end optimization path from semantics to geometry. Further, the local structure transformation is modeled by combining the Jacobian field, and the local gradient information is restored to the global vertex coordinates by solving the Poisson equation, so that the optimized mesh has structural continuity and deformation rationality.
[0060] It should be noted that the third embodiment of the present application can also be an improvement based on any one or more of the first embodiment to the second embodiment.
[0061] Fourth Embodiment
[0062] The fourth embodiment of the present application relates to a 3D data synthesis method. The fourth embodiment is an improvement based on the first embodiment. The specific improvement lies in: in the fourth embodiment of the present application, a specific implementation manner of texture generation based on a multi-view Figure 1 consistency mechanism is provided. That is, step S4 can further include the following steps: Step S401: Input the deformed three-dimensional mesh model into the texture generation model. The texture generation model automatically renders the input three-dimensional mesh based on multiple preset viewpoints to generate initial two-dimensional texture images from multiple viewpoints; Step S402: The texture generation model processes the multiple initial two-dimensional texture images through a viewpoint alignment mechanism, extracts local texture features from each viewpoint, and constructs cross-viewpoint correspondence relationships; Step S403: Based on the local texture features, generate a consistent texture representation in the latent space through a multi-view feature fusion strategy, and further generate a final two-dimensional texture image with consistent style and continuous structure; Step S404: Paste the final two-dimensional texture image onto the vertices or faces of the deformed three-dimensional mesh model according to the mapping relationship with the three-dimensional mesh surface to obtain synthetic three-dimensional data with a complete texture map.
[0063] Exemplarily, referring to Figure 4 As shown, input the deformed three-dimensional mesh model into the texture generation model, such as a deep network structure based on the MvPaint framework. The texture generation model will automatically render the input mesh based on multiple preset viewpoints to generate multiple initial two-dimensional texture images. Each two-dimensional texture image records the texture information of the mesh surface from that viewpoint.
[0064] Furthermore, the texture generation model processes the initial two-dimensional texture images from multiple viewpoints by introducing a viewpoint alignment mechanism, extracts local texture features in each two-dimensional texture image, such as edge contours, texture directionality, surface details, etc. By constructing spatial or semantic correspondence relationships between viewpoints, the matching and mapping of local features are realized, avoiding problems such as inconsistent texture style or structural breakage caused by viewpoint changes. Subsequently, the texture generation model performs multi-level fusion on the local texture features extracted from different viewpoints in the latent space. The fusion strategy can include methods such as feature weighting guided by an attention mechanism, feature stitching after alignment, or residual fusion. Based on this consistent latent space representation, a final two-dimensional texture image with consistent style and continuous texture is further generated. Finally, the generated final two-dimensional texture image will be accurately pasted onto the surface of the deformed three-dimensional mesh model according to the UV mapping or vertex attribute mapping mechanism in the three-dimensional mesh topology. This mapping process can be unfolded for vertices or faces to ensure that the generated result is visually complete and coherent, thus obtaining the final synthetic three-dimensional data with complete texture information .
[0065] It is not difficult to find that in the solution provided by the embodiments of the present application, while maintaining a reasonable three-dimensional deformation model structure, by introducing a multi-view texture feature modeling mechanism, the consistency of texture expression and the ability to retain details are enhanced, avoiding problems such as inconsistent multi-view texture mapping, texture jumps or blurring, and significantly improving the overall quality and realism of the synthesized three-dimensional data.
[0066] It should be noted that the fourth embodiment of the present application can also be an improvement based on any one or more of the first to third embodiments.
[0067] The Fifth Embodiment
[0068] The fifth embodiment of the present application relates to a three-dimensional data synthesis method. The fifth embodiment is an improvement based on the first embodiment. The specific improvement lies in: in the fifth embodiment of the present application, a specific implementation manner for evaluating the effectiveness of the synthesized three-dimensional data is provided, which specifically includes the following steps: Step A1: Mix the synthesized three-dimensional data and the real three-dimensional data to construct a training dataset, and fine-tune the three-dimensional generation model.
[0069] Further, step A1 specifically includes: Step A101: Render each synthesized three-dimensional data and real three-dimensional data in the training dataset according to multiple preset orthogonal views to generate two corresponding groups of two-dimensional images; Step A102: Input the two groups of two-dimensional images as supervised images into the three-dimensional generation model to guide the three-dimensional generation model to predict the corresponding three-dimensional mesh reconstruction result under the input view condition; Step A103: Perform differentiable rendering on the three-dimensional mesh reconstruction result to generate a predicted two-dimensional image, and construct a loss function based on the error at the pixel level between the predicted two-dimensional image and the supervised image; Step A104: Update the network weight parameters of the three-dimensional generation model through the backpropagation mechanism based on the loss function to improve the geometric accuracy, shape diversity, and generalization ability of the three-dimensional generation model under different views.
[0070] Specifically, referring to Figure 5 as shown, several three-dimensional object samples are respectively selected from the synthesized three-dimensional data and the real three-dimensional data D. For each three-dimensional object sample, differentiable rendering is respectively performed based on multiple preset orthogonal views (such as front view, rear view, left view, right view) to generate two corresponding groups of two-dimensional images, which are respectively denoted as the synthesized image group and the real image group V. The above two groups of images ( (V) As the supervised image samples are input into the 3D generation model (such as the LGM model), along with the corresponding view conditions, the model is guided to perform 3D structure reconstruction prediction at different viewing angles to obtain the 3D mesh reconstruction results. The 3D reconstruction results generated in step A102 are again passed through the differentiable rendering module to generate predicted 2D images from the corresponding perspectives, and are compared pixel by pixel with the real supervised images. A training loss function is constructed based on the pixel error and perceptual loss between the predicted image and the real image:
[0071] where represents the 2D image differentiably rendered by the LGM model based on the input view conditions; represents the real supervised image at the corresponding perspective, which is sourced from the synthetic image and the real image V; represents the pixel-level mean squared error used to measure the direct difference between the image color and structure; represents the perceptual loss, which is used to measure the similarity of the image at the visual perception level; λ represents the adjustment coefficient, which is used to balance the contribution weights of the perceptual loss and the pixel loss in the total loss.
[0072] Step A2: Based on the publicly available 3D test dataset, the performance of the 3D generation model before and after fine-tuning is evaluated. The performance evaluation metrics include: geometric accuracy metrics, shape diversity metrics, and generalization ability metrics.
[0073] Step A3: Based on the performance evaluation results, the performance differences of the 3D generation model before and after fine-tuning are compared to verify the effectiveness and applicability of the synthetic 3D data in improving the performance of the 3D generation model.
[0074] Based on the publicly available standard 3D test dataset (such as the Objaverse test set or the ShapeNet test set), the performance of the 3D generation model before and after fine-tuning is compared and evaluated respectively, aiming to verify the improvement effect of the model trained by introducing synthetic 3D data in terms of structure recovery and generalization ability.
[0075] Exemplarily, the geometric accuracy metrics evaluation can adopt two metrics: Chamfer Distance (CD) and F-score. Chamfer Distance is used to calculate the average point-to-point distance between the predicted mesh and the real mesh, reflecting the accuracy of the reconstructed geometric shape; F-score combines the matching degree and accuracy rate between the predicted point set and the real point set, and is used to comprehensively measure the performance of the reconstruction effect in terms of fineness and overall structure.
[0076] The evaluation of the shape diversity index can count the number of categories covered by the statistical model during the generation stage, the range of deformation distributions within the categories, or evaluate the model's coverage ability of object shapes through clustering analysis of the shape distributions after reconstructing multiple different object categories, reflecting its modeling diversity and expression ability.
[0077] The evaluation of the generalization ability index can apply the model to perform reconstruction tasks under new categories or new perspectives that do not appear in the training set and evaluate the changes in its output quality. For example, the degradation degree of accuracy can be calculated by predicting inputs for unseen categories in the Zero-shot scenario and combining metrics such as CD or LPIPS to measure the generalization robustness and semantic transfer ability of the model.
[0078] In the publicly available standard 3D test dataset or typical application scenarios, a systematic evaluation of the 3D generation model was carried out before and after introducing synthetic 3D data. The experimental results show that under the same model structure and training strategy, using the synthetic 3D data constructed in this application for mixed training with real 3D data can significantly improve the comprehensive performance of the model in the multi-view 3D reconstruction task. After being applied to the mainstream 3D generation model, the generated synthetic data can effectively improve the performance of the output model in terms of structural continuity, texture consistency, and semantic alignment while maintaining a realistic generation effect, helping to generate high-quality and structurally diverse 3D content.
[0079] It is not difficult to find that in the solution provided by the embodiments of this application, by constructing a training process that mixes synthetic 3D data and real 3D data and combining a standardized evaluation index system, the performance differences of the 3D generation model before and after fine-tuning are systematically compared, and the improvement effect of the synthetic 3D data on the model performance is quantitatively evaluated, which can objectively verify the training value of the synthetic data. Moreover, this embodiment constructs a closed-loop mechanism from data construction, model training to performance evaluation, providing a clear path for subsequent model optimization and data quality feedback.
[0080] It should be noted that the fifth embodiment of this application can also be an improvement based on any one or more of the first to fourth embodiments.
[0081] As can be seen from the above embodiments, compared with the high-cost real data obtained by relying on manual annotation or scanning, the method of the present application effectively expands the scale of the training set by automatically generating a large amount of synthetic 3D data, alleviates the problem of lack of 3D data, and provides more sufficient learning samples for the deep model. This method combines semantic-driven geometric deformation and texture reconstruction technology with multi-view consistency, enabling each synthetic sample to have high variability in morphology and appearance, thus significantly improving the coverage and representativeness of the training data set. The joint training of synthetic data and real data, guided by multi-view input conditions, improves the robustness and reconstruction accuracy of the model under unseen categories, unknown viewpoints or complex shapes, and enhances the generalization ability of the model to real scenes.
[0082] The step division of the above various methods is only for clear description. When implemented, they can be combined into one step or some steps can be split into multiple steps. As long as the same logical relationship is included, it is within the protection scope of the present application; adding insignificant modifications or introducing insignificant designs to the algorithm or process, but not changing the core design of its algorithm and process, are all within the protection scope of this application.
[0083] In addition, some embodiments of the present application also provide an electronic device. The electronic device can be various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and so on. The electronic device can also be various forms of mobile devices, such as personal digital processors, cellular phones, smart phones, wearable devices, and other similar computing devices.
[0084] The electronic device includes: one or more processors; and a memory storing computer program instructions, which when executed cause the processor to execute a 3D data synthesis method provided by any one or more of the above embodiments. Figure 6An exemplary structural diagram of the electronic device is disclosed. The electronic device includes: one or more processors 1101, a memory 1102, and interfaces for connecting various components, including a high-speed interface and a low-speed interface. Each component is interconnected using different buses and can be mounted on a common motherboard or otherwise installed as required. The processor can process instructions executed within the electronic device, including instructions stored in the memory or on the memory for displaying graphical information of a GUI on an external input / output device (such as a display device coupled to the interface). In some other embodiments, multiple processors and / or multiple buses can be used together with multiple memories and multiple memories if needed. Similarly, multiple electronic devices can be connected, and each device provides part of the necessary operations. Among them, the components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the present application described and / or claimed herein.
[0085] The electronic device may further include: an input device 1103 and an output device 1104. The processor 1101, the memory 1102, the input device 1103, and the output device 1104 can be connected via a bus or other means. Figure 6 Taking connection via a bus as an example.
[0086] The input device 1103 can receive input digital or character information and generate key signal inputs related to user settings and function controls of the electronic device, such as input devices like a touch screen, a keypad, a mouse, a trackpad, a touchpad, a pointing stick, one or more mouse buttons, a trackball, a joystick, etc. The output device 1104 can include a display device, an auxiliary lighting device (such as an LED), and a haptic feedback device (such as a vibration motor), etc. The display device can include, but is not limited to, a liquid crystal display, a light-emitting diode display, and a plasma display. In some embodiments, the display device can be a touch screen.
[0087] To provide interaction with the user, the electronic device can be a computer. The computer has: a display device for displaying information to the user (such as a cathode ray tube or an LCD monitor); and a keyboard and a pointing device (such as a mouse), through which the user can provide input to the computer. Other types of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (such as visual feedback, auditory feedback); and input from the user can be received in any form (such as voice input or tactile input).
[0088] In the embodiments of the present application, a computer program / instruction is stored on a computer-readable medium. When the computer program / instruction is executed by a processor, it implements a three-dimensional data synthesis method provided in any one or more of the above embodiments. The computer-readable medium may be included in the electronic device described in the above embodiments; or it may exist separately without being assembled into the device. The above computer-readable medium carries one or more computer-readable instructions.
[0089] The memory 1102 may be used as a non-transitory computer-readable storage medium for storing non-transitory software programs, non-transitory computer-executable programs, and modules. The processor 1101 executes various functional applications and data processing of the server by running the non-transitory software programs, instructions, and modules stored in the memory 1102, so as to implement the program instructions / modules corresponding to the methods provided in any one or more of the above embodiments of the present application.
[0090] The memory 1102 may include a program storage area and a data storage area. Among them, the program storage area may store an operating system and application programs required for at least one function; the data storage area may store data created according to the use of the electronic device. In addition, the memory 1102 may include a high-speed random access memory, and may also include a non-transitory memory, such as at least one magnetic disk storage device, a flash memory device, or other non-transitory solid-state storage devices. In some embodiments, the memory 1102 may optionally include a memory remotely disposed relative to the processor 1101, and these remote memories may be connected to the electronic device through a network. Examples of the above network include but are not limited to the Internet, an enterprise intranet, a local area network, a mobile communication network, and combinations thereof.
[0091] It should be noted that the computer-readable medium described in the present application may be a computer-readable signal medium, a computer-readable storage medium, or any combination of the above two. The computer-readable medium may be, for example, but not limited to: an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples of the computer-readable storage medium may include but are not limited to: an electrical connection having one or more wires, a portable computer disk, a hard disk, a random access memory, a read-only memory, an erasable programmable read-only memory, an optical fiber, a portable compact disk read-only memory, an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present application, the computer-readable medium may be any tangible medium that contains or stores a program, and this program may be used by or in combination with an instruction execution system, apparatus, or device.
[0092] A computer-readable medium includes permanent and non-permanent, removable and non-removable media, and information storage can be implemented by any method or technology. The information can be computer-readable instructions, data structures, program modules, or other data. Examples of computer storage media include, but are not limited to, phase change memory, static random access memory, dynamic random access memory, other types of random access memory, read-only memory, electrically erasable programmable read-only memory, flash memory, or other memory technologies, compact disc read-only memory, digital versatile disc, or other optical storage, magnetic cassettes, magnetic tape disk storage, or other magnetic storage devices, or any other non-transmission medium that can be used to store information accessible by a computing device.
[0093] Computer program code for performing the operations of this application can be written in one or more programming languages or combinations thereof. The programming languages include object-oriented programming languages such as Java, Smalltalk, C++, and also include conventional procedural programming languages such as the C language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, executed as a stand-alone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the case of a remote computer, the remote computer can be connected to the user's computer through any type of network, including a local area network or a wide area network, or can be connected to an external computer (for example, by using an Internet service provider to connect through the Internet).
[0094] In the above embodiments, it can be implemented in whole or in part by software, hardware, firmware, or any combination thereof. For example, dedicated integrated circuits, general-purpose computers, or any other similar hardware devices can be used. In some embodiments, the software program of this application can be executed by a processor to implement the above steps or functions. Similarly, the software program of this application (including related data structures) can be stored in a computer-readable recording medium, such as a RAM memory, a magnetic or optical drive, or a floppy disk and similar devices. In addition, some steps or functions of this application can be implemented by hardware, for example, as a circuit that cooperates with a processor to execute each step or function.
[0095] The computer program product provided by the embodiments of the present application includes one or more computer programs / instructions. When the computer programs / instructions are executed by a processor, they wholly or partly generate the processes or functions described in the embodiments of the present application. The computer may be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions may be stored in a computer-readable storage medium, or transmitted from one computer-readable storage medium to another computer-readable storage medium. For example, the computer instructions may be transmitted from a website, a computer, a server, or a data center to another website, a computer, a server, or a data center in a wired manner (such as coaxial cable, optical fiber, digital subscriber line) or a wireless manner (such as infrared, wireless, microwave, etc.). The computer-readable storage medium may be any available medium that can be accessed by a computer or a data storage device such as a server or a data center that includes one or more integrated available media. The available medium may be a magnetic medium (such as a floppy disk, a hard disk, a magnetic tape), an optical medium (such as a DVD), or a semiconductor medium (such as a solid-state drive), etc.
[0096] The flowcharts or block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of devices, methods, and computer program products according to various embodiments of the present application. In this regard, each block in the flowchart or block diagram may represent a module, a program segment, or a part of code that contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order from that marked in the accompanying drawings. For example, two consecutive blocks shown may actually be executed substantially in parallel, and they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram and / or flowchart, as well as the combination of blocks in the block diagram and / or flowchart, may be implemented by a dedicated hardware-based system for performing the specified functions or operations, or may be implemented by a combination of dedicated hardware and computer instructions.
[0097] The scope of the present application is defined by the appended claims rather than the above description. Therefore, all changes that fall within the meaning and scope of the equivalent elements of the claims are intended to be included in the present application. Any reference signs in the claims should not be construed as limiting the claims involved. In addition, it is obvious that the word "comprising" does not exclude other elements or steps, and the singular does not exclude the plural. The multiple elements or devices stated in the apparatus claims may also be implemented by one element or device through software or hardware. The words "first", "second", etc. are only used for distinguishing descriptions and do not represent any specific order, nor can they be understood as indicating or implying relative importance.
[0098] As described above, the above are only specific embodiments of the present application, but the protection scope of the present application is not limited thereto. Any person skilled in the art within the technical scope disclosed in the present application can easily make changes or substitutions, which should all be covered within the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims, and the above embodiments should be regarded as exemplary and non-limiting.
Claims
1. A three-dimensional data synthesis method, characterized in that, Including: Obtain multiple textureless 3D mesh models, render each of the 3D mesh models based on multiple preset perspectives to obtain multiple perspective images, and construct a training sample pair of images and texts in combination with semantic description texts; Based on the training sample pair, use the low-rank adaptation technique to fine-tune the pre-trained first text-to-image model to obtain a second text-to-image model with the ability to understand multi-perspective consistency; Based on the second text-to-image model, model the latent space semantic difference between the rendered image of the to-be-deformed 3D mesh model and the target semantic text, construct an optimization constraint and update the parameters of the to-be-deformed 3D mesh model to obtain a deformed 3D mesh model that conforms to the target semantics; Based on the deformed 3D mesh model, use a texture generation model to generate a texture image and attach it to the surface of the deformed 3D mesh model to obtain textured synthetic 3D data.
2. The three-dimensional data synthesis method according to claim 1, characterized in that, The step of obtaining multiple textureless 3D mesh models includes: selecting multiple 3D mesh models with texture and material attributes from a public dataset containing multi-category 3D objects, and preprocessing the 3D mesh models. The preprocessing includes deleting texture maps, clearing material parameter bindings, and resetting mesh surface attributes to obtain white film 3D meshes that only retain the geometric structure.
3. The three-dimensional data synthesis method according to claim 1, wherein, The step of, based on the training sample pair, using the low-rank adaptation technique to fine-tune the pre-trained first text-to-image model to obtain a second text-to-image model with the ability to understand multi-perspective consistency includes: On the basis of freezing the main parameters of the first text-to-image model, insert low-rank adaptation modules into the attention mapping sub-layer and the feed-forward sub-layer in the multi-layer Transformer structure of the first text-to-image model respectively; Add random noise obeying a Gaussian distribution to the combined image of the training sample pair for constructing a diffusion prediction task. The combined image is formed by splicing multiple perspective images in a predetermined order; Input the training sample pair into the first text-to-image model and predict the noise added to the combined image through a diffusion modeling process; Construct a first mean square error loss function based on the difference between the predicted noise and the random noise and perform backpropagation gradient, and only optimize the trainable parameters in the low-rank adaptation module on the premise of keeping the main weights of the model unchanged; Under the condition that the first mean square error loss function reaches the convergence condition, complete the fine-tuning training of the low-rank adaptation module, and the second text-to-image model is jointly composed of the first text-to-image model and the optimized low-rank adaptation module.
4. The three-dimensional data synthesis method according to claim 1, characterized in that, The step of, based on the second text-to-image model, modeling the latent space semantic difference between the rendered image of the to-be-deformed 3D mesh model and the target semantic text, constructing an optimization constraint and updating the parameters of the to-be-deformed 3D mesh model to obtain a deformed 3D mesh model that conforms to the target semantics includes: Render the to-be-deformed 3D mesh model through a differentiable rendering module to obtain multiple rendered images under the preset perspectives and splice them into a combined image; Construct a Jacobi field based on the topological structure of the three-dimensional mesh model to be deformed, which is used to describe the local spatial transformation relationship in the three-dimensional mesh, and introduce a Jacobi regularization term to constrain the change of the vertex positions of the three-dimensional mesh; Add Gaussian-distributed random noise to the combined image and input it into the image encoder to obtain an image latent space representation. At the same time, input the target semantic text into the text encoder to obtain a text latent space representation; Use the image latent space representation and the text latent space representation as conditions to input into the second text-to-image model, and output the predicted value of the noise added to the combined image; Construct a second mean squared error loss function based on the difference between the predicted value and the random noise, and rely on the gradient backpropagation mechanism supported by the differentiable rendering module to conduct the gradient of the second mean squared error loss function with respect to the image pixels from the image domain to the corresponding vertex positions in the Jacobi field; Jointly construct a total optimization objective function by combining the second mean squared error loss function and the Jacobi regularization term, and perform multiple rounds of iterative optimization on the vertex positions of the Jacobi field until the loss value of the total optimization objective function converges, and obtain a deformed three-dimensional mesh model consistent with the target semantic text.
5. The three-dimensional data synthesis method according to claim 1, characterized in that The step of generating a texture image using a texture generation model based on the deformed three-dimensional mesh model and attaching it to the surface of the deformed three-dimensional mesh model to obtain synthetic three-dimensional data with texture includes: Input the deformed three-dimensional mesh model into the texture generation model, and the texture generation model automatically renders the input three-dimensional mesh based on multiple preset viewpoints to generate initial two-dimensional texture images from multiple viewpoints; The texture generation model processes multiple initial two-dimensional texture images through a viewpoint alignment mechanism, extracts local texture features from each viewpoint, and constructs a cross-view correspondence relationship; Based on the local texture features, generate a consistent texture representation in the latent space through a multi-view feature fusion strategy, and then generate a final two-dimensional texture image with consistent style and continuous structure; Paste the final two-dimensional texture image onto the vertices or faces of the deformed three-dimensional mesh model according to the mapping relationship with the three-dimensional mesh surface to obtain synthetic three-dimensional data with a complete texture map.
6. The three-dimensional data synthesis method according to claim 1, wherein It also includes: Conduct a validity evaluation on the synthetic three-dimensional data, specifically including: Mix the synthetic three-dimensional data with real three-dimensional data to construct a training data set and fine-tune the three-dimensional generation model; On the basis of a public three-dimensional test data set, conduct a performance evaluation on the three-dimensional generation model before and after fine-tuning. The performance evaluation indicators include: geometric accuracy indicators, shape diversity indicators, and generalization ability indicators; Based on the performance evaluation results, compare the performance differences of the three-dimensional generation model before and after fine-tuning to verify the effectiveness and applicability of the synthetic three-dimensional data in improving the performance of the three-dimensional generation model.
7. The three-dimensional data synthesis method according to claim 6, wherein The step of mixing the synthetic three-dimensional data with real three-dimensional data to construct a training data set and fine-tuning the three-dimensional generation model includes: Render each synthetic 3D data and real 3D data in the training dataset according to multiple preset orthogonal views respectively to generate two corresponding sets of 2D images; Input the two sets of 2D images as supervised images into the 3D generation model to guide the 3D generation model to predict the corresponding 3D mesh reconstruction result under the input view condition; Perform differentiable rendering on the 3D mesh reconstruction result to generate predicted 2D images, and construct a loss function based on the error at the pixel level between the predicted 2D images and the supervised images; Update the network weight parameters of the 3D generation model based on the loss function through the backpropagation mechanism to improve the geometric accuracy, shape diversity, and generalization ability of the 3D generation model under different views.
8. An electronic device, characterized in that, The electronic device includes: One or more processors; and a memory storing computer program instructions, which when executed cause the processors to execute the 3D data synthesis method according to any one of claims 1-7.
9. A computer-readable storage medium having a computer program and / or instructions stored thereon, characterized in that, The computer program and / or instructions, when executed by the processor, implement the 3D data synthesis method according to any one of claims 1-7.
10. A computer program product, comprising a computer program and / or instructions, characterized in that, The computer program and / or instructions, when executed by the processor, implement the 3D data synthesis method according to any one of claims 1-7.
Citation Information
Patent Citations
Three-dimensional generation method based on semantic enhancement hybrid reconstruction
CN119273871A
Universal 3D intelligent modeling method for multi-view coupling constraint
CN120088427A
Method for generating training data, image semantic segmentation method and electronic device
US20200160114A1
Cited By
Data-driven UV mapping generation method, electronic device and program product
CN120976397A
Image generation method, device and equipment, computer readable storage medium and computer program product
CN121170051A
An image generation method, device, equipment, computer readable storage medium and computer program product
CN121170051B
Multi-source data fusion collaborative completion method and system for digital intelligent material yard
CN121256720A
Three-dimensional model reconstruction method and device
CN121962535A