Three-dimensional data synthesis method, electronic device, storage medium and program product

By obtaining a textureless three-dimensional mesh model, constructing training sample pairs using multi-view rendering and semantic description, and fine-tuning the text graph model with low-rank adaptation technology, a model with multi-view consistent understanding capabilities is generated. This solves the problems of scarcity, insufficient diversity and limited model generalization capabilities in existing technologies, and achieves high-quality and high-diversity three-dimensional data synthesis.

CN120259590BActive Publication Date: 2025-09-19SHANG HAI JIE YUE XING CHEN ZHI NENG KE JI YOU XIAN GONG SI
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202510751286.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-06-06
Publication Date
2025-09-19
Estimated Expiration
2045-06-06

AI Technical Summary

Technical Problem

Existing technologies face problems such as data scarcity, insufficient diversity, and limited model generalization ability when generating high-quality and diverse 3D data.

Method used

By acquiring a textureless 3D mesh model, constructing training sample pairs using multi-view rendering and semantic description, and fine-tuning the text-based graph model with low-rank adaptation techniques, the team generated a model with multi-view consistent understanding capabilities. This model was then used to perform geometric deformation and texture generation on the 3D mesh, ensuring that the generated 3D data exhibited excellent semantic consistency, geometric structure recovery accuracy, and texture continuity.

Benefits of technology

It significantly improves the overall performance of synthetic 3D data, provides high-quality and highly diverse training data support, and enhances the reconstruction accuracy, diverse expression capabilities, and cross-category generalization capabilities of downstream 3D generation models.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120259590B_ABST
    Figure CN120259590B_ABST
Patent Text Reader

Abstract

This application provides a 3D data synthesis method, electronic device, storage medium, and program product. The method includes: obtaining multiple untextured 3D mesh models, rendering each 3D mesh model to obtain multiple perspective images, and constructing training sample pairs in combination with semantic description text; based on the training sample pairs, fine-tuning a first text-based image model using low-rank adaptation technology to obtain a second text-based image model with multi-perspective consistent understanding capabilities; modeling the latent space semantic difference between the rendered image of the deformable 3D mesh model and the target semantic text based on the second text-based image model, constructing optimization constraints and updating parameters to obtain a deformable 3D mesh model that conforms to the target semantics; and generating a texture image using a texture generation model and attaching it to the surface of the deformable 3D mesh model to obtain textured synthetic 3D data. This significantly enhances the reconstruction accuracy, diverse expression capabilities, and cross-category generalization capabilities of the synthetic 3D data.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of computer technology, and in particular to a three-dimensional data synthesis method, electronic equipment, storage medium, and program product. Background Art

[0002] With the continuous advancement of deep learning and computer graphics technologies, 3D generative models and multi-view diffusion models are showing broad application prospects in fields such as virtual reality, game development, and robotic navigation. These models can learn complex geometric structures and texture information from limited training samples, enabling the automated reconstruction and generation of 3D objects or scenes, becoming important tools for promoting digital content generation and improving 3D perception capabilities.

[0003] However, existing technologies still face the following significant challenges in generating high-quality and diverse 3D data:

[0004] Scarcity of 3D data: Compared with 2D data, the acquisition cost of 3D data is higher and the process is more complicated, resulting in a limited number of high-quality, well-structured 3D datasets, which are difficult to support the training needs of large-scale models.

[0005] Insufficient data diversity: Current mainstream public 3D datasets (such as ShapeNet) typically focus on limited object categories or standard scenes, lacking comprehensive coverage of the diverse morphologies, detailed variations, and complex backgrounds found in the real world. This limited sample structure restricts the model's ability to learn sufficient representation of variability during training, which in turn affects its performance on complex or novel objects.

[0006] Limited model generalization: Existing models generally rely on fixed real-world datasets for training, making them difficult to adapt to new categories or scenarios outside the training data distribution. When faced with unseen object structures or environmental conditions, generated results often exhibit distortion and loss of detail, indicating insufficient generalization performance and difficulty meeting the versatility and adaptability requirements of real-world applications. Summary of the Invention

[0007] In response to the shortcomings of the existing technology, the present application provides a three-dimensional data synthesis method, electronic device, storage medium and program product, which are at least used to solve the problems of scarcity of three-dimensional image data, insufficient data diversity and limited generalization ability of the generation model in the existing technology.

[0008] In order to achieve the above objectives and other advantages, some embodiments of the present application provide the following aspects:

[0009] In a first aspect, some embodiments of the present application provide a three-dimensional data synthesis method, including:

[0010] Acquire multiple texture-free three-dimensional mesh models, render each of the three-dimensional mesh models based on multiple preset perspectives to obtain multiple perspective images, and construct training sample pairs of images and texts in combination with semantic description texts;

[0011] Based on the training sample pairs, the pre-trained first text-based graph model is fine-tuned using a low-rank adaptation technique to obtain a second text-based graph model with multi-perspective consistent understanding capability;

[0012] Modeling the latent space semantic difference between the rendered image of the deformable 3D mesh model and the target semantic text based on the second text-generated graph model, constructing optimization constraints and updating the parameters of the deformable 3D mesh model to obtain a deformable 3D mesh model that conforms to the target semantics;

[0013] Based on the deformable three-dimensional mesh model, a texture image is generated using a texture generation model and attached to the surface of the deformable three-dimensional mesh model to obtain synthetic three-dimensional data with texture.

[0014] In a second aspect, some embodiments of the present application further provide an electronic device, comprising:

[0015] One or more processors; and a memory storing computer program instructions, wherein when the computer program instructions are executed, the processors execute any one of the three-dimensional data synthesis methods described above.

[0016] In a third aspect, some embodiments of the present application further provide a computer-readable storage medium having stored thereon a computer program and / or instructions, which, when executed by a processor, implements any of the three-dimensional data synthesis methods described above.

[0017] In a fourth aspect, some embodiments of the present application further provide a computer program product, comprising a computer program and / or instructions, which, when executed by a processor, implements any of the three-dimensional data synthesis methods described above.

[0018] Compared with the related art, the solution provided in the embodiment of the present application constructs training sample pairs of images and texts and uses low-rank adaptation technology to fine-tune the first text-based image model to obtain a second text-based image model with multi-view consistency understanding ability, thereby guiding the geometric deformation of the three-dimensional mesh model. Figure 1A consistent texture generation strategy ensures that the generated texture images have good continuity and style consistency under multiple viewing angles. Therefore, the 3D data synthesis method provided by this application effectively improves the comprehensive performance of synthesized 3D data in terms of semantic consistency, geometric structure recovery accuracy, and texture expression continuity. It provides high-quality and highly diverse training data support for downstream 3D generation models, thereby significantly enhancing their reconstruction accuracy, diverse expression capabilities, and cross-category generalization capabilities. BRIEF DESCRIPTION OF THE DRAWINGS

[0019] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other implementation methods can be obtained based on these drawings without paying any creative work.

[0020] Figure 1 This is a flow chart of a three-dimensional data synthesis method provided in an embodiment of the present application;

[0021] Figure 2 1 is a schematic diagram of a process for fine-tuning a cultural graph model based on low-rank adaptation provided in an embodiment of the present application;

[0022] Figure 3 This is a schematic diagram of a semantically guided three-dimensional mesh deformation optimization process provided by an embodiment of the present application;

[0023] Figure 4 The embodiment of the present application provides a multi-view Figure 1 Schematic diagram of the 3D mesh mapping process for consistent texture generation;

[0024] Figure 5 This is a schematic diagram of the training process of fine-tuning a 3D generation model based on mixed 3D data provided in an embodiment of the present application.

[0025] Figure 6 It is a structural diagram of an electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION

[0026] To make the purpose, technical solutions, and advantages of the embodiments of this application more clear, the technical solutions in the embodiments of this application will be clearly and completely described below in conjunction with the drawings in the embodiments of this application. Obviously, the described embodiments are part of the embodiments of this application, not all of the embodiments. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.

[0027] The Objaverse dataset, one of the most popular open-source 3D data resources, is primarily generated through manual capture, model reconstruction, or graphic synthesis, aiming to approximate the appearance of real-world 3D objects. While this dataset covers a wide range of object categories and exhibits good annotation consistency, due to the simplified construction process for some 3D models, it still suffers from certain deficiencies in geometric complexity and texture fidelity, making it difficult to meet the requirements for training generative models that require high 3D data.

[0028] The Diffusion Transformer (DiT) architecture is a graph-based architecture that integrates the Transformer network structure into the diffusion model. It combines the gradual denoising capabilities of diffusion probability modeling with the Transformer's strengths in modeling global context, enabling the generation and modeling of high-quality images or feature maps.

[0029] First embodiment

[0030] The first embodiment of the present application relates to a three-dimensional data synthesis method, referring to Figure 1 As shown, the method may include the following steps:

[0031] Step S1: Acquire multiple texture-free three-dimensional mesh models, render each three-dimensional mesh model based on multiple preset perspectives to obtain multiple perspective images, and construct training sample pairs of images and texts in combination with semantic description texts.

[0032] Specifically, in step S1, the original 3D mesh model in this embodiment can be from a public 3D dataset (such as the Objaverse dataset). To ensure the uniformity of the training data and avoid external texture interference, all texture information in the original 3D mesh model is removed, retaining only the geometric structure data.

[0033] Furthermore, in order to obtain multiple texture-free three-dimensional mesh models, the following steps may be included:

[0034] A plurality of three-dimensional mesh models with texture and material attributes are selected from a public dataset containing multiple categories of three-dimensional objects, and the three-dimensional mesh models are preprocessed. The preprocessing includes deleting texture maps, clearing material parameter bindings, and resetting mesh surface properties to obtain a white film three-dimensional mesh that only retains the geometric structure.

[0035] Specifically, multiple 3D mesh models with texture maps and material parameters are selected from a public 3D dataset, which may include but is not limited to Objaverse, ShapeNet, etc. The selected 3D meshes cover multiple object categories, such as furniture, tools, animals, vehicles, etc. A preprocessing operation is then performed on the selected 3D mesh models to obtain a white film 3D mesh with a unified style. The preprocessing includes: deleting the texture map files bound to the model (such as .jpg or .png format textures), clearing the material parameter binding information (such as PBR material, metalness, roughness, etc.), resetting the mesh surface properties (such as unifying the surface color, removing normal perturbations or highlight effects to make the surface appear neutral and textureless), etc. After the above processing, the obtained 3D mesh model only retains the original geometric structure of vertices, edges, and faces, thereby forming a structurally standardized and visually consistent texture-free 3D white film mesh.

[0036] Through the above preprocessing method, the surface style of the 3D model is unified during the training data construction stage, effectively avoiding the rendering style deviation caused by differences in material complexity. This helps the subsequent training model focus more on learning the relationship between structural modeling and semantic expression, thereby improving the model's ability to model 3D geometric features and cross-category generalization performance.

[0037] Utilize 3D modeling and rendering software (e.g., Blender) to render each de-textured 3D mesh model from multiple preset perspectives. For example, the preset perspectives may include: front view, back view, left view, right view, left front view, right front view, left rear view, right rear view, and top view, for a total of nine directions, to comprehensively capture the spatial geometric information of the object. The number of preset perspectives may also be four (e.g., front, rear, left, right) or six (e.g., front, rear, left, right, top, bottom), though this embodiment does not limit this.

[0038] The nine view images are stitched together in a predetermined order, arranged in a 3×3 grid, to generate a composite image containing complete view information, which is used as the training input image. The predetermined order can be view order, spatial position order, two-dimensional arrangement order, or any stitching order configured by the user or program.

[0039] For each stitched image, a corresponding semantic description is constructed. This description can be automatically generated based on manually defined language template rules, such as view order and object structural features (e.g., "front view shows the overall shape of the object," "top view shows the outline structure," etc.). Alternatively, the combined image can be used as input and guided by a large multimodal language model (such as GPT-4o). This guidance can include object category, typical structural features, and detailed expression requirements for each viewpoint. This approach can construct highly aligned and semantically complete training sample pairs, supporting the subsequent low-rank fine-tuning of the text-to-graph model.

[0040] Step S2: Based on the training sample pairs, the pre-trained first text-based graph model is fine-tuned using low-rank adaptation technology to obtain a second text-based graph model with multi-perspective consistent understanding capabilities.

[0041] Specifically, regarding step S2, the first text-based graph model of this embodiment employs a diffusion-based Transformer architecture (e.g., DiT). To achieve efficient parameter fine-tuning, while keeping the main weight parameters of the first text-based graph model frozen, low-rank adaptation modules (LoRA modules) are inserted into the attention mapping sublayer and feedforward sublayer of its multi-layer Transformer network.

[0042] During training, random noise conforming to a Gaussian distribution is first added to the composite image input to construct the denoising prediction task of the diffusion model. The model takes the noisy image at the current time step as input and learns to predict the residual of the added noise in the original composite image. The model training uses the L2 norm as the loss function to measure the difference between the predicted noise and the actual noise. The gradient is calculated based on this loss. Through backpropagation, only the trainable parameters in the LoRA module are updated, achieving lightweight fine-tuning of the first-generation image model.

[0043] After training is completed, the fine-tuned first text-based image model and the LoRA module with inserted and updated parameters are combined to form the second text-based image model. This model has the ability to model multi-perspective consistency and can accurately model the potential association between images with different perspective combinations and semantic text; at the same time, the second text-based image model learns the visual style features of the training images in the latent space and can generate images that are consistent or similar in style to the training images based on the input text.

[0044] Step S3: Based on the second text-based image model, the latent space semantic difference between the rendered image of the deformable 3D mesh model and the target semantic text is modeled, optimization constraints are constructed, and the parameters of the deformable 3D mesh model are updated to obtain a deformable 3D mesh model that conforms to the target semantics.

[0045] Specifically, step S3 involves rendering the deformed 3D mesh model from multiple perspectives into a composite image, and adding Gaussian noise. The encoded target semantic context is then fed into a fine-tuned second-context graph model along with the image latent vector. The noise is then predicted and the loss is back-propagated. Combining the Jacobian field structure with a regularization term, the gradient is propagated back to each vertex parameter, achieving geometric structure optimization driven by the target semantics.

[0046] For example, for a 3D mesh model representing a chair, the system first renders the model from multiple perspectives, including the front, back, left, right, left front, right front, left back, right back, and top, to obtain corresponding 2D perspective images. These images are then stitched together to form a nine-grid composite image. Subsequently, noise conforming to a Gaussian distribution is added to this composite image to construct the diffusion prediction input.

[0047] The target semantic text can be an automatically generated semantic description, such as "The nine-square grid shows different views of a chair. The chair consists of a square seat, four vertical legs, and a backrest extending from the rear edge." This text is input into the text encoder to obtain a latent space semantic vector. This, along with the latent vector extracted by the image encoder, is then input into the fine-tuned second text-to-image model to predict the noise added to the combined image.

[0048] A loss function is calculated based on the difference between the predicted noise and the actual noise. The gradient of this loss function is then transferred from the image space back to the vertex positions of the 3D mesh via a differentiable rendering mechanism. Leveraging a pre-constructed Jacobian field, this optimization signal is used to iteratively adjust the local geometric structure of the 3D mesh. Ultimately, while maintaining structural rationality, the 3D model of the "chair" is transformed into a new model structure that better aligns with the target semantics, such as a higher backrest and thicker legs, thus completing the target semantics-driven geometric optimization process.

[0049] Step S4: Based on the deformable three-dimensional mesh model, a texture image is generated using a texture generation model, and attached to the surface of the deformable three-dimensional mesh model to obtain synthetic three-dimensional data with texture.

[0050] Specifically, in step S4, after the geometric structure is deformed, the deformed 3D mesh model is fed into a texture generation model (e.g., the MvPaint model) to generate a texture image appropriate for the 3D mesh structure. The texture generation model can generate high-quality, continuous texture maps based on multiple perspective input images while maintaining stylistic consistency. This avoids the local jumps, edge misalignment, and stylistic discontinuities that can occur with traditional texture mapping methods.

[0051] In this embodiment, the MvPaint network receives as input 2D images rendered from multiple viewpoints of a deformable 3D mesh. A view alignment module extracts local texture features from these multiple viewpoints and fuses them into a consistent global texture representation. The network ultimately outputs one or more texture images, which can be attached to the 3D mesh surface using UV mapping, resulting in textured synthetic 3D data with consistent detail and style.

[0052] It is not difficult to find that compared with the related art, the solution provided by the embodiment of the present application, by constructing training sample pairs of images and texts and using low-rank adaptation technology to fine-tune the first text-based image model, obtains a second text-based image model with multi-view consistency understanding ability, thereby guiding the geometric deformation of the three-dimensional mesh model. Figure 1 A consistent texture generation strategy ensures that the generated texture images have good continuity and style consistency under multiple viewing angles. Therefore, the 3D data synthesis method provided by this application effectively improves the comprehensive performance of synthesized 3D data in terms of semantic consistency, geometric structure recovery accuracy, and texture expression continuity. It provides high-quality and highly diverse training data support for downstream 3D generation models, thereby significantly enhancing their reconstruction accuracy, diverse expression capabilities, and cross-category generalization capabilities.

[0053] Second embodiment

[0054] The second embodiment of this application relates to a method for three-dimensional data synthesis. The second embodiment is an improvement on the first embodiment. The specific improvement is that: in the second embodiment of this application, a specific implementation method for low-rank adaptation and fine-tuning of the first cultural graph model is provided. That is, step S2 can further include the following steps:

[0055] Step S201: On the basis of freezing the main parameters of the first text-generated graph model, inserting a low-rank adaptation module into the attention mapping sublayer and the feedforward sublayer in the multi-layer Transformer structure of the first text-generated graph model respectively;

[0056] Step S202: adding random noise obeying Gaussian distribution to the combined image of the training sample pair to construct a diffusion prediction task, wherein the combined image is formed by splicing multiple perspective images in a predetermined order;

[0057] Step S203: input the training sample pair into the first text-image model, and predict the noise added to the combined image through the diffusion modeling process;

[0058] Step S204: constructing a first mean square error loss function based on the difference between the predicted noise and the random noise and performing backpropagation of the gradient, and optimizing only the trainable parameters in the low-rank adaptation module while keeping the main weights of the model unchanged;

[0059] Step S205: When the first mean square error loss function reaches convergence, complete the fine-tuning training of the low-rank adaptation module, and construct a second cultural graph model based on the first cultural graph model and the optimized low-rank adaptation module.

[0060] Specifically, refer to Figure 2 As shown in the figure, on the basis of freezing the main parameters of the first-generation graph model (such as the diffusion transformer model based on the DiT architecture), the LoRA module is inserted into the attention mapping sublayer and feedforward sublayer of multiple transformer layers in the model. The LoRA module achieves efficient parameter reconstruction by introducing two sets of low-rank matrices, such as a pair of low-rank matrices with rank r and The matrix A and B are constructed and inserted into the linear mapping path of the specified layer in the first cultural graph model. Since r is much smaller than n, the degree of freedom after multiplying the matrices A and B is much smaller than a full-parameter matrix. In this way, only a few parameters need to be trained, and the entire large model structure can be fine-tuned, effectively adjusting the original model behavior with minimal parameter overhead.

[0061] For example, the first text-to-graph model is a pre-trained DiT-based image-text diffusion generative network with excellent image-text alignment capabilities. Building on this model architecture, the LoRA module is introduced for parameter fine-tuning, thereby constructing a second text-to-graph model with multi-view consistency modeling capabilities.

[0062] For each training sample, a de-textured 3D mesh model is rendered differently from multiple preset viewpoints (nine viewing directions) to obtain corresponding 2D images. These images are then concatenated into a composite image in a predetermined order, which can be arranged in a nine-square grid. To implement the diffusion modeling task, random noise conforming to a Gaussian distribution is added to the composite image as a perturbation to the image channels in the training input. This allows the model to learn how to denoise and restore the original image based on the image-text conditions during training, thereby establishing a latent space mapping capability between semantics and images.

[0063] The combined image with random noise and the corresponding semantic description text together form the conditional input of the diffusion model. The first text-based image model performs diffusion modeling on this basis and predicts the residual of the Gaussian noise added to the combined image, thereby establishing a semantic alignment mechanism between the image and text inputs. The first mean square error loss function is constructed based on the difference between the predicted noise output by the first text-based image model and the actual random noise added to the combined image. The L2 norm can also be used to measure the prediction error.

[0064] After calculating the loss function, gradient calculation and weight updates are performed through the backpropagation mechanism. To ensure the stability of the model's core structure, the Transformer backbone of the first-text image model is kept frozen during training, and only the trainable parameters in the LoRA module are updated with gradients. This significantly reduces parameter tuning overhead and enables effective transfer learning even with a small number of samples (only approximately 20-60 image and text pairs), thereby improving the model's modeling capabilities for textureless white film images.

[0065] When the first mean square error loss function converges to below the preset error threshold, or the training process reaches the set maximum number of iterations, the fine-tuning process is considered to be completed. The trained first text-based graph model and the optimized LoRA module together constitute the second text-based graph model described in this application. Based on the original visual generation capability of the pre-trained model, the text-based graph model can perceive and align the structural associations of the same three-dimensional object in multiple perspective images; at the same time, it can also generate images consistent with the style of the training samples based on the input semantic description text, thereby improving the quality of the latent space mapping from text to image. It improves the structural understanding and visual discrimination capabilities of white film style images, and provides a stable image semantic prior for downstream tasks such as three-dimensional geometry optimization.

[0066] It's easy to see that the solution provided in the embodiments of this application utilizes only a small number of image and text sample pairs during fine-tuning of the text graph model, resulting in a lightweight diffusion model that performs robustly in multi-view consistency modeling tasks. Compared to existing methods that require fine-tuning on a large number of RGB images, this application fine-tunes the model on detexturized white film data, resulting in a model with stronger generalization and semantic expression capabilities for real-world 3D geometric modeling tasks, providing stable image semantic prior support for subsequent 3D geometric deformation modeling.

[0067] Third embodiment

[0068] The third embodiment of the present application relates to a method for 3D data synthesis. The third embodiment is an improvement on the first embodiment. The specific improvement is that: in the third embodiment of the present application, a specific implementation method combining differentiable rendering, semantically guided optimization, and Jacobian field regularization constraints is provided for the geometric deformation modeling process. That is, step S3 can further include the following steps:

[0069] Step S301: Rendering the 3D mesh model to be deformed by a microscopic rendering module to obtain multiple rendered images at preset viewing angles and stitching them into a combined image;

[0070] Step S302: constructing a Jacobian field based on the topological structure of the 3D mesh model to be deformed, which is used to describe the local spatial transformation relationship in the 3D mesh, and introducing a Jacobian regularization term to constrain the change of the vertex position of the 3D mesh;

[0071] Step S303: adding random noise that follows a Gaussian distribution to the combined image and inputting it into an image encoder to obtain an image latent space representation, while simultaneously inputting the target semantic text into a text encoder to obtain a text latent space representation;

[0072] Step S304: Input the image latent space representation and the text latent space representation as conditions into the second text-image model, and output a predicted value of the noise added to the combined image;

[0073] Step S305: construct a second mean square error loss function based on the difference between the predicted value and the random noise, and rely on the gradient backpropagation mechanism supported by the differentiable rendering module to transfer the gradient of the second mean square error loss function with respect to the image pixel from the image domain to the corresponding vertex position in the Jacobian field;

[0074] Step S306: Combine the second mean square error loss function and the Jacobian regularization term to jointly construct the overall optimization objective function, perform multiple rounds of iterative optimization on the vertex positions of the Jacobian field until the loss value of the overall optimization objective function converges, and obtain a deformable three-dimensional mesh model consistent with the target semantic text.

[0075] Specifically, refer to Figure 3 As shown, the original 3D mesh model to be deformed is fed into a differentiable rendering module (e.g., the nvdiffrast.torch library, which provides differentiable rasterization and interpolation operations) and rendered from multiple preset viewpoints. The multiple viewpoint images are arranged and concatenated in a predetermined order to form a combined image, which serves as the input image for the fine-tuned second-level image generation model.

[0076] It's important to explain that differentiable rasterization projects 3D vertices onto a 2D plane, recording which triangles and vertices each pixel contributes most to, and retaining the gradient channel. Differentiable interpolation interpolates image attributes like color, normals, and UVs from vertices to pixels, and can propagate the error gradient back to the vertices.

[0077] It should be noted that the same viewing angle configuration and image arrangement order are used in the training sample construction stage and the geometric deformation optimization stage. For example, nine fixed viewing angle directions can be selected, including front, back, left, right, left front, right front, left back, right back and top, and spliced ​​into a combined image in a fixed order. The above-mentioned combined image structure remains consistent during the training of the cultural graph model and the three-dimensional mesh deformation process, which is conducive to the model's consistent perception ability and optimization constraints between different stages. This can significantly reduce the risk of semantic drift caused by input changes and improve the generalization ability and accuracy of the fine-tuned second cultural graph model in the geometric optimization stage.

[0078] Based on the topological structure of the original three-dimensional mesh model, a Jacobian field is constructed. The Jacobian field is used to describe the geometric transformation relationship between vertices in the local neighborhood of the three-dimensional mesh. Specifically, the Jacobian field expresses the local gradient information of the mesh during the deformation process by calculating the coordinate difference between adjacent patches or vertex pairs in the mesh. On this basis, the Jacobian regularization term is introduced as an optimization constraint to maintain geometric continuity and smoothness during the deformation process. This regularization term can prevent the three-dimensional mesh from having abnormal structures such as irregular deformation, topological fracture or local interpenetration during the optimization process by constraining the norm change or gradient smoothness of the Jacobian matrix, thereby ensuring that the final generated deformed mesh has good renderability and geometric rationality while maintaining semantic consistency.

[0079] Random noise following a Gaussian distribution is added to the combined image generated in step S301 to construct an image corruption input, simulating the noise prediction task of the diffusion model during training. This noise addition process is achieved by performing a Gaussian perturbation with a standard deviation of 𝜎 on the image pixel values. Subsequently, the noisy combined image is input into an image encoder (e.g., an image backbone network based on the ViT architecture) to extract the corresponding latent space representation vector for the image, which is used to describe the structure and style information in the multi-view combined image. Simultaneously, the target semantic text is input into a text encoder (e.g., the text branch in the CLIP model) to extract the corresponding text latent space representation, which is used to represent the user-specified target semantics.

[0080] The image latent space representation and text latent space representation obtained in step S303 are fed as conditional inputs to the fine-tuned second text-based image model. The model then predicts the random noise added to the image using a diffusion modeling mechanism, outputting a predicted value of the noise added to the combined image. This predicted value is then subtracted from the actual random noise added to the combined image to construct a loss function (such as mean square error loss or L2 norm loss) to measure the second text-based image model's ability to model image-text consistency under the current conditions.

[0081] Thanks to the use of a differentiable rendering module that supports gradient backpropagation, the loss function calculated by the model not only reflects the difference between the combined image and the predicted image, but also transmits the gradient of the loss function with respect to the image pixels layer by layer to the structural parameters of the 3D model during the backpropagation process. Specifically, the pixel error in the image first acts on the rendered image itself, then traces back to the vertex attributes on which the image depends through differentiable interpolation operations, and finally transmits it to the position parameters of each vertex in the 3D mesh model. Through this mechanism, the model is able to adjust the mesh structure based on the image error, achieving 3D deformation optimization guided by semantic goals.

[0082] The second loss function constructed in step S305 and the Jacobian regularization term introduced in step S302 are combined to form an overall optimization objective function. This optimization objective comprehensively considers the semantic consistency error in the image domain and the spatial smoothness constraints of the three-dimensional structural deformation, ensuring that the model maintains geometric stability while satisfying semantic guidance. Under this optimization objective, an iterative optimization algorithm (such as the Adam or Adagrad optimizer) is used to perform multiple rounds of updates to the position parameters of each vertex in the Jacobian field. During the optimization process, the total loss function, consisting of the prediction error of the second text-generated graph model and the Jacobian regularization term constraints, is gradually minimized until the loss value converges or the preset number of iterations is reached. After the loss function converges, the optimization process terminates. At this point, the 3D mesh model satisfies the requirement that rendered images from multiple perspectives are closer to the image described by the target semantic text. Furthermore, the 3D mesh structure is continuous and distortion-free, exhibiting good rendering and trainability.

[0083] Since the constructed Jacobian field records the transformation gradient information between vertices in the local area, in order to convert these local gradient information into a globally consistent vertex coordinate distribution, the Poisson equation solution mechanism is further introduced in the optimization process, that is, the gradient field extracted from the Jacobian field is restored to the deformed 3D mesh coordinates by solving the Poisson equation. The final output 3D mesh is the optimized deformed mesh model that meets the target semantic description. The model will be exported in the standard .ply format for subsequent texture generation or model training.

[0084] It is not difficult to find that in the solution provided by the embodiment of the present application, by introducing a diffusion text graph model for joint understanding of images and texts and a differentiable rendering mechanism that supports gradient backpropagation, the optimization of the three-dimensional mesh structure driven by semantic information is realized. With the help of the differentiable rendering module, the loss gradient is effectively transmitted from the image domain to the three-dimensional mesh structure, thereby forming an end-to-end optimization path from semantics to geometry. Furthermore, the local structural transformation is modeled in combination with the Jacobian field, and the local gradient information is restored to the global vertex coordinates by solving the Poisson equation, so that the optimized mesh has structural continuity and deformation rationality.

[0085] It should be noted that the third embodiment of the present application may also be an improvement based on any one or more of the first to second embodiments.

[0086] Fourth embodiment

[0087] The fourth embodiment of the present application relates to a three-dimensional data synthesis method. The fourth embodiment is an improvement on the first embodiment. The specific improvement is: In the fourth embodiment of the present application, a method based on multi-view Figure 1 The specific implementation of the texture generation of the consistency mechanism. That is, step S4 can further include the following steps:

[0088] Step S401: Inputting the deformable 3D mesh model into the texture generation model, the texture generation model automatically renders the input 3D mesh based on multiple preset view angles to generate initial 2D texture images under multiple view angles;

[0089] Step S402: The texture generation model processes the multiple initial two-dimensional texture images through a perspective alignment mechanism, extracts local texture features at each perspective, and constructs cross-perspective correspondences.

[0090] Step S403: Based on the local texture features, a consistent texture representation is generated in the latent space through a multi-view feature fusion strategy, thereby generating a final two-dimensional texture image with consistent style and continuous structure;

[0091] Step S404: Paste the final two-dimensional texture image to the vertices or facets of the deformable three-dimensional mesh model according to the mapping relationship between the final two-dimensional texture image and the three-dimensional mesh surface, so as to obtain synthetic three-dimensional data with a complete texture map.

[0092] For example, referring to Figure 4 As shown, the deformed 3D mesh model is input into a texture generation model, such as a deep network architecture based on the MvPaint framework. The texture generation model automatically renders the input mesh based on multiple preset viewpoints, generating multiple initial 2D texture images. Each 2D texture image records the texture information of the mesh surface at that viewpoint.

[0093] Furthermore, the texture generation model processes the initial two-dimensional texture images from multiple perspectives by introducing a perspective alignment mechanism, and extracts local texture features from each two-dimensional texture image, such as edge contours, texture directionality, surface details, etc. By constructing spatial or semantic correspondences between perspectives, local feature matching and mapping are achieved, avoiding texture style inconsistency or structural fracture problems caused by perspective changes. Subsequently, the texture generation model performs multi-level fusion of local texture features extracted from different perspectives in the latent space. The fusion strategy may include feature weighting guided by the attention mechanism, feature splicing after alignment, or residual fusion. Based on the consistent latent space representation, the final two-dimensional texture image with consistent style and continuous texture is further generated. Finally, the generated final two-dimensional texture image will be accurately pasted to the deformed three-dimensional mesh model surface according to the UV mapping or vertex attribute mapping mechanism in the three-dimensional mesh topology structure. The mapping process can be expanded for vertices or patches to ensure that the generated results are visually complete and coherent, thereby obtaining synthetic three-dimensional data with complete texture information. .

[0094] It is not difficult to find that in the solution provided in the embodiment of the present application, on the basis of maintaining the reasonable structure of the three-dimensional deformation model, by introducing a multi-view texture feature modeling mechanism, the consistency of texture expression and the ability to retain details are enhanced, and problems such as inconsistent multi-view mapping, texture jumps or blurring are avoided, thereby significantly improving the overall quality and realism of the synthesized three-dimensional data.

[0095] It should be noted that the fourth embodiment of the present application may also be an improvement based on any one or more of the first to third embodiments.

[0096] Fifth embodiment

[0097] The fifth embodiment of the present application relates to a method for synthesizing three-dimensional data. The fifth embodiment is an improvement on the first embodiment. The specific improvement is that: in the fifth embodiment of the present application, a specific implementation method for evaluating the effectiveness of synthesized three-dimensional data is provided, which specifically includes the following steps:

[0098] Step A1: Mix synthetic 3D data with real 3D data to construct a training dataset and fine-tune the 3D generation model.

[0099] Furthermore, step A1 specifically includes:

[0100] Step A101: Rendering each synthetic 3D data and real 3D data in the training data set according to a plurality of preset orthogonal perspectives to generate two corresponding sets of 2D images;

[0101] Step A102: Input the two sets of 2D images as supervision images into the 3D generative model, guiding the 3D generative model to predict corresponding 3D mesh reconstruction results under the input view condition;

[0102] Step A103: Perform differentiable rendering on the 3D mesh reconstruction result to generate a predicted 2D image, and construct a loss function based on the pixel-level error between the predicted 2D image and the supervision image;

[0103] Step A104: Based on the loss function, the network weight parameters of the 3D generative model are updated through a back-propagation mechanism to improve the geometric accuracy, shape diversity, and generalization ability of the 3D generative model under different perspectives.

[0104] Specifically, refer to Figure 5 As shown, from the synthetic 3D data A number of 3D object samples are selected from the real 3D data D. For each 3D object sample, differentiable rendering is performed based on multiple preset orthogonal perspectives (such as front view, back view, left view, and right view) to generate two sets of corresponding 2D images, which are respectively recorded as synthetic image groups and the real image group V. The above two groups of images ( The 3D reconstruction result generated in step A102 is again passed through the differentiable rendering module to generate a predicted 2D image from the corresponding viewpoint. This image is then compared pixel by pixel with the true supervisory image. A training loss function is constructed based on the pixel error between the predicted and true images and the perceptual loss:

[0105]

[0106] in, Represents a 2D image generated by the LGM model based on the input view conditional differentiable rendering; Represents the real supervision image under the corresponding perspective, which comes from the synthetic image and the real image V; The pixel-level mean square error is used to measure the direct difference between image color and structure; represents the perceptual loss, which is used to measure the similarity of images at the visual perception level; λ represents the adjustment coefficient, which is used to balance the contribution weights of perceptual loss and pixel loss in the total loss.

[0107] Step A2: Based on the public 3D test dataset, the performance of the 3D generative model before and after fine-tuning is evaluated. The performance evaluation indicators include: geometric accuracy index, shape diversity index, and generalization ability index.

[0108] Step A3: Based on the performance evaluation results, compare the performance differences of the 3D generative model before and after fine-tuning to verify the effectiveness and applicability of synthetic 3D data in improving the performance of the 3D generative model.

[0109] Based on publicly available standard 3D test datasets (such as the Objaverse test set or the ShapeNet test set), a performance comparison evaluation was conducted on the 3D generative models before and after fine-tuning, aiming to verify the improvement in structural recovery and generalization capabilities of the models trained by introducing synthetic 3D data.

[0110] For example, geometric accuracy can be evaluated using two metrics: Chamfer Distance (CD) and F-score. Chamfer Distance calculates the average point-to-point distance between the predicted and true meshes, reflecting the accuracy of the reconstructed geometry. F-score combines the matching degree and accuracy between the predicted and true point sets to comprehensively measure the reconstruction's performance in terms of detail and overall structure.

[0111] Shape diversity index evaluation can count the number of categories covered by the model during the generation phase, the range of deformation distribution within the category, or perform cluster analysis on the shape distribution after reconstruction of multiple different object categories to evaluate the model's coverage of object morphology and reflect its modeling diversity and expression ability.

[0112] Generalization metrics can be used to evaluate the model's generalization robustness and semantic transfer capabilities by applying it to reconstruction tasks using new categories or viewpoints not seen in the training set and assessing the change in output quality. For example, the model can be evaluated by predicting inputs for unseen categories in a zero-shot scenario and calculating the degree of accuracy degradation using metrics such as CD or LPIPS.

[0113] In public standard 3D test datasets or typical application scenarios, a systematic evaluation of the 3D generation model was conducted before and after the introduction of synthetic 3D data. Experimental results show that under the same model structure and training strategy, the mixed training of synthetic 3D data constructed by this application and real 3D data can significantly improve the overall performance of the model in multi-view 3D reconstruction tasks. After being applied to mainstream 3D generation models, the synthetic data generated can effectively improve the performance of the output model in terms of structural continuity, texture consistency and semantic alignment while maintaining realistic generation effects, helping to generate high-quality and structurally diverse 3D content.

[0114] It's easy to see that the solution provided in this embodiment of the application, by constructing a training process that uses a mix of synthetic and real 3D data and incorporating a standardized evaluation metric system, systematically compares the performance differences of the 3D generation model before and after fine-tuning. This quantitatively evaluates the improvement in model performance achieved by synthetic 3D data, objectively verifying the training value of synthetic data. Furthermore, this embodiment establishes a closed-loop mechanism from data construction, model training, to performance evaluation, providing a clear path for subsequent model tuning and data quality feedback.

[0115] It should be noted that the fifth embodiment of the present application may also be an improvement based on any one or more of the first to fourth embodiments.

[0116] From the above embodiments, it can be seen that compared with the high-cost real data obtained by manual annotation or scanning, the method of the present application effectively expands the size of the training set by automatically generating large-scale synthetic three-dimensional data, alleviates the problem of lack of three-dimensional data, and provides more sufficient learning samples for the deep model. This method combines semantically driven geometric deformation with multi-view consistency texture reconstruction technology, so that each synthetic sample has high variability in morphology and appearance, thereby significantly improving the coverage and representativeness of the training data set. The joint training of synthetic data and real data, guided by multi-view input conditions, improves the robustness and reconstruction accuracy of the model under unseen categories, unknown perspectives or complex shapes, and enhances the model's generalization ability for real scenes.

[0117] The step division of the above various methods is only for the purpose of clear description. During implementation, they can be combined into one step or some steps can be split and decomposed into multiple steps. As long as they include the same logical relationship, they are all within the scope of protection of this application; adding insignificant modifications or introducing insignificant designs to the algorithm or process without changing the core design of the algorithm and process are all within the scope of protection of this application.

[0118] In addition, some embodiments of the present application further provide an electronic device. The electronic device may be various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, etc. The electronic device may also be various forms of mobile devices, such as personal digital assistants, cellular phones, smartphones, wearable devices, and other similar computing devices.

[0119] The electronic device includes: one or more processors; and a memory storing computer program instructions, wherein when the computer program instructions are executed, the processor executes a three-dimensional data synthesis method provided by any one or more of the above embodiments. Figure 6An exemplary structural diagram of the electronic device is disclosed. The electronic device includes: one or more processors 1101, a memory 1102, and interfaces for connecting various components, including high-speed interfaces and low-speed interfaces. The various components are connected to each other using different buses and can be installed on a common motherboard or installed in other ways as needed. The processor can process instructions executed within the electronic device, including instructions stored in or on the memory to display graphical information of a GUI on an external input / output device (such as a display device coupled to the interface). In some other embodiments, if necessary, multiple processors and / or multiple buses can be used with multiple memories and multiple memories. Similarly, multiple electronic devices can be connected, with each device providing some of the necessary operations. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the present application described and / or required herein.

[0120] The electronic device may further include: an input device 1103 and an output device 1104. The processor 1101, the memory 1102, the input device 1103 and the output device 1104 may be connected via a bus or other means. Figure 6 The bus connection is taken as an example.

[0121] Input device 1103 can receive input digital or character information and generate key signal input related to user settings and function control of the electronic device. Examples include a touch screen, keypad, mouse, trackpad, touchpad, pointing stick, one or more mouse buttons, trackball, joystick, and other input devices. Output device 1104 may include a display device, auxiliary lighting devices (e.g., LEDs), and tactile feedback devices (e.g., vibration motors). The display device may include, but is not limited to, a liquid crystal display, a light emitting diode display, and a plasma display. In some embodiments, the display device may be a touch screen.

[0122] To provide user interaction, the electronic device may be a computer. The computer includes a display device (e.g., a cathode ray tube or LCD monitor) for displaying information to the user, and a keyboard and pointing device (e.g., a mouse) through which the user can provide input to the computer. Other types of devices may also be used to provide user interaction; for example, feedback provided to the user may be any form of sensory feedback (e.g., visual feedback, auditory feedback), and input from the user may be received in any form (e.g., voice input or tactile input).

[0123] In the embodiments of the present application, a computer-readable medium stores a computer program / instructions that, when executed by a processor, implements a three-dimensional data synthesis method provided by any one or more of the above-described embodiments. The computer-readable medium may be included in the electronic device described in the above embodiments, or it may exist independently and not be incorporated into the device. The computer-readable medium carries one or more computer-readable instructions.

[0124] The memory 1102 can be used as a non-transitory computer-readable storage medium to store non-transitory software programs, non-transitory computer executable programs, and modules. The processor 1101 executes the non-transitory software programs, instructions, and modules stored in the memory 1102 to execute various functional applications and data processing of the server, thereby implementing the program instructions / modules corresponding to the method provided in any one or more of the above embodiments of the present application.

[0125] The memory 1102 may include a program storage area and a data storage area, wherein the program storage area may store an operating system and applications required for at least one function; the data storage area may store data created based on the use of the electronic device, etc. In addition, the memory 1102 may include a high-speed random access memory, and may also include a non-transient memory, such as at least one disk storage device, a flash memory device, or other non-transient solid-state storage device. In some embodiments, the memory 1102 may optionally include a memory remotely located relative to the processor 1101, and these remote memories may be connected to the electronic device via a network. Examples of the above-mentioned network include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.

[0126] It should be noted that the computer-readable medium described in this application may be a computer-readable signal medium or a computer-readable storage medium or any combination of the above. Computer-readable media may be, for example, but not limited to: electrical, magnetic, optical, electromagnetic, infrared or semiconductor systems, devices or components, or any combination of the above. More specific examples of computer-readable storage media may include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory, a read-only memory, an erasable programmable read-only memory, an optical fiber, a portable compact disk read-only memory, an optical storage device, a magnetic storage device, or any suitable combination of the above. In this application, a computer-readable medium may be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, device or device.

[0127] Computer-readable media includes both permanent and non-permanent, removable and non-removable media, and can be implemented using any method or technology for information storage. The information can be computer-readable instructions, data structures, program modules, or other data. Examples of computer storage media include, but are not limited to, phase-change memory, static random access memory, dynamic random access memory, other types of random access memory, read-only memory, electrically erasable programmable read-only memory, flash memory or other memory technology, compact discs, digital versatile discs or other optical storage, magnetic cassettes, magnetic disk storage or other magnetic storage devices, or any other non-transmission medium that can be used to store information that can be accessed by a computing device.

[0128] Computer program code for performing the operations of the present application may be written in one or more programming languages, or a combination thereof, including object-oriented programming languages ​​such as Java, Smalltalk, C++, and conventional procedural programming languages ​​such as C or similar programming languages. The program code may be executed entirely on the user's computer, partially on the user's computer, as a stand-alone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving a remote computer, the remote computer may be connected to the user's computer via any type of network, including a local area network or a wide area network, or may be connected to an external computer (e.g., via the Internet using an Internet service provider).

[0129] In the above embodiments, all or part of the steps or functions of the present invention may be implemented using software, hardware, firmware, or any combination thereof. For example, implementation may be achieved using a dedicated integrated circuit, a general-purpose computer, or any other similar hardware device. In some embodiments, the software program of the present application may be executed by a processor to implement the above steps or functions. Similarly, the software program of the present application (including related data structures) may be stored in a computer-readable recording medium, such as a RAM memory, a magnetic or optical drive, a floppy disk, or the like. In addition, some steps or functions of the present application may be implemented using hardware, for example, as a circuit that cooperates with a processor to perform the various steps or functions.

[0130] The computer program product provided in the embodiments of the present application includes one or more computer programs / instructions that, when executed by a processor, fully or partially produce the processes or functions described in accordance with the embodiments of the present application. The computer may be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions may be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions may be transmitted from one website, computer, server, or data center to another website, computer, server, or data center via a wired (e.g., coaxial cable, optical fiber, digital subscriber line) or wireless (e.g., infrared, wireless, microwave, etc.) method. The computer-readable storage medium may be any available medium that can be accessed by a computer or a data storage device such as a server or data center that includes one or more available media. The available medium may be a magnetic medium (e.g., a floppy disk, a hard disk, a magnetic tape), an optical medium (e.g., a DVD), or a semiconductor medium (e.g., a solid-state drive).

[0131] The flowcharts or block diagrams in the accompanying drawings illustrate the possible architectures, functions and operations of the devices, methods and computer program products according to various embodiments of the present application. In this regard, each box in the flowchart or block diagram can represent a module, program segment or part of code, and the module, program segment or part of code contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the box can also occur in an order different from that marked in the accompanying drawings. For example, two boxes represented in succession can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram and / or flowchart, as well as the combination of boxes in the block diagram and / or flowchart, can be implemented with a dedicated hardware-specific system that performs the specified function or operation, or can be implemented with a combination of dedicated hardware and computer instructions.

[0132] The scope of this application is defined by the appended claims rather than the foregoing description and is therefore intended to encompass within this application all changes that come within the meaning and range of equivalents of the claims. Any reference signs in the claims should not be construed as limiting the claims to which they relate. In addition, it is clear that the word "comprising" does not exclude other units or steps, and the singular does not exclude the plural. Multiple units or devices stated in a device claim may also be implemented by one unit or device through software or hardware. Words such as "first" and "second" are only used to distinguish the description and do not indicate any particular order, nor should they be understood as indicating or implying relative importance.

[0133] The above descriptions are merely specific embodiments of the present application, but the scope of protection of the present application is not limited thereto. Any person skilled in the art may easily propose variations or substitutions within the technical scope disclosed in the present application, and such variations or substitutions shall be encompassed within the scope of protection of the present application. Therefore, the scope of protection of the present application shall be subject to the scope of protection of the claims, and the above descriptions shall be regarded as exemplary and non-limiting.

Claims

1. A three-dimensional data synthesis method, characterized in that: include: Acquire multiple texture-free three-dimensional mesh models, render each of the three-dimensional mesh models based on multiple preset perspectives to obtain multiple perspective images, and construct training sample pairs of images and texts in combination with semantic description texts; Based on the training sample pairs, the pre-trained first text-based graph model is fine-tuned using a low-rank adaptation technique to obtain a second text-based graph model with multi-perspective consistent understanding capability; Modeling the latent space semantic difference between the rendered image of the deformable three-dimensional mesh model and the target semantic text based on the second text-generated graph model, constructing optimization constraints and updating the parameters of the deformable three-dimensional mesh model to obtain a deformable three-dimensional mesh model that conforms to the target semantics, including: Rendering the three-dimensional mesh model to be deformed by a microscopic rendering module to obtain a plurality of rendered images under the preset viewing angles and stitching them into a combined image; Constructing a Jacobian field based on the topological structure of the three-dimensional mesh model to be deformed, for describing the local spatial transformation relationship in the three-dimensional mesh, and introducing a Jacobian regularization term to constrain the change of the vertex position of the three-dimensional mesh; Relying on the gradient backpropagation mechanism supported by the differentiable rendering module, the gradient of the image pixels is transmitted to the vertex position of the Jacobian field to adjust the structure of the three-dimensional network, thereby achieving three-dimensional deformation optimization guided by semantic goals; Based on the deformable three-dimensional mesh model, a texture image is generated using a texture generation model and attached to the surface of the deformable three-dimensional mesh model to obtain synthetic three-dimensional data with texture.

2. The three-dimensional data synthesis method according to claim 1, characterized in that: The step of obtaining multiple texture-free three-dimensional mesh models includes: selecting multiple three-dimensional mesh models with texture and material attributes from a public data set containing multiple categories of three-dimensional objects, and preprocessing the three-dimensional mesh models, wherein the preprocessing includes deleting texture maps, clearing material parameter bindings, and resetting mesh surface properties to obtain a white film three-dimensional mesh that only retains the geometric structure.

3. The three-dimensional data synthesis method according to claim 1, characterized in that: The step of fine-tuning the pre-trained first text-based graph model using a low-rank adaptation technique based on the training sample pairs to obtain a second text-based graph model with multi-perspective consistent understanding capability includes: On the basis of freezing the main parameters of the first cultural graph model, inserting low-rank adaptation modules into the attention mapping sublayer and the feedforward sublayer of the multi-layer Transformer structure of the first cultural graph model respectively; Adding random noise obeying a Gaussian distribution to a combined image of the training sample pair for constructing a diffusion prediction task, wherein the combined image is formed by splicing multiple perspective images in a predetermined order; Inputting the training sample pairs into the first Vincent graph model, and predicting the noise added to the combined image through a diffusion modeling process; Constructing a first mean square error loss function based on the difference between the predicted noise and the random noise and performing backpropagation of the gradient, and optimizing only the trainable parameters in the low-rank adaptation module while keeping the main weight of the model unchanged; Under the condition that the first mean square error loss function reaches convergence, fine-tuning training of the low-rank adaptation module is completed, and the second cultural graph model is jointly constructed based on the first cultural graph model and the optimized low-rank adaptation module.

4. The three-dimensional data synthesis method according to claim 1, characterized in that: The steps of modeling the latent space semantic difference between the rendered image of the deformable three-dimensional mesh model and the target semantic text based on the second text-based graph model, constructing optimization constraints, and updating the parameters of the deformable three-dimensional mesh model to obtain a deformable three-dimensional mesh model that conforms to the target semantics specifically include: Adding random noise obeying a Gaussian distribution to the combined image and inputting the noise into an image encoder to obtain an image latent space representation, and simultaneously inputting the target semantic text into a text encoder to obtain a text latent space representation; Inputting the image latent space representation and the text latent space representation as conditions into the second text-generated graph model, and outputting a predicted value of the noise added to the combined image; constructing a second mean square error loss function based on the difference between the predicted value and the random noise, and transmitting the gradient of the second mean square error loss function with respect to the image pixel from the image domain to the corresponding vertex positions in the Jacobian field; The second mean square error loss function and the Jacobian regularization term are combined to jointly construct the total optimization objective function, and the vertex positions of the Jacobian field are iteratively optimized for multiple rounds until the loss value of the total optimization objective function converges, thereby obtaining a deformable three-dimensional mesh model consistent with the target semantic text.

5. The three-dimensional data synthesis method according to claim 1, characterized in that: The step of generating a texture image based on the deformable three-dimensional mesh model using a texture generation model and attaching the texture image to the surface of the deformable three-dimensional mesh model to obtain textured synthetic three-dimensional data includes: Inputting the deformed three-dimensional mesh model into the texture generation model, the texture generation model automatically rendering the input three-dimensional mesh based on multiple preset perspectives to generate initial two-dimensional texture images under multiple perspectives; The texture generation model processes the multiple initial two-dimensional texture images through a perspective alignment mechanism, extracts local texture features under each perspective, and constructs cross-perspective correspondences; Based on the local texture features, a consistent texture representation is generated in the latent space through a multi-view feature fusion strategy, thereby generating a final two-dimensional texture image with consistent style and continuous structure; The final two-dimensional texture image is pasted to the vertices or facets of the deformable three-dimensional mesh model according to the mapping relationship between the final two-dimensional texture image and the three-dimensional mesh surface, so as to obtain synthetic three-dimensional data with complete texture mapping.

6. The three-dimensional data synthesis method according to claim 1, characterized in that: Also includes: Performing a validity evaluation on the synthesized three-dimensional data, specifically including: Mixing the synthetic 3D data with real 3D data to construct a training dataset, and fine-tuning the 3D generation model; Based on a public 3D test dataset, the performance of the 3D generative model before and after fine-tuning is evaluated, using performance evaluation metrics including geometric accuracy, shape diversity, and generalization capability. Based on the performance evaluation results, the performance difference of the 3D generative model before and after fine-tuning is compared to verify the effectiveness and applicability of the synthetic 3D data in improving the performance of the 3D generative model.

7. The three-dimensional data synthesis method according to claim 6, characterized in that: The step of mixing the synthetic 3D data with real 3D data to construct a training dataset and fine-tuning the 3D generation model includes: Rendering each synthetic 3D data and real 3D data in the training data set according to a plurality of preset orthogonal perspectives to generate two corresponding sets of 2D images; Inputting the two sets of two-dimensional images as supervision images into a three-dimensional generation model, guiding the three-dimensional generation model to predict corresponding three-dimensional mesh reconstruction results under input view conditions; Performing differentiable rendering on the three-dimensional mesh reconstruction result to generate a predicted two-dimensional image, and constructing a loss function based on the pixel-level error between the predicted two-dimensional image and the supervision image; Based on the loss function, the network weight parameters of the three-dimensional generative model are updated through a back-propagation mechanism to improve the geometric accuracy, shape diversity and generalization ability of the three-dimensional generative model under different perspectives.

8. An electronic device, characterized in that: The electronic device comprises: One or more processors; and a memory storing computer program instructions, wherein when the computer program instructions are executed, the processor executes the three-dimensional data synthesis method according to any one of claims 1 to 7.

9. A computer-readable storage medium having a computer program and / or instructions stored thereon, characterized in that: When the computer program and / or instructions are executed by a processor, the three-dimensional data synthesis method according to any one of claims 1 to 7 is implemented.

10. A computer program product comprising a computer program and / or instructions, characterized in that When the computer program and / or the instructions are executed by a processor, the three-dimensional data synthesis method according to any one of claims 1 to 7 is implemented.

Citation Information

Patent Citations

  • Universal 3D intelligent modeling method for multi-view coupling constraint

    CN120088427A