System and method for generating three-dimensional model from virtual reality / augmented reality three-dimensional sketch, processing system for three-dimensional model, editing method, and diffusion model training method

Drawing three-dimensional sketches through AR/VR devices and combining AI deep learning and diffusion models, the complex problems of existing 3D modeling methods are solved, and ordinary user-friendly three-dimensional model generation and editing are realized.

WO2025140611A1PCT designated stage expired Publication Date: 2025-07-03MOXIN (HUZHOU) TECH CO LTD

Patent Information

Application Number
PCT/CN2024/143325
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-12-27
Filing Date
2024-12-27
Publication Date
2025-07-03

AI Technical Summary

Technical Problem

The existing 3D modeling methods are too complex and time-consuming for ordinary users. The traditional 3D model editing methods are not intuitive and difficult to meet the needs of a wide range of creators.

Method used

It provides a three-dimensional model processing system, which uses AR/VR equipment to draw three-dimensional sketches, converts the sketches into three-dimensional models through computing devices, and generates physical models using a 3D printer or engraving machine, and combines AI deep learning and diffusion models for editing and generating.

Benefits of technology

The user-friendly three-dimensional model creation process is implemented, reducing the professional knowledge requirements, and allowing ordinary users to intuitively edit and generate high-quality 3D models.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2024143325_03072025_PF_FP_ABST
    Figure CN2024143325_03072025_PF_FP_ABST
Patent Text Reader

Abstract

Provided are a processing system for a three-dimensional model, a method for obtaining a physical model on the basis of three-dimensional printing of a three-dimensional processing system, a system for generating a three-dimensional model from a virtual reality / augmented reality three-dimensional sketch, a diffusion model training method, and a method for editing a three-dimensional model on the basis of a virtual reality / augmented reality three-dimensional sketch. The method for obtaining a physical model on the basis of three-dimensional printing of a three-dimensional processing system comprises: step one, at an input apparatus, a user using a hand or a handle to draw three-dimensional sketch content in the air, wherein a three-dimensional sketch is stored in the form of a three-dimensional point cloud, and when the user draws a trajectory, a corresponding point cloud is generated in the drawn trajectory and is displayed in a three-dimensional virtual space in a highlighted manner, and displacement and deletion operations of the user for the three-dimensional sketch are overall displacement and deletion operations for points in a region corresponding to the point cloud of the three-dimensional sketch; step two, a computing apparatus converting the three-dimensional sketch content into a three-dimensional model; step three, converting the three-dimensional model into a readable file for an execution apparatus, wherein the conversion is implemented by using a slicing algorithm; and step four, a three-dimensional printer executing the readable file to perform a printing operation, so as to obtain a physical model.
Need to check novelty before this filing date? Find Prior Art

Description

A system and method for generating a 3D model from a virtual reality / augmented reality 3D sketch, a 3D model processing system, an editing method, and a training method for a diffusion model Technical Field

[0001] The present invention relates to the technical field of computer graphics and computer-aided graphic design, and in particular to a method for generating a three-dimensional model from a virtual reality / augmented reality three-dimensional sketch, a three-dimensional processing method and a system. Background Art

[0002] In recent years, the demand for 3D content has increased significantly, driven by the rise of virtual and augmented reality technologies and the need for immersive experiences in various fields such as entertainment, education, and product design. However, creating professional 3D models is still not suitable for everyone. The widely adopted 3D modeling method using computer-aided design (CAD) software is a challenging and time-consuming process that requires a high level of expertise. CAD software often requires users to have extensive knowledge of 3D modeling techniques and the ability to use complex interfaces and tools. This creates a barrier for many potential creators, who may have great ideas but lack the necessary skills to bring them to life. Traditional 3D model editing relies on users manually dragging and dropping meshes, which is not conducive to intuitive content editing. Summary of the Invention

[0003] In order to solve the above problems, the present invention provides the following technical solutions:

[0004] The present invention provides a three-dimensional model processing system, including an execution device, a computing device and an input device, wherein the execution device is a 3D printer and / or a milling machine and / or an engraving machine, and the input device is an AR (augmented reality) glasses or a VR (virtual reality) helmet. The AR / VR device is electrically connected to the computing device, and the user draws the three-dimensional sketch content in the air with his hands or a handle. The computing device converts the three-dimensional sketch content into a three-dimensional model, and then converts the three-dimensional model into a readable file of the execution device, such as a file in GCode format. The file is input into the execution device to obtain a three-dimensional physical object, wherein the three-dimensional sketch is stored in the form of a three-dimensional point cloud. When the user draws a trajectory, a corresponding point cloud is generated in the drawn trajectory and highlighted in the three-dimensional virtual space; the user's displacement and deletion operations on the three-dimensional sketch are the overall displacement and deletion operations of the points in the corresponding area of ​​the point cloud of the three-dimensional sketch.

[0005] The present invention provides a method for obtaining a physical model by three-dimensional printing based on a three-dimensional processing system, comprising:

[0006] In the first step, the user uses their hand or a controller to draw a 3D sketch in the air. The 3D sketch is stored as a 3D point cloud. When the user draws a trajectory, the corresponding point cloud is generated and highlighted in the 3D virtual space. The user can also perform displacement and deletion operations on the 3D sketch, which are operations on the entire point in the corresponding area of ​​the 3D sketch point cloud.

[0007] In the second step, the computing device converts the three-dimensional sketch content into a three-dimensional model;

[0008] The third step is to convert the 3D model into a file that can be read by the execution device, such as a file in GCode format, and the conversion is achieved using a slicing algorithm;

[0009] In the fourth step, the 3D printer executes the readable file to perform printing operations to obtain a physical model.

[0010] Furthermore, the specific process of the first step is:

[0011] 1.1 The user specifies the drawing pose, such as "front view", "top view", "side view", or the rotation pose direction angle or the cube box or coordinate axis used to indicate the direction, or draws according to the given coordinate axis;

[0012] 1.2 The user adds a text description of the drawn content, such as "a horse with wings";

[0013] 1.3 The user selects an existing model and draws on the existing model in the drawn cube box;

[0014] 1.4 The user selects an existing model or picture, which can be displayed for reference in another floating window.

[0015] Furthermore, in the second step, the conversion uses an AI deep learning model, wherein the AI ​​deep learning model is an encoder-decoder structure model pre-trained on sketch-3D model pairs or a diffusion model pre-trained on a large amount of 3D model data with a sketch-based constraint.

[0016] Alternatively, the conversion process is an optimization-based face filling algorithm that fills the faces between the sketch strokes and then optimizes the shape.

[0017] Furthermore, the method further includes: preprocessing the three-dimensional sketch drawn by the user, and the preprocessing step includes:

[0018] 2.1 Point Sampling,The 3D sketch drawn by the user is represented by a point cloud, and N points are uniformly sampled on the trajectory of the sketch as input;

[0019] 2.2 Position alignment: move the center point of the sketch to the center of the bounding box of size M*N*P;

[0020] 2.3 Size alignment: calculate the spatial size of the drawn 3D sketch and scale it to a fixed size M*N*P; preferably, M=N=P=1

[0021] 2.4 Pose estimation: Use the pose estimation algorithm to estimate the pose of the sketch, expressed as pitch angle and rotation angle;

[0022] 2.5 Pose Alignment, using the pose estimated in 2.4, rotate the sketch to a specific pre-defined pitch angle and rotation angle.

[0023] Furthermore, the pose estimation algorithm is specifically as follows:

[0024] A multi-layer perceptron is used to encode the sketch, and the multi-layer perceptron outputs two pose data through an activation function; the multi-layer perceptron is trained using a large number of sketches with different poses and real pose results, so that the neural network outputs a pose that is closer to the real pose;

[0025] Alternatively, a self-supervised learning framework is designed, which uses SE(3) (special Euclidean group) equivariance to decouple the object shape and posture, including an invariant shape reconstruction module and an equivariant posture estimation module, and a self-supervised loss function is designed to achieve self-supervised learning through the consistency loss between the input point cloud and the transformed canonical shape reconstruction;

[0026] The invariant shape reconstruction module is used to learn to reconstruct the object shape aligned with the category-level reference frame, ensuring that the reconstructed shape remains unchanged under SE(3) transformation. The features are processed using an element-level multi-layer perceptron (MLP), and then max-pooling is used to obtain global SE(3) invariant features. Finally, these features are converted into point clouds through an MLP-based point generator.

[0027] The equivariant pose estimation module is used to estimate the pose of the object so that it is equivariant to the SE(3) transformation. An element-wise MLP is applied to each feature independently to generate equivariant global features through maximum pooling. Then, another MLP is used to process each global feature to predict the residual quaternion and residual translation to achieve pose estimation.

[0028] Among them, Equivariant Point Network (EPN) is used to extract features from point clouds to ensure that the features remain equivariant under SE(3) transformation.

[0029] Furthermore, the method also includes the user editing the obtained three-dimensional model, including adjusting the size, position, or local geometric information; or using a sketch to edit and adjust the obtained three-dimensional model.

[0030] The present invention provides a system for generating a 3D model from a virtual reality / augmented reality 3D sketch, comprising:

[0031] A sketch encoding module, comprising a Transformer-based point cloud encoder, for encoding a 3D sketch S into a latent code, wherein the latent code is a feature vector in a high-dimensional space; preferably, the sketch encoding module can be pre-trained on a large number of point cloud and text and / or image pairs using a contrastive loss to align the point cloud to the text and / or image space;

[0032] Diffusion model, a given sample Z is gradually introduced into Gaussian noise, and Gaussian noise z is generated according to the predetermined variance distribution t (t∈[1,T]); then a time-conditioned 3D Unet is used for denoising, outputting discrete 3D voxels - signed distance fields (SDFs). SDF voxels calculate the signed distance from the center of each grid cell to the nearest shape surface; a grid, i.e., a zero isosurface, can be derived from these grids using the Marching Cube algorithm; the SDF voxels can be converted into discrete occupied voxels, where each grid cell stores binary occupancy based on whether the absolute value of its SDF is below a predefined threshold; alternatively, the SDF can also be an unsigned distance field (UDF) voxel; the 3D voxels are then passed through an octree-based convolutional neural network to obtain a 3D model;

[0033] 3D Unet consists of a 5-level UNet and a 4-level UNet. This is because high-fidelity 3D shape representation requires the use of high-resolution discrete signed distance fields (SDFs) to model details; however, fully generating dense SDF grids will incur unacceptable computational and memory costs due to cubic complexity; to avoid the huge computational burden while still maintaining high-quality model generation, a two-stage diffusion framework is used, utilizing a self-conditioning continuous diffusion model; a low-resolution 3D occupancy voxel is generated through a 5-level UNet A preliminary approximation of the 3D shape; a high-resolution sparse voxel V∈R is constructed through a 4-level UNet N*N*N*4 ; Both the 5-level UNet and the 4-level UNet are trained with denoising loss, specifically:

[0034] Where N(0,1) represents Gaussian distribution, Represents a condition applied to the build process; a condition is injected into the diffusion-based generation process to accept user input; the conditional signal c (in this embodiment, the conditional signal C is the implicit code encoded by the sketch encoding module) is converted into a 1024-dimensional latent code I as the conditional signal by using the task-specific encoder; then, I is injected into the UNet using multi-head cross attention; given the conditional input, the diffusion model may generate samples with limited diversity; to alleviate this problem, a method without classifier guidance is adopted;

[0035] During use, the diffusion model is first pre-trained based on a large number of 3D model sample pairs corresponding to sketches, point clouds, pictures or texts. After the pre-training is completed, the user inputs a 3D sketch, which is encoded into a latent code by the sketch encoding module and then injected into the diffusion model through multi-head cross attention, guiding the diffusion model to output 3D model content that conforms to the guidance of the 3D sketch.

[0036] The present invention provides a method for training a diffusion model, which includes first training a generative model using a large amount of existing data (such as image-3D model pairs or point cloud-3D model pairs), and then fine-tuning the encoder of the generative model using a small amount of desired data (3D sketch-3D model in this example), while keeping the entire generative model unchanged during the fine-tuning.

[0037] Furthermore, the encoder and the generative model are trained together.

[0038] Furthermore, the training method specifically includes:

[0039] Generative pre-training: The goal of this stage is to train a diffusion model to generate 3D shapes from the latent space; use a large number of high-quality 3D shapes to train the diffusion model and use rendered images to regulate the diffusion process; the rendered images are generated by a pre-trained encoder E pretrain (·) encoding, where the pre-trained encoder maps the image condition to the CLIP latent space with a fixed pre-trained CLIP encoder;

[0040] Both 4-level UNet and 5-level UNet are trained using denoising loss:

[0041] Among them, N(0,1) represents Gaussian distribution, Represents the conditions applied to the generation process. In this stage, it is a picture constraint. The picture can also be replaced by text and / or point cloud. The pre-trained encoder can also be Uni3D or BLIP or other pre-trained encoders.

[0042] 3D Sketch Mapping: Map the 3D sketch input to the CLIP latent space, keeping the pre-trained diffusion model in generative pre-training unchanged, and then fine-tune the weights of the sketch encoder;

[0043] Joint Fine-tuning: Simultaneously fine-tuning the 3D sketch encoder E sketch (·) and diffusion model to ensure significantly improved spatial semantic alignment. This step also fine-tunes the 3D sketch encoder E sketch (·) and diffusion models to ensure significantly improved spatial semantic alignment; this approach has demonstrated its benefits in maximizing the utilization of pre-trained knowledge and is crucial for achieving better quality;

[0044] Furthermore, because there is only limited training data of 3D sketch and 3D model pairs, the network has difficulty generalizing across a wide range of sketch styles and geometric variations when mapping abstract 3D sketches to a latent space to condition the generative model; experiments have shown that the network occasionally fails to accurately reconstruct complex local details; in these areas, the network has difficulty in performing precise parameterization due to insufficient smoothness in the implicit function; even small errors can lead to sign errors, resulting in inaccurate surface reconstructions; in course learning, careful selection and sequencing of samples from simple to complex are crucial for effective and gradual skill development; the choice is based on the sample difficulty score;

[0045] Furthermore, in the second and / or third steps, additional curriculum learning is introduced into the training process. During the training process, the difficulty of each sample in the current training state is measured, and then the number of sample pools participating in the training and the speed of adding new samples are dynamically adjusted based on the difficulty of the samples, and difficult samples are gradually added in order from simple to difficult samples.

[0046] Furthermore, the training process of the additional course learning is:

[0047] Points with incorrect estimates are considered hard samples, points with correct estimates are considered easy samples, and points between 0 and the ground truth are considered semi-hard samples, using the following difficulty score:

[0048] Among them, y is the SDF value of the real sample, is the predicted SDF value, 0≤λ≤1 controls the importance of hard and semi-hard samples; if v≥0, sgn(v)=1, otherwise, sgn(v)=-1; where v represents the input SDF value;

[0049] Adaptive course: First, use the trained 3D model to generate the network to obtain the initial difficulty score, and sort the pre-sorted dataset X in ascending order according to the current difficulty score s; then, use the rhythm function p(·) to form a sample pool X′. Specifically, the sorted dataset is divided according to the sample pool size of each small batch obtained by the rhythm function to form a new sample pool; a small batch B′=[B′1,...,B′ M ] to train the target network; the rhythm function p(·) is a monotonically increasing function that determines the speed of learning from simple to complex samples; finally, at the end of the forward propagation, the difficulty score s is updated and the new sample pool X′ is calculated; where B′ represents the set of all small batches; B′ M Indicates a small batch;

[0050] Difficulty scores also apply to different training durations, the (k+1)th th The difficulty score of a position can be expressed as: s k+1 =(1-α)s k +αs cur

[0051] in, m represents the mth mini-batch; inv controls the frequency of difficulty score updates, α controls the speed of difficulty score updates; Scur represents the difficulty score during training;

[0052] During the training process, the difficulty of each sample in the current training state is measured, and then the number of sample pools participating in the training and the speed of adding new samples are dynamically adjusted based on the difficulty of the samples, and difficult samples are gradually added in order from simple to difficult samples.

[0053] Furthermore, in order to manage the speed of sample learning, a monotonically increasing rhythm function p(·) is set to limit the size of the sample pool X′. The function is expressed as:

[0054] Where n is the total number of samples; p0 is the sample ratio of the initial step; q is the speed at which the sample ratio increases; r0 is the frequency at which the sample ratio increases, and i is the current step.

[0055] The present invention provides a method for editing a three-dimensional model based on a virtual reality / augmented reality three-dimensional sketch, comprising:

[0056] In the first step, a suitable viewing angle is selected. The user selects the area M to be edited on the existing model;

[0057] In the second step, the user uses an AR or VR device to draw the rough shape of the desired modified model in the area;

[0058] The third step is to use the three-dimensional diffusion model to perform editing based on the original model and the area M to be edited;

[0059] The fourth step is to encode the text description into a latent code, where the text description comes from user input or from multiple sets of prompts generated by a large language model (LLM) based on user prompts, where the prompts are used to guide the diffusion model to generate content that conforms to the text description.

[0060] Furthermore, the specific process of the third step is:

[0061] The region to be edited M guides the diffusion by specifying which areas of the model should be modified or edited; the diffusion model follows the guidance of the region to be edited M and selectively applies modifications to the corresponding areas. In each diffusion step, the three-dimensional information is modified according to the mask guidance; the modification can involve changing the color, texture or other properties of the selected area; the diffusion model ensures that these edits are applied gradually and are consistent with the surrounding model content; specifically, this process is expressed by the following equation:

[0062] Among them, z t represents the implicit code after adding noise, y t It represents the result after adding noise continuously during DDIM sampling; M is the mask (the area to be edited);

[0063] This step is performed N times, N>=1, until the required diffusion steps are completed; with each iteration, the 3D model is further improved and the editing effect becomes more obvious; the diffusion process ensures that the modifications are smoothly integrated with the rest of the 3D model.

[0064] Furthermore, in the third step, all the layers of the diffusion model are traversed and the cross attention and self-attention maps M are operated. attn ;

[0065] The new mask m is obtained by downsampling the mask M (the area to be edited) according to the resolution of the cross-attention layer or the self-attention layer; for the cross-attention map, the input and output column indices J are determined according to the token pointing to the content to be edited and the content not to be edited. in and J out ; For the self-attention map, the column index J is determined according to whether the three-dimensional primitive belongs to in and J out , using J in and J out Get the new attention map:

[0066] Using the new attention map, we can get the constrained noise z in each denoising stage cond , the noise without mask constraint is recorded as zuncond , calculate z t-1 ←z cond +w g *(z cond -z uncond )

[0067] Among them, w g represents the weighted value of unconditional noise;

[0068] Furthermore, the first to fourth steps are repeated, with different Ri, Ti selected in the first step and different combinations of text prompts T selected in the fourth step, thereby obtaining diversified three-dimensional models under different editing conditions.

[0069] The present invention has the following beneficial effects:

[0070] The present invention pre-trains the diffusion model based on a large number of three-dimensional model sample pairs corresponding to sketches or point clouds or pictures or texts. After the pre-training is completed, the user inputs a three-dimensional sketch, which is encoded into a latent code by the sketch encoding module and then injected into the first stage of the diffusion model through multi-head cross attention, guiding the diffusion model to output three-dimensional model content that conforms to the guidance of the three-dimensional sketch. The present invention can edit the content more intuitively. BRIEF DESCRIPTION OF THE DRAWINGS

[0071] FIG1 is a schematic diagram of a system for generating a three-dimensional model in the present invention. DETAILED DESCRIPTION

[0072] The specific embodiments of the present invention are described in detail below with reference to the accompanying drawings. It should be pointed out that the embodiments are only specific explanations of the invention and should not be regarded as limitations of the invention. The purpose of the embodiments is to enable those skilled in the art to better understand and reproduce the technical solutions of the present invention. The scope of protection of the present invention shall still be based on the scope defined by the claims.

[0073] Example 1

[0074] This embodiment provides a three-dimensional model processing system, comprising an execution device, a computing device, and an input device. The execution device is a 3D printer and / or milling machine and / or engraving machine, and the input device is an AR (augmented reality) pair of glasses or a VR (virtual reality) helmet. The AR / VR device is electrically connected to the computing device. A user draws a three-dimensional sketch in the air using their hands or a controller. The computing device converts the sketch into a three-dimensional model, which is then converted into a file readable by the execution device, such as a GCode file. This file is input into the execution device to produce a three-dimensional object. The three-dimensional sketch is stored as a three-dimensional point cloud. When a user draws a trajectory, a corresponding point cloud is generated and highlighted in the three-dimensional virtual space. User displacement and deletion operations on the three-dimensional sketch are equivalent to overall displacement and deletion operations on the points in the corresponding area of ​​the three-dimensional sketch's point cloud.

[0075] Example 2

[0076] This embodiment provides a method for obtaining a physical model by three-dimensional printing based on a three-dimensional processing system, comprising:

[0077] In the first step, the user uses his hand or handle on the input device to draw the three-dimensional sketch content in the air; the three-dimensional sketch is stored in the form of a three-dimensional point cloud. When the user draws a trajectory, the corresponding point cloud is generated in the drawn trajectory and highlighted in the three-dimensional virtual space; the user's displacement and deletion operations on the three-dimensional sketch are the overall displacement and deletion operations of the points in the corresponding area of ​​the point cloud of the three-dimensional sketch.

[0078] In the second step, the computing device converts the three-dimensional sketch content into a three-dimensional model;

[0079] The third step is to convert the 3D model into a file that can be read by the execution device, such as a file in GCode format, and the conversion is achieved using a slicing algorithm;

[0080] In the fourth step, the 3D printer executes the readable file to perform printing operations to obtain a physical model.

[0081] Furthermore, the specific process of the first step is:

[0082] 1.1 The user specifies the drawing pose, such as "front view", "top view", "side view", or the rotation pose direction angle or the cube box or coordinate axis used to indicate the direction, or draws according to the given coordinate axis;

[0083] 1.2 The user adds a text description of the drawn content, such as "a horse with wings";

[0084] 1.3 The user selects an existing model and draws on the existing model in the drawn cube box;

[0085] 1.4 The user selects an existing model or picture, which can be displayed for reference in another floating window.

[0086] Furthermore, in the second step, the conversion uses an AI deep learning model, wherein the AI ​​deep learning model is an encoder-decoder structure model pre-trained on sketch-3D model pairs or a diffusion model pre-trained on a large amount of 3D model data with a sketch-based constraint.

[0087] Alternatively, the conversion process is an optimization-based face filling algorithm that fills the faces between the sketch strokes and then optimizes the shape.

[0088] Furthermore, the method further includes: preprocessing the three-dimensional sketch drawn by the user, and the preprocessing step includes:

[0089] 2.1 Point Sampling,The 3D sketch drawn by the user is represented by a point cloud, and N points are uniformly sampled on the trajectory of the sketch as input;

[0090] 2.2 Position alignment: move the center point of the sketch to the center of the bounding box of size M*N*P;

[0091] 2.3 Size alignment: calculate the spatial size of the drawn 3D sketch and scale it to a fixed size M*N*P; preferably, M=N=P=1

[0092] 2.4 Pose estimation: Use the pose estimation algorithm to estimate the pose of the sketch, expressed as pitch angle and rotation angle;

[0093] 2.5 Pose Alignment, using the pose estimated in 2.4, rotate the sketch to a specific pre-defined pitch angle and rotation angle.

[0094] Furthermore, the pose estimation algorithm is specifically as follows:

[0095] A multi-layer perceptron is used to encode the sketch, and the multi-layer perceptron outputs two pose data through an activation function; the multi-layer perceptron is trained using a large number of sketches with different poses and real pose results, so that the neural network outputs a pose that is closer to the real pose;

[0096] Alternatively, a self-supervised learning framework is designed, which uses SE(3) (special Euclidean group) equivariance to decouple the object shape and posture, including an invariant shape reconstruction module and an equivariant posture estimation module, and a self-supervised loss function is designed to achieve self-supervised learning through the consistency loss between the input point cloud and the transformed canonical shape reconstruction;

[0097] The invariant shape reconstruction module is used to learn to reconstruct the object shape aligned with the category-level reference frame, ensuring that the reconstructed shape remains unchanged under SE(3) transformation. The features are processed using an element-level multi-layer perceptron (MLP), and then max-pooling is used to obtain global SE(3) invariant features. Finally, these features are converted into point clouds through an MLP-based point generator.

[0098] The equivariant pose estimation module is used to estimate the pose of the object so that it is equivariant to the SE(3) transformation. An element-wise MLP is applied to each feature independently to generate equivariant global features through maximum pooling. Then, another MLP is used to process each global feature to predict the residual quaternion and residual translation to achieve pose estimation.

[0099] Among them, Equivariant Point Network (EPN) is used to extract features from point clouds to ensure that the features remain equivariant under SE(3) transformation.

[0100] Furthermore, the method also includes the user editing the obtained three-dimensional model, including adjusting the size, position, or local geometric information; or using a sketch to edit and adjust the obtained three-dimensional model.

[0101] Example 3

[0102] This embodiment provides a system for generating a 3D model from a virtual reality / augmented reality 3D sketch, including:

[0103] A sketch encoding module, comprising a Transformer-based point cloud encoder, for encoding a 3D sketch S into a latent code, wherein the latent code is a feature vector in a high-dimensional space; preferably, the sketch encoding module can be pre-trained on a large number of point cloud and text and / or image pairs using a contrastive loss to align the point cloud to the text and / or image space;

[0104] Diffusion model, a given sample Z is gradually introduced into Gaussian noise, and Gaussian noise z is generated according to the predetermined variance distribution t(t∈[1,T]); then a time-conditioned 3D Unet is used for denoising, outputting discrete 3D voxels - signed distance fields (SDFs). SDF voxels calculate the signed distance from the center of each grid cell to the nearest shape surface; a grid, i.e., a zero isosurface, can be derived from these grids using the Marching Cube algorithm; the SDF voxels can be converted into discrete occupied voxels, where each grid cell stores binary occupancy based on whether the absolute value of its SDF is below a predefined threshold; alternatively, the SDF can also be an unsigned distance field (UDF) voxel; the 3D voxels are then passed through an octree-based convolutional neural network to obtain a 3D model;

[0105] 3D Unet consists of a 5-level UNet and a 4-level UNet. This is because high-fidelity 3D shape representation requires the use of high-resolution discrete signed distance fields (SDFs) to model details; however, fully generating dense SDF grids will incur unacceptable computational and memory costs due to cubic complexity; to avoid the huge computational burden while still maintaining high-quality model generation, a two-stage diffusion framework is used, utilizing a self-conditioning continuous diffusion model; a low-resolution 3D occupancy voxel is generated through a 5-level UNet A preliminary approximation of the 3D shape; a high-resolution sparse voxel V∈R is constructed through a 4-level UNet N*N*N*4 ; Both the 5-level UNet and the 4-level UNet are trained with denoising loss, specifically:

[0106] Where N(0,1) represents Gaussian distribution, Represents a condition applied to the build process; a condition is injected into the diffusion-based generation process to accept user input; the conditional signal c (in this embodiment, the conditional signal C is the implicit code encoded by the sketch encoding module) is converted into a 1024-dimensional latent code I as the conditional signal by using the task-specific encoder; then, I is injected into the UNet using multi-head cross attention; given the conditional input, the diffusion model may generate samples with limited diversity; to alleviate this problem, a method without classifier guidance is adopted;

[0107] During use, the diffusion model is first pre-trained based on a large number of 3D model sample pairs corresponding to sketches, point clouds, pictures or texts. After the pre-training is completed, the user inputs a 3D sketch, which is encoded into a latent code by the sketch encoding module and then injected into the diffusion model through multi-head cross attention, guiding the diffusion model to output 3D model content that conforms to the guidance of the 3D sketch.

[0108] Example 4

[0109] This embodiment provides a method for training a diffusion model, including first training a generative model using a large amount of existing data (e.g., image-3D model pairs or point cloud-3D model pairs), then fine-tuning the encoder of the generative model using a small amount of desired data (in this example, 3D sketch-3D model pairs), while keeping the entire generative model unchanged during the fine-tuning. The encoder and generative model are trained together.

[0110] The training method specifically includes:

[0111] Generative pre-training: The goal of this stage is to train a diffusion model to generate 3D shapes from the latent space; use a large number of high-quality 3D shapes to train the diffusion model and use rendered images to regulate the diffusion process; the rendered images are generated by a pre-trained encoder E pretrain (·) encoding, where the pre-trained encoder maps the image condition to the CLIP latent space with a fixed pre-trained CLIP encoder;

[0112] Both 4-level UNet and 5-level UNet are trained using denoising loss:

[0113] Among them, N(0,1) represents Gaussian distribution, Represents the conditions applied to the generation process. In this stage, it is a picture constraint. The picture can also be replaced by text and / or point cloud. The pre-trained encoder can also be Uni3D or BLIP or other pre-trained encoders.

[0114] 3D Sketch Mapping: Map the 3D sketch input to the CLIP latent space, keeping the pre-trained diffusion model in generative pre-training unchanged, and then fine-tune the weights of the sketch encoder;

[0115] Joint Fine-tuning: Simultaneously fine-tuning the 3D sketch encoder E sketch (·) and diffusion model to ensure significantly improved spatial semantic alignment. This step also fine-tunes the 3D sketch encoder E sketch (·) and diffusion models to ensure significantly improved spatial semantic alignment; this approach has demonstrated its benefits in maximizing the utilization of pre-trained knowledge and is crucial for achieving better quality;

[0116] Furthermore, because there is only limited training data of 3D sketch and 3D model pairs, the network has difficulty generalizing across a wide range of sketch styles and geometric variations when mapping abstract 3D sketches to a latent space to condition the generative model; experiments have shown that the network occasionally fails to accurately reconstruct complex local details; in these areas, the network has difficulty in performing precise parameterization due to insufficient smoothness in the implicit function; even small errors can lead to sign errors, resulting in inaccurate surface reconstructions; in course learning, careful selection and sequencing of samples from simple to complex are crucial for effective and gradual skill development; the choice is based on the sample difficulty score;

[0117] In the second and / or third steps, additional curriculum learning is introduced into the training process. During the training process, the difficulty of each sample in the current training state is measured. Then, based on the difficulty of the sample, the number of sample pools participating in the training and the speed of adding new samples are dynamically adjusted, and difficult samples are gradually added in the order from simple to difficult samples.

[0118] The training process is as follows: points with incorrect estimates are considered hard samples, points with correct estimates are considered easy samples, and points between 0 and the ground truth are considered semi-hard samples, using the following difficulty score:

[0119] Among them, y is the SDF value of the real sample, is the predicted SDF value, 0≤λ≤1 controls the importance of hard and semi-hard samples; if v≥0, sgn(v)=1, otherwise, sgn(v)=-1; where v represents the input SDF value;

[0120] Adaptive course: First, use the trained 3D model to generate the network to obtain the initial difficulty score, and sort the pre-sorted dataset X in ascending order according to the current difficulty score s; then, use the rhythm function p(·) to form a sample pool X′. Specifically, the sorted dataset is divided according to the sample pool size of each small batch obtained by the rhythm function to form a new sample pool; a small batch B′=[B′1,...,B′ M ] to train the target network; the rhythm function p(·) is a monotonically increasing function that determines the speed of learning from simple to complex samples; finally, at the end of the forward propagation, the difficulty score s is updated and the new sample pool X′ is calculated; where B′ represents the set of all small batches; B′ M Indicates a small batch;

[0121] Difficulty scores also apply to different training durations, the (k+1)th th The difficulty score of a position can be expressed as: s k+1 =(1-α)s k +αscur

[0122] in, m represents the mth mini-batch; inv controls the frequency of difficulty score updates, α controls the speed of difficulty score updates; Scur represents the difficulty score during training.

[0123] Furthermore, in order to manage the speed of sample learning, a monotonically increasing rhythm function p(·) is set to limit the size of the sample pool X′. The function is expressed as:

[0124] Where n is the total number of samples; p0 is the sample ratio of the initial step; q is the speed at which the sample ratio increases; r0 is the frequency at which the sample ratio increases, and i is the current step.

[0125] Example 5

[0126] This embodiment provides a method for editing a 3D model based on a virtual reality / augmented reality 3D sketch, including:

[0127] In the first step, a suitable viewing angle is selected. The user selects the area M to be edited on the existing model;

[0128] In the second step, the user uses an AR or VR device to draw the rough shape of the desired modified model in the area;

[0129] The third step is to use the three-dimensional diffusion model to perform editing based on the original model and the area M to be edited;

[0130] The fourth step is to encode the text description into a latent code, where the text description comes from user input or from multiple sets of prompts generated by a large language model (LLM) based on user prompts, where the prompts are used to guide the diffusion model to generate content that conforms to the text description.

[0131] Furthermore, the specific process of the third step is:

[0132] The region to be edited M guides the diffusion by specifying which areas of the model should be modified or edited; the diffusion model follows the guidance of the region to be edited M and selectively applies modifications to the corresponding areas. In each diffusion step, the three-dimensional information is modified according to the mask guidance; the modification can involve changing the color, texture or other properties of the selected area; the diffusion model ensures that these edits are applied gradually and are consistent with the surrounding model content; specifically, this process is expressed by the following equation:

[0133] Among them, z t represents the implicit code after adding noise, y tIt represents the result after adding noise continuously during DDIM sampling; M is the mask (the area to be edited);

[0134] This step is performed N times, N>=1, until the required diffusion steps are completed; with each iteration, the 3D model is further improved and the editing effect becomes more obvious; the diffusion process ensures that the modifications are smoothly integrated with the rest of the 3D model.

[0135] Furthermore, in the third step, all the layers of the diffusion model are traversed and the cross attention and self-attention maps M are operated. attn ;

[0136] The new mask m is obtained by downsampling the mask M (the area to be edited) according to the resolution of the cross-attention layer or the self-attention layer; for the cross-attention map, the input and output column indices J are determined according to the token pointing to the content to be edited and the content not to be edited. in and J out ; For the self-attention map, the column index J is determined according to whether the three-dimensional primitive belongs to in and J out , using J in and J out Get the new attention map:

[0137] Using the new attention map, we can get the constrained noise z in each denoising stage cond , the noise without mask constraint is recorded as z uncond , calculate z t-1 ←z cond +w g *(z cond -z uncond )

[0138] Among them, w g represents the weighted value of unconditional noise;

[0139] Furthermore, the first to fourth steps are repeated, with different Ri, Ti selected in the first step and different combinations of text prompts T selected in the fourth step, thereby obtaining diversified three-dimensional models under different editing conditions.

[0140] Although the preferred embodiments of the present application have been described, those skilled in the art may make additional changes and modifications to these embodiments once they have learned the basic creative concept. Therefore, the appended claims are intended to be interpreted as including the preferred embodiments and all changes and modifications that fall within the scope of the present application.

[0141] It should be noted that the technical features not described in detail in the present invention can be implemented by any existing technology.

Claims

1. A processing system for a three-dimensional model, characterized in that, It includes an execution device, a computing device and an input device. The execution device is a 3D printer and / or a milling machine and / or a engraving machine. The input device is an AR (augmented reality) glasses or a VR (virtual reality) helmet. The AR / VR device is electrically connected to the computing device. The user draws three-dimensional sketch content in the air with a hand or a handle. The computing device converts the three-dimensional sketch content into a three-dimensional model, and then converts the three-dimensional model into a readable file for the execution device, and inputs the file into the execution device to obtain a three-dimensional physical object. Among them, the three-dimensional sketch is stored in the form of a three-dimensional point cloud. When the user draws a trajectory, corresponding point clouds are generated in the drawn trajectory and highlighted in the three-dimensional virtual space. The displacement and deletion operations of the user on the three-dimensional sketch are the overall displacement and deletion operations of the points in the corresponding area of the point cloud of the three-dimensional sketch.

2. A method for obtaining a physical model by three-dimensional printing based on a three-dimensional processing system, characterized in that, It includes: The first step: The user uses a hand or a handle in the input device to draw three-dimensional sketch content in the air. Among them, the three-dimensional sketch is stored in the form of a three-dimensional point cloud. When the user draws a trajectory, corresponding point clouds are generated in the drawn trajectory and highlighted in the three-dimensional virtual space. The displacement and deletion operations of the user on the three-dimensional sketch are the overall displacement and deletion operations of the points in the corresponding area of the point cloud of the three-dimensional sketch. The second step: The computing device converts the three-dimensional sketch content into a three-dimensional model. The third step: Convert the three-dimensional model into a readable file for the execution device, and the conversion is implemented by a slicing algorithm. The fourth step: The 3D printer executes the readable file for printing operation to obtain a physical model.

3. A method for obtaining a physical model by three-dimensional printing based on a three-dimensional processing system according to claim 2, characterized in that, The specific process of the first step is: 1.1 The user specifies the drawing pose. 1.2 The user adds a text description of the drawn content. 1.3 The user selects an existing model and depicts it on the existing model in the drawn cube frame. 1.4 The user selects an existing model or a picture, and the existing model or picture can be referenced and displayed in another floating window.

4. A method for obtaining a physical model by three-dimensional printing based on a three-dimensional processing system according to claim 2, characterized in that, In the second step, the conversion uses an AI deep learning model. The AI deep learning model is an encoder-decoder structure model pre-trained on a sketch-three-dimensional model pair or a diffusion model pre-trained on a large amount of three-dimensional model data based on constraints (conditions) on the sketch. Or, the conversion process is an optimized face filling algorithm that fills the faces between the sketch strokes and optimizes the shape.

5. A method for obtaining a physical model by three-dimensional printing based on a three-dimensional processing system according to claim 2, characterized in that, The method further includes: preprocessing the three-dimensional sketch drawn by the user. The preprocessing steps include: 2.1 Point sampling. The three-dimensional sketch drawn by the user is represented by a point cloud, and N points are evenly sampled on the trajectory of the sketch as input. 2.2 Position alignment. Move the center point of the drawn sketch to the center of the surrounding box (Bounding Box) with the size of M*N*P. 2.3 Size alignment. Calculate the spatial size of the drawn three-dimensional sketch and scale it to a fixed size of M*N*P. 2.4 Pose estimation. Use a pose estimation algorithm to estimate the pose of the sketch, represented by the pitch angle and the rotation angle. 2.5 Pose alignment. Use the pose estimated in 2.4 to rotate the sketch to a specific predefined pitch angle and rotation angle.

6. The method for obtaining a physical model by three-dimensional printing based on a three-dimensional processing system according to claim 5, characterized in that, The specific pose estimation algorithm is as follows: Use a multi-layer perceptron to encode the sketch, and the multi-layer perceptron outputs two pose data via an activation function; the multi-layer perceptron is trained using sketches with a large number of different poses and real pose results to prompt the neural network to output a more approximate real pose; Alternatively, design a self-supervised learning framework, use SE(3) (Special Euclidean Group) equivariance to decouple the object shape and pose, including an invariant shape reconstruction module and an equivariant pose estimation module, design a self-supervised loss function, and achieve self-supervised learning through the consistency loss between the input point cloud and the transformed canonical shape reconstruction; Among them: the invariant shape reconstruction module is used to learn to reconstruct the object shape aligned with the category-level reference frame, ensuring that the reconstructed shape remains invariant under SE(3) transformation. Use an element-wise multi-layer perceptron (MLP) to process the features, and then use max-pooling to obtain global SE(3) invariant features; finally, convert these features into point cloud form through an MLP-based point generator; The equivariant pose estimation module is used to estimate the object pose to make it equivariant to SE(3) transformation; apply an element-wise MLP to each feature independently, generate equivariant global features through max-pooling; then, use another MLP to process each global feature to predict the residual quaternion and residual translation to achieve pose estimation; Among them, use Equivariant Point Network (EPN) to extract features from the point cloud to ensure that the features remain equivariant under SE(3) transformation.

7. A method for obtaining a physical model by three-dimensional printing based on a three-dimensional processing system according to claim 2, characterized in that, The method further includes the user editing the obtained 3D model, including size adjustment, position adjustment, or adjustment of local geometric information; or using a sketch to edit and adjust the obtained 3D model.

8. A generation system from virtual reality / augmented reality three-dimensional sketches to three-dimensional models, characterized in that, Including: A sketch encoding module, including a Transformer-based point cloud encoder for encoding the 3D sketch S into a latent code, and the latent code is a feature vector in a high-dimensional space; Diffusion model, which generates Gaussian noise by gradually introducing Gaussian noise to a given sample Z following a predetermined variance distribution Then use a 3D Unet denoising based on time conditions to output discrete 3D voxels, and then obtain a 3D model through an octree-based convolutional neural network for the 3D voxels; The 3D Unet includes a 5-level UNet and a 4-level UNet, and a low-resolution 3D occupancy voxel is generated by the 5-level UNet Initial approximation of the 3D shape; constructing a high-resolution sparse voxel V ∈ R through a 4-level UNet N*N*N*4 ; The user inputs a 3D sketch, encodes it into a latent code through the sketch encoding module, and then injects it into the diffusion model through multi-head cross-attention to guide the diffusion model to output 3D model content that conforms to the 3D sketch guidance.

9. A training method for a diffusion model, characterized in that, Including: First, train a generative model using a large amount of existing data, and then fine-tune the encoder of the generative model with a small amount of desired data. When performing fine-tuning, the entire generative model remains unchanged.

10. A training method for a diffusion model according to claim 9, characterized in that, The encoder and the generative model are trained together.

11. The training method of a diffusion model according to claim 9, characterized in that, The specific training method includes: Generative Pretraining: The goal of this stage is to train a diffusion model to generate 3D shapes from the latent space; a large number of high-quality 3D shapes are used to train the diffusion model, and rendered images are used to condition the diffusion process; the rendered images are encoded by a pre-trained encoder E pretrain (·), and the pre-trained encoder maps the image conditions to the CLIP latent space with a fixed pre-trained CLIP encoder; Both the 4-level UNet and the 5-level UNet are trained using denoising loss: where N(0, 1) represents a Gaussian distribution, Denote the condition applied to the generation process. In this stage, it is a constraint for the picture. 3D sketch mapping: Map the input of the 3D sketch to the CLIP latent space, keep the pre-trained diffusion model in the generative pre-training unchanged, and then fine-tune and optimize the weights of the sketch encoder; Joint fine-tuning: Fine-tuning the 3D sketch encoder E sketch (·) and the diffusion model simultaneously to ensure a significant improvement in spatial semantic alignment.

12. The training method of a diffusion model according to claim 11, characterized in that, In the second step and / or the third step, additional curriculum learning is introduced to measure the difficulty of each sample in the current training state during the training process, and then the number of samples in the sample pool participating in the training and the speed of adding new samples are dynamically adjusted based on the difficulty of the samples, and difficult samples are gradually added in the order of samples from easy to difficult.

13. A training method for a diffusion model according to claim 12, characterized in that, The training process of the additional curriculum learning is as follows: Points with incorrect estimates are considered hard samples, points with correct estimates are considered easy samples, and points between 0 and the ground truth are considered semi-hard samples, using the following difficulty scores: where y is the SDF value of the real sample, is the predicted SDF value, 0 ≤ λ ≤ 1 controls the importance of difficult and semi-difficult samples; if v ≥ 0, sgn(v) = 1, otherwise, sgn(v) = -1; where v represents the input sdf value; Adaptive Course: First, use the pre-trained 3D model generation network to obtain the initial difficulty score, and sort the dataset X before sorting in ascending order according to the current difficulty score s; Subsequently, use the rhythm function p(·) to form the sample pool X′, specifically, divide the sorted dataset according to the sample pool size of each small batch obtained by the rhythm function to form a new sample pool; Extract a small batch B′ = [B′1,..., B′ M to train the target network; The rhythm function p(·) is a monotonically increasing function that determines the learning speed from simple to complex samples; Finally, update the difficulty score s and calculate the new sample pool X′ at the end of the forward propagation; where B′ represents the set of all small batches; B′ M represents a small batch; The difficulty score is also applicable to different training durations. The difficulty score at the (k + 1) th position can be expressed as: s k+1 = (1 - α)s k + αs cuγ Among them, m represents the m-th mini-batch; inv controls the frequency of difficulty score update, and α controls the speed of difficulty score update; Scur represents the difficulty score during the training process.

14. A training method for a diffusion model according to claim 13, wherein To manage the speed of sample learning, a monotonically increasing rhythm function p(·) is set to limit the size of the sample pool X′, and this function is expressed as: Among them, n is the total number of samples; p0 is the sample proportion in the initial step; q is the speed of controlling the growth of the sample proportion; r0 is the frequency of controlling the growth of the sample proportion, and i is the current step.

15. A method for editing a 3D model based on a virtual reality / augmented reality 3D sketch, characterized in that, including: In the first step, select a suitable perspective, and the user selects the area M to be edited on the existing model; In the second step, the user uses an AR or VR device to draw the approximate shape of the expected modified model within this area; In the third step, based on the original model and the area M to be edited, a three-dimensional diffusion model is used for editing; In the fourth step, the text description is encoded into a latent code, and the text description comes from user input or multiple sets of prompts generated by a large language model (LLM) according to user prompts, where the prompts are used to guide the diffusion model to generate content that conforms to the text description.

16. A method for editing a 3D model based on a virtual reality / augmented reality three-dimensional sketch according to claim 15, characterized in that, The specific process of the third step is as follows: The region M to be edited guides the diffusion by specifying which regions of the model should be modified or edited; the diffusion model follows the guidance of the region M to be edited and selectively applies the modifications to the corresponding regions. In each diffusion step, the three-dimensional information is modified according to the mask guidance; the modifications can involve changing the color, texture, or other properties of the selected regions; the diffusion model ensures that these edits are gradually applied and are consistent with the surrounding model content; specifically, this process is represented by the following equation: Among them, z t represents the implicit encoding after adding noise, and y t represents the result after continuously adding noise during the DDIM sampling process; M is a mask (the area to be edited); This step is executed N times, N >= 1, until the required diffusion steps are completed; in each iteration, the three-dimensional model will be further improved, and the editing effect becomes more obvious; the diffusion process ensures that the modification smoothly integrates with the rest of the three-dimensional model.

17. A method for editing a 3D model based on a virtual reality / augmented reality three-dimensional sketch according to claim 15, characterized in that, In the third step, traverse all the layers of the diffusion model's diffusion and operate on the cross-attention and self-attention map M atta ; Downsample the mask M (the area to be edited) according to the resolution of the cross-attention layer or the self-attention layer to obtain a new mask m; for the cross-attention map, determine the input and output column metrics J according to the tokens pointing to the content to be edited and the content not to be edited ia and J oot ; for the self-attention map, determine the column metrics J according to whether the three-dimensional primitive belongs to ia and J oot , and use J in and J oot to obtain a new attention map: Using the new attention map, obtain the constrained noise z at each denoising stage cond , denote the unmasked constrained noise as z uncond , calculate z t-1 ←z cond +w g *(Z cond -z uncond ) where, w g represents the weighted value of unconditional noise; 18. A method for editing a 3D model based on a virtual reality / augmented reality 3D sketch according to claim 15, characterized in that, Repeat the first step to the fourth step, select different Ri, Ti in the first step, and select different combinations of text prompts T in the fourth step, so as to obtain diversified three-dimensional models under different editing conditions.

Citation Information

Patent Citations

  • Posture determination method, tray loading method and device, medium and electronic equipment

    CN109829947A

  • Image enhancement method and device, equipment, storage medium and program product

    CN117115023A

Cited By

  • End-to-end underwater three-dimensional reconstruction method and system based on underwater imaging model

    CN120635333A

  • Non-rigid three-dimensional editing method and system based on cross-modal attention guidance

    CN120655802A

  • Grouping point cloud driven arm operation system

    CN120949936A

  • Performance index-based three-dimensional automobile model generation network, generation method and system

    CN121052146A

  • Indoor scene layout generation method and device and electronic equipment

    CN121147440A