Text-to-3D content generation methods aligned with human preferences

By training a generation framework using DreamAlign and D-3DPO algorithms, the problem of inconsistency between 3D content and human aesthetic preferences in existing technologies is solved, and high-quality 3D content that conforms to human aesthetics is generated.

CN119417981BActive Publication Date: 2026-01-06SHANGHAI JIAOTONG UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411437977.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-10-15
Publication Date
2026-01-06
Estimated Expiration
2044-10-15

AI Technical Summary

Technical Problem

Existing text-to-3D content generation models struggle to maintain consistency with human preferences across multiple perspectives and attributes, resulting in generated 3D content that does not match human aesthetics in terms of style, shadows, geometry, and appearance.

Method used

The generation framework DreamAlign, based on a multi-view diffusion model, is adopted. It combines the Direct 3D Preference Optimization (D-3DPO) algorithm and preference contrast feedback training. The training is conducted using the HP3D dataset to ensure that the generated 3D content conforms to human aesthetic preferences. An implicit neural network is used for representation.

Benefits of technology

The generated 3D content is highly consistent with human preferences in multiple perspectives and attributes, enhancing the visual appeal and detail of the content, and improving user satisfaction and acceptance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119417981B_ABST
    Figure CN119417981B_ABST
Patent Text Reader

Abstract

A text-to-3D content generation method aligned with human preferences is proposed. In the offline stage, a directly 3D preference optimization (D-3DPO) algorithm is trained using a constructed generative framework (DreamAlign) based on a multi-view diffusion model on a text-to-3D dataset (HP3D) containing expert preference annotations. In the online stage, the trained generative framework is used to generate 3D content. This invention can generate 3D content highly consistent with the input text throughout the entire text-to-3D generation process, improving user satisfaction and acceptance of the 3D content. It better addresses the problem of mismatch with human aesthetic preferences in existing technologies, thus possessing higher practical value in real-world applications.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to a technique in the field of image processing, specifically a method for generating text-to-3D content that aligns with human preferences. Background Technology

[0002] 3D asset creation is widely used in numerous fields such as games, animation, simulation environments, chatbots, and art. Despite the high demand, creating high-quality 3D content requires both artistic creativity and professional 3D modeling skills. Recent research has leveraged successful generative models and massive 3D datasets to efficiently generate 3D content. Optimized 2D augmentation methods have attracted significant attention due to their ability to create 3D assets based on textual cues.

[0003] While existing models can generate highly complex, viewpoint-consistent 3D content, it is difficult for this generated content to align with human preferences. This consistency includes not only multi-viewpoint consistency and 3D text consistency, but also the consistency of attributes with human preferences, such as style, shadows, geometry, and appearance. Summary of the Invention

[0004] To address the aforementioned shortcomings of existing technologies, this invention proposes a text-to-3D content generation method aligned with human preferences. Throughout the entire text-to-3D generation process, it not only generates 3D content highly consistent with the input text but also ensures that the generated 3D content conforms to human aesthetic preferences in multiple key attributes. This ensures that the final generated 3D content not only meets technical requirements but also better aligns with human visual and aesthetic habits, improving user satisfaction and acceptance of 3D content. It also better solves the problem of mismatch between existing technologies and human aesthetic preferences, thus possessing higher practical value in real-world applications.

[0005] This invention is achieved through the following technical solution:

[0006] This invention relates to a text-to-3D content generation method aligned with human preferences. In the offline stage, a direct 3D preference optimization (D-3DPO) algorithm is used to train a constructed generative framework (DreamAlign) based on a multi-view diffusion model using a preference comparison feedback method on a text-to-3D dataset (HP3D) containing expert preference annotations. In the online stage, based on a fractional distillation method, the trained DreamAlign generates 3D content represented using an implicit neural network.

[0007] The text-to-3D dataset is a text-3D object dataset annotated with human preferences, which alleviates the limitation of insufficient data in human preference learning. Each sample in this dataset includes a text description and multiple sets of multi-view images, and each multi-view image is labeled with the degree of human preference.

[0008] The multi-view diffusion model, based on the input text prompts and view coordinates, generates multiple orthogonal view images. Its framework includes: a text encoder, a U-Net module, and an image autoencoder. Specifically, the text encoder encodes the input text information into a feature matrix, the U-Net module denoises the latent space noise features based on the text feature matrix and view information, and the image autoencoder decodes the latent space noise features into multi-view images.

[0009] The aforementioned generative framework based on a multi-view diffusion model generates multiple orthogonal view images that conform to human aesthetics based on text. It adds a Lora adapter to the U-Net module of the multi-view diffusion model to learn human aesthetic preferences.

[0010] The fractional distillation method described converts text into 3D content and uses a pre-trained image generation model to guide the training of another model.

[0011] The implicit neural network records information about 3D objects through the parameters of the neural network and generates a 2D image of the 3D object from a specified viewpoint to represent 3D content.

[0012] The Direct 3D Preference Optimization (D-3DPO) algorithm specifically includes:

[0013] Step 1: Sample a data set from HP3D, including a cue text D, a camera viewpoint c, and two sets of multi-view images {I}. w ,I l}, where: I w Than I l It is more in line with human aesthetics.

[0014] Step 2: Add Gaussian noise ∈ to the two sets of multi-view images to create an image I that conforms to human aesthetics. w Inputting images into DreamAlign will result in images that do not conform to human aesthetics. l Input it into the original diffusion model.

[0015] Step 3, using L D-3DPO The model was fine-tuned to ensure that DreamAlign aligns with data distributions that conform to human aesthetics and avoids those that do not. Specifically:

[0016] in: The parameter is The noise predicted by DreamAlign, ∈ ref This represents the noise predicted by the original diffusion model.

[0017] The aforementioned preference comparison feedback training specifically includes:

[0018] Step i: Sample the 3D content represented by the implicit neural network with parameter θ to obtain a multi-view image x0, add noise ∈ and input it into DreamAlign and the original multi-view diffusion model.

[0019] Step ii, based on the idea of ​​contrastive learning, aims to obtain multi-view images that are close to the images reconstructed by DreamAlign and far from the images reconstructed by the original diffusion model. This is achieved using L... PCL The training will be conducted as follows:

[0020]

[0021] Technical effect

[0022] This invention addresses the technical problem in existing text-to-3D content generation methods where inconsistencies arise between generated 3D content and human preferences in multiple attributes such as text, shadows, and shape. This is due to the complexity and noise of datasets and the algorithm's neglect of human preference information. Compared to existing technologies, this invention introduces the Direct 3D Preference Optimization (D-3DPO) algorithm to fine-tune DreamAlign on the HP3D dataset. This ensures that multi-view generated images not only perform well in a single viewpoint but also maintain high consistency and aesthetic appeal across viewpoints, resulting in more coherent and realistic 3D content. Furthermore, the introduced preference contrast feedback training mechanism further enhances the visual appeal and detail of the 3D content by optimizing the distribution of generated content. Attached Figure Description

[0023] Figure 1 This is a flowchart of the present invention;

[0024] Figure 2 This is a schematic diagram illustrating the creation of the HP3D dataset for an example.

[0025] Figure 3 This is a schematic diagram of DreamAlign-2D as an example.

[0026] Figure 4 This is a schematic diagram of DreamAlign-3D as an example.

[0027] Figure 5 This is a schematic diagram illustrating the effect of the present invention. Detailed Implementation

[0028] like Figure 1 As shown, this embodiment relates to a text-to-3D content generation method aligned with human preferences, including:

[0029] Step 1, as follows Figure 2 As shown, the HP3D dataset is constructed, specifically including:

[0030] 1.1 Raw Data Filtering: To ensure the diversity of the selected 3D objects, the K-center algorithm and CLIP feature were used to select 10,000 text-3D object pairs from the Cap3D dataset. After manual screening, 376 high-quality objects that highly matched the text descriptions were obtained.

[0031] 1.2 Rendering Process: The selected 3D objects are subjected to intensive rendering, generating 32 multi-view images with a resolution of 512x512 for each object. During the camera's 360-degree orbital motion, the azimuth angle starts from the frontal view (90°), and the elevation angle is randomly selected from 0 to 30 degrees, while the camera's external parameters corresponding to each view are saved.

[0032] 1.3 Generating Comparative Samples: After generating text descriptions with the same semantics but different attributes using an advanced multimodal model (LLava), a multi-view diffusion model is used to generate 3D content with clearly distinguishable preferences based on the text descriptions. Specifically, after intensive rendering of a single 3D object, multi-view images and their corresponding text descriptions are obtained. LLava is then used to generate 2 to 3 text descriptions with different attributes in appearance, geometry, and style but the same semantics. Based on these text descriptions with the same semantics but different attributes, a multi-view diffusion model is finally used to generate multi-view images with obvious preference differences to facilitate subsequent expert evaluation.

[0033] 1.4, Preference Ranking: Based on the consistency between text and 3D objects, multi-view Figure 1 3D content is scored based on consistency and aesthetics, ranging from 0 to 7. Multi-view images are then ranked according to these scores to obtain the HP3D dataset with added preferences.

[0034] Step 2: Construct the generative framework for the multi-view diffusion model (DreamAlign): In the U-Net module of the diffusion model, a LoRA adapter is introduced. During preference optimization, only the parameters of the LoRA adapter are trained, while other parameters are frozen. With relatively low training cost, DreamAlign can generate multi-view images consistent with human preferences.

[0035] Step 3, as follows Figure 3 As shown, based on the HP3D dataset obtained in step 1, direct 3D preference optimization (D-

[0036] The 3DPO algorithm trains the DreamAlign algorithm obtained in step 2, specifically including:

[0037] 3.1 Randomly select an image from the HP3D dataset that includes a text description D, a camera viewpoint c, and two corresponding sets of multi-view images I with preference ranking. w ,I l The sample {I w ,I l ,D,c}.

[0038] 3.2 Add random Gaussian noise ε to the two sets of multi-view images.

[0039] 3.3 Input the noisy samples, camera viewpoint, and text prompts into DreamAlign and the multi-view diffusion model respectively, and the predicted noise is as follows: and ∈ ref Then, the loss function is calculated, and the network parameters are optimized with the minimum loss function as the optimization objective. Specifically:

[0040] Where β is a hyperparameter used for regularization. ω(λ) represents the signal-to-noise ratio. t ) is a predefined weighting function.

[0041] 3.4 Execute step 3.3 iteratively on two V100 GPUs, with a total batch size of 256, a local batch size of 1, a gradient accumulation count of 128, and a learning rate of 1e-5. For the D-3DPO algorithm, β is set to 5000.

[0042] Step 4, as follows Figure 4 As shown, based on DreamAlign obtained in step 3, preference contrastive feedback learning is used to generate 3D content from text that conforms to human aesthetics, specifically including:

[0043] 4.1 Randomly sample orthogonal camera views from the 3D content represented by the implicit neural network, then obtain orthogonal multi-view images and add random Gaussian noise ∈.

[0044] 4.2 Input the noisy image from step 4.2 into DreamAlign and the multi-view diffusion model used as a reference, respectively, to perform noise prediction and generate new images.

[0045] 4.3 Based on the DreamAlign algorithm obtained in step 4.2, which can generate images that conform to human aesthetics, a fractional distillation method is used to generate text-to-3D content. To further improve the aesthetics of the generated results, a preference contrast loss function is introduced. The parameters of the implicit neural network are optimized with the minimum loss function as the optimization objective. Specifically:

[0046] Where: x = g(θ) represents the image x sampled from the 3D implicit neural network representation g with parameter θ.

[0047] Step 5: In the online phase, text-to-3D content is generated using the methods described in Step 4, resulting in a similar... Figure 5 The output results were then analyzed. CLIP and ImageReward were used as quantitative metrics to evaluate the textual consistency and aesthetic appeal of the 3D content. CLIP was used to extract text and image features and calculate cosine similarity to evaluate the textual consistency and multi-view consistency of the 3D content, yielding scores of 35.85 and 78.29, respectively. The scalar results from ImageReward were used to evaluate the aesthetic appeal of the 3D content, with a score of 0.29. This indicates that the 3D content generated by this method better meets user preferences in terms of visual appeal and aesthetics.

[0048] The above-described specific implementations can be partially adjusted by those skilled in the art in different ways without departing from the principles and purpose of the present invention. The scope of protection of the present invention is defined by the claims and is not limited to the above-described specific implementations. All implementation schemes within the scope of the claims are bound by the present invention.

Claims

1. A text-to-3D content generation method aligned with human preferences, characterized by, In the offline stage, the generated framework DreamAlign based on the multi-view diffusion model is directly optimized by the D-3DPO algorithm based on the text-to-3D dataset HP3D containing expert preference annotations for preference comparison feedback training, and in the online stage, the method based on score distillation is used to generate 3D content represented by an implicit neural network using the trained DreamAlign; The multi-view diffusion model generates a diffusion model of multiple orthogonal view images according to the input text prompt and view coordinates, and the framework includes a text encoder, a U-Net module, and an image autoencoder, wherein: the text encoder encodes the input text information into a feature matrix, the U-Net module denoises the input hidden space noise features according to the text feature matrix and view information, and the image autoencoder decodes the hidden space noise features into multiple view images. In the preference optimization process, only the parameters of the LoRA adapter are trained, and other parameters are frozen. The D-3DPO algorithm specifically includes: Step 1, sample a sample from HP3D, including a prompt text , camera view , two sets of multi-view images , wherein: more in line with human aesthetics than ​ Step 2, add Gaussian noise to two groups of multi-view pictures pictures that match human aesthetics input DreamAlign, pictures that do not match human aesthetics input multi-view diffusion model as reference; Step 3, using Fine-tuning the generation framework DreamAlign based on the multi-view diffusion model, so that DreamAlign is consistent with the data distribution conforming to human aesthetics, and far away from the data distribution not conforming to human aesthetics, specifically: Wherein: Indicates the noise predicted by the DreamAlign with parameters Indicates the noise predicted by the original diffusion model;​ The preference comparison feedback training specifically includes: Step i, to parameters 3D content sampling of implicit neural network representations to multi-view images , adding noise into the DreamAlign and original multi-view diffusion model; Step ii, based on the idea of contrastive learning, the multi-view pictures sampled are expected to be close to the DreamAlign reconstructed images and far away from the original diffusion model reconstructed images, and the following loss function is used is trained, specifically: .

2. The text-to-3D content generation method aligned with human preferences of claim 1, wherein, The text-to-3D dataset is a human preference annotated text-3D object dataset, which alleviates the limitation of insufficient data in human preference learning, and each sample in the dataset includes a text description and multiple sets of multi-view pictures, each of which is annotated with human preference degree.

3. The text-to-3D content generation method aligned with human preferences of claim 1, wherein, The generation framework based on the multi-view diffusion model generates a diffusion model of multiple orthogonal view images that meet human aesthetic preferences based on text, and adds a Lora adapter to the U-Net module of the multi-view diffusion model to learn human aesthetic preferences.

4. The text-to-3D content generation method aligned with human preferences of claim 1, wherein, The score distillation method converts text into 3D content and uses a pre-trained image generation model to guide the training of another model.

5. The text-to-3D content generation method aligned with human preferences of claim 1, wherein, The implicit neural network records the information of the 3D object through the parameters of the neural network and generates a 2D image of the 3D object from a specified view to represent the 3D content.

Citation Information

Patent Citations

  • Text-image generation method, system and device and storage medium

    CN117095083A

  • Text-driven end-to-end 3D face rapid generation and editing method

    CN117853638A