Method, device and electronic equipment for generating three-dimensional scene images based on large models

By obtaining the text description and two-dimensional position information of the object, generating a reference image and performing instance segmentation and three-dimensional reconstruction, the problem of insufficient accuracy and authenticity of three-dimensional scene image generation in the existing technology is solved, and high-quality three-dimensional scene image generation is achieved.

CN119784940BActive Publication Date: 2025-09-16BEIJING BAIDU NETCOM SCI & TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411834354.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-12-12
Publication Date
2025-09-16
Estimated Expiration
2044-12-12

AI Technical Summary

Technical Problem

Existing technologies have difficulty in efficiently generating high-quality three-dimensional scene images based on text description information, especially in terms of object position and appearance accuracy.

Method used

By obtaining the text description information and two-dimensional position information of the objects in the target scene, a reference image is generated, and high-quality three-dimensional scene images are generated through instance segmentation and three-dimensional reconstruction combined with depth estimation and visual angle adjustment.

Benefits of technology

It improves the accuracy and authenticity of 3D scene images, ensures the reasonable layout of objects in space and their appearance conforms to the description, and supports personalized editing and dynamic object insertion.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119784940B_ABST
    Figure CN119784940B_ABST
Patent Text Reader

Abstract

This application discloses a method, device, and electronic device for generating a three-dimensional scene image based on a large model, relating to the fields of computer technology, particularly large models, deep learning, and computer vision. The specific implementation scheme comprises: obtaining textual description information and two-dimensional position information of at least one object in a target scene; generating a reference image based on the textual description information and two-dimensional position information; performing instance segmentation on the reference image to obtain an instance image of the object; performing three-dimensional reconstruction of the object based on the instance image to obtain a first three-dimensional image of the object; and generating a three-dimensional scene image of the target scene based on the first three-dimensional image.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of computer technology, in particular to artificial intelligence fields such as large models, deep learning, and computer vision, and specifically to a method, device, and electronic device for generating three-dimensional scene images based on large models. Background Art

[0002] With the development of artificial intelligence technology, many intelligent technologies have also entered the public's field of vision, such as generating three-dimensional images through text input by users. Summary of the Invention

[0003] The present application provides a method, device and electronic device for generating a three-dimensional scene image based on a large model.

[0004] According to one aspect of the present application, a method for generating a three-dimensional scene image based on a large model is provided, comprising:

[0005] Acquire text description information of at least one object in a target scene and two-dimensional position information of the object;

[0006] generating a reference image according to the text description information and the two-dimensional position information;

[0007] Performing instance segmentation on the reference image to obtain an instance image of the object;

[0008] Performing three-dimensional reconstruction on the object based on the instance image to obtain a first three-dimensional image of the object;

[0009] A three-dimensional scene image of the target scene is generated according to the first three-dimensional image.

[0010] According to another aspect of the present application, a device for generating a three-dimensional scene image based on a large model is provided, comprising:

[0011] an acquisition module, configured to acquire text description information of at least one object in a target scene and two-dimensional position information of the object;

[0012] A first generating module, configured to generate a reference image according to the text description information and the two-dimensional position information;

[0013] An instance segmentation module, configured to perform instance segmentation on the reference image to obtain an instance image of the object;

[0014] a three-dimensional reconstruction module, configured to perform three-dimensional reconstruction of the object based on the instance image to obtain a first three-dimensional image of the object;

[0015] The second generating module is configured to generate a three-dimensional scene image of the target scene according to the first three-dimensional image.

[0016] According to another aspect of the present application, an electronic device is provided, including:

[0017] at least one processor; and

[0018] a memory communicatively connected to the at least one processor; wherein,

[0019] The memory stores instructions that can be executed by the at least one processor. The instructions are executed by the at least one processor to enable the at least one processor to perform the method described in the above embodiment.

[0020] According to another aspect of the present application, a non-transitory computer-readable storage medium storing computer instructions is provided, wherein the computer instructions are used to enable the computer to execute the method according to the above embodiment.

[0021] According to another aspect of the present application, a computer program product is provided, including a computer program, which implements the steps of the method described in the above embodiment when executed by a processor.

[0022] It should be understood that the content described in this section is not intended to identify the key or important features of the embodiments of the present application, nor is it intended to limit the scope of the present application. Other features of the present application will become easily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS

[0023] The accompanying drawings are provided to facilitate a better understanding of the present invention and do not constitute a limitation of the present application.

[0024] Figure 1 A schematic flow chart of a method for generating a three-dimensional scene image based on a large model provided in one embodiment of the present application;

[0025] Figure 2 A schematic flow chart of a method for generating a three-dimensional scene image based on a large model provided in another embodiment of the present application;

[0026] Figure 3 A schematic flow chart of a method for generating a three-dimensional scene image based on a large model provided in another embodiment of the present application;

[0027] Figure 4 A schematic structural diagram of a large-model-based three-dimensional scene image generation device provided in one embodiment of the present application;

[0028] Figure 5 It is a block diagram of an electronic device used to implement the large model-based three-dimensional scene image generation method of an embodiment of the present application. DETAILED DESCRIPTION

[0029] The following description of exemplary embodiments of the present application is made in conjunction with the accompanying drawings, including various details of the embodiments of the present application to facilitate understanding. These details should be considered as merely exemplary. Therefore, those skilled in the art will recognize that various changes and modifications may be made to the embodiments described herein without departing from the scope and spirit of the present application. Similarly, for the sake of clarity and conciseness, descriptions of well-known functions and structures are omitted in the following description.

[0030] The following describes the large model-based three-dimensional scene image generation method, device, electronic device and storage medium of the embodiments of the present application with reference to the accompanying drawings.

[0031] Figure 1 A flowchart of a method for generating a three-dimensional scene image based on a large model provided in one embodiment of the present application.

[0032] The large model-based three-dimensional scene image generation method of the embodiment of the present application can be executed by the large model-based three-dimensional scene image generation device of the embodiment of the present application, and the device can be configured in an electronic device.

[0033] Among them, the electronic device can be any device with computing capabilities, such as a personal computer, mobile terminal, server, etc. The mobile terminal can be, for example, a vehicle-mounted device, a mobile phone, a tablet computer, a personal digital assistant, a wearable device, and other hardware devices with various operating systems, touch screens and / or display screens.

[0034] like Figure 1 As shown, the large model-based three-dimensional scene image generation method includes:

[0035] Step 101: Acquire text description information of at least one object in a target scene and the two-dimensional position information of the object.

[0036] In this application, the user can input text description information of at least one object in the target scene and two-dimensional position information of each object.

[0037] The text description information can be used to describe features such as the color and appearance of objects in the target scene. The text description information can include semantic information such as the appearance and attributes of the object. The text description information can be used to control the appearance and geometric structure of the object. For example, the text description information of object A is "a dog wearing a blue tie."

[0038] The two-dimensional position information of the object can be used to guide the generation of the object at the two-dimensional position information. For example, the two-dimensional position information of the object can be the position information of the bounding box corresponding to the object, and can include the coordinates of the four vertices of the bounding box.

[0039] It should be noted that if there are multiple objects in the target scene, the two-dimensional position information of each object is position information in the same coordinate system.

[0040] Step 102: Generate a reference image based on the text description information and the two-dimensional position information.

[0041] For example, prompt information for guiding the generation of a two-dimensional scene image can be generated based on the text description information and the two-dimensional position information, and the prompt information can be input into a large model to generate a two-dimensional reference image using the large model. The large model can be a text-based image model.

[0042] Step 103: perform instance segmentation on the reference image to obtain an instance image of the object.

[0043] The instance image may be an object region segmented from a reference image. It is understood that the instance image is a two-dimensional image.

[0044] As a possible implementation method, a pre-trained segmentation model can be used to perform instance segmentation on the reference image to obtain an instance image of each object.

[0045] As a possible implementation method, a pre-trained detection model can be used to detect the reference image to determine the detection frame of each object in the reference image, and then extract the instance image of the object from the reference image according to the detection frame.

[0046] Step 104 : Perform three-dimensional reconstruction on the object based on the instance image to obtain a first three-dimensional image of the object.

[0047] In the present application, depth estimation can be used to obtain a depth map of a reference image, and based on the depth map of the reference image, a depth map of an instance image can be obtained, and then based on the depth map of the instance image, a first three-dimensional image of the object can be obtained.

[0048] It should be noted that, if there are multiple objects in the target scene, the first three-dimensional image of each object is a three-dimensional image in the same coordinate system.

[0049] Step 105: Generate a three-dimensional scene image of the target scene based on the first three-dimensional image.

[0050] Exemplarily, if there is only one object in the target scene, the first three-dimensional image may be used as the three-dimensional scene image.

[0051] For example, if there are multiple objects in the target scene, the first 3D images of the multiple objects may be combined according to the 3D position information of each pixel in the first 3D image to generate a 3D scene image of the target scene.

[0052] For example, there are two objects in the target scene, object A and object B. The text description information of object A is "a puppy wearing a blue tie", and the text description information of object B is "a white dove". Based on the text description information of object A and object B, and the two-dimensional position information of object A and the two-dimensional position information of object B, the above method can be used to generate a three-dimensional scene image containing a puppy wearing a blue tie and a white dove.

[0053] In an embodiment of the present application, a reference image is generated based on the text description information and two-dimensional position information of the objects in the target scene, and the reference image is instance-segmented to obtain instance images of each object. The object is then three-dimensionally reconstructed based on the instance image to obtain a first three-dimensional image of the object. Then, based on the first three-dimensional images of each object in the target scene, a three-dimensional scene image of the target scene is generated. Thus, a three-dimensional scene image is generated based on the text description information and two-dimensional position information of the object, achieving a deep fusion of space and semantics. Based on the two-dimensional position information, the objects in the three-dimensional scene can be reasonably arranged, improving the accuracy and authenticity of the three-dimensional scene image, thereby improving the quality of the three-dimensional scene image. In addition, the appearance of objects can be adjusted or new objects can be added through text to expand the application scenarios.

[0054] Figure 2 A flowchart of a method for generating a three-dimensional scene image based on a large model provided in another embodiment of the present application.

[0055] like Figure 2 As shown, the method for generating a three-dimensional scene image based on a large model includes:

[0056] Step 201: Acquire text description information of at least one object in a target scene and the two-dimensional position information of the object.

[0057] In the present application, step 201 can be implemented in any of the embodiments of the present application, so it will not be described in detail here.

[0058] Step 202: Generate a reference image based on the text description information and the two-dimensional position information.

[0059] As a possible implementation method, the two-dimensional layout information of the target scene can be determined based on the text description information and the two-dimensional position information, and the reference image can be generated using the second largest model based on the two-dimensional layout information and the text description information.

[0060] The two-dimensional layout information may include two-dimensional position information and text description information of all objects in the target scene.

[0061] Exemplarily, the second largest model may be used to convert text description information into a corresponding two-dimensional image. For example, the second largest model may be a text-image model.

[0062] For example, the text description information Y_B of all objects in the target scene = {y_1, y_2, ..., y_N}, where y_1 represents the text description information of object 1, y_2 represents the text description information of object 2, and y_N represents the text description information of object N. The two-dimensional layout information B = {(b_1, y_1), (b_2, y_2), ..., (b_N, y_N)}, where b_1 represents the two-dimensional position information of the bounding box of object 1. The text description information Y_B of all objects and the two-dimensional layout information B can be input into the second largest model, and the reference image can be generated using the second largest model.

[0063] Therefore, by utilizing text description information and two-dimensional layout information and using a large model to generate a reference image, the accuracy of the reference image can be improved.

[0064] Step 203: perform instance segmentation on the reference image to obtain an instance image of the object.

[0065] In the present application, step 203 can be implemented in any of the embodiments of the present application, so it will not be described in detail here.

[0066] Step 204 : Perform three-dimensional reconstruction on the object based on the instance image to obtain a first three-dimensional image of the object.

[0067] In the present application, a mask image and a depth map of a reference image of an object can be obtained, and the three-dimensional position information of the object can be determined based on the two-dimensional position information of the object, the mask image and the depth map. Based on the text description information and the instance image, a second three-dimensional image of the object can be generated using a first large model. Based on the three-dimensional position information, the three-dimensional position of the object in the second three-dimensional image can be updated to obtain a third three-dimensional image of the object, and the angle of the object in the third three-dimensional image can be adjusted to obtain the first three-dimensional image.

[0068] Thus, based on the mask image of the object, the depth map of the reference image and the two-dimensional position information of the object, the object is three-dimensionally reconstructed, thereby improving the accuracy of the first three-dimensional image.

[0069] For example, the reference image may be segmented using a segmentation model, which may output an instance image of the object and a mask image. The mask image may be a binary mask image, in which the pixel values ​​of the object may be 1 and the pixel values ​​of the background may be 0.

[0070] Exemplarily, a depth estimation model may be used to extract depth information from a reference image to obtain a depth map of the reference image.

[0071] For example, the depth information of the object may be determined based on the mask image and the depth map, and the three-dimensional position information of the object may be determined based on the two-dimensional position information and the depth information of the object.

[0072] The depth information of the object may include the depth of the object in a three-dimensional space, and the depth information of the object may be used to locate the front and back positions of the object in a three-dimensional scene.

[0073] Illustratively, the center position of the object in the reference image may be determined based on the two-dimensional position information, and the three-dimensional position information of the object may be determined based on the center position and depth information of the object.

[0074] For example, the depth information of an object can be determined by the following formula (1):

[0075] z_i=mean(M_i*Depth(I_ref))(1)

[0076] Among them, z_i represents the depth of the object in three-dimensional space; Mi_i represents the mask image; I_ref represents the reference image; Depth(I_ref) represents the depth map of the reference image, and the value of each pixel in the depth map corresponds to the depth information of the corresponding position in the target scene; mean() represents the averaging operation.

[0077] For example, the two-dimensional position information of an object is the two-dimensional position information of the bounding box where the object is located (x_1, y_1, x_2, y_2), then the center position of the object in the reference image can be (x_i, y_i), where x_i = (x_1 + x_2) / 2, y_i = (y_1 + y_2) / 2, then the three-dimensional position information of the object is (x_i, y_i, z_i).

[0078] Therefore, by determining the depth information of the object from the depth map of the reference image based on the mask image of the object, the accuracy of the object's depth information can be improved, and thus the three-dimensional position information of the object can be determined based on the depth information of the object, which can improve the accuracy of the object's position in the three-dimensional scene.

[0079] Since the angle of an object in three-dimensional space may affect the quality of a three-dimensional scene image, the angle of the object in the three-dimensional image can be adjusted based on this.

[0080] Exemplarily, a second view image of the object in the first three-dimensional image at a different perspective can be obtained, feature extraction can be performed on the second view image to obtain first image features, feature extraction can be performed on the instance image to obtain second image features, and based on the similarity between the first image features and the second image features, a three-dimensional rotation angle can be determined. Then, based on the three-dimensional rotation angle, the object in the third three-dimensional image can be rotated to obtain the first three-dimensional image.

[0081] The rotation angle can be used to adjust the direction of the object so that it conforms to the visual angle of the reference image. For example, the rotation angle can include pitch angle, azimuth angle, etc.

[0082] For example, the first three-dimensional image may be captured from multiple angles by a virtual camera to obtain second view images at different viewing angles.

[0083] Exemplarily, the first image feature may include high-level semantic information of the object, low-level geometric information, etc. For example, the high-level semantic information may include color distribution, object shape, etc., and the low-level geometric information may include edge features.

[0084] Exemplarily, the rotation angle may be determined according to the viewing angle corresponding to the second view image when the similarity is maximum.

[0085] For example, the rendering angle can be determined using the following formula (2):

[0086] r_i=argmax(e,a)(cos(f_Ii,F))(2)

[0087] Among them, r_i represents the three-dimensional rotation angle of the object; argmax(e,a) indicates that the three-dimensional rotation angle pair is selected by calculating the similarity. argmax(e,a) can be used to indicate under which (e,a) combination the value of cos(f_Ii,F) reaches the maximum. Here (e,a) can represent different rotation angle pairs, such as the pitch angle e and the azimuth angle a; f_Ii represents the second image feature; F represents the set of first image features.

[0088] Thus, a 3D rotation angle is determined based on the feature similarity between the view images of the object in the 3D scene from different perspectives and the instance image. The object's orientation is rotated according to the 3D rotation angle to align with the visual angle of the reference image, improving the accuracy and authenticity of the 3D image of the object. Furthermore, determining the 3D rotation angle based on the features of the view images from multiple perspectives ensures a more comprehensive description of the 3D object.

[0089] Step 205 : In response to the mismatch between the first three-dimensional image and the text description information, adjusting the visual attributes of the object in the first three-dimensional image to obtain a second three-dimensional image.

[0090] In this application, information such as the color and shape of an object in a first 3D image can be determined and compared with the text description to determine whether the first 3D image matches the text description. If the first 3D image does not match the text description, the visual attributes of the object in the first 3D image can be adjusted to obtain a second 3D image.

[0091] Among them, visual attributes may include attributes such as size, geometric shape, and texture of the object.

[0092] In order to reduce unreasonable overlap between objects, as a possible implementation method, the initial size of the objects in the first three-dimensional image can be determined, the sparsity of the three-dimensional point cloud can be determined based on the three-dimensional point cloud distribution of the objects in the first three-dimensional image, and the collision perception loss can be determined based on the three-dimensional point cloud sparsity. Then, based on the collision perception loss, the initial size of each object can be adjusted to obtain a second three-dimensional image.

[0093] Among them, collision-aware loss can be used to reduce unreasonable overlap between objects in three-dimensional scenes.

[0094] The initial size may refer to the relative size of the object in the three-dimensional scene.

[0095] For example, the scale information of the bounding box of the object in the instance image may be determined, and the initial size may be determined based on the two-dimensional position information and the scale information.

[0096] For example, the two-dimensional position information is the position information of the specified bounding box, which can be understood as the position information of the bounding box of the object in the reference image. The ratio of the width of the specified bounding box to the width of the bounding box of the object in the instance image can be determined as the initial size of the object.

[0097] For example, the following formula (3) can be used to determine the relative size of an object in a three-dimensional scene:

[0098] s_i=W_bi / W_bi_prime(3)

[0099] Among them, s_i represents the relative size, which is used to describe the size ratio of the object; W_bi represents the width of the bounding box of the object in the reference image; W_bi_prime represents the width of the bounding box of the object in the instance image.

[0100] For example, the relative size of the pigeon should be smaller than the relative size of the table to maintain visual plausibility in the three-dimensional scene.

[0101] Therefore, based on the two-dimensional position information of the object and the scale information of the bounding box of the object in the instance image, the proportion of the two-dimensional object can be mapped to three dimensions, which facilitates the subsequent optimization of the size of the object in the three-dimensional scene.

[0102] For example, the collision perception loss can be calculated using the following formula (4):

[0103] C_overlap=k_c*sum_k(ReLU(R_diff-||q_2k-q_1center||))(4)

[0104] Among them, C_overlap represents the collision perception loss; k_c represents the weight factor of the collision perception loss; sum_k represents the summation operation on the index k of the Gaussian point, which is used to calculate the distribution difference of all Gaussian points. The Gaussian point represents the three-dimensional point cloud distribution of the object, and each point contains spatial coordinate information; R_diff represents the average sparsity of the Gaussian point, which is used to measure the geometric distribution of the instance; q_2k represents the coordinates of the Gaussian point, and q_1center can represent the coordinates of the center point of the target object. For example, the target object can be any object, or the first object, etc.

[0105] Therefore, based on the sparsity of the three-dimensional point cloud of the object in the first three-dimensional image, the collision perception loss is determined, and the size of the object is adjusted using the collision perception loss, which can reduce unreasonable overlap between objects, improve the rationality of object interaction, and thus improve the rationality of the three-dimensional scene image.

[0106] Optionally, the three-dimensional position information of the objects and the target distance between each object can be determined, and the spatial perception loss can be determined based on the three-dimensional position information and target distance between each object. Then, the initial size of each object can be adjusted based on the collision perception loss and the spatial perception loss to obtain a second three-dimensional image.

[0107] Among them, spatial perception loss can be used to constrain the spatial position relationship between objects, such as distance, context position, etc.

[0108] The target distance between any two objects can be set according to actual needs and is not limited thereto.

[0109] For example, the spatial perception loss can be calculated using the following formula (5):

[0110]

[0111] Among them, C_Spatial represents the spatial perception loss, t i and t j Represents the three-dimensional position information of object i and object j respectively, d ij represents the target distance between object i and object j.

[0112] Exemplarily, the collision perception loss and the space perception loss may be weighted, and based on the weighted loss, the initial size may be adjusted to obtain the second three-dimensional image.

[0113] Therefore, combined with spatial perception loss, the size of objects in three-dimensional space can be adjusted to reduce unreasonable overlap between objects and constrain the spatial position relationship between objects, thereby improving the rationality of three-dimensional scene layout.

[0114] Optionally, an initial scene image of the target scene can be generated based on the first three-dimensional image of all objects in the target scene, and feature comparison can be performed between the initial scene image and the reference image to determine the feature alignment loss. Then, the initial size can be adjusted based on the collision perception loss and the feature alignment loss to obtain a second three-dimensional image.

[0115] Among them, feature alignment loss can be used to measure the similarity between the initial scene image and the reference image.

[0116] For example, first view images of the initial scene image at different perspectives can be obtained, and the feature alignment loss can be determined based on the similarity between the first view image and the reference image. Thus, determining the feature alignment loss based on the similarity between the view images at multiple perspectives and the reference image can improve the accuracy of the feature alignment loss.

[0117] Exemplarily, the collision perception loss and the feature alignment loss may be weighted, and based on the weighted loss, the initial size of each object may be adjusted to obtain the second three-dimensional image.

[0118] Therefore, combined with feature alignment loss, the size of the object in the three-dimensional space is adjusted, which can reduce the unreasonable overlap between objects while ensuring that the objects in the three-dimensional scene conform to the description.

[0119] In order to improve the guarantee that the geometric details and textures of objects in the three-dimensional scene image conform to the description, as a possible implementation method, for each object, a diffusion model can be used to obtain predicted noise based on the first three-dimensional image and text description information of the object, and the diffusion loss is determined based on the predicted noise and the real noise added to the first three-dimensional image. Then, based on the diffusion loss, the initial geometric parameters and initial texture parameters of the object in the first three-dimensional image are adjusted until the diffusion loss is less than a preset threshold, and a second three-dimensional image of the object is obtained.

[0120] Among them, the diffusion loss can reflect the gap between the model prediction noise and the actual noise, and is used to guide the generation of three-dimensional images of objects that are closer to the real semantics.

[0121] The initial geometric parameters may include vertex positions of the object in the first three-dimensional image, and the initial texture parameters may include color, texture mapping, etc. of the object in the first three-dimensional image.

[0122] For example, the first 3D image, text description information, and time step can be input into the diffusion model for noise prediction to obtain predicted noise, wherein the first 3D image is the 3D image input into the diffusion model at the current time step.

[0123] Since geometric parameters and texture parameters may need to be adjusted multiple times, the diffusion model can be used to obtain predicted noise based on the current three-dimensional image of the object to be optimized, the text description information of the object and the current time step. The diffusion loss is determined based on the predicted noise and the actual noise, and the current geometric parameters and texture parameters are adjusted based on the diffusion loss.

[0124] For example, the diffusion loss can be calculated using the following formula (6):

[0125] C_Diffusion=||ε_hat_phi(x_t;y_i,t)-ε||^2(6)

[0126] Where C_Diffusion represents the diffusion loss; ε_hat_phi(x_t; y_i, t) represents the predicted noise of the diffusion model, x_t is the 3D image to be adjusted of object i at the current time step t, y_i is the text description information of object i; ε represents the real noise added to the 3D image to be adjusted; ||^2 represents the square norm.

[0127] Therefore, combined with text description information, the diffusion model is used for instance-level optimization. Object details are optimized through gradual denoising, and fine geometric structures and high-resolution textures are generated from rough shapes. Each step optimizes and adjusts the appearance of the object to make it more consistent with the semantic description provided by the text description information, thereby improving the accuracy and quality of the object's three-dimensional image.

[0128] Optionally, the normal vector of the object surface in the first three-dimensional image can be determined, and the normal vector smoothing loss can be determined based on the normal vector and the local average of the normal vector and adjacent normal vectors. Then, the initial geometric parameters and initial texture parameters can be adjusted based on the diffusion loss and the normal vector smoothing loss to obtain the second three-dimensional image.

[0129] Among them, the normal vector smoothness loss can be used to ensure the continuity of the normal vector distribution on the object surface and reduce sharp changes. The normal vector smoothness loss can encourage the normal vector distribution on the object surface to be smoother and avoid unnatural sharp mutations or bumps.

[0130] For example, the normal vector smoothing loss can be calculated using the following formula (7):

[0131] C_smooth=Σ_i||n_i-avg(n_i)||^2(7)

[0132] Among them, C_smooth represents the normal vector smoothing loss, n_i represents the normal vector of the i-th surface, avg(n_i) represents the local average of the i-th normal vector, that is, the average of adjacent normal vectors, Σ_i represents the sum of all normal vectors, and ||^2 represents the square norm, that is, the sum of the squares of the differences.

[0133] For example, the diffusion loss and the normal vector smoothing loss may be weighted, and based on the weighted losses, the geometric parameters and texture parameters of the object may be adjusted to obtain the second three-dimensional image.

[0134] Therefore, by combining the normal vector smoothing loss and adjusting the geometric parameters and texture parameters, the surface of the three-dimensional scene object can be made smoother, avoiding unnatural sharp mutations or bumps, thereby improving the quality of the three-dimensional image of the object.

[0135] Optionally, a first total variation regularization loss can be obtained based on the initial geometric parameters, and a second total variation regularization loss can be obtained based on the square of the difference between adjacent normal vectors on the surface of the object in the first three-dimensional image. Then, the initial geometric parameters and the initial texture parameters are adjusted based on the diffusion loss, the first total variation regularization loss and the second total variation regularization loss to obtain a second three-dimensional image.

[0136] Among them, the first total variation regularization loss can be used to reduce the gradient noise of geometric surfaces such as depth information, making the object surface smoother.

[0137] Among them, the second total variation regularization loss can also make the object surface smoother.

[0138] The initial geometric parameters may include vertex positions of the object in the first three-dimensional image. Exemplarily, the first total variation regularization loss may be determined based on differences between adjacent vertices.

[0139] For example, the first total variation regularization loss and the second total variation regularization loss may be added together, and then weighted with the diffusion loss. Based on the obtained loss sum, the initial geometric parameters and the initial texture parameters may be adjusted.

[0140] Therefore, combined with the total variation regularization loss, the geometric parameters and texture parameters of the object in the three-dimensional space are adjusted, which can reduce the unreasonable overlap between objects while ensuring that the objects in the three-dimensional scene conform to the description.

[0141] Step 206: Generate a three-dimensional scene image based on the second three-dimensional image.

[0142] Exemplarily, if there is only one object in the target scene, the second three-dimensional image may be used as the three-dimensional scene image.

[0143] For example, if there are multiple objects in the target scene, the second three-dimensional images of the multiple objects may be combined according to the three-dimensional position information of each pixel in the second three-dimensional image to generate a three-dimensional scene image of the target scene.

[0144] In an embodiment of the present application, if the first three-dimensional image does not match the text description information, the visual attributes of the object in the first three-dimensional image can be adjusted, and the visual attributes of the object in the three-dimensional image can be optimized to make the three-dimensional image more consistent with the description, thereby improving the quality of the three-dimensional image of the object, and further improving the quality of the three-dimensional scene image.

[0145] Figure 3 A flowchart of a method for generating a three-dimensional scene image based on a large model provided in another embodiment of the present application.

[0146] like Figure 3 As shown, the method for generating a three-dimensional scene image based on a large model includes:

[0147] Step 301: Acquire text description information of at least one object in a target scene and the two-dimensional position information of the object.

[0148] Step 302: Generate a reference image based on the text description information and the two-dimensional position information.

[0149] Step 303: perform instance segmentation on the reference image to obtain an instance image of the object.

[0150] Step 304: reconstruct the object in three dimensions based on the instance image to obtain a first three-dimensional image of the object.

[0151] Step 305: Generate a three-dimensional scene image of the target scene based on the first three-dimensional image.

[0152] In the present application, steps 301 to 305 can be implemented in any of the embodiments of the present application and will not be described in detail here.

[0153] Step 306 : In response to the mismatch between the first 3D image and the text description information, the initial size is adjusted according to the collision perception loss, the spatial position loss, and the feature alignment loss.

[0154] In this application, if the first three-dimensional image does not match the text description information, the collision perception loss, spatial perception loss and feature alignment loss can be weighted to obtain the total layout optimization loss. According to the total layout optimization loss, the initial size of each object is adjusted until the total layout optimization loss is less than the preset threshold to obtain the second three-dimensional image.

[0155] Among them, the total loss of layout optimization can be used to guide the adjustment of the three-dimensional position and relationship of objects to generate a more reasonable three-dimensional scene layout.

[0156] For example, the total loss of layout optimization can be calculated using the following formula (8):

[0157] C_layout=C_Spatial+k_f*C_feat+k_c*C_overlap(8)

[0158] Among them, C_layout represents the total loss of layout optimization, C_Spatial represents the spatial perception loss, C_feat represents the feature alignment loss, C_overlap represents the collision perception loss, k_f represents the weight factor of feature alignment loss, and k_c represents the weight factor of collision perception loss.

[0159] Therefore, according to the collision perception loss, combined with the spatial perception loss and feature alignment loss, the size of the object is adjusted, which can not only reduce the unreasonable overlap between objects, but also constrain the spatial position relationship between objects and ensure that the objects in the three-dimensional scene conform to the description, generating a more reasonable three-dimensional scene layout, thereby improving the authenticity of the three-dimensional image of the object, and further improving the quality of the three-dimensional scene image.

[0160] Step 307 : Adjust the initial geometric parameters and the initial texture parameters according to the diffusion loss, the normal vector smoothing loss, the first total variation regularization loss, and the second total variation regularization loss to obtain a second three-dimensional image.

[0161] Exemplarily, the diffusion loss, normal vector smoothing loss, first total variation regularization loss and second total variation regularization loss can be weighted to obtain instance-level optimization loss, and the instance-level optimization loss is used to adjust the initial geometric parameters and initial texture parameters until the instance-level optimization loss is less than a preset threshold to obtain a second three-dimensional image.

[0162] Among them, instance-level optimization loss can be used to constrain geometry and texture, adjusting the geometric shape and texture details of the object.

[0163] For example, the instance-level optimization loss can be calculated using the following formula (9):

[0164] C_Instance=C_Diffusion+k_s*C_smooth+k_TV*(C_TV_d+C_TV_n)(9)

[0165] Among them, C_Instance represents the instance-level optimization loss, C_Diffusion represents the diffusion loss, C_smooth represents the normal vector smoothing loss, C_TV_d represents the first total variational regularization loss, C_TV_n represents the second total variational regularization loss, and the k_s and k_TV weight factors control the contributions of the normal vector smoothing loss and the total variational regularization loss to the instance-level optimization loss, respectively.

[0166] Therefore, by optimizing the geometric shape and texture of the object based on the diffusion loss, combined with the normal vector smoothing loss and the total variation regularization loss, the appearance and structure of the three-dimensional object meet the visual and semantic requirements, thereby improving the accuracy and authenticity of the three-dimensional image of the object, thereby improving the quality of the three-dimensional image of the object, and further improving the quality of the three-dimensional scene image.

[0167] Step 308: Generate a 3D scene image based on the second 3D image.

[0168] In this application, step 308 can be implemented in any of the embodiments of this application, and will not be described in detail here.

[0169] In an embodiment of the present application, if the first three-dimensional image does not match the text description information, the size of the object in the three-dimensional scene can be adjusted based on collision perception loss, spatial perception loss and feature alignment loss, which can improve the rationality of the three-dimensional scene layout, and based on diffusion loss, combined with normal vector smoothing loss and total variation regularization loss, the geometric shape and details of the object can be optimized, so that the appearance and structure of the three-dimensional object can meet the visual and semantic requirements, thereby improving the accuracy and authenticity of the three-dimensional scene image, and further improving the quality of the three-dimensional scene image.

[0170] To facilitate understanding of the large model-based three-dimensional scene image generation method of the present application embodiment, the following embodiments are described below:

[0171] The solution for generating a 3D scene image based on a large model in this application is as follows:

[0172] 1. System input

[0173] Inputs include:

[0174] The text hint Y_B = {y_1, y_2, ..., y_N} describes the semantic information of the objects in the scene, where y_N represents the text description of object N.

[0175] The two-dimensional layout information B = {(b_1, y1), (b_2, y_2), ..., (b_N, y_N)}, consists of the bounding box b_i and the corresponding object description y_i.

[0176] The bounding box b_i=(x1, y1, x2, y2) represents the two-dimensional location of each object.

[0177] 2. Preliminary 3D generation stage

[0178] (1) Reference image generation

[0179] The reference image is generated by combining the text hint Y_B and layout B through MIGC (Multi-Instance Generation Controller):

[0180] I_ref=MIGC(Y_B,B)(10)

[0181] Among them, text prompts can control the appearance and semantics of objects, and layout ensures spatial consistency.

[0182] (2) 3D reconstruction of a single object

[0183] Use a segmentation model to separate instances from a reference image:

[0184] I_i,M_i=SAM(I_ref,b_i)(11)

[0185] Where SAM represents the segmentation model, I_i is the instance image of object i, and M_i is the mask image of object i.

[0186] Using the above formula (1), the depth estimation model is used to extract depth information and calculate the depth information of the object.

[0187] Generate a rough 3D instance using LGM (LGM Large Multi-View Gaussian Model):

[0188] A_i=LGM(I_i,y_i)(12)

[0189] (3) 3D layout initialization

[0190] According to the ratio of the 2D bounding box to 3D, the size s_i is calculated using the above formula (3).

[0191] The three-dimensional position information t_i=(x_i, y_i, z_i) is determined by combining the depth information and the two-dimensional position information of the object.

[0192] Using the above formula (2), the three-dimensional rotation angle is determined based on feature similarity, and the object is rotated.

[0193] 3. Separation optimization stage

[0194] (1) Collision perception layout optimization

[0195] The above formula (9) is used to calculate the total loss of layout optimization, and the total loss of layout optimization is used to adjust the s_i of each object.

[0196] (2) Instance-level optimization

[0197] The instantiation loss is calculated using the above formula (9), and the geometry and texture of the object are optimized using the instantiation loss.

[0198] In the initial 3D generation phase, the solution of this application utilizes an efficient layout control mechanism to generate a reference image and a 3D reconstruction method to generate a rough 3D scene. In the separation optimization phase, collision-aware optimization and instance-level optimization are used to optimize the scene's geometry and texture, making the appearance more consistent with the description and improving the accuracy and realism of the generated 3D scene image. Furthermore, object appearance can be adjusted or new objects can be added using text, supporting personalized editing and dynamic object insertion.

[0199] In order to implement the above embodiment, the embodiment of the present application also proposes a three-dimensional scene image generation device based on a large model. Figure 4 A schematic structural diagram of a large-model-based three-dimensional scene image generation device provided in one embodiment of the present application.

[0200] like Figure 4 As shown, the large model-based 3D scene image generation device 400 includes:

[0201] An acquisition module 410 is configured to acquire text description information of at least one object in a target scene and two-dimensional position information of the object;

[0202] A first generating module 420 is configured to generate a reference image based on the text description information and the two-dimensional position information;

[0203] An instance segmentation module 430 is configured to perform instance segmentation on the reference image to obtain an instance image of the object;

[0204] A three-dimensional reconstruction module 440 is configured to perform three-dimensional reconstruction of the object based on the instance image to obtain a first three-dimensional image of the object;

[0205] The second generating module 450 is configured to generate a three-dimensional scene image of the target scene according to the first three-dimensional image.

[0206] Optionally, the second generating module 450 is configured to: in response to the first three-dimensional image not matching the text description information, adjust the visual attributes of the object in the first three-dimensional image to obtain a second three-dimensional image;

[0207] The three-dimensional scene image is generated according to the second three-dimensional image.

[0208] Optionally, the second generating module 450 is configured to:

[0209] determining an initial size of the object in the first three-dimensional image;

[0210] determining a 3D point cloud sparsity based on a 3D point cloud distribution of the object in the first 3D image;

[0211] determining a collision perception loss based on the sparsity of the three-dimensional point cloud;

[0212] The initial size is adjusted according to the collision perception loss to obtain the second three-dimensional image.

[0213] Optionally, the second generating module 450 is configured to:

[0214] Determining three-dimensional position information of the objects and target distances between any two of the objects;

[0215] Determining a spatial perception loss based on the three-dimensional position information of the two objects and the target distance;

[0216] The initial size is adjusted according to the collision perception loss and the space perception loss to obtain the second three-dimensional image.

[0217] Optionally, the second generating module 450 is configured to:

[0218] generating an initial scene image of the target scene according to the first three-dimensional image;

[0219] performing feature comparison between the initial scene image and the reference image to determine a feature alignment loss;

[0220] The initial size is adjusted according to the collision perception loss and the feature alignment loss to obtain the second three-dimensional image.

[0221] Optionally, the second generating module 450 is configured to:

[0222] Acquire first view images of the initial scene image at different viewing angles;

[0223] The feature alignment loss is determined according to a similarity between the first view image and the reference image.

[0224] Optionally, the second generating module 450 is configured to:

[0225] determining scale information of a bounding box of the object in the instance image;

[0226] The initial size is determined according to the two-dimensional position information and the scale information.

[0227] Optionally, the second generating module 450 is configured to:

[0228] Obtaining predicted noise using a diffusion model according to the first three-dimensional image and the text description information;

[0229] determining a diffusion loss based on the predicted noise and actual noise added to the first three-dimensional image;

[0230] According to the diffusion loss, initial geometric parameters and initial texture parameters of the object in the first three-dimensional image are adjusted to obtain the second three-dimensional image.

[0231] Optionally, the second generating module 450 is configured to:

[0232] determining a normal vector of a surface of the object in the first three-dimensional image;

[0233] determining a normal vector smoothing loss based on the normal vector and a local average of the normal vector and adjacent normal vectors;

[0234] The initial geometric parameters and the initial texture parameters are adjusted according to the diffusion loss and the normal vector smoothing loss to obtain the second three-dimensional image.

[0235] Optionally, the second generating module 450 is configured to:

[0236] Obtaining a first total variation regularization loss according to the initial geometric parameters;

[0237] Obtaining a second total variation regularization loss based on squares of differences between adjacent normal vectors of the surface of the object in the first three-dimensional image;

[0238] The initial geometric parameters and the initial texture parameters are adjusted according to the diffusion loss, the first total variation regularization loss, and the second total variation regularization loss to obtain the second three-dimensional image.

[0239] Optionally, the 3D reconstruction module 440 is configured to:

[0240] Acquire a mask image of the object and a depth map of the reference image;

[0241] determining three-dimensional position information of the object based on the two-dimensional position information, the mask image, and the depth map;

[0242] Based on the text description information and the instance image, using the first large model, generating a second three-dimensional image of the object;

[0243] updating the three-dimensional position of the object in the second three-dimensional image according to the three-dimensional position information to obtain a third three-dimensional image of the object;

[0244] Angle adjustment is performed on the object in the third three-dimensional image to obtain the first three-dimensional image.

[0245] Optionally, the 3D reconstruction module 440 is configured to:

[0246] Determining depth information of the object according to the mask image and the depth map;

[0247] The three-dimensional position information is determined based on the two-dimensional position information and the depth information.

[0248] Optionally, the 3D reconstruction module 440 is configured to:

[0249] Acquire a second view image of the object in the first three-dimensional image at a different viewing angle;

[0250] performing feature extraction on the second view image to obtain first image features;

[0251] Performing feature extraction on the instance image to obtain a second image feature;

[0252] determining a three-dimensional rotation angle based on a similarity between the first image feature and the second image feature;

[0253] The object in the third three-dimensional image is rotated according to the three-dimensional rotation angle to obtain the first three-dimensional image.

[0254] Optionally, the first generating module 420 is configured to:

[0255] Determining two-dimensional layout information of the target scene according to the text description information and the two-dimensional position information; wherein the two-dimensional layout information includes the two-dimensional position information and the text description information;

[0256] The reference image is generated using a second large model according to the two-dimensional layout information and the text description information.

[0257] It should be noted that the explanation of the embodiment of the large model-based three-dimensional scene image generation method is also applicable to the large model-based three-dimensional scene image generation device of this embodiment, so it will not be repeated here.

[0258] In an embodiment of the present application, a reference image is generated based on the text description information and two-dimensional position information of the objects in the target scene, and the reference image is instance-segmented to obtain instance images of each object. The object is then three-dimensionally reconstructed based on the instance image to obtain a first three-dimensional image of the object. Then, based on the first three-dimensional images of each object in the target scene, a three-dimensional scene image of the target scene is generated. Thus, a three-dimensional scene image is generated based on the text description information and two-dimensional position information of the object, achieving a deep fusion of space and semantics. Based on the two-dimensional position information, the objects in the three-dimensional scene can be reasonably arranged, improving the accuracy and authenticity of the three-dimensional scene image, thereby improving the quality of the three-dimensional scene image. In addition, the appearance of objects can be adjusted or new objects can be added through text to expand the application scenarios.

[0259] According to an embodiment of the present application, the present application also provides an electronic device, a readable storage medium and a computer program product.

[0260] Figure 5 A schematic block diagram of an example electronic device 500 that can be used to implement embodiments of the present application is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital assistants, cellular phones, smartphones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the present application described and / or claimed herein.

[0261] like Figure 5 As shown, the device 500 includes a computing unit 501, which can perform various appropriate actions and processes according to a computer program stored in a ROM (Read-Only Memory) 502 or a computer program loaded from a storage unit 508 into a RAM (Random Access Memory) 503. Various programs and data required for the operation of the device 500 can also be stored in the RAM 503. The computing unit 501, the ROM 502, and the RAM 503 are connected to each other via a bus 504. An I / O (Input / Output) interface 505 is also connected to the bus 504.

[0262] Various components in device 500 are connected to I / O interface 505, including: an input unit 506, such as a keyboard, mouse, etc.; an output unit 507, such as various types of displays, speakers, etc.; a storage unit 508, such as a magnetic disk, optical disk, etc.; and a communication unit 509, such as a network card, modem, wireless communication transceiver, etc. The communication unit 509 allows device 500 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks.

[0263] The computing unit 501 can be any general-purpose and / or specialized processing component with processing and computing capabilities. Some examples of the computing unit 501 include, but are not limited to, a CPU (Central Processing Unit), a GPU (Graphic Processing Unit), various specialized AI (Artificial Intelligence) computing chips, various computing units that run machine learning model algorithms, a DSP (Digital Signal Processor), and any suitable processor, controller, microcontroller, etc. The computing unit 501 performs the various methods and processes described above, such as the large-scale model-based 3D scene image generation method. For example, in some embodiments, the large-scale model-based 3D scene image generation method can be implemented as a computer software program tangibly embodied in a machine-readable medium, such as the storage unit 508. In some embodiments, part or all of the computer program can be loaded and / or installed on the device 500 via the ROM 502 and / or the communication unit 509. When the computer program is loaded into the RAM 503 and executed by the computing unit 501, one or more steps of the large-scale model-based 3D scene image generation method described above can be performed. Alternatively, in other embodiments, the computing unit 501 may be configured to execute the large model-based three-dimensional scene image generation method in any other appropriate manner (for example, by means of firmware).

[0264] Various embodiments of the systems and techniques described herein can be implemented in digital electronic circuit systems, integrated circuit systems, FPGAs (Field Programmable Gate Arrays), ASICs (Application-Specific Integrated Circuits), ASSPs (Application-Specific Standard Products), SOCs (System on Chips), CPLDs (Complex Programmable Logic Devices), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include being implemented in one or more computer programs that are executable and / or interpreted on a programmable system that includes at least one programmable processor, which can be a special-purpose or general-purpose programmable processor that can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit data and instructions to the storage system, the at least one input device, and the at least one output device.

[0265] The program code for implementing the methods of the present application can be written in any combination of one or more programming languages. Such program code can be provided to a processor or controller of a general-purpose computer, a special-purpose computer, or other programmable data processing device, so that when the program code is executed by the processor or controller, the functions / operations specified in the flow charts and / or block diagrams are implemented. The program code can be executed entirely on the machine, partially on the machine, as a stand-alone software package, partially on the machine and partially on a remote machine, or entirely on a remote machine or server.

[0266] In the context of the present application, a machine-readable medium may be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, device, or apparatus. A machine-readable medium may be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium may include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, device, or apparatus, or any suitable combination of the foregoing. More specific examples of machine-readable storage media may include an electrical connection based on one or more wires, a portable computer disk, a hard disk, RAM, ROM, EPROM (Electrically Programmable Read-Only-Memory) or flash memory, optical fiber, CD-ROM (Compact Disc Read-Only Memory), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0267] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device (e.g., a CRT (Cathode-Ray Tube) or LCD (Liquid Crystal Display) monitor) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user can provide input to the computer. Other types of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, voice input, or tactile input).

[0268] The systems and techniques described herein can be implemented in a computing system that includes backend components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes frontend components (e.g., a user computer with a graphical user interface or web browser through which a user can interact with implementations of the systems and techniques described herein), or a computing system that includes any combination of such backend components, middleware components, or frontend components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include: LAN (Local Area Network), WAN (Wide Area Network), the Internet, and blockchain networks.

[0269] A computer system may include a client and a server. The client and server are generally remote from each other and typically interact via a communication network. The client-server relationship is established by computer programs running on the respective computers and establishing a client-server relationship with each other. The server may be a cloud server, also known as a cloud computing server or cloud host, a host product within the cloud computing service ecosystem that addresses the management difficulties and poor scalability of traditional physical hosts and VPS services. The server may also be a server in a distributed system or a server integrated with blockchain.

[0270] According to an embodiment of the present application, the present application also provides a computer program product, which, when an instruction processor in the computer program product executes, executes the large model-based three-dimensional scene image generation method proposed in the above embodiment of the present application.

[0271] It should be understood that the various forms of the processes shown above can be used to reorder, add, or delete steps. For example, the steps described in this application can be performed in parallel, sequentially, or in a different order, as long as the desired results of the technical solutions disclosed in this application can be achieved. This is not a limitation herein.

[0272] The above specific embodiments do not constitute a limitation on the scope of protection of this application. Those skilled in the art will appreciate that various modifications, combinations, sub-combinations, and substitutions may be made based on design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this application shall be included within the scope of protection of this application.

Claims

1. A method for generating a three-dimensional scene image based on a large model, comprising: Acquire text description information of at least one object in a target scene and two-dimensional position information of the object; generating a reference image according to the text description information and the two-dimensional position information; Performing instance segmentation on the reference image to obtain an instance image of the object; Performing three-dimensional reconstruction on the object based on the example image to obtain a first three-dimensional image of the object; obtaining a mask image of the object and a depth map of the reference image; determining three-dimensional position information of the object based on the two-dimensional position information, the mask image, and the depth map; generating a three-dimensional image of the object using a first large model based on the text description information and the example image; updating the three-dimensional position of the object in the three-dimensional image of the object generated by the first large model based on the three-dimensional position information to obtain a third three-dimensional image of the object; and adjusting the angle of the object in the third three-dimensional image to obtain the first three-dimensional image; Based on the first three-dimensional image, a three-dimensional scene image of the target scene is generated; wherein, in response to the first three-dimensional image not matching the text description information, the visual attributes of the object in the first three-dimensional image are adjusted to obtain a second three-dimensional image; and based on the second three-dimensional image, the three-dimensional scene image is generated.

2. The method according to claim 1, wherein The adjusting the visual attributes of the object in the first three-dimensional image to obtain a second three-dimensional image includes: determining an initial size of the object in the first three-dimensional image; determining a 3D point cloud sparsity based on a 3D point cloud distribution of the object in the first 3D image; determining a collision perception loss based on the sparsity of the three-dimensional point cloud; The initial size is adjusted according to the collision perception loss to obtain the second three-dimensional image.

3. The method according to claim 2, wherein: The adjusting the initial size according to the collision perception loss to obtain the second three-dimensional image includes: Determining three-dimensional position information of the objects and target distances between any two of the objects; Determining a spatial perception loss based on the three-dimensional position information of the two objects and the target distance; The initial size is adjusted according to the collision perception loss and the space perception loss to obtain the second three-dimensional image.

4. The method according to claim 2, wherein: The adjusting the initial size according to the collision perception loss to obtain the second three-dimensional image includes: generating an initial scene image of the target scene according to the first three-dimensional image; performing feature comparison between the initial scene image and the reference image to determine a feature alignment loss; The initial size is adjusted according to the collision perception loss and the feature alignment loss to obtain the second three-dimensional image.

5. The method according to claim 4, wherein: The determining the feature alignment loss by performing feature comparison on the initial scene image and the reference image includes: Acquire first view images of the initial scene image at different viewing angles; The feature alignment loss is determined according to a similarity between the first view image and the reference image.

6. The method of claim 2, wherein: Determining the initial size of the object in the first three-dimensional image includes: determining scale information of a bounding box of the object in the instance image; The initial size is determined according to the two-dimensional position information and the scale information.

7. The method of claim 1, wherein: The adjusting the visual attributes of the object in the first three-dimensional image to obtain a second three-dimensional image includes: Obtaining predicted noise using a diffusion model according to the first three-dimensional image and the text description information; determining a diffusion loss based on the predicted noise and actual noise added to the first three-dimensional image; According to the diffusion loss, initial geometric parameters and initial texture parameters of the object in the first three-dimensional image are adjusted to obtain the second three-dimensional image.

8. The method of claim 7, wherein: The adjusting, based on the diffusion loss, initial geometric parameters and initial texture parameters of the object in the first three-dimensional image to obtain the second three-dimensional image includes: determining a normal vector of a surface of the object in the first three-dimensional image; determining a normal vector smoothing loss based on the normal vector and a local average of the normal vector and adjacent normal vectors; The initial geometric parameters and the initial texture parameters are adjusted according to the diffusion loss and the normal vector smoothing loss to obtain the second three-dimensional image.

9. The method of claim 7, wherein: The adjusting, based on the diffusion loss, initial geometric parameters and initial texture parameters of the object in the first three-dimensional image to obtain the second three-dimensional image includes: Obtaining a first total variation regularization loss according to the initial geometric parameters; Obtaining a second total variation regularization loss based on squares of differences between adjacent normal vectors of the surface of the object in the first three-dimensional image; The initial geometric parameters and the initial texture parameters are adjusted according to the diffusion loss, the first total variation regularization loss, and the second total variation regularization loss to obtain the second three-dimensional image.

10. The method of claim 1, wherein: Determining the three-dimensional position information of the object according to the two-dimensional position information, the mask image, and the depth map includes: Determining depth information of the object according to the mask image and the depth map; The three-dimensional position information is determined based on the two-dimensional position information and the depth information.

11. The method of claim 1, wherein: The step of adjusting the angle of the object in the third three-dimensional image to obtain the first three-dimensional image includes: Acquire a second view image of the object in the first three-dimensional image at a different viewing angle; performing feature extraction on the second view image to obtain first image features; Performing feature extraction on the instance image to obtain a second image feature; determining a three-dimensional rotation angle based on a similarity between the first image feature and the second image feature; The object in the third three-dimensional image is rotated according to the three-dimensional rotation angle to obtain the first three-dimensional image.

12. The method according to any one of claims 1 to 9, wherein Generating a reference image according to the text description information and the two-dimensional position information includes: Determining two-dimensional layout information of the target scene according to the text description information and the two-dimensional position information; wherein the two-dimensional layout information includes the two-dimensional position information and the text description information; The reference image is generated using a second large model according to the two-dimensional layout information and the text description information.

13. A device for generating a three-dimensional scene image based on a large model, comprising: an acquisition module, configured to acquire text description information of at least one object in a target scene and two-dimensional position information of the object; A first generating module, configured to generate a reference image according to the text description information and the two-dimensional position information; An instance segmentation module, configured to perform instance segmentation on the reference image to obtain an instance image of the object; a 3D reconstruction module configured to perform 3D reconstruction of the object based on the example image to obtain a first 3D image of the object; wherein the module obtains a mask image of the object and a depth map of the reference image; determines 3D position information of the object based on the 2D position information, the mask image, and the depth map; generates a 3D image of the object using a first large model based on the text description information and the example image; updates the 3D position of the object in the 3D image of the object generated by the first large model based on the 3D position information to obtain a third 3D image of the object; and adjusts the angle of the object in the third 3D image to obtain the first 3D image; A second generating module is configured to generate a three-dimensional scene image of the target scene based on the first three-dimensional image; wherein, in response to a mismatch between the first three-dimensional image and the text description information, the visual attributes of the object in the first three-dimensional image are adjusted to obtain a second three-dimensional image; and the three-dimensional scene image is generated based on the second three-dimensional image.

14. The apparatus of claim 13, wherein: The second generating module is used to: determining an initial size of the object in the first three-dimensional image; determining a 3D point cloud sparsity based on a 3D point cloud distribution of the object in the first 3D image; determining a collision perception loss based on the sparsity of the three-dimensional point cloud; The initial size is adjusted according to the collision perception loss to obtain the second three-dimensional image.

15. The apparatus of claim 14, wherein: The second generating module is used to: Determining three-dimensional position information of the objects and target distances between any two of the objects; Determining a spatial perception loss based on the three-dimensional position information of the two objects and the target distance; The initial size is adjusted according to the collision perception loss and the space perception loss to obtain the second three-dimensional image.

16. The apparatus of claim 14, wherein: The second generating module is used to: generating an initial scene image of the target scene according to the first three-dimensional image; performing feature comparison between the initial scene image and the reference image to determine a feature alignment loss; The initial size is adjusted according to the collision perception loss and the feature alignment loss to obtain the second three-dimensional image.

17. The apparatus of claim 16, wherein: The second generating module is used to: Acquire first view images of the initial scene image at different viewing angles; The feature alignment loss is determined according to a similarity between the first view image and the reference image.

18. The apparatus of claim 14, wherein: The second generating module is used to: determining scale information of a bounding box of the object in the instance image; The initial size is determined according to the two-dimensional position information and the scale information.

19. The apparatus of claim 13, wherein: The second generating module is used to: Obtaining predicted noise using a diffusion model according to the first three-dimensional image and the text description information; determining a diffusion loss based on the predicted noise and actual noise added to the first three-dimensional image; According to the diffusion loss, initial geometric parameters and initial texture parameters of the object in the first three-dimensional image are adjusted to obtain the second three-dimensional image.

20. The apparatus of claim 19, wherein: The second generating module is used to: determining a normal vector of a surface of the object in the first three-dimensional image; determining a normal vector smoothing loss based on the normal vector and a local average of the normal vector and adjacent normal vectors; The initial geometric parameters and the initial texture parameters are adjusted according to the diffusion loss and the normal vector smoothing loss to obtain the second three-dimensional image.

21. The apparatus of claim 19, wherein: The second generating module is used to: Obtaining a first total variation regularization loss according to the initial geometric parameters; Obtaining a second total variation regularization loss based on squares of differences between adjacent normal vectors of the surface of the object in the first three-dimensional image; The initial geometric parameters and the initial texture parameters are adjusted according to the diffusion loss, the first total variation regularization loss, and the second total variation regularization loss to obtain the second three-dimensional image.

22. The apparatus of claim 13, wherein: The three-dimensional reconstruction module is used to: Determining depth information of the object according to the mask image and the depth map; The three-dimensional position information is determined based on the two-dimensional position information and the depth information.

23. The apparatus of claim 13, wherein: The three-dimensional reconstruction module is used to: Acquire a second view image of the object in the first three-dimensional image at a different viewing angle; performing feature extraction on the second view image to obtain first image features; Performing feature extraction on the instance image to obtain a second image feature; determining a three-dimensional rotation angle based on a similarity between the first image feature and the second image feature; The object in the third three-dimensional image is rotated according to the three-dimensional rotation angle to obtain the first three-dimensional image.

24. The device according to any one of claims 13 to 21, wherein The first generating module is configured to: Determining two-dimensional layout information of the target scene according to the text description information and the two-dimensional position information; wherein the two-dimensional layout information includes the two-dimensional position information and the text description information; The reference image is generated using a second large model according to the two-dimensional layout information and the text description information.

25. An electronic device comprising: at least one processor; as well as a memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the method according to any one of claims 1 to 12.

26. A non-transitory computer-readable storage medium storing computer instructions, wherein: The computer instructions are used to cause the computer to execute the method according to any one of claims 1-12.

27. A computer program product comprising a computer program, which, when executed by a processor, implements the steps of the method according to any one of claims 1 to 12.

Citation Information

Patent Citations

  • Three-dimensional character modeling method and device, storage medium and computer equipment

    CN118628673A

  • Flexible manufacturing system target component three-dimensional reconstruction method based on multi-modal large model

    CN119091062A