Three-dimensional scene generation method and device and electronic equipment
By generating a panoramic image and combining multi-perspective information and sparse point clouds, a stable three-dimensional scene model is generated using a diffusion model and three-dimensional reconstruction technology, which solves the problems of multi-perspective consistency and three-dimensional scene stability in existing technologies and achieves high-quality 3D scene generation.
Patent Information
- Application Number
- CN202410339067.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-03-22
- Publication Date
- 2025-09-23
AI Technical Summary
When generating high-quality 3D scenes, existing technologies have difficulty in achieving multi-perspective consistency and stability of three-dimensional scene models.
By acquiring the target text, a panoramic image is generated. Multi-view information and depth estimation are used to determine the sparse point cloud. The multi-view image and sparse point cloud are combined to generate a three-dimensional scene model. Diffusion models and three-dimensional reconstruction technologies such as NeRF and NeuS are used to optimize parameters to improve the generation effect.
It achieves the consistency of multiple perspectives and the stability of three-dimensional scene models, generates high-quality 3D scenes, and improves user experience.
Smart Images

Figure CN120689489A_ABST
Abstract
Description
Technical Field
[0001] The embodiments of the present disclosure relate to the field of computer technology, and more particularly to a method, device, and electronic device for generating a three-dimensional scene. Background Art
[0002] Artificial Intelligence Generated Content (AIGC) refers to content generated by artificial intelligence. In the context of 3D scene generation, AIGC can be used to automatically create realistic background environments. With the emergence of commercial mixed reality platforms and the rapid innovation of 3D graphics technology, high-quality 3D scene generation has become one of the most important problems in computer vision. Using AIGC to generate 3D scene backgrounds offers advantages such as speed, efficiency, customizability, creativity, and diversity. Summary of the Invention
[0003] This disclosure section is provided to briefly introduce concepts that will be described in detail in the detailed description section below. This disclosure section is not intended to identify key features or essential features of the claimed technical solution, nor is it intended to limit the scope of the claimed technical solution.
[0004] In a first aspect, an embodiment of the present disclosure provides a three-dimensional scene generation method, including: obtaining target text and generating a panoramic image of the target text description; obtaining multi-perspective information of preset multi-perspectives, and using the panoramic image to generate a multi-perspective picture under multiple perspectives; performing depth estimation on the panoramic image and determining a sparse point cloud corresponding to the panoramic image; and generating a three-dimensional scene model of the target text description based on the multi-perspective picture, multi-perspective information and sparse point cloud.
[0005] In the second aspect, an embodiment of the present disclosure provides a three-dimensional scene generation device, including: an acquisition unit, used to acquire target text and generate a panoramic image described by the target text; a first generation unit, used to acquire multi-perspective information of preset multi-perspectives, and use the panoramic image to generate a multi-perspective picture under multiple perspectives; a determination unit, used to perform depth estimation on the panoramic image and determine the sparse point cloud corresponding to the panoramic image; a second generation unit, used to generate a three-dimensional scene model described by the target text based on the multi-perspective picture, multi-perspective information and sparse point cloud.
[0006] In a third aspect, an embodiment of the present disclosure provides an electronic device comprising: one or more processors; a storage device for storing one or more programs, wherein when the one or more programs are executed by the one or more processors, the one or more processors implement the three-dimensional scene generation method as in the first aspect.
[0007] In a fourth aspect, an embodiment of the present disclosure provides a computer-readable medium having a computer program stored thereon, which, when executed by a processor, implements the steps of the three-dimensional scene generation method of the first aspect.
[0008] The 3D scene generation method, device, and electronic device provided by the embodiments of the present disclosure obtain target text to generate a panoramic image of the target text description; then, obtain multi-perspective information of preset multiple perspectives and, using the panoramic image, generate a multi-perspective image under the multiple perspectives; then, perform depth estimation on the panoramic image to determine the sparse point cloud corresponding to the panoramic image; finally, based on the multi-perspective image, the multi-perspective information, and the sparse point cloud, generate a 3D scene model of the target text description. In this way, a panoramic image of the text description can be generated first and then the corresponding 3D scene model can be generated, ensuring the consistency of multiple perspectives and the stability of the 3D scene model. BRIEF DESCRIPTION OF THE DRAWINGS
[0009] The above and other features, advantages, and aspects of the various embodiments of the present disclosure will become more apparent with reference to the following detailed description in conjunction with the accompanying drawings. Throughout the drawings, the same or similar reference numerals represent the same or similar elements. It should be understood that the drawings are schematic and that the originals and elements are not necessarily drawn to scale.
[0010] Figure 1 is a flow chart of an embodiment of a three-dimensional scene generation method according to the present disclosure;
[0011] Figure 2A and 2B is a schematic diagram of generating a panoramic image in the three-dimensional scene generation method according to the present disclosure;
[0012] Figure 3 is a schematic diagram of an application scenario of the three-dimensional scene generation method according to the present disclosure;
[0013] Figure 4 is a schematic diagram of an embodiment of generating a panoramic image by fine-tuning an original diffusion model in a three-dimensional scene generation method according to the present disclosure;
[0014] Figure 5 is a schematic diagram of another embodiment of generating a panoramic image by fine-tuning an original diffusion model in a three-dimensional scene generation method according to the present disclosure;
[0015] Figure 6 is a flowchart of another embodiment of a three-dimensional scene generation method according to the present disclosure;
[0016] Figure 7 is a flowchart of another embodiment of the three-dimensional scene generation method according to the present disclosure;
[0017] Figure 8is a structural diagram of an embodiment of a three-dimensional scene generating device according to the present disclosure;
[0018] Figure 9 is an exemplary system architecture diagram in which various embodiments of the present disclosure may be applied;
[0019] Figure 10 It is a structural diagram of a computer system suitable for implementing the electronic device of the embodiment of the present disclosure. DETAILED DESCRIPTION
[0020] The following describes embodiments of the present disclosure in more detail with reference to the accompanying drawings. Although certain embodiments of the present disclosure are shown in the accompanying drawings, it should be understood that the present disclosure can be implemented in various forms and should not be construed as limited to the embodiments described herein. Rather, these embodiments are provided to provide a more thorough and complete understanding of the present disclosure. It should be understood that the drawings and embodiments of the present disclosure are for illustrative purposes only and are not intended to limit the scope of protection of the present disclosure.
[0021] It should be understood that the various steps described in the method embodiments of the present disclosure may be performed in different orders and / or in parallel. In addition, the method embodiments may include additional steps and / or omit the steps shown. The scope of the present disclosure is not limited in this respect.
[0022] As used herein, the term "including" and its variations are open-ended, i.e., "including but not limited to." The term "based on" means "based, at least in part, on." The term "one embodiment" means "at least one embodiment," the term "another embodiment" means "at least one additional embodiment," and the term "some embodiments" means "at least some embodiments." Other terms are defined in the following description.
[0023] It should be noted that the concepts of "first" and "second" mentioned in this disclosure are only used to distinguish different devices, modules or units, and are not used to limit the order or interdependence of the functions performed by these devices, modules or units.
[0024] It should be noted that the modifications of "one" and "multiple" mentioned in the present disclosure are illustrative rather than restrictive, and those skilled in the art should understand that unless otherwise clearly indicated in the context, they should be understood as "one or more".
[0025] The names of the messages or information exchanged between multiple devices in the embodiments of the present disclosure are only used for illustrative purposes and are not used to limit the scope of these messages or information.
[0026] Please refer to Figure 1, shows a process 100 of an embodiment of a 3D scene generation method according to the present disclosure. The 3D scene generation method includes the following steps:
[0027] Step 101: Obtain target text and generate a panoramic image described by the target text.
[0028] In this embodiment, the execution subject of the three-dimensional scene generation method can obtain the target text. The above-mentioned target text is usually descriptive text, and the above-mentioned target text is usually text determined based on the user's input operation. As an example, the above-mentioned target text can be text manually entered by the user, or it can be text obtained by converting the user's input voice, or it can be determined by the user triggering a preset control corresponding to the text. For example, multiple preset text controls can be presented to the user, such as sunset, sea, snowflakes, etc. If the user triggers the "sunset" control, it is determined that the target text includes sunset.
[0029] Afterwards, the execution entity may generate a panoramic image of the target text description. Specifically, the target text may be input into a pre-trained image generation model to obtain a panoramic image of the target text description. The image generation model is used to characterize the correspondence between the text and the panoramic image described by the text. The image generation model may include, but is not limited to, a Generative Adversarial Network (GAN) and a Variational Autoencoder (VAE).
[0030] like Figure 2A and 2B As shown, Figure 2A and 2B This is a schematic diagram of generating a panoramic image in the three-dimensional scene generation method according to this embodiment. Figure 2A In the example, the user inputs the text “crowded alleys, cherry blossom trees, traditional lanterns”, as shown in 201, and a panorama as shown in 202 is generated. The user inputs the text “winding streets, antique shops, old-fashioned lampposts”, as shown in 203, and a panorama as shown in 204 is generated.
[0031] Step 102 : obtaining multi-view information of preset multi-views, and generating multi-view pictures under the multi-views by using the panoramic image.
[0032] In this embodiment, the execution entity may obtain multi-view information of a preset multi-view and generate a multi-view image from the multi-view using the panoramic image. Since one view corresponds to one camera pose, and multiple views can correspond to multiple camera poses, the multi-view information may also be understood as camera pose information. Therefore, the multi-view information may include camera intrinsic parameters and camera extrinsic parameters.
[0033] Here, multiple viewing angles may be preset, multi-viewing information of the preset multiple viewing angles may be obtained, and the panoramic image may be used to determine images at various viewing angles, thereby generating pictures at various viewing angles.
[0034] Step 103 : perform depth estimation on the panoramic image to determine a sparse point cloud corresponding to the panoramic image.
[0035] In this embodiment, the execution entity may perform depth estimation on the panoramic image to determine a sparse point cloud corresponding to the panoramic image.
[0036] Specifically, the panoramic depth D(x, y) can be estimated through the panoramic image I(x, y), and the camera intrinsic parameter K and camera extrinsic parameter R, t are used to project and obtain a three-dimensional sparse point cloud.
[0037] First, the pixel coordinates (x, y) can be converted to the coordinates in the camera coordinate system (X c ,Y c ,Z c ).
[0038] After that, the coordinates in the camera coordinate system (X c ,Y c ,Z c ) is converted to the coordinates in the world coordinate system (X w ,Y w , Z w ).
[0039] The sparse point cloud can then be scaled based on the depth value D(x,y). This way, the estimated depth of the panorama can be used to generate a 3D point corresponding to each pixel.
[0040] Here, depth estimation methods such as ZoeDepth (Zero-shot Transfer by Combining Relative and Metric Depth) or MVSNet (an end-to-end depth estimation framework based on deep learning) can be used to estimate the depth of the panorama.
[0041] Step 104 : Generate a three-dimensional scene model of the target text description based on the multi-view image, multi-view information, and sparse point cloud.
[0042] In this embodiment, the execution entity may generate a three-dimensional scene model described by the target text based on the multi-view image, the multi-view information, and the sparse point cloud.
[0043] Specifically, the above-mentioned execution entity can use three-dimensional reconstruction methods such as SFM (Structure From Motion) reconstruction, NeRF (Neural Radiance Field) reconstruction and NeuS (Neural Implicit Surfaces) / NeuS2 to generate a three-dimensional scene model described by the above-mentioned target text.
[0044] The method provided by the above-mentioned embodiment of the present disclosure obtains the target text and generates a panoramic image of the target text description; then obtains multi-perspective information of preset multiple perspectives and uses the panoramic image to generate a multi-perspective image under the above-mentioned multiple perspectives; then, depth estimation is performed on the panoramic image to determine the sparse point cloud corresponding to the panoramic image; finally, based on the multi-perspective image, the multi-perspective information, and the sparse point cloud, a three-dimensional scene model of the target text description is generated. In this way, a panoramic image of the text description is first generated and then the corresponding three-dimensional scene model is generated, ensuring the consistency of multiple perspectives and the stability of the three-dimensional scene model.
[0045] Continue to see Figure 3 , Figure 3 FIG. 1 is a schematic diagram of an application scenario of the three-dimensional scene generation method according to this embodiment. Figure 3 In an application scenario, a user enters the text "beach, blue sky, ocean, coconut trees, sunset," as shown in icon 301. A panoramic image 302 describing the text is then generated. Pre-set multi-perspective information 304 is then obtained, and a multi-perspective image is generated from the panoramic image 302, as shown in icon 303. Depth estimation is then performed on the panoramic image 302 to determine the sparse point cloud corresponding to the panoramic image 302, as shown in icon 305. Finally, a 3D scene model describing the text 301 is generated based on the multi-perspective image 303, the multi-perspective information 304, and the sparse point cloud 305, as shown in icon 306.
[0046] In some optional implementations, the execution entity may generate a panoramic image of the target text description in the following manner: using a pre-trained target diffusion model (StableDiffusionModel) to generate a panoramic image of the target text description, wherein the target diffusion model is used to characterize the correspondence between the text and the panoramic image. A diffusion model may also be referred to as a generative diffusion model. A diffusion model is a type of generative model, which is a type of model that can generate synthetic images. The generation of a diffusion model starts with random noise and is gradually refined through multiple steps until an output image appears. At each step, the model estimates how to get from the current input to a denoised version.
[0047] Diffusion models outperform networks like GANs and VAEs in generating new images. Specifically, they outperform GANs and VAEs in terms of memory capacity, image freedom, smooth transitions between images, and the variety of generated images. Diffusion models are effective and easy to implement, producing high-quality images. Therefore, combining diffusion models with 3D reconstruction techniques can better generate the 3D images or scenes required for AR / VR.
[0048] In some optional implementations, the target diffusion model is a model obtained by performing a target operation on an original diffusion model. The original diffusion model is typically used to represent the correspondence between text and a two-dimensional image. The target operation typically includes freezing parameters of the original diffusion model and inserting a learnable module into the original diffusion model. The learnable module can be used to convert the two-dimensional image into a panoramic image.
[0049] Generally speaking, neural networks have forward propagation and backpropagation. Freezing the parameters of a neural network means only forward propagation is performed on the parameters, without backpropagation, and the parameters are not optimized. Instead, the parameters of the inserted learnable modules are optimized and learned, thereby adjusting the network's generation results to enable the network to complete a specific task. In this case, the learnable module is responsible for converting a two-dimensional image into a panoramic image.
[0050] Here, the learnable module can use the controlnet method to copy the parameters of the original diffusion model, learn the copied parameters, and make the copied parameters complete specific tasks.
[0051] like Figure 4 As shown, Figure 4 A schematic diagram of an embodiment of a three-dimensional scene generation method for generating a panoramic image by fine-tuning the original diffusion model is shown. Figure 4 In [1], inputting a text description into the original generative diffusion model generates a regular 2D image. By freezing the parameters of the original generative diffusion model and inserting a parameter fine-tuning module, a panoramic image of the text description can be output. The parameters in the parameter fine-tuning module are learnable. By inserting the parameter fine-tuning module into the original generative diffusion model, the model's generation results can be adjusted, allowing the model to generate panoramic images.
[0052] Diffusion models are a powerful technique for generating text samples, but they require a large number of network parameters and require extensive training. To reduce training time, we propose freezing the parameters of the original generative diffusion model and inserting a learnable module into the model to adjust the model's generation results.
[0053] In some optional implementation manners, the above-mentioned learnable module may include a low-rank matrix obtained by decomposing the parameter matrix of the above-mentioned original diffusion model using the Low-Rank Adaptation (LORA) technique.
[0054] As Figure 5 shown, Figure 5 it shows a schematic diagram of another embodiment of generating a panoramic view by fine-tuning an original diffusion model in a three-dimensional scene generation method. In Figure 5 this, the text x 501 is input into the target diffusion model 502, and a panoramic view h 503 can be obtained. The target diffusion model 502 may be composed of an original extended model and a learnable module with learnable parameters.
[0055] The mathematical expression of low-rank adaptation can be represented by matrix decomposition. Suppose there is a parameter matrix W of an original generative diffusion model, whose shape is d×d, where the dimensions of the input features and the output features are both d. The goal of low-rank adaptation is to achieve parameter compression and simplification by decomposing the parameter matrix W into the product of two lower-rank matrices. This decomposition is usually achieved using Singular Value Decomposition (SVD) or other low-rank approximation algorithms. Suppose the parameter matrix W is decomposed into the product of two lower-rank matrices A and B, that is, W = A×B, where the shape of A is d×r, the shape of B is r×d, r is the lower rank, and r << d. In this solution, A can be initialized with a standard normal distribution and B with 0.
[0056] This way uses the low-rank adaptation technique and utilizes the properties of low-rank matrices to perform parameter compression and simplification on the generative diffusion model. By decomposing the parameter matrix of the original model into a low-rank approximate representation, the storage requirements and computational complexity of the model can be significantly reduced. By appropriately adjusting and updating the parameters of the low-rank matrix, effective tuning of the model can be achieved.
[0057] In some optional implementation manners, the above-mentioned three-dimensional scene model may include 3D-Gaussian Splatting. 3D-Gaussian Splatting is an explicit 3D scene representation method that uses a set of differentiable 3D Gaussian functions. Each Gaussian function is defined by a center position, a covariance matrix, a color, and an opacity. Specifically, the positions and covariance matrices of 3D Gaussian spheres can be initialized first through the positions of sparse point clouds, and the colors and opacities of 3D Gaussian spheres can be integrated through multi-view pictures and multi-view information. Since 3D-Gaussian Splatting has high rendering quality and high rendering speed at the same time, thus, the solution described in this embodiment can generate high-quality rendering results at a relatively fast speed, improving the real-time performance of the system and the user experience.
[0058] In some optional implementations, after initializing the three-dimensional Gaussian radiation field, the above-mentioned execution entity can project the above-mentioned three-dimensional Gaussian radiation field into each perspective in the multi-perspectives, compare the projected image of the perspective with the multi-perspective image corresponding to the perspective, obtain a loss value, and use the above-mentioned loss value to optimize the parameters of the above-mentioned three-dimensional Gaussian radiation field until the above-mentioned three-dimensional Gaussian radiation field converges, that is, optimize the center position, covariance matrix, color and opacity, etc.
[0059] The above execution entity can determine the loss value through the following formula (1):
[0060]
[0061] in, Represents the total loss value, y i represents the picture from the i-th perspective, Represents the image at the i-th perspective rendered by the Gaussian radiation field, and n represents the number of perspectives.
[0062] In this way, the parameters of the three-dimensional Gaussian radiation field can be optimized and the accuracy of the three-dimensional Gaussian radiation field can be improved.
[0063] Continue to refer Figure 6 , which shows a process 600 of another embodiment of a three-dimensional scene generation method. Figure 6 In this paper, the target text is first input into the text-generated graph model to obtain a panoramic image; the camera poses under multiple perspectives are obtained, and the panoramic image is used to generate multi-perspective images under multiple perspectives; the depth of the panoramic image is estimated to determine the sparse point cloud; then, based on the multi-perspective images, the corresponding camera poses and the sparse point cloud, the three-dimensional Gaussian radiation field described by the target text can be output.
[0064] Further references Figure 7 , which shows a process 700 of another embodiment of a three-dimensional scene generation method. The process 700 of the three-dimensional scene generation method includes the following steps:
[0065] Step 701: Obtain target text and generate a panoramic image described by the target text.
[0066] Step 702: Obtain multi-view information of preset multi-views, and generate multi-view images under multiple viewpoints using the panoramic image.
[0067] Step 703: perform depth estimation on the panoramic image to determine a sparse point cloud corresponding to the panoramic image.
[0068] Step 704: Generate a three-dimensional scene model described by the target text based on the multi-view image, multi-view information, and sparse point cloud.
[0069] In this embodiment, steps 701-704 may be performed in a manner similar to steps 101-104, and are not described again herein.
[0070] Step 705: Determine the current viewing angle, and output scene information of the current viewing angle based on the current viewing angle and the three-dimensional scene model.
[0071] In this embodiment, the execution subject of the three-dimensional scene generation method can determine the current viewing angle, and output the scene information of the current viewing angle according to the current viewing angle and the three-dimensional scene model.
[0072] Here, the current perspective may be a perspective specified by a user, and the execution entity may determine scene information corresponding to the current perspective in the three-dimensional scene model and output the scene information corresponding to the current perspective. The scene information may include, but is not limited to, an image and depth corresponding to the current perspective.
[0073] from Figure 7 It can be seen that Figure 1 Compared to the corresponding embodiment, process 700 of the 3D scene generation method in this embodiment embodies the step of outputting scene information for the current perspective based on the current perspective and the 3D scene model. Thus, the solution described in this embodiment can output scene information for the perspective of interest to the user, providing the user with a more realistic 3D scene and improving the user experience.
[0074] Further references Figure 8 As an implementation of the methods shown in the above figures, the present application provides an embodiment of a three-dimensional scene generation device. Figure 1 Corresponding to the method embodiment shown, the device can be specifically applied to various electronic devices.
[0075] like Figure 8 As shown, the three-dimensional scene generation device 800 of this embodiment includes: an acquisition unit 801, a first generation unit 802, a determination unit 803, and a second generation unit 804. The acquisition unit 801 is used to acquire target text and generate a panoramic image described by the target text; the first generation unit 802 is used to acquire multi-view information of preset multi-view angles and, using the panoramic image, generate a multi-view picture under the multi-view angles; the determination unit 803 is used to perform depth estimation on the panoramic image and determine a sparse point cloud corresponding to the panoramic image; and the second generation unit 804 is used to generate a three-dimensional scene model described by the target text based on the multi-view picture, the multi-view information, and the sparse point cloud.
[0076] In this embodiment, the specific processing of the acquisition unit 801, the first generation unit 802, the determination unit 803 and the second generation unit 804 of the three-dimensional scene generation device 800 can refer to Figure 1 This corresponds to step 101, step 102, step 103 and step 104 in the embodiment.
[0077] In some optional implementations, the acquisition unit 801 may be further used to generate a panoramic image of the target text description in the following manner: using a pre-trained target diffusion model to generate a panoramic image of the target text description, wherein the target diffusion model is used to characterize the correspondence between the text and the panoramic image.
[0078] In some optional implementations, the target diffusion model may be a model obtained by performing a target operation on the original diffusion model, wherein the original diffusion model is used to characterize the correspondence between text and two-dimensional images, and the target operation includes: freezing the parameters of the original diffusion model, inserting a learnable module into the original diffusion model, and the learnable module is used to convert the two-dimensional image into a panoramic image.
[0079] In some optional implementations, the learnable module includes a low-rank matrix obtained by decomposing the parameter matrix of the original diffusion model using a low-rank adaptation technique.
[0080] In some optional implementations, the 3D scene generation device 800 may further include an output unit (not shown in the figure). The output unit is configured to determine a current viewing angle and output scene information of the current viewing angle based on the current viewing angle and the 3D scene model.
[0081] In some optional implementations, the three-dimensional scene model includes a three-dimensional Gaussian radiation field.
[0082] In some optional implementations, the 3D scene generation device 800 may further include an optimization unit (not shown). The optimization unit may be configured to project the 3D Gaussian radiation field onto each of the multiple perspectives, compare the projected image of the perspective with the multi-perspective images corresponding to the perspective, obtain a loss value, and optimize the parameters of the 3D Gaussian radiation field using the loss value.
[0083] Figure 9 An exemplary system architecture 900 is shown to which an embodiment of the three-dimensional scene generation method of the present disclosure can be applied.
[0084] like Figure 9 As shown, system architecture 900 may include terminal devices 9011, 9012, and 9013, a network 902, and a server 903. Network 902 is used as a medium for providing communication links between terminal devices 9011, 9012, and 9013 and server 903. Network 902 may include various connection types, such as wired or wireless communication links or fiber optic cables.
[0085] Users can use terminal devices 9011, 9012, and 9013 to interact with server 903 via network 902 to send or receive messages. For example, terminal devices 9011, 9012, and 9013 can obtain pre-trained target diffusion models from server 903. Various communication client applications can be installed on terminal devices 9011, 9012, and 9013, such as game applications, image capture applications, video processing applications, video playback applications, and instant messaging software.
[0086] Terminal devices 9011, 9012, and 9013 can obtain the target text and generate a panoramic image described by the above target text; then, obtain multi-perspective information of the preset multi-perspectives, and use the above panoramic image to generate a multi-perspective picture under the above multi-perspectives; then, perform depth estimation on the above panoramic image to determine the sparse point cloud corresponding to the above panoramic image; finally, based on the above multi-perspective picture, the above multi-perspective information and the above sparse point cloud, generate a three-dimensional scene model described by the above target text.
[0087] Terminal devices 9011, 9012, and 9013 can be hardware or software. When terminal devices 9011, 9012, and 9013 are hardware, they can be various electronic devices with cameras, display screens, and support information interaction, including but not limited to extended reality devices, smart phones, tablet computers, laptop portable computers, etc. When terminal devices 9011, 9012, and 9013 are software, they can be installed in the electronic devices listed above. They can be implemented as multiple software or software modules (for example, multiple software or software modules for providing distributed services), or they can be implemented as a single software or software module. No specific limitation is made here.
[0088] The server 903 may be a server that provides various services, for example, it may be a background server that provides a pre-trained target diffusion model to the terminal devices 9011 , 9012 , and 9013 .
[0089] It should be noted that the server 903 can be hardware or software. When the server 903 is hardware, it can be implemented as a distributed server cluster consisting of multiple servers, or it can be implemented as a single server. When the server 903 is software, it can be implemented as multiple software or software modules (for example, for providing distributed services), or it can be implemented as a single software or software module. No specific limitation is made here.
[0090] It should also be noted that the three-dimensional scene generation method provided in the embodiment of the present disclosure is usually executed by the terminal devices 9011, 9012, and 9013, and the three-dimensional scene generation device is usually set in the terminal devices 9011, 9012, and 9013.
[0091] It should be understood that Figure 9 The number of terminal devices, networks and servers in the embodiment is merely illustrative. Any number of terminal devices, networks and servers may be provided as required.
[0092] Reference below Figure 10 , which shows an electronic device (eg, Figure 9 The electronic devices in the embodiments of the present disclosure may include, but are not limited to, mobile terminals such as extended reality devices, mobile phones, laptop computers, digital broadcast receivers, PDAs (personal digital assistants), PADs (tablet computers), PMPs (portable multimedia players), in-vehicle terminals (such as in-vehicle navigation terminals), and fixed terminals such as digital TVs and desktop computers. Figure 10 The electronic device shown is only an example and should not bring any limitation to the functions and scope of use of the embodiments of the present disclosure.
[0093] like Figure 10 As shown, the electronic device 1000 may include a processing device (e.g., a central processing unit, a graphics processing unit, etc.) 1001, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 1002 or a program loaded from a storage device 1008 into a random access memory (RAM) 1003. Various programs and data required for the operation of the electronic device 1000 are also stored in the RAM 1003. The processing device 1001, the ROM 1002, and the RAM 1003 are connected to each other via a bus 1004. An input / output (I / O) interface 1005 is also connected to the bus 1004.
[0094] Typically, the following devices may be connected to the I / O interface 1005: an input device 1006 including, for example, a touch screen, a touchpad, a keyboard, a mouse, a camera, a microphone, an accelerometer, a gyroscope, etc.; an output device 1007 including, for example, a liquid crystal display (LCD), a speaker, a vibrator, etc.; a storage device 1008 including, for example, a magnetic tape, a hard disk, etc.; and a communication device 1009. The communication device 1009 may allow the electronic device 1000 to communicate with other devices wirelessly or by wire to exchange data. Figure 10 The electronic device 1000 is shown with various devices, but it should be understood that it is not required to implement or possess all of the devices shown. More or fewer devices may be implemented or possessed instead. Figure 10 Each block shown in the figure may represent one device, or may represent multiple devices as needed.
[0095] In particular, according to an embodiment of the present disclosure, the process described above with reference to the flowchart can be implemented as a computer software program. For example, an embodiment of the present disclosure includes a computer program product, which includes a computer program carried on a computer-readable medium, and the computer program includes a program code for executing the method shown in the flowchart. In such an embodiment, the computer program can be downloaded and installed from the network through the communication device 1009, or installed from the storage device 1008, or installed from the ROM 1002. When the computer program is executed by the processing device 1001, the above-mentioned functions defined in the method of the embodiment of the present disclosure are executed. It should be noted that the computer-readable medium described in the embodiment of the present disclosure can be a computer-readable signal medium or a computer-readable storage medium or any combination of the above two. The computer-readable storage medium can be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, device or device, or any combination of the above. More specific examples of computer-readable storage media may include, but are not limited to, an electrical connection having one or more conductors, a portable computer disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination thereof. In embodiments of the present disclosure, a computer-readable storage medium may be any tangible medium that contains or stores a program that can be used by or in conjunction with an instruction execution system, apparatus, or device. In embodiments of the present disclosure, a computer-readable signal medium may include a data signal propagated in baseband or as part of a carrier wave, which carries computer-readable program code. Such a propagated data signal may take a variety of forms, including, but not limited to, an electromagnetic signal, an optical signal, or any suitable combination thereof. A computer-readable signal medium may also be any computer-readable medium other than a computer-readable storage medium that can transmit, propagate, or transfer a program for use by or in conjunction with an instruction execution system, apparatus, or device. Program code embodied on a computer readable medium may be transmitted using any appropriate medium, including but not limited to wire, optical cable, RF (radio frequency), etc., or any suitable combination thereof.
[0096] The computer-readable medium may be included in the electronic device, or may exist independently without being incorporated into the electronic device. The computer-readable medium carries one or more programs. When executed by the electronic device, the electronic device: obtains target text and generates a panoramic image described by the target text; obtains multi-perspective information of preset multi-perspectives and, using the panoramic image, generates a multi-perspective image under the multi-perspectives; performs depth estimation on the panoramic image and determines a sparse point cloud corresponding to the panoramic image; and generates a three-dimensional scene model described by the target text based on the multi-perspective image, the multi-perspective information, and the sparse point cloud.
[0097] Computer program code for performing the operations of embodiments of the present disclosure may be written in one or more programming languages, or a combination thereof, including object-oriented programming languages such as Java, Smalltalk, C++, and conventional procedural programming languages such as "C" or similar programming languages. The program code may be executed entirely on the user's computer, partially on the user's computer, as a stand-alone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving a remote computer, the remote computer may be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., via the Internet using an Internet service provider).
[0098] The flowcharts and block diagrams in the accompanying drawings illustrate the possible implementation architecture, functions and operations of the systems, methods and computer program products according to various embodiments of the present disclosure. In this regard, each box in the flowchart or block diagram can represent a module, program segment, or a part of code, and the module, program segment, or a part of code contains one or more executable instructions for realizing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the box can also occur in a different order than that marked in the accompanying drawings. For example, two boxes represented in succession can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram and / or flowchart, and the combination of the boxes in the block diagram and / or flowchart, can be implemented with a dedicated hardware-based system that performs the specified function or operation, or can be implemented with a combination of dedicated hardware and computer instructions.
[0099] The units involved in the embodiments described in the present disclosure may be implemented by software or by hardware. The described units may also be provided in a processor, for example, may be described as: a processor including an acquisition unit, a first generation unit, a determination unit, and a second generation unit. The names of these units do not, in some cases, constitute a limitation on the units themselves. For example, the acquisition unit may also be described as a "unit that acquires the target text and generates a panoramic view of the target text description."
[0100] The above description is merely a preferred embodiment of the present disclosure and an explanation of the technical principles employed. Those skilled in the art should understand that the scope of the invention involved in the embodiments of the present disclosure is not limited to the technical solutions formed by a specific combination of the above-mentioned technical features, but should also encompass other technical solutions formed by any combination of the above-mentioned technical features or their equivalents without departing from the above-mentioned inventive concept. For example, a technical solution formed by mutually replacing the above-mentioned features with (but not limited to) technical features with similar functions disclosed in the embodiments of the present disclosure.
Claims
1. A three-dimensional scene generation method, characterized in that: include: Obtaining a target text and generating a panoramic image described by the target text; Acquire multi-view information of preset multi-views, and generate multi-view pictures under the multi-views using the panoramic image; performing depth estimation on the panoramic image to determine a sparse point cloud corresponding to the panoramic image; Based on the multi-view image, the multi-view information and the sparse point cloud, a three-dimensional scene model of the target text description is generated.
2. The method according to claim 1, characterized in that Generating a panoramic view of the target text description includes: A pre-trained target diffusion model is used to generate a panoramic image described by the target text, wherein the target diffusion model is used to characterize the correspondence between the text and the panoramic image.
3. The method according to claim 2, characterized in that The target diffusion model is a model obtained by performing a target operation on the original diffusion model, wherein the original diffusion model is used to characterize the correspondence between text and two-dimensional images. The target operation includes: freezing the parameters of the original diffusion model, inserting a learnable module into the original diffusion model, and the learnable module is used to convert the two-dimensional image into a panoramic image.
4. The method according to claim 3, characterized in that The learnable module includes a low-rank matrix obtained by decomposing the parameter matrix of the original diffusion model using a low-rank adaptation technology.
5. The method according to claim 1, wherein The method further comprises: A current viewing angle is determined, and scene information of the current viewing angle is output based on the current viewing angle and the three-dimensional scene model.
6. The method according to any one of claims 1 to 5, characterized in that The three-dimensional scene model includes a three-dimensional Gaussian radiation field.
7. The method according to claim 6, characterized in that The method further comprises: For each perspective in the multi-perspectives, the three-dimensional Gaussian radiation field is projected into the perspective, and the projected image of the perspective is compared with the multi-perspective image corresponding to the perspective to obtain a loss value, and the loss value is used to optimize the parameters of the three-dimensional Gaussian radiation field.
8. A three-dimensional scene generation device, characterized in that: include: An acquisition unit, configured to acquire a target text and generate a panoramic image described by the target text; A first generating unit is configured to obtain multi-view information of preset multi-views and generate multi-view images under the multi-views using the panoramic image; a determining unit, configured to perform depth estimation on the panoramic image and determine a sparse point cloud corresponding to the panoramic image; The second generating unit is used to generate a three-dimensional scene model described by the target text based on the multi-view image, the multi-view information and the sparse point cloud.
9. An electronic device, characterized in that: include: one or more processors; a storage device having one or more programs stored thereon, When the one or more programs are executed by the one or more processors, the one or more processors implement the method according to any one of claims 1 to 7.
10. A computer-readable medium having a computer program stored thereon, characterized in that: When the program is executed by a processor, the method according to any one of claims 1 to 7 is implemented.