Three-dimensional scene generation method and device, equipment and storage medium
By generating two-dimensional scene images based on scene description text and using a generative AI model to generate three-dimensional scene objects and ground, the problems of tedious manual object combination by users and poor applicability of PCG technology are solved, achieving the effect of simplifying the three-dimensional scene generation process and improving applicability.
Patent Information
- Application Number
- CN202410285715.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-03-12
- Publication Date
- 2025-09-12
AI Technical Summary
In the existing technology, it is cumbersome for users to manually combine objects to generate three-dimensional scenes. The PCG technology that requires professional knowledge has poor applicability and is difficult to output results in real time on the user side.
Generate a two-dimensional scene image based on the scene description text, and generate three-dimensional scene objects and ground through the generative AI model to combine them to obtain a three-dimensional scene.
It simplifies the process of users generating three-dimensional scenes, improves the convenience and applicability of generation, and is suitable for UGC games, film and television, architectural design and VR fields.
Smart Images

Figure CN120635340A_ABST
Abstract
Description
Technical Field
[0001] The embodiments of the present application relate to the field of computer vision technology, and in particular to a three-dimensional scene generation method, apparatus, device, and storage medium. Background Art
[0002] Nowadays, as game content becomes increasingly rich, the demand for building 3D game scenes is also gradually increasing.
[0003] Related technologies offer users the ability to edit game scenes in UGC (User Generated Content) games. Users can manually combine object components, search for objects by category in a toolbox, and manually select and place them in the scene. Furthermore, related technologies employ PCG (Professionally Content Generation) technology, which uses mathematical models to describe object geometry, textures, and relationships between objects and the background, thereby enabling the generation of 3D game scenes.
[0004] However, in the related art, when users manually combine objects, they need to manually place multiple objects, which is rather cumbersome. PCG technology requires users to have strong professional knowledge and has poor applicability in UGC games. Summary of the Invention
[0005] The present invention provides a method, apparatus, device, and storage medium for generating a three-dimensional scene. The technical solution is as follows:
[0006] In one aspect, an embodiment of the present application provides a method for generating a three-dimensional scene, the method comprising:
[0007] generating a two-dimensional scene image based on the scene description text, wherein image content of the two-dimensional scene image conforms to the scene description text, and the image content includes a scene ground and scene objects located on the scene ground;
[0008] generating a three-dimensional scene object that conforms to the scene object image based on the scene object image of the scene object in the two-dimensional scene image;
[0009] generating a three-dimensional scene ground based on the two-dimensional scene image after removing the scene object image, wherein the terrain of the three-dimensional scene ground is consistent with the terrain of the scene ground in the two-dimensional scene image;
[0010] The three-dimensional scene objects and the three-dimensional scene ground are combined to obtain a three-dimensional scene.
[0011] On the other hand, an embodiment of the present application provides a three-dimensional scene generation device, the device comprising:
[0012] A first generating module is configured to generate a two-dimensional scene image based on the scene description text, wherein image content of the two-dimensional scene image conforms to the scene description text, and the image content includes a scene ground and scene objects located on the scene ground;
[0013] a second generating module, configured to generate a three-dimensional scene object conforming to the scene object image based on the scene object image of the scene object in the two-dimensional scene image;
[0014] a third generating module, configured to generate a three-dimensional scene ground based on the two-dimensional scene image after removing the scene object image, wherein the terrain of the three-dimensional scene ground is consistent with the terrain of the scene ground in the two-dimensional scene image;
[0015] The combination module is used to combine the three-dimensional scene objects and the three-dimensional scene ground to obtain a three-dimensional scene.
[0016] On the other hand, an embodiment of the present application provides a computer device, which includes a processor and a memory, wherein the memory stores at least one instruction, at least one program, a code set or an instruction set, and the at least one instruction, the at least one program, the code set or the instruction set is loaded and executed by the processor to implement the three-dimensional scene generation method as described in the above aspects.
[0017] On the other hand, a computer-readable storage medium is provided, wherein the computer-readable storage medium stores at least one instruction, at least one program, a code set, or an instruction set, and the at least one instruction, the at least one program, the code set, or the instruction set is loaded and executed by a processor to implement the three-dimensional scene generation method as described in the above aspects.
[0018] In another aspect, embodiments of the present application provide a computer program product or computer program, which includes computer instructions stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium and executes the computer instructions, causing the computer device to perform the three-dimensional scene generation method provided in the above aspects.
[0019] In an embodiment of the present application, a two-dimensional scene image is first generated based on the scene description text, and then a three-dimensional scene is obtained based on the two-dimensional scene image. The three-dimensional scene corresponding to the text description can be directly generated, thereby improving the convenience of generating the three-dimensional scene. For example, in a UGC game, the player can be directly guided to input a description text to generate a three-dimensional scene corresponding to the description text, which is conducive to increasing the playability of the UGC game. In the process of obtaining a three-dimensional scene based on a two-dimensional scene image, a three-dimensional scene object is first generated based on the scene object image in the two-dimensional scene image, and then a three-dimensional scene ground is generated based on the two-dimensional scene image with the scene object image removed, and finally the three-dimensional scene object is combined with the three-dimensional scene ground to obtain a three-dimensional scene. By generating a three-dimensional scene through the solution provided in the embodiment of the present application, it is only necessary to obtain the scene description text provided by the user, which simplifies the process for the user to complete the three-dimensional scene creation and has a wider applicability. BRIEF DESCRIPTION OF THE DRAWINGS
[0020] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.
[0021] Figure 1 A schematic diagram of a three-dimensional scene construction interface is shown;
[0022] Figure 2 A schematic diagram of a virtual scene generated using PCG technology is shown;
[0023] Figure 3 A schematic diagram showing an implementation environment provided by an exemplary embodiment of the present application is shown;
[0024] Figure 4 A schematic diagram of a three-dimensional scene generation interface provided by an exemplary embodiment of the present application is shown;
[0025] Figure 5 A flowchart of a three-dimensional scene generation method provided by an exemplary embodiment of the present application is shown;
[0026] Figure 6 A schematic diagram showing the correspondence between scene description text and two-dimensional scene images provided by an exemplary embodiment of the present application is shown;
[0027] Figure 7 A schematic diagram showing a two-dimensional scene image and a mask corresponding to a scene object image provided by an exemplary embodiment of the present application is shown;
[0028] Figure 8A flowchart of a three-dimensional scene ground generation process provided by an exemplary embodiment of the present application is shown;
[0029] Figure 9 A schematic diagram showing a two-dimensional scene image from a bird's-eye view and a bird's-eye view mask scene image provided by an exemplary embodiment of the present application is shown;
[0030] Figure 10 A schematic diagram showing a two-dimensional scene image and a bird's-eye view complementary ground image provided by an exemplary embodiment of the present application is shown;
[0031] Figure 11 A schematic diagram showing a two-dimensional scene image, a masked scene image, and a completed ground image provided by an exemplary embodiment of the present application is shown;
[0032] Figure 12 A schematic diagram of a three-dimensional scene ground provided by an exemplary embodiment of the present application is shown;
[0033] Figure 13 A flowchart of a process for acquiring a three-dimensional scene object provided by an exemplary embodiment of the present application is shown;
[0034] Figure 14 A schematic diagram of a scene object image provided by an exemplary embodiment of the present application is shown;
[0035] Figure 15 A schematic diagram of determining a target angle image sequence number provided by an exemplary embodiment of the present application is shown;
[0036] Figure 16 A flowchart of a process for generating a three-dimensional scene provided by an exemplary embodiment of the present application is shown;
[0037] Figure 17 A schematic diagram of mapping a center point and an image range provided by an exemplary embodiment of the present application is shown;
[0038] Figure 18 A schematic diagram showing a three-dimensional scene generation process provided by an exemplary embodiment of the present application is shown;
[0039] Figure 19 This is a structural block diagram of a three-dimensional scene generation device provided by an exemplary embodiment of the present application;
[0040] Figure 20 A schematic structural diagram of a computer device provided by an exemplary embodiment of the present application is shown. DETAILED DESCRIPTION
[0041] In order to make the objectives, technical solutions and advantages of this application clearer, the implementation methods of this application will be further described in detail below with reference to the accompanying drawings.
[0042] UGC games refer to user-generated content games. They are based on user needs and users can publish content on the game platform to show it to other users or for other users to play.
[0043] PCG is a toolset for procedural content and tools used to create virtual content in Unreal Engine. With PCG technology, technical artists, designers, and programmers can build fast-iterative tools and content of any complexity, from asset tools (such as building or biome generation) to entire worlds, meaning they can build three-dimensional scenes through PCG.
[0044] Nowadays, three-dimensional scene generation plays an important role in film and television and game design, industrial design, architectural design, interior design and VR (Virtual Reality).
[0045] In the related art, during the process of generating a three-dimensional scene, there may be the following two ways to achieve the construction of the three-dimensional scene.
[0046] First, 3D scenes can be constructed by manually combining object components. This involves creating a 3D object library and a terrain library. Users can manually select terrain and 3D objects from the library, then drag or click to place the objects within the terrain, thereby creating a single 3D object.
[0047] This method is widely used in UGC games, please refer to Figure 1 , which shows a schematic diagram of a 3D scene construction in the present application. The figure shows a toolbox 101 and a 3D scene 102. The toolbox contains different 3D objects. Users can search for objects in the toolbox. For example, searching for houses will display a variety of houses. Users can select (click, drag, etc.) 3D objects in the toolbox 101 and place them in the 3D scene 102, thereby constructing the 3D scene.
[0048] Second, based on automated PCG technology, various mathematical models are used to describe collections of objects, the relationships between textured objects, and the relationships between objects and their backgrounds, thereby constructing three-dimensional scenes based on these relationships. Wave function collapse is widely used in PCG technology. It can programmatically generate virtual scenes through wave function collapse algorithms, increasing system stability and ease of design and construction, and is commonly used in various gaming scenarios.
[0049] For illustration, please refer to Figure 2 , which shows a schematic diagram of a virtual scene generated using PCG technology.
[0050] However, in solutions that manually combine object components to generate 3D scenes, the user must manually position each object, resulting in a lengthy creation process, high aesthetic standards, and a complex production process. Solutions that use PCG technology to generate 3D scenes often require a certain theoretical foundation, can incur significant computational costs, and place high demands on the device's computing performance. This makes it difficult to output results in real time on the user's end, resulting in poor applicability.
[0051] Therefore, the embodiment of the present application provides a three-dimensional scene generation method that can directly generate a three-dimensional scene based on a natural language description text. It has low requirements for users, simplifies the three-dimensional scene generation process on the user side, and has strong applicability in multiple scenarios.
[0052] Please refer to Figure 3 , which shows a schematic diagram of an implementation environment provided by an exemplary embodiment of the present application. The implementation environment includes a terminal 310 and a server 320. Data communication between the terminal 310 and the server 320 is performed via a communication network. Optionally, the communication network can be a wired network or a wireless network, and the communication network can be at least one of a local area network, a metropolitan area network, and a wide area network.
[0053] Terminal 310 is a computer device installed with an application program having a 3D scene generation function. The 3D scene generation function may be a function of a remote application in terminal 310 or a function of a third-party application. Terminal 310 may be a smartphone, tablet computer, laptop computer, desktop computer, smart TV, wearable device, or vehicle-mounted terminal, etc. Figure 3 In the description, the terminal 310 is taken as a desktop computer as an example, but this is not a limitation.
[0054] Server 320 can be an independent physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, content delivery networks (CDNs), and big data and artificial intelligence platforms. In the embodiment of the present application, server 320 can be a backend server for an application with a three-dimensional scene generation function.
[0055] In one possible implementation, Figure 1As shown, there is data interaction between the terminal 310 and the server 320. Taking the execution of the three-dimensional scene generation method provided in the embodiment of the present application by the server 320 as an example, when receiving the scene description text and the three-dimensional scene generation instruction, the terminal 310 sends the scene description text to the server 320, and the server 320 generates a two-dimensional scene image based on the scene description text, generates a three-dimensional scene object based on the scene object image in the two-dimensional scene image, and generates a three-dimensional scene ground based on the two-dimensional scene image after removing the scene object image. Finally, the three-dimensional scene ground is combined with the three-dimensional scene object to obtain a three-dimensional scene. After obtaining the three-dimensional scene, the server 320 returns the three-dimensional scene to the terminal 310 and displays the three-dimensional scene to the user through the terminal 310.
[0056] In another possible implementation, the terminal 310 can independently execute the three-dimensional scene generation method provided in the embodiment of the present application, that is, the terminal does not need to interact with the server 320. Upon receiving the scene description text and the three-dimensional scene generation instruction, the terminal 310 generates a two-dimensional scene image based on the scene description text, and generates a three-dimensional scene based on the two-dimensional scene image.
[0057] In another possible embodiment, the terminal 310 and the server 320 collaborate to generate a two-dimensional scene image based on the scene description text when the terminal 310 receives the scene description text and the three-dimensional scene generation instruction, and then send the two-dimensional scene image to the server 320, so that the server 320 generates a three-dimensional scene based on the two-dimensional scene image and returns it to the terminal 310.
[0058] For the convenience of description, the following embodiments are described by taking the three-dimensional scene generation method executed by a computer device as an example.
[0059] The 3D scene generation solution provided in the embodiments of the present application can be applied to at least the following scenarios:
[0060] 1. In the field of gaming, during the game development process, the solution provided by this application can be used to create game scenes based on scene description text entered by the developer and then apply the game scenes to the game. Optionally, applicable game types include Multiplayer Online Battle Arena (MOBA), Simulation Game (SLG), and party games, among others.
[0061] The UGC game client provides users with a game scene generation interface. There is a description text input box in the scene generation interface, allowing users to enter scene description text from the text input box and generate a three-dimensional scene based on the scene description text in the background, which can provide users with a convenient three-dimensional scene generation solution.
[0062] For illustration, please refer to Figure 4 , which shows a schematic diagram of a 3D scene generation interface provided by an exemplary embodiment of the present application. The 3D scene generation interface includes a description text input box 401 and a 3D scene display area 402. After receiving the scene description text entered into description text input box 401, a 2D scene image corresponding to the scene description text is first generated. Then, corresponding 3D scene objects and a 3D scene ground are generated based on the 2D scene image. Finally, the 3D scene objects and the 3D scene ground are combined to generate a 3D scene, and the 3D scene effect is displayed in 3D scene display area 402.
[0063] 2. In the field of film and television, in the scenes of movies and TV series, two-dimensional scene images can be generated based on the descriptive text of the environment in the script, and then a virtual three-dimensional scene can be generated based on the two-dimensional scene image. Combined with the captured video content, it is possible to add a virtual scene to the video screen.
[0064] 3. In the field of architectural design, according to the descriptive text of the design requirements provided by the designer, the corresponding two-dimensional scene image is generated, and then the three-dimensional scene is further generated. The designer can generate three-dimensional scenes of different styles by changing the descriptive text.
[0065] 4. VR virtual reality field. In the VR field, the solution provided by the embodiment of the present application can generate a 360-degree virtual three-dimensional space based on the description text provided by the developer, and enable users to obtain a real visual experience through special imaging technology.
[0066] It should be noted that the solution provided in this application can also be used in fields such as interior design, and does not limit the application scope of the embodiments of this application.
[0067] Please refer to Figure 5 , which shows a flowchart of a three-dimensional scene generation method provided by an exemplary embodiment of the present application, the method comprising the following steps:
[0068] Step 501: Generate a two-dimensional scene image based on the scene description text.
[0069] The image content of the two-dimensional scene image conforms to the scene description text, and the image content includes the scene ground and scene objects located on the scene ground.
[0070] The scene description text is used to describe the content of the 3D scene to be generated, and may include the terrain, buildings, and landforms of the 3D scene. For example, the 3D description text may be "a 2D image of a town in the game style." Furthermore, during the generation of the 2D scene image, some inverse descriptions may be used to avoid generating the inverse description content. For example, the inverse description content may be "bird's-eye view, top-down perspective, ugly, blurry," which is used to prevent the computer device from generating a bird's-eye view from a top-down perspective, and to avoid generating ugly and blurry 2D scene images.
[0071] For illustration, please refer to Figure 6 , which shows a schematic diagram of the correspondence between scene description text and a two-dimensional scene image provided by an exemplary embodiment of the present application. The scene description text is: Generate a two-dimensional image of a modern city with many buildings, and the computer device generates a two-dimensional scene image 601 based on the scene description text.
[0072] In a possible implementation, the computer device may generate a two-dimensional scene image based on the scene description text by using SDXL (Stable Diffusion XL, XL version of the stable diffusion model).
[0073] Optionally, before using SDXL to generate a two-dimensional scene image based on scene description text, the SDXL model can be fine-tuned based on sample scene description text and sample two-dimensional scene images, so that after the fine-tuned SDXL model is deployed in a computer device, it can generate a two-dimensional scene image that conforms to the current application scenario based on the scene description text. For example, if the solution provided in the embodiment of the present application is used to generate a cartoon-style three-dimensional scene, it is necessary to fine-tune the SDXL model using multiple scene description texts and cartoon-style sample two-dimensional scene images, so that the fine-tuned SDXL model can generate a cartoon-style two-dimensional scene image based on the input scene description text, and the computer device can further generate a cartoon-style three-dimensional scene. For another example, the SDXL model can be fine-tuned so that the model can generate a two-dimensional scene image of a specific perspective. For example, if the SDXL model is to be able to generate a high-quality isometric view after being deployed in a computer device, the SDXL model can be fine-tuned using sample description text and sample two-dimensional scene images under the isometric view.
[0074] Step 502 : generating a three-dimensional scene object that conforms to the scene object image based on the scene object image of the scene object in the two-dimensional scene image.
[0075] Optionally, scene objects are all objects above the ground, which may include buildings, vehicles, trees, streetlights, pedestrians, bridges, etc. Furthermore, the scene ground is not necessarily flat, and the scene ground may be the ground in different terrains, such as mountains, hills, plains, and deserts, etc. This embodiment does not limit this.
[0076] In one possible implementation, a computer device uses a generative AI (artificial intelligence) model to generate a three-dimensional scene object corresponding to a scene object image based on the scene object image. The computer device inputs the scene object image into the generative AI model to obtain a three-dimensional scene object generated by the generative AI model.
[0077] In another possible implementation, the computer device acquires the three-dimensional object component corresponding to the scene object image by a three-dimensional object retrieval method.
[0078] Step 503 : Generate a three-dimensional scene ground based on the two-dimensional scene image after removing the scene object images.
[0079] The terrain of the ground in the three-dimensional scene is consistent with the terrain of the ground in the two-dimensional scene image.
[0080] When generating 3D scene objects, a 3D scene ground needs to be generated. The 3D scene ground is obtained by removing the 2D scene image from the scene object image. Optionally, image segmentation technology can be used to segment the scene object image to obtain the 2D scene ground after removing the 2D scene image.
[0081] Optionally, a generative AI model is used to generate a 3D scene ground from a 2D scene image after removing scene object images. Optionally, the ground texture in the 2D scene image is used as the texture of the 3D scene ground, thereby obtaining the 3D scene ground through mapping.
[0082] Step 504: Combine the three-dimensional scene objects and the three-dimensional scene ground to obtain a three-dimensional scene.
[0083] In the process of combining the three-dimensional scene objects and the three-dimensional scene ground, the computer device places the three-dimensional scene objects on the three-dimensional scene ground. It is also possible that the three-dimensional scene objects are suspended above the three-dimensional scene ground.
[0084] Optionally, the final 3D scene can be exported to various 3D model formats, such as fbx format, glb format, glttf format, and obj format, or the 3D scene can be directly loaded into a 3D game engine.
[0085] In summary, in the embodiment of the present application, a two-dimensional scene image is first generated based on the scene description text, and then a three-dimensional scene is obtained based on the two-dimensional scene image. The three-dimensional scene corresponding to the text description can be directly generated, thereby improving the convenience of generating the three-dimensional scene. For example, in a UGC game, the player can be directly guided to input a description text to generate a three-dimensional scene corresponding to the description text, which is conducive to increasing the playability of the UGC game. In the process of obtaining a three-dimensional scene based on a two-dimensional scene image, a three-dimensional scene object is first generated based on the scene object image in the two-dimensional scene image, and then a three-dimensional scene ground is generated based on the two-dimensional scene image with the scene object image removed, and finally the three-dimensional scene object is combined with the three-dimensional scene ground to obtain a three-dimensional scene. By generating a three-dimensional scene through the solution provided in the embodiment of the present application, it is only necessary to obtain the scene description text provided by the user, which simplifies the process for the user to complete the three-dimensional scene creation and has a wider applicability.
[0086] In a possible implementation, image segmentation is performed on the two-dimensional scene image to obtain a scene object image and a two-dimensional scene image (masked scene image) after the scene object image is removed.
[0087] The computer device performs image segmentation on the two-dimensional scene image to obtain a scene object image, and masks the two-dimensional scene image with the scene object image to obtain a masked scene image. Optionally, the scene object image can be segmented using a SAM (SegmentAnything Model) to obtain an object image mask.
[0088] For illustration, please refer to Figure 7 , which shows a schematic diagram of a two-dimensional scene image and a scene object image mask provided by an exemplary embodiment of the present application. The two-dimensional scene image 701 is a scene image of urban buildings, and the scene object image mask 702 is an image mask corresponding to the scene object image in the two-dimensional scene image 701.
[0089] Since the 2D scene image after removing the scene object images only contains a portion of the scene ground texture, to ensure that the 3D scene ground has complete ground texture, the computer device can first perform ground completion on the masked scene image to obtain a completed ground image. The completed ground image has complete ground image texture, and then the 3D scene ground is generated based on the completed ground image, where the topography of the 3D scene ground is consistent with that of the completed ground image.
[0090] In one possible implementation, the quality of the completed ground image obtained by ground completion on the masked scene image may be low, for example, with unnatural transitions at the edges of the completed region. This can be further optimized using an image optimization network, resulting in a higher level of detail than the unoptimized ground image. The 3D scene ground is then generated based on the optimized ground image. Optionally, a StableDiffusion XL refiner network can be used to optimize the completed ground image.
[0091] Optionally, the computer generates an isometric image of the scene based on the scene description text. Using isometric images, which are readily available and readily available for game scene isometrics, can fine-tune the SDXL model, enabling it to generate high-quality two-dimensional scene images based on the scene description text. Furthermore, because isometric images are parallel projections, they can be easily stitched together from multiple isometric images taken from different camera positions to form a large scene image. Furthermore, the computer can easily convert the isometric image into a bird's-eye view.
[0092] The process of generating a three-dimensional scene ground will be described below through an exemplary embodiment.
[0093] Please refer to Figure 8 , which shows a flowchart of a three-dimensional scene ground generation process provided by an exemplary embodiment of the present application, the process includes the following steps.
[0094] Step 801 : converting the mask scene image into a bird's-eye view mask scene image.
[0095] Since a scene image with a bird's-eye view is more conducive to ground completion, the computer device converts the mask scene image into a bird's-eye view mask scene image before performing ground completion.
[0096] Optionally, since converting the mask scene image into a BEV (Bird's Eye View) image with a standard orientation will produce some blank areas around it, the mask scene image may be converted into a 30° oriented bird's eye view image.
[0097] The computer device considers the depths of objects in the mask scene image to be on the same plane. A homography transformation can be used to convert the mask scene image to a bird's-eye view mask scene image. The computer device multiplies the width of the mask scene image by 0.57735 to obtain a bird's-eye view mask scene image rotated 30°.
[0098] For illustration, please refer to Figure 9 , which shows a schematic diagram of a two-dimensional scene image from a bird's-eye view and a bird's-eye view mask scene image provided by an exemplary embodiment of the present application. The bird's-eye view mask scene image 902 is obtained by performing homography transformation based on the two-dimensional scene image 901.
[0099] Step 802 : Perform ground completion on the bird's-eye view masked scene image through a ground completion network to obtain a completed ground image.
[0100] The ground completion network may be a GAN (Generative Adversarial Networks), an Inpainting (image restoration) network, etc., which is not limited in this embodiment.
[0101] The computer device inputs the bird's-eye view mask image into the ground completion network to obtain a completed ground image output by the ground completion network. Both the completed ground image and the bird's-eye view mask scene image are bird's-eye view images.
[0102] In a possible implementation, a diffusion-based inpainting network may be used for ground completion, specifically in the following two ways.
[0103] Method 1: Fine-tune the stable diffusion model with image restoration capabilities, deploy the fine-tuned Inpainting network on a computer device, and input the bird's-eye view mask image into the fine-tuned Inpainting network to obtain the completed ground image.
[0104] During the training of the Inpainting network, assuming that the size of the sample ground-completing image to be generated is (H, W), the input variable of the denoising network (U-net) in the diffusion model is the latent variable obtained by encoding the sample 2D scene image from a bird's-eye view through an encoder, with a size of (H / 8, W / 8, 4). The encoder can be a VAE (Variational Auto-Encoder). Correspondingly, the size of the sample scene object mask image is also (H, W). In order to incorporate the mask information of the sample scene object into the input information of the model, the sample scene object mask image needs to be compressed through the encoder, and the compressed size is (H / 8, W / 8, 4). The final input variable of the Inpainting network is x′=cat(f(xM),f(M)), where f is the image encoder, M is the mask (the part to be supplemented is 0, and the other parts are 1), x represents the sample 2D scene image, and cat combines the two variables in the feature dimension.
[0105] In one possible implementation, when training the Inpainting network, the sample scene ground image is first masked to obtain a sample masked scene image. This mask can be a random mask or a mask that masks specific objects in the sample scene ground image. The sample masked scene image is then input into the ground completion network to obtain the sample completed ground image as output.
[0106] The diffusion-based Inpainting network includes two processes: noise addition and denoising. The noise addition process is Among them, x0 represents the input sample mask scene image, x t It represents the result of adding noise to the sample mask scene image at time t, is the hyperparameter of noise addition, q(x t |x0) represents the x in the noise adding process t The distribution of which the mean is The variance is Gaussian distribution.
[0107] Optionally, in order to use the reverse gradient transfer algorithm to q(x t |x0) to optimize, x t Reparameterize and get Among them, ε~N(0,1) is the standard normal distribution.
[0108] After the noisy sample mask scene image is denoised by the denoising network and then decoded, the sample completed ground image output by the ground completion network can be obtained.
[0109] The completion loss can be determined based on the difference between the sample completed ground image and the sample scene ground image.
[0110] Optionally, the loss function of the completion loss is:
[0111]
[0112] Among them, θ is the parameter to be optimized, ∈ θ is the denoising function, and E is the variational lower bound.
[0113] Method 2: Perform inpainting at test time, that is, use the image object mask to guide the generation of the completed ground image during the test phase of the inpainting network.
[0114] This approach does not require retraining the Inpainting network. Assuming that the Inpainting network needs to undergo K steps of denoising during its test phase, a common diffusion model is used for unconditional denoising in the first K-1 steps. During the first K-1 steps, only the portion of the image to be completed, i.e., the mask, is updated. In the final denoising step, the entire image is denoised to reduce the blurring of the mask boundary, thereby obtaining a completed ground image.
[0115] For illustration, please refer to Figure 10 , which shows a schematic diagram of a two-dimensional scene image and a completed ground image from a bird's-eye view, provided by an exemplary embodiment of the present application. The two-dimensional scene image 1001 is an isometric view image. The computer device converts the two-dimensional scene image, which extracts scene object images, into a bird's-eye view mask scene image. The computer then completes the bird's-eye view mask image using a ground completion network, resulting in a completed ground image 1002 from a bird's-eye view.
[0116] Indicative, Figure 11 A schematic diagram illustrating a two-dimensional scene image, a masked scene image, and a completed ground image provided by an exemplary embodiment of the present application is shown. Masked scene image 1102 is a two-dimensional scene image obtained by removing scene object images from two-dimensional scene image 1101, and completed ground image 1103 is a complete ground image obtained by performing ground completion on the masked scene image.
[0117] Step 803: Generate a three-dimensional ground based on the size of the completed ground image.
[0118] The terrain of the three-dimensional ground is consistent with the terrain of the completed ground image.
[0119] First, the computer device determines the size of the completed ground image, thereby determining the size of the unit horizontal ground, and the size of the three-dimensional horizontal ground matches the size of the completed ground image.
[0120] For example, if a completed ground image from a bird's-eye view is 100 meters long and 60 meters wide, representing horizontal terrain, a computer can use a 3D engine such as Blender or Unity to generate a 3D horizontal ground surface that matches the dimensions of the completed ground image. The 3D engine can then map the completed ground image onto the horizontal ground surface, resulting in a complete 3D scene with horizontal terrain.
[0121] For another example, assuming that the terrain of the three-dimensional ground is hilly, the computer device generates a hilly terrain through a 3D engine such as Blender or Unity. The size of the terrain matches the completed ground image, and then the completed ground image is mapped onto the generated hilly terrain to obtain the three-dimensional ground.
[0122] Step 804 : performing mapping processing on the three-dimensional ground based on the completed ground image to obtain a three-dimensional scene ground.
[0123] Optionally, the computer device pastes the completed ground image onto the three-dimensional ground through a 3D engine such as Blender or Unity, thereby obtaining a three-dimensional scene ground.
[0124] For illustration, please refer to Figure 12 , which shows a schematic diagram of a three-dimensional scene ground provided by an exemplary embodiment of the present application. The ground is a horizontal terrain and there are no three-dimensional objects on the ground.
[0125] In the embodiments of the present application, a computer device uses a segmentation model to segment objects and the ground within a scene object image. The computer device then converts the two-dimensional scene image, after removing the scene object image, to a bird's-eye view, facilitating ground completion and obtaining a complete ground texture. Furthermore, the computer device generates a three-dimensional scene ground based on the completed ground image, ensuring that the generated three-dimensional scene ground has the same ground effect as the two-dimensional scene image and complies with the scene description text.
[0126] In the embodiments of this application, a 3D scene is generated by combining 3D scene objects and a 3D scene ground. Therefore, before generating the 3D scene, the 3D scene objects must be acquired. Alternatively, a 3D object search can be performed in a 3D asset library based on the scene object image to obtain the 3D scene object corresponding to the scene object image. The following describes the process of acquiring 3D scene objects using an exemplary embodiment.
[0127] Please refer to Figure 13 , which shows a flowchart of a process for obtaining a three-dimensional scene object provided by an exemplary embodiment of the present application, the process includes the following steps:
[0128] Step 1301: Encode the scene object image to obtain scene object image features.
[0129] Optionally, a CLIP (Contrastive Language-Image Pre-Training) model may be used to encode scene object images to obtain scene object image features. Alternatively, scene object images may be encoded using other image encoders, which are not limited in this embodiment.
[0130] For illustration, please refer to Figure 14 , which shows a schematic diagram of scene object images provided by an exemplary embodiment of the present application, including a first scene object image 1401 , a second scene object image 1402 and a third scene object image 1403 .
[0131] Step 1302 : Match the scene object image features with the angle image features in the 3D asset library to obtain a matching result.
[0132] Among them, there are at least two angle images of candidate three-dimensional objects at different angles and angle image features corresponding to each angle image in the three-dimensional asset library, and the matching result is used to represent the similarity between the scene object image and each angle image.
[0133] When building a 3D asset library, a computer first renders angled images of multiple candidate 3D objects offline. For each candidate object, N angled images are rendered at different angles, where N can be 12. The CLIP model then extracts image features from each angled image to obtain angled image features. Assuming that the angled image features are M-dimensional vectors, and that the features of each candidate 3D object are represented by a combination of 12 angled images, the candidate object features are an N*12*M matrix.
[0134] In the process of searching for three-dimensional objects, the scene object image is first encoded by CLIP to obtain the scene object image features with a size of 1*M. Then the cosine feature similarity between the angle image features and the scene object image features is calculated to obtain the matching result.
[0135] Step 1303 : Determine a three-dimensional scene object from the candidate three-dimensional objects based on the matching result.
[0136] In a possible implementation, the computer device determines an angle image indicated by the matching result that has the highest similarity to the scene object image, and determines the candidate three-dimensional object corresponding to the angle image as the three-dimensional scene object.
[0137] In another possible implementation, in order to more accurately determine the three-dimensional scene object corresponding to the scene object image, the three-dimensional scene object is determined from the candidate three-dimensional objects by voting.
[0138] First, based on the matching result, the computer device determines an angle image whose similarity with the scene object image is higher than a similarity threshold as a target angle image.
[0139] Optionally, the computer device determines K angle images with the highest similarity to the scene object image as target angle images based on the matching results, and the value range of K can be 10 to 50.
[0140] The computer device first selects a portion of angle images with a high similarity to the scene object image from the angle images. For example, if the similarity threshold is 0.75, the angle images with a similarity greater than 0.75 are determined as target angle images.
[0141] Optionally, the computer device calculates the cosine feature similarity between the scene object image and all angle images in the 3D asset library. Assuming that the image sequence numbers of images at different angles are different, the computer device determines the image sequence number of the target angle image based on the similarity.
[0142] For illustration, please refer to Figure 15 , which shows a schematic diagram of determining the target angle image sequence number provided by an exemplary embodiment of the present application. Angle image rendering is performed on candidate 3D objects in the 3D asset library to obtain angle images. Features of the angle images are extracted using a CLIP image encoder to obtain angle image features. Feature extraction is also performed on scene object images to obtain scene object image features. Cosine feature similarity is then calculated between the angle image features and the scene object image features to obtain matching results. This results in the image sequence numbers of the K target angle images with the highest similarity, i.e., the Top-K asset sequence numbers.
[0143] Subsequently, based on the correspondence between the target angle image and the candidate three-dimensional object, the candidate three-dimensional object is voted to obtain a voting result.
[0144] Voting for candidate 3D objects involves determining the number of target angle images corresponding to each 3D object. Since each angle image corresponds to a single candidate 3D object, the computer can determine the candidate 3D object corresponding to each selected target angle image. Multiple target angle images may correspond to the same candidate 3D object, resulting in that candidate receiving multiple votes. For example, if there are five target angle images, three of which correspond to the first candidate 3D object and two correspond to the second candidate 3D object, the first candidate 3D object will receive three votes, while the second candidate 3D object will receive two votes.
[0145] Finally, the candidate three-dimensional object that has obtained the most votes as indicated by the voting results is determined as the three-dimensional scene object.
[0146] The more votes a candidate 3D object receives, the more angles its images share a high degree of similarity with the scene object image. The closer the candidate 3D object is to the scene object image, the more likely it is to be the 3D scene object corresponding to the scene object image. Therefore, the candidate 3D object with the most votes is determined as the 3D scene object.
[0147] In an embodiment of the present application, a computer device can select a corresponding three-dimensional scene object from the candidate three-dimensional objects by calculating the similarity between the object scene image and the angle image corresponding to the candidate three-dimensional objects in the three-dimensional asset library. Compared with the method of generating three-dimensional scene objects based on scene object images through neural networks, the scheme adopted in this embodiment can determine a three-dimensional scene object that is closer to the scene object image, and has lower computing resource requirements during the application process.
[0148] After acquiring the three-dimensional scene objects and the three-dimensional scene ground, the three-dimensional scene objects and the three-dimensional scene ground need to be combined to obtain a three-dimensional scene. During the combination process, the computer device first needs to determine the target placement position of the three-dimensional scene objects on the three-dimensional scene ground based on the two-dimensional scene image, and then place the three-dimensional scene objects at the target placement position on the three-dimensional scene ground to obtain the three-dimensional scene. The target placement position should be consistent with the position of the scene object image in the two-dimensional scene image. The following will illustrate the process of combining the three-dimensional scene ground and the three-dimensional scene objects through an exemplary embodiment.
[0149] Please refer to Figure 16 , which shows a flowchart of a process for generating a three-dimensional scene provided by an exemplary embodiment of the present application, the process comprising the following steps:
[0150] Step 1601 : Determine the center point of the location of the scene object image in the two-dimensional scene image and the image range of the scene object image.
[0151] During the image segmentation process, the computer device can determine the edge coordinates of the scene object image in the two-dimensional scene image, and then determine the center point coordinates and the image range based on the edge coordinates.
[0152] Step 1602: Determine the mapping position of the center point and the image range in the completed ground image as the target placement position.
[0153] Among them, the center point of the target placement position corresponds to the center point of the position of the scene object image, the range of the target placement position corresponds to the image range, and the completed ground image is obtained by completing the ground based on the two-dimensional scene image after removing the scene object image.
[0154] In one possible implementation, the two-dimensional scene image is an isometric view image. In the process of determining the target placement position, the computer device performs homography transformation mapping on the center point and the image range, determines the mapping position of the center point and the image range in the completed ground image, and determines the mapping position as the target placement position.
[0155] If the completed ground image is a bird's-eye view image, it is necessary to map the center point of the isometric view image to the bird's-eye view image and map the image range of the isometric view image to the bird's-eye view image through homography transformation.
[0156] Optionally, if the mask scene image is transformed into a bird's-eye view mask scene image at a 30° viewing angle, the image width of the mask scene image needs to be multiplied by 0.57735. Correspondingly, when using homography to map the center point, the horizontal coordinate of the center point coordinate of the two-dimensional scene image needs to be multiplied by 0.57735 to obtain the horizontal coordinate in the completed ground image after mapping. When mapping the image range, the width of the image range needs to be multiplied by 0.57735 to obtain the width of the mapped image range. For example, if the center point coordinates in the two-dimensional scene image are (x, y), then the center point coordinates in the completed ground image after mapping are (0.57735x, y).
[0157] For illustration, please refer to Figure 17 , which shows a schematic diagram of mapping the center point and image range provided by an exemplary embodiment of the present application. In the two-dimensional scene image 1710, the coordinates of the initial center point 1712 of the scene object image 1711 are (x, y), and the corresponding initial image range width is w and length is h. After being mapped to the completed ground image 1720 at a 30° bird's-eye view through homography transformation, the corresponding mapping center point 1721 coordinates are (0.57735x, y), the mapping image range width is 0.57735w, and the length remains unchanged.
[0158] Step 1603: scaling the three-dimensional scene object so that the size of the scaled three-dimensional scene object complies with the range of the target placement position.
[0159] The size of the 3D scene objects retrieved from the 3D asset library by the computer device may not match the target placement range. If the 3D scene object is too large or too small, the overall visual effect of the 3D scene may be poor after it is placed on the 3D scene floor. Therefore, it is necessary to scale the 3D scene object before combining it with the 3D scene floor to achieve a better 3D scene effect.
[0160] Step 1604 : Place the scaled three-dimensional scene object at the target placement position on the ground of the three-dimensional scene to obtain a three-dimensional scene.
[0161] The computer device places the scaled different three-dimensional scene objects at their corresponding center points to obtain a three-dimensional scene.
[0162] In an embodiment of the present application, the placement position and placement range of three-dimensional scene objects on the three-dimensional scene ground are determined based on the two-dimensional scene image by mapping, so that the finally generated three-dimensional scene can have a better visual effect.
[0163] Schematic, reference Figure 18 , which shows a schematic diagram of the three-dimensional scene generation process provided by an exemplary embodiment of the present application. First, the SDXL model is used to generate a two-dimensional scene image based on the scene description text, corresponding to the IsometricRGB image in the figure, and then the SAM model is used to segment the two-dimensional scene image to obtain a mask scene image and a scene object image. The computer device performs ground completion on the mask scene image through the ground completion network to obtain a completed ground image, and generates a three-dimensional scene ground based on the ground completion image. At the same time, the computer device searches for three-dimensional objects in the three-dimensional asset library to obtain a three-dimensional scene object corresponding to the scene object image. The computer device determines the target placement position of the three-dimensional scene object in the three-dimensional scene ground based on the two-dimensional scene image, and then combines the three-dimensional scene ground and the three-dimensional scene object based on the target placement position to obtain a three-dimensional scene.
[0164] In addition to the three-dimensional scene generation solution shown in the above embodiment, the three-dimensional scene can also be generated in the following manner.
[0165] In some embodiments, a three-dimensional scene can be generated by incremental RGB-D (Red, Green, and Blue Depth map, RG) inpainting. The computer device generates a single-frame RGB-D image based on scene description text. Subsequently, by changing the description text and thus the camera position observing the scene, RGB-D images from different perspectives are drawn. For example, if the scene description text is "a forward-facing image of a town scene," then when the camera position observing the scene is changed, the scene description text may be "a top-down perspective image of the town scene" or "generating images of a town scene from different perspectives."
[0166] Finally, a 3D scene is generated from the RGB-D images from different perspectives through stereo merging. RGB-D images contain depth information about the scene, and each pixel in an image frame has information about its distance from the camera. However, visual cues in a single image are limited. Therefore, to generate a higher level of detail in the generated 3D scene, the number of single-frame RGB-D images can be increased, resulting in a higher-quality 3D scene.
[0167] In other embodiments, a computer device can directly generate a 3D scene based on scene description text. This method requires obtaining sample scene description text and sample 3D scenes in advance, training a 3D scene generation model based on the sample scene description text and sample 3D scenes, and deploying the trained 3D scene generation model on the computer device.
[0168] In the embodiments of the present application, two three-dimensional scene generation schemes are provided, which can also directly generate corresponding three-dimensional scenes based on scene description texts and have strong applicability.
[0169] Figure 19 This is a structural block diagram of a three-dimensional scene generation device provided by an exemplary embodiment of the present application. Figure 19 As shown, the device includes:
[0170] A first generating module 1901 is configured to generate a two-dimensional scene image based on the scene description text, wherein the image content of the two-dimensional scene image conforms to the scene description text, and the image content includes a scene ground and scene objects located on the scene ground;
[0171] A second generating module 1902 is configured to generate a three-dimensional scene object that conforms to the scene object image based on the scene object image of the scene object in the two-dimensional scene image;
[0172] A third generating module 1903 is configured to generate a three-dimensional scene ground based on the two-dimensional scene image after removing the scene object image, wherein the terrain of the three-dimensional scene ground is consistent with the terrain of the scene ground in the two-dimensional scene image;
[0173] The combining module 1904 is configured to combine the three-dimensional scene objects and the three-dimensional scene ground to obtain a three-dimensional scene.
[0174] Optionally, the device further includes:
[0175] A segmentation module is used to perform image segmentation on the two-dimensional scene image to obtain the scene object image and the mask scene image, wherein the mask scene image is the two-dimensional scene image after removing the scene object image, and the mask scene image is obtained by masking the two-dimensional scene image with the object image mask.
[0176] Optionally, the third generating module 1903 is configured to:
[0177] Performing ground completion on the masked scene image to obtain a completed ground image;
[0178] The three-dimensional scene ground is generated based on the completed ground image, and the terrain of the three-dimensional scene ground is consistent with the terrain of the completed ground image.
[0179] Optionally, the two-dimensional scene image and the mask scene image are isometric view images;
[0180] The third generating module 1903 is configured to:
[0181] Converting the mask scene image into a bird's-eye view mask scene image;
[0182] The ground completion network is used to perform ground completion on the bird's-eye view mask scene image to obtain the completed ground image.
[0183] Optionally, the device further includes:
[0184] A training module is used to mask the sample scene ground image from a bird's-eye view to obtain a sample masked scene image;
[0185] The training module is further configured to input the sample masked scene image into the ground completion network to obtain a sample completed ground image output by the ground completion network;
[0186] The training module is further configured to determine a completion loss based on a difference between the sample completed ground image and the sample scene ground image;
[0187] The training module is further configured to train the ground completion network based on the completion loss to obtain the trained ground completion network.
[0188] Optionally, the third generating module 1903 is configured to:
[0189] generating a three-dimensional ground surface based on a size of the completed ground surface image, wherein a topography of the three-dimensional ground surface is consistent with a topography of the completed ground surface image;
[0190] Mapping is performed on the three-dimensional ground based on the completed ground image to obtain the three-dimensional scene ground.
[0191] Optionally, the device further includes:
[0192] an optimization module, configured to optimize the image quality of the completed ground image through an image optimization network to obtain an optimized completed ground image, wherein the optimized completed ground image has a higher level of detail than the unoptimized completed ground image;
[0193] The third generating module 1903 is further configured to generate the three-dimensional scene ground based on the optimized completed ground image.
[0194] Optionally, the second generating module 1902 is configured to:
[0195] A three-dimensional object search is performed in a three-dimensional asset library based on the scene object image to obtain a three-dimensional scene object corresponding to the scene object image.
[0196] Optionally, the second generating module 1902 is configured to:
[0197] Encoding the scene object image to obtain scene object image features;
[0198] Matching the scene object image features with the angle image features in the 3D asset library to obtain a matching result, where the 3D asset library contains angle images of at least two candidate 3D objects at different angles and angle image features corresponding to each angle image, and the matching result is used to represent the similarity between the scene object image and each angle image;
[0199] Based on the matching result, the three-dimensional scene object is determined from the candidate three-dimensional objects.
[0200] Optionally, the second generating module 1902 is configured to:
[0201] Based on the matching result, determining the angle image having a similarity with the scene object image higher than a similarity threshold as a target angle image;
[0202] Voting on the candidate three-dimensional objects based on the correspondence between the target angle image and the candidate three-dimensional objects to obtain a voting result;
[0203] The candidate three-dimensional object that has obtained the most votes as indicated by the voting result is determined as the three-dimensional scene object.
[0204] Optionally, the combining module 1904 is configured to:
[0205] Determining a target placement position of the three-dimensional scene object on the ground of the three-dimensional scene based on the two-dimensional scene image;
[0206] The three-dimensional scene object is placed at the target placement position on the three-dimensional scene ground to obtain the three-dimensional scene.
[0207] Optionally, the combining module 1904 is configured to:
[0208] Determining a center point of a position of the scene object image in the two-dimensional scene image and an image range of the scene object image;
[0209] The mapping position of the center point and the image range in the completed ground image is determined as the target placement position, the center point of the target placement position corresponds to the center point of the position of the scene object image, and the range of the target placement position corresponds to the image range. The completed ground image is obtained by ground completion based on the two-dimensional scene image after removing the scene object image.
[0210] Optionally, the two-dimensional scene image is an isometric view image;
[0211] The combination module 1904 is used to:
[0212] Performing homography transformation mapping on the center point and the image range to determine mapping positions of the center point and the image range in the completed ground image, where the completed ground image is a bird's-eye view image;
[0213] The mapped position is determined as the target placement position.
[0214] Optionally, the combining module 1904 is configured to:
[0215] Scaling the three-dimensional scene object so that the size of the scaled three-dimensional scene object meets the range of the target placement position;
[0216] The scaled three-dimensional scene object is placed at the target placement position on the ground of the three-dimensional scene to obtain the three-dimensional scene.
[0217] In summary, in the embodiment of the present application, a two-dimensional scene image is first generated based on the scene description text, and then a three-dimensional scene is obtained based on the two-dimensional scene image. The three-dimensional scene corresponding to the text description can be directly generated, thereby improving the convenience of generating the three-dimensional scene. For example, in a UGC game, the player can be directly guided to input a description text to generate a three-dimensional scene corresponding to the description text, which is conducive to increasing the playability of the UGC game. In the process of obtaining a three-dimensional scene based on a two-dimensional scene image, a three-dimensional scene object is first generated based on the scene object image in the two-dimensional scene image, and then a three-dimensional scene ground is generated based on the two-dimensional scene image with the scene object image removed, and finally the three-dimensional scene object is combined with the three-dimensional scene ground to obtain a three-dimensional scene. By generating a three-dimensional scene through the solution provided in the embodiment of the present application, it is only necessary to obtain the scene description text provided by the user, which simplifies the process for the user to complete the three-dimensional scene creation and has a wider applicability.
[0218] It should be noted that the apparatus provided in the above embodiments is merely exemplified by the division of the above functional modules. In actual applications, the above functions can be distributed among different functional modules as needed, that is, the internal structure of the apparatus can be divided into different functional modules to complete all or part of the functions described above. In addition, the apparatus and method embodiments provided in the above embodiments are based on the same concept. The implementation process is detailed in the method embodiments and will not be repeated here.
[0219] Please refer to Figure 20 , which shows a schematic diagram of the structure of a computer device provided by an exemplary embodiment of the present application. The computer device can be implemented as a terminal or server in the above-mentioned embodiment. Specifically, the computer device 2000 includes a central processing unit (CPU) 2001, a system memory 2004 including a random access memory 2002 and a read-only memory 2003, and a system bus 2005 connecting the system memory 2004 and the central processing unit 2001. The computer device 2000 also includes a basic input / output system (I / O system) 2006 that helps transmit information between various components within the computer, and a large-capacity storage device 2007 for storing an operating system 2013, application programs 2014, and other program modules 2015.
[0220] In some embodiments, the basic input / output system 2006 includes a display 2008 for displaying information and an input device 2009, such as a mouse or keyboard, for user input. Both the display 2008 and the input device 2009 are connected to the central processing unit 2001 via an input / output controller 2010 connected to the system bus 2005. The basic input / output system 2006 may also include an input / output controller 2010 for receiving and processing input from a variety of other devices, such as a keyboard, mouse, or electronic stylus. Similarly, the input / output controller 2010 also provides output to a display screen, printer, or other types of output devices.
[0221] The mass storage device 2007 is connected to the central processing unit 2001 via a mass storage controller (not shown) connected to the system bus 2005. The mass storage device 2007 and its associated computer-readable media provide non-volatile storage for the computer device 2000. In other words, the mass storage device 2007 may include a computer-readable medium (not shown) such as a hard disk or drive.
[0222] Without loss of generality, the computer-readable medium may include computer storage media and communication media. Computer storage media include volatile and non-volatile, removable and non-removable media implemented by any method or technology for storing information such as computer-readable instructions, data structures, program modules or other data. Computer storage media include random access memory (RAM), read-only memory (ROM), flash memory or other solid-state storage technologies, compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical storage, magnetic tape cassettes, magnetic tapes, disk storage or other magnetic storage devices. Of course, those skilled in the art will appreciate that the computer storage media are not limited to the above-mentioned ones. The above-mentioned system memory 2004 and mass storage device 2007 can be collectively referred to as memory.
[0223] The memory stores one or more programs, and the one or more programs are configured to be executed by one or more central processing units 2001. The one or more programs contain instructions for implementing the above-mentioned methods. The central processing unit 2001 executes the one or more programs to implement the methods provided by the above-mentioned various method embodiments.
[0224] According to various embodiments of the present application, the computer device 2000 can also be connected to a remote computer on a network such as the Internet for operation. That is, the computer device 2000 can be connected to the network 2011 via the network interface unit 2012 connected to the system bus 2005, or the network interface unit 2012 can be used to connect to other types of networks or remote computer systems (not shown).
[0225] The memory also includes one or more programs, which are stored in the memory and include steps executed by a computer device in the method provided in the embodiment of the present application.
[0226] An embodiment of the present application also provides a computer-readable storage medium, which stores at least one instruction, at least one program, code set or instruction set, and the at least one instruction, at least one program, code set or instruction set is loaded and executed by a processor to implement the three-dimensional model generation described in any of the above embodiments.
[0227] The present application provides a computer program product or computer program, which includes computer instructions stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium and executes the computer instructions, causing the computer device to perform the three-dimensional model generation provided in the above aspects.
[0228] Those skilled in the art will appreciate that all or part of the steps in the various methods of the above embodiments can be completed by instructing the relevant hardware through a program. The program can be stored in a computer-readable storage medium, which can be the computer-readable storage medium included in the memory in the above embodiments; or it can be a separate computer-readable storage medium not incorporated into the terminal. The computer-readable storage medium stores at least one instruction, at least one program, code set, or instruction set. The at least one instruction, at least one program, code set, or instruction set is loaded and executed by a processor to implement the three-dimensional model generation described in any of the above method embodiments.
[0229] Optionally, the computer-readable storage medium may include: ROM, RAM, solid-state drives (SSDs), or optical disks. Among them, RAM may include resistance random access memory (ReRAM) and dynamic random access memory (DRAM). The serial numbers of the above embodiments of the present application are for description only and do not represent the advantages or disadvantages of the embodiments.
[0230] Those skilled in the art will understand that all or part of the steps to implement the above embodiments may be accomplished by hardware, or by a program to instruct the relevant hardware, and the program may be stored in a computer-readable storage medium, which may be a read-only memory, a disk, or an optical disk, etc.
[0231] It should be noted that the information (including but not limited to user device information, user personal information, etc.), data (including but not limited to data used for analysis, stored data, displayed data, etc.) and signals involved in this application are all authorized by the user or fully authorized by all parties, and the collection, use and processing of relevant data must comply with the relevant laws, regulations and standards of relevant countries and regions.
[0232] Before collecting the user's relevant data and during the process of collecting the user's relevant data, this application can display a prompt interface, pop-up window or output voice prompt information. The prompt interface, pop-up window or voice prompt information is used to remind the user that his or her relevant data is currently being collected, so that this application only starts to execute the relevant steps of obtaining the user's relevant data after obtaining the user's confirmation operation on the prompt interface or pop-up window. Otherwise (that is, when the user's confirmation operation on the prompt interface or pop-up window is not obtained), the relevant steps of obtaining the user's relevant data are terminated, that is, the user's relevant data is not obtained.
[0233] It should be understood that the "multiple" mentioned in this article refers to two or more. "And / or" describes the association relationship of associated objects, indicating that three relationships may exist. For example, A and / or B can represent three situations: A exists alone, A and B exist at the same time, and B exists alone. The character " / " generally indicates that the previous and next associated objects are in an "or" relationship. And the "first", "second", etc. mentioned in this article are used to distinguish similar objects, and are not used to limit a specific order or sequence. In addition, the step numbers described in this article only exemplify a possible execution sequence between the steps. In some other embodiments, the above steps may not be executed in the order of the numbers, such as two steps with different numbers are executed at the same time, or two steps with different numbers are executed in the opposite order to the diagram. The embodiments of the present application do not limit this.
[0234] The above description is merely an optional embodiment of the present application and is not intended to limit the present application. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of the present application shall be included in the scope of protection of the present application.
Claims
1. A three-dimensional scene generation method, characterized in that: The method comprises: generating a two-dimensional scene image based on the scene description text, wherein image content of the two-dimensional scene image conforms to the scene description text, and the image content includes a scene ground and scene objects located on the scene ground; generating a three-dimensional scene object that conforms to the scene object image based on the scene object image of the scene object in the two-dimensional scene image; generating a three-dimensional scene ground based on the two-dimensional scene image after removing the scene object image, wherein the terrain of the three-dimensional scene ground is consistent with the terrain of the scene ground in the two-dimensional scene image; The three-dimensional scene objects and the three-dimensional scene ground are combined to obtain a three-dimensional scene.
2. The method according to claim 1, characterized in that The method further comprises: The two-dimensional scene image is segmented to obtain the scene object image and the mask scene image, wherein the mask scene image is the two-dimensional scene image after removing the scene object image, and the mask scene image is obtained by masking the two-dimensional scene image with an object image mask.
3. The method according to claim 2, characterized in that Generating a three-dimensional scene ground based on the two-dimensional scene image after removing the scene object image includes: Performing ground completion on the masked scene image to obtain a completed ground image; The three-dimensional scene ground is generated based on the completed ground image, and the terrain of the three-dimensional scene ground is consistent with the terrain of the completed ground image.
4. The method according to claim 3, characterized in that The two-dimensional scene image and the mask scene image are isometric view images; The performing ground completion on the mask scene image to obtain a completed ground image includes: Converting the mask scene image into a bird's-eye view mask scene image; The ground completion network is used to perform ground completion on the bird's-eye view mask scene image to obtain the completed ground image.
5. The method according to claim 3, characterized in that The method further comprises: Masking the sample scene ground image from a bird's-eye view to obtain a sample masked scene image; Inputting the sample mask scene image into the ground completion network to obtain a sample completed ground image output by the ground completion network; determining a completion loss based on a difference between the sample completed ground image and the sample scene ground image; The ground completion network is trained based on the completion loss to obtain the trained ground completion network.
6. The method according to claim 3, characterized in that Generating a three-dimensional scene ground based on the completed ground image includes: generating a three-dimensional ground surface based on a size of the completed ground surface image, wherein a topography of the three-dimensional ground surface is consistent with a topography of the completed ground surface image; Mapping is performed on the three-dimensional ground based on the completed ground image to obtain the three-dimensional scene ground.
7. The method according to claim 3, characterized in that The method further comprises: performing image quality optimization on the completed ground image through an image optimization network to obtain an optimized completed ground image, wherein the optimized completed ground image has a higher level of detail than the unoptimized completed ground image; The generating the three-dimensional scene ground based on the completed ground image includes: The three-dimensional scene ground is generated based on the optimized completed ground image.
8. The method according to claim 1, characterized in that The generating, based on the scene object image of the scene object in the two-dimensional scene image, a three-dimensional scene object that conforms to the scene object image, comprises: A three-dimensional object search is performed in a three-dimensional asset library based on the scene object image to obtain a three-dimensional scene object corresponding to the scene object image.
9. The method according to claim 8, characterized in that The performing a three-dimensional object search in a three-dimensional asset library based on the scene object image to obtain a three-dimensional scene object corresponding to the scene object image includes: Encoding the scene object image to obtain scene object image features; Matching the scene object image features with the angle image features in the 3D asset library to obtain a matching result, where the 3D asset library contains angle images of at least two candidate 3D objects at different angles and angle image features corresponding to each angle image, and the matching result is used to represent the similarity between the scene object image and each angle image; Based on the matching result, the three-dimensional scene object is determined from the candidate three-dimensional objects.
10. The method according to claim 9, characterized in that Determining the three-dimensional scene object from the candidate three-dimensional objects based on the matching result includes: Based on the matching result, determining the angle image having a similarity with the scene object image higher than a similarity threshold as a target angle image; Voting on the candidate three-dimensional objects based on the correspondence between the target angle image and the candidate three-dimensional objects to obtain a voting result; The candidate three-dimensional object that has obtained the most votes as indicated by the voting result is determined as the three-dimensional scene object.
11. The method according to claim 1, wherein The combining of the three-dimensional scene objects and the three-dimensional scene ground to obtain a three-dimensional scene includes: Determining a target placement position of the three-dimensional scene object on the ground of the three-dimensional scene based on the two-dimensional scene image; The three-dimensional scene object is placed at the target placement position on the three-dimensional scene ground to obtain the three-dimensional scene.
12. The method according to claim 11, characterized in that The determining, based on the two-dimensional scene image, a target placement position of the three-dimensional scene object on the three-dimensional scene ground includes: Determining a center point of a position of the scene object image in the two-dimensional scene image and an image range of the scene object image; The mapping position of the center point and the image range in the completed ground image is determined as the target placement position, the center point of the target placement position corresponds to the center point of the position of the scene object image, the range of the target placement position corresponds to the image range, and the completed ground image is obtained by ground completion based on the two-dimensional scene image after removing the scene object image.
13. The method according to claim 12, characterized in that The two-dimensional scene image is an isometric view image; Determining the mapping position of the center point and the image range in the completed ground image as the target placement position includes: Performing homography transformation mapping on the center point and the image range to determine mapping positions of the center point and the image range in the completed ground image, where the completed ground image is a bird's-eye view image; The mapped position is determined as the target placement position.
14. The method according to claim 11, characterized in that Placing the three-dimensional scene object at the target placement position on the ground of the three-dimensional scene to obtain the three-dimensional scene includes: Scaling the three-dimensional scene object so that the size of the scaled three-dimensional scene object meets the range of the target placement position; The scaled three-dimensional scene object is placed at the target placement position on the ground of the three-dimensional scene to obtain the three-dimensional scene.
15. A three-dimensional scene generation device, characterized in that: The device comprises: A first generating module is configured to generate a two-dimensional scene image based on the scene description text, wherein image content of the two-dimensional scene image conforms to the scene description text, and the image content includes a scene ground and scene objects located on the scene ground; a second generating module, configured to generate a three-dimensional scene object conforming to the scene object image based on the scene object image of the scene object in the two-dimensional scene image; a third generating module, configured to generate a three-dimensional scene ground based on the two-dimensional scene image after removing the scene object image, wherein the terrain of the three-dimensional scene ground is consistent with the terrain of the scene ground in the two-dimensional scene image; The combination module is used to combine the three-dimensional scene objects and the three-dimensional scene ground to obtain a three-dimensional scene.
16. A computer device, characterized in that: The computer device includes a processor and a memory, wherein the memory stores at least one program, and the at least one program is loaded and executed by the processor to implement the three-dimensional model generation method according to any one of claims 1 to 14.
17. A computer-readable storage medium, characterized in that The readable storage medium stores at least one program, and the at least one program is loaded and executed by the processor to implement the three-dimensional model generation method according to any one of claims 1 to 14.
18. A computer program product, characterized in that The computer program product includes computer instructions, which are stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions to implement the three-dimensional model generation method according to any one of claims 1 to 14.