Method, apparatus, device, and storage medium for generating a three-dimensional scene based on a large-scale model
By using a large language model to obtain a target asset set from description information, the method automates the generation of 3D scenes, addressing inefficiencies and inaccuracies in existing manual design processes.
Patent Information
- Application Number
- JP2024100030
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2023-10-16
- Filing Date
- 2024-06-20
- Publication Date
- 2025-06-18
- Estimated Expiration
- 2044-06-20
AI Technical Summary
Existing methods for creating 3D scenes are inefficient and lack accuracy, as they require manual design by artists using rendering engines.
A method that processes description information to extract label information, generates query prompts for a large language model (LLM) to obtain a target asset set including assets, material information, and scene attributes, and uses this set to generate a 3D scene.
This approach simplifies and enhances the efficiency and accuracy of 3D scene generation by automating the process and providing comprehensive asset and attribute information.
Smart Images

Figure 0007694861000001 
Figure 0007694861000002 
Figure 0007694861000003
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of artificial intelligence technology, specifically to technical fields such as 3D modeling and large models, and particularly relates to a method, apparatus, device, and storage medium for generating a 3D scene based on a large model.
Background Art
[0002] With the development of artificial intelligence technology, 3D (three-dimensional) scene creation has been widely applied in fields such as the game industry, the animation industry, and the digital human industry.
[0003] In related technologies, 3D scenes are usually artificially created by designers using a rendering engine.
Summary of the Invention
Problems to be Solved by the Invention
[0004] The present disclosure provides a method, apparatus, device, and storage medium for generating a 3D scene based on a large model.
Means for Solving the Problems
[0005] According to one aspect of the present disclosure, a method for generating a 3D scene based on a large model is provided, including the steps of processing description information of a target 3D scene to obtain label information in the description information; generating query operation prompt information for a large language model (LLM) based on the label information, and using the LLM to obtain a target asset set that matches the label information based on the query operation prompt information, where the target asset set includes target assets in the target 3D scene, target material information of the target assets, and target scene attribute information of the target assets; and generating the target 3D scene based on the target asset set.
[0006] According to another aspect of the present disclosure, there is provided a device for generating a 3D scene based on a large-scale model, including a processing module for processing description information of a target 3D scene to obtain label information in the description information, and an acquisition module for generating query operation prompt information of a large language model (LLM) based on the label information and using the LLM to obtain a target asset set that matches the label information based on the query operation prompt information, where the target asset set includes target assets in the target 3D scene, target material information of the target assets, and target scene attribute information of the target assets, and a generation module for generating the target 3D scene based on the target asset set.
[0007] According to another aspect of the present disclosure, there is provided an electronic device including at least one processor and a memory communicatively connected to the at least one processor, where instructions executable by the at least one processor are stored in the memory, and when the instructions are executed by the at least one processor, the at least one processor can execute the method of this aspect and any possible implementation manner.
[0008] According to another aspect of the present disclosure, there is provided a non-transitory computer-readable storage medium storing computer instructions, where the computer instructions cause the computer to execute the method of this aspect and any possible implementation manner.
[0009] According to another aspect of the present disclosure, there is provided a computer program product including a computer program, where when the computer program is executed by a processor, the method of this aspect and any possible implementation manner are realized.
[0010] According to the technical solution of the present disclosure, a 3D scene can be generated simply and efficiently.
[0011] The content described in this specification is not intended to identify the key or important features of the embodiments of the present disclosure, nor is it used to limit the scope of the present disclosure. Other features of the present disclosure can be easily understood throughout the following specification.
Brief Description of the Drawings
[0012] The drawings are for better understanding of the present application and do not limit the present application.
Figure 1
Figure 2
Figure 3
Figure 4
Figure 5
Figure 6
Modes for Carrying Out the Invention
[0013] Hereinafter, exemplary embodiments of the present application will be described with reference to the drawings. For ease of understanding, various details of the embodiments of the present application are included, and they should be regarded as merely illustrative. Therefore, those skilled in the art should recognize that various changes and modifications can be made to the embodiments described in this specification without departing from the scope and spirit of the present application. Similarly, for the sake of brevity, descriptions of well-known functions and structures are omitted in the following description.
[0014] In the related art, since the 3D scene is artificially created, there are problems in all aspects such as efficiency and accuracy.
[0015] To improve the accuracy and efficiency of 3D scene generation, the present disclosure provides the following examples.
[0016] FIG. 1 is a schematic diagram according to the first embodiment of the present disclosure. This embodiment provides a method for generating a 3D scene based on a large model, and the method includes the following steps: 101. Process the description information of the target 3D scene to obtain the label information in the description information.
[0017] 102. Generate query operation prompt information for the LLM based on the label information, and use the LLM to obtain a target asset set that matches the label information based on the query operation prompt information. The target asset set includes target assets in the target 3D scene, target material information of the target assets, and target scene attribute information of the target assets.
[0018] 103. Generate the target 3D scene based on the target asset set.
[0019] Among them, the description information refers to the description information of the target 3D scene input by the user. For example, the description information is "I want to generate a scene of playing in a forest in northern China in winter. And there is a dirt road and a house made of logs in the scene. White smoke is coming out of the chimney on the roof of the house, and there is a warm yellow light on in the window of the house. It has a cold color tone and is realistic."
[0020] The label information is the key information for the description information. For example, based on the above example, the label information includes one scene, winter season, northern China, forest, for games, one dirt road, a house made of logs, the chimney on the roof of the house, white smoke coming out of the chimney, there is a warm yellow light in the window of the house, having a cold color tone, and being realistic.
[0021] After obtaining the label information, a target asset set that matches the label information can be obtained using a Large Language Model (LLM). The input to the LLM includes a prompt, and the LLM generates corresponding output information based on the prompt.
[0022] To obtain the target asset set, the prompt of the LLM can be called the query operation prompt, and the LLM queries and obtains the target asset set that matches the label information based on the query operation prompt. Among them, the query operation prompt can include the above-mentioned label information, and can also include instruction information. The instruction information is used to instruct the LLM to execute the query operation, and the specific content of the instruction information can be set. After receiving the query operation prompt, the LLM can obtain the target asset set based on the prompt.
[0023] The 3D scene is composed of assets, and the assets include, for example, models, texture balls, maps, particles, etc. The models include, for example, models of houses, flowers, grass, etc.
[0024] The assets that match the target 3D scene are called target assets. For example, when the label information includes a house, the house model is obtained as the target asset.
[0025] The target scene attribute information refers to the relevant information within the target 3D scene of the target asset, including, for example, size, position, rotation angle, etc.
[0026] The target texture information refers to the texture-related information of the target asset, such as color, roughness, etc.
[0027] After obtaining the target asset set, a target 3D scene is generated based on the available target asset set. For example, a target asset corresponding to the size is generated based on the size information in the scene attribute information, and the generated target asset is placed at the specified position in the scene attribute information.
[0028] In related technologies, when only the target asset is obtained, the scene attribute information of the target asset is undetermined. Therefore, for example, the position of the model is randomly determined, and then the user needs to further adjust the target asset, increasing the complexity.
[0029] In this embodiment, the target asset set not only includes the target asset, but also further includes the scene attribute information and material information of the target asset, which is equivalent to obtaining the target asset and its material, as well as scene attribute information such as position and size at one time. Compared with the method of only obtaining the target asset, a 3D scene can be generated simply and efficiently. In addition, by generating the query operation prompt information of the LLM, the LLM can be used to obtain the target asset set, improving the processing efficiency and accuracy.
[0030] To better understand the embodiments of the present disclosure, the application scenarios to which the embodiments of the present disclosure can be applied will be described.
[0031] Figure 2 is a schematic diagram of an application scenario provided by an embodiment of the present disclosure. In this scenario, it includes a user terminal 201 and a server 202. The user terminal 201 includes a personal computer (PC), a mobile device (such as a mobile phone), a tablet, a notebook computer, a smart wearable device, etc. The server 202 may be a local server or a cloud server, etc., and the server may be a single server or a server cluster.
[0032] As shown in FIG. 3, a three-dimensional scene generation system based on a large-scale model includes a rendering engine 301, a plugin 302, a large language model (LLM) 303, and an asset repository. The rendering engine 301 is set on the user terminal, and the rendering engine is, for example, Unity software, Unreal Engine (UE) software, etc. Both Unity software and UE software (specifically, it may be UE4 or UE5) are real-time rendering engines that can process dynamic images in real time and perform three-dimensional rendering. The plugin 302 is pre-installed in the rendering engine. The LLM 303 is set on the server. The rendering engine 301 can call the LLM 303 through a preset interface to realize communication between the rendering engine 301 and the LLM 303. The asset repository may be provided by the plugin, or in order to generate a three-dimensional scene of the user's personal style, the asset repository may be a locally customized asset repository by the user. The asset repository in FIG. 3 is displayed as the local asset repository 304.
[0033] LLM has been a hot topic in the field of artificial intelligence in recent years. By performing pre-training on a large amount of text data, LLM can learn rich language knowledge and world knowledge, thereby obtaining amazing effects in various natural language processing (NLP) tasks. It is a pre-trained language model. Wenxin Yiyan, ChatGPT, etc. are all applications developed based on LLM, which can generate fluent, logical, and creative text content and can even interact with humans naturally. Specifically, the large-scale model may also be a general pre-trained (Generative Pre-trained Transformer, GPT) model based on Transformer, an enhanced representation (Enhanced Representation through Knowledge Integration, ERNIE) model based on knowledge integration, etc.
[0034] The asset repository includes assets, scene attribute information of assets, and material information of assets. Each piece of information corresponds to a folder that can be called an asset folder, a scene attribute folder, and a material folder. Each folder may be in a hierarchical format. For example, for the asset folder, the first level is the model, the second level includes vegetation, rocks, buildings, props, etc., and the third level corresponding to vegetation includes trees, flowers, etc. For the scene attribute folder, the first level includes light, weather, screen style, asset position, etc., the second level corresponding to light includes light intensity, light color, light position, etc., and the second level corresponding to weather includes sunny, rainy, snowy, etc. For the material folder, material-adjustable parameters such as color, material attributes, and snow / lichen height can be recorded.
[0035] Taking the local asset repository as an example, the above-mentioned various types of information (assets, material information, scene attribute information) are created by the user in advance. Regarding the scene attribute information, it can be created using an asset set tool, which can be an asset management tool or a level animation sequence tool. The asset management tool is, for example, a Variant Manager tool, and the level animation sequence tool is, for example, a LevelSequence tool. Both the Variant Manager tool and the LevelSequence tool are asset set tools within UE. The Variant Manager tool sets combinations such as asset positions, material balls, and various attributes of the scene in a set file, and the LevelSequence tool can achieve dynamic switching of assets / scene attributes in a way of setting keyframes.
[0036] Also, the asset-related information in the local asset repository can be recorded in the asset repository list. The information recorded in the asset repository list can be called candidate information.
[0037] The rendering engine provides an interactive interface for the user. The interactive interface includes an input box, and the user can input the description information of the target 3D scene in the input box. For example, "I want to generate a scene of playing in a forest in northern China in winter. And there is a house composed of dirt roads and piles in the scene, white smoke is coming out of the chimney on the roof of the house, warm yellow lights are on in the windows of the house, with a cold color tone and being realistic."
[0038] The rendering engine calls the LLM via a pre-set interface, uses the LLM to perform extraction processing on the above-described description information, and obtains label information. For example, based on the above description information, the extracted label information includes one scene, winter season, northern China, forest, for games, one dirt road, a house made of piles, a chimney on the roof of the house, white smoke coming out of the chimney, warm yellow light inside the windows of the house, with a cold color tone and being realistic.
[0039] In this embodiment, when using the LLM to extract label information in the description information, more concise and accurate label information can be obtained based on the description information, the accuracy of the target asset set obtained based on the label information can be further improved, and the generation accuracy and efficiency of the 3D scene can be improved.
[0040] After obtaining the LLM label information, match the label information with the candidate information (candidate asset information, candidate material information, and candidate scene attribute information) in the asset repository list to obtain the target information (target asset information, target material information, and target scene attribute information) that matches the label information. Then, send the target information to the plugin.
[0041] The plugin obtains a target asset set (target assets and their corresponding target material information, and target scene attribute information) based on the target information, and generates a target 3D scene based on the target asset set.
[0042] After the rendering engine renders the target 3D scene generated by the plugin, it displays it to the user.
[0043] Combining the above application scenarios, the present disclosure further provides the following embodiments.
[0044] Figure 4 is a schematic diagram according to the second embodiment of the present disclosure. This embodiment provides a method for generating a three-dimensional scene based on a large-scale model, and the method includes the following steps: 401. Generate extraction operation prompt information for the LLM based on the description information of the target three-dimensional scene, and use the LLM to process the description information based on the extraction operation prompt information to obtain label information in the description information.
[0045] For example, the user inputs description information, such as "I want to generate a scene of playing in a forest in northern China in the winter season. And there is a dirt road and a house made of logs in the scene, white smoke is coming out of the chimney on the roof of the house, warm yellow lights are on in the windows of the house, with a cold color tone and being realistic", through an interface provided by the rendering engine. The rendering engine calls the LLM through a preset interface, and uses the LLM to extract label information in the description information. For example, the label information includes one scene, winter season, northern China, forest, for games, one dirt road, a house made of logs, the chimney on the roof of the house, white smoke is coming out of the chimney, there is a warm yellow light in the window of the house, with a cold color tone and being realistic.
[0046] Specifically, the rendering engine can send prompt information to the LLM, and the prompt information can be called extraction operation prompt information. The extraction operation prompt information includes description information and further includes instruction information. The instruction information is used to instruct the LLM to perform an information extraction operation, and the specific content of the instruction information can be set. The LLM can extract label information in the description information based on the extraction operation prompt information.
[0047] After the LLM obtains the label information, the label information can further be displayed to the user.
[0048] Furthermore, the LLM can further determine the importance of each label information and display each label information in order based on the importance.
[0049] In this embodiment, when using an LLM to extract label information in the description information, more concise and accurate label information can be obtained based on the description information, the accuracy of the target asset set obtained based on the label information can be further improved, and the generation accuracy and efficiency of the 3D scene can be improved.
[0050] 402. Generate query operation prompt information for the LLM based on the label information, and use the LLM to match the label information with a plurality of pre-recorded candidate information based on the query operation prompt information to obtain target information of the target asset set.
[0051] The plurality of candidate information can be divided into three categories. The first category is asset information, such as vegetation - north - trees and ground - forest - land, architecture - house - trees, particle - smoke - chimney, etc.; the second category is scene attribute information, such as season - winter season, layout - forest, color tone - cold color tone, style - realistic, light - window; the third category is material information, such as snow height, color hue, normal intensity, roughness intensity, etc. Taking the display of the above asset information and scene attribute information in a hierarchical form as an example, for example, for vegetation, the first layer is "vegetation", the second layer is north, the third layer is "trees and ground", and so on until the last layer is recorded. For example, the above "vegetation - north - trees and ground - forest - land"; similarly, assuming that for the "vegetation" in the first layer, the vegetation information with the second layer being "south" is further included, the relevant asset information can be displayed as vegetation - south -....
[0052] The plurality of candidate information can be recorded in the asset repository list, and the LLM can obtain target information that matches the label information based on the pre - learned correspondence. For example, "winter season" in the label information matches information such as "winter", "snow", "ice", "cold" in the asset repository.
[0053] In this embodiment, by matching the label information with the pre-recorded candidate information, target information can be obtained, and the target information can be obtained simply and efficiently, improving the processing efficiency.
[0054] 403. Use a pre-set plugin to obtain a target asset set within the locally customized local asset repository based on the target information.
[0055] After obtaining the target information, the LLM can send the target information to a plugin pre-installed in the rendering engine, and the plugin can obtain a target asset set within the local asset repository based on the target information.
[0056] In this embodiment, a target asset set is obtained within the locally customized local asset repository by the user, and the user can set candidate assets and their related information according to their own needs, and a personalized target 3D scene can be generated.
[0057] 404. Use a pre-set plugin to generate a target 3D scene based on the target asset set.
[0058] Among them, the target asset set includes target assets, corresponding target material information, and target scene attribute information. Therefore, the plugin can directly generate a target 3D scene based on this information, eliminating the need for the user to set attributes such as asset positions and simplifying user operations.
[0059] Furthermore, the plugin can further generate an initial 3D scene based on the target asset set and adjust the initial 3D scene based on the scene function information in the label information to generate the target 3D scene.
[0060] For example, the above label information includes "for games", and this information is scene function information. The plugin can preset operations corresponding to each function, and then can adjust the initial 3D scene based on the preset operations. For example, based on the preset adjustment rules for games, adjustments can be made to aspects such as the overall level of detail (LOD), lighting, resolution, and rendering quality of the initial 3D scene to obtain the target 3D scene.
[0061] In this embodiment, based on the scene function information, the initial 3D scene can be adjusted to obtain the target 3D scene, improving the accuracy of the target 3D scene and the rendering effect.
[0062] 405. Use the rendering engine to render the target 3D scene.
[0063] After generating the target 3D scene using the plugin, the rendering engine renders the target 3D scene and displays the rendering result to the user.
[0064] Furthermore, after the user obtains the rendering result, the target 3D scene can be modified as needed. It can be modified based on the label information, or directly modify the assets or scene attribute information (such as position) within the target 3D scene.
[0065] Specifically, display the label information to the user, obtain the label information modified by the user, obtain the modified target asset set based on the modified label information, and generate the modified target 3D scene based on the modified target asset set.
[0066] For example, the above label information includes "winter", and the user can modify the label information to "summer", and then the system can regenerate the scene related to summer.
[0067] In this embodiment, the target 3D scene is modified based on the modified label information, so that the user does not need to perform specific modification operations on the scene, which simplifies the user operation and can improve the processing efficiency.
[0068] In addition, the target 3D scene can be displayed to the user, and a modified target 3D scene can be generated based on the modification instruction of the user.
[0069] Among them, the modification instruction is, for example, to modify the model in the target 3D scene or to modify the position of the model, etc., and a modified target 3D scene can be obtained based on the modification instruction. For example, there is a circular table in the target 3D scene, and the user can directly modify the circular table into a square table in the scene.
[0070] In this embodiment, when the target 3D scene is modified based on the modification instruction, a personalized modified target 3D scene required by the user can be obtained.
[0071] FIG. 5 is a schematic diagram according to the third embodiment of the present disclosure. This embodiment provides a 3D scene generation device based on a large-scale model. As shown in FIG. 5, the device 500 includes a processing module 501, an acquisition module 502, and a generation module 503.
[0072] The processing module 501 is used to process the description information of the target 3D scene to obtain the label information in the description information. The acquisition module 502 is used to generate the query operation prompt information of the LLM based on the label information, and use the LLM to obtain a target asset set that matches the label information based on the query operation prompt information. The target asset set includes the target assets in the target 3D scene, the target material information of the target assets, and the target scene attribute information of the target assets. The generation module 503 is used to generate the target 3D scene based on the target asset set.
[0073] In this embodiment, the target asset set not only includes the target assets, but also further includes the scene attribute information and material information of the target assets, which is equivalent to obtaining the target assets and their material and scene attribute information such as position and size at one time. Compared with the method of only obtaining the target assets, the 3D scene can be generated simply and efficiently.
[0074] In some embodiments, the processing module 501 is further used to generate the extraction operation prompt information of the LLM based on the description information, and use the LLM to process the description information based on the extraction operation prompt information to obtain the label information.
[0075] In this embodiment, when using the LLM to extract the label information in the description information, more concise and accurate label information can be obtained based on the description information, further improving the accuracy of the target asset set obtained based on the label information, and improving the generation accuracy and efficiency of the 3D scene.
[0076] In some embodiments, the acquisition module 502 is further configured to use the LLM to match the label information with a plurality of pre-recorded candidate information based on the query operation prompt information, so as to obtain the target information of the target asset set, and is used to obtain the target asset set based on the target information.
[0077] In this embodiment, by matching the label information with the pre-recorded candidate information to obtain the target information, the target information can be obtained simply and efficiently, and the processing efficiency can be improved.
[0078] In some embodiments, the acquisition module 502 is further configured to obtain the target asset set in the locally defined local asset repository by the user based on the target information.
[0079] In this embodiment, by obtaining the target asset set in the locally defined local asset repository by the user, the user can set candidate assets and their related information according to his own needs, and a personalized target 3D scene can be generated.
[0080] In some embodiments, the generation module 503 is further configured to generate an initial 3D scene based on the target asset set, and adjust the initial 3D scene based on the scene function information in the label information to generate the target 3D scene.
[0081] In this embodiment, by adjusting the initial 3D scene based on the scene function information to obtain the target 3D scene, the accuracy of the target 3D scene can be improved, and the rendering effect can be improved.
[0082] In some embodiments, the apparatus 500 Further include a first correction module for displaying the label information to the user, obtaining the label information corrected by the user, obtaining a corrected target asset set based on the corrected label information, and generating a corrected target 3D scene based on the corrected target asset set.
[0083] In this embodiment, the target 3D scene is corrected based on the corrected label information, so that the user does not need to perform specific correction operations on the scene, which simplifies the user operation and improves the processing efficiency.
[0084] In some embodiments, the apparatus 500 Further include a second correction module for displaying the target 3D scene to the user and generating a corrected target 3D scene based on the correction instruction of the user.
[0085] In this embodiment, when the target 3D scene is corrected based on the correction instruction, a personalized corrected target 3D scene required by the user can be obtained.
[0086] It should be understood that in the embodiments of the present disclosure, the same or similar content in different embodiments can be referred to each other.
[0087] It can be understood that the "first", "second", etc. in the embodiments of the present disclosure are only for distinction and do not represent the importance level, the order of timing, etc.
[0088] In the technical solution of the present disclosure, all processes such as the collection, storage, use, processing, transfer, provision, and disclosure of relevant user personal information comply with relevant laws and regulations and do not violate public order and good customs.
[0089] According to the embodiments of the present disclosure, the present disclosure further provides an electronic device, a readable storage medium, and a computer program product.
[0090] As shown in FIG. 6, it is a block diagram of an electronic device 600 for implementing an embodiment of the present disclosure. The electronic device is intended to represent various forms of digital computers such as laptop computers, desktop computers, workstations, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices such as personal digital assistants, mobile phones, smartphones, wearable devices, and other similar computing devices. The components shown in this specification, their connections and relationships, and their functions are merely examples and are not intended to limit the description of this specification and / or the implementation of the present disclosure required.
[0091] As shown in FIG. 6, the device 600 includes a computing unit 601, and the computing unit 601 can execute various appropriate operations and processes based on a computer program stored in a read-only memory (ROM) 602 or a computer program loaded from a storage unit 608 into a random access memory (RAM) 603. The RAM 603 can also store various programs and data necessary for the device 600 to operate. The computing unit 601, the ROM 602, and the RAM 603 are connected to each other via a bus 604. An input / output (I / O) interface 605 is also connected to the bus 604.
[0092] A plurality of components within the electronic device 600 are connected to the I / O interface 605, and include an input unit 606 such as a keyboard and a mouse, an output unit 607 such as various types of displays and speakers, a storage unit 608 such as a disk and an optical disk, and a communication unit 609 such as a network card, a modem, and a wireless communication transceiver. The communication unit 609 enables the device 600 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks.
[0093] The computing unit 601 is a general-purpose and / or dedicated processing component with various processing and computing capabilities. Some examples of the computing unit 601 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, computing units that execute various machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The computing unit 601 executes the various methods and processes described above, such as the method for generating a 3D scene based on a large-scale model. For example, in some embodiments, the method for generating a 3D scene based on a large-scale model can be realized as a computer software program tangibly included in a machine-readable medium such as the storage unit 608. In some embodiments, part or all of the computer program is loaded and / or installed on the device 600 via the ROM 602 and / or the communication unit 609. When the computer program is loaded into the RAM 603 and executed by the computing unit 601, one or more steps of the method for generating a 3D scene based on a large-scale model can be executed. Alternatively, in other embodiments, the computing unit 601 can be configured to execute the method for generating a 3D scene based on a large-scale model via any other suitable means (e.g., by firmware).
[0094] The various embodiments of the systems and techniques described in this specification can be implemented in digital electronic circuit systems, integrated circuit systems, field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), system on chip systems (SOCs), loading programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include being implemented by one or more computer programs, which can be executed and / or interpreted in a programmable system including at least one programmable processor, and the programmable processor can be an application specific or general purpose programmable processor, receiving data and instructions from a storage system, at least one input device, and at least one output device, and transmitting the data and instructions to the storage system, the at least one input device, and the at least one output device.
[0095] The program code for implementing the methods of the present disclosure can be created using any combination of one or more programming languages. These program codes can be provided to the processor or controller of a general purpose computer, a dedicated computer, or other programmable data processing device so that when the program code is executed by the processor or controller, the functions / operations defined in the flowchart and / or block diagram are implemented. The program code can be executed entirely on a machine, partially on a machine, partially on a machine as an independent software package, partially remotely on a machine, or entirely remotely on a machine or server.
[0096] In the context of the present disclosure, a machine-readable medium may be a tangible medium that includes or can store a program for use in or in combination with an instruction execution system, apparatus, or device. The machine-readable medium may be a machine-readable signal medium or a machine-readable storage medium. The machine-readable medium includes, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of the machine-readable storage medium include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fibers, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.
[0097] To provide for interaction with a user, the systems and techniques described herein can be implemented on a computer, which includes a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user, and a keyboard and a pointing device (e.g., a mouse or a trackball), by which the user can provide input to the computer. Other kinds of devices can also be used to provide for interaction with a user, for example, feedback provided to the user can be any form of sensing feedback (e.g., visual feedback, auditory feedback, or tactile feedback), and input from the user can be received in any form (including acoustic input, voice input, and tactile input).
[0098] The systems and techniques described herein can be implemented in a computing system that includes backend components (such as a data server), or a computing system that includes middleware components (such as an application server), or a computing system that includes frontend components (such as a user computer having a graphical user interface or a web browser, and the user interacts with implementations of the systems and techniques described herein via the graphical user interface or the web browser), or a computing system that includes any combination of such backend components, middleware components, and frontend components. The components of the system can be interconnected with each other by digital data communication in any form or medium (such as a communication network). Examples of communication networks include local area networks (LANs), wide area networks (WANs), the Internet, and blockchain networks.
[0099] A computer system can include a client and a server. The client and the server are generally remote from each other and typically interact via a communication network. The relationship between the client and the server is created by computer programs that are executed on corresponding computers and have a client-server relationship with each other. The server can be a cloud server, also called cloud computing or cloud hosting, which is one of the host products in a cloud computing service system and solves the problems of high management difficulty and weak business scalability existing in traditional physical hosts and VPS services (abbreviated as "Virtual Private Server" or "VPS"). The server can also be a server of a distributed system or a server that combines blockchain.
[0100] It should be understood that the steps can be rearranged, added, or deleted using the various forms of flow shown above. For example, each step described in this disclosure may be executed in parallel, sequentially, or in a different order, but is not limited herein as long as the technical solutions disclosed in this disclosure can achieve the desired results.
[0101] The above specific embodiments do not constitute a limitation on the protection scope of this disclosure. Those skilled in the art can make various modifications, combinations, sub - combinations, and substitutions based on design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of any of this disclosure must be included within the protection scope of this disclosure.
Claims
1. A method for generating a three-dimensional scene based on a large scale model, comprising the steps of: processing a description of a target 3D scene to obtain label information in the description; generating query operation presentation information of an LLM according to the label information, and obtaining a target asset set matching the label information according to the query operation presentation information using the LLM, the target asset set including target assets in the target 3D scene, target material information of the target assets, and target scene attribute information of the target assets; generating the target three-dimensional scene based on the target asset set. A method for generating 3D scenes based on large-scale models.
2. The step of processing description information of the target 3D scene to obtain label information in the description information includes: generating extraction operation presentation information for the LLM based on the description information; and processing the description information based on the extraction operation presentation information using the LLM to obtain the label information. The method for generating a three-dimensional scene based on a large scale model according to claim 1.
3. The step of obtaining a target asset set matching the label information based on the query operation presentation information using the LLM includes: Using the LLM, based on the query operation presentation information, matching the label information with a plurality of pre-recorded candidate information to obtain target information of the target asset set; obtaining the target asset set based on the target information. The method for generating a three-dimensional scene based on a large scale model according to claim 1.
4. The step of obtaining the target asset set based on the target information includes: obtaining the target asset set in a user-defined local asset repository based on the target information; The method for generating a 3D scene based on a large scale model according to claim 3.
5. generating the target three-dimensional scene based on the target asset set, generating an initial three-dimensional scene based on the target asset set; and adjusting the initial 3D scene based on scene feature information in the label information to generate the target 3D scene. The method for generating a three-dimensional scene based on a large scale model according to claim 1.
6. displaying the label information to a user; obtaining label information modified by the user; obtaining a modified target asset set based on the modified label information; and generating a modified target 3D scene based on the modified target asset set. A method for generating a three-dimensional scene based on a large-scale model according to any one of claims 1 to 5.
7. displaying the target three-dimensional scene to a user; generating a modified target three-dimensional scene based on the user's modification instructions. A method for generating a three-dimensional scene based on a large-scale model according to any one of claims 1 to 5.
8. An apparatus for generating a three-dimensional scene based on a large-scale model, comprising: a processing module for processing a description of a target 3D scene to obtain label information in the description; an acquisition module for generating query operation presentation information of an LLM according to the label information, and using the LLM to acquire a target asset set matching the label information according to the query operation presentation information, the target asset set including target assets in the target 3D scene, target material information of the target assets, and target scene attribute information of the target assets; a generation module for generating the target three-dimensional scene based on the target asset set. A device for generating 3D scenes based on large-scale models.
9. The processing module further comprises: generating extraction operation presentation information of the LLM based on the description information; and processing the description information based on the extraction operation presentation information using the LLM to obtain the label information; The apparatus for generating a three-dimensional scene based on a large scale model according to claim 8.
10. The acquisition module further comprises: Using the LLM, based on the query operation presentation information, match the label information with a plurality of pre-recorded candidate information to obtain target information of the target asset set; used to obtain the target asset set based on the target information. The apparatus for generating a three-dimensional scene based on a large scale model according to claim 8.
11. The acquisition module further comprises: used to obtain the target asset set in a user self-defined local asset repository based on the target information; The apparatus for generating a three-dimensional scene based on a large scale model according to claim 10.
12. The generating module further comprises: generating an initial three-dimensional scene based on the target asset set; and adjusting the initial 3D scene based on scene feature information in the label information to generate the target 3D scene. The apparatus for generating a three-dimensional scene based on a large scale model according to claim 8.
13. a first modification module for displaying the label information to a user, obtaining label information modified by the user, obtaining a modified target asset set based on the modified label information, and generating a modified target three-dimensional scene based on the modified target asset set. An apparatus for generating a three-dimensional scene based on a large-scale model according to any one of claims 8 to 12.
14. a second modification module for displaying the target three-dimensional scene to a user and generating a modified target three-dimensional scene based on a modification instruction of the user. An apparatus for generating a three-dimensional scene based on a large-scale model according to any one of claims 8 to 12.
15. An electronic device, At least one processor; a memory in communication with the at least one processor; The memory stores instructions executable by the at least one processor, the instructions, when executed by the at least one processor, causing the at least one processor to perform a method according to any one of claims 1 to 5. electronic equipment.
16. A non-transitory computer-readable storage medium having computer instructions stored thereon, comprising: The computer instructions cause the computer to carry out a method according to any one of claims 1 to 5. A non-transitory computer-readable storage medium.
17. A computer program comprising: A computer program, which, when executed by a processor, implements the method according to any one of claims 1 to 5. Computer program.
Citation Information
Patent Citations
Digital twin three-dimensional scene construction method, device, equipment and medium
CN116310148A