Scene generation method and apparatus, and medium, device and program product

By receiving scene description information, generating candidate scene graphs and rendering them into three-dimensional scenes, the problem of complex scene construction in voxel games is solved, a fast and simplified scene generation process is achieved, and the user experience is improved.

WO2025200212A1PCT designated stage Publication Date: 2025-10-02BEIJING ZITIAO NETWORK TECH CO LTD
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
PCT/CN2024/108808
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-03-28
Filing Date
2024-07-31
Publication Date
2025-10-02

AI Technical Summary

Technical Problem

The scene construction in existing voxel games requires players to manually design and build, which has high technical requirements and a long cycle, and cannot meet the needs of users to quickly generate scenes.

Method used

By receiving scene description information, generating and displaying candidate scene graphs, and then selecting the target scene graph, voxel data is generated for rendering to achieve 3D scene generation.

Benefits of technology

It reduces the technical requirements for scene generation, simplifies the operation process, improves the efficiency of scene generation and its matching degree with user needs, reduces resource waste and enhances interactivity.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2024108808_02102025_PF_FP_ABST
    Figure CN2024108808_02102025_PF_FP_ABST
Patent Text Reader

Abstract

A scene generation method and apparatus, and a medium, a device and a program product. The scene generation method comprises: receiving scene description information for a target scene to be generated (11); on the basis of the scene description information, generating candidate scene graphs of a plurality of candidate scenes, wherein the candidate scene graphs are two-dimensional images (12); displaying the candidate scene graphs (13); in response to having received a selection operation of a user for a candidate scene graph, generating voxel data corresponding to the selected target scene graph (14); and performing rendering on the basis of the voxel data, so as to obtain a target scene, wherein the target scene is a three-dimensional scene (15). When a user intends to generate a scene, the user only needs to describe the scene he / she wants to construct, without the need for manual design and construction, thereby simplifying a scene generation operation, and effectively shortening a scene building cycle.
Need to check novelty before this filing date? Find Prior Art

Description

Scene generation method, device, medium, equipment and program product

[0001] This application claims priority to Chinese patent application No. 202410371010.X filed on March 28, 2024, and the contents of the above-mentioned Chinese patent application disclosure are hereby incorporated by reference in their entirety as a part of this application. Technical Field

[0002] Embodiments of the present disclosure relate to a scene generation method, apparatus, medium, device, and program product. Background Art

[0003] Current voxel gaming technology typically requires users to build scenes for interactive operation. Players often need to manually construct scenes to achieve their desired environment, or they can copy and share other players' scenes to build or modify them. Both of these construction methods rely on the player's full commitment to design and hands-on production skills, requiring high technical skills and a long construction cycle.

[0004] Summary of the Invention

[0005] In a first aspect, the present disclosure provides a scene generation method, the method comprising:

[0006] Receiving scene description information of a target scene to be generated;

[0007] Based on the scene description information, generating a candidate scene graph of a plurality of candidate scenes, wherein the candidate scene graph is a two-dimensional image;

[0008] Displaying the candidate scene graph;

[0009] In response to receiving a user selection operation on the candidate scene graph, generating voxel data corresponding to the selected target scene graph;

[0010] The target scene is obtained by rendering based on the voxel data, and the target scene is a three-dimensional scene.

[0011] In a second aspect, the present disclosure provides a scene generation device, the device comprising:

[0012] A receiving module is configured to receive scene description information of a target scene to be generated;

[0013] a first generating module configured to generate a candidate scene graph of a plurality of candidate scenes based on the scene description information, wherein the candidate scene graph is a two-dimensional image;

[0014] A first display module is configured to display the candidate scene graph;

[0015] a second generating module configured to generate voxel data corresponding to the selected target scene graph in response to receiving a user selection operation on the candidate scene graph;

[0016] The first processing module is configured to perform rendering based on the voxel data to obtain the target scene, where the target scene is a three-dimensional scene.

[0017] In a third aspect, the present disclosure provides a computer-readable medium having a computer program stored thereon, which implements the steps of the method described in the first aspect when executed by a processing device.

[0018] In a fourth aspect, the present disclosure provides an electronic device, comprising:

[0019] a storage device having a computer program stored thereon;

[0020] A processing device is configured to execute the computer program in the storage device to implement the steps of the method of the first aspect.

[0021] In a fifth aspect, the present disclosure provides a computer program product, comprising a computer program, which implements the steps of the method described in the first aspect when executed by a processor. BRIEF DESCRIPTION OF THE DRAWINGS

[0022] The above and other features, advantages and aspects of the various embodiments of the present disclosure will become more apparent with reference to the following detailed description in conjunction with the accompanying drawings. Throughout the drawings, the same or similar reference numerals represent the same or similar elements. It should be understood that the drawings are schematic and that the originals and elements are not necessarily drawn to scale. In the drawings:

[0023] FIG1 is a flowchart of a scene generation method provided according to an embodiment of the present disclosure.

[0024] FIG2 is a schematic diagram of a scene graph display interface provided according to an embodiment of the present disclosure.

[0025] FIG3 is a schematic diagram of receiving scene description information of a target scene to be generated according to an embodiment of the present disclosure.

[0026] FIG4 is a schematic diagram of a candidate scene graph display provided according to an embodiment of the present disclosure.

[0027] FIG5 is a schematic diagram of a grid display diagram provided according to an embodiment of the present disclosure.

[0028] FIG6 is a schematic diagram of target area selection according to an embodiment of the present disclosure.

[0029] FIG7 is a block diagram of a scene generation device according to an embodiment of the present disclosure.

[0030] FIG8 shows a schematic structural diagram of an electronic device suitable for implementing the embodiments of the present disclosure. DETAILED DESCRIPTION

[0031] The following describes embodiments of the present disclosure in more detail with reference to the accompanying drawings. Although certain embodiments of the present disclosure are shown in the accompanying drawings, it should be understood that the present disclosure can be implemented in various forms and should not be construed as limited to the embodiments described herein. Rather, these embodiments are provided to provide a more thorough and complete understanding of the present disclosure. It should be understood that the drawings and embodiments of the present disclosure are for illustrative purposes only and are not intended to limit the scope of protection of the present disclosure.

[0032] It should be understood that the various steps described in the method embodiments of the present disclosure may be performed in different orders and / or in parallel. In addition, the method embodiments may include additional steps and / or omit the steps shown. The scope of the present disclosure is not limited in this respect.

[0033] As used herein, the term "including" and its variations are open-ended, i.e., "including but not limited to." The term "based on" means "based, at least in part, on." The term "one embodiment" means "at least one embodiment," the term "another embodiment" means "at least one additional embodiment," and the term "some embodiments" means "at least some embodiments." Other terms are defined in the following description.

[0034] It should be noted that the concepts of "first" and "second" mentioned in this disclosure are only used to distinguish different devices, modules or units, and are not used to limit the order or interdependence of the functions performed by these devices, modules or units.

[0035] It should be noted that the modifications of "one" and "plurality" mentioned in the present disclosure are illustrative rather than restrictive, and those skilled in the art should understand that unless otherwise clearly indicated in the context, they should be understood as "one or more".

[0036] The names of the messages or information exchanged between multiple devices in the embodiments of the present disclosure are only used for illustrative purposes and are not used to limit the scope of these messages or information.

[0037] It is understandable that before using the technical solutions disclosed in the various embodiments of this disclosure, the type, scope of use, usage scenarios, etc. of the personal information involved in this disclosure should be informed to the user and the user's authorization should be obtained in an appropriate manner in accordance with relevant laws and regulations.

[0038] For example, in response to a user's active request, a prompt message is sent to the user to clearly inform the user that the operation requested will require the acquisition and use of the user's personal information. This allows the user to independently choose whether to provide personal information to the electronic device, application, server, storage medium, or other software or hardware that performs the operations of the disclosed technical solution based on the prompt message.

[0039] As an optional but non-limiting implementation, in response to receiving a user's active request, the prompt information may be sent to the user in the form of a pop-up window, in which the prompt information may be presented in text form. Furthermore, the pop-up window may also contain a selection control for the user to select "agree" or "disagree" to provide personal information to the electronic device.

[0040] It is understandable that the above notification and user authorization process are merely illustrative and do not limit the implementation of the present disclosure. Other methods that comply with relevant laws and regulations may also be applied to the implementation of the present disclosure.

[0041] At the same time, it is understandable that the data involved in this technical solution (including but not limited to the data itself, the acquisition or use of the data) must comply with the requirements of relevant laws, regulations and relevant provisions.

[0042] FIG1 is a flowchart of a scene generation method according to an embodiment of the present disclosure. As shown in FIG1 , the scene generation method may include:

[0043] In step 11, scene description information of a target scene to be generated is received.

[0044] The scene description information may be information provided by the user describing the target scene they wish to build. For example, the scene description information may be a natural language description text. For example, an information input interface may be displayed, such as a text input box, so that the user can directly enter the description according to their needs. For example, the user can enter parameters such as scene size and height in the corresponding interface based on the scene they wish to obtain, and enter keywords for the corresponding plot, such as style, color, holiday atmosphere, current events, etc. As another example, in order to ensure the comprehensiveness of the scene description information, multiple attributes may be pre-set for the user to fill in.

[0045] In step 12, based on the scene description information, a candidate scene graph of a plurality of candidate scenes is generated, wherein the candidate scene graph is a two-dimensional image.

[0046] In this step, the scene generation can be performed using a pre-trained scene generation model. For example, the scene generation model can be a large language model. Then, the scene description information can be input into the scene generation model to obtain the output image as the candidate scene graph. As another example, a portion of the image output by the scene generation model can be selected as the candidate scene graph, for example, by randomly selecting a portion from the output image.

[0047] In step 13, the candidate scene graph is displayed.

[0048] As an example, each generated candidate scene graph can be displayed separately in the scene graph display interface. If the current interface is not fully displayed, the scene graph display interface can be moved by sliding up and down or moving the scroll bar to display the candidate scene graphs, making it easier for users to select a scene graph that meets their building needs. As shown in Figure 2, the scene graph display interface can be moved by moving the scroll bar at A1.

[0049] In step 14 , in response to receiving a user selection operation on a candidate scene graph, voxel data corresponding to the selected target scene graph is generated.

[0050] Voxel is short for Volume Pixel. A volume containing voxels can be represented by stereo rendering or by extracting polygonal isosurfaces with a given threshold contour. In this embodiment, the voxel data corresponding to the candidate scene includes the voxel blocks used to construct the candidate scene and the attribute information of the voxel blocks. For example, if the candidate scene is a villa, it may include wall voxel blocks, ceiling voxel blocks, door frame voxel blocks, etc.

[0051] For example, based on the candidate scene graph determined in step 12 to be a two-dimensional image, this step can achieve a three-dimensional representation of the selected target scene graph by generating voxel data corresponding to the scene graph, so as to subsequently generate a three-dimensional scene corresponding to the target scene graph.

[0052] In step 15, rendering is performed based on the voxel data to obtain a target scene, which is a three-dimensional scene.

[0053] Among them, if the currently displayed candidate scene graph meets the user's construction requirements, the user can select the candidate scene graph and confirm. As shown in Figure 2, if the user selects the first candidate scene graph, the first candidate scene graph can be used as the target scene graph, and further rendered based on the voxel data corresponding to the first candidate scene graph to obtain the three-dimensional structure corresponding to the voxel data, that is, the target scene. Among them, the method of rendering the voxel data can be based on the common implementation method in this field, which will not be repeated here.

[0054] Therefore, through the above technical solution, when the user wants to generate a scene, the user only needs to describe the scene he wants to build, without the need for manual design and construction, which effectively reduces the technical requirements for the user to generate and build the scene, simplifies the scene generation operation, and effectively reduces the scene building cycle. In addition, the user's scene description information of the target scene can be used to generate a candidate scene graph for display to the user. The user can then preview a variety of possible scenes based on the candidate scene graph to determine the scene that meets his building needs. After the user selects the scene, the scene can be built. On the one hand, it can effectively improve the matching degree between the generated target scene and the user's needs, and at the same time, it can effectively avoid the waste of resources caused by generating a large number of candidate scenes, thereby improving the efficiency of scene generation and increasing the diversity of interactions between users.

[0055] In a possible embodiment, an exemplary implementation manner of receiving scene description information of a target scene to be generated may include:

[0056] A first dialogue area is displayed, wherein a plurality of candidate scene labels are displayed in the first dialogue area.

[0057] As shown in Q1 in Figure 3, it is the first dialogue area, wherein the first dialogue area can be used to prompt the user for input, such as "Please tell me the architectural style and characteristics you want, or you can choose from the keywords below", and accordingly, the candidate scene tags can be displayed in the first dialogue area. The candidate scene tags can be selected based on the tags in the tag library. For example, the topN can be selected as candidate scene tags in the order of selection frequency from high to low in the recent historical period. It can also be selected as candidate scene tags in the order of similarity from high to low with the scene corresponding to the current user. The specific selection method can be configured based on actual needs, and this disclosure does not limit this.

[0058] In response to the user's selection of a candidate scene label, the candidate scene label selected by the user is displayed in the second dialogue area, where the second dialogue area may be an area for the user to input information.

[0059] The second dialog area is shown as Q2 in Figure 3. When the user selects three candidate scene labels, namely, "European", "City", and "Street", the three selected candidate scene labels are automatically added and displayed in the second dialog area.

[0060] Thereafter, the scene description information may be determined based on the display information in the second dialog area.

[0061] As an example, if the user only selects a candidate scene tag, the tag information corresponding to the candidate scene tag displayed in the second dialogue area can be used as the scene description information.

[0062] Therefore, through the above technical solution, multiple candidate scene labels can be displayed in the first dialogue area to prompt the user, and multiple optional labels can be provided for the user to determine the scene description information, so as to refine the prompts of the scene that the user wants to generate to a certain extent, simplify the generation process of the scene description information, and further simplify the user operation process.

[0063] As another example, receiving scene description information of a target scene to be generated further includes:

[0064] The user's input information in the second dialogue area is received, where the input information includes text and / or images.

[0065] In order to further improve the comprehensiveness and diversity of scene description information, users can also enter information in the second dialogue area, such as text or images. For example, after selecting the "Street" label, the user can further describe what the street scene in the scene they want to build is like, or they can query and upload the corresponding street image from a preset image library or the Internet.

[0066] Accordingly, the label information corresponding to the input information and the candidate scene label selected by the user is used as the scene description information.

[0067] In this embodiment, the tag information and input information corresponding to the candidate scene tags displayed in the second dialogue area can be further used as the scene description information to facilitate subsequent accurate scene generation based on the scene description information.

[0068] Therefore, through the above technical solution, the scene description information can be determined by combining the candidate scene labels and the user's input information. The user can then intuitively describe the scene he wants to build through text description or image description, so that the corresponding scene can be automatically generated without the user having to design and manually build the scene building process, thereby improving the efficiency of scene generation.

[0069] As an example, an exemplary implementation of receiving scene description information of a target scene to be generated may include:

[0070] Receive user input information in the second dialogue area, the input information including text and / or images. Further, the input information can be directly used as scene description information. In this embodiment, the user can directly enter the description information of the scene they want to build in the second dialogue area.

[0071] In a possible embodiment, an exemplary implementation of generating a candidate scene graph of multiple candidate scenes based on scene description information may include:

[0072] Construct prompt text based on scene description information.

[0073] The prompt text is input into a scene generation model, and a candidate scene graph of multiple candidate scenes is obtained according to the output of the scene generation model, wherein the scene generation model is implemented based on a large language model.

[0074] As an example, the input format of the scene generation model can be pre-set. For example, the user's input information and label information can be spliced ​​together as a prompt text prompt. The scene generation model can be implemented based on a large language model (LLM). For example, the large language model can be trained and learned based on the voxel data of the generated scene and its description information. In the scene generation process, multiple candidate scene graphs can be generated based on the prompt text.

[0075] As an example, voxel data corresponding to a candidate scene graph can be generated directly based on a scene generation model. If only a two-dimensional surface graph is displayed in the candidate scene graph, the scene image corresponding to the voxel data can be screenshotted to obtain a candidate scene graph, so that the user can preview the scene based on the candidate scene graph and have an intuitive understanding of the external display of the scene before the scene rendering is generated. As another example, a two-dimensional candidate scene graph can be directly generated by a scene generation model so as to prompt the user based on the candidate scene graph, and then further generate the corresponding three-dimensional scene after the user confirms, thereby improving the user experience. The generation and recommendation of voxel data of scenes based on a large language model can, to a certain extent, break through the scope limitations of manual design and broaden the scope of use of the disclosed method.

[0076] In a possible embodiment, an exemplary implementation of displaying a candidate scene graph may include:

[0077] For the candidate scene graph, the editing control and the selection control corresponding to the candidate scene graph are displayed in the first display area corresponding to the candidate scene graph.

[0078] As an example, each candidate scene graph may be displayed, along with its corresponding editing control and selection control.

[0079] As another example, a portion of the scene graphs for display in the current round can be selected from the candidate scene graphs. As shown in Figure 4, four candidate scene graphs can be selected for display: candidate scene graphs S1-S4, corresponding to the first display areas P1-P4, respectively. The user can edit the candidate scene graphs using the edit control and select the candidate scene graphs using the select control. As shown in Figure 4, the control K1 can represent a select control, and the control K2 can be used to represent an edit control.

[0080] Accordingly, the scene generation method further includes:

[0081] In response to receiving a selection operation on an editing control, displaying a candidate scene graph corresponding to the selected target editing control in the editing interface;

[0082] In response to the editing information input by the user in the editing interface, a new candidate scene graph is generated and displayed according to the editing information and the candidate scene graph corresponding to the target editing control.

[0083] In this embodiment, when the generated candidate scene graph is displayed to the user, the user may be satisfied with the overall appearance of one of the candidate scene graphs, but only the local features thereof do not meet expectations. In this case, the user can edit and modify the single candidate scene graph through the editing control corresponding to the candidate scene graph.

[0084] As an example, the user can edit the candidate scene graph S1 by clicking the editing control K2 in the first display area P1. After the user selects the editing control, the editing interface can be displayed. As an example, the corresponding candidate scene graph and its editable properties can be displayed in the editing interface, such as building color, single-story height, number of floors and other properties. For example, if the user wants to change the number of floors of the candidate scene graph S1 to 5 floors, the user can directly enter the editing information at the corresponding number of floors attribute. After the user submits and confirms, a new candidate scene graph can be generated and displayed based on the editing information and the candidate scene graph corresponding to the target editing control.

[0085] As an example, when displaying candidate scene graphs, a new candidate scene graph S1' can be displayed to replace candidate scene graph S1, that is, S1', S2, S3, and S4 can be displayed. As another example, when a user edits candidate scene graph S1, it can be considered that the user is most interested in S1 among the currently displayed candidate scene graphs. After generating a new candidate display graph S1', only the newly generated candidate scene graph S1' can be displayed to facilitate the user to quickly confirm whether the scene corresponding to the new candidate scene graph meets their expectations.

[0086] As another example, the user can also select a partial area in the candidate scene graph that he wants to adjust. In response to the user's selection, the area location information of the area selected by the user can be used as editing information, and the editing information and the candidate scene graph can be input into the scene generation model to perform image adjustment and update to generate a new candidate scene graph.

[0087] Therefore, through the above technical solution, users can select a scene graph that meets their expectations from the candidate scene graphs, and can also realize secondary editing of a single candidate scene graph, further improving the matching degree between the candidate scene graph and user needs, while effectively reducing the data calculation and processing amount corresponding to the scene generation, and can provide effective and accurate data support for subsequent rendering of the target scene.

[0088] In a possible embodiment, another implementation of displaying the candidate scene graph may include:

[0089] A presentation scene graph is determined from the candidate scene graphs. As an example, M scene graphs may be randomly selected as the presentation scene graph, i.e., the scene graph to be presented in the current round.

[0090] The display scene graph is displayed. As an example, it can be displayed in the manner shown in FIG4 .

[0091] Accordingly, the scene generation method may further include:

[0092] When displaying the display scene graph, a scene update control is displayed, such as the control shown by B1 in FIG4 . The scene update control is used to trigger the update of the displayed scene graph.

[0093] In response to receiving a selection operation on a scene update control, determining a new presentation scene graph;

[0094] Display the new display scene graph.

[0095] In one possible embodiment, the currently displayed scene graphs may not meet the user's building requirements. In this embodiment, the user can click a scene update control. In response to receiving a selection operation on the scene update control, as an example, a new scene graph can be selected from the candidate scene graphs that have not yet been displayed as the display scene graph. As another example, if all scene graphs of the candidate scene graphs have been displayed, a batch of candidate scene graphs can be regenerated based on the scene description information, and a scene graph can be further selected from the regenerated candidate scene graphs as the display scene graph.

[0096] Therefore, in this embodiment, when there is no scene graph that meets the user's building needs in the currently displayed scene graph, multiple new display scene graphs can be re-displayed through scene update, and the user can choose from the displayed new display scene graphs, providing the user with more options. At the same time, a new display scene graph can be selected from the candidate scene graphs, reducing the call to the scene generation model during the scene generation process, and reducing the generation resource occupancy of the scene graph to a certain extent.

[0097] In a possible embodiment, the scene generation method may further include:

[0098] When displaying the candidate scene graph, a scene editing control is displayed, such as the control shown as B2 in FIG4 . The scene editing control is used to trigger the update of the scene description information.

[0099] As an example, when there is no scene graph that meets the user's building requirements among the displayed candidate scene graphs, the user can further adjust the scene description information to further clarify and clarify his or her own building requirements for the target scene. In this embodiment, the user can update the scene description information by clicking the scene editing control.

[0100] In response to receiving a selection operation on the scene editing control, a scene editing interface is displayed.

[0101] As an example, the scene editing interface may be the interface shown in FIG3 , or may be an input box interface.

[0102] In response to the user's editing operation in the scene editing interface, new scene description information is obtained, and the operation of generating a candidate scene graph of multiple candidate scenes based on the scene description information is returned.

[0103] As an example, the user may input new description information in the scene editing interface, and the existing scene description information and the new description information may be concatenated to serve as the new scene description information.

[0104] As another example, the scene editing interface may display existing scene description information, and the user may add or modify the current scene description information. In this example, the information submitted in the scene editing interface may be used as the new scene description information.

[0105] After obtaining new scene description information, multiple candidate scene graphs corresponding to the target scene can be generated based on the new scene description information, and the candidate scene graphs can be displayed; in response to receiving the user's selection operation for the candidate scene graph, the step of rendering the target scene is performed according to the voxel data corresponding to the selected target scene graph.

[0106] Therefore, through the above technical solution, when displaying candidate scene graphs to users, users can edit and adjust the scene description information multiple times to increase the amount of information in the scene description information, thereby obtaining a candidate scene graph that meets the user's building needs. In this process, a two-dimensional preview image is generated, and there is no need to render the three-dimensional scene multiple times. While ensuring that the generated target scene meets the user's building needs, the efficiency of scene generation is improved. At the same time, the richness of UGC (User Generated Content) is improved to provide data reference and inspiration for user-generated content, making it easier for users to obtain the target scene they want.

[0107] In a possible embodiment, in response to receiving a user selection operation on a candidate scene graph, generating voxel data corresponding to the selected target scene graph may include:

[0108] In response to the selection operation, the selected target scene graph is gridded to obtain a grid display graph corresponding to the target scene graph.

[0109] As shown in Figure 2, if the user believes that the first image meets their needs, they can select the first image. After the user confirms the selection, the selected target scene image can be gridded. As an example, the target scene image can be pixelated to obtain a grid display image corresponding to the target scene image. As shown in Figure 5, it is a grid display image corresponding to the target scene image selected by the user. In Figure 5, only the gridding processing of the roof slant eaves is shown as an example. Other parts are directly gridded, such as the gridding processing of surface B, which is not shown.

[0110] The grid display image is then displayed; in response to receiving a confirmation operation on the grid display image, voxel data corresponding to the confirmed target scene image is generated, and the confirmed target scene image is the target scene image corresponding to the grid display image indicated by the confirmation operation.

[0111] Among them, after the scene graph is gridded, its display is slightly different from the two-dimensional image. In this case, when the scene graph is gridded, the user can select one or more candidate scene graphs for gridding. If the user selects multiple scene graphs, each scene graph selected by the user can be gridded and output for display, so that the user can further determine the scene graph that meets their needs from the grid display graphs. The user can select the grid display graph they want from the displayed grid display graphs, and in response to the confirmation operation, the scene graph that the user finally determined to meet their needs can be determined.

[0112] Therefore, through the above technical solution, the candidate scene graphs can be displayed to allow the user to have a rough preview of the scene graphs that can be generated, and the user can make a preliminary selection. The scene graph selected by the user can then be gridded, so that the user can further understand the overview of the corresponding three-dimensional scene graph based on the grid display graph to determine the scene graph used to generate voxel data later, thereby effectively reducing the processing volume of generating voxel data, improving the usability of the generated voxel data, and also improving the diversity of interaction with the user.

[0113] In a possible embodiment, in response to receiving a user selection operation on a candidate scene graph, generating voxel data corresponding to the selected target scene graph may include:

[0114] The target scene graph is input into a voxel data model to obtain voxel data corresponding to the target scene graph, wherein the voxel data model is obtained by training based on voxel data corresponding to the three-dimensional scene and scene screenshot data corresponding to the three-dimensional scene.

[0115] As an example, a voxel data model can be trained based on data from existing scenes. For example, training can be performed based on voxel data corresponding to an already constructed scene and scene screenshot data of the scene. For example, the scene screenshot data can be used as the input of the model, and the voxel data corresponding to the scene can be used as the target input of the model to train the model to obtain the voxel data model. The voxel data model can be implemented based on a large language model or a neural network model.

[0116] Accordingly, the target scene graph can be input into the voxel data model to obtain the corresponding voxel data. If the target scene graph has a corresponding grid representation, the grid representation can be input into the voxel data model; if the target scene graph does not generate a corresponding grid representation, the target scene graph can be input into the voxel data model to obtain the voxel data.

[0117] Therefore, through the above technical solution, voxel data corresponding to the scene graph can be quickly generated through the voxel data model. In this process, the correlation between the display graph of the existing scene and the voxel data in its construction process can be learned, thereby improving the effectiveness and accuracy of the voxel data, and improving the consistency between the image rendered based on the voxel data and the target scene graph, thereby ensuring the user experience.

[0118] In a possible embodiment, in response to receiving a user selection operation on a candidate scene graph, an exemplary implementation of generating voxel data corresponding to the selected target scene graph may include:

[0119] Get the grid display image corresponding to the target scene graph.

[0120] Among them, if the target scene graph has a corresponding grid display graph, it can be directly obtained. If the target scene graph has not yet generated a corresponding grid display graph, the target scene graph can be pixelated to generate the grid display graph.

[0121] For each grid in the grid display image, determine the voxel block corresponding to the grid and obtain display voxel data.

[0122] As an example, various scene parts in the target scene graph can be identified, such as door frames, walls, roofs, foundations, and the like. The properties of the meshes therein can then be determined based on the types of the identified parts. The properties of each mesh can then be matched with the types of voxel blocks in a voxel block library to determine the voxel block corresponding to the mesh. Voxel blocks with the same properties as the mesh or a similarity greater than a threshold can be used as the voxel blocks corresponding to the mesh, thereby obtaining display voxel data corresponding to the portion of the target scene graph displayed to the user.

[0123] The portion of the target scene graph that is not displayed to the user can then be further filled based on the target scene and the display voxel data. For example, the size information corresponding to the target scene graph can be determined based on the grid display graph.

[0124] As an example, the size information includes the length and width of the target scene on the horizontal plane. Accordingly, an exemplary implementation method for determining the size information corresponding to the target scene graph based on the grid display graph may include:

[0125] If the grid display image includes display images of two adjacent vertical surfaces of the target scene, the length and width of the target scene on the horizontal plane are determined according to the display images of the two adjacent surfaces.

[0126] As shown in Figure 5, the grid display includes two adjacent vertical surfaces of the target scene, namely, surface A and surface B in Figure 5. For example, the length of the intersection of surface A and the horizontal plane can be used as the length of the target scene, and the length of the intersection of surface B and the horizontal plane can be used as the width of the target scene. The length of the intersection can be adjusted based on the tilt angle of surface B to determine the width of the target scene. The adjustment method can be based on the dimensional change under perspective angle commonly used in the art, which will not be further described here.

[0127] If the grid display image only includes a display image of a vertical surface of the target scene, the length corresponding to the vertical surface is determined based on the display image of the vertical surface, and the width is determined based on the length corresponding to the vertical surface.

[0128] If the grid display image only contains a vertical surface of the target scene, that is, the target scene image is displayed directly on the display screen, the length of the intersection of the display image and the horizontal surface can be used as the length of the target scene. As an example, a default ratio of length and width can be pre-set. After determining the length, the corresponding width can be determined based on the length and the default ratio to obtain the size information.

[0129] As another example, the user can pre-set the area for scene generation in the scene map. The length and width need to be within the area selected by the user. The size information can then be adjusted based on the area selected by the user. That is, the size information corresponding to the target scene graph is constrained using the user-selected area as the maximum range of the size information. If the determined width exceeds the width of the user-selected area, the width of the user-selected area is used as the width of the target scene graph. Thus, through the above technical solution, the corresponding size information of the generated target scene graph can be further determined, providing reliable data support for the subsequent generation of its three-dimensional structure.

[0130] Afterwards, the display voxel data is filled based on the grid display graph and size information to obtain the voxel data corresponding to the target scene graph.

[0131] As an example, as shown in Figure 5, if the width determined based on the B surface is 10, then the grid on the front of the grid display image can be filled in the S direction shown in Figure 5. The filling can be based on the default voxel block or the association between different voxel blocks. Multiple filling rules can be pre-set. For example, if the voxel block type in the display voxel data is a wall, then the four voxel blocks connected in the S direction must also be of the wall type. When filling the display voxel data, the corresponding filling can be performed based on this rule. The internal structure of the scene that is not displayed can be matched based on the part that has been displayed on the surface, and then filled according to the matched structure. If the display voxel data indicates that the scene is a three-story building structure, then the internal structure can match the building structure, and the combination of multiple voxel blocks corresponding to the staircase structure can be filled when filling the interior. For example, if the display voxel data indicates that the top floor of the scene is a spire structure, then filling can be performed in the S direction corresponding to the top according to the matched spire structure, thereby obtaining the voxel data corresponding to the target scene graph. Among them, default voxels can be filled in during internal filling. Users can subsequently change and adjust the type of internal voxel data to achieve personalized scene building. This can also further improve the user's interactive method for scene building and simplify the user's scene building process.

[0132] In a possible embodiment, the scene generation method may further include:

[0133] Display a scene map. The scene map may be a map of areas in the virtual scene where scene construction is allowed, which may be determined based on pre-configuration.

[0134] Determining a target area in response to area information selected by a user in a scene map;

[0135] Display the target scene in the target area.

[0136] As an example, the user can select the area in the scene map where he wants to build the scene. As shown in Figure 6, the area selection can be performed in the form of a parallelogram. As an example, you can first move to the starting position A, and then move to the end position A'. The area information can include the starting position and the end position. After determining the starting and end positions, the middle area formed by the straight line from the starting point to the end point as the diagonal line can be used as the target area. The target area can be generated based on the default height, or the user can further determine an auxiliary third point, with the third point as the vertex and the middle area formed by the straight line from the starting point to the end point as the diagonal line as the target area.

[0137] Thus, the target area can be further determined in the scene map to quickly determine the environment in which the constructed scene is located. At the same time, the environment in which the target area is located can also be used as a reference for users to determine scene description information.

[0138] In a possible embodiment, generating a candidate scene graph of multiple candidate scenes based on the scene description information may include:

[0139] The scene description information and the target area are input into the scene generation model to obtain multiple candidate scenes corresponding to the target area.

[0140] Among them, the difference in regions may also have an impact on the construction of the scene. For example, if the target scene is a street, the smaller the target area is, the fewer buildings the street scene usually contains. In this case, the target area is further combined with the candidate scene graph to generate the candidate scene to improve the consistency between the candidate scene and the target area.

[0141] In a possible embodiment, displaying a target scene in a target area includes:

[0142] If the target scene contains multiple scene units, the multiple scene units are displayed in the target area.

[0143] Among them, the scene unit can be an inseparable scene building that is pre-set according to the actual application scenario, such as a house or a building.

[0144] If the target scene is a street, it can be composed of multiple houses and buildings, which can include multiple scene units. When generating the street scene, the scene units are combined based on the selected target area to generate the target scene. The scene units in the combination are generated in the target area. In this embodiment, the multiple scene units can be directly displayed in the target area.

[0145] If the target scene includes a scene unit, the target scene is displayed at the target position in response to the target position being selected in the target area.

[0146] As an example, the target scene is a building, which includes a scene unit, which can be an independent scene, and can be located anywhere in the target area. In this example, the user can specify the target location corresponding to the target scene in the target area, and then display it at the target location.

[0147] Therefore, through the above technical solution, target scenes such as building complexes or combined scenes can be displayed directly in the target area without the user having to configure the location, thereby improving the adaptation between the target scene and the target area. For a single scene unit, the user can adjust its display position, and to a certain extent, the diversity of the scene display position can be improved, further improving the content richness and interactivity of the scene generation, and enhancing the user experience.

[0148] Based on the same inventive concept, the present disclosure also provides a scene generation device. As shown in FIG7 , the scene generation device 10 includes: a receiving module 100 , a first generation module 200 , a first display module 300 , a second generation module 400 and a first processing module 500 .

[0149] The receiving module 100 is configured to receive scene description information of a target scene to be generated;

[0150] A first generating module 200 is configured to generate a candidate scene graph of a plurality of candidate scenes based on the scene description information, wherein the candidate scene graph is a two-dimensional image;

[0151] A first display module 300 is configured to display a candidate scene graph;

[0152] The second generating module 400 is configured to generate voxel data corresponding to the selected target scene graph in response to receiving a user selection operation on the candidate scene graph;

[0153] The first processing module 500 is configured to perform rendering based on the voxel data to obtain a target scene, where the target scene is a three-dimensional scene.

[0154] Optionally, the receiving module includes:

[0155] A first display submodule is configured to display a first dialogue area, wherein a plurality of candidate scene labels are displayed in the first dialogue area;

[0156] a first processing submodule configured to, in response to a user selecting a candidate scene tag, display the candidate scene tag selected by the user in the second dialogue area;

[0157] The first determining submodule is configured to determine scene description information according to display information in the second dialogue area.

[0158] Optionally, the receiving module further includes:

[0159] a receiving submodule, configured to receive input information from a user in a second dialogue area, the input information including text and / or images;

[0160] The second determining submodule is configured to use the input information and the tag information corresponding to the candidate scene tag selected by the user as the scene description information.

[0161] Optionally, the second generation module includes:

[0162] The second processing submodule is configured to, in response to the selection operation, perform gridding processing on the selected target scene graph to obtain a grid display graph corresponding to the target scene graph;

[0163] The second display submodule is configured to display a grid display image;

[0164] In response to receiving a confirmation operation on the grid presentation graph, voxel data corresponding to the confirmed target scene graph is generated, and the confirmed target scene graph is the target scene graph corresponding to the grid presentation graph indicated by the confirmation operation.

[0165] Optionally, the second generation module includes:

[0166] The third processing submodule is configured to input the target scene graph into a voxel data model to obtain voxel data corresponding to the target scene graph, wherein the voxel data model is obtained by training based on voxel data corresponding to the three-dimensional scene and scene screenshot data corresponding to the three-dimensional scene.

[0167] Optionally, the second generation module includes:

[0168] An acquisition submodule, configured to acquire a grid display image corresponding to a target scene image;

[0169] A third determining submodule is configured to determine, for each grid in the grid display image, a voxel block corresponding to the grid, and obtain display voxel data;

[0170] A fourth determining submodule is configured to determine size information corresponding to the target scene graph based on the grid display graph;

[0171] The fourth processing submodule is configured to fill the display voxel data based on the grid display graph and the size information to obtain voxel data corresponding to the target scene graph.

[0172] Optionally, the size information includes the length and width of the target scene on the horizontal plane;

[0173] The fourth determination submodule includes:

[0174] a fifth determining submodule configured to determine the length and width of the target scene on the horizontal plane based on the display images of the two adjacent vertical surfaces if the grid display image includes the display images of the two adjacent surfaces;

[0175] The sixth determining submodule is configured to determine the length corresponding to the vertical surface based on the display image of the vertical surface if the grid display image only includes a display image of a vertical surface of the target scene, and determine the width based on the length corresponding to the vertical surface.

[0176] Optionally, the scene generating device further includes:

[0177] A second display module is configured to display a scene map;

[0178] A first determining module is configured to determine a target area in response to area information selected by a user in the scene map;

[0179] The third display module is configured to display the target scene in the target area.

[0180] Optionally, the third display module includes:

[0181] a third display submodule, configured to display the multiple scene units in the target area if the target scene includes multiple scene units;

[0182] The fourth display submodule is configured to display the target scene at the target position in response to the target position selected in the target area if the target scene includes a scene unit.

[0183] Optionally, the first display module includes:

[0184] A fifth display submodule is configured to display, for a candidate scene graph, an edit control and a selection control corresponding to the candidate scene graph in a first display area corresponding to the candidate scene graph;

[0185] The scene generating device also includes:

[0186] a fourth display module configured to display, in response to receiving a selection operation on the editing control, a candidate scene graph corresponding to the selected target editing control in the editing interface;

[0187] The second processing module is configured to generate and display a new candidate scene graph in response to editing information input by the user in the editing interface according to the editing information and the candidate scene graph corresponding to the target editing control.

[0188] Optionally, the first display module includes:

[0189] a seventh determination submodule, configured to determine a presentation scene graph from the candidate scene graphs;

[0190] a sixth display submodule, configured to display the display scene graph;

[0191] The scene generating device also includes:

[0192] a fifth display module, configured to display a scene update control when displaying the display scene graph;

[0193] a second determining module, configured to determine a new presentation scene graph in response to receiving a selection operation on the scene update control;

[0194] The sixth display module is configured to display a new display scene graph.

[0195] Optionally, the scene generating device further includes:

[0196] a seventh display module, configured to display a scene editing control when displaying the candidate scene graph;

[0197] an eighth display module, configured to display a scene editing interface in response to receiving a selection operation on the scene editing control;

[0198] The third processing module is configured to obtain new scene description information in response to the user's editing operation in the scene editing interface, and trigger the first generation module to execute the operation of generating a candidate scene graph of multiple candidate scenes based on the scene description information.

[0199] Optionally, the first generating module includes:

[0200] a fifth processing submodule, configured to construct a prompt text according to the scene description information;

[0201] The generation submodule is configured to input the prompt text into a scene generation model, and obtain a candidate scene graph of multiple candidate scenes according to the output of the scene generation model, wherein the scene generation model is implemented based on a large language model.

[0202] Reference is now made to FIG8 , which illustrates a schematic diagram of the structure of an electronic device (e.g., a terminal device or server) 600 suitable for implementing embodiments of the present disclosure. The terminal device in the embodiments of the present disclosure may include, but is not limited to, mobile terminals such as mobile phones, laptop computers, digital broadcast receivers, PDAs (personal digital assistants), PADs (tablet computers), PMPs (portable multimedia players), in-vehicle terminals (e.g., in-vehicle navigation terminals), and fixed terminals such as digital TVs and desktop computers. The electronic device illustrated in FIG8 is merely an example and should not limit the functionality or scope of use of the embodiments of the present disclosure.

[0203] As shown in Figure 8, the electronic device 600 may include a processing device (e.g., a central processing unit, a graphics processing unit, etc.) 601, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 602 or a program loaded from a storage device 608 into a random access memory (RAM) 603. Various programs and data required for the operation of the electronic device 600 are also stored in the RAM 603. The processing device 601, the ROM 602, and the RAM 603 are connected to each other via a bus 604. An input / output (I / O) interface 605 is also connected to the bus 604.

[0204] Typically, the following devices may be connected to the I / O interface 605: an input device 606 including, for example, a touch screen, a touchpad, a keyboard, a mouse, a camera, a microphone, an accelerometer, a gyroscope, etc.; an output device 607 including, for example, a liquid crystal display (LCD), a speaker, a vibrator, etc.; a storage device 608 including, for example, a magnetic tape, a hard disk, etc.; and a communication device 609. The communication device 609 may allow the electronic device 600 to communicate with other devices wirelessly or by wire to exchange data. Although FIG8 shows the electronic device 600 with various devices, it should be understood that not all of the devices shown are required to be implemented or present. More or fewer devices may be implemented or present instead.

[0205] In particular, according to an embodiment of the present disclosure, the process described above with reference to the flowchart can be implemented as a computer software program. For example, an embodiment of the present disclosure includes a computer program product, which includes a computer program carried on a non-transitory computer-readable medium, and the computer program includes a program code for executing the method shown in the flowchart. In such an embodiment, the computer program can be downloaded and installed from the network through the communication device 609, or installed from the storage device 608, or installed from the ROM 602. When the computer program is executed by the processing device 601, the above-mentioned functions defined in the method of the embodiment of the present disclosure are performed.

[0206] It should be noted that the computer-readable medium mentioned above in the present disclosure may be a computer-readable signal medium or a computer-readable storage medium, or any combination of the two. A computer-readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, device, or component, or any combination of the above. More specific examples of computer-readable storage media may include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present disclosure, a computer-readable storage medium may be any tangible medium that contains or stores a program that can be used by or in conjunction with an instruction execution system, device, or component. In the present disclosure, a computer-readable signal medium may include a data signal propagated in baseband or as part of a carrier wave, which carries computer-readable program code. Such a propagated data signal may take a variety of forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. A computer-readable signal medium may also be any computer-readable medium other than a computer-readable storage medium that can transmit, propagate, or transport a program for use by or in conjunction with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium may be transmitted using any suitable medium, including but not limited to wires, optical cables, RF (radio frequency), etc., or any suitable combination thereof.

[0207] In some embodiments, the client and server can communicate using any currently known or later developed network protocol, such as HTTP (HyperText Transfer Protocol), and can be interconnected with any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include a local area network ("LAN"), a wide area network ("WAN"), an internet (e.g., the Internet), and a peer-to-peer network (e.g., an ad hoc peer-to-peer network), as well as any currently known or later developed network.

[0208] The computer-readable medium may be included in the electronic device, or may exist independently without being incorporated into the electronic device.

[0209] The above-mentioned computer-readable medium carries one or more programs. When the above-mentioned one or more programs are executed by the electronic device, the electronic device: receives scene description information of the target scene to be generated; generates candidate scene graphs of multiple candidate scenes based on the scene description information, wherein the candidate scene graphs are two-dimensional images; displays the candidate scene graphs; in response to receiving a user's selection operation for the candidate scene graph, generates voxel data corresponding to the selected target scene graph; and renders based on the voxel data to obtain the target scene, wherein the target scene is a three-dimensional scene.

[0210] Computer program code for performing the operations of the present disclosure may be written in one or more programming languages, or a combination thereof, including, but not limited to, object-oriented programming languages ​​such as Java, Smalltalk, C++, and conventional procedural programming languages ​​such as "C" or similar programming languages. The program code may be executed entirely on the user's computer, partially on the user's computer, as a stand-alone software package, partially on the user's computer and partially on a remote computer, or entirely on the remote computer or server. In cases involving a remote computer, the remote computer may be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., through the Internet using an Internet service provider).

[0211] The flowcharts and block diagrams in the accompanying drawings illustrate the possible implementation architecture, functions and operations of the systems, methods and computer program products according to various embodiments of the present disclosure. In this regard, each box in the flowchart or block diagram can represent a module, program segment, or a part of code, and the module, program segment, or a part of code contains one or more executable instructions for realizing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the box can also occur in a different order than that marked in the accompanying drawings. For example, two boxes represented in succession can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram and / or flowchart, and the combination of the boxes in the block diagram and / or flowchart, can be implemented with a dedicated hardware-based system that performs the specified function or operation, or can be implemented with a combination of dedicated hardware and computer instructions.

[0212] The modules described in the embodiments of the present disclosure may be implemented in software or hardware. In some cases, the name of a module does not limit the module itself. For example, a receiving module may also be described as a "module that receives scene description information of a target scene to be generated."

[0213] The functions described above herein may be performed, at least in part, by one or more hardware logic components. For example, and without limitation, exemplary types of hardware logic components that may be used include: field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), systems on chip (SOCs), complex programmable logic devices (CPLDs), and the like.

[0214] In the context of the present disclosure, a machine-readable medium can be a tangible medium that can contain or store a program for use by or in conjunction with an instruction execution system, device or equipment. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, device or equipment, or any suitable combination of the foregoing. A more specific example of a machine-readable storage medium can include an electrical connection based on one or more lines, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0215] According to one or more embodiments of the present disclosure, Example 1 provides a scene generation method, the scene generation method including:

[0216] Receiving scene description information of a target scene to be generated;

[0217] Based on the scene description information, generating a candidate scene graph for a plurality of candidate scenes, wherein the candidate scene graph is a two-dimensional image generated based on voxel data corresponding to the candidate scenes;

[0218] Display candidate scene graphs;

[0219] In response to receiving a user's selection operation on a candidate scene graph, rendering is performed according to voxel data corresponding to the selected target scene graph to obtain a target scene, where the target scene is a three-dimensional scene.

[0220] According to one or more embodiments of the present disclosure, Example 2 provides the method of Example 1, wherein receiving scene description information of a target scene to be generated includes:

[0221] Displaying a first dialogue area, wherein a plurality of candidate scene labels are displayed in the first dialogue area;

[0222] In response to a user selecting a candidate scene label, displaying the candidate scene label selected by the user in the second dialogue area;

[0223] The scene description information is determined according to the display information in the second dialog area.

[0224] According to one or more embodiments of the present disclosure, Example 3 provides the method of Example 2, wherein receiving scene description information of a target scene to be generated further includes:

[0225] receiving user input information in the second dialogue area, where the input information includes text and / or images;

[0226] The label information corresponding to the input information and the candidate scene label selected by the user is used as the scene description information.

[0227] According to one or more embodiments of the present disclosure, Example 4 provides the method of Example 1, wherein, in response to receiving a user selection operation for a candidate scene graph, generating voxel data corresponding to the selected target scene graph includes:

[0228] In response to the selection operation, the selected target scene graph is gridded to obtain a grid display graph corresponding to the target scene graph;

[0229] Display grid display diagram;

[0230] In response to receiving a confirmation operation on the grid presentation graph, voxel data corresponding to the confirmed target scene graph is generated, and the confirmed target scene graph is the target scene graph corresponding to the grid presentation graph indicated by the confirmation operation.

[0231] According to one or more embodiments of the present disclosure, Example 5 provides the method of Example 1, wherein, in response to receiving a user selection operation for a candidate scene graph, generating voxel data corresponding to the selected target scene graph includes:

[0232] The target scene graph is input into a voxel data model to obtain voxel data corresponding to the target scene graph, wherein the voxel data model is obtained by training based on voxel data corresponding to the three-dimensional scene and scene screenshot data corresponding to the three-dimensional scene.

[0233] According to one or more embodiments of the present disclosure, Example 6 provides the method of Example 1, wherein, in response to receiving a user selection operation for a candidate scene graph, generating voxel data corresponding to the selected target scene graph includes:

[0234] Get the grid display image corresponding to the target scene graph;

[0235] For each grid in the grid display image, determine the voxel block corresponding to the grid and obtain display voxel data;

[0236] Determine the size information corresponding to the target scene graph based on the grid display graph;

[0237] The display voxel data is filled based on the grid display image and size information to obtain the voxel data corresponding to the target scene image.

[0238] According to one or more embodiments of the present disclosure, Example 7 provides the method of Example 6, wherein the size information includes a length and a width corresponding to the target scene on a horizontal plane;

[0239] Determine the size information corresponding to the target scene graph based on the grid display graph, including:

[0240] If the grid display image includes display images of two adjacent vertical surfaces of the target scene, then the length and width of the target scene on the horizontal plane are determined according to the display images of the two adjacent surfaces respectively;

[0241] If the grid display image only includes a display image of a vertical surface of the target scene, the length corresponding to the vertical surface is determined based on the display image of the vertical surface, and the width is determined based on the length corresponding to the vertical surface.

[0242] According to one or more embodiments of the present disclosure, Example 8 provides the method of Example 1, wherein the scene generation method further includes:

[0243] Display scene map;

[0244] Determining a target area in response to area information selected by a user in a scene map;

[0245] Display the target scene in the target area.

[0246] According to one or more embodiments of the present disclosure, Example 9 provides the method of Example 8, wherein presenting the target scene in the target area includes:

[0247] If the target scene contains multiple scene units, then the multiple scene units are displayed in the target area;

[0248] If the target scene includes a scene unit, the target scene is displayed at the target position in response to the target position being selected in the target area.

[0249] According to one or more embodiments of the present disclosure, Example 10 provides the method of Example 1, wherein presenting the candidate scene graph includes:

[0250] For the candidate scene graph, displaying the editing control and the selection control corresponding to the candidate scene graph in the first display area corresponding to the candidate scene graph;

[0251] The scene generation method further includes:

[0252] In response to receiving a selection operation on an editing control, displaying a candidate scene graph corresponding to the selected target editing control in the editing interface;

[0253] In response to the editing information input by the user in the editing interface, a new candidate scene graph is generated and displayed according to the editing information and the candidate scene graph corresponding to the target editing control.

[0254] According to one or more embodiments of the present disclosure, Example 11 provides the method of Example 1, wherein presenting a candidate scene graph includes:

[0255] Determine a display scene graph from the candidate scene graphs;

[0256] Displaying the display scene graph;

[0257] The scene generation method further includes:

[0258] When displaying the display scene graph, display scene update controls;

[0259] In response to receiving a selection operation on a scene update control, determining a new presentation scene graph;

[0260] Display the new display scene graph.

[0261] According to one or more embodiments of the present disclosure, Example 12 provides the method of Example 1, wherein the scene generation method further includes:

[0262] Display scene editing controls when displaying candidate scene graphs;

[0263] In response to receiving a selection operation on the scene editing control, displaying a scene editing interface;

[0264] In response to the user's editing operation in the scene editing interface, new scene description information is obtained, and the operation of generating a candidate scene graph of multiple candidate scenes based on the scene description information is returned.

[0265] According to one or more embodiments of the present disclosure, Example 13 provides the method of Example 1, wherein generating a candidate scene graph of multiple candidate scenes based on scene description information includes:

[0266] Construct prompt text based on scene description information;

[0267] The prompt text is input into a scene generation model, and a candidate scene graph of multiple candidate scenes is obtained according to the output of the scene generation model, wherein the scene generation model is implemented based on a large language model.

[0268] According to one or more embodiments of the present disclosure, Example 14 provides a scene generation device, which includes: a receiving module, configured to receive scene description information of a target scene to be generated; a first generation module, configured to generate a candidate scene graph of multiple candidate scenes based on the scene description information, wherein the candidate scene graph is a two-dimensional image; a first display module, configured to display the candidate scene graph; a second generation module, configured to generate voxel data corresponding to the selected target scene graph in response to receiving a user's selection operation for the candidate scene graph; and a first processing module, configured to render based on the voxel data to obtain a target scene, wherein the target scene is a three-dimensional scene.

[0269] According to one or more embodiments of the present disclosure, Example 15 provides a computer-readable medium having a computer program stored thereon, which, when executed by a processing device, implements the steps of the scene generation method involved in any one of Examples 1-13.

[0270] According to one or more embodiments of the present disclosure, Example 16 provides an electronic device, comprising: a storage device on which a computer program is stored; and a processing device configured to execute the computer program in the storage device to implement the steps of the scene generation method involved in any one of Examples 1-13.

[0271] According to one or more embodiments of the present disclosure, Example 17 provides a computer program product, including a computer program, which implements the steps of the method of any one of Examples 1-13 when executed by a processor.

[0272] The above description is merely a preferred embodiment of the present disclosure and an illustration of the technical principles employed. Those skilled in the art should understand that the scope of disclosure involved in the present disclosure is not limited to the technical solutions formed by the specific combination of the above-mentioned technical features, but also includes other technical solutions formed by any combination of the above-mentioned technical features or their equivalents without departing from the above-mentioned disclosed concepts. For example, a technical solution formed by replacing the above-mentioned features with (but not limited to) technical features with similar functions disclosed in this disclosure.

[0273] In addition, although each operation is described in a specific order, this should not be understood as requiring these operations to be performed in the specific order shown or in a sequential order. Under certain circumstances, multitasking and parallel processing may be advantageous. Similarly, although some specific implementation details have been included in the above discussion, these should not be interpreted as limiting the scope of the present disclosure. Some features described in the context of a separate embodiment can also be implemented in a single embodiment in combination. On the contrary, the various features described in the context of a single embodiment can also be implemented in multiple embodiments individually or in any suitable sub-combination mode.

[0274] Although the subject matter has been described using language specific to structural features and / or methodological logical acts, it should be understood that the subject matter defined in the appended claims is not necessarily limited to the specific features or acts described above. Rather, the specific features and acts described above are merely example forms of implementing the claims. Regarding the apparatus in the above-described embodiments, the specific manner in which each module performs operations has been described in detail in the embodiments related to the method and will not be elaborated upon here.

Claims

1. A scene generation method, comprising: Receiving scene description information of a target scene to be generated; Based on the scene description information, generating a candidate scene graph of a plurality of candidate scenes, wherein the candidate scene graph is a two-dimensional image; Displaying the candidate scene graph; In response to receiving a user selection operation on the candidate scene graph, generating voxel data corresponding to the selected target scene graph; The target scene is obtained by rendering based on the voxel data, wherein the target scene is a three-dimensional scene.

2. The method according to claim 1, wherein The receiving scene description information of the target scene to be generated includes: Displaying a first dialogue area, wherein a plurality of candidate scene labels are displayed in the first dialogue area; In response to the user selecting the candidate scene label, displaying the candidate scene label selected by the user in the second dialogue area; The scene description information is determined according to display information in the second dialog area.

3. The method according to claim 2, wherein: The receiving scene description information of the target scene to be generated further includes: receiving input information from the user in the second dialogue area, wherein the input information includes text and / or images; The input information and the tag information corresponding to the candidate scene tag selected by the user are used as the scene description information.

4. The method according to claim 1, wherein The step of generating voxel data corresponding to the selected target scene graph in response to receiving a user selection operation on the candidate scene graph includes: In response to the selection operation, gridding the selected target scene graph to obtain a grid display graph corresponding to the target scene graph; displaying the grid display diagram; In response to receiving a confirmation operation on the grid presentation graph, voxel data corresponding to the confirmed target scene graph is generated, wherein the confirmed target scene graph is the target scene graph corresponding to the grid presentation graph indicated by the confirmation operation.

5. The method according to claim 1, wherein The step of generating voxel data corresponding to the selected target scene graph in response to receiving a user selection operation on the candidate scene graph includes: The target scene graph is input into a voxel data model to obtain voxel data corresponding to the target scene graph, wherein the voxel data model is obtained by training based on voxel data corresponding to a three-dimensional scene and scene screenshot data corresponding to the three-dimensional scene.

6. The method according to claim 1, wherein The step of generating voxel data corresponding to the selected target scene graph in response to receiving a user selection operation on the candidate scene graph includes: Obtaining a grid display image corresponding to the target scene image; For each grid in the grid display image, determining a voxel block corresponding to the grid to obtain display voxel data; Determine size information corresponding to the target scene graph based on the grid display graph; The display voxel data is filled based on the grid display graph and the size information to obtain voxel data corresponding to the target scene graph.

7. The method according to claim 6, wherein: The size information includes the length and width of the target scene on the horizontal plane; The determining the size information corresponding to the target scene graph based on the grid presentation graph includes: If the grid display image includes display images of two adjacent vertical surfaces of the target scene, determining the length and width of the target scene on the horizontal plane according to the display images of the two adjacent surfaces respectively; If the grid display image only includes a display image of a vertical surface of the target scene, the length corresponding to the vertical surface is determined based on the display image of the vertical surface, and the width is determined based on the length corresponding to the vertical surface.

8. The method according to claim 1, further comprising: Display scene map; determining a target area in response to area information selected by the user in the scene map; The target scene is displayed in the target area.

9. The method according to claim 8, wherein Displaying the target scene in the target area includes: If the target scene includes multiple scene units, displaying the multiple scene units in the target area; If the target scene includes a scene unit, the target scene is displayed at the target position in response to the target position selected in the target area.

10. The method according to claim 1, wherein The displaying of the candidate scene graph includes: For the candidate scene graph, displaying an editing control and a selection control corresponding to the candidate scene graph in a first display area corresponding to the candidate scene graph; The method further comprises: In response to receiving a selection operation on the editing control, displaying a candidate scene graph corresponding to the selected target editing control in the editing interface; In response to the editing information input by the user in the editing interface, a new candidate scene graph is generated and displayed according to the editing information and the candidate scene graph corresponding to the target editing control.

11. The method according to claim 1, wherein The displaying of the candidate scene graph includes: Determine a presentation scene graph from the candidate scene graphs; Displaying the display scene graph; The method further comprises: When presenting the presentation scene graph, presenting a scene update control; In response to receiving a selection operation on the scene update control, determining a new presentation scene graph; The new presentation scene graph is presented.

12. The method according to claim 1, further comprising: displaying a scene editing control when displaying the candidate scene graph; In response to receiving a selection operation on the scene editing control, displaying a scene editing interface; In response to the user's editing operation in the scene editing interface, new scene description information is obtained, and the operation of generating a candidate scene graph of multiple candidate scenes based on the scene description information is returned to be executed.

13. The method according to claim 1, wherein Generating a candidate scene graph of multiple candidate scenes based on the scene description information includes: Constructing a prompt text according to the scene description information; The prompt text is input into a scene generation model, and a candidate scene graph of the multiple candidate scenes is obtained according to an output of the scene generation model, wherein the scene generation model is implemented based on a large language model.

14. A scene generation device, comprising: A receiving module is configured to receive scene description information of a target scene to be generated; a first generating module configured to generate a candidate scene graph of a plurality of candidate scenes based on the scene description information, wherein the candidate scene graph is a two-dimensional image; A first display module is configured to display the candidate scene graph; A second generating module is configured to generate voxel data corresponding to the selected target scene graph in response to receiving a user selection operation on the candidate scene graph; and The first processing module is configured to perform rendering based on the voxel data to obtain the target scene, wherein the target scene is a three-dimensional scene.

15. A computer-readable medium having a computer program stored thereon, wherein: When the program is executed by a processing device, the steps of the method according to any one of claims 1 to 13 are implemented.

16. An electronic device comprising: a storage device having a computer program stored thereon; as well as A processing device is configured to execute the computer program in the storage device to implement the steps of the method according to any one of claims 1 to 13.

17. A computer program product comprising a computer program, wherein When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 13 are implemented.

Citation Information

Patent Citations

  • Virtual scene generation method and device, readable medium and electronic equipment

    CN116416404A

  • Building scene rendering method and device and storage medium

    CN117008795A

  • Scene generation method and device, computer equipment and storage medium

    CN117274489A

  • Scene generation method and device, medium, equipment and program product

    CN118079378A