Virtual scene processing method, device and equipment, computer readable storage medium and computer program product
By editing text or picture requirements in the virtual scene generation interface and generating reference scene maps, the problem of time-consuming manual adjustment by users in the prior art is solved, and efficient personalized virtual scene generation is achieved.
Patent Information
- Application Number
- CN202510565794.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-29
- Publication Date
- 2025-08-05
AI Technical Summary
In the prior art, users need to manually adjust the three-dimensional virtual scene module, which is time-consuming and dependent on experience, and cannot meet personalized needs, and the generation complexity is limited by the richness of the module library.
Provide a virtual scene generation method, edit scene requirements through the requirements editing entrance, support text and image input, generate reference scene maps, and generate target virtual scenes according to user selection.
It improves the pertinence and accuracy of virtual scene generation, enhances user creative freedom, and simplifies the scene generation process.
Smart Images

Figure CN120428892A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to human-computer interaction technology, and in particular to a method, device, equipment, storage medium and program product for processing a virtual scene. Background Art
[0002] In the related technologies, most of the template-based tools are used to construct three-dimensional virtual scenes. These tools provide preset terrain modules (such as mountains, rivers, buildings, etc.), and users manually drag the modules into the scene in a "building block" manner, and repeatedly adjust the module position, rotation angle and scaling ratio to adapt to the overall layout; however, in this method, users need to make micro-adjustments to each module (such as the direction of trees, the curvature of roads, etc.), which is time-consuming and relies on empirical aesthetic design capabilities. In addition, the complexity of generating three-dimensional virtual scenes is limited by the richness of the prefabricated module library and cannot meet personalized needs. Summary of the Invention
[0003] The embodiments of the present application provide a method, device, electronic device, computer-readable storage medium, and computer program product for processing a virtual scene, which can specifically generate a virtual scene that meets user needs.
[0004] The technical solution of the embodiment of the present application is implemented as follows:
[0005] The present invention provides a method for processing a virtual scene, including:
[0006] Displaying a requirement editing portal in the virtual scene generation interface, wherein the requirement editing portal is used to edit the scene requirements required for generating the virtual scene;
[0007] In response to a requirement editing operation triggered based on the requirement editing portal, displaying the edited target scenario requirement, wherein the type of the target scenario requirement includes at least one of the following: text, image;
[0008] In response to a preview generation operation, displaying at least one reference scene graph generated based on the target scene requirement;
[0009] In response to a scene generation operation triggered based on the target scene graph, a target virtual scene generated based on the target scene graph is displayed.
[0010] The present invention provides a virtual scene processing device, including:
[0011] An entry display module is used to display a requirement editing entry in the virtual scene generation interface, wherein the requirement editing entry is used to edit the scene requirements required for generating the virtual scene;
[0012] A requirement display module, configured to display the edited target scenario requirement in response to a requirement editing operation triggered based on the requirement editing entry, wherein the type of the target scenario requirement includes at least one of the following: text, image;
[0013] a reference display module, configured to display at least one reference scene graph generated based on the target scene requirement in response to a preview generation operation;
[0014] The scene display module is used to display a target virtual scene generated based on the target scene graph in response to a scene generation operation triggered based on the target scene graph.
[0015] In the above scheme, the generation interface includes a requirement editing area and a scene display area; wherein, the requirement editing area is used to display the requirement editing entrance, the target scene requirement and the at least one reference scene graph, and the scene display area is used to display the target virtual scene.
[0016] In the above scheme, the demand editing area is a conversation area for the target account and the smart assistant to conduct a conversation. The demand display module is also used to display the edited target scenario requirements through a first conversation message sent to the smart assistant by the target account; correspondingly, the reference display module is also used to display at least one reference scene graph generated based on the target scenario requirements through a second conversation message replied to the target account by the smart assistant.
[0017] In the above scheme, the demand display module is also used to, when the demand editing entrance is a text editing box, display the text content based on the input of the text editing box in the text editing box in response to a text editing operation triggered based on the text editing box; and display the text content in response to a determination operation on the text content, and use the text content as the target scenario demand.
[0018] In the above scheme, the demand display module is also used to display the text content obtained by text conversion based on the recorded voice in response to the voice recording operation triggered by the voice recording icon when the demand editing entrance is a voice recording icon, and use the text content as the target scenario demand.
[0019] In the above scheme, the demand display module is also used to display a picture drawing interface in response to a triggering operation on the picture drawing entrance when the demand editing entrance is a picture drawing entrance, wherein the picture drawing interface includes a picture drawing canvas and a drawing toolbar, and the drawing toolbar includes multiple brushes, and brushes of different colors represent different scene elements; in response to a picture drawing operation triggered by a brush based on a target color, an outline picture of the target scene element drawn by the brush of the target color is displayed in the picture drawing canvas; in response to a determination operation on the outline picture, a target picture determined based on the outline picture is displayed, and the target picture is used as the target scene demand.
[0020] In the above scheme, the demand display module is also used to control the outline picture to be in a state to be published in response to a determination operation on the outline picture, and display the outline picture in a state to be published at the associated position of the picture drawing entrance; in response to a text editing operation triggered based on a text editing box, display the text description content based on the text editing box input in the text editing box; in response to a determination operation on the text description content, publish the text description content and the outline picture in a state to be published, and display the target picture including the text description content and the outline picture.
[0021] In the above scheme, the demand display module is also used to display the target image obtained by supplementing the outline image when the outline image meets the supplementary conditions; wherein, the supplementary conditions include at least one of the following: there is a shape to be supplemented in the outline image, and the associated elements of the target scene element are missing in the outline image, and the associated elements include at least one of the following: some elements in the target scene element, and environmental elements associated with the target scene element.
[0022] In the above solution, the demand display module is further used to display supplementary prompt information, and the supplementary prompt information is used to prompt the outline image to be supplemented; in response to the confirmation operation on the supplementary prompt information, the target image obtained by supplementing the outline image is displayed.
[0023] In the above scheme, the demand display module is also used to display supplementary prompt information and a supplementary editing area in response to a confirmation operation on the contour image; in response to a supplementary editing operation triggered in the supplementary editing area based on the supplementary prompt information, the target image obtained by supplementing the contour image based on the edited supplementary description content is displayed.
[0024] In the above scheme, the demand display module is also used to respond to a picture drawing operation triggered by a brush based on the target color, and when the picture drawing operation indicates that the initial picture of the target scene element drawn by the brush of the target color meets the supplementary condition, display the outline picture obtained by supplementing the initial picture; wherein, the supplementary condition includes at least one of the following: there is a shape to be supplemented in the initial picture, and the associated elements of the target scene element are missing in the initial picture, and the associated elements include at least one of the following: some elements in the target scene element, and environmental elements associated with the target scene element.
[0025] In the above scheme, the demand display module is also used to display a picture selection interface in response to a trigger operation on the picture import entrance when the demand editing entrance is a picture import entrance, and the picture selection interface includes multiple pictures to choose from; in response to a selection operation on a target picture, the target picture is displayed, and the target picture is used as the target scene demand.
[0026] In the above scheme, the demand display module is also used to control the target image to be in a pending release state in response to a selection operation on the target image; display the text description content based on the input in the text editing box in the text editing box; respond to a confirmation operation on the text description content, publish the text description content and the target image in the pending release state, and display the text description content and the target image as the target scene demand.
[0027] In the above scheme, after the display of at least one reference scene graph generated based on the target scene requirements, the device also includes: an adjustment module, used for at least one of the following: in response to an adjustment operation for the target scene requirements, updating and displaying the reference scene graph based on the adjusted target scene requirements; in response to a trigger operation for a variant control, updating the reference scene graph to another reference scene graph generated based on the target scene requirements, wherein the other reference scene graph is different from the reference scene graph.
[0028] In the above scheme, before displaying at least one reference scene graph generated based on the target scene requirement, the device also includes: a reference generation module, which is used to extract the semantic vector of the target scene requirement of the text type when the type of the target scene requirement includes text; initialize the noise image and the iteration time step, and extract the noise tensor of the noise image and the time step vector of the iteration time step; perform attention adjustment on the semantic vector based on the noise tensor to obtain an attention feature, and perform residual prediction based on the attention feature and the time step vector to obtain a noise residual for the noise image; and iteratively denoise the noise image based on the residual noise to obtain the reference scene graph.
[0029] In the above scheme, before displaying at least one reference scene graph generated based on the target scene requirement, the reference generation module is also used to, when the type of the target scene requirement includes pictures and texts, perform semantic region segmentation on the target scene requirement of the picture type to obtain a semantic mask graph, and perform feature extraction on the target scene requirement of the text type to obtain text prompt features; extract spatial condition features from the semantic heat map encoding corresponding to the semantic mask graph, and extract multi-scale spatial features of the target scene requirement of the picture type; fuse the spatial condition features and the multi-scale spatial features to obtain spatial fusion features; perform attention adjustment on the spatial fusion features based on the text prompt features to obtain attention features; and perform layered and refined synthesis on the semantic mask graph based on the attention features to obtain the reference scene graph.
[0030] In the above scheme, before displaying the target virtual scene generated based on the target scene graph, the device also includes: a scene generation module, which is used to perform scene semantic segmentation on the target scene graph to obtain a background mask and a foreground mask; based on the perspective of the target scene graph, the background mask is corrected for perspective to obtain a first scene depth map; for the foreground pixels in the foreground mask, the depth values of the foreground pixels are replaced by the depth values of the background areas surrounding the foreground pixels to obtain a second scene depth map; the first scene depth map and the second scene depth map are fused to obtain a target scene depth map, and the target virtual scene is constructed based on the target scene depth map.
[0031] An embodiment of the present application provides an electronic device, including:
[0032] a memory for storing computer-executable instructions or computer programs;
[0033] The processor is used to implement the virtual scene processing method provided in the embodiment of the present application when executing the computer executable instructions or computer program stored in the memory.
[0034] An embodiment of the present application provides a computer-readable storage medium storing computer-executable instructions or a computer program for implementing a virtual scene processing method provided in an embodiment of the present application when executed by a processor.
[0035] An embodiment of the present application provides a computer program product, including a computer program or computer-executable instructions. When the computer program or computer-executable instructions are executed by a processor, the virtual scene processing method provided by the embodiment of the present application is implemented.
[0036] The embodiments of the present application have the following beneficial effects:
[0037] By applying the embodiments of the present application, when generating a virtual scene, the user can edit the required target scene requirements based on the requirement editing entrance, and can also preview at least one reference scene graph generated based on the target scene requirements, and can select the target scene graph from at least one reference scene graph according to actual needs. In this way, the terminal can generate the corresponding target virtual scene based on the target scene graph selected by the user. During the entire virtual scene generation process, the user is allowed to define the scene requirements in the form of text or pictures according to actual needs, thereby increasing the user's freedom to create virtual scenes; in addition, since the text description allows the user to abstractly describe the scene elements, the picture description can provide visual priors such as composition and color, which can better understand the user's intentions, thereby facilitating the generation of virtual scenes that meet the user's needs, thereby improving the pertinence and accuracy of virtual scene generation. BRIEF DESCRIPTION OF THE DRAWINGS
[0038] Figure 1 1 is a schematic diagram of the architecture of a virtual scene processing system 100 provided in an embodiment of the present application;
[0039] Figure 2 is a structural diagram of an electronic device 500 provided in an embodiment of the present application;
[0040] Figure 3 1 is a flow chart of a method for processing a virtual scene provided in an embodiment of the present application;
[0041] Figure 4 This is a demand interaction diagram provided by an embodiment of the present application;
[0042] Figure 5 This is a demand interaction diagram provided by an embodiment of the present application;
[0043] Figure 6 This is a schematic diagram of editing the scenario requirements provided by the embodiment of the present application;
[0044] Figure 7 This is a schematic diagram of the scenario requirements provided by the embodiment of the present application;
[0045] Figure 8 This is a schematic diagram of editing the scenario requirements provided by the embodiment of the present application;
[0046] Figure 9 This is a schematic diagram of adjusting the scene graph provided in an embodiment of the present application;
[0047] Figure 10 is a schematic diagram of a display of a virtual scene provided in an embodiment of the present application;
[0048] Figure 11 Schematic diagram of the architecture of a virtual scene processing system provided in an embodiment of the present application;
[0049] Figure 12 This is a diagram of the implementation framework of the client provided in the embodiment of the present application;
[0050] Figure 13 This is a diagram of the implementation framework of the server provided in the embodiment of the present application;
[0051] Figure 14 This is a diagram of the implementation framework of the algorithm provided in the embodiment of the present application;
[0052] Figure 15 This is a training diagram of the scene semantic segmentation model provided in an embodiment of the present application. DETAILED DESCRIPTION
[0053] In order to make the purpose, technical solutions and advantages of this application clearer, the application will be further described in detail below with reference to the accompanying drawings. The described embodiments should not be regarded as limiting this application. All other embodiments obtained by ordinary technicians in this field without making creative work are within the scope of protection of this application.
[0054] It is understandable that in the embodiments of the present application, when user information and other related data (such as user trigger operations, team attributes or role characteristics, etc.) are involved, when the embodiments of the present application are applied to specific products or technologies, user permission or consent is required, and the collection, use and processing of relevant data must comply with relevant laws, regulations and standards.
[0055] In the following description, reference is made to “some embodiments”, which describes a subset of all possible embodiments, but it will be understood that “some embodiments” may be the same subset or different subsets of all possible embodiments and may be combined with each other without conflict.
[0056] In the following description, the terms "first\second..." are merely used to distinguish similar objects and do not represent a specific ordering of the objects. It is understandable that "first\second..." can be interchanged with a specific order or sequence where permitted, so that the embodiments of the present application described herein can be implemented in an order other than that illustrated or described herein.
[0057] In the embodiments of the present application, the term "module" or "unit" refers to a computer program or a part of a computer program that has a predetermined function and works together with other related parts to achieve a predetermined goal, and can be implemented in whole or in part by using software, hardware (such as processing circuits or memories) or a combination thereof. Similarly, a processor (or multiple processors or memories) can be used to implement one or more modules or units. In addition, each module or unit can be part of an overall module or unit that includes the function of the module or unit.
[0058] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the art to which this application pertains. The terms used herein are for the purpose of describing the embodiments of this application only and are not intended to limit this application.
[0059] Before further explaining the embodiments of the present application in detail, the nouns and terms involved in the embodiments of the present application are explained. The nouns and terms involved in the embodiments of the present application are subject to the following interpretations.
[0060] 1) Client: An application running in a terminal to provide various services, such as a game client, a virtual scene generation client, etc.
[0061] 2) In response, it is used to indicate the conditions or states on which the executed operations depend. When the dependent conditions or states are met, one or more operations executed can be real-time or have a set delay. Unless otherwise specified, there is no restriction on the order in which the multiple operations executed are executed.
[0062] 3) A virtual scene is a virtual scene displayed (or provided) when an application is running on a terminal. The virtual scene can be a simulation of the real world, a semi-simulation and semi-fictitious virtual environment, or a purely fictitious virtual environment. The virtual scene can be any of a two-dimensional virtual scene, a 2.5-dimensional virtual scene, or a three-dimensional virtual scene. The embodiments of the present application do not limit the dimensions of the virtual scene. For example, the virtual scene may include the sky, land, ocean, etc. The land may include environmental elements such as deserts and cities, and users can control virtual objects to move in the virtual scene.
[0063] The embodiments of the present application provide a method, device, electronic device, computer-readable storage medium and computer program product for processing a virtual scene, which can generate a virtual scene that meets the needs of the user in a targeted manner. The exemplary application of the electronic device provided by the embodiment of the present application is described below. The electronic device provided by the embodiment of the present application can be implemented as various types of user terminals such as laptop computers, tablet computers, desktop computers, set-top boxes, mobile devices (for example, mobile phones, portable music players, personal digital assistants, dedicated messaging devices, portable gaming devices), smart phones, smart speakers, smart watches, smart TVs, vehicle-mounted terminals, augmented reality (AR, Augmented Reality) devices and virtual reality (VR, Virtual Reality) devices, and can also be implemented as servers. Below, an exemplary application when the device is implemented as a terminal will be described.
[0064] See also Figure 1 , Figure 1 This is an architectural diagram of a virtual scene processing system 100 provided in an embodiment of the present application. To support an exemplary application, terminals (terminal 400-1 and terminal 400-2 are shown as examples) are connected to a server 200 via a network 300. The network 300 may be a wide area network or a local area network, or a combination of the two.
[0065] In some embodiments, the terminal is provided with a client having a virtual scene processing function, and the server 200 is a background server corresponding to the client, which can be an independent physical server, or a server cluster or distributed system composed of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, content delivery networks (CDNs), and big data and artificial intelligence platforms. The terminal and the server can be directly or indirectly connected via wired or wireless communication, which is not limited in the embodiments of the present application.
[0066] In actual applications, the terminal displays a requirement editing entry in the virtual scene generation interface. In response to a requirement editing operation triggered by the requirement editing entry, the terminal displays the edited target scene requirement, where the target scene requirement type includes at least one of the following: text or image. In response to a preview generation operation, the terminal sends a preview request to server 200. Based on the preview request, server 200 generates at least one reference scene graph based on the target scene requirement and returns the generated reference scene graph to the terminal for display. When a user selects a target scene graph, the terminal sends a scene creation request to server 200 in response to a scene generation operation triggered by the target scene graph. Based on the scene creation request, server 200 generates a target virtual scene based on the target scene graph and returns the generated target virtual scene to the terminal for display.
[0067] See also Figure 2 , Figure 2 This is a structural diagram of an electronic device 500 provided in an embodiment of the present application, with the electronic device 500 as an example. Figure 1 For example, Figure 2 The electronic device 500 shown includes: at least one processor 510, a memory 550, at least one network interface 520 and a user interface 530. The various components in the electronic device 500 are coupled together via a bus system 540. It is understood that the bus system 540 is used to achieve connection and communication between these components. In addition to including a data bus, the bus system 540 also includes a power bus, a control bus and a status signal bus. However, for the sake of clarity, the bus system 540 is not shown in FIG. Figure 2 Various buses are labeled as bus system 540 .
[0068] The processor 510 can be an integrated circuit chip with signal processing capabilities, such as a general-purpose processor, a digital signal processor (DSP), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc., where the general-purpose processor can be a microprocessor or any conventional processor, etc.
[0069] The memory 550 includes a volatile memory or a non-volatile memory, or may include both volatile and non-volatile memories. The non-volatile memory may be a read-only memory (ROM), and the volatile memory may be a random access memory (RAM). The memory 550 described in the embodiments of the present application is intended to include any suitable type of memory. The memory 550 may optionally include one or more storage devices physically remote from the processor 510.
[0070] In some embodiments, the memory 550 can store data to support various operations, examples of which include programs, modules, and data structures, or a subset or superset thereof, as exemplified below.
[0071] The operating system 551 includes system programs for processing various basic system services and performing hardware-related tasks, such as a framework layer, a core library layer, a driver layer, etc., which are used to implement various basic businesses and process hardware-based tasks; the network communication module 552 is used to reach other electronic devices via one or more (wired or wireless) network interfaces 520. Exemplary network interfaces 520 include: Bluetooth, Wireless Compatibility Certification (WiFi), and Universal Serial Bus (USB), etc.
[0072] In some embodiments, the virtual scene processing device provided in the embodiments of the present application can be implemented in software. The virtual scene processing device provided in the embodiments of the present application can be provided as various software embodiments, including various forms including applications, software, software modules, scripts or codes. Figure 2 A processing device 555 of a virtual scene stored in a memory 550 is shown, which can be software in the form of a program and a plug-in, and includes a series of modules, including an entry display module 5551, a demand display module 5552, a reference display module 5553 and a scene display module 5554. These modules are logical, and therefore can be arbitrarily combined or further split according to the functions implemented. The functions of each module will be explained below.
[0073] In other embodiments, the apparatus provided in the embodiments of the present application may be implemented in hardware. As an example, the apparatus provided in the embodiments of the present application may be a processor in the form of a hardware decoding processor, which is programmed to execute the virtual scene processing method provided in the embodiments of the present application. For example, the processor in the form of a hardware decoding processor may be one or more application-specific integrated circuits (ASICs), digital signal processors (DSPs), programmable logic devices (PLDs), complex programmable logic devices (CPLDs), field-programmable gate arrays (FPGAs), or other electronic components.
[0074] In some embodiments, the terminal or server can implement the processing method of the virtual scene provided by the embodiment of the present application by running various computer executable instructions or computer programs. For example, computer executable instructions can be commands, machine instructions or software instructions at the microprogram level. The computer program can be a native program or software module in the operating system; it can be a local (Native) application (APPlication, APP), that is, a program that needs to be installed in the operating system to run, such as instant messaging APP, live broadcast APP; it can also be a small program that can be embedded in any APP, that is, a program that can be run only by downloading it to a browser environment. In short, the above-mentioned computer executable instructions can be instructions in any form, and the above-mentioned computer program can be an application, module or plug-in in any form.
[0075] As mentioned above, the virtual scene processing method provided in the embodiment of the present application can be implemented by various types of electronic devices, for example, Figure 1 The terminal and the server 200 in the embodiment can be executed by either one of them separately or by Figure 1 The terminal and server 200 in the embodiment cooperate to execute. Figure 1 The terminal in the embodiment of the present application independently executes the processing method of the virtual scene provided by the embodiment of the present application as an example. Figure 3 , Figure 3 This is a flow chart of the virtual scene processing method provided by the embodiment of the present application, which will be combined with Figure 3 The steps shown are explained.
[0076] Step 101: The terminal displays a requirement editing entry in a virtual scene generation interface.
[0077] Among them, the requirement editing entrance is used to edit the scene requirements required to generate the virtual scene. The types of requirement editing entrances can be various, such as a text editing box and a voice input entrance for editing text-type scene requirements, or a picture drawing entrance and a picture import entrance for editing picture-type scene requirements. The generation interface for generating virtual scenes may include a requirement editing area and a scene display area, wherein the virtual editing area is used to display the relevant content required to generate the virtual scene, such as the requirement editing entrance, the scene requirements edited by the user, and the reference scene graph for the user to select, which are all displayed in the requirement editing area, that is, the content displayed in the virtual editing area includes the key factors such as the basic properties, environment settings, object configuration and interaction logic required to generate the virtual scene. The scene display area is used to display the virtual scene generated by real-time rendering, and supports multi-view display and interactive operations, such as zooming, rotating, and even AR / VR preview, so that users can clearly edit and view the scene requirements, as well as view the reference scene graph and finally generate the virtual scene.
[0078] Step 102: In response to a requirement editing operation triggered by a requirement editing portal, the edited target scenario requirement is displayed, wherein the type of the target scenario requirement includes at least one of the following: text, image.
[0079] In actual applications, users can edit the target scene requirements required to generate virtual scenes based on the demand editing entrance. The target scene requirements can be scene requirements of text type edited by the user, or scene requirements of picture type edited by the user, or scene requirements that combine text type and picture type. They will be explained one by one below.
[0080] In some embodiments, when the requirement editing area is a conversation area for the target account and the smart assistant to conduct a conversation, the edited target scenario requirement can be displayed in the following manner: a first conversation message sent by the target account to the smart assistant displays the edited target scenario requirement; correspondingly, at least one reference scene graph generated based on the target scenario requirement can be displayed in the following manner: a second conversation message replied to the target account by the smart assistant displays at least one reference scene graph generated based on the target scenario requirement.
[0081] In actual applications, the target account is the account that triggers the relevant operations of generating a virtual scene, and the smart assistant is an artificial intelligence object that provides virtual scene generation services. The target account can convey what kind of virtual scene the user wants to generate and the scene requirements required to generate the virtual scene through a conversation with the smart assistant. That is, after the user completes editing the target scene requirements, the target scene requirements are conveyed to the smart assistant in the form of a conversation message; correspondingly, after receiving the conversation message carrying the target scene requirements, the smart assistant generates a reference scene graph based on the target scene requirements, and also feeds the reference scene graph back to the target account in the form of a conversation message.
[0082] For example, see Figure 4 , Figure 4 This is a demand interaction diagram provided by an embodiment of the present application. Taking the target scene demand type as text as an example, after the user completes the text editing of the target scene demand, a conversation message 402 carrying the target scene demand input by the user is displayed in the conversation interface 401 between the target account and the smart assistant. After receiving the conversation message 402 carrying the target scene demand, the smart assistant regards the user's sending operation for the target scene demand as a preview generation operation, and automatically feeds back the reference scene graph generated based on the target scene demand to the target account in the form of a conversation message, that is, displays a conversation message 403 carrying the reference scene graph and corresponding prompt information, so that the user can select a suitable target scene graph for the generation of the virtual scene.
[0083] It should be noted that in applications that use conversational messaging to convey user scenario requirements and reference scenario graphs fed back by intelligent assistants, a progressive input paradigm can be used to gradually input user scenario requirements. That is, users do not need to input all scenario requirements at once, but can use step-by-step interactions to gradually improve their scenario requirements. For example, see Figure 5 , Figure 5 This is a demand interaction diagram provided by an embodiment of the present application. The target account inputs a scene requirement (such as "I need a cyberpunk-style city") 501, and the intelligent assistant displays a reference scene graph 502 generated based on the scene requirement 501. The target account adds a scene requirement (such as "add holographic billboards and flying cars") 503 on the basis of the scene requirement 501. The intelligent assistant adjusts the relevant scene elements in the reference scene graph 503, such as adding architectural elements of billboards and traffic elements of flying cars, and displays the adjusted reference scene graph 504. The user can gradually improve the scene requirements on this basis. Based on the scene requirements gradually completed by the user, the intelligent assistant provides a reference scene graph that meets the scene requirements of each step, so that the user can select from it to generate a target scene graph for the virtual scene.
[0084] Through the above method, the scene requirements and the corresponding reference scene graph are conveyed in the form of conversation messages, allowing the scene requirements to be improved once or gradually to generate the required virtual scene, thereby improving the pertinence and accuracy of virtual scene generation.
[0085] In some embodiments, the terminal can display the edited target scenario requirements in response to the requirement editing operation triggered by the requirement editing entry in the following manner: when the requirement editing entry is a text editing box, in response to the text editing operation triggered by the text editing box, the text content based on the input of the text editing box is displayed in the text editing box; in response to the determination operation on the text content, the text content is displayed, and the text content is used as the target scenario requirement.
[0086] In actual applications, when the requirement editing entry is a text edit box, users can use natural language to edit the scene requirements for the virtual scene to be generated in the text edit box, and the input and edited text content is displayed in real time in the text edit box. When the text edit box displays the text content entered by the user, the terminal supports the input of hierarchical descriptions in multiple formats (such as mountains [altitude: 2000 meters]), and can perform semantic understanding or intent analysis of the text content entered by the user in real time, and highlight the keyword part in the text content (which reflects the user's intention). For example, if the text content entered by the user is "a forest in the sun, with a stream and a red cabin", the keywords such as "forest", "stream" and "cabin" used to describe the physical objects are highlighted in red, and the modifiers such as "under the sun" and "red" that modify the physical objects are highlighted in blue to indicate the user's intention. Grammar prompts can also be displayed to automatically complete some content based on the user's intention.
[0087] See also Figure 6 , Figure 6 It is an editing diagram of the scenario requirements provided in an embodiment of the present application. A text editing box 602 is displayed in the requirement editing area 601. When the user enters the edited text content in the text editing box 602, the terminal displays the text content edited by the user in the text editing box 602 in real time (such as "I want to generate a scenario, which requires xxxxxxx"). When the user clicks the send control, the terminal responds to the trigger operation, publishes the text content, and displays the published target content in the requirement editing area 601 as the target scenario requirement 603. After the terminal receives the target scenario requirement, the user can trigger a preview generation operation. The terminal responds to the preview generation operation and displays multiple reference scene graphs 604 generated based on the target scenario requirement.
[0088] Through the above method, users can input the required scenario requirements in natural language according to actual needs, which improves the convenience of inputting scenario requirements.
[0089] In some embodiments, the terminal can respond to the requirement editing operation triggered by the requirement editing entry in the following manner and display the edited target scenario requirement: when the requirement editing entry is a voice input icon, in response to the voice input operation triggered by the voice input icon, the text content obtained by text conversion based on the input voice is displayed, and the text content is used as the target scenario requirement.
[0090] In actual applications, the demand editing entrance can also be a voice input icon. The user uses the voice input icon to input voice to describe the virtual scene he wants to generate. The voice input by the user can obtain the corresponding text content through text conversion. The voice input by the user can be a simple description, such as the text obtained by voice conversion can be: "A sunny beach", or it can be a more detailed description, such as the text obtained by voice conversion: "A tropical beach with sand, coconut trees and blue sky".
[0091] Similarly, this method of conveying scene requirements through voice can also be used to gradually improve scene requirements through step-by-step interaction. In this way, users can express the virtual scene they want to generate in one or more voice messages based on their actual needs, improving the convenience and efficiency of scene requirement editing and enhancing user creative freedom.
[0092] In some embodiments, the terminal can display the edited target scene requirements in response to the requirement editing operation triggered by the requirement editing entrance in the following manner: in a case where the requirement editing entrance is a picture drawing entrance, in response to the triggering operation for the picture drawing entrance, a picture drawing interface is displayed, wherein the picture drawing interface includes a picture drawing canvas and a drawing toolbar, and the drawing toolbar includes multiple brushes, and brushes of different colors represent different scene elements; in response to the picture drawing operation triggered by the brush based on the target color, an outline picture of the target scene element drawn by the brush of the target color is displayed in the picture drawing canvas; in response to the determination operation for the outline picture, a target picture determined based on the outline picture is displayed, and the target picture is used as the target scene requirement.
[0093] In actual applications, when the requirement editing entrance is a picture drawing entrance, the user can draw the scene requirements of the virtual scene that he wants to generate through the picture drawing entrance. When the user triggers the picture drawing entrance, the terminal responds to the trigger operation, displays the picture drawing interface, and displays the drawing toolbar and the picture drawing canvas in the picture drawing interface, wherein the drawing toolbar includes a variety of brushes of different colors, and brushes of different colors represent different scene elements. In addition, the brush can also be associated with a setting control, through which the transparency of the color indicated by the brush can be adjusted. The user can select a brush of one color or multiple colors of brushes according to actual needs to edit the outline picture of the scene element corresponding to the selected brush, and can directly use the outline picture as the target scene requirement, or use the target picture obtained by supplementing or expanding the outline picture as the target scene requirement.
[0094] See also Figure 7 , Figure 7This is a schematic diagram of the scene requirement drawing provided by an embodiment of the present application. When a user triggers the image drawing entry 701, the terminal responds to the trigger operation by displaying the image drawing interface, and displays a drawing toolbar 702 and a picture drawing canvas 703 in the picture drawing interface. The drawing toolbar includes a variety of brushes of different colors. Different colored brushes represent different scene elements, such as a red brush for land, a green brush for forest, and a gray brush for water. The user can view more brush colors by sliding the brush wheel. When the user selects a brush of a target color to draw an image on the picture drawing canvas 703, the image is the outline image of the target scene element corresponding to the target color brush. For example, if the target color is red, the image drawn on the picture drawing canvas using a red brush is the outline image of the target scene element "land". In this way, the user can select the desired brush color to draw the outline image of the corresponding scene element. When the user completes the drawing, the user can click the send button to display the drawn outline image as the target image or the target image expanded based on the outline image as the target scene requirement in the requirement editing area. In this way, the terminal can return the reference scene graph 704 generated based on the target image based on the target scene requirements.
[0095] Through the above method, users can draw and express the virtual scenes to be generated in the form of simple sketches according to actual needs, which improves the diversity and convenience of obtaining scene requirements and enhances the user's creative freedom.
[0096] In some embodiments, in response to a determination operation on the outline picture, a target picture drawn based on the outline picture can be displayed in the following manner: in response to a determination operation on the outline picture, the outline picture is controlled to be in a to-be-published state, and the outline picture in the to-be-published state is displayed at the associated position of the picture drawing entrance; in response to a text editing operation triggered based on the text editing box, the text description content based on the input in the text editing box is displayed in the text editing box; in response to a determination operation on the text description content, the text description content and the outline picture in the to-be-published state are published, and the target picture including the text description content and the outline picture is displayed.
[0097] See also Figure 8 , Figure 8This is an editing diagram of the scenario requirements provided by an embodiment of the present application. After the user completes the outline image 801 by drawing it with a brush, when a confirmation operation is triggered for the outline image 801 (such as triggering a completion button), the terminal responds to the trigger operation, controls the outline image 801 to be published, and displays the outline image 801 in the pending publication state at the associated position (such as above) of the image drawing entrance 800. When the user triggers a text editing operation in the text editing box, the terminal responds to the text editing operation and displays the text description content (such as a summer island, a dense forest) 802 entered by the user in the text editing box. When the user triggers a confirmation operation for the text description content (such as triggering a send button), the terminal responds to the confirmation operation and publishes the outline image 801 and the text description content 802 together, and displays a target image 803 including the text description content and the outline image, or displays the text description content and the outline image in the form of a conversation message, that is, this multimodal description consisting of text description content of text type and outline image of picture type is used as the target scenario requirement. In this way, the terminal can return the reference scene graph 804 generated based on the text description content and the outline image based on the target scene requirements.
[0098] Through the above method, users can describe the scene requirements of the virtual scene to be generated in text description or drawing according to actual needs, thereby improving the diversity and convenience of obtaining scene requirements and enhancing the user's creative freedom.
[0099] In some embodiments, the target image determined based on the contour image can be displayed in the following manner: when the contour image meets the supplementary conditions, the target image obtained by supplementing the contour image is displayed; wherein the supplementary conditions include at least one of the following: the presence of shapes to be supplemented in the contour image, and associated elements of the target scene elements are missing in the contour image, and the associated elements include at least one of the following: partial elements in the target scene elements, and environmental elements associated with the target scene elements.
[0100] In actual applications, after the user triggers the confirmation operation for the contour image, the terminal detects whether the contour image meets the supplementary conditions, and if the contour image meets the supplementary conditions, automatically supplements the contour image accordingly to obtain the target image, and displays the target image as the target scene requirement. Among them, there can be multiple supplementary conditions. For example, the supplementary condition can be that there is a shape to be completed in the contour image. In this case, the supplement is the supplement of shape and structure. For example, the user draws a curve representing a water area, but does not finish it (that is, the curve is not closed), that is, the water boundary in the contour image is incomplete. In this case, the contour image is considered to meet the supplementary conditions, and the terminal automatically completes the unclosed curve into a complete water boundary. In addition to the above-mentioned curve closure, there may be other shapes that need to be completed, such as straight lines, polygons, etc. If the user draws a partial geometric shape (such as an unclosed triangle edge, a discontinuous line or a blurred outline), the terminal automatically completes the complete shape; for example, if the user draws a semicircular arc, the system infers it as an "arch" and completes the top arc structure.
[0101] In addition, the supplementary conditions can also be partial elements of the target scene elements that are missing in the outline picture. For example, if the user draws a part of the target scene element (an object), the terminal infers and supplements the other part of the target scene element based on the context to obtain a picture of the complete object. That is, the user draws the local features of the object, and the terminal predicts and infers the complete object based on the drawn content and fills in the details, or uses a geometric algorithm to automatically close the shape to maintain proportion and symmetry. For example, the user draws "a dot", and the terminal infers it as a "street light" and completes the lamp pole and halo effect. The user draws "two parallel lines", and the terminal recognizes it as a "road" and completes the lane line and roadside vegetation. The user draws a wavy line, and the terminal recognizes it as a "river" and completes the details of the opposite shoreline and water flow.
[0102] Furthermore, supplementary conditions can also involve missing environmental elements associated with the target scene elements in the outline image. For example, if a user draws an outline image of the sun, the terminal automatically adds clouds or a glowing effect to the outline image. Alternatively, the terminal can supplement missing parts of the outline image based on scene logic. For example, if a user draws an outline image of a wall, the terminal automatically completes the outline image with "roof" or "doors and windows." Alternatively, the terminal can automatically fill the outline image with appropriate colors and textures. For example, if a user sketches an object, the terminal automatically fills the corresponding outline image with appropriate colors or textures, such as green for "grass" and blue for "sky." Alternatively, the terminal can supplement the outline image with necessary scene elements, such as path connections, based on logical requirements. For example, if a user draws an outline image describing "a bridge but no piers," the terminal automatically completes the pier structure. Alternatively, if a user draws an outline image describing "a river flowing through a mountain," the terminal automatically corrects the outline image to "a river winding around a mountain" or adds a "waterfall" transition.
[0103] Moreover, in the case of multimodal input, such as combining text description and pictures as target scene requirements, the user enters "a forest", and the terminal automatically completes the trees, animals and other elements in the outline picture, or dynamically adjusts the relevant scene elements in the outline picture according to time. For example, if the time described in the outline picture is daytime, the sun is completed in the outline picture; if the time described in the outline picture is night, the moon and stars are completed in the outline picture.
[0104] Through the above-mentioned automatic supplementation method, the outline image can be automatically supplemented, reducing the trouble of manual adjustment and improving the efficiency of scene requirement editing.
[0105] In some embodiments, the target image obtained by supplementing the outline image can be displayed in the following manner: displaying supplementary prompt information, the supplementary prompt information is used to prompt the outline image to be supplemented; in response to a confirmation operation on the supplementary prompt information, displaying the target image obtained by supplementing the outline image.
[0106] In actual applications, in order to ensure a balance between the accuracy of supplementation and the user's creative freedom, the terminal can also provide supplementary prompt information before automatically supplementing the outline image to prompt the user what aspects of the outline image will be supplemented, and the outline image will only be supplemented after the user agrees. The content automatically supplemented by the terminal can also be modified to avoid excessive automation leading to results that do not meet the user's expectations.
[0107] In some embodiments, a target image determined based on the contour image can be displayed in response to a determination operation on the contour image in the following manner: in response to a determination operation on the contour image, supplementary prompt information and a supplementary editing area are displayed; in response to a supplementary editing operation triggered in the supplementary editing area based on the supplementary prompt information, a target image obtained by supplementing the contour image based on the edited supplementary description content is displayed.
[0108] Here, in order to ensure a balance between the accuracy of supplementation and the user's creative freedom, after the user triggers the confirmation operation on the outline picture, the terminal detects whether the outline picture meets the supplementation conditions, and if the outline picture meets the supplementation conditions, displays supplementary prompt information and supplementary editing area, that is, the supplementary prompt information prompts the user which parts of the outline picture need to be supplemented, and the supplementary editing area provides the user with a supplementary entrance. The user can edit the supplementary description content in the supplementary editing area. In this way, the terminal can supplement the outline picture based on the supplementary description content edited by the user to obtain the target picture.
[0109] Through the above methods, not only the accuracy and completeness of scenario requirements are improved, but also the initiative and freedom of users to create based on scenario requirements are improved.
[0110] In some embodiments, the terminal may respond to a picture drawing operation triggered by a brush based on the target color by displaying an outline picture of the target scene element drawn by the brush of the target color in the picture drawing canvas in the following manner: in response to a picture drawing operation triggered by a brush based on the target color, when the picture drawing operation indicates that the initial picture of the target scene element drawn by the brush of the target color meets the supplementary condition, the outline picture obtained by supplementing the initial picture is displayed; wherein the supplementary condition includes at least one of the following: the existence of a shape to be supplemented in the initial picture, the missing associated elements of the target scene element in the initial picture, and the associated elements include at least one of the following: partial elements in the target scene element, and environmental elements associated with the target scene element.
[0111] In actual applications, when drawing a picture on a picture drawing canvas with a brush, if the initial picture is drawn, the terminal detects whether the drawn initial picture meets the supplementary conditions, and if the initial picture meets the supplementary conditions, the terminal automatically supplements the initial picture accordingly to obtain a contour picture, and uses the contour picture as the target scene requirement. Among them, there can be multiple supplementary conditions. For example, the supplementary condition can be that there is a shape to be completed in the drawn initial picture. In this case, the supplement is the supplement of shape and structure. For example, the user draws a curve representing a water area, but does not finish it (that is, the curve is not closed), that is, the water boundary in the drawn initial picture is incomplete. In this case, it is considered that the drawn initial picture meets the supplementary conditions, and the terminal automatically completes the unclosed curve into a complete water boundary. In addition to the above-mentioned closed curves, there may be other shapes that need to be completed, such as straight lines, polygons, etc. If the user draws a partial geometric shape (such as an unclosed triangle edge, a discontinuous line or a blurred outline), the terminal automatically completes the complete shape; for example, if the user draws a semicircular arc, the system infers it as an "arch" and completes the top arc structure.
[0112] In addition, the supplementary conditions may also be partial elements of the target scene elements that are missing in the initial image that has been drawn. For example, if the user draws a part of the target scene element (an object), the terminal infers and supplements the other part of the target scene element based on the context to obtain the outline image of the complete object. That is, the user draws the local features of the object, and the terminal predicts and infers the complete object based on the drawn content and fills in the details, or uses a geometric algorithm to automatically close the shape to maintain proportion and symmetry. For example, the user draws "a dot", and the terminal infers it as a "street light" and completes the lamp pole and halo effect. The user draws "two parallel lines", which the terminal recognizes as a "road" and completes the lane line and roadside vegetation. The user draws a wavy line, which the terminal recognizes as a "river" and completes the shoreline and water flow details on the opposite bank.
[0113] Furthermore, the supplementary condition can also be that the initial image is missing environmental elements associated with the target scene element. For example, if the user draws an initial image of the "sun," the terminal automatically adds clouds or a glowing effect to the initial image. Alternatively, the terminal can supplement missing parts of the initial image based on scene logic. For example, if the user draws an initial image of a "wall," the terminal automatically completes the "roof" or "doors and windows" in the outline image. Alternatively, the terminal can automatically fill the outline image with appropriate colors and textures. For example, if the user sketches an object, the terminal automatically fills the corresponding initial image with appropriate colors or textures, such as green for "grass" and blue for "sky." Alternatively, the terminal can supplement the initial image with necessary scene elements, such as path connections, based on logical requirements. For example, if the user's initial image depicts "a bridge but no piers," the terminal automatically completes the pier structure in the initial image. For another example, if the user's initial image depicts "a river flowing through a mountain," the terminal automatically corrects the initial image to "a river winding around a mountain" or adds a "waterfall" transition.
[0114] Moreover, in the case of multimodal input, such as combining text description and pictures as target scene requirements, the user enters "a forest", and the terminal automatically completes the trees, animals and other elements in the initial picture, or dynamically adjusts the relevant scene elements in the initial picture according to time. For example, if the time described in the initial picture is daytime, the sun will be completed in the initial picture; if the time described in the initial picture is night, the moon and stars will be completed in the initial picture.
[0115] It should be noted that in order to ensure a balance between the accuracy of the supplement and the user's creative freedom, the terminal can also provide supplementary prompt information before automatically supplementing the initial picture to prompt the user what aspects of the supplement will be made to the initial picture, and the initial picture will only be supplemented after the user agrees. The content automatically supplemented by the terminal can also be modified to avoid excessive automation leading to results that do not meet user expectations.
[0116] Through the above-mentioned automatic supplementation method, the initial picture drawn by the user can be automatically supplemented, reducing the trouble of manual adjustment and improving the efficiency of scene requirement editing.
[0117] In some embodiments, the terminal can respond to the requirement editing operation triggered by the requirement editing entrance in the following manner to display the edited target scene requirement: when the requirement editing entrance is a picture import entrance, in response to the triggering operation on the picture import entrance, a picture selection interface is displayed, and the picture selection interface includes multiple pictures to choose from; in response to the selection operation on the target picture, the target picture is displayed, and the target picture is used as the target scene requirement.
[0118] In actual applications, when the requirement editing entrance is a picture import entrance, the user can also import ready-made pictures through the terminal as target scene requirements, where these ready-made pictures can be pictures stored in the system album of the terminal, or pictures taken recently, or pictures downloaded from the Internet. The embodiment of this application does not limit the source of the ready-made pictures.
[0119] Through the above method, users can conveniently use ready-made pictures as target scene requirements, which improves the efficiency of editing target scene requirements.
[0120] In some embodiments, in response to a selection operation on the target image, the target image can be displayed and used as a target scene requirement in the following manner: in response to a selection operation on the target image, the target image is controlled to be in a pending release state; the text description content based on the input in the text edit box is displayed in the text edit box; in response to a confirmation operation on the text description content, the text description content and the target image in the pending release state are released, and the text description content and the target image are displayed as target scene requirements.
[0121] In actual application, when the user triggers a selection operation for the target image, the terminal responds to the selection operation, controls the target image to be in a pending state, and displays the target image in a pending state at the associated position of the image import entry (such as above); when the user triggers a text editing operation in the text editing box, the terminal responds to the text editing operation, and displays the text description content entered by the user in the text editing box; when the user triggers a determination operation for the text description content (such as triggering a send button), the terminal responds to the determination operation, publishes the target image and the text description content together, and displays the text description content and the target image as target scene requirements, or displays the text description content and the target image as target scene requirements in the form of a conversation message, that is, a multimodal description consisting of a text description content of a text type and a target image of a picture type is used as a target scene requirement. In this way, the terminal can return a reference scene graph generated based on the text description content and the target image based on the target scene requirements.
[0122] Through the above method, users can describe the scene requirements of the virtual scene to be generated in text description or drawing according to actual needs, thereby improving the diversity and convenience of obtaining scene requirements and enhancing the user's creative freedom.
[0123] Step 103: In response to the preview generation operation, display at least one reference scene graph generated based on the target scene requirements.
[0124] In actual applications, after obtaining the target scene requirements for generating a virtual scene, the terminal generates a corresponding reference scene graph based on the target scene requirements and displays the generated reference scene graph in the scene editing area for the user to view and select. When the user selects the desired target scene graph, the final 3D virtual scene is generated based on the selected target scene graph. The following describes the specific implementation of generating a reference scene graph based on the target scene requirements.
[0125] In some embodiments, before displaying at least one reference scene graph generated based on the target scene requirements, the reference scene graph can be generated based on the target scene requirements in the following manner: when the type of the target scene requirements includes text, extracting the semantic vector of the target scene requirements of the text type; initializing the noise image and the iteration time step, and extracting the noise tensor of the noise image and the time step vector of the iteration time step; performing attention adjustment on the semantic vector based on the noise tensor to obtain an attention feature, and performing residual prediction based on the attention feature and the time step vector to obtain a noise residual for the noise image; and iteratively denoising the noise image based on the residual noise to obtain a reference scene graph.
[0126] In actual applications, when the target scene requirement is textual, taking the textual target scene requirement of "a forest under the sun, with streams and red cabins" as an example, when the terminal generates a reference scene graph based on the textual target scene requirement, it first performs text cleaning on the textual target scene requirement, such as removing redundant symbols, unifying capitalization, correcting spelling errors, etc. Then, it extracts key information from it, such as identifying the core elements in the target scene requirement (such as "sunshine", "forest", "stream", "red cabin") through a natural language processing model, and parsing attributes (such as color, position relationship). Then, it performs intent parsing on the extracted key information to obtain the corresponding semantic vector, that is, converting the text into a high-dimensional semantic vector indicating the scene elements required to generate the virtual scene. The high-dimensional semantic vector is used as the conditional input of the diffusion model to ensure that the generated reference scene graph is aligned with the text semantics.
[0127] Then, the noisy image is initialized in the latent space and the iterative timestep is initialized. The iterative timestep defines the process of gradually recovering the desired reference scene graph from the noisy image. The noisy image is initialized by randomly sampling a noise tensor in the latent space (typically smaller than the output reference scene graph, such as 64×64). A subsequent diffusion model is used to gradually denoise the image according to the timestep vector, transforming the noisy image into structured information that matches the text (i.e., the reference scene graph). During the back-diffusion (i.e., iterative denoising) process, the diffusion model iteratively corrects the noisy image to gradually generate clear latent features. Specifically, at each step, the current noise tensor, semantic vector, and timestep vector are input into the diffusion model. The diffusion model then adjusts the semantic vector based on the noise tensor to obtain an attention feature. A cross-attention mechanism associates the semantic vector of the target scene requirement with the noise tensor of the noisy image, ensuring that the generated scene graph meets the target scene requirement. Residual prediction is then performed based on the attention feature and timestep vector. This predicts the portion of the current noisy image that is unrelated to the semantic vector (i.e., the target scene requirement), calculates the noise residual, and updates the latent tensor of the noise tensor based on the residual noise, gradually reducing the noise in the noisy image.
[0128] Finally, the denoised latent tensor is decoded to restore the low-dimensional latent variable to a high-resolution two-dimensional image (such as 512×512 pixels). Optional post-processing steps (such as super-resolution reconstruction, brightness adjustment, contrast adjustment, etc.) can be used to enhance the image clarity and ensure that the final reference scene graph meets the lighting description (such as "sunlight") required by the target scene.
[0129] Through the above method, since the diffusion model is good at generating detailed and natural images, by using the semantic vector required by the target scene as the conditional input of the diffusion model, the content of the generated reference scene graph (such as object position and color) can be accurately controlled to meet the requirements of the target scene; at the same time, the latent space operation greatly reduces the amount of calculation, improves the generation efficiency of the reference scene graph, and supports real-time generation.
[0130] In some embodiments, before displaying at least one reference scene graph generated based on the target scene requirements, a reference scene graph can be generated based on the target scene requirements in the following manner: when the types of target scene requirements include pictures and texts, semantic region segmentation is performed on the target scene requirements of the picture type to obtain a semantic mask graph, and feature extraction is performed on the target scene requirements of the text type to obtain text prompt features; spatial condition features are extracted from the semantic heat map encoding corresponding to the semantic mask graph, and multi-scale spatial features of the target scene requirements of the picture type are extracted; the spatial condition features and the multi-scale spatial features are fused to obtain spatial fusion features; attention adjustment is performed on the spatial fusion features based on the text prompt features to obtain attention features; and the semantic mask graph is layered and refined based on the attention features to obtain a reference scene graph.
[0131] In practical applications, when the target scene requirements are text and pictures, for the target scene requirements of the picture type (i.e., a picture), the edge enhancement processing can be performed on the target scene requirements of the picture type first, such as extracting the main outline of the target scene requirements of the picture type through the edge detection algorithm to eliminate the interference of jitter lines; then, multi-scale spatial feature extraction is performed on the target scene requirements of the picture type after edge enhancement to obtain multi-scale spatial features, and semantic region segmentation is performed on the target scene requirements of the picture type after edge enhancement to obtain a semantic mask map. The semantic mask map is marked with category labels of each segmented area for identifying the scene category therein (such as houses, trees, etc.). The semantic mask is then encoded and converted to a semantic heatmap encoding. This involves converting the index color of each pixel in the semantic mask into a corresponding scene category identification matrix, generating a multi-channel (e.g., 32-channel) semantic heatmap encoding. The resulting semantic heatmap encoding is then fed into a network control model (e.g., ControlNet). After the zero-convolution layer of the network control model, spatial conditional features are extracted using multiple residual blocks (each containing a 3×3 convolution, GroupNorm, and SiLU). The spatial conditional features are then fused with multi-scale spatial features to generate spatial fusion features. Since spatial conditional features reflect the mapping relationship between scene categories and locations, such as the regional distribution corresponding to the "house" category, and multi-scale spatial features are related to texture and structure generation rules (e.g., the generation pattern of wooden wall materials), the two are fused to achieve semantically guided detail generation.
[0132] Finally, the attention of the spatial fusion feature is adjusted based on the text prompt feature to obtain the attention feature, that is, the spatial fusion feature is aligned with the text prompt through the cross-attention mechanism. The adjusted attention feature can reflect the constraints imposed on the specific area annotated by the semantic mask (such as the area corresponding to the "house" logo), such as superimposing the gradient correction term on the feature map of the corresponding spatial position to ensure that the sampling result of the area meets the material characteristics of the text prompt (such as wood grain wall). In this way, the reference scene graph obtained by layered refinement and synthesis of the semantic mask based on the attention feature can ensure that the generated reference scene graph meets the target scene requirements input by the user.
[0133] After the terminal generates a reference scene graph that meets the requirements of the target scene, the generated reference scene graph can be directly displayed in the requirements editing area for the user to view and select, or the generated reference scene graph can be displayed in the requirements editing area when the user triggers the preview generation operation (such as triggering the "Generate Preview" button).
[0134] Through the above method, users are allowed to define scene requirements in the form of text and pictures according to actual needs, which increases the freedom of users to create virtual scenes; in addition, since text descriptions allow users to abstractly describe scene elements, and picture descriptions can provide visual priors such as composition and color, the multimodal requirements input can better understand user intentions, which is conducive to generating virtual scenes that meet user needs, improving the pertinence and accuracy of virtual scene generation.
[0135] In some embodiments, after the terminal displays at least one reference scene graph generated based on the target scene requirements, it may also perform at least one of the following processing: in response to an adjustment operation for the target scene requirements, updating and displaying the reference scene graph based on the adjusted target scene requirements; in response to a trigger operation for a variant control, updating the reference scene graph to another reference scene graph generated based on the target scene requirements, where the other reference scene graph is different from the reference scene graph.
[0136] In actual applications, after the terminal displays the reference scene graph generated based on the target scene requirements to the user, the user can select the target scene graph to generate the corresponding target virtual scene, and can also update one or all of the reference scene graphs. For example, see Figure 9 , Figure 9 It is a schematic diagram of the adjustment of the scene graph provided in an embodiment of the present application. By triggering the update control 901, all reference scene graphs displayed to the user can be updated with one click; after selecting one of the reference scene graphs (such as reference scene graph 903 is selected), the entire content of the selected reference scene graph 903 can be previewed through the preview control, and the variant control (such as re-variant) 902 can be triggered to update the selected reference scene graph 903 to other reference scene graphs generated based on the target scene requirements.
[0137] It is understandable that the updated reference scene graph is still generated based on the target scene requirements. In this way, users can update the reference scene graph according to their own needs until they are satisfied with the reference scene graph, which improves the diversity of the reference scene graph and the convenience of updating.
[0138] Step 104: In response to the scene generation operation triggered based on the target scene graph, display the target virtual scene generated based on the target scene graph.
[0139] In some embodiments, before displaying the target virtual scene generated based on the target scene graph, the terminal may generate the target virtual scene based on the target scene graph in the following manner: perform scene semantic segmentation on the target scene graph to obtain a background mask and a foreground mask; based on the perspective of the target scene graph, perform perspective correction on the background mask to obtain a first scene depth map; for the foreground pixels in the foreground mask, replace the depth values of the foreground pixels with the depth values of the surrounding background areas of the foreground pixels to obtain a second scene depth map; fuse the first scene depth map and the second scene depth map to obtain a target scene depth map, and construct the target virtual scene based on the target scene depth map.
[0140] In actual applications, after generating each reference scene graph, the terminal can extract the initial depth map relative to the camera from each reference scene graph (including the depth information of each pixel in the scene). The initial depth map mainly has the following two problems: perspective deviation: the perspective of the reference scene graph is not a completely top-down view, and the camera angle is inconsistent; foreground interference: the depth information of the foreground targets (such as trees, buildings, etc.) in the reference scene graph will be retained in the initial depth map and needs to be excluded from the terrain; therefore, after generating the reference scene graph (including the target scene graph), the reference scene graph is subjected to scene semantic segmentation through the scene semantic segmentation model to obtain a background mask and a foreground mask; based on the perspective of the reference scene graph, the background mask is perspective-corrected to obtain a first scene depth map; for the foreground pixels in the foreground mask, the depth values of the background areas surrounding the foreground pixels are used to replace the depth values of the foreground pixels to obtain a second scene depth map; the first scene depth map and the second scene depth map are fused to obtain the target scene depth map (i.e., terrain height map) corresponding to the reference scene graph.
[0141] It should be noted that in actual applications, after the terminal generates multiple reference scene graphs based on the target scene requirements, it can perform the above-mentioned scene semantic segmentation, perspective correction, depth map fusion and other operations on each reference scene graph to obtain the target scene depth map (i.e., terrain height map) corresponding to each reference scene graph; the user can also select the target scene graph from multiple reference scene graphs, and perform the above-mentioned scene semantic segmentation, perspective correction, depth map fusion and other operations on the target scene graph to obtain the target scene depth map (i.e., terrain height map) corresponding to the target scene graph. This can reduce the amount of calculation and thus improve processing efficiency.
[0142] After obtaining the target scene depth map (i.e., terrain height map) corresponding to the target scene graph, the required target virtual scene can be rendered based on the target scene depth map (i.e., terrain height map), thereby improving the accuracy of generating the virtual scene.
[0143] See also Figure 10 , Figure 10 This is a display diagram of the virtual scene provided in an embodiment of the present application. After the user selects the target scene requirement 1001, if the button for generating the scene is triggered, the terminal responds to the trigger operation, generates a corresponding three-dimensional target virtual scene based on the target scene requirement, and displays the generated target virtual scene 1002 in the scene display area.
[0144] In actual applications, when users input their target scene requirements via voice, they can express their emotional state (such as happiness, sadness, excitement) through the timbre or rhythm of their voice. These emotional states influence the generation of the virtual scene, and the terminal can adjust the atmosphere and elements of the generated virtual scene based on these emotions. When users generate the target scene requirements by drawing an image, the rendering style of the target virtual scene is adjusted based on the user's drawing style (such as cartoon or realistic).
[0145] Through the above method, when generating a virtual scene, the user can edit the required target scene requirements based on the requirement editing entrance, and can also preview at least one reference scene graph generated based on the target scene requirements, and can select the target scene graph from at least one reference scene graph according to actual needs. In this way, the terminal can generate the corresponding target virtual scene based on the target scene graph selected by the user. During the entire virtual scene generation process, the user is allowed to define the scene requirements in the form of text and pictures according to actual needs, thereby increasing the user's freedom to create virtual scenes; in addition, since the text description allows the user to abstractly describe the scene elements, the picture description can provide visual priors such as composition and color, which can better understand the user's intentions, and thus facilitate the generation of virtual scenes that meet user needs, thereby improving the pertinence and accuracy of virtual scene generation.
[0146] Below, an exemplary application of the embodiment of the present application in a practical application scenario will be described.
[0147] See also Figure 11 , Figure 11 This is a schematic diagram of the architecture of the virtual scene processing system provided in an embodiment of the present application. The system consists of three modules: client, server, and algorithm. The functions of each module are as follows:
[0148] Server: Used to execute multiple agent services, including a main agent and multiple sub-agents. The main agent's main function is to analyze user intent and split the user intent into corresponding sub-tasks; the sub-agents are responsible for completing specific sub-tasks.
[0149] Client: Responsible for obtaining user input data (supporting natural language text input and stick figure image input), calling the corresponding functional interfaces of the server and algorithm based on the results returned by the proxy, and creating a virtual scene based on the metadata generated by the algorithm.
[0150] Algorithm side: Generate a scene preview image (i.e., the above-mentioned reference scene image) based on the user's input data (i.e., the above-mentioned target scene requirements) for the user to select a suitable scene preview image, and generate the metadata required to create a virtual scene based on the scene preview image selected by the user (i.e., the above-mentioned target scene image) and return it to the client.
[0151] Specifically: the user describes the scene he wants to create in natural language or simple drawings in the client. After the client obtains the user's description (the context data input by the user, i.e., the target scene requirements), it sends a preview request with the target scene requirements to the main agent in the server. The main agent analyzes the user's intention based on the target scene requirements carried in the preview request. When it determines that the user needs to create the target virtual scene, it returns the sub-agent name and parameters for creating the target virtual scene. The client uses the parameters returned by the main agent to request the creation of the sub-agent corresponding to the target virtual scene. When the sub-agent analyzes the input data and determines that the scene preview image (i.e., the reference scene image) should be created first, it generates a prompt (i.e., Prompt) for creating the scene preview image based on the target scene requirements, and returns the prompt to the client. The client uses Prompt to call the algorithm-side interface to generate a scene preview image. After the user selects a suitable scene preview image (i.e., the target scene image mentioned above) from the scene preview image, the scene creation request is sent to the main agent. The main agent analyzes the user's intention based on the target scene requirements carried in the scene creation request. When it determines that the user wants to create the target virtual scene, it returns the sub-agent name and parameters corresponding to the target virtual scene. The client uses the parameters returned by the main agent to request the sub-agent. The sub-agent analyzes the target scene requirements input by the user and determines that the target virtual scene should be created based on the scene preview image selected by the user. It returns the interface name and parameters for creating the target virtual scene to the client. The client uses the interface name and parameters returned by the sub-agent to call the algorithm-side interface. The algorithm-side interface generates scene data corresponding to the target virtual scene, such as terrain height data and landform data of different regions, and returns it to the client. The client uses scene data such as terrain height data and landform data of different regions to generate the final target virtual scene.
[0152] Next, Figure 11 Each module in is described below.
[0153] See also Figure 12 , Figure 12This is an implementation framework diagram of the client provided in an embodiment of the present application. The client mainly includes the following modules: a navigation webpage (WebBrowser) plug-in, a programmatic content generation (PCG, Procedural Content Generation) plug-in, and a network module and a terrain module in a virtual engine (such as a UE engine). Among them, the WebBrowser plug-in is mainly used to implement the overall page layout and user interaction logic, including the user's natural language input and simple drawing input, and calls the server and algorithm interfaces according to the user input. The PCG plug-in is mainly used to implement the creation of scene terrain and vegetation generation. The specific functions include terrain generation: various complex terrains such as mountains, plains, rivers, etc. are generated through the terrain height data and landform data of different regions returned by the algorithm interface. Vegetation generation: Generate corresponding vegetation such as trees, grass, etc. based on the landform data, terrain features of different regions returned by the algorithm, and predefined vegetation generation rules. The network module is mainly responsible for processing the communication between the WebBrower plug-in and the server and algorithm ends, and calling the functional interfaces of the server and algorithm ends. The terrain module is mainly responsible for handling the dynamic changes of terrain and the dynamic rendering of landform effects. When the terrain height data and landform data of different areas change, the corresponding terrain and landforms can be dynamically created.
[0154] In an embodiment of the present application, the PCG plug-in converts the grayscale value of each pixel into the geometric height of the 3D terrain based on the terrain height map (grayscale image, i.e. the target scene depth map mentioned above), constructs terrain features such as mountains, plains, and rivers, and optimizes the transition effect of the terrain through a smoothing algorithm to make it more natural. At the same time, corresponding elements are dynamically generated based on the semantic information in the scene semantic segmentation map (such as vegetation, water bodies, roads and other areas). For example, different types of vegetation are generated in the vegetation area according to predefined rules, and the density is adjusted according to the area size and user needs; a flat water surface is generated in the water area, and a dynamic ripple effect is added; and a three-dimensional road model that conforms to the terrain is generated in the road area. The terrain module is responsible for processing the dynamic changes and rendering effects of the terrain. When the input terrain height data or semantic segmentation data changes, the geometric structure and texture effects of the terrain can be updated in real time.
[0155] See also Figure 13 , Figure 13It is an implementation framework diagram of the server provided in the embodiment of the present application, which mainly includes a configuration service, a main agent, multiple sub-agents and an AI model. The configuration service is used to dynamically update the configuration of each agent, including a description of the functions supported by the agent, a prompt template for requesting the AI model, etc. The main agent is responsible for analyzing the user's intention according to the target scene requirements input by the user, and splitting it into sub-tasks. Each sub-agent is responsible for completing a specific task, and the functions of each sub-agent are as follows: The scene creation sub-agent is used to complete the function of creating a scene, including generating a scene preview image based on the user's text description / simple drawing, generating a three-dimensional virtual scene based on the scene preview image, etc.; the function of the preview image prompt sub-agent is to generate appropriate prompts based on the target scene requirements input by the user, describing the virtual scene that the user wants to create; the function of the prompt translation sub-agent is to translate the preview image prompt into English.
[0156] The AI model is the final processing unit, which generates return results based on the prompts configured by each agent and the target scenario requirements input by the user. The return results may be ordinary text content (for example, when the AI determines that the user's question can be solved by a simple text answer, it will return a clear text reply, such as an explanation of a certain function, operation suggestions or status confirmation. The user may ask about a specific parameter or function of the scenario generation, and the AI model can directly return an explanatory text answer). It may also be a function interface call. The client parses the return result. If it is a function interface call, the corresponding function interface is called.
[0157] See also Figure 14 , Figure 14 This is an implementation framework diagram of the algorithm side provided in the embodiment of the present application. The algorithm side mainly includes the following modules: a two-dimensional scene preview image generation module, a terrain height map extraction module, and a scene semantic parsing module. Among them, the two-dimensional scene preview image generation module takes the scene text description or simple sketch (i.e., the above-mentioned target scene requirements) input by the user as input and a two-dimensional scene preview image as output. The scene preview image generation module uses a diffusion model to realize the function of generating a two-dimensional scene preview image from text, and combines the control network (such as ControlNet) model to realize precise control of the scene layout through fixed simple sketches.
[0158] The text description of the scene (i.e., the target scene requirements of the above text type) is converted into a semantic vector through a text encoder. The semantic vector is input into the diffusion model as a control signal, providing the diffusion model with the semantic information required to generate the scene preview image, such as the style, elements and layout requirements of the scene. The control network model intervenes in the forward reasoning process of the diffusion model to convert the user's stick figure input (i.e., the target scene requirements of the picture type) into a structured control signal, which directly acts on the intermediate layer output of the diffusion model, thereby accurately guiding the layout and structure of the generated image. For example, when the user draws the outline of a lake in a stick figure, the control network model ensures that the diffusion model places the lake in the corresponding position when generating the image, and keeps it consistent with the grass. Figure 1 Exquisite shape and proportion.
[0159] This solution is based on the Flux pre-trained diffusion model and uses low-rank adaptation (LoRA) technology for customized training (inserting low-rank matrices A and B into the key layers of the Flux model and initializing these matrices. The LoRA module is fine-tuned using the prepared training dataset. During training, only the low-rank matrices A and B are updated, while the other weights of the pre-trained Flux model remain unchanged) to improve the controllability and stability of the generated scene content and camera perspective.
[0160] In order to minimize the information loss caused by scene self-occlusion, the closer the target perspective of the image generated by this module is to the top-down perspective, the more conducive it is for the subsequent model to extract the three-dimensional information of the scene. In addition, the pre-trained model is prone to generating situations where the proportions of different elements are not coordinated and the layout is unreasonable, and the generated scene is difficult to match the text well. In order to optimize these inherent problems of the model, the embodiment of the present application uses a semi-automatic data generation and screening process to construct a high-quality scene graph dataset, and its construction process is as follows:
[0161] Scene description generation: Use a large language model to automatically generate diverse scene descriptions (i.e., the target scene requirements mentioned above). For example, set some basic scene categories or themes (such as "forest lake," "snow mountain town," "desert oasis," etc.), and use the large language model to generate diverse scene descriptions related to these themes. For example, for "forest lake," the large language model can generate the description: "A tranquil lake surrounded by a dense pine forest, with several small hills in the distance."
[0162] Scene graph generation: Use the generated scene description to generate a two-dimensional scene preview image. Use the pre-trained diffusion model and input the scene description generated in the previous step into the diffusion model. The diffusion model generates the corresponding two-dimensional scene preview image based on the text description.
[0163] Perspective filtering: Calculates the perspective of the scene preview image and automatically filters out data with non-top-down perspectives. Use computer vision techniques (such as OpenCV or deep learning models) to analyze the generated scene preview image and estimate the perspective of the scene preview image by detecting horizontal lines, vanishing points, or other geometric features in the scene preview image. For example, the closer the horizontal line is to the bottom of the scene preview image, the closer the perspective is to a top-down view. If a scene preview image is detected to have a non-top-down perspective (for example, with a large tilt angle or non-orthogonal projection), the scene preview image is automatically filtered out.
[0164] Multimodal initial screening: Use a multimodal model to preliminarily screen unreasonable scene preview images. Using a pre-trained multimodal model (such as CLIP or Flamingo), the generated scene preview image and the corresponding scene description are input into the multimodal model. The model evaluates whether the content in the scene preview image matches the text description and whether the layout of the elements in the scene preview image is reasonable. Based on the model's score or confidence, scene preview images with unreasonable layout or disproportionate element proportions are screened out. For example, if the lake in the scene preview image is too large or the mountain is too small, the model may mark the scene preview image as unreasonable.
[0165] Manual screening: Manually score the scene previews and filter out unqualified scene previews. Based on the manual scoring, filter out the scene previews with low scores and retain the high-quality scene previews as the final scene previews.
[0166] In order to provide more intuitive scene layout control, an embodiment of the present application provides a network control model for stick figures. Users can draw layouts of categories such as grass, snow, water, and roads through stick figures. The training data set of the network control model is constructed through a semi-automated process, in which the real annotation data of the stick figures are generated by the scene parsing model. The scene parsing model automatically analyzes the target scene requirements of the input image type, identifies the distribution of different elements (such as grass, water, roads, etc.), and generates corresponding annotation data. These annotation data serve as the real layout information of the stick figures, and together with the stick figures drawn by the user, constitute a training data set. In this way, the network control model can learn the mapping relationship between stick figures and scene layouts, that is, there is a one-to-one correspondence between stick figures and annotation data.
[0167] The terrain height map extraction module takes a 2D scene preview image as input and a terrain height map (i.e., the target scene depth map mentioned above). The terrain height map extraction module uses a pre-trained depth prediction model (such as Monodepth2 or DPT) to extract an initial depth map relative to the camera from the 2D preview scene image (containing the depth information of each pixel in the scene). However, the initial depth map has the following two main problems: perspective bias: the perspective of the input preview scene image is not a complete bird's-eye view, and the camera angle is inconsistent; foreground interference: the depth information of foreground objects (such as trees and buildings) in the preview image is retained in the initial depth map and needs to be excluded from the terrain. Therefore, after obtaining the 2D scene preview image (including the target scene image), this module uses the scene semantic segmentation model to perform scene semantic segmentation on the scene preview image to obtain a background mask and a foreground mask. Based on the perspective of the scene preview image, the background mask is perspective-corrected to obtain a first scene depth map. For foreground pixels in the foreground mask, the depth values of the surrounding background area are used to replace the foreground pixel depth values to obtain a second scene depth map. The first scene depth map and the second scene depth map are fused to obtain the target scene depth map.
[0168] Taking the background mask as a water mask as an example, the water mask is a binary image or data layer used to extract and mark the water area. It is usually composed of 0 and 1, where 1 represents the water area and 0 represents the non-water area. This module obtains the water mask fitting horizontal plane equation and uses it to correct the overall depth map. For example, according to the water mask, the corresponding pixel coordinates (x i ,y i ) and depth value z i , the fitting horizontal plane equation is: z=ax+by+c, and the deviation from the actual terrain height can be expressed as: min a,b,c ∑ i (z i -(ax i +by i +c)) 2 , adjust the overall depth value to: Z`=Z-(ax+by+c).
[0169] For the foreground mask obtained by the scene semantic segmentation model, the position of the foreground target (such as trees and buildings) is marked, and the foreground pixel depth is replaced by the nearest background depth. For example, if a pixel is marked as a foreground target, it is replaced by the depth value of the surrounding background pixels, and the foreground depth is excluded by smoothing the transition area through bilinear filtering to avoid obvious boundaries in the depth map.
[0170] After these two steps of post-processing, the depth results are normalized and can be used as a terrain depth map or a terrain height map.
[0171] The scene semantic segmentation model takes a two-dimensional scene preview image as input and a scene semantic mask as output. During training, the scene semantic segmentation module uses a custom-trained semantic segmentation model to segment the scene layout and each category of targets, and outputs a semantic segmentation mask. The scene segmentation dataset construction and model training process are shown in the figure below: In order to train a customized scene semantic segmentation model, the embodiment of the present application generates scene graph data in batches by combining a large language model and a scene graph generation model, and uses an unsupervised scene segmentation model to pre-segment it. After manual screening of the pre-segmentation results, they can be used for preliminary training of the segmentation model. The initial training dataset can be relatively small, such as about one hundred images. Through iterative optimization of the model and data, a semantic segmentation model with good accuracy and stability can eventually be trained.
[0172] See also Figure 15 , Figure 15 This is a training diagram of the scene semantic segmentation model provided by the embodiment of the present application. During training, a large language model is used to generate a variety of scene text descriptions, and then a scene graph generation module is used to create a random scene graph based on the text description; an unsupervised scene segmentation model is used to pre-segment the scene graph to generate preliminary semantic segmentation results; the quality of the data set is ensured by manual screening, and accurate pre-segmentation results and scene Figure 1 The training dataset is composed of a few datasets; a customized semantic segmentation model is trained. Initially, the dataset is small, but through continuous iteration to optimize model parameters and expand the dataset, a segmentation model that is both accurate and stable is eventually trained. This model can output high-quality semantic segmentation masks to identify different landform areas in the scene, such as water surfaces, roads, and vegetation, providing key terrain information for the construction of 3D virtual scenes. That is, after obtaining the terrain height map, the required 3D virtual scene is constructed based on the terrain height map.
[0173] That is, the embodiment of the present application utilizes a control network model to analyze the layout of stick figures, and generates a two-dimensional scene preview image with a top-down perspective optimization through a diffusion model + LoRA fine-tuning technology, thereby solving the problems of perspective deviation and element proportion imbalance in traditional generation technologies; subsequently, based on the scene preview image selected by the user, the main agent dynamically disassembles the task flow, coordinates the sub-agent and the algorithm-end module to complete the depth map correction (using the water mask to fit the horizontal plane equation to eliminate perspective interference), semantic analysis (customized segmentation model trained by semi-automatic data sets) and terrain generation, and finally generates a three-dimensional virtual scene that meets the user's intentions and is interactive through dynamic rendering through the PCG plug-in.
[0174] Through the above-mentioned approach, the present application embodiment provides a multi-modal driven dynamic task decomposition content scene generation system. The core of this system is to achieve an end-to-end closed loop from user intent understanding to high-precision three-dimensional scene generation through dual-channel input of natural language and stick figures combined with an agent dynamic collaboration architecture. The present application embodiment breaks through the limitations of existing UGC tools with a single input method and rigid generation process. Through multi-stage generation optimization and agent collaboration mechanism, it significantly improves the accuracy and controllability of scene construction while lowering the user operation threshold.
[0175] So far, the method for processing virtual scenes provided by the embodiments of the present application has been described in combination with the exemplary application and implementation of the electronic device provided by the embodiments of the present application. The following will continue to describe the processing scheme for implementing virtual scenes by cooperating with each module in the virtual scene processing device 555 provided by the embodiments of the present application.
[0176] An entry display module 5551 is used to display a requirement editing entry in the virtual scene generation interface, where the requirement editing entry is used to edit the scene requirements required for generating the virtual scene; a requirement display module 5552 is used to display the edited target scene requirements in response to a requirement editing operation triggered based on the requirement editing entry, where the type of the target scene requirements includes at least one of the following: text, image; a reference display module 5553 is used to display at least one reference scene graph generated based on the target scene requirements in response to a preview generation operation; a scene display module 5554 is used to display a target virtual scene generated based on the target scene graph in response to a scene generation operation triggered based on the target scene graph.
[0177] In some embodiments, the generation interface includes a requirement editing area and a scene display area; wherein, the requirement editing area is used to display the requirement editing entrance, the target scene requirement and the at least one reference scene graph, and the scene display area is used to display the target virtual scene.
[0178] In some embodiments, the demand editing area is a conversation area for the target account and the smart assistant to conduct a conversation. The demand display module is also used to display the edited target scenario requirements through a first conversation message sent to the smart assistant by the target account; accordingly, the reference display module is also used to display at least one reference scene graph generated based on the target scenario requirements through a second conversation message replied to the target account by the smart assistant.
[0179] In some embodiments, the demand display module is also used to, when the demand editing entry is a text editing box, display the text content based on the input in the text editing box in the text editing box in response to a text editing operation triggered based on the text editing box; and display the text content in response to a determination operation on the text content, and use the text content as the target scenario demand.
[0180] In some embodiments, the demand display module is also used to display the text content obtained by text conversion based on the recorded voice in response to a voice recording operation triggered by the voice recording icon when the demand editing entrance is a voice recording icon, and use the text content as the target scenario demand.
[0181] In some embodiments, the demand display module is also used to display a picture drawing interface in response to a triggering operation on the picture drawing entrance when the demand editing entrance is a picture drawing entrance, wherein the picture drawing interface includes a picture drawing canvas and a drawing toolbar, and the drawing toolbar includes multiple brushes, and brushes of different colors represent different scene elements; in response to a picture drawing operation triggered by a brush based on a target color, an outline picture of the target scene element drawn by the brush of the target color is displayed in the picture drawing canvas; in response to a determination operation on the outline picture, a target picture determined based on the outline picture is displayed, and the target picture is used as the target scene demand.
[0182] In some embodiments, the demand display module is also used to control the outline picture to be in a to-be-published state in response to a determination operation on the outline picture, and display the outline picture in a to-be-published state at an associated position of the picture drawing entrance; in response to a text editing operation triggered based on a text editing box, display the text description content based on the text editing box input in the text editing box; in response to a determination operation on the text description content, publish the text description content and the outline picture in a to-be-published state, and display a target picture including the text description content and the outline picture.
[0183] In some embodiments, the demand display module is also used to display the target image obtained by supplementing the outline image when the outline image meets the supplementary conditions; wherein the supplementary conditions include at least one of the following: there is a shape to be supplemented in the outline image, and the associated elements of the target scene element are missing in the outline image, and the associated elements include at least one of the following: some elements in the target scene element, and environmental elements associated with the target scene element.
[0184] In some embodiments, the demand display module is further used to display supplementary prompt information, where the supplementary prompt information is used to prompt the outline image to be supplemented; in response to a confirmation operation on the supplementary prompt information, a target image obtained by supplementing the outline image is displayed.
[0185] In some embodiments, the demand display module is also used to display supplementary prompt information and a supplementary editing area in response to a determination operation on the contour image; and in response to a supplementary editing operation triggered in the supplementary editing area based on the supplementary prompt information, display a target image obtained by supplementing the contour image based on the edited supplementary description content.
[0186] In some embodiments, the demand display module is also used to respond to a picture drawing operation triggered by a brush based on a target color, and when the picture drawing operation indicates that the initial picture of the target scene element drawn by the brush of the target color meets the supplementary condition, display the outline picture obtained by supplementing the initial picture; wherein the supplementary condition includes at least one of the following: there is a shape to be supplemented in the initial picture, and the associated elements of the target scene element are missing in the initial picture, and the associated elements include at least one of the following: some elements in the target scene element, and environmental elements associated with the target scene element.
[0187] In some embodiments, the demand display module is also used to display a picture selection interface in response to a trigger operation on the picture import entrance when the demand editing entrance is a picture import entrance, and the picture selection interface includes multiple pictures to choose from; in response to a selection operation on a target picture, display the target picture and use the target picture as the target scene demand.
[0188] In some embodiments, the demand display module is also used to control the target image to be in a pending state in response to a selection operation on the target image; display the text description content based on the input in the text editing box in the text editing box; and publish the text description content and the target image in a pending state in response to a confirmation operation on the text description content, and display the text description content and the target image as the target scene demand.
[0189] In some embodiments, after the display of at least one reference scene graph generated based on the target scene requirement, the device further includes: an adjustment module, used for at least one of the following: in response to an adjustment operation for the target scene requirement, updating and displaying the reference scene graph based on the adjusted target scene requirement; in response to a trigger operation for a variant control, updating the reference scene graph to another reference scene graph generated based on the target scene requirement, wherein the other reference scene graph is different from the reference scene graph.
[0190] In some embodiments, before displaying at least one reference scene graph generated based on the target scene requirement, the device also includes: a reference generation module for extracting a semantic vector of the target scene requirement of the text type when the type of the target scene requirement includes text; initializing a noise image and an iteration time step, and extracting a noise tensor of the noise image and a time step vector of the iteration time step; performing attention adjustment on the semantic vector based on the noise tensor to obtain an attention feature, and performing residual prediction based on the attention feature and the time step vector to obtain a noise residual for the noise image; and iteratively denoising the noise image based on the residual noise to obtain the reference scene graph.
[0191] In some embodiments, before displaying at least one reference scene graph generated based on the target scene requirement, the reference generation module is further used to, when the type of the target scene requirement includes pictures and texts, perform semantic region segmentation on the target scene requirement of the picture type to obtain a semantic mask graph, and perform feature extraction on the target scene requirement of the text type to obtain text prompt features; extract spatial condition features from the semantic heat map encoding corresponding to the semantic mask graph, and extract multi-scale spatial features of the target scene requirement of the picture type; fuse the spatial condition features and the multi-scale spatial features to obtain spatial fusion features; perform attention adjustment on the spatial fusion features based on the text prompt features to obtain attention features; and perform layered and refined synthesis on the semantic mask graph based on the attention features to obtain the reference scene graph.
[0192] In some embodiments, before displaying the target virtual scene generated based on the target scene graph, the device also includes: a scene generation module, which is used to perform scene semantic segmentation on the target scene graph to obtain a background mask and a foreground mask; based on the perspective of the target scene graph, the background mask is corrected for perspective to obtain a first scene depth map; for the foreground pixels in the foreground mask, the depth values of the foreground pixels are replaced by the depth values of the background areas surrounding the foreground pixels to obtain a second scene depth map; the first scene depth map and the second scene depth map are fused to obtain a target scene depth map, and the target virtual scene is constructed based on the target scene depth map.
[0193] The present invention provides a computer program product comprising a computer program or computer-executable instructions stored in a computer-readable storage medium. A processor of an electronic device reads the computer-executable instructions from the computer-readable storage medium and executes the computer-executable instructions, causing the electronic device to perform the virtual scene processing method described in the present invention.
[0194] The embodiment of the present application provides a computer-readable storage medium storing computer-executable instructions, wherein the computer-executable instructions or computer programs are stored. When the computer-executable instructions or computer programs are executed by a processor, the processor will execute the virtual scene processing method provided by the embodiment of the present application, for example, Figure 3 The processing method of the virtual scene is shown.
[0195] In some embodiments, the computer-readable storage medium may be a memory such as FRAM, ROM, PROM, EPROM, EEPROM, flash memory, magnetic surface storage, optical disk, or CD-ROM; or various devices including one or any combination of the above memories.
[0196] In some embodiments, computer-executable instructions may be in the form of a program, software, software module, script, or code, written in any form of programming language (including compiled or interpreted languages, or declarative or procedural languages), and may be deployed in any form, including as a stand-alone program or as a module, component, subroutine, or other unit suitable for use in a computing environment.
[0197] As an example, computer-executable instructions may, but need not, correspond to a file in a file system, may be stored as part of a file that stores other programs or data, such as, for example, in one or more scripts in a HyperText Markup Language (HTML) document, in a single file dedicated to the program in question, or in multiple coordinating files (e.g., files storing one or more modules, subroutines, or code portions).
[0198] By way of example, computer-executable instructions may be deployed to be executed on one electronic device, or on multiple electronic devices located at one site, or on multiple electronic devices distributed across multiple sites and interconnected by a communication network.
[0199] The above description is merely an embodiment of the present application and is not intended to limit the scope of protection of the present application. Any modifications, equivalent replacements, and improvements made within the spirit and scope of the present application are included in the scope of protection of the present application.
Claims
1. A method for processing a virtual scene, characterized in that: The method comprises: Displaying a requirement editing portal in the virtual scene generation interface, wherein the requirement editing portal is used to edit the scene requirements required for generating the virtual scene; In response to a requirement editing operation triggered based on the requirement editing portal, displaying the edited target scenario requirement, wherein the type of the target scenario requirement includes at least one of the following: text, image; In response to a preview generation operation, displaying at least one reference scene graph generated based on the target scene requirement; In response to a scene generation operation triggered based on the target scene graph, a target virtual scene generated based on the target scene graph is displayed.
2. The method according to claim 1, characterized in that The generation interface includes a demand editing area and a scenario display area; The requirement editing area is used to display the requirement editing entrance, the target scene requirement and the at least one reference scene graph, and the scene display area is used to display the target virtual scene.
3. The method according to claim 2, characterized in that The demand editing area is a conversation area for the target account and the smart assistant to communicate. The target scenario requirements to be edited include: A first conversation message sent to the smart assistant via the target account displays the edited target scenario requirements; The displaying of at least one reference scene graph generated based on the target scene requirement includes: The smart assistant replies to the second conversation message of the target account, and displays at least one reference scene graph generated based on the target scene requirements.
4. The method according to claim 1, wherein The displaying of the edited target scenario requirements in response to the requirement editing operation triggered by the requirement editing entry includes: In a case where the requirement editing entry is a text editing box, in response to a text editing operation triggered based on the text editing box, displaying text content input into the text editing box in the text editing box; In response to a determination operation on the text content, the text content is displayed, and the text content is used as the target scenario requirement.
5. The method according to claim 1, wherein The displaying of the edited target scenario requirements in response to the requirement editing operation triggered by the requirement editing entry includes: In the case where the requirement editing entry is a voice input icon, in response to the voice input operation triggered by the voice input icon, the text content obtained by text conversion based on the input voice is displayed, and the text content is used as the target scenario requirement.
6. The method according to claim 1, characterized in that The displaying of the edited target scenario requirements in response to the requirement editing operation triggered by the requirement editing entry includes: In a case where the required editing entry is a picture drawing entry, in response to a triggering operation on the picture drawing entry, a picture drawing interface is displayed, wherein the picture drawing interface includes a picture drawing canvas and a drawing toolbar, and the drawing toolbar includes a plurality of brushes, and brushes of different colors represent different scene elements; In response to a picture drawing operation triggered by a brush of a target color, displaying an outline picture of a target scene element drawn by the brush of the target color in the picture drawing canvas; In response to a determination operation on the outline picture, a target picture determined based on the outline picture is displayed, and the target picture is used as the target scene requirement.
7. The method according to claim 6, characterized in that The step of displaying a target image drawn based on the outline image in response to a determination operation on the outline image comprises: In response to a determination operation on the outline picture, controlling the outline picture to be in a to-be-published state, and displaying the outline picture in the to-be-published state at a location associated with the picture drawing entry; In response to a text editing operation triggered based on the text editing box, displaying text description content based on input into the text editing box in the text editing box; In response to a determination operation on the text description content, the text description content and the outline picture in a pending publication state are published, and a target picture including the text description content and the outline picture is displayed.
8. The method according to claim 6, characterized in that The displaying of the target image determined based on the outline image includes: If the outline image satisfies the supplementation condition, displaying a target image obtained by supplementing the outline image; Among them, the supplementary conditions include at least one of the following: there is a shape to be completed in the outline image, and the associated elements of the target scene element are missing in the outline image. The associated elements include at least one of the following: some elements in the target scene element, and environmental elements associated with the target scene element.
9. The method according to claim 8, characterized in that The displaying of the target image obtained by supplementing the outline image includes: Displaying supplementary prompt information, wherein the supplementary prompt information is used to prompt the user to supplement the outline image; In response to a determination operation on the supplementary prompt information, a target image obtained by supplementing the outline image is displayed.
10. The method according to claim 6, characterized in that The step of displaying a target image determined based on the outline image in response to a determination operation on the outline image comprises: In response to a determination operation on the outline picture, displaying supplementary prompt information and a supplementary editing area; In response to a supplementary editing operation triggered in the supplementary editing area based on the supplementary prompt information, a target image obtained by supplementing the outline image based on the edited supplementary description content is displayed.
11. The method according to claim 6, characterized in that The step of displaying, in response to a picture drawing operation triggered by a brush of a target color, an outline picture of a target scene element drawn by a brush of the target color in the picture drawing canvas comprises: In response to a picture drawing operation triggered by a brush of a target color, if the picture drawing operation indicates that an initial picture of a target scene element drawn by the brush of the target color satisfies a supplementation condition, displaying an outline picture obtained by supplementing the initial picture; Among them, the supplementary conditions include at least one of the following: there is a shape to be completed in the initial picture, and the associated elements of the target scene element are missing in the initial picture. The associated elements include at least one of the following: some elements in the target scene element, and environmental elements associated with the target scene element.
12. The method according to claim 1, characterized in that The displaying of the edited target scenario requirements in response to the requirement editing operation triggered by the requirement editing entry includes: In a case where the required editing entry is a picture import entry, in response to a triggering operation on the picture import entry, a picture selection interface is displayed, wherein the picture selection interface includes a plurality of pictures for selection; In response to a selection operation on a target picture, the target picture is displayed, and the target picture is used as the target scene requirement.
13. The method according to claim 11, characterized in that In response to the selection operation on the target picture, displaying the target picture and using the target picture as the target scene requirement includes: In response to a selection operation on a target image, controlling the target image to be in a ready-to-publish state; Display the text description content based on the text edit box input in the text edit box; In response to a determination operation on the text description content, the text description content and the target image in a to-be-published state are published, and the text description content and the target image are displayed as the target scene requirement.
14. The method according to claim 1, wherein After displaying at least one reference scene graph generated based on the target scene requirement, the method further includes at least one of the following: In response to an adjustment operation for the target scene requirement, updating and displaying the reference scene graph based on the adjusted target scene requirement; In response to a trigger operation on a variant control, the reference scene graph is updated to another reference scene graph generated based on the target scene requirement, where the other reference scene graph is different from the reference scene graph.
15. The method according to claim 1, wherein Before displaying at least one reference scene graph generated based on the target scene requirement, the method further includes: In a case where the type of the target scenario requirement includes text, extracting semantic features of the target scenario requirement of the text type; Initializing a noise image and an iteration time step, and extracting a noise tensor of the noise image and a time step vector of the iteration time step; Performing attention adjustment on the semantic vector based on the noise tensor to obtain an attention feature, and performing residual prediction based on the attention feature and the time step vector to obtain a noise residual for the noisy image; The noisy image is iteratively denoised based on the residual noise to obtain the reference scene graph.
16. The method according to claim 1, wherein Before displaying at least one reference scene graph generated based on the target scene requirement, the method further includes: In a case where the target scene requirements include images and texts, semantic region segmentation is performed on the image-type target scene requirements to obtain a semantic mask image, and feature extraction is performed on the text-type target scene requirements to obtain a text prompt feature; Extracting spatial condition features from the semantic heat map encoding corresponding to the semantic mask image, and extracting multi-scale spatial features required by the target scene of the image type; fusing the spatial condition feature and the multi-scale spatial feature to obtain a spatial fusion feature; Performing attention adjustment on the spatial fusion feature based on the text prompt feature to obtain an attention feature; The semantic mask graph is hierarchically refined and synthesized based on the attention features to obtain the reference scene graph.
17. The method according to claim 1, wherein Before displaying the target virtual scene generated based on the target scene graph, the method further includes: Performing scene semantic segmentation on the target scene graph to obtain a background mask and a foreground mask; Based on the viewing angle of the target scene image, the background mask is corrected to obtain a first scene depth map; For a foreground pixel in the foreground mask, replacing the depth value of the foreground pixel with the depth value of the background area surrounding the foreground pixel to obtain a second scene depth map; The first scene depth map and the second scene depth map are fused to obtain a target scene depth map, and the target virtual scene is constructed based on the target scene depth map.
18. A virtual scene processing device, characterized in that: The device comprises: An entry display module is used to display a requirement editing entry in the virtual scene generation interface, wherein the requirement editing entry is used to edit the scene requirements required for generating the virtual scene; A requirement display module, configured to display the edited target scenario requirement in response to a requirement editing operation triggered based on the requirement editing entry, wherein the type of the target scenario requirement includes at least one of the following: text, image; a reference display module, configured to display at least one reference scene graph generated based on the target scene requirement in response to a preview generation operation; The scene display module is used to display a target virtual scene generated based on the target scene graph in response to a scene generation operation triggered based on the target scene graph.
19. An electronic device, characterized in that: include: a memory for storing computer-executable instructions or computer programs; The processor is configured to implement the method for processing a virtual scene according to any one of claims 1 to 17 when executing the computer executable instructions or computer program stored in the memory.
20. A computer-readable storage medium, characterized in that Computer executable instructions or computer programs are stored, and when the computer executable instructions or computer programs are executed by a processor, the virtual scene processing method according to any one of claims 1 to 17 is implemented.
21. A computer program product comprising a computer program or computer executable instructions, characterized in that When the computer program or computer executable instructions are executed by a processor, the virtual scene processing method according to any one of claims 1 to 17 is implemented.
Citation Information
Cited By
Intelligent generation method and system for territorial space planning effect picture
CN121304848A