Methods, apparatus, devices, storage media, and program products for constructing virtual scenes

CN122569804APending Publication Date: 2026-08-14BEIJING ZITIAO NETWORK TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-05-22
Publication Date
2026-08-14

AI Technical Summary

Benefits of technology

[0008]以此方式,丰富了虚拟场景的构建方式,降低了初始构建的门槛,又减少了反复调整的迭代成本,从而在保证场景质量的同时,大幅提高了虚拟场景的构建效率。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122569804A_ABST
    Figure CN122569804A_ABST
Patent Text Reader

Abstract

A method, apparatus, device, storage medium, and program product for constructing virtual scenes are provided. The proposed method includes: receiving a first input related to terrain information of a three-dimensional virtual scene; presenting an editing component for editing a first image generated based on the first input, the first image representing the two-dimensional terrain distribution of the three-dimensional virtual scene; and constructing a three-dimensional virtual scene based on a second image obtained from the editing component. This approach enriches the methods for constructing virtual scenes and improves the efficiency of virtual scene construction.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The examples in this article generally relate to the field of computer science, and in particular to methods, apparatus, devices, computer storage media, and computer program products for constructing virtual scenes. Background Technology

[0002] With the increasing maturity of internet technology, more and more users are engaging in interactive activities on online platforms. For example, users can interact through virtual scenes. Users can also utilize online platforms to construct virtual scenes. Therefore, how to enrich the ways in which virtual scenes are constructed is worthy of attention. Summary of the Invention

[0003] In a first aspect, a method for constructing a virtual scene is provided. The method includes: receiving a first input related to terrain information of a three-dimensional virtual scene; presenting an editing component for editing a first image generated based on the first input, the first image representing a two-dimensional terrain distribution of the three-dimensional virtual scene; and constructing a three-dimensional virtual scene based on a second image obtained from the editing component.

[0004] In a second aspect, an apparatus for constructing a virtual scene is provided. The apparatus includes: a first receiving module configured to receive a first input related to terrain information of a three-dimensional virtual scene; a first rendering module configured to render an editing component for editing a first image generated based on the first input, the first image representing a two-dimensional terrain distribution of the three-dimensional virtual scene; and a construction module configured to construct a three-dimensional virtual scene based on a second image obtained from the editing component.

[0005] In a third aspect, an electronic device is provided. The device includes at least one processor; and at least one memory coupled to the at least one processor and storing instructions for execution by the at least one processor. When executed by the at least one processor, the instructions cause the device to perform the method of the first aspect.

[0006] In a fourth aspect, a computer-readable storage medium is provided. The computer-readable storage medium stores computer-executable instructions that can be executed by a processor to implement the method of the first aspect.

[0007] In a fifth aspect, a computer program product is provided, which is tangibly stored in a computer storage medium and includes computer-executable instructions that, when executed by a device, cause the device to perform the method of the first aspect.

[0008] This approach enriches the methods for constructing virtual scenes, lowers the threshold for initial construction, and reduces the iterative costs of repeated adjustments, thereby significantly improving the construction efficiency of virtual scenes while ensuring scene quality.

[0009] It should be understood that the content described in this section is not intended to limit the key or important features of the examples in this article, nor is it intended to restrict the scope of the solution. Other features will become readily apparent from the following description. Attached Figure Description

[0010] The above and other features, advantages, and aspects of the various examples herein will become more apparent when taken in conjunction with the accompanying drawings and the following detailed description. In the accompanying drawings, the same or similar reference numerals denote the same or similar elements, wherein: Figure 1 A schematic diagram of the example environment is shown; Figures 2A to 2F Example interfaces for some scenarios are shown; Figures 3A to 3G Example interfaces for some scenarios are shown; Figure 4 Flowcharts illustrating example processes for generating virtual scenes in several scenarios are shown; Figure 5 The flowcharts show example processes for constructing virtual scenes in several scenarios; Figure 6 Schematic structural block diagrams of example devices for constructing virtual scenes in several scenarios are shown; and Figure 7 A block diagram of an electronic device capable of implementing multiple illustrative scenarios is shown. Detailed Implementation

[0011] The examples in this document will now be described in more detail with reference to the accompanying drawings. While some examples are shown in the drawings, it should be understood that solutions can be implemented in various forms and should not be construed as limited to the examples presented herein. Rather, these examples are provided to provide a more thorough and complete understanding of the solutions. It should be understood that the drawings and examples in this document are for illustrative purposes only and are not intended to limit the scope of protection of the solutions.

[0012] It should be noted that the headings of any section / subsection provided herein are not restrictive. Various examples are described throughout this document, and examples of any type may be included under any section / subsection. Furthermore, examples described in any section / subsection may be combined in any way with any other examples described in the same section / subsection and / or different sections / subsections.

[0013] In the description of the examples in this document, the term "including" and similar terms should be understood as open inclusion, i.e., "including but not limited to". The term "based on" should be understood as "at least partially based on". The term "an example" or "the example" should be understood as "at least one example". The term "some examples" should be understood as "at least some examples". Other explicit and implicit definitions may also be included below. The terms "first", "second", etc., may refer to different or the same objects. Other explicit and implicit definitions may also be included below.

[0014] The examples in this article may involve user data, data acquisition, and / or use. All of these aspects comply with relevant laws, regulations, and rules. In the examples presented here, all data collection, acquisition, processing, manipulation, forwarding, and use are conducted with the user's knowledge and confirmation. Accordingly, when implementing each example, the type, scope of use, and usage scenarios of any data or information that may be involved should be communicated to the user and their authorization obtained through appropriate means, in accordance with relevant laws and regulations. The specific methods of notification and / or authorization can vary depending on the actual situation and application scenario; the scope of the solution is not limited in this regard.

[0015] In this manual and the sample solutions, any processing of personal information will be conducted only on a legal basis (such as with the consent of the data subject or as necessary for the performance of a contract) and will only be carried out within the scope stipulated or agreed upon. A user's refusal to process personal information beyond what is necessary for basic functions will not affect the user's use of basic functions.

[0016] As mentioned above, with the increasing maturity of network technology, more and more users are engaging in interactive activities on online platforms. For example, users can interact through virtual scenes. Users can also use online platforms to build virtual scenes. On the one hand, building virtual scenes using modeling software requires users to possess strong professional skills, which presents a high barrier to entry. On the other hand, virtual scenes directly generated using existing generative models are of poor quality and cannot meet interactive needs. Therefore, how to enrich the methods of building virtual scenes and improve the quality and efficiency of virtual scene construction is worthy of attention.

[0017] A scheme for constructing a virtual scene is proposed. The scheme includes: receiving a first input, which is related to the terrain information of a 3D virtual scene; presenting an editing component for editing a first image, generated based on the first input; and representing the 2D terrain distribution of the 3D virtual scene. Further, a 3D virtual scene can be constructed based on a second image, obtained using the editing component.

[0018] This approach enables the generation of a 2D image representing the 2D terrain distribution (e.g., a first image) based on a first input, and allows for flexible editing of the first image using editing components to obtain a second image. This second image can then be used to construct a virtual scene. This provides a virtual scene construction method that combines input generation and editing adjustments, thus enriching the ways to construct virtual scenes. On one hand, it allows users to intuitively modify the first image through editing components, making the final constructed 3D virtual scene more tailored to the user's needs, thereby significantly improving the customization and visual quality of the virtual scene. On the other hand, compared to purely manual modeling or fully automated generation, this solution organically integrates input generation and interactive editing, lowering the initial construction threshold and reducing the iterative costs of repeated adjustments, thereby significantly improving the construction efficiency of the virtual scene while ensuring scene quality. Furthermore, since the editing operation directly affects the 2D image, users do not need to deal with a complex 3D modeling process, further reducing the operational difficulty and the user's learning cost.

[0019] The following describes various examples of this scheme in further detail with reference to the accompanying drawings.

[0020] Example Environment Figure 1 A schematic diagram of example environment 100 is shown. (e.g.) Figure 1 As shown, example environment 100 may include electronic device 110.

[0021] In this example environment 100, electronic device 110 may run an application 120 that supports the construction of virtual scenes. Application 120 may be any suitable type of application for constructing virtual scenes, including but not limited to: media applications, social applications, or other suitable applications. User 140 may interact with application 120 via electronic device 110 and / or its attached devices.

[0022] exist Figure 1 In environment 100, if application 120 is active, electronic device 110 can use application 120 to present interface 150 for supporting the construction of virtual scenes.

[0023] In some cases, electronic device 110 communicates with server 130 to provide services to application 120. Electronic device 110 can be any type of mobile terminal, fixed terminal, or portable terminal, including mobile phones, desktop computers, laptop computers, notebook computers, netbook computers, tablet computers, media computers, multimedia tablets, handheld computers, portable gaming terminals, VR / AR devices, personal communication system (PCS) devices, personal navigation devices, personal digital assistants (PDAs), audio / video players, digital cameras / camcorders, positioning devices, television receivers, radio receivers, e-book devices, gaming devices, or any combination of the foregoing, including accessories and peripherals of these devices or any combination thereof. In some cases, electronic device 110 can also support any type of user-facing interface (such as "wearable" circuitry).

[0024] Server 130 can be a standalone physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, content delivery networks, and big data and artificial intelligence platforms. Server 130 may include, for example, computing systems / servers such as mainframes, edge computing nodes, computing devices in a cloud environment, etc. Server 130 can provide backend services for applications 120 in electronic devices 110 that support the construction of virtual scenes.

[0025] A communication connection can be established between server 130 and electronic device 110. This communication connection can be established via wired or wireless means. The communication connection can include, but is not limited to, Bluetooth, mobile network, Universal Serial Bus (USB), and Wireless Fidelity (WiFi) connections. In some cases, server 130 and electronic device 110 can exchange signaling information through their communication connection.

[0026] It should be understood that the structure and function of the various elements in environment 100 are described for illustrative purposes only and do not imply any limitation on the scope of the scheme.

[0027] The following description of the example will continue with reference to the accompanying drawings.

[0028] Example Interaction Figures 2A to 2FExample interfaces 200A to 200F are shown for constructing virtual scenes according to certain scenarios. Interfaces 200A to 200F can, for example, be constructed by... Figure 1 The electronic device 110 shown is provided.

[0029] In some cases, the electronic device 110 may receive a first input. The first input relates to terrain information of a three-dimensional virtual scene. As an example, the first input corresponds to user input. Such input may include natural language input. For example, natural language input may include text content described in natural language. Alternatively, such input may also include at least one piece of media content. As an example, such at least one piece of media content is provided by the user. For example, at least one piece of media content is selected by the user from a media library (e.g., a local photo album or an online media library associated with the user). As an example, media content may include image content, audio content, or video content, or a combination thereof.

[0030] In some cases, such as Figure 2A As shown, electronic device 110 can present interface 200A. Interface 200A can be implemented as an interface for constructing a virtual scene. As an example, such an interface can be provided by appropriate application software (e.g., software that supports constructing virtual scenes). As an example, electronic device 110 can present input component 202 in interface 200A. As an example, input component 202 is configured to receive a first input (e.g., terrain description) associated with the three-dimensional virtual scene to be constructed. For example, electronic device 110 can receive the first input via input component 202. For example, the first input can include natural language input. For example, such first input can include text content 204 (e.g., "Generate a map of an island with beaches, rivers, and volcanoes"). For example, the first input is related to terrain information of the three-dimensional virtual scene. As an example, terrain information can indicate the morphological features of the surface of the three-dimensional virtual scene. Such morphological features can, for example, indicate elevation, aspect, surface material composition, and the distribution of natural or artificial features. For example, such morphological features are composed of one or more terrain elements. For example, terrain elements may include islands, beaches, rivers, volcanoes, basins, valleys, plains, etc. For example, text content 204 may indicate terrain information of the 3D virtual scene to be constructed. For example, text content 204 may indicate that the terrain elements included in the 3D virtual scene to be constructed are islands, beaches, rivers, and volcanoes. For example, text content 204 may also indicate that beaches, rivers, and volcanoes are distributed on islands.

[0031] In some cases, electronic device 110 may present control 206 (e.g., "Generate terrain structure") in interface 200A. Electronic device 110 may receive first input (e.g., text content 204) in response to a triggering operation on control 206 (e.g., a click operation).

[0032] In some scenarios, electronic device 110 or server 130 can expand upon the first input (e.g., a terrain description input by a user). For example, electronic device 110 or server 130 can utilize a machine learning model to expand upon the first input. Such a machine learning model has the ability to expand text content, generating richer and more coherent natural language text based on concise descriptions or keywords. For instance, such a machine learning model can semantically expand the original terrain description, making the expanded terrain description more complete and suitable for subsequent construction of a 3D virtual scene. This approach is not intended to limit the training process or specific implementation of this machine learning model.

[0033] Alternatively, the expansion operation of the first input can be triggered manually by the user or automatically by the electronic device 110 or the server 130. For example, the electronic device 110 or the server 130 can automatically invoke a machine learning model to expand the terrain description in response to the first input meeting preset conditions. For example, such preset conditions may include the text length of the first input being less than a preset threshold (e.g., 50 characters). Alternatively, such preset conditions may include the first input lacking key terrain features (e.g., terrain distribution, etc.).

[0034] Alternatively, the expansion operation of the first input can be manually triggered by the user. For example, electronic device 110 can present control 208 (e.g., word association enhancement) in interface 200A. Electronic device 110 can respond to a triggering operation on control 208 (e.g., a click operation, etc.) to trigger the invocation of a machine learning model to expand the input content obtained by input component 202. As an example, electronic device 110 can provide control 208 in response to the text length of the input content received by input component 202 being greater than a preset threshold (e.g., 10 characters).

[0035] In some cases, such as Figure 2BAs shown, electronic device 110 can present interface 200B. Interface 200B can be implemented as an interface for constructing a three-dimensional virtual scene. As an example, electronic device 110 can present first descriptive text 210 about terrain information on interface 200B. As an example, the first descriptive text 210 is obtained based on a first input. For example, the first descriptive text 210 can correspond to the first input (e.g., a terrain description). Alternatively, the first descriptive text 210 can also correspond to an expanded first input (e.g., an expanded terrain description).

[0036] For example, electronic device 110 can display text content 212 (e.g., "parsed terrain: beach, river, valley, volcano") on interface 200B. For example, text content 212 can represent one or more terrain elements of the 3D virtual scene to be constructed (e.g., "beach, river, valley, volcano"). Through this visual presentation, during the construction of the 3D virtual scene, the specific terrain type identified after semantic parsing of the original input (e.g., the terrain description entered by the user) can be intuitively shown to the user. This helps improve the efficiency of users obtaining key information, enabling them to quickly confirm the terrain type understood by the system; on the other hand, it also provides a clear reference for users to further construct the 3D virtual scene or supplement and modify the original terrain description.

[0037] In some scenarios, electronic device 110 or server 130 may generate a first image based on the first input in response to receiving the first input. For example, electronic device 110 or server 130 may generate the first image by invoking a pre-trained machine learning model (e.g., a generative model). The first image may also be referred to as a terrain semantic map. As an example, the input to the machine learning model may include the first input, the terrain type obtained through semantic parsing (e.g., "beach," "river," "valley," "volcano"), and optional spatial relationship constraints (e.g., "rivers flow from high to low," "beach should not be in a volcano," etc.). As an example, the input to the machine learning model may also include mappings between various terrain elements and various styles. For example, different types of terrain elements may correspond to different styles (e.g., different colors or texture fills). For example, different styles indicate different colors: "beach" may correspond to yellow, "river" to blue, "valley" to green, "volcano" to red, etc. These mappings may be provided to the machine learning model in a structured manner as part of cue words. As an example, the output of the machine learning model may include the first image. As an example, the first image may represent the two-dimensional terrain distribution of a three-dimensional virtual scene.

[0038] As an example, the first image can be a two-dimensional raster image. The first image can include various graphic elements (e.g., pixels or pixel blocks). As an example, these various graphic elements correspond to different types of terrain elements. These various graphic elements have different styles (e.g., different colors or texture fills). For example, different graphic elements can correspond to different colors, thus allowing the spatial distribution of various terrain elements to be visually indicated through color distribution. For example, the first image can clearly show the location, extent, and spatial relationships of various terrain elements in a three-dimensional virtual scene by using yellow areas to represent beaches, blue areas to represent rivers, green areas to represent valleys, and red areas to represent volcanoes.

[0039] In some cases, electronic device 110 or server 130 can generate a scene preview. For example, the scene preview can represent the three-dimensional terrain of a three-dimensional virtual scene. For example, the three-dimensional terrain can be obtained based on a first image. In some cases, electronic device 110 can present a scene preview (e.g., image 214) on interface 200B. This scene preview can be used to visually demonstrate the three-dimensional terrain of the three-dimensional virtual scene to be constructed to the user. In some cases, electronic device 110 can present a first preview image of the three-dimensional terrain (e.g., image 214). Further, electronic device 110 can present a second preview image of the three-dimensional terrain (not shown) in response to receiving an adjustment operation for the scene preview. For example, such adjustment operations can include rotation operations (for viewing the three-dimensional terrain from different angles), zoom operations (for viewing an overall overview or local details of the three-dimensional terrain), and / or drag operations (for panning the view), etc., of the first preview image. In some cases, electronic device 110 may also provide additional view control (e.g., direction control, view reset control, top / side view quick switch control, etc.) to assist the user in quickly adjusting the viewing angle. The first and second preview images can correspond to different viewing angles, allowing users to view the 3D terrain of the 3D virtual scene from all angles, which facilitates subsequent editing.

[0040] In some cases, the first image is generated based on the first descriptive text 210. Alternatively, the electronic device 110 can receive editing operations from the user on the first descriptive text 210. For example, the electronic device 110 can present the first descriptive text 210 in an input component in response to a triggering operation on control 216 (e.g., "Modify description"). Further, the electronic device 110 can receive editing operations on the first descriptive text 210 via the input component. For example, such editing operations may include adding, deleting, or modifying text content in the first descriptive text 210 (e.g., changing "there are volcanoes" to "there are two active volcanoes," or adding "and a swamp," etc.). For example, the electronic device 110 can edit the first descriptive text 210 based on the editing operation to obtain a second descriptive text. Further, the electronic device 110 can present a third image generated based on the second descriptive text. For example, the third image may include a terrain semantic map regenerated based on the second descriptive text. Alternatively, the electronic device 110 can trigger the regeneration of the three-dimensional terrain of the three-dimensional virtual scene based on the regenerated terrain semantic map in response to the terrain semantic map being regenerated.

[0041] In some cases, electronic device 110 may present an editing component. As an example, the editing component may be used to edit a first image. As an example, electronic device 110 may present the editing component in response to receiving a preset operation. As an example, electronic device 110 may present a control 218 (e.g., "Edit Terrain") on interface 200B. For example, the preset operation may include a trigger operation on control 218 (e.g., a click operation, etc.). Alternatively, the preset operation may also include a swipe operation on interface 200B, etc.

[0042] like Figure 2CAs shown, electronic device 110 can present interface 200C. For example, interface 200C can be implemented as an interface for constructing a three-dimensional virtual scene. As an example, electronic device 110 can present editing component 220 on interface 200C. Electronic device 110 can present a first image (e.g., image 222) in the editing component. The first image can represent the two-dimensional terrain distribution of the three-dimensional virtual scene. As an example, the first image includes various graphic elements, such as graphic element 224, graphic element 226, graphic element 228, and graphic element 220. The various graphic elements correspond to different types of terrain elements. For example, graphic element 224 can correspond to a "beach". For example, graphic element 226 can correspond to a "valley". For example, graphic element 228 can correspond to a "river". For example, graphic element 230 can correspond to a "volcano". For example, the various graphic elements have different styles (e.g., different colors or texture fills, etc.). Through these graphic elements and their style differences, the first image can intuitively and clearly represent the distribution of various terrain elements in the three-dimensional virtual scene on a two-dimensional plane, thereby improving the efficiency of users obtaining information.

[0043] In some cases, electronic device 110 can receive editing operations on a first image (e.g., image 222) through editing component 220 to edit the two-dimensional terrain distribution. As an example, electronic device 110 can present multiple controls in the editing component. These multiple controls may include, for example, controls 232, 234, 236, and 238. The multiple controls correspond to various terrain elements. For example, control 232 may correspond to a "beach." For example, control 234 may correspond to a "volcano." For example, control 236 may correspond to a "river." For example, control 238 may correspond to a "valley." Furthermore, electronic device 110 can receive a trigger (e.g., a click operation) on the first control among the multiple controls. As an example, the first control may correspond to a first terrain element. Taking control 234 as an example, the first terrain element may correspond to a "volcano."

[0044] In some cases, electronic device 110 may, in response to a triggering of a first control, present a prompt message in editing component 220. For example, the prompt message may represent a second region in a first image. For instance, the second region may correspond to a first terrain element. As an example, such a prompt message may include a visual effect corresponding to the second region. Taking the first terrain element as “volcano” as an example, the second region may be the region in the first image (e.g., image 222) corresponding to graphic element 230. For example, the visual effect corresponding to the second region may include highlighting the region corresponding to graphic element 230. Alternatively, the visual effect corresponding to the second region may include bolding the outline of the region corresponding to graphic element 230.

[0045] In some cases, electronic device 110 can receive editing operations based on triggering a first control. For example... Figure 2D As shown, electronic device 110 can present interface 200D. For example, interface 200D can be implemented as an interface for constructing a three-dimensional virtual scene. For example, electronic device 110 can present editing component 220 and a first image (e.g., image 222) in interface 200D. As an example, electronic device 110 can receive editing operations. For example, editing operations can include selection of a first region in the first image. As an example, electronic device 110 can present at least one editing control in the editing component. For example, at least one editing control can include editing control 240 (e.g., drawing control), editing control 242 (e.g., erasing control), and editing control 244 (e.g., brush size control). As an example, electronic device 110 can receive selection of editing control 240 in response to a triggering operation on editing control 240. While editing control 240 is in a selected state, electronic device 110 can select a first region in the first image based on the user's editing operation. For example, the first region can be determined based on input signals (e.g., a sequence of touch points or a sequence of mouse sampling points) received by the input unit (e.g., a touchscreen, mouse, or touchpad) deployed on the electronic device 110. For example, the electronic device 110 can present a first indicator element (e.g., indicator element 246) in a first image (e.g., image 222) while the editing control 240 is in a selection state to provide real-time feedback on the currently selected region or the brush's landing point position. As an example, the first region may include region 248. Further, the electronic device 110 can replace graphic elements (e.g., a graphic element corresponding to "valley") in the first region with a first graphic element. For example, the first graphic element corresponds to a first terrain element (e.g., "volcano").

[0046] As an example, editing control 242 (e.g., an erase control) can be used to remove one or more graphic elements drawn while editing control 240 is in the selected state, such as a first graphic element in a first region. Specifically, when the user triggers editing control 242, a click or swipe operation on image 222 can erase the pixel markers originally generated by the drawing control (e.g., the first graphic element in the first region), thereby enabling local adjustments to the first image. Alternatively, editing control 244 (e.g., a brush size control) can be used to adjust the brush size during drawing or erasing operations, for example, by setting the brush radius (e.g., in pixels) via a slider, numerical input, or continuous increment / decrement buttons. By combining the erase control and the brush size control, the user can flexibly adjust the two-dimensional terrain distribution indicated by the first image, thereby improving the precision and controllability of terrain editing.

[0047] Additionally, the electronic device 110 may present a second image in an editing component 220 or other suitable location (e.g., a pop-up component, etc.). For example, the second image corresponds to an edited first image (e.g., Figure 2D (Image 222 in the image). As an example, the second image can represent the edited two-dimensional terrain distribution. By presenting the second image, the electronic device 110 can provide the user with immediate visual feedback, helping the user confirm whether the editing effect meets expectations and providing an accurate semantic basis for subsequent regeneration of the three-dimensional terrain.

[0048] In some cases, electronic device 110 can receive confirmation of the editing operation. For example, electronic device 110 can receive confirmation of the editing operation in response to a trigger (e.g., a click operation) on control 250 (e.g., "Update Preview Image"). The scene preview corresponding to the first image is also referred to as the first scene preview. Further, electronic device 110 can present a second scene preview in response to editing the first image into a second image via the editing component. The second scene preview corresponds to the second image.

[0049] like Figure 2E As shown, electronic device 110 can present interface 200E. Interface 200E can be implemented as an interface for constructing a virtual scene. As an example, electronic device 110 can present a second scene preview (e.g., image 252) corresponding to a second image in interface 200E. This second scene preview is generated based on the edited terrain semantic map (i.e., the second image) and is used to show the user the effect of the 3D terrain after editing. As an example, the 3D terrain represented by the second scene preview may include a portion (e.g., the portion indicated by identifier 254) determined based on user editing operations (e.g., adding a first graphic element to a first region using drawing controls). For example, the portion indicated by identifier 254 is the first newly added terrain element (e.g., "volcano") in the 3D virtual scene. In this way, the user can intuitively see how the modifications they made on the terrain semantic map are mapped to the 3D scene, thereby confirming the correctness of the editing operation and providing a visual basis for further adjustments.

[0050] In some cases, electronic device 110 or server 130 can construct a three-dimensional virtual scene based on a second image. For example, the second image may be obtained based on an editing component. For example, if no editing operation is received, the second image used to construct the three-dimensional virtual scene is the first image. Alternatively, after receiving an editing operation, the second image used to construct the three-dimensional virtual scene may be an image obtained by editing the first image.

[0051] As an example, electronic device 110 can trigger the construction of a 3D virtual scene based on a second image, based on a triggering action (e.g., a click operation) on control 256 (e.g., a composite map). As an example, the process of constructing a 3D virtual scene can be referred to later in the text regarding... Figure 4 An exemplary description is provided, which will not be repeated here.

[0052] like Figure 2F As shown, electronic device 110 can present interface 200F. For example, interface 200F can be implemented as an interface for constructing a three-dimensional virtual scene. For example, electronic device 110 can present a three-dimensional virtual scene 258 constructed based on three-dimensional terrain in interface 200F. As an example, electronic device 110 can publish the three-dimensional virtual scene 258 in response to a publishing operation. For example, the three-dimensional virtual scene 258 can be published to a server or online platform. As an example, the published three-dimensional virtual scene 258 can be accessed or interacted with by other users. For example, the three-dimensional virtual scene 258 can be used to add virtual characters associated with users and control the virtual characters to walk, jump, explore, or perform specific tasks on the three-dimensional terrain in the three-dimensional virtual scene.

[0053] In some cases, the process of constructing a 3D virtual scene mentioned above only involves the construction of terrain elements within the 3D virtual scene. In other cases, scene elements within the 3D virtual scene can also be constructed. The following is based on... Figures 3A to 3G The process of constructing scene elements in a 3D virtual scene will be described exemplarily.

[0054] like Figure 3A As shown, electronic device 110 can present interface 300A. Interface 300A can be implemented, for example, as an interface for constructing a three-dimensional virtual scene. For example, electronic device 110 can provide an add control 302 in interface 300A. For example, electronic device 110 can present at least one candidate image in response to a triggering operation of add control 302. For example, the at least one candidate image may include at least one terrain semantic map. Electronic device 110 can receive a selection of a first candidate image from the at least one candidate image.

[0055] like Figure 3B As shown, the electronic device 110 can present an interface 300B. The interface 300B can, for example, be implemented as an interface for constructing a three-dimensional virtual scene. For example, the electronic device 110 can present a selected first candidate image (e.g., image 304) on the interface 300B. For example, the first candidate image can correspond to the first image mentioned above. The first image can represent the terrain information of the three-dimensional virtual scene.

[0056] In some cases, electronic device 110 may present input component 306 in interface 300B. As an example, input component 306 is configured to receive a second input (e.g., a description of surface objects) associated with the 3D virtual scene to be constructed. For example, electronic device 110 may receive the second input via input component 306. For example, the second input may include natural language input. For example, such second input may include text content 308 (e.g., "There are coconut trees and rocks in the valley"). For example, the second input describes scene elements included in the 3D virtual scene in natural language. For example, text content 308 may indicate the scene elements included in the 3D virtual scene to be constructed. For example, text content 308 may indicate that the scene elements included in the 3D virtual scene to be constructed include coconut trees and rocks. For example, text content 308 may also indicate that coconut trees and rocks are distributed in the valley.

[0057] In some cases, electronic device 110 may present control 310 in interface 300B (e.g., "Generate Now"). Electronic device 110 may receive a second input (e.g., text content 308) in response to a triggering operation on control 310 (e.g., a click operation).

[0058] Alternatively, electronic device 110 or server 130 can expand the second input (e.g., a user-input description of a surface object). For example, electronic device 110 or server 130 can expand the second input using a machine learning model. Such a machine learning model has the ability to expand text content, generating richer and more coherent natural language text based on concise descriptions or keywords. For example, such a machine learning model can semantically expand the original surface object description, making the expanded surface object description more complete and suitable for subsequent construction of a 3D virtual scene. This approach is not intended to limit the training process or specific implementation of this machine learning model.

[0059] Alternatively, the expansion operation of the second input can be triggered manually by the user or automatically by the electronic device 110 or the server 130. For example, the electronic device 110 or the server 130 can automatically invoke a machine learning model to expand the terrain description in response to the second input meeting preset conditions. For example, such preset conditions may include the text length of the second input being less than a preset threshold (e.g., 50 characters). Alternatively, such preset conditions may include the second input lacking key scene element information (e.g., the distribution of scene elements, etc.).

[0060] Alternatively, the expansion operation on the second input can be triggered manually by the user. For example, see [link to relevant documentation]. Figure 3BThe electronic device 110 can display control 312 (e.g., word association enhancement) on interface 300B. The electronic device 110 can respond to a triggering operation on control 312 (e.g., a click operation) to trigger a machine learning model to expand the input content received by input component 306. As an example, the electronic device 110 can provide control 312 in response to the input content received by input component 306 having a text length greater than a preset threshold (e.g., 10 characters).

[0061] In some cases, electronic device 110 or server 130 can decompose the scene elements represented by the second input (or an expanded version of the second input) to determine one or more scene elements. For example, electronic device 110 or server 130 can perform semantic analysis on the second input to determine one or more scene elements included in the 3D virtual scene to be constructed. Taking the second input as including text content (e.g., "There are coconut trees and stones in the valley") as an example, the scene elements corresponding to this second input may include "coconut trees" and "stones," etc.

[0062] As an example, a user can associate with a pre-defined set of objects. This pre-defined set of objects can include at least one virtual object associated with the user. For example, such a virtual object could include virtual objects in the user's associated asset library (or resource library). As an example, a virtual object refers to a specific material or 3D model (e.g., a coconut tree model, a stone model, a house model, etc.) that can be used for rendering in a 3D virtual scene, and it can contain meshes, textures, materials, and metadata. For example, such virtual objects can have various sources: those created and uploaded by the user, general assets provided by the platform, or those obtained by the user through redemption or other means. As an example, the virtual objects associated with the user are also referred to as the user's assets or resources.

[0063] In some cases, virtual objects can correspond to scene elements. As an example, a scene element (such as a "coconut tree") is a semantically abstract concept representing the type of object (such as vegetation, buildings, natural objects, etc.) that needs to be placed in a 3D virtual scene. Virtual objects are the actual model files used for rendering. For example, a virtual object corresponding to a scene element can be implemented in multiple forms. For instance, the scene element "coconut tree" can be implemented as different 3D models of "coconut trees." The same scene element can be associated with multiple virtual objects of different forms. For example, when generating a single instance, one form can be selected for instantiation. For example, the scene element "coconut tree" can be associated with a tall coconut tree model, a short coconut tree model, a coconut tree model with fruit, etc. When placing a coconut tree in a 3D virtual scene, one of the models can be selected for generation based on preset rules or constraints indicated by a second input.

[0064] In some scenarios, electronic device 110 or server 130 can semantically retrieve assets corresponding to the second input. For example, electronic device 110 or server 130 can semantically match multiple scene elements (e.g., "coconut tree," "stone," etc.) described by the second input with at least one virtual object from a predetermined set of objects. As an example, electronic device 110 or server 130 can utilize natural language understanding technology to calculate the semantic similarity between the scene element and the descriptive information (e.g., tags, names, etc.) of the at least one virtual object. As an example, electronic device 110 or server 130 can determine that the scene element matches the corresponding virtual object in response to a semantic similarity greater than a preset threshold. Conversely, electronic device 110 or server 130 can determine that the scene element does not match the corresponding virtual object in response to a semantic similarity less than a preset threshold.

[0065] In some cases, such as Figure 3C As shown, the electronic device 110 can present an interface 300C. The interface 300C can be implemented as an interface for constructing a three-dimensional virtual scene. As an example, the electronic device 110 can present third descriptive text 314 about multiple scene information on the interface 300C. As an example, the third descriptive text 314 is obtained based on a first input. For example, the third descriptive text 314 can correspond to a second input (e.g., a description of surface objects). Alternatively, the third descriptive text 314 can also correspond to an expanded second input (e.g., an expanded description of surface objects).

[0066] In some cases, the electronic device 110 may display instruction text corresponding to multiple scene elements included in the three-dimensional virtual scene on the interface 300C, such as instruction text 316 (e.g., "coconut tree") and instruction text 318 (e.g., "stone").

[0067] In some cases, electronic device 110 may present at least one graphic element, such as graphic element 320 and graphic element 322, in response to receiving a second input. For example, a graphic element may also be referred to as a legend. For example, different graphic elements may correspond to different styles (e.g., different colors or texture fills, etc.). At least one graphic element may represent at least one scene element. For example, at least one scene element is determined based on the second input. For example, at least one scene element is obtained by semantic analysis of the second input. For example, at least one graphic element may include a second graphic element. The second graphic element may be used to represent a first scene element among at least one scene element. Taking the second graphic element corresponding to graphic element 320 as an example, graphic element 320 may represent the first scene element (e.g., "coconut tree").

[0068] Additionally or alternatively, the electronic device 110 may, in response to receiving a first operation, replace a second graphic element with a third graphic element, such that the third graphic element represents a first scene element. The electronic device 110 may, in response to a triggering action on the second graphic element (e.g., a click operation), present multiple candidate graphic elements. The third operation may include a selection operation of the third graphic element from among the multiple candidate graphic elements. For example, the second graphic element may be different from the third graphic element. For example, the second graphic element may correspond to a first style. The third graphic element may correspond to a second style. For example, the first style and the second style may correspond to different colors or texture fills. For example, the first style may correspond to a first color (e.g., green), and the second style may correspond to a second color (e.g., brown).

[0069] As an example, electronic device 110 can trigger the generation of a fourth image. For instance, electronic device 110 can trigger the generation of a fourth image in response to a trigger on control 324 (e.g., "Generate surface objects") (e.g., a click operation, etc.).

[0070] As an example, electronic device 110 or server 130 can generate a fourth image by invoking a pre-trained machine learning model (e.g., a generative model). For example, the fourth image can also be referred to as a scene element semantic map. As an example, the input to this machine learning model can include a second input, multiple scene elements obtained through semantic parsing (e.g., "coconut tree," "stone"), and optional spatial relationship constraints (e.g., "coconut trees cannot grow on stones," "coconut trees should not be in rivers," etc.). As an example, the input to the machine learning model can also include mapping relationships between multiple scene elements and multiple graphical elements. For example, different scene elements can correspond to different styles (e.g., different colors or texture fills, etc.). For example, different styles indicate different colors; "coconut tree" can correspond to green, "stone" can correspond to black, etc. These mapping relationships can be provided to the machine learning model in a structured manner as part of cue words. As an example, the output of the machine learning model can include the fourth image. As an example, the fourth image can represent the distribution of multiple scene elements in a two-dimensional plane. The two-dimensional plane corresponds to a three-dimensional virtual scene. For example, the two-dimensional plane can correspond to a top-down view of a three-dimensional virtual scene.

[0071] As an example, the fourth image can be a two-dimensional raster image. The fourth image can include various graphic elements (e.g., pixels or pixel blocks). As an example, these various graphic elements correspond to multiple scene elements. These various graphic elements have different styles (e.g., different colors or texture fills). For example, different graphic elements can correspond to different colors, thus allowing the spatial distribution of various scene elements to be visually indicated through color distribution. For example, the fourth image can use green areas to represent coconut trees and black areas to represent stones, thus clearly showing the distribution of various scene elements in a two-dimensional plane (e.g., location, extent, and positional relationships).

[0072] like Figure 3D As shown, electronic device 110 can present interface 300D. For example, electronic device 110 can present a third scene preview (e.g., preview image 326) on interface 300D. For example, the third scene preview is generated based on a fourth image. This third scene preview can be used to visually display the three-dimensional terrain of the three-dimensional virtual scene to be constructed to the user. For example, the third scene preview is used to display multiple virtual objects (i.e., specific three-dimensional model instances) placed in the three-dimensional virtual scene according to the fourth image (semantic map of surface objects). Each virtual object corresponds to a scene element (such as a coconut tree, a rock, etc.). The distribution of multiple virtual objects in the three-dimensional virtual scene is determined based on the distribution of multiple scene elements in a two-dimensional plane.

[0073] In some cases, the electronic device 110 may present a third preview image of the three-dimensional terrain (e.g., preview image 326). Further, the electronic device 110 may, in response to receiving an adjustment operation on the third scene preview, present a fourth preview image of the three-dimensional terrain (not shown). For example, such adjustment operations may include rotation operations on the third preview image (for viewing the three-dimensional terrain from different angles), zoom operations (for viewing an overall overview or local details of multiple virtual objects), and / or drag operations (for panning the view). In some cases, the electronic device 110 may also provide additional viewpoint controls (e.g., direction controls, viewpoint reset controls, top / side view quick switch controls, etc.) to assist the user in quickly adjusting the viewing angle. The third and fourth preview images may correspond to different viewing angles, allowing the user to view multiple virtual objects in the three-dimensional virtual scene from all angles, facilitating subsequent editing.

[0074] In some cases, continue to refer to Figure 3CThe fourth image is generated based on the third descriptive text 314. Alternatively, the electronic device 110 can receive editing operations from the user on the third descriptive text 314. For example, the electronic device 110 can present the third descriptive text 314 in an input component in response to a triggering operation on control 328 (e.g., "Modify description"). Further, the electronic device 110 can receive editing operations on the third descriptive text 314 via the input component. For example, such editing operations may include adding, deleting, or modifying text content in the third descriptive text 314 (e.g., changing "there are coconut trees" to "there are apple trees," or adding "there are fish in the river," etc.). For example, the electronic device 110 can edit the third descriptive text 314 based on the editing operations to obtain the fourth descriptive text. Further, the electronic device 110 can present an image generated based on the fourth descriptive text (also referred to as the sixth image). For example, the image may include a semantic map of surface objects regenerated based on the fourth descriptive text. Alternatively, the electronic device 110 may, in response to the regeneration of the semantic map of the ground objects, trigger a third scene preview of the three-dimensional virtual scene to be regenerated based on the regenerated semantic map of the ground objects.

[0075] In some cases, the electronic device 110 may present an editing component. As an example, the editing component may be used to edit a fourth image. As an example, the electronic device 110 may present the editing component in response to receiving a preset operation. As an example, the electronic device 110 may present a control 330 (e.g., "Edit Object Distribution") in the interface 300D. For example, the preset operation may include a trigger operation on the control 330 (e.g., a click operation, etc.). Alternatively, the preset operation may also include a swipe operation on the interface 300D, etc.

[0076] like Figure 3E As shown, electronic device 110 can present interface 300E. For example, interface 300E can be implemented as an interface for constructing a three-dimensional virtual scene. As an example, electronic device 110 can present editing component 332 on interface 300E. Electronic device 110 can present a fourth image (e.g., image 334) in the editing component. The fourth image can represent the distribution of multiple scene elements on a two-dimensional plane. As an example, the fourth image includes multiple graphic elements. Multiple graphic elements correspond to different scene elements. For example, graphic elements 336, 338, 340, and 342 can correspond to the scene element "coconut tree". For example, graphic elements 344 and 346 can correspond to the scene element "stone". For example, multiple graphic elements have different styles (e.g., different colors or texture fills). Through these graphic elements and their style differences, the fourth image can intuitively and clearly represent the distribution of multiple scene elements on a two-dimensional plane, thereby improving the efficiency of users obtaining information.

[0077] As an example, taking a first scene element (e.g., a "coconut tree") as an example, the first scene element may correspond to a second graphic element. For example, the fourth image may include multiple graphic element instances corresponding to the second graphic element, such as graphic element 336, graphic element 338, graphic element 340, and graphic element 342. These multiple graphic elements can represent the first distribution (i.e., the positional arrangement of the style on a two-dimensional plane) of the second graphic element (i.e., the style corresponding to the first scene element) in the 3D virtual scene. For example, the position of each graphic element corresponds to a location where a coconut tree model will be placed, and its density, distribution range, and relative positional relationships can be intuitively read from the diagram. The first distribution can represent the second distribution of the first scene element in the 3D virtual scene. For example, the second distribution can indicate the arrangement of virtual objects corresponding to the first scene element in the 3D virtual scene. Through this two-layer mapping, abstract semantic elements are transformed into a visual graphic distribution, facilitating users' intuitive understanding and editing of the layout of scene elements.

[0078] In some cases, electronic device 110 can construct a three-dimensional virtual scene based on a fourth image. In some cases, electronic device 110 can receive editing operations on the fourth image (e.g., image 334) through editing component 332 to obtain a fifth image. As an example, electronic device 110 can present multiple controls in the editing component. The multiple controls may include, for example, controls 348 and 350. The multiple controls correspond to multiple scene elements. For example, control 348 may correspond to a "coconut tree". For example, control 350 may correspond to a "stone". Further, electronic device 110 can receive a trigger (e.g., a click operation, etc.) on a second control among the multiple controls. As an example, the second control may correspond to a fourth scene element. Taking the second control corresponding to control 348 as an example, the fourth scene element may correspond to a "coconut tree".

[0079] In some cases, electronic device 110 may, in response to a triggering of a second control, present a prompt message in editing component 332. For example, the prompt message may characterize a fourth region in a fourth image. For instance, the fourth region may correspond to a fourth scene element. As an example, such a prompt message may include a visual effect corresponding to the fourth region. Taking the fourth scene element as "coconut tree" as an example, the fourth region may be the region in the fourth image (e.g., image 334) corresponding to graphic elements 336, 338, 340, and 342. For example, the visual effect corresponding to the fourth region may include highlighting the region corresponding to graphic elements 336, 338, 340, and 342. Alternatively, the visual effect corresponding to the fourth region may include bolding the outline of the region corresponding to graphic elements 336, 338, 340, and 342.

[0080] In some cases, electronic device 110 can receive editing operations based on triggering a first control. For example... Figure 3F As shown, electronic device 110 can present interface 300F. For example, interface 300F can be implemented as an interface for constructing a three-dimensional virtual scene. For example, electronic device 110 can present editing component 332 and a fourth image (e.g., image 334) in interface 300F. As an example, electronic device 110 can receive editing operations. For example, editing operations can include selection of a third region in the fourth image. As an example, electronic device 110 can present at least one editing control in the editing component. For example, at least one editing control can include editing control 352 (e.g., a drawing control), editing control 354 (e.g., an erasing control), and editing control 356 (e.g., a brush size control). As an example, electronic device 110 can receive selection of editing control 352 in response to a triggering operation on editing control 352. While editing control 352 is in a selected state, electronic device 110 can select a third region in the fourth image based on the user's editing operation. For example, the third region can be determined based on input signals (e.g., a sequence of touch points or a sequence of mouse sampling points) received by the input unit (e.g., a touchscreen, mouse, or touchpad) deployed on the electronic device 110. For example, while the editing control 352 is in a selected state, the electronic device 110 can present a second indicator element (e.g., indicator element 358) in a fourth image (e.g., image 334) to provide real-time feedback on the currently selected area or the brush's landing point position. As an example, the third region may include region 360. Furthermore, the electronic device 110 can replace graphic elements in the third region (e.g., a graphic element corresponding to "valley") with a fifth graphic element. For example, the fifth graphic element corresponds to a fourth scene element (e.g., "coconut tree").

[0081] As an example, editing control 354 (e.g., an erase control) can be used to remove one or more graphic elements drawn while editing control 352 is in the selected state, such as the first graphic element in the first region. Specifically, when the user triggers editing control 354, a click or swipe operation on image 334 can erase the pixel markers originally generated by the drawing control (e.g., the fifth graphic element in the third region), thereby enabling local adjustments to the fourth image. Alternatively, editing control 356 (e.g., a brush size control) can be used to adjust the brush size during drawing or erasing operations, for example, by setting the brush radius (e.g., in pixels) via a slider, numerical input, or continuous increment / decrement buttons. By combining the erase control and the brush size control, the user can flexibly adjust the two-dimensional terrain distribution indicated by the fourth image, thereby improving the precision and controllability of terrain editing.

[0082] Additionally, the electronic device 110 can present a fifth image in the editing component 332. For example, the fifth image corresponds to the edited fourth image (e.g., Figure 3F (Image 334 in the image). As an example, the fifth image can represent the distribution of multiple scene elements in a two-dimensional image. By presenting the fifth image, the electronic device 110 can provide the user with immediate visual feedback, helping the user confirm whether the editing effect meets expectations, and providing an accurate semantic basis for subsequently regenerating scene renderings or constructing three-dimensional virtual scenes.

[0083] In some cases, electronic device 110 can receive confirmation of the editing operation. For example, electronic device 110 can receive confirmation of the editing operation in response to a trigger (e.g., a click operation) on control 362 (e.g., "Update Preview Image"). The scene preview corresponding to the fourth image is also referred to as the third scene preview. Further, electronic device 110 can present a fourth scene preview in response to editing the fourth image into a fifth image via the editing component. The fourth scene preview corresponds to the fifth image.

[0084] In some cases, continue to refer to Figure 3CElectronic device 110 can present a first virtual object in response to receiving a first input. For example, electronic device 110 can present the first virtual object (e.g., virtual object 364) in interface 300C. The first virtual object corresponds to a second scene element. Exemplarily, the second scene element may be, for example, a "coconut tree". As an example, the second scene element is determined based on a second input. For example, the second scene element is obtained by semantic analysis of the second input. In some cases, the first virtual object is determined based on a user's selection operation. As an example, electronic device 110 can present at least one virtual object corresponding to the second scene element. As an example, electronic device 110 can present at least one virtual object corresponding to the second scene element in response to a second operation. For example, electronic device 110 can present control 366. The second operation may include a trigger operation on control 366 (e.g., a click operation, etc.). Alternatively, the second operation may also include a trigger operation on the location of virtual object 364.

[0085] In some scenarios, electronic device 110 may present at least one virtual object corresponding to a second scene element in interface 300C, such as virtual object 368, virtual object 370, virtual object 372, and virtual object 374. For example, such at least one virtual object is determined from a predetermined set of objects. For example, the predetermined set of objects may include multiple virtual objects in an asset library associated with the user. For example, electronic device 110 or server 130 may perform semantic matching between the second scene element (e.g., "coconut tree") and multiple virtual objects in the predetermined set of objects. As an example, electronic device 110 or server 130 may utilize natural language understanding technology to calculate the semantic similarity between the second scene element and the descriptive information (e.g., tags, names, etc.) of multiple virtual objects. As an example, electronic device 110 or server 130 may determine that the second scene element matches a corresponding virtual object in response to a semantic similarity greater than a preset threshold, thereby determining at least one virtual object corresponding to the second scene element.

[0086] Additionally, the electronic device 110 can present the first virtual object in response to the selection of a first virtual object among at least one virtual object. For example, the first virtual object can be used to construct a three-dimensional virtual scene. For example, the three-dimensional virtual scene places at least one first virtual object based on a third distribution. For example, the third distribution can include the actual arrangement and density of virtual objects corresponding to second scene elements (such as "coconut trees") in three-dimensional space. For example, the third distribution is determined based on a fourth distribution of fourth graphic elements in a fourth image. For example, the fourth graphic element is used to represent the second scene element. For example, the fourth graphic element is a legend in the fourth image used to represent the second scene element.

[0087] Alternatively, the electronic device 110 can, in response to receiving a third operation, replace the first virtual object with a second virtual object, such that at least one second virtual object is placed in the 3D virtual scene based on a third distribution. As an example, the electronic device 110 can, in response to selecting a second virtual object from at least one virtual object corresponding to a second scene element, replace the first virtual object with the second virtual object. Through this replacement operation, the user can adjust the specific model form corresponding to the second scene element without changing the spatial layout, thereby enabling rapid adjustment of the overall visual style of the 3D virtual scene and meeting local detail requirements.

[0088] Additionally or alternatively, electronic device 110 may present a prompt message in response to receiving the first input. The prompt message indicates that the third scene element is not associated with a virtual object. For example, electronic device 110 or server 130 may determine that the third scene element is not associated with a virtual object in response to multiple virtual objects in a predetermined object set having a semantic similarity to the third scene element that is less than a threshold. As an example, taking the third scene element corresponding to "stone" as an example, electronic device 110 may present prompt message 376. For example, electronic device 110 may present prompt message 376 at the associated location of the instruction text (e.g., "stone") corresponding to the third scene element. For example, prompt message 376 may include a preset visual element. Such a preset visual element may indicate that the third scene element is not associated with a virtual object.

[0089] In some cases, the final constructed 3D virtual scene may include a third virtual object corresponding to a third scene element. For example, such a third virtual object is generated based on a second input. For instance, electronic device 110 or server 130 may utilize a generative model to generate a third virtual object (e.g., a 3D model of a "stone") based on the second input. This generative model may have the ability to generate a 3D model based on input text. This approach is not intended to limit the training process or specific implementation of the generative model.

[0090] In some cases, electronic device 110 or server 130 can construct a three-dimensional virtual scene based on a fourth image (or an edited fourth image). See also, as an example. Figure 3D The electronic device 110 can trigger the construction of a three-dimensional virtual scene based on a fourth image, based on a trigger on the control 378 (e.g., a composite map, a click operation, etc.).

[0091] like Figure 3GAs shown, electronic device 110 can present interface 300G. For example, interface 300G can be implemented as an interface for constructing a three-dimensional virtual scene. As an example, electronic device 110 can present a three-dimensional virtual scene 380 including three-dimensional terrain and multiple virtual objects in interface 300G. As an example, electronic device 110 can publish the three-dimensional virtual scene 380 in response to a publishing operation. For example, the three-dimensional virtual scene 380 can be published to a server or online platform. As an example, the published three-dimensional virtual scene 380 can be accessed or interacted with by other users. For example, the three-dimensional virtual scene 380 can be used to add virtual characters associated with users and control the virtual characters to walk, jump, explore, or perform specific tasks on the three-dimensional terrain in the three-dimensional virtual scene.

[0092] Figure 4 A flowchart illustrating an example process 400 for generating a virtual scene is shown. Process 400 can be implemented at electronic device 110 and / or server 130. See below for reference. Figure 1 To describe process 400.

[0093] like Figure 4 As shown, in box 402, electronic device 110 can receive a first input. For example, the first input may include a terrain description input by a user. For example, such a terrain description may correspond to the first input. Such a terrain description can describe the terrain information of the 3D virtual scene to be constructed.

[0094] In box 404, electronic device 110 or server 130 can expand the first input (e.g., a terrain description input by the user). For example, electronic device 110 or server 130 can expand the first input using a machine learning model. Such a machine learning model has the ability to expand text content, generating richer and more coherent natural language text based on brief descriptions or keywords. For example, such a machine learning model can semantically expand the original terrain description, making the expanded terrain description more complete and suitable for subsequent construction of a 3D virtual scene. This approach is not intended to limit the training process or specific implementation of this machine learning model.

[0095] Alternatively, the expansion operation of the first input can be triggered manually by the user or automatically by the electronic device 110 or the server 130. For example, the electronic device 110 or the server 130 can automatically invoke a machine learning model to expand the terrain description in response to the first input meeting preset conditions. For example, such preset conditions may include the text length of the first input being less than a preset threshold (e.g., 50 characters). Alternatively, such preset conditions may include the first input lacking key terrain features (e.g., terrain distribution, etc.). Alternatively, the expansion operation of the first input can be triggered manually by the user.

[0096] In box 406, electronic device 110 or server 130 can decompose the terrain type of the first input (or an expanded version of the first input) to determine one or more terrain elements. For example, electronic device 110 or server 130 can perform semantic parsing on the first input to obtain terrain information corresponding to the first input. For example, such terrain information can characterize one or more terrain elements associated with the 3D virtual scene to be constructed. Taking a first input that includes text content (e.g., "Generate a map of an island with beaches, rivers, and volcanoes") as an example, the terrain elements corresponding to this first input can include "beach," "river," "valley," and "volcano," etc.

[0097] In box 408, electronic device 110 or server 130 can generate a first image. For example, electronic device 110 or server 130 can generate the first image by invoking a pre-trained machine learning model (e.g., a generative model). For example, the first image can also be referred to as a terrain semantic map. As an example, the input to the machine learning model can include the first input, the terrain type obtained through semantic parsing (e.g., "beach", "river", "valley", "volcano"), and optional spatial relationship constraints (e.g., "rivers flow from high to low", "beach should not be in a volcano", etc.). As an example, the input to the machine learning model can also include mappings between various terrain elements and various styles. For example, different types of terrain elements can correspond to different styles (e.g., different colors or texture fills, etc.). Taking different styles indicating different colors as an example, "beach" can correspond to yellow, "river" can correspond to blue, "valley" can correspond to green, "volcano" can correspond to red, etc. These mappings can be provided to the machine learning model in a structured manner as part of cue words. As an example, the output of the machine learning model can include the first image. As an example, the first image can represent the two-dimensional terrain distribution of a three-dimensional virtual scene.

[0098] As an example, the first image can be a two-dimensional raster image. The first image can include various graphic elements (e.g., pixels or pixel blocks). As an example, these various graphic elements correspond to different types of terrain elements. These various graphic elements have different styles (e.g., different colors or texture fills). For example, different graphic elements can correspond to different colors, thus allowing the spatial distribution of various terrain elements to be visually indicated through color distribution. For example, the first image can clearly show the location, extent, and spatial relationships of various terrain elements in a three-dimensional virtual scene by using yellow areas to represent beaches, blue areas to represent rivers, green areas to represent valleys, and red areas to represent volcanoes.

[0099] In box 410, electronic device 110 or server 130 can also generate a scene rendering map corresponding to the first image (e.g., a terrain semantic map). For example, electronic device 110 or server 130 can utilize a pre-trained generative model to generate a scene rendering map corresponding to the first image based on the first input (or an expanded first input), the terrain type obtained through semantic parsing, and / or the first image. As an example, the scene rendering map can correspond to a top-down view of a 3D virtual scene. In some cases, the scene rendering map can be a visual enhancement or stylized representation of the first image (e.g., a terrain semantic map) to make it closer to a real geographical environment. For example, the "beach" area marked in yellow in the terrain semantic map can be rendered as a sandy texture with a grainy feel in the top-down rendering map. The "river" area marked in blue can be rendered as a flowing water surface effect. The "valley" area marked in green can be presented as a slope covered with vegetation. The "volcano" area marked in red can be given details such as crater shadows or lava colors.

[0100] In box 412, electronic device 110 or server 130 can generate a depth map (also known as a terrain height map) corresponding to the first image. As an example, the depth map can characterize the height of different terrain elements corresponding to the first image in a 3D virtual scene. As an example, the depth map can be represented as a single-channel grayscale image or a two-dimensional numerical matrix. For example, each pixel in the depth map corresponds to a numerical value used to quantify the height of the terrain element at the pixel's location. For example, this value can be normalized to the range of 0 to 1. The closer the value is to 1, the higher the height of the terrain element in the area corresponding to the pixel (e.g., ridge, peak). The closer the value is to 0, the lower the height of the terrain element in the area corresponding to the pixel (e.g., valley, basin, lake bottom). In this way, the depth map can visually express vertical dimension information that cannot be represented by a two-dimensional terrain semantic map, providing a reference for subsequent 3D virtual scene construction.

[0101] Alternatively, electronic device 110 or server 130 can generate a depth map based on the scene rendering map. For example, electronic device 110 or server 130 can perform depth inference on the scene rendering map to obtain a depth map. For example, electronic device 110 or server 130 can use a pre-trained machine learning model (e.g., a depth estimation model) to output a normalized depth value corresponding to each pixel in the scene rendering map, thereby constructing a depth map (i.e., a terrain height map). Compared to directly generating a depth map based on the terrain semantic map, this alternative approach can infer more detailed height information of surface elements by utilizing rich lighting, shadow, and / or texture cues in the scene rendering map, thus helping to maintain consistency between the depth map and the visual representation.

[0102] In box 414, electronic device 110 or server 130 can generate a scene preview. For example, the scene preview can represent the 3D terrain of a 3D virtual scene. For example, the 3D terrain can be obtained based on a first image. Alternatively, the 3D terrain can be obtained by fusing the first image, a scene rendering, and / or a depth map. As an example, the process of constructing the 3D terrain may include: using the terrain elements indicated by the first image (such as beaches, rivers, volcanoes, valleys, etc.) as a semantic basis, overlaying texture details and stylized presentations (such as sand texture, lava sheen, etc.) from the scene rendering, and then assigning vertical dimension information corresponding to each location based on the height values ​​in the depth map, thereby generating the 3D terrain. This scene preview can be presented on a user interface (e.g., Figure 3B The interface (300B) allows users to intuitively observe the overall shape, material distribution, and lighting effects of the 3D virtual scene before finalizing the terrain plan or making subsequent edits, thereby improving the interactive experience and construction accuracy.

[0103] In box 416, electronic device 110 or server 130 can generate first structured data based on the second image. As an example, electronic device 110 or server 130 can utilize a pre-trained machine learning model to identify and semantically understand terrain elements in the second image, scene rendering, and / or depth map. This machine learning model can determine the terrain elements (e.g., valleys, deserts, rivers, etc.) included in the 3D virtual scene based on the distribution, contours, and spatial relationships of each terrain element, and decide on the procedural content generation templates to be applied. Such templates may include pre-packaged, reusable procedural content generation algorithms or combinations of algorithms. Each template may indicate the processing logic for the corresponding terrain element (e.g., water erosion, noise overlay, wind erosion, etc.). As an example, the first structured data may indicate one or more templates corresponding to one or more terrain elements. Alternatively, the first structured data may also include parameter configurations required for each template to facilitate subsequent construction of the 3D virtual scene. As an example, if the machine learning model determines that the mountains in the 3D virtual scene are high and have a steep slope, the "water erosion" template can be enabled in the first structured data to perform fine-grained terrain sculpting. Alternatively, if the machine learning model determines that there are large flat areas in the 3D virtual scene, the "noise overlay" template can be enabled in the first structured data to increase the sense of natural undulation.

[0104] In box 418, electronic device 110 or server 130 can utilize a terrain generation tool to generate a first 3D terrain based on the first structured data. As an example, based on the first structured data, a template type (e.g., a template indicating the terrain element type) and parameter configuration corresponding to the terrain element are determined. Further, at least one submap corresponding to the terrain element can be generated based on the template type and parameter configuration. As an example, the terrain generation tool may include a software module or subsystem encapsulating a terrain construction algorithm from a procedural content generation (PCG) workflow, whose inputs are the first structured data (containing terrain element types, corresponding templates, and parameters) and a depth map (terrain height map), and whose output may include 3D terrain. As an example, the terrain generation tool may sequentially perform steps such as dynamic submap generation, detail sculpting, mask calculation (bump mapping and slope mapping), terrain blending map calculation, and pixel-wise weighted blending of multiple repeatable textures, ultimately generating a 3D terrain with geometric details and realistic texture.

[0105] As an example, electronic device 110 or server 130 can utilize terrain generation tools to dynamically generate at least one sub-map for multiple terrain elements based on specified terrain elements (such as valleys, volcanoes, beaches, rivers, etc.) in the first structured data, corresponding templates (such as water erosion, noise overlay, etc.), and a depth map. Furthermore, a first three-dimensional terrain can be produced based on at least one sub-map. For example, electronic device 110 or server 130 can perform detail enhancement processing on at least one sub-map. For example, on the sub-map corresponding to the water erosion template, gullies and riverbed markings can be enhanced. Radial folds or lava flow traces can be added to volcanic areas. The steepness of ridges can be strengthened. In this way, the visual features of the terrain elements can be highlighted at a geometric level.

[0106] Alternatively, mask information corresponding to at least one submap can be generated (e.g., the mask information can be associated with a normal map and / or a slope distribution map) to generate a first 3D terrain based on at least one submap and the mask information. Additionally, during the mask calculation stage, the terrain generation tool can perform mask calculation operations based on a depth map (i.e., a terrain height map) and the submap after detail sculpting. For example, this operation includes at least two types: bump mapping and slope mapping. Bump mapping generates a corresponding normal map by resolving the normal direction of each pixel on the height map, which is used to simulate the microscopic bumps and depressions of the surface during subsequent lighting rendering. Slope mapping calculates the slope value at each location based on the height gradient, and then outputs a slope distribution map (e.g., represented as a grayscale image, with higher grayscale values ​​for higher slopes).

[0107] Additionally, in the step of generating the terrain blending map, the terrain generation tool integrates the masks obtained in the previous stages (including bump mapping normal maps and slope distribution maps), the sub-maps corresponding to each terrain element, and the semantic information in the first structured data to calculate the terrain blending map. This terrain blending map is stored, for example, in the form of a two-dimensional image, where each pixel location can record a set of weight vectors, which correspond to the blending ratio of various repeatable surface maps (such as volcanoes, rivers, valleys, beaches, etc.) at that location.

[0108] Additionally, the terrain generation tool can generate a first 3D terrain based on a terrain blending map. As an example, the terrain generation tool can perform pixel-by-pixel weighted blending of multiple sets of repeatable surface maps (such as maps corresponding to volcanoes, maps corresponding to valleys, etc.) based on the weight vector stored in each pixel of the terrain blending map. The blending result is combined with the geometric mesh defined by the depth map (terrain height map), and surface lighting details are enhanced through normal maps generated by bump mapping, ultimately outputting a 3D terrain that can be rendered and interacted with in real time.

[0109] In some cases, electronic device 110 or server 130 can construct a three-dimensional virtual scene based on the first three-dimensional terrain. In box 420, electronic device 110 or server 130 can determine whether the first three-dimensional terrain meets the requirements. Alternatively, electronic device 110 or server 130 can acquire at least one scene image based on the first three-dimensional terrain. For example, at least one scene image can correspond to at least one viewpoint of the first three-dimensional terrain. For example, at least one scene image can correspond to a screenshot of at least one viewpoint of the first three-dimensional terrain. For example, at least one viewpoint can include, but is not limited to, a top-down viewpoint, a side viewpoint, an oblique viewpoint, and any surrounding viewpoint, thereby being able to cover the overall outline and local details of the three-dimensional terrain. Further, electronic device 110 or server 130 can generate first information (also known as evaluation information, e.g., a score) for the first three-dimensional terrain based on at least one scene image. As an example, this first information can take the form of a quantitative score (e.g., an overall quality score from 0 to 100), or it can include multiple dimensions of subdivided indicators, such as terrain realism, texture clarity, terrain semantic conformity (whether it matches the description entered by the user), visual diversity, etc. This first information can be automatically generated by analyzing multi-view scene images through a pre-trained evaluation model, or it can be weighted and adjusted based on user feedback.

[0110] Additionally, the electronic device 110 or the server 130 can construct a three-dimensional virtual scene corresponding to the first three-dimensional terrain in response to the first information meeting preset conditions. As an example, the conditions may include: the quantitative score of the first information exceeding a certain threshold (e.g., the overall score is greater than 90 points), or the various subdivided dimension indicators (e.g., terrain realism, texture clarity, semantic conformity) are not lower than the corresponding thresholds.

[0111] Alternatively, electronic device 110 or server 130 may, in response to the first information not meeting the conditions, generate second structured data based on the second image. For example, electronic device 110 or server 130 may re-execute the steps in block 416 to generate the second structured data. Further, electronic device 110 or server 130 may re-execute the steps in block 418, using a terrain generation tool to generate a second three-dimensional terrain based on the second structured data. Further, electronic device 110 or server 130 may construct a three-dimensional virtual scene based on the second three-dimensional terrain. For example, electronic device 110 or server 130 may generate second first information corresponding to the second three-dimensional terrain. Further, electronic device 110 or server 130 may, in response to the second first information meeting the conditions, construct a three-dimensional virtual scene based on the second three-dimensional terrain. If the second first information still does not meet the conditions, iteration may continue until the requirements are met or the maximum number of attempts is reached.

[0112] In box 422, electronic device 110 or server 130 can construct a 3D virtual scene with 3D terrain. For example, electronic device 110 can construct a 3D virtual scene based on 3D terrain that meets certain conditions (e.g., a first 3D terrain or a second 3D terrain). As an example, electronic device 110 or server 130 can convert a depth map into a 3D mesh and generate colliders for it. Electronic device 110 or server 130 can construct multi-layered materials (e.g., beach, valley, volcano, river, etc.) using terrain blending maps and normal maps. Alternatively, electronic device 110 or server 130 can add environmental components such as directional lights, sky spheres, and water bodies. Alternatively, electronic device 110 or server 130 can configure physical materials for corresponding terrain elements (e.g., low friction for beach, high damage for lava), thereby enabling the generated 3D virtual scene to have a certain degree of interactivity.

[0113] In some cases, the process of constructing a 3D virtual scene mentioned above only involves the construction of terrain elements within the 3D virtual scene. In other cases, scene elements within the 3D virtual scene can also be constructed.

[0114] In box 424, electronic device 110 can receive a second input. As an example, the second input may include a description of surface objects input by the user. For instance, the second input may describe scene elements included in a 3D virtual scene in natural language. Unlike the first input, which corresponds to terrain (such as islands, rivers, volcanoes, and other macroscopic landforms), the second input corresponds to specific objects on the surface. As an example, such scene elements may include: vegetation (such as trees, flowerbeds, grasslands, shrubs, etc.), man-made structures (such as houses, bridges, roads, squares, lighthouses, etc.), and / or natural objects (such as scattered rocks, fallen logs, pebbles, etc.).

[0115] In box 426, electronic device 110 or server 130 can expand the second input (e.g., a user-input description of a surface object). For example, electronic device 110 or server 130 can expand the second input using a machine learning model. Such a machine learning model has the ability to expand text content, generating richer and more coherent natural language text based on brief descriptions or keywords. For example, such a machine learning model can semantically expand the original surface object description, making the expanded surface object description more complete and suitable for subsequent construction of a 3D virtual scene. This approach is not intended to limit the training process or specific implementation of this machine learning model.

[0116] Alternatively, the expansion operation of the second input can be triggered manually by the user or automatically by the electronic device 110 or the server 130. For example, the electronic device 110 or the server 130 can automatically invoke a machine learning model to expand the terrain description in response to the second input meeting preset conditions. For example, such preset conditions may include the text length of the second input being less than a preset threshold (e.g., 50 characters). Alternatively, such preset conditions may include the second input lacking key scene element information (e.g., the distribution of scene elements, etc.). Alternatively, the expansion operation of the second input can be triggered manually by the user.

[0117] In box 428, electronic device 110 or server 130 can decompose the scene elements represented by the second input (or an expanded version of the second input) to determine one or more scene elements. For example, electronic device 110 or server 130 can perform semantic analysis on the second input to determine one or more scene elements included in the 3D virtual scene to be constructed. Taking the second input as including text content (e.g., "There are coconut trees and stones in the valley") as an example, the scene elements corresponding to this second input may include "coconut trees" and "stones," etc.

[0118] As an example, a user can associate with a pre-defined set of objects. This pre-defined set of objects can include at least one virtual object associated with the user. For example, such a virtual object could include virtual objects in the user's associated asset library (or resource library). As an example, a virtual object refers to a specific material or 3D model (e.g., a coconut tree model, a stone model, a house model, etc.) that can be used for rendering in a 3D virtual scene, and it can contain meshes, textures, materials, and metadata. For example, such virtual objects can have various sources: those created and uploaded by the user, general assets provided by the platform, or those obtained by the user through redemption or other means. As an example, the virtual objects associated with the user are also referred to as the user's assets or resources.

[0119] In some cases, virtual objects can correspond to scene elements. As an example, a scene element (such as a "coconut tree") is a semantically abstract concept representing the type of object (such as vegetation, buildings, natural objects, etc.) that needs to be placed in a 3D virtual scene. Virtual objects are the actual model files used for rendering. For example, a virtual object corresponding to a scene element can be implemented in multiple forms. For instance, the scene element "coconut tree" can be implemented as different 3D models of "coconut trees." The same scene element can be associated with multiple virtual objects of different forms. For example, when generating a single instance, one form can be selected for instantiation. For example, the scene element "coconut tree" can be associated with a tall coconut tree model, a short coconut tree model, a coconut tree model with fruit, etc. When placing a coconut tree in a 3D virtual scene, one of the models can be selected for generation based on preset rules or constraints indicated by a second input.

[0120] In box 430, electronic device 110 or server 130 can semantically retrieve assets corresponding to the second input. For example, electronic device 110 or server 130 can semantically match multiple scene elements (e.g., "coconut tree," "stone," etc.) described by the second input with at least one virtual object from a predetermined set of objects. As an example, electronic device 110 or server 130 can utilize natural language understanding technology to calculate the semantic similarity between the scene element and the descriptive information (e.g., tags, names, etc.) of the at least one virtual object. As an example, electronic device 110 or server 130 can determine that the scene element matches the corresponding virtual object in response to a semantic similarity greater than a preset threshold. Conversely, electronic device 110 or server 130 can determine that the scene element does not match the corresponding virtual object in response to a semantic similarity less than a preset threshold.

[0121] In box 432, electronic device 110 or server 130 can generate a fourth image. For example, electronic device 110 or server 130 can generate the fourth image by invoking a pre-trained machine learning model (e.g., a generative model). For example, the fourth image can also be referred to as a scene element semantic map. As an example, the input to the machine learning model can include a second input, multiple scene elements obtained through semantic parsing (e.g., "coconut tree," "stone"), and optional spatial relationship constraints (e.g., "coconut trees cannot grow on stones," "coconut trees should not be in rivers," etc.). As an example, the input to the machine learning model can also include mapping relationships between multiple scene elements and multiple graphic elements. For example, different scene elements can correspond to different styles (e.g., different colors or texture fills, etc.). Taking different styles indicating different colors as an example, "coconut tree" can correspond to green, "stone" can correspond to black, etc. These mapping relationships can be provided to the machine learning model in a structured manner as part of cue words. As an example, the output of the machine learning model can include the fourth image. As an example, the fourth image can represent the distribution of multiple scene elements in a two-dimensional plane. The two-dimensional plane corresponds to a three-dimensional virtual scene. For example, the two-dimensional plane can correspond to the top-down view of the three-dimensional virtual scene.

[0122] As an example, the fourth image can be a two-dimensional raster image. The fourth image can include various graphic elements (e.g., pixels or pixel blocks). As an example, these various graphic elements correspond to multiple scene elements. These various graphic elements have different styles (e.g., different colors or texture fills). For example, different graphic elements can correspond to different colors, thus allowing the spatial distribution of various scene elements to be visually indicated through color distribution. For example, the fourth image can use green areas to represent coconut trees and black areas to represent stones, thus clearly showing the distribution of various scene elements in a two-dimensional plane (e.g., location, extent, and positional relationships).

[0123] In box 434, electronic device 110 or server 130 can also generate a scene rendering map corresponding to the fourth image (e.g., a semantic map of surface objects). For example, electronic device 110 or server 130 can utilize a pre-trained generative model to generate a scene rendering map corresponding to the fourth image based on a second input (or an expanded second input), multiple scene elements obtained through semantic parsing, a first image (e.g., a terrain semantic map), and / or the fourth image. As an example, the scene rendering map can correspond to a top-down view of a 3D virtual scene. In some cases, the scene rendering map can be a visual enhancement or stylized representation of the fourth image (e.g., a semantic map of surface objects) to make it closer to a real geographical environment. For example, graphic elements in the fourth image (e.g., graphic elements representing coconut trees) can be replaced with images with lighting and texture details (e.g., an image of a coconut tree with shadows).

[0124] In box 436, electronic device 110 or server 130 can generate a third scene preview. For example, this third scene preview is used to display multiple virtual objects (i.e., specific 3D model instances) placed in a 3D virtual scene based on a fourth image (a semantic map of surface objects). Each virtual object corresponds to a scene element (such as a coconut tree, a rock, etc.). The distribution of multiple virtual objects in the 3D virtual scene is determined based on the distribution of multiple scene elements in a two-dimensional plane.

[0125] In box 438, electronic device 110 or server 130 can generate third structured data based on the fourth image. As an example, the third structured data can guarantee how to place virtual objects corresponding to multiple scene elements. As an example, electronic device 110 or server 130 can utilize a pre-trained machine learning model to identify and semantically understand scene elements in the fourth image and / or scene rendering. Based on the scene element types (such as coconut trees, stones, etc.) represented by different color or texture regions in the fourth image, the types of virtual objects to be instantiated in the 3D virtual scene are determined, and the procedural content generation template to be applied is decided. Such templates may include pre-packaged, reusable procedural content generation algorithms or combinations of algorithms. Each template can indicate the processing logic for the corresponding scene element (e.g., the "forest spotting" template can be used to randomly place trees within a specified area, the "building spotting" template can be used to place houses according to a grid or road layout, etc.). As an example, the third structured data can indicate one or more templates corresponding to one or more scene elements. For example, the scene element "coconut trees" can use the "forest spotting" template (e.g., setting a density of 0.05 trees per square meter). Alternatively, the third structured data can also include the parameter configurations required for each template (such as placement area, density, disabling overlap, etc.), thereby providing precise automated guidance for the subsequent construction of a complete 3D virtual scene. Furthermore, a 3D virtual scene can be constructed based on the first structured data.

[0126] Alternatively, in box 440, electronic device 110 or server 130 can obtain a first three-dimensional virtual scene based on third structured data. For example, electronic device 110 or server 130 can obtain the first three-dimensional virtual scene through a procedural content generation process based on the third structured data. For example, electronic device 110 or server 130 can parse the template type (such as "forest sprinkle", "building sprinkle", etc.) and its parameter configuration (such as density, placement area, disabling overlap, etc.) specified in the third structured data. Further, electronic device 110 or server 130 can generate a corresponding sprinkle submap for each scene element based on a fourth image (surface object semantic map), template type and its parameter configuration. For example, the sprinkle submap may include the spatial coordinates of each instantiated location. Further, electronic device 110 or server 130 can place virtual objects corresponding to the corresponding scene elements in the three-dimensional virtual scene based on the spatial coordinates indicated by the sprinkle submap, thereby obtaining the first three-dimensional virtual scene.

[0127] In box 442, electronic device 110 or server 130 can generate first information (e.g., a rating) based on at least one scene image of the first three-dimensional virtual scene. For example, electronic device 110 or server 130 can acquire at least one scene image based on the first three-dimensional virtual scene. For example, at least one scene image can correspond to at least one viewpoint of the first three-dimensional virtual scene. For example, at least one scene image can correspond to a screenshot of at least one viewpoint of the first three-dimensional terrain. For example, at least one viewpoint can include, but is not limited to, a top-down viewpoint, a side viewpoint, an oblique viewpoint, and an arbitrary surround viewpoint, thereby being able to cover the overall outline and local details of the first three-dimensional virtual scene. Further, electronic device 110 or server 130 can generate first information (e.g., a rating) for the first three-dimensional virtual scene based on at least one scene image. As an example, the first information can take the form of a quantitative rating (e.g., an overall quality score from 0 to 100), or it can include multiple dimensions of subdivided indicators, such as visual realism (evaluating lighting, texture, and model details), semantic conformity (checking whether the object type and distribution conform to the description of the second input), spatial plausibility (verifying whether there is penetration or violation of constraints between objects and between objects and terrain), etc. This initial information can be automatically generated by analyzing multi-view scene images through a pre-trained evaluation model, or it can be weighted and adjusted based on user feedback.

[0128] Additionally, the electronic device 110 or the server 130 may output a first three-dimensional virtual scene in response to the first information meeting a condition. As an example, the condition may include: the quantitative score of the first information exceeds a certain threshold (e.g., the overall score is greater than 90 points), or the subdivided dimension indicators (e.g., visual realism, semantic conformity, spatial rationality) are not lower than the corresponding thresholds.

[0129] Alternatively, electronic device 110 or server 130 may generate fourth structured data based on the fourth image in response to the first information not meeting the condition. For example, electronic device 110 or server 130 may re-execute the steps in block 238 to generate the fourth structured data. Further, electronic device 110 or server 130 may re-execute the steps in block 240 to construct a three-dimensional virtual scene (e.g., a second three-dimensional virtual scene) using the fourth structured data. Alternatively, electronic device 110 or server 130 may generate first information corresponding to the second three-dimensional virtual scene. Further, electronic device 110 or server 130 may output the second three-dimensional virtual scene in response to the first information meeting the condition. If the first information still does not meet the condition, iteration may continue until the requirement is met or the maximum number of attempts is reached.

[0130] In box 444, electronic device 110 or server 130 can construct the final 3D virtual scene. This 3D virtual scene can be a first 3D virtual scene constructed based on third structured data, or a second 3D virtual scene constructed based on fourth structured data after iterative optimization (e.g., regeneration in response to the first information not meeting conditions). In some cases, the constructed 3D virtual scene includes 3D terrain and multiple virtual objects (e.g., surface objects) distributed on the 3D terrain. As an example, the 3D terrain is generated based on the first input mentioned above. Multiple virtual objects (such as coconut trees, rocks, etc.) are generated based on the second input and its corresponding structured data and a point-scattering process. By integrating the terrain generation and surface object generation chains, a complete 3D virtual scene with 3D terrain and virtual objects can be constructed.

[0131] Based on the process described above, this system supports generating a two-dimensional image (e.g., a first image) representing the two-dimensional terrain distribution based on a first input. It also supports flexible editing of the first image using editing components to obtain a second image. This second image can be used to construct a three-dimensional virtual scene with three-dimensional terrain. Furthermore, it supports generating a fourth image representing the distribution of multiple scene elements on a two-dimensional plane based on the second input. This fourth image can be used to place virtual objects in the three-dimensional virtual scene. This provides a virtual scene construction method that combines input generation and editing adjustments, thus enriching the ways to construct virtual scenes. On the one hand, it allows users to intuitively modify the first or fourth image through editing components, making the final constructed three-dimensional virtual scene more tailored to the user's needs, thereby significantly improving the customization and visual quality of the virtual scene. On the other hand, compared to purely manual modeling or fully automated generation, this solution organically integrates input generation and interactive editing, lowering the initial construction threshold and reducing the iterative costs of repeated adjustments, thereby significantly improving the construction efficiency of virtual scenes while ensuring scene quality. Furthermore, since the editing operations are applied directly to the two-dimensional image, users do not need to deal with the complex three-dimensional modeling process, which further reduces the difficulty of operation and the learning cost for users.

[0132] Example process Figure 5 A flowchart of an example process 500 for constructing a virtual scene based on certain scenarios is shown. Process 500 can be implemented at electronic device 110. See below for reference. Figure 1 To describe process 500.

[0133] like Figure 5 As shown, in box 510, electronic device 110 can receive a first input, which is related to the terrain information of the three-dimensional virtual scene.

[0134] In frame 520, electronic device 110 can present an editing component for editing a first image, which is generated based on a first input and represents the two-dimensional terrain distribution of a three-dimensional virtual scene.

[0135] In box 530, electronic device 110 can construct a three-dimensional virtual scene based on a second image, which is obtained based on an editing component.

[0136] In some cases, process 500 also includes: in the editing component, presenting a first image, the first image including multiple graphic elements, the multiple graphic elements corresponding to different types of terrain elements, the multiple graphic elements having different styles.

[0137] In this way, since the first image already contains a variety of graphic elements corresponding to different types of terrain elements, and these graphic elements have different styles (such as different colors or texture fills), users can quickly identify the spatial distribution and boundaries of various terrain elements through different graphic elements, thereby reducing the cognitive burden on users and improving the efficiency of terrain editing.

[0138] In some cases, process 500 further includes: receiving an editing operation on a first image via an editing component to edit a two-dimensional terrain distribution; and presenting a second image representing the edited two-dimensional terrain distribution.

[0139] In this way, users can directly modify the two-dimensional terrain distribution in the editing component and see a second image reflecting the editing results in real time, thus realizing a WYSIWYG interactive experience. This effectively reduces repeated trial and error caused by the inability to preview in real time, and improves the intuitiveness and efficiency of terrain editing.

[0140] In some cases, receiving an editing operation on a first image via an editing component to edit a two-dimensional terrain distribution includes: receiving an editing operation, the editing operation including selecting a first region in the first image; and replacing a graphic element in the first region with a first graphic element corresponding to a first terrain element.

[0141] In this way, users can modify the two-dimensional terrain distribution by selecting a specific area in the first image and replacing the graphic elements therein, thereby indirectly adjusting the corresponding area of ​​the three-dimensional terrain. This provides an intuitive editing method that uses a two-dimensional semantic map to drive changes in the three-dimensional terrain, avoiding the complexity of directly manipulating the three-dimensional model and improving the feasibility and efficiency of terrain editing.

[0142] In some cases, receiving an edit operation includes: presenting multiple controls in an edit component, the multiple controls corresponding to multiple terrain elements; and receiving an edit operation based on a triggering of a first control among the multiple controls, the first control corresponding to a first terrain element.

[0143] In this way, users can quickly specify the terrain type to be used for editing by triggering the first control corresponding to the first terrain element, thereby avoiding the tedious steps of repeatedly setting the terrain type manually during the drawing process and improving the efficiency and intuitiveness of terrain editing.

[0144] In some cases, process 500 further includes: in response to a triggering of the first control, presenting a prompt message in the editing component, the prompt message representing a second region in the first image, the second region corresponding to the first terrain element.

[0145] In this way, triggering the first control displays a prompt message in the editing component, which visually shows the second area in the first image corresponding to the first terrain element. This allows users to quickly understand the existing distribution of the terrain element without manual searching, providing a clear reference for subsequent editing operations (such as adding, modifying, or deleting specific areas), thus improving the targeting and efficiency of terrain editing.

[0146] In some cases, process 500 further includes: presenting first descriptive text about terrain information, wherein the first image is generated based on the first descriptive text, and the first descriptive text is obtained based on the first input.

[0147] In this way, the user's first input can be transformed into readable first descriptive text, and a first image representing the two-dimensional terrain distribution can be generated based on the descriptive text. This allows the user to intuitively understand the natural language description corresponding to the current terrain, which is convenient for subsequent iteration and adjustment of the terrain by modifying the text. This achieves a two-way association between text and image, and improves the interpretability and controllability of terrain generation.

[0148] In some cases, process 500 further includes: editing the first descriptive text to obtain the second descriptive text; and presenting a third image generated based on the second descriptive text.

[0149] In some cases, process 500 may also include: presenting a scene preview, which represents the three-dimensional terrain of the three-dimensional virtual scene, the three-dimensional terrain being obtained based on the first image.

[0150] In this way, a 3D terrain preview generated based on the first image can be presented, allowing users to intuitively observe the undulations, material distribution, and lighting effects of the terrain in 3D space. This eliminates the need to switch to a separate rendering module to verify the editing results, improving the feedback efficiency and decision-making accuracy in the terrain design process.

[0151] In some cases, presenting a scene preview includes: presenting a first preview image of the three-dimensional terrain; and, in response to receiving an adjustment operation for the scene preview, presenting a second preview image of the three-dimensional terrain, wherein the first preview image and the second preview image correspond to different viewing angles.

[0152] In some cases, 3D terrain is constructed based on the following process: generating a depth map corresponding to a first image, the depth map representing the height of different terrain elements corresponding to the first image in a 3D virtual scene; and generating 3D terrain based on the first image and the depth map.

[0153] In this way, a depth map containing height information can be generated based on the first image (two-dimensional terrain semantic map), and the two can be merged to construct a three-dimensional terrain. This automatically transforms the planar terrain distribution into a three-dimensional form with realistic undulations, ensuring that the representation of various terrain elements (such as rivers, volcanoes, and beaches) in the height dimension is consistent with the semantic labels, thus improving the automation and geometric accuracy of terrain generation.

[0154] In some cases, generating a depth map corresponding to the first image based on the first image includes: generating a scene rendering image based on the first image, the scene rendering image corresponding to a top-down view of the three-dimensional virtual scene; and generating a depth map based on the scene rendering image.

[0155] In this way, the scene rendering map can be used as an intermediate bridge to infer depth information from the top-down rendering map, thereby generating a terrain height map without relying on additional height annotations, which enhances the visual consistency of depth estimation.

[0156] In some cases, the scene preview is a first scene preview, and process 500 further includes: in response to editing the first image into a second image through an editing component, presenting a second scene preview, the second scene preview corresponding to the second image.

[0157] In this way, after the user modifies the first image into the second image through the editing component, the corresponding second scene preview is presented, so that the user can observe in real time how the changes in the two-dimensional terrain distribution affect the shape and details of the three-dimensional terrain. This achieves a seamless connection between editing and feedback, and significantly improves the intuitiveness and iteration efficiency of terrain adjustment.

[0158] In some cases, a three-dimensional virtual scene is obtained based on the following process: generating first structured data based on a second image; generating first three-dimensional terrain based on the first structured data using a terrain generation tool; and constructing a three-dimensional virtual scene based on the first three-dimensional terrain.

[0159] In this way, the edited second image can be converted into structured procedural generation instructions, and the first 3D terrain can be automatically constructed using terrain generation tools, thereby completing the creation of the entire 3D virtual scene. In this way, users can directly drive the procedural generation of 3D terrain from a 2D semantic map without manual modeling or pixel-by-pixel adjustments, achieving automation and reproducibility of terrain construction.

[0160] In some cases, generating a first three-dimensional terrain based on the first structured data includes: determining a template type and parameter configuration corresponding to multiple terrain elements based on the first structured data; generating at least one submap corresponding to the multiple terrain elements based on the template type and parameter configuration; and generating the first three-dimensional terrain based on the at least one submap.

[0161] In some cases, generating a first three-dimensional terrain based on at least one sub-map includes: generating mask information corresponding to at least one sub-map; and generating the first three-dimensional terrain based on at least one sub-map and the mask information.

[0162] In some cases, constructing a three-dimensional virtual scene based on a first three-dimensional terrain includes: acquiring at least one scene image based on the first three-dimensional terrain; generating first information for the first three-dimensional terrain based on the at least one scene image; and constructing a three-dimensional virtual scene corresponding to the first three-dimensional terrain in response to the first information satisfying a condition.

[0163] In this way, multi-view images of the first three-dimensional terrain can be automatically acquired and quantified first information can be generated. The final three-dimensional virtual scene is only constructed when the first information meets the preset conditions. This achieves the automatic assessment of terrain quality by using the first information, avoiding invalid construction and resource waste caused by terrain not meeting the requirements, and improving the reliability and efficiency of the scene construction process.

[0164] In some cases, constructing a three-dimensional virtual scene based on the first three-dimensional terrain further includes: generating second structured data based on the second image in response to the first information not meeting the conditions; generating second three-dimensional terrain based on the second structured data using a terrain generation tool; and constructing a three-dimensional virtual scene based on the second three-dimensional terrain.

[0165] In this way, when the first information does not meet the conditions, the generation of second structured data and the reconstruction of the second three-dimensional terrain can be automatically triggered, thus forming a closed-loop iterative optimization mechanism. The terrain quality can be continuously improved without manual intervention by the user until the preset requirements are met, which effectively improves the robustness of terrain generation and the reliability of the final scene.

[0166] In some cases, the first input includes at least one of the following: natural language input; at least one piece of media content.

[0167] In some cases, the first image is generated by the following process: semantic parsing of the first input to determine terrain information, including terrain type and spatial relationship constraints; and generating the first image based on the terrain information.

[0168] In some cases, process 500 further includes: expanding the first input to obtain an expanded first input; and generating a first image based on the expanded first input.

[0169] In some cases, different styles may include different colors or different texture fills.

[0170] Example devices and equipment A corresponding apparatus for implementing the above methods or processes is also provided. Figure 6 A schematic structural block diagram of an example device 600 for constructing virtual scenes according to some scenarios is shown. Device 600 can be implemented as or included in electronic device 110. The various modules / components in device 600 can be implemented by hardware, software, firmware, or any combination thereof.

[0171] like Figure 6 As shown, the device 600 includes: a first receiving module 610 configured to receive a first input, the first input being related to terrain information of a three-dimensional virtual scene; a first rendering module 620 configured to render an editing component, the editing component being used to edit a first image, the first image being generated based on the first input, the first image representing the two-dimensional terrain distribution of the three-dimensional virtual scene; and a construction module 630 configured to construct a three-dimensional virtual scene based on a second image, the second image being obtained based on the editing component.

[0172] In some cases, the device 600 also includes a second presentation module configured to present a first image in an editing component. The first image includes multiple graphic elements corresponding to different types of terrain elements, and the multiple graphic elements have different styles.

[0173] In some cases, the device 600 further includes a second receiving module configured to: receive, via an editing component, an editing operation on the first image to edit the two-dimensional terrain distribution; and present a second image representing the edited two-dimensional terrain distribution.

[0174] In some cases, the second receiving module is also configured to: receive an editing operation, the editing operation including selecting a first region in the first image; and replace a graphic element in the first region with a first graphic element corresponding to a first terrain element.

[0175] In some cases, the second receiving module is also configured to: present multiple controls in the editing component, the multiple controls corresponding to multiple terrain elements; and receive editing operations based on the triggering of a first control among the multiple controls, the first control corresponding to a first terrain element.

[0176] In some cases, the device 600 also includes a prompting module configured to: in response to a triggering of the first control, present a prompting message in the editing component, the prompting message representing a second region in the first image, the second region corresponding to the first terrain element.

[0177] In some cases, the device 600 also includes a third presentation module configured to present first descriptive text about terrain information, wherein the first image is generated based on the first descriptive text, and the first descriptive text is obtained based on the first input.

[0178] In some cases, the device 600 also includes a text editing module configured to: edit the first descriptive text to obtain the second descriptive text; and present a third image generated based on the second descriptive text.

[0179] In some cases, the device 600 also includes a first preview module configured to: present a scene preview, the scene preview representing the three-dimensional terrain of the three-dimensional virtual scene, the three-dimensional terrain being obtained based on the first image.

[0180] In some cases, the first preview module is also configured to: present a first preview image of the three-dimensional terrain; and, in response to receiving an adjustment operation for the scene preview, present a second preview image of the three-dimensional terrain, the first preview image and the second preview image corresponding to different viewing angles.

[0181] In some cases, 3D terrain is constructed based on the following process: generating a depth map corresponding to a first image, the depth map representing the height of different terrain elements corresponding to the first image in a 3D virtual scene; and generating 3D terrain based on the first image and the depth map.

[0182] In some cases, generating a depth map corresponding to the first image based on the first image includes: generating a scene rendering image based on the first image, the scene rendering image corresponding to a top-down view of the three-dimensional virtual scene; and generating a depth map based on the scene rendering image.

[0183] In some cases, the scene preview is a first scene preview, and the device 600 also includes a second preview module configured to: in response to editing the first image into a second image via an editing component, present a second scene preview, the second scene preview corresponding to the second image.

[0184] In some cases, a three-dimensional virtual scene is obtained based on the following process: generating first structured data based on a second image; generating first three-dimensional terrain based on the first structured data using a terrain generation tool; and constructing a three-dimensional virtual scene based on the first three-dimensional terrain.

[0185] In some cases, generating a first three-dimensional terrain based on the first structured data includes: determining a template type and parameter configuration corresponding to multiple terrain elements based on the first structured data; generating at least one submap corresponding to the multiple terrain elements based on the template type and parameter configuration; and generating the first three-dimensional terrain based on the at least one submap.

[0186] In some cases, generating a first three-dimensional terrain based on at least one sub-map includes: generating mask information corresponding to at least one sub-map; and generating the first three-dimensional terrain based on at least one sub-map and the mask information.

[0187] In some cases, constructing a three-dimensional virtual scene based on a first three-dimensional terrain includes: acquiring at least one scene image based on the first three-dimensional terrain; generating first information for the first three-dimensional terrain based on the at least one scene image; and constructing a three-dimensional virtual scene corresponding to the first three-dimensional terrain in response to the first information satisfying a condition.

[0188] In some cases, constructing a three-dimensional virtual scene based on the first three-dimensional terrain further includes: generating second structured data based on the second image in response to the first information not meeting the conditions; generating second three-dimensional terrain based on the second structured data using a terrain generation tool; and constructing a three-dimensional virtual scene based on the second three-dimensional terrain.

[0189] In some cases, the first input includes at least one of the following: natural language input; at least one piece of media content.

[0190] In some cases, the first image is generated by the following process: semantic parsing of the first input to determine terrain information, including terrain type and spatial relationship constraints; and generating the first image based on the terrain information.

[0191] In some cases, the device 600 also includes an amplification module configured to: amplify the first input to obtain an amplified first input; and generate a first image based on the amplified first input.

[0192] In some cases, different styles may include different colors or different texture fills.

[0193] The modules included in device 600 can be implemented in various ways, including software, hardware, firmware, or any combination thereof. In some cases, one or more modules can be implemented using software and / or firmware, such as machine-executable instructions stored on a storage medium. In addition to or as an alternative to machine-executable instructions, some or all of the units in device 600 can be implemented at least partially by one or more hardware logic components. By way of example, and not limitation, exemplary types of hardware logic components that can be used include field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard parts (ASSPs), systems on a chip (SOCs), complex programmable logic devices (CPLDs), and so on.

[0194] Figure 7 A block diagram of an electronic device 700 in which one or more examples may be implemented is shown. It should be understood that... Figure 7 The electronic device 700 shown is merely exemplary and should not be construed as limiting the functionality and scope of the examples described herein. Figure 7 The illustrated electronic device 700 can be used to implement the electronic device 110 discussed above.

[0195] like Figure 7 As shown, electronic device 700 is in the form of a general-purpose electronic device. Components of electronic device 700 may include, but are not limited to, one or more processing units or processors 710, memory 720, storage devices 730, one or more communication units 740, one or more input devices 750, and one or more output devices 760. Processor 710 may be a physical or virtual processor and is capable of performing various processes according to programs stored in memory 720. In a multiprocessor system, multiple processors execute computer-executable instructions in parallel to improve the parallel processing capability of electronic device 700.

[0196] Electronic device 700 typically includes multiple computer storage media. Such media can be any accessible media that is accessible to electronic device 700, including but not limited to volatile and non-volatile media, removable and non-removable media. Memory 720 can be volatile memory (e.g., registers, cache, random access memory (RAM)), non-volatile memory (e.g., read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory), or some combination thereof. Storage device 730 can be removable or non-removable media and can include machine-readable media, such as flash drives, disks, or any other media that can be used to store information and / or data and can be accessed within electronic device 700.

[0197] Electronic device 700 may further include additional removable / non-removable, volatile / non-volatile storage media. Although not explicitly stated... Figure 7 As shown, disk drives for reading from or writing to removable, non-volatile disks (e.g., "floppy disks") and optical disk drives for reading from or writing to removable, non-volatile optical disks can be provided. In these cases, each drive can be connected to a bus (not shown) via one or more data media interfaces. Memory 720 may include computer program product 725 having one or more program modules configured to perform various methods or actions of various examples.

[0198] The communication unit 740 enables communication with other electronic devices via a communication medium. Additionally, the functionality of the components of the electronic device 700 can be implemented using a single computing cluster or multiple computing machines capable of communicating via communication connections. Therefore, the electronic device 700 can operate in a networked environment using logical connections to one or more other servers, networked personal computers, or another network node.

[0199] Input device 750 can be one or more input devices, such as a mouse, keyboard, trackball, etc. Output device 760 can be one or more output devices, such as a monitor, speaker, printer, etc. Electronic device 700 can also communicate with one or more external devices (not shown) via communication unit 740 as needed. These external devices include storage devices, display devices, etc., and can communicate with one or more devices that enable user interaction with electronic device 700, or with any device that enables electronic device 700 to communicate with one or more other electronic devices (e.g., network card, modem, etc.). Such communication can be performed via input / output (I / O) interface (not shown).

[0200] A computer-readable storage medium is provided that stores computer-executable instructions thereon, wherein the computer-executable instructions are executed by a processor to implement the methods described above. A computer program product is also provided, which is tangibly stored on a non-transitory computer-readable medium and includes computer-executable instructions, which are executed by a processor to implement the methods described above.

[0201] The flowcharts and / or block diagrams of the methods, apparatus, devices, and computer program products referred to herein describe various aspects. It should be understood that each block of the flowcharts and / or block diagrams, as well as combinations of blocks in the flowcharts and / or block diagrams, can be implemented by computer-readable program instructions.

[0202] These computer-readable program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing apparatus to produce a machine such that, when executed by the processor of the computer or other programmable data processing apparatus, they create means for implementing the functions / actions specified in one or more blocks of the flowchart and / or block diagram. These computer-readable program instructions can also be stored in a computer-readable storage medium that causes a computer, programmable data processing apparatus, and / or other device to operate in a particular manner; thus, the computer-readable medium storing the instructions comprises an article of manufacture that includes instructions for implementing aspects of the functions / actions specified in one or more blocks of the flowchart and / or block diagram.

[0203] Computer-readable program instructions can be loaded onto a computer, other programmable data processing apparatus, or other device to cause a series of operational steps to be performed on the computer, other programmable data processing apparatus, or other device to produce a computer-implemented process, thereby causing the instructions that execute on the computer, other programmable data processing apparatus, or other device to perform the functions / actions specified in one or more boxes of a flowchart and / or block diagram.

[0204] The flowcharts and block diagrams in the accompanying figures illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products under various scenarios. In this respect, each block in a flowchart or block diagram may represent a module, segment, or portion of an instruction, which contains one or more executable instructions for implementing the specified logical function. In some alternative implementations, the functions marked in the blocks may occur in a different order than those shown in the figures. For example, two consecutive blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented using a dedicated hardware-based system that performs the specified function or action, or using a combination of dedicated hardware and computer instructions.

[0205] Various examples have been described above. The foregoing descriptions are exemplary and not exhaustive, nor are they limited to the disclosed implementations. Many modifications and variations will be apparent to those skilled in the art without departing from the scope and spirit of the described implementations. The terminology used herein is chosen to best explain the principles, practical applications, or improvements to technology in the market, or to enable others skilled in the art to understand the various implementations disclosed herein.

Claims

1. A method for constructing a virtual scene, comprising: Receive a first input, which is related to the terrain information of the three-dimensional virtual scene; An editing component is presented, which is used to edit a first image, the first image being generated based on the first input, and the first image representing the two-dimensional terrain distribution of the three-dimensional virtual scene; as well as The three-dimensional virtual scene is constructed based on the second image, which is obtained based on the editing component.

2. The method according to claim 1, further comprising: In the editing component, the first image is presented, which includes a variety of graphic elements, each corresponding to a different type of terrain element, and each graphic element having a different style.

3. The method according to claim 2, further comprising: The editing component receives editing operations on the first image to edit the two-dimensional terrain distribution; as well as The second image is presented, which represents the edited two-dimensional terrain distribution.

4. The method of claim 3, wherein receiving an editing operation on the first image via the editing component to edit the two-dimensional terrain distribution comprises: Receive the editing operation, the editing operation including the selection of a first region in the first image; as well as Replace the graphic elements in the first region with a first graphic element, which corresponds to a first terrain element.

5. The method of claim 4, wherein receiving the editing operation comprises: The editing component presents multiple controls, which correspond to various terrain elements; as well as The editing operation is received based on the triggering of the first control among the plurality of controls, wherein the first control corresponds to the first terrain element.

6. The method according to claim 5, further comprising: In response to the triggering of the first control, a prompt message is presented in the editing component, the prompt message representing a second region in the first image, the second region corresponding to the first terrain element.

7. The method according to claim 1, further comprising: The first image is generated based on the first descriptive text, which is obtained based on the first input.

8. The method of claim 7, further comprising: Edit the first description text to obtain the second description text; as well as A third image is generated based on the second descriptive text.

9. The method according to claim 1, further comprising: A scene preview is presented, which represents the three-dimensional terrain of the three-dimensional virtual scene, and the three-dimensional terrain is obtained based on the first image.

10. The method of claim 9, wherein presenting the scene preview includes: Presenting a first preview image of the three-dimensional terrain; as well as In response to receiving an adjustment operation for the scene preview, a second preview image of the three-dimensional terrain is presented, wherein the first preview image and the second preview image correspond to different viewing angles.

11. The method of claim 9, wherein the three-dimensional terrain is constructed based on the following process: Based on the first image, a depth map corresponding to the first image is generated, wherein the depth map represents the height of different terrain elements corresponding to the first image in the three-dimensional virtual scene; and The three-dimensional terrain is generated based on the first image and the depth map.

12. The method of claim 11, wherein generating a depth map corresponding to the first image based on the first image comprises: Based on the first image, a scene rendering map is generated, the scene rendering map corresponding to the top view of the three-dimensional virtual scene; as well as The depth map is generated based on the scene rendering map.

13. The method according to claim 9, wherein the scene preview is a first scene preview, and the method further includes: In response to editing the first image into the second image via the editing component, a second scene preview is presented, the second scene preview corresponding to the second image.

14. The method according to claim 1, wherein the three-dimensional virtual scene is obtained based on the following process: Based on the second image, generate the first structured data; Using a terrain generation tool, a first three-dimensional terrain is generated based on the first structured data; and Based on the first three-dimensional terrain, the three-dimensional virtual scene is constructed.

15. The method of claim 14, wherein constructing the three-dimensional virtual scene based on the first three-dimensional terrain comprises: Based on the first three-dimensional terrain, acquire at least one scene image; Based on the at least one scene image, generate first information for the first three-dimensional terrain; as well as In response to the condition being met by the first information, a three-dimensional virtual scene corresponding to the first three-dimensional terrain is constructed.

16. The method according to claim 15, wherein constructing the three-dimensional virtual scene based on the first three-dimensional terrain further comprises: In response to the first information not satisfying the condition, second structured data is generated based on the second image; Using terrain generation tools, a second three-dimensional terrain is generated based on the second structured data; as well as Based on the second three-dimensional terrain, the three-dimensional virtual scene is constructed.

17. An apparatus for constructing a virtual scene, comprising: The first receiving module is configured to receive a first input, which is related to the terrain information of the three-dimensional virtual scene; A first presentation module is configured to present an editing component, the editing component being used to edit a first image, the first image being generated based on the first input, the first image representing the two-dimensional terrain distribution of the three-dimensional virtual scene; as well as The building module is configured to construct the three-dimensional virtual scene based on a second image obtained from the editing component.

18. An electronic device, comprising: At least one processor; as well as At least one memory coupled to the at least one processor and storing instructions for execution by the at least one processor, the instructions causing the electronic device to perform the method according to any one of claims 1 to 16 when executed by the at least one processor.

19. A computer-readable storage medium having stored thereon computer-executable instructions that can be executed by a processor to implement the method according to any one of claims 1 to 16.

20. A computer program product tangibly stored in a computer storage medium and comprising computer-executable instructions that, when executed by a device, cause the device to perform the method according to any one of claims 1 to 16.