Method, device and equipment for generating three-dimensional scene, medium and product

By generating depth maps and three-dimensional grid models corresponding to the target image, the three-dimensional scenes are quickly constructed, which solves the problems of complex production and long periods in the existing technology, and realizes efficient three-dimensional scene generation.

CN120495524APending Publication Date: 2025-08-15BEIJING ZITIAO NETWORK TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510592879.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-08
Publication Date
2025-08-15

AI Technical Summary

Technical Problem

When generating three-dimensional scenes, especially three-dimensional scenes with live broadcast backgrounds, the prior art has problems such as complex production, long cycles and low efficiency, making it difficult to produce quickly in batches.

Method used

By determining the target image, generate corresponding depth maps and three-dimensional grid models, combine the image and model to generate three-dimensional scenes, and use the image depth map and deep learning algorithm to quickly build three-dimensional scenes.

Benefits of technology

It significantly improves the mass production capacity of three-dimensional scenes, shortens the production cycle, and improves the user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120495524A_ABST
    Figure CN120495524A_ABST
Patent Text Reader

Abstract

The embodiment of the invention relates to a method and device for generating a three-dimensional scene, equipment, a medium and a product. The method includes determining a target image available to generate a three-dimensional scene. The method further includes determining a depth map corresponding to the target image based on the target image. The method further includes determining a three-dimensional mesh model corresponding to the target image based on the depth map and the target image. The method further comprises the step of determining a three-dimensional scene corresponding to the target image based on the target image and the three-dimensional grid model. Through the method, the three-dimensional grid model is generated by using the depth map of the image and the image, and the image and the three-dimensional grid model are combined to obtain the three-dimensional scene corresponding to the image, so that the three-dimensional scene can be generated more quickly, the batch production capacity of the three-dimensional scene is improved, and the user experience is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Embodiments of the present disclosure generally relate to the field of information processing, and more particularly to methods, devices, equipment, media, and products for generating three-dimensional scenes. Background Art

[0002] Artificial Intelligence Generated Content (AIGC) technology is rapidly developing in the field of digital content creation. AI-driven automated content generation is demonstrating strong potential in areas such as virtual scene construction, digital human creation, and interactive media production. This technology leverages deep learning algorithms to rapidly produce high-quality content, significantly improving digital content production efficiency and bringing new possibilities to the creative industry.

[0003] With the continuous advancement of AIGC technology, its capabilities in real-time rendering and multimodal fusion are constantly improving. In particular, in the field of three-dimensional (3D) content generation, technical solutions based on neural rendering and generative adversarial networks are becoming increasingly mature, enabling end-to-end generation from text and images to 3D models. These technological advances are driving the rapid development of emerging fields such as virtual image production, injecting new vitality into the digital content industry. Summary of the Invention

[0004] Embodiments of the present disclosure provide a method, apparatus, device, medium, and product for generating a three-dimensional scene.

[0005] According to a first aspect of the present disclosure, a method for generating a three-dimensional scene is provided. The method includes determining a target image that can be used to generate the three-dimensional scene. The method also includes determining, based on the target image, a depth map corresponding to the target image. The method also includes determining, based on the depth map and the target image, a three-dimensional mesh model corresponding to the target image. The method also includes determining, based on the target image and the three-dimensional mesh model, a three-dimensional scene corresponding to the target image.

[0006] According to a second aspect of the present disclosure, an apparatus for generating a three-dimensional scene is provided. The apparatus includes a target image determination module configured to determine a target image that can be used to generate the three-dimensional scene; a depth map determination module configured to determine a depth map corresponding to the target image based on the target image; a three-dimensional mesh model determination module configured to determine a three-dimensional mesh model corresponding to the target image based on the depth map and the target image; and a three-dimensional scene determination module configured to determine a three-dimensional scene corresponding to the target image based on the target image and the three-dimensional mesh model.

[0007] In a third aspect of the present disclosure, an electronic device is provided, comprising at least one processor; and a storage device for storing at least one program, wherein when the at least one program is executed by the at least one processor, the at least one processor implements the method according to the first aspect of the present disclosure.

[0008] In a fourth aspect of the present disclosure, a computer-readable storage medium is provided, on which a computer program is stored. When the program is executed by a processor, the method according to the first aspect of the present disclosure is implemented.

[0009] In a fifth aspect of the present disclosure, a computer program product is provided, comprising a computer program, which implements the method according to the first aspect of the present disclosure when executed by a processor.

[0010] It should be understood that the content described in this content section is not intended to limit the key or important features of the embodiments of the present disclosure, nor is it intended to limit the scope of the present disclosure. Other features of the present disclosure will become easy to understand through the following description. BRIEF DESCRIPTION OF THE DRAWINGS

[0011] The above and other objects, features and advantages of the present disclosure will become more apparent through a more detailed description of exemplary embodiments of the present disclosure with reference to the accompanying drawings, wherein like reference numerals generally represent like components throughout the exemplary embodiments of the present disclosure.

[0012] Figure 1 A schematic diagram illustrating an example environment in which the apparatus and / or methods of some embodiments of the present disclosure may be implemented;

[0013] Figure 2 A schematic diagram illustrating a method for generating a three-dimensional scene according to some embodiments of the present disclosure is illustrated;

[0014] Figure 3 A schematic diagram illustrating another example method for generating a three-dimensional scene according to some embodiments of the present disclosure is illustrated;

[0015] Figure 4 A schematic diagram illustrating text generation of an image according to some embodiments of the present disclosure is illustrated;

[0016] Figure 5 FIG2 illustrates a schematic diagram of extracting image depth information according to some embodiments of the present disclosure;

[0017] Figure 6 A schematic diagram illustrating three-dimensional mesh model conversion according to some embodiments of the present disclosure is shown;

[0018] Figure 7 A schematic diagram illustrating generating a three-dimensional scene based on a model and a texture or a dynamic texture according to some embodiments of the present disclosure

[0019] Figure 8 A schematic diagram illustrating a workflow of image processing according to some embodiments of the present disclosure is shown;

[0020] Figure 9 A schematic diagram illustrating a workflow of generating a video from an image according to some embodiments of the present disclosure is provided;

[0021] Figure 10 A schematic diagram illustrating a workflow for image quality improvement according to some embodiments of the present disclosure is shown;

[0022] Figure 11 A schematic diagram illustrating a workflow for extracting depth information according to some embodiments of the present disclosure is illustrated;

[0023] Figure 12 A schematic diagram illustrating an interface for generating and exporting a model according to some embodiments of the present disclosure is shown;

[0024] Figure 13 A schematic block diagram of an apparatus for generating a three-dimensional scene according to some embodiments of the present disclosure is illustrated;

[0025] Figure 14 Illustrated is a schematic block diagram of an example device suitable for implementing various embodiments of the present disclosure. DETAILED DESCRIPTION

[0026] It is understandable that the data involved in this technical solution (including but not limited to the data itself, the acquisition or use of the data) must comply with the requirements of relevant laws, regulations and relevant provisions.

[0027] It is understandable that before using the technical solutions disclosed in the various embodiments of this disclosure, the type, scope of use, usage scenarios, etc. of the personal information involved in this disclosure should be informed to the user and the user's authorization should be obtained in an appropriate manner in accordance with relevant laws and regulations.

[0028] For example, upon receiving a user's active request, a prompt message is sent to the user to clearly inform the user that the requested operation will require the acquisition and use of the user's personal information. This allows the user to independently choose whether to provide personal information to the electronic device, application, server, storage medium, or other software or hardware that performs the operations of the disclosed technical solution based on the prompt message.

[0029] As an optional but non-limiting implementation, in response to receiving a user's active request, the prompt information may be sent to the user in the form of a pop-up window, in which the prompt information may be presented in text form. Furthermore, the pop-up window may also contain a selection control for the user to select "agree" or "disagree" to provide personal information to the electronic device.

[0030] It is understandable that the above notification and user authorization process are merely illustrative and do not limit the implementation of the present disclosure. Other methods that comply with relevant laws and regulations may also be applied to the implementation of the present disclosure.

[0031] The following describes embodiments of the present disclosure in more detail with reference to the accompanying drawings. Although certain embodiments of the present disclosure are shown in the accompanying drawings, it should be understood that the present disclosure can be implemented in various forms and should not be construed as limited to the embodiments described herein. Rather, these embodiments are provided to provide a more thorough and complete understanding of the present disclosure. It should be understood that the drawings and embodiments of the present disclosure are for illustrative purposes only and are not intended to limit the scope of protection of the present disclosure.

[0032] In the description of the embodiments of the present disclosure, the term "including" and similar terms should be understood as open inclusion, that is, "including but not limited to." The term "based on" should be understood as "based at least in part on." The term "one embodiment" or "the embodiment" should be understood as "at least one embodiment." The terms "first," "second," etc. may refer to different or the same objects. Other explicit and implicit definitions may also be included below.

[0033] Generally, in live broadcast scenarios, the combination of virtual and real (real-life footage + virtual background) is a relatively new live broadcast method. By creating different two-dimensional (2D) and three-dimensional (3D) virtual backgrounds, the host's live broadcast effect can be enriched and the number of viewers can be increased. However, there are some problems with using 2D virtual backgrounds for live broadcasts. For example, the user's observation angle is limited, only simple forward and backward camera movements can be made, and the presentation format is relatively simple. Using virtual 3D backgrounds can avoid these problems. However, using 3D virtual backgrounds for live broadcasts also has some problems. For example, although the scene expression is relatively good, the production of a single 3D scene is relatively complex and takes weeks or even months to produce. Therefore, the production cycle is long, resulting in low production efficiency and the inability to quickly mass-produce.

[0034] To this end, an embodiment of the present disclosure proposes a scheme for generating a three-dimensional scene. In this scheme, a computing device can first determine a target image that can be used to generate a three-dimensional scene. Then, using the target image, a depth map corresponding to the target image can be determined. Then, the computing device can further process the depth information and the target image in the depth map to generate a three-dimensional mesh model corresponding to the target image. Finally, the target scene corresponding to the target image is generated by further processing the target image and the three-dimensional mesh model. Through this method, the depth map of the image and the image are used to generate a three-dimensional mesh model, and the image and the three-dimensional mesh model are combined to obtain a three-dimensional scene corresponding to the image, so that the three-dimensional scene can be generated more quickly, the mass production capacity of the three-dimensional scene is improved, and the user experience is improved.

[0035] The embodiments of the present disclosure will be described in detail below with reference to the accompanying drawings. Figure 1 An example environment in which the devices and / or methods of embodiments of the present disclosure may be implemented is shown. In environment 100 , a computing device 102 may be configured to generate a three-dimensional scene 110 corresponding to a target image 104 .

[0036] Examples of computing device 102 include, but are not limited to, personal computers, server computers, handheld or laptop devices, mobile devices (such as mobile phones, personal digital assistants (PDAs), media players, etc.), multi-processor systems, consumer electronics, minicomputers, mainframe computers, distributed computing environments that include any of the above systems or devices, and the like.

[0037] like Figure 1 To construct the three-dimensional scene 110, the computing device 104 needs to obtain a two-dimensional target image 104. This target image 104 can be an image that a user wants to use as a live broadcast background. For example, the target image is an image of a live broadcast stage. The user wishes to use the image of the live broadcast stage to generate a corresponding three-dimensional stereoscopic stage for their use.

[0038] Next, computing device 102 may process target image 104 to extract depth information corresponding to the image, thereby forming a corresponding depth map 106. The depth information for each pixel in depth map 106 represents the distance between an object in the image and the camera or observer, adding a third dimension of spatial information to the two-dimensional image. The depth map provides additional depth information for the target image.

[0039] After obtaining the depth map of the target image, computing device 102 can further use depth map 106 to generate a three-dimensional mesh model corresponding to the target image. At this point, computing device 102 can further use target image 104 and depth map 106 to construct a three-dimensional mesh model 108 for objects in target image 104. Each object in target image 104 can be represented by a three-dimensional mesh model 108. For example, if target image 104 shows a stage, a three-dimensional mesh model 108 of the stage can be constructed using the depth map corresponding to target image 104 and the target image.

[0040] Next, the computing device 102 further processes the target image 104 and the generated three-dimensional mesh model 108 to generate a three-dimensional scene corresponding to the target image 104. In one example, the target image 104 and the three-dimensional mesh model 108 can be directly combined to generate a static three-dimensional scene. For example, the target image 104 for the stage and the three-dimensional mesh model corresponding to the stage are combined to form a three-dimensional stage as a live scene. In another example, a video can be generated based on the target image 104, and then the video is combined with the three-dimensional mesh model to form a dynamic three-dimensional scene. For example, a video of the stage is generated using the target image 104 for the stage, and the video dynamically displays the stage, including rotating the stage. Then, the video is combined with the three-dimensional mesh model of the stage to generate a dynamic display of the three-dimensional stage as a live scene.

[0041] Through this method, the depth map of the image and the image are used to generate a three-dimensional mesh model, and the image and the three-dimensional mesh model are combined to obtain a three-dimensional scene corresponding to the image, so that the three-dimensional scene can be generated more quickly, the mass production capacity of the three-dimensional scene is improved, and the user experience is improved.

[0042] Combined with the above Figure 1 A schematic diagram of an example environment in which the devices and / or methods of some embodiments of the present disclosure may be implemented is described below. Figure 2 A schematic diagram illustrating a method for generating a three-dimensional scene according to some embodiments of the present disclosure. Figure 2 The method in can be Figure 1 Executed by computing device 102 or any suitable computing device.

[0043] like Figure 2As described above, in the example method 200, at box 202, the computing device 102 determines a target image that can be used to generate a three-dimensional scene. In order to generate a three-dimensional scene, the computing device 102 needs to obtain a target image. The target image can be obtained in a variety of ways. In one example, an image can be provided by a user who is broadcasting live, such as an image taken with a camera. In another example, the computing device 102 can use a machine learning model to process input text to generate a target image. For example, the computing device 102 can receive text input by a user, such as a prompt word for describing a target image. The computing device 102 then inputs the input text into the machine learning model to generate a target image. The machine learning model can also be referred to as a first machine learning model. For example, the first machine learning model can be a text-based image model.

[0044] After acquiring the target image 104, the computing device 102 may further process the target image 104. For example, the user-provided image may be captured by a camera, but some cameras may have suboptimal hardware parameters, resulting in blurry images, low quality, unclear regions of interest, and other issues. To address these issues, the computing device 102 may perform image super-resolution processing on the target image to update the target image or increase its resolution. Image super-resolution processing involves restoring a given low-resolution target image to a corresponding high-resolution image through predetermined operations. For example, digital image processing, computer vision, and other technologies, combined with predetermined processing algorithms, can be used to reconstruct a low-resolution image into a high-resolution image.

[0045] At block 204 , the computing device 102 determines a depth map corresponding to the target image based on the target image. To generate a 3D scene, the computing device 102 needs to further obtain a depth map corresponding to the target image to increase depth information related to the 3D scene.

[0046] There are multiple implementations for obtaining a depth map of a target image. In one example, the computing device 102 may first obtain a pre-trained machine learning model for generating a depth map from an image. For ease of description, this machine learning model is also referred to as a second machine learning model. The second machine learning model may be trained using pre-labeled sample images and sample depth maps. The computing device 102 may then input the target image 104 into the second machine learning model to obtain a depth map. The second machine learning model may be any suitable machine learning model.

[0047] In another example, when determining a depth map corresponding to a target image, the computing device 102 may first extract depth information from the target image 104. For example, the depth information for each pixel is extracted from the target image. The computing device 102 then uses this depth information to generate a depth map for the target image. For example, the target image is input into a loaded model, which extracts and analyzes features of the image, learns the texture, edge, and context information in the image, and predicts the depth value corresponding to each pixel. Afterwards, the depth value output by the model is converted to a suitable data type and range, and is usually normalized to the range of 0-255 to facilitate subsequent processing and display. Finally, a depth map is created based on these normalized depth values. The grayscale value of each pixel in the map represents the depth information of the corresponding position of the pixel in the original image. Brighter pixels indicate closer distances, and darker pixels indicate farther distances.

[0048] At block 206, computing device 102 determines a three-dimensional mesh model corresponding to the target image based on the depth map and the target image. After obtaining the depth map, computing device 102 further combines the depth map and the target image to generate a three-dimensional mesh model of the object in the target image.

[0049] In some embodiments, the computing device 102 may input the depth map and the target image into a pre-trained machine learning model to generate a three-dimensional mesh model corresponding to the target image. The machine learning model may be any suitable model.

[0050] In some embodiments, computing device 102 may first extract depth information for each pixel in the depth map and determine the position of each point on the object's surface in three-dimensional space based on the depth value. For example, brighter pixels in the depth map correspond to points closer to the camera and positioned forward in three-dimensional space, while darker pixels correspond to points farther from the camera and positioned backward in three-dimensional space. Then, based on the color and texture information of the target image, the points in three-dimensional space are assigned corresponding color and material attributes. For example, if a certain area in the target image is red, the points at the corresponding location in three-dimensional space are assigned a red material. Next, based on the position information of each point in the depth map, these points can be connected to form a triangular mesh. For example, adjacent points can be connected according to certain rules to construct the surface of the object. For example, adjacent points with similar depths in the depth map are also adjacent in three-dimensional space. These points are connected to form triangles, and the combination of multiple triangles forms the surface mesh of the object. During the process of constructing the triangular mesh, texture mapping can also be performed on the mesh based on the texture information of the target image, so that the generated three-dimensional mesh model has an appearance similar to the original image.

[0051] At block 208, computing device 102 determines a three-dimensional scene corresponding to the target image based on the target image and the three-dimensional mesh model. After obtaining the three-dimensional mesh model corresponding to the target image, computing device 102 further utilizes the target image and the three-dimensional mesh model to generate a three-dimensional scene corresponding to the target image.

[0052] In some embodiments, computing device 102 directly combines the target image into the 3D mesh model to generate a 3D scene corresponding to the target image. For example, the target image may be combined into the 3D mesh model as a texture to generate a static 3D scene corresponding to the target image.

[0053] In some embodiments, the computing device 102 may further utilize the target image to generate a video corresponding to the target image. In one example, the computing device 102 applies the target image to a third machine learning model to generate a video corresponding to the target image. In another example, the computing device 102 may also obtain another image corresponding to the target image, and then perform an interpolation operation on the target image and the other image to generate multiple video frames to form a video corresponding to the target image. The computing device 102 then uses the video and the three-dimensional mesh model to determine a three-dimensional scene corresponding to the target image. For example, each video frame in the video is combined with the three-dimensional mesh model to generate a dynamic three-dimensional scene. Additionally, the computing device 102 may also perform image super-resolution processing on the video frames in the obtained video to update the video.

[0054] Through this method, the depth map of the image and the image are used to generate a three-dimensional mesh model, and the image and the three-dimensional mesh model are combined to obtain a three-dimensional scene corresponding to the image, so that the three-dimensional scene can be generated more quickly, the mass production capacity of the three-dimensional scene is improved, and the user experience is improved.

[0055] Combined with the above Figure 2 A schematic diagram of a method for generating a three-dimensional scene according to some embodiments of the present disclosure is described below. Figure 3 A flowchart depicting another example method for generating a three-dimensional scene according to some embodiments of the present disclosure. Figure 3 The example method 300 in Figure 1 The computing device 102 or any suitable computing device in the system may process the data.

[0056] First, the computing device 102 receives image material in common formats such as jpg or png as input at box 302. These image materials can be generated using AIGC technology or obtained from existing image libraries, such as photographic images or digital drawings. These images have good versatility and controllability, and their preparation time is usually within minutes. Subsequently, the computing device 102 processes the input image using any suitable depth estimation model. Such a model can accurately predict the depth information of each pixel from a single or multiple images, and then generate a corresponding depth map at box 304. This step is computationally efficient, with processing time controlled to seconds, providing an accurate geometric foundation for subsequent 3D reconstruction.

[0057] After obtaining the depth map, the computing device uses a 3D reconstruction tool at block 306 to convert the depth information into a 3D mesh model. This conversion process can also be completed in seconds. The generated mesh model is output in a predetermined format for subsequent editing and use. Finally, at block 308, the computing device imports the processed 3D model into a mainstream rendering engine. The entire import and scene construction process takes approximately ten minutes, significantly shortening the production cycle compared to traditional manual modeling methods.

[0058] Furthermore, computing device 102 integrates AIGC technology at block 310 to implement image-to-video functionality. By inputting the first and last keyframes, the computing device automatically generates intermediate frames and synthesizes continuous, natural dynamic video content, simulating complex animation effects such as camera roll, perspective rotation, and lighting changes, greatly enriching the expressiveness and immersiveness of the virtual live broadcast scene. By combining the generated video with a 3D mesh model at block 308, a dynamic 3D scene can be generated.

[0059] Through this method, a fully automated process from static images to complete three-dimensional virtual scenes is achieved, covering key links such as image input, depth prediction, three-dimensional reconstruction, model engine import, and video generation. The cost and time of the entire process are compressed to a few minutes, significantly improving the efficiency and scalability of content generation and improving the user experience.

[0060] Combined with the above Figure 3 A flowchart of another exemplary method for generating a three-dimensional scene according to some embodiments of the present disclosure is described below; Figure 4 A schematic diagram illustrating text-generated images according to some embodiments of the present disclosure. Figure 4 The example process 400 may be performed by Figure 1 The system may be executed by the computing device 102 shown in FIG. 1 or any suitable device.

[0061] Example 400 demonstrates a method for generating a high-tech stage scene using an AI drawing tool. As shown, computing device 102 uses precise prompts to achieve targeted scene generation. A positive prompt can specify a "high-tech stage, cool tones, with a holographic screen in the center" or "cyberpunk style, no people, cool tones, primarily blue-purple colors, with a logo in the center," ensuring high consistency with the target style. This prompt mechanism effectively guides the AI model to generate background material that meets the requirements of the live broadcast scenario, providing a high-quality 2D flat material foundation for subsequent 3D scene conversion. The images generated by the AI tool undergo a unified screening process, and the optimal results are used for depth map generation and 3D conversion. This solution significantly improves efficiency compared to traditional manual background drawing, with a single generation time of just minutes. It also supports rapid iteration of different styles by adjusting the prompts, providing a large number of easily accessible and consistent basic materials.

[0062] The following combination Figure 5 A schematic diagram describing the extraction of image depth information according to some embodiments of the present disclosure. In example 500, the computing device 102 uses a workflow for generating a depth map to extract image depth information. Specifically, the computing device first builds a customized workflow through a visual node editor and integrates the depth estimation model into the processing pipeline. When a two-dimensional image is input, the depth estimation model performs real-time depth information calculation and generates a corresponding depth map, in which the near-view area of the image is represented by a lighter grayscale value and the distant view area is presented with a darker grayscale value. This AI-based depth extraction method has significant advantages over traditional stereo vision algorithms. It can not only process monocular image input, but also maintain high accuracy and robustness. The generation of the depth map serves as the basic data for subsequent three-dimensional mesh reconstruction, and its quality directly affects the realism and stereoscopic effect of the final 3D scene. Through the graphical interface, users can intuitively monitor the processing process and adjust the parameter configuration as needed to ensure that the depth extraction results meet the needs of different application scenarios.

[0063] The following combination Figure 6A schematic diagram describing the conversion of a three-dimensional mesh model according to some embodiments of the present disclosure. In example 600, the computing device 102 converts the depth map generated previously into an editable three-dimensional mesh model. Specifically, the computing device first imports the depth map output by the depth estimation model and the original color image into the conversion model, and automatically generates a three-dimensional mesh surface with a spatial topological structure based on the depth information of the image through the depth-to-mesh function of the three-dimensional modeling auxiliary plug-in. The generated three-dimensional mesh model not only retains the visual features of the original image, but also has a reasonable geometric structure and can support secondary editing and optimization in the creative software. The algorithm optimization of this step achieves a conversion speed of seconds, which is greatly improved compared to the traditional manual modeling method, and provides key technical support for the rapid construction of three-dimensional scenes. Finally, a model in a predetermined format is output, which can be directly imported into a real-time rendering engine to complete subsequent scene construction and dynamic mapping production.

[0064] The following combination Figure 7 A schematic diagram describing the generation of a three-dimensional scene based on a model and a map or a dynamic map according to some embodiments of the present disclosure. In example 700, the computing device 102 uses a model and map intelligent integration technology to implement an efficient three-dimensional scene construction process. The computing device first imports the generated optimized three-dimensional model in a predetermined format to ensure that the performance requirements can be met while maintaining visual accuracy, and uses the dynamic video sequence generated by AIGC as a texture input to achieve real-time rendering and playback of the video map, and supports the adjustment of a full range of related parameters. At the same time, static maps can also be directly used as texture input to complete the production of the scene. In the scene synthesis stage, a visual programming system is used to implement the control of the interactive logic, so that non-technical personnel can quickly complete the adjustment of the scene parameters, and finally output a complete virtual scene that supports high-resolution live streaming.

[0065] The following combination Figure 8 A schematic diagram describing the workflow of image processing according to some embodiments of the present disclosure. In example 800, it is mainly divided into key processing steps such as loading a super-resolution model, loading an image, super-resolution image processing, image magnification, and storing an image. At 802, the computing device loads a specified super-resolution model, which can be optimized for the scene image first, supports four times magnification, and effectively suppresses shadows while maintaining sharpness. At box 804, the uploaded image is loaded, and mask local processing is supported to achieve selective enhancement. Then, at box 806, the super-resolution model is received together with the loaded image and input to 808, and then the target image is proportionally magnified according to the custom magnification ratio at 808, and finally the magnified image is stored at 810.

[0066] The following combination Figure 9The diagram illustrates a workflow for generating videos from images in some embodiments of the present disclosure. In example 900, the computing device 102 first loads the clip encoder and configures basic parameters, including selecting a predefined super-resolution algorithm, setting a 0.25x scaling factor, and specifying a primary processing device. The workflow uses a 1920*1080 resolution image as the input source and implements content parsing through a multimodal encoding processing system. The visual encoding uses a corresponding model in conjunction with a variational autoencoder with bf16 precision, while the text encoding integrates a predefined model to process scene description text. The computing device is equipped with a sophisticated parameter control system, including defining a 768×432 generation size, configuring a loop parameter set, and a fragment block processing strategy for the semantic segmentation model. The video sampler optimizes content by changing multiple parameters. The workflow also supports a rich set of post-processing functions, including visual enhancement options such as camera motion effects, spatial distortion, and color adjustment, as well as stylized processing such as static images and subtitle overlays. The entire workflow uses bf16 precision to balance computational efficiency and output quality, and ensures content consistency through the collaborative work of multiple models. Its modular design supports full-process processing from material input to 4K super-resolution output.

[0067] The following combination Figure 10 A schematic diagram depicting a workflow for image quality enhancement according to some embodiments of the present disclosure. In example 1000, the computing device 102 first loads a pre-trained super-resolution model specifically designed to achieve high-quality four-fold image magnification. The processing flow uses a video file as the input source, and the video specification is a standard frame rate of 30 frames per second. The system deframes the video through a processing module, converting the continuous video stream into a single-frame image sequence, and then calls a variational autoencoder to perform feature extraction and enhancement processing on each frame. In the core processing link, the super-resolution model performs four-fold resolution enhancement on each frame while maintaining the details and clarity of the original picture. The system supports user-defined processing parameters, including options such as setting processing intensity (initial value is 0) and loop count. After processing is completed, the system can re-synthesize the enhanced image sequence into a high-resolution video output and support resynchronization with the audio track. The entire workflow realizes the automated conversion from standard resolution video to high-definition video, while maintaining the original frame rate, and can also increase the video resolution to four times the original size, effectively enhancing the video effect.

[0068] The following combination Figure 11The figure illustrates a schematic diagram of a workflow for extracting depth information according to some embodiments of the present disclosure. In example 1100, the computing device 102 first loads a depth estimation model, which is specifically used to extract accurate depth information from a single image. The processing flow uses a 1920*1080 resolution image as the input source, and supports users to select the image file to be processed through the file upload interface. The computing device directly uses the depth estimation model to perform depth estimation. During the processing, the system supports a mask function, allowing users to specify specific areas for focus processing or exclusion. After completing the depth estimation, the computing device 102 will generate a corresponding depth map, and the output file will automatically add a prefix for easy identification and management, while maintaining the 1920*1080 resolution specification of the original image. This workflow significantly improves the efficiency of the traditional depth estimation process through automated processing, and can provide a higher quality depth data foundation for subsequent three-dimensional modeling and virtual scene synthesis.

[0069] The following combination Figure 12 A schematic diagram illustrates an interface for generating and exporting models according to some embodiments of the present disclosure. In example 1200, box 1202 is the panel for generating a 3D model, and box 1204 is the panel for exporting a model. These two panels primarily contain multiple functional modules, including file operations, depth map settings, mesh view control, and export parameter configuration. Computing device 102 first provides file management options, supporting expanded browsing of images, videos, and directories. The user can select the primary input source using the "Open File" button. In the depth map settings area, the user can specify the target object and adjust depth parameters. The mesh view section contains mesh properties and modifier application options, allowing the user to make final adjustments to the 3D mesh. The operation presets area provides a variety of processing mode options, including path mode, copy functionality, and batch processing. Computing device 102 also provides granular export scope control, supporting export of selected objects, visible objects, or active collections. In the object type filter, the user can select to export various 3D elements, including empty objects, cameras, lights, skeletons, meshes, and more, and supports preserving custom properties. At the bottom of the interface, there are two action buttons, "Export" and "Cancel," for confirming or aborting the export process. The entire panel adopts a layered design, arranging complex export parameters by logical functional groupings. This not only ensures precise control of export settings for professional users, but also reduces operational complexity through a clear interface layout. This functional module is particularly suitable for workflows that require exporting 3D scenes to a predetermined format for cross-platform collaboration, providing a simple and convenient interactive operation scenario for the production of 3D backgrounds.

[0070] Figure 13 A schematic block diagram of an apparatus for generating a three-dimensional scene according to some embodiments of the present disclosure is illustrated. Figure 13 The device 1300 shown can be Figure 1 is implemented in the computing device 102. Figure 13 As shown, the device 1300 includes a target image determination module 1302, configured to determine a target image that can be used to generate a three-dimensional scene; a depth map determination module 1304, configured to determine a depth map corresponding to the target image based on the target image; a three-dimensional mesh model determination module 1306, configured to determine a three-dimensional mesh model corresponding to the target image based on the depth map and the target image; and a three-dimensional scene determination module 1308, configured to determine a three-dimensional scene corresponding to the target image based on the target image and the three-dimensional mesh model.

[0071] In some embodiments, the target image determination module 1302 includes: a prompt word receiving module, configured to receive prompt words for describing the target image; and a first image generation module, configured to generate the target image by inputting the prompt words into a first machine learning model.

[0072] In some embodiments, the apparatus 1300 further includes a resolution processing module configured to update the target image by performing image super-resolution processing on the target image.

[0073] In some embodiments, the depth map determination module 1304 includes: a depth information extraction module configured to extract depth information from the target image; and a first depth map module configured to determine a depth map for the target image based on the depth information.

[0074] In some embodiments, the depth map determination module 1304 includes: a second depth map module configured to obtain a depth map by inputting the target image into a second machine learning model.

[0075] In some embodiments, the 3D scene determination module 1308 includes: a first scene determination module configured to generate a 3D scene corresponding to the target image by combining the target image with a 3D mesh model.

[0076] In some embodiments, the three-dimensional scene determination module 1308 includes: a video determination module, configured to generate a video corresponding to the target image based on the target image; and a second scene determination module, configured to determine the three-dimensional scene corresponding to the target image based on the video and the three-dimensional mesh model.

[0077] In some embodiments, the video determination module includes: a first video determination module configured to generate a video corresponding to the target image by applying the target image to a third machine learning model.

[0078] In some embodiments, the apparatus 1300 further includes: an updating module configured to update the video by performing image super-resolution processing on video frames in the video.

[0079] In some embodiments, the second scene determination module includes: a second combining module configured to generate a three-dimensional scene by combining each video frame in the video into a three-dimensional mesh model.

[0080] Figure 14 A schematic block diagram of an example device 1400 is shown, which may be used to implement embodiments of the present disclosure. Figure 1 The computing device 102 in the embodiment of the present invention can be implemented using device 1400. As shown in the figure, device 1400 includes a central processing unit (CPU) 1401, which can perform various appropriate actions and processes according to computer program instructions stored in a read-only memory (ROM) 1402 or computer program instructions loaded from a storage unit 1408 into a random access memory (RAM) 1403. Various programs and data required for the operation of device 1400 can also be stored in RAM 1403. CPU 1401, ROM 1402, and RAM 1403 are connected to each other via a bus 1404. An input / output (I / O) interface 1405 is also connected to bus 1404.

[0081] Various components in device 1400 are connected to I / O interface 1405, including: an input unit 1406, such as a keyboard, mouse, etc.; an output unit 1407, such as various types of displays, speakers, etc.; a storage unit 1408, such as a magnetic disk, optical disk, etc.; and a communication unit 1409, such as a network card, modem, wireless communication transceiver, etc. Communication unit 1409 allows device 1400 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks.

[0082] The various processes and processing described above, such as methods 200 and 300, may be performed by the processing unit 1401. For example, in some embodiments, the methods 200 and 300 may be implemented as a computer software program that is tangibly contained in a machine-readable medium, such as the storage unit 1408. In some embodiments, part or all of the computer program may be loaded and / or installed on the device 1400 via the ROM 1402 and / or the communication unit 1409. When the computer program is loaded into the RAM 1403 and executed by the CPU 1401, one or more actions of the example methods 200 and 300 described above may be performed.

[0083] The present disclosure may be a method, an apparatus, a system and / or a computer program product. The computer program product may include a computer-readable storage medium carrying computer-readable program instructions for executing various aspects of the present disclosure.

[0084] Computer-readable storage media can be a tangible device that can hold and store instructions used by an instruction execution device. Computer-readable storage media can be, for example, but not limited to, an electrical storage device, a magnetic storage device, an optical storage device, an electromagnetic storage device, a semiconductor storage device, or any suitable combination thereof. More specific examples (non-exhaustive list) of computer-readable storage media include: a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), a static random access memory (SRAM), a portable compact disc read-only memory (CD-ROM), a digital versatile disk (DVD), a memory stick, a floppy disk, a mechanical encoding device, a punch card or a raised structure in a groove on which instructions are stored, and any suitable combination thereof. The computer-readable storage media used herein is not to be interpreted as a transient signal itself, such as a radio wave or other freely propagating electromagnetic wave, an electromagnetic wave propagated by a waveguide or other transmission medium (for example, a light pulse by a fiber optic cable), or an electrical signal transmitted by a wire.

[0085] The computer-readable program instructions described herein can be downloaded from a computer-readable storage medium to each computing / processing device, or downloaded to an external computer or external storage device via a network, such as the Internet, a local area network, a wide area network, and / or a wireless network. The network can include copper transmission cables, fiber optic transmission, wireless transmission, routers, firewalls, switches, gateway computers, and / or edge servers. The network adapter card or network interface in each computing / processing device receives the computer-readable program instructions from the network and forwards the computer-readable program instructions to be stored in the computer-readable storage medium in each computing / processing device.

[0086] The computer program instructions for performing the operations of the present disclosure may be assembly instructions, instruction set architecture (ISA) instructions, machine instructions, machine-dependent instructions, microcode, firmware instructions, state setting data, or source code or object code written in any combination of one or more programming languages, including object-oriented programming languages such as Smalltalk, C++, and conventional procedural programming languages such as "C" language or similar programming languages. Computer-readable program instructions may be executed entirely on a user's computer, partially on a user's computer, as an independent software package, partially on a user's computer, partially on a remote computer, or entirely on a remote computer or server. In the case of a remote computer, the remote computer may be connected to the user's computer via any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., utilizing an Internet service provider to connect via the Internet). In some embodiments, an electronic circuit, such as a programmable logic circuit, a field programmable gate array (FPGA), or a programmable logic array (PLA), may be personalized by utilizing the state information of the computer-readable program instructions. The electronic circuit may execute the computer-readable program instructions, thereby realizing various aspects of the present disclosure.

[0087] Various aspects of the present disclosure are described herein with reference to flowcharts and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the present disclosure. It should be understood that each block of the flowcharts and / or block diagrams, and combinations of blocks in the flowcharts and / or block diagrams, can be implemented by computer-readable program instructions.

[0088] These computer-readable program instructions can be provided to a processing unit of a general-purpose computer, a special-purpose computer, or other programmable data processing device, thereby producing a machine such that when these instructions are executed by the processing unit of the computer or other programmable data processing device, a device is generated that implements the functions / actions specified in one or more blocks in the flowchart and / or block diagram. These computer-readable program instructions can also be stored in a computer-readable storage medium, where these instructions cause the computer, programmable data processing device, and / or other device to operate in a specific manner. Thus, the computer-readable medium storing the instructions comprises an article of manufacture that includes instructions for implementing various aspects of the functions / actions specified in one or more blocks in the flowchart and / or block diagram.

[0089] Computer-readable program instructions may also be loaded onto a computer, other programmable data processing apparatus, or other device so that a series of operational steps are performed on the computer, other programmable data processing apparatus, or other device to produce a computer-implemented process, thereby causing the instructions executed on the computer, other programmable data processing apparatus, or other device to implement the functions / actions specified in one or more blocks in the flowchart and / or block diagram.

[0090] The flow charts and block diagrams in the accompanying drawings show the possible architecture, functions and operations of the systems, methods and computer program products according to multiple embodiments of the present disclosure. In this regard, each box in the flow chart or block diagram can represent a part of a module, program segment or instruction, and a part of a module, program segment or instruction includes one or more executable instructions for realizing the prescribed logical function. In some alternative implementations, the functions marked in the box can also occur in a sequence different from that marked in the accompanying drawings. For example, two consecutive boxes can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram and / or flow chart, and the combination of the boxes in the block diagram and / or flow chart, can be implemented by a dedicated hardware-based system that performs the prescribed function or action, or can be implemented by a combination of dedicated hardware and computer instructions.

[0091] While various embodiments of the present disclosure have been described above, the foregoing description is intended to be illustrative, non-exhaustive, and not limited to the disclosed embodiments. Many modifications and variations will be apparent to those skilled in the art without departing from the scope and spirit of the described embodiments. The terminology used herein is selected to best explain the principles of the embodiments, their practical applications, or technical improvements to existing technologies, or to enable others skilled in the art to understand the embodiments disclosed herein.

Claims

1. A method for generating a three-dimensional scene, comprising: determining a target image that can be used to generate a three-dimensional scene; Based on the target image, determining a depth map corresponding to the target image; Determining a three-dimensional mesh model corresponding to the target image based on the depth map and the target image; and Based on the target image and the three-dimensional mesh model, a three-dimensional scene corresponding to the target image is determined.

2. The method according to claim 1, wherein determining a target image that can be used to generate a three-dimensional scene comprises: receiving a prompt word for describing the target image; The target image is generated by inputting the prompt word into a first machine learning model.

3. The method according to claim 2, further comprising The target image is updated by performing image super-resolution processing on the target image.

4. The method of claim 1 , wherein determining a depth map corresponding to the target image comprises: extracting depth information from the target image; as well as Based on the depth information, a depth map for the target image is determined.

5. The method of claim 1 , wherein determining a depth map corresponding to the target image comprises: The depth map is obtained by inputting the target image into a second machine learning model.

6. The method according to claim 1 , wherein determining a three-dimensional scene corresponding to the target image based on the target image and the three-dimensional mesh model comprises: A three-dimensional scene corresponding to the target image is generated by combining the target image with the three-dimensional mesh model.

7. The method according to claim 1 , wherein determining a three-dimensional scene corresponding to the target image based on the target image and the three-dimensional mesh model comprises: Based on the target image, generating a video corresponding to the target image; as well as The three-dimensional scene corresponding to the target image is determined based on the video and the three-dimensional mesh model.

8. The method of claim 7, wherein generating a video corresponding to the target image comprises: A video corresponding to the target image is generated by applying the target image to a third machine learning model.

9. The method according to claim 8, further comprising: The video is updated by performing image super-resolution processing on video frames in the video.

10. The method according to claim 7, wherein determining a three-dimensional scene corresponding to the target image based on the video and the three-dimensional mesh model comprises: The three-dimensional scene is generated by combining each video frame in the video into the three-dimensional mesh model.

11. A device for generating a three-dimensional scene, comprising: a target image determination module, configured to determine a target image that can be used to generate a three-dimensional scene; a depth map determining module, configured to determine a depth map corresponding to the target image based on the target image; a three-dimensional mesh model determination module configured to determine a three-dimensional mesh model corresponding to the target image based on the depth map and the target image; as well as The three-dimensional scene determination module is configured to determine a three-dimensional scene corresponding to the target image based on the target image and the three-dimensional mesh model.

12. An electronic device comprising: at least one processor; as well as A storage device for storing at least one program, wherein when the at least one program is executed by the at least one processor, the at least one processor implements the method according to any one of claims 1 to 10.

13. A computer-readable storage medium having a computer program stored thereon, wherein the computer program implements the method according to any one of claims 1 to 10 when executed by a processor.

14. A computer program product comprising a computer program, which, when executed by a processor, implements the method according to any one of claims 1 to 10.