Three-dimensional scene generation method and device, electronic equipment and storage medium

By generating and correcting the physical relationships of target objects in the 3D scene generation model, the problem of poor 3D scene effects in existing technologies is solved, and high-quality 3D asset construction and correct object placement are achieved.

CN120747362APending Publication Date: 2025-10-03SHADOW EYE TECH SHANGHAI CO LTD
View PDF 7 Cites 0 Cited by

Patent Information

Application Number
CN202510857795.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-25
Publication Date
2025-10-03

AI Technical Summary

Technical Problem

When using scene images to construct three-dimensional scenes in existing technologies, the scene effects are poor, the quality of individual three-dimensional assets is limited, and objects are prone to incorrect placement such as interpenetration and floating.

Method used

By determining the target objects and target point cloud data in the scene image, the 3D scene generation model is used to generate the 3D assets of each target object, and the physical relationship of the target object assets in the initial 3D scene is corrected based on the target physical relationship to prevent incorrect placement.

Benefits of technology

Improved the 3D asset quality of individual target objects, ensuring correct placement of objects and improving the overall quality of the generated 3D scene.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120747362A_ABST
    Figure CN120747362A_ABST
Patent Text Reader

Abstract

The invention provides a three-dimensional scene generation method and device, electronic equipment and a storage medium, and relates to the technical field of computers. The three-dimensional scene generation method comprises the following steps: determining target objects in a scene picture and target point cloud data corresponding to each target object; inputting the scene picture and the target point cloud data into a three-dimensional scene generation model to obtain an initial three-dimensional scene corresponding to the scene picture output by the three-dimensional scene generation model; and obtaining a target physical relationship among the target objects in the scene picture, and performing physical relationship correction on the three-dimensional assets corresponding to the target objects in the initial three-dimensional scene based on the target physical relationship to obtain a target three-dimensional scene corresponding to the scene picture. According to the technical scheme, multiple high-quality three-dimensional assets can be generated according to the input scene picture, the three-dimensional scene can be formed according to the correct physical relationship, the quality of the single three-dimensional asset and the effect of the three-dimensional scene are improved, and the accuracy of the positions of the three-dimensional assets in the three-dimensional scene is ensured.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of computer technology, and in particular to a three-dimensional scene generation method, device, electronic device and storage medium. Background Art

[0002] Three-dimensional scene construction refers to a virtual environment or collection of objects with a three-dimensional sense of space built through computer technology. It can simulate the spatial relationships, light and shadow effects, and object interactions in the real world, and has been widely used in film and television production, animation production, games, virtual reality and other fields.

[0003] With the development of computer technology, constructing 3D scenes using scene images has become a trend. However, in related technologies, constructing 3D scenes using scene images mostly uses point clouds of the entire scene, resulting in poor scene effects. Summary of the Invention

[0004] The present application provides a three-dimensional scene generation method, device, electronic device and storage medium to solve the problem of poor scene effect in the prior art of constructing three-dimensional scenes using scene images.

[0005] In a first aspect, the present application provides a three-dimensional scene generation method, comprising:

[0006] Determine target objects in the scene image and target point cloud data corresponding to each target object;

[0007] Inputting the scene image and the target point cloud data into a 3D scene generation model to obtain an initial 3D scene corresponding to the scene image output by the 3D scene generation model; the 3D scene generation model is used to generate each target object in the scene image into a 3D asset based on the target point cloud data corresponding to each target object, and construct the 3D assets into the initial 3D scene;

[0008] A target physical relationship between each target object in the scene image is obtained, and based on the target physical relationship, a physical relationship correction is performed on the three-dimensional assets corresponding to each target object in the initial three-dimensional scene to obtain a target three-dimensional scene corresponding to the scene image.

[0009] In one embodiment, inputting the scene image and the target point cloud data into a three-dimensional scene generation model to obtain an initial three-dimensional scene corresponding to the scene image output by the three-dimensional scene generation model includes:

[0010] Inputting the scene image into a feature extraction module of the 3D scene generation model, and obtaining a 3D reconstruction latent vector corresponding to each target object output by the feature extraction module; wherein the feature extraction module is used to extract information describing the 3D asset of each target object in the scene image; and the 3D reconstruction latent vector is used to represent the 3D asset information of the target object;

[0011] For each target object, the three-dimensional reconstructed latent vector and the corresponding target point cloud data corresponding to each target object are input into the three-dimensional asset generation module of the three-dimensional scene generation model to obtain the initial three-dimensional scene output by the three-dimensional asset generation module; wherein the three-dimensional asset generation module is used to generate a three-dimensional asset corresponding to each target object, and generate each three-dimensional asset into a three-dimensional scene.

[0012] In one embodiment, the 3D asset generation module of the 3D scene generation model includes a point cloud image generation layer, a point cloud alignment layer, and a scene generation layer; inputting the 3D reconstructed latent vector corresponding to each target object and the corresponding target point cloud data into the 3D asset generation module of the 3D scene generation model to obtain the initial 3D scene output by the 3D asset generation module includes:

[0013] For each target object, the 3D reconstruction latent vector and the target point cloud data corresponding to each target object are input into the point cloud image generation layer, and a 3D asset latent vector corresponding to each target object is output by the point cloud image generation layer; wherein the 3D asset latent vector is used to represent the 3D asset of the target object; and the point cloud image generation layer is used to construct the 3D asset based on the 3D reconstruction latent vector and the target point cloud data.

[0014] Inputting the three-dimensional asset latent vector and the target point cloud data corresponding to each target object into the point cloud alignment layer, obtaining an aligned point cloud corresponding to each target object output by the point cloud alignment layer; the point cloud alignment layer is used to align the target point cloud data with the complete point cloud of the target object in the real scene corresponding to the scene image based on the three-dimensional asset latent vector;

[0015] The initial three-dimensional scene is generated based on the scene generation layer and the aligned point cloud corresponding to each target object; the scene generation layer is used to generate the input point cloud into a three-dimensional scene.

[0016] In one embodiment, generating the initial three-dimensional scene based on the scene generation layer and the aligned point cloud corresponding to each target object includes:

[0017] The aligned point cloud corresponding to each target object is used as the target point cloud data and re-inputted into the point cloud image generation layer together with the 3D reconstruction latent vector. The 3D asset construction based on the point cloud image generation layer and the alignment operation based on the point cloud alignment layer are repeated until the similarity between the aligned point cloud output by the point cloud alignment layer and the complete point cloud is greater than or equal to a preset similarity threshold.

[0018] The aligned point cloud finally output by the point cloud alignment layer is input into the scene generation layer to obtain the initial three-dimensional scene output by the scene generation layer.

[0019] In one embodiment, the physical relationship correction of the three-dimensional assets corresponding to the target objects in the initial three-dimensional scene based on the target physical relationship to obtain the target three-dimensional scene corresponding to the scene image includes:

[0020] Determining an object to be corrected among all the target objects in the initial three-dimensional scene according to the target physical relationship;

[0021] For two target objects to be corrected among the objects to be corrected, determining a target moving object and a corresponding target correction loss function among the two target objects to be corrected based on the target physical relationship between the two target objects to be corrected;

[0022] By moving the three-dimensional asset corresponding to the target moving object, iteratively optimizing the loss value of the target correction loss function until the target correction loss function converges, thereby obtaining the initial three-dimensional scene after physical relationship correction;

[0023] The initial three-dimensional scene after the physical relationship correction is determined as the target three-dimensional scene.

[0024] In one embodiment, determining the target moving object and the corresponding target correction loss function of the two target objects to be corrected based on the target physical relationship between the two target objects to be corrected includes:

[0025] In a case where the target physical relationship between the two target objects to be corrected is a placement relationship, determining the placed object among the two target objects to be corrected as the target moving object, and determining the correction loss function corresponding to the placed object as the target correction loss function;

[0026] When the target physical relationship between the two target objects to be corrected is a non-placement contact relationship, both the target objects to be corrected are determined as the target moving objects, and the sum of the correction loss functions corresponding to each target object to be corrected is determined as the target correction loss function;

[0027] The target correction loss function is used to characterize the distances of all three-dimensional points of one of the two target moving objects to be corrected relative to the other target moving object to be corrected.

[0028] In one embodiment, determining the target objects in the scene image and the target point cloud data corresponding to each target object includes:

[0029] Identify target objects in the scene image and mark the edges of each target object;

[0030] Extracting scene point cloud data of the scene image, and determining point cloud data in the scene point cloud data that is aligned with pixels of each of the target objects;

[0031] The point cloud data of each target object is matched with the edge to obtain the target point cloud data corresponding to each target object.

[0032] In a second aspect, the present application provides a three-dimensional scene generation device, comprising:

[0033] a determination unit, configured to determine target objects in a scene image and target point cloud data corresponding to each target object;

[0034] a 3D scene generation unit, configured to input the scene image and the target point cloud data into a 3D scene generation model to obtain an initial 3D scene corresponding to the scene image output by the 3D scene generation model; the 3D scene generation model is configured to generate each target object in the scene image into a 3D asset based on the target point cloud data corresponding to each target object, and construct the 3D assets into the initial 3D scene;

[0035] A correction unit is used to obtain the target physical relationship between each target object in the scene image, and based on the target physical relationship, perform physical relationship correction on the three-dimensional assets corresponding to each target object in the initial three-dimensional scene to obtain the target three-dimensional scene corresponding to the scene image.

[0036] In a third aspect, the present application provides an electronic device comprising a memory and a processor, wherein the memory stores a computer program that can be run on the processor, and when the processor executes the program, the steps of the three-dimensional scene generation method as described in any one of the first aspects above are implemented.

[0037] In a third aspect, the present application provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps of the three-dimensional scene generation method as described in any one of the first aspects above.

[0038] The three-dimensional scene generation method, device, electronic device and storage medium provided in the present application determine the target objects in the scene image and the target point cloud data corresponding to each target object, input the scene image and the target point cloud data into the three-dimensional scene generation model, and obtain the initial three-dimensional scene corresponding to the scene image output by the three-dimensional scene generation model, wherein the three-dimensional scene generation model is used to generate each target object in the scene image into a three-dimensional asset based on the target point cloud data corresponding to each target object, and construct the three-dimensional assets into an initial three-dimensional scene. In this way, the three-dimensional assets of each target object can be generated for the target point cloud data of each target object in the scene image, and then the initial three-dimensional scene corresponding to the scene image can be constructed, thereby improving the quality of the three-dimensional assets of a single target object in the scene; further, by obtaining the target physical relationship between each target object in the scene image, and correcting the physical relationship of the three-dimensional assets corresponding to each target object in the initial three-dimensional scene based on the target physical relationship, the target three-dimensional scene corresponding to the scene image is obtained, so that each target object in the target three-dimensional scene can be placed in the correct physical relationship, further improving the quality of the generated target three-dimensional scene. BRIEF DESCRIPTION OF THE DRAWINGS

[0039] Figure 1 A schematic diagram of a flow chart of a three-dimensional scene generation method provided in an embodiment of the present application;

[0040] Figure 2 This is a schematic diagram of the structure and working principle of the three-dimensional scene generation model provided in an embodiment of the present application;

[0041] Figure 3 The second schematic diagram of the structure and working principle of the three-dimensional scene generation model provided in the embodiment of the present application;

[0042] Figure 4 A schematic diagram of the structure of a three-dimensional scene generation device provided in an embodiment of the present application. DETAILED DESCRIPTION

[0043] In this application, "at least one" means one or more, and "plurality" means two or more. "And / or" describes the association relationship of associated objects, indicating that three relationships may exist. For example, A and / or B can mean: A exists alone, A and B exist at the same time, and B exists alone. Among them, A and B can be singular or plural. The character " / " generally indicates that the previous and next associated objects are in an "or" relationship. "At least one of the following items" or similar expressions refers to any combination of these items, including any combination of single items or plural items. For example, at least one of a alone, b alone, or c alone can mean: a alone, b alone, c alone, a and b in combination, a and c in combination, b and c in combination, or a, b, and c in combination. Among them, a, b, and c can be single or multiple. In addition, the terms "first" and "second" are used for descriptive purposes only and are not to be understood as indicating or implying relative importance.

[0044] The directions or positional relationships indicated by terms such as "center", "longitudinal", "lateral", "up", "down", "left", "right", "front", and "back" are based on the directions or positional relationships shown in the accompanying drawings and are only for the convenience of describing the present application and simplifying the description. They do not indicate or imply that the device or element referred to must have a specific direction, be constructed and operated in a specific direction. Therefore, they should not be understood as limiting the present application.

[0045] The terms "connected" and "connect" should be interpreted broadly. For example, "connected" or "connected" in a circuit structure can refer not only to a physical connection, but also to an electrical connection or a signal connection. For example, it can be a direct connection, i.e., a physical connection, or an indirect connection through at least one intermediate component, as long as the circuit is interconnected. It can also refer to internal connectivity between two components. Signal connection can refer not only to signal connection through circuits but also to signal connection through media, such as radio waves. Those skilled in the art will understand the specific meanings of the above terms in this application on a case-by-case basis.

[0046] The terms involved in the embodiments of this application are explained as follows:

[0047] (1) Three-dimensional assets, also known as three-dimensional digital assets, refer to virtual objects or resources created in three-dimensional space through digital technology (such as three-dimensional modeling), including elements such as models, textures and animations.

[0048] (2) 3D scene construction refers to a virtual environment or collection of objects with a three-dimensional sense of space, constructed through computer technology, which can simulate the spatial relationships, light and shadow effects, and object interactions in the real world. It has been widely used in film and television production, animation production, games, virtual reality and other fields.

[0049] (3) Hidden vector is a vector representation method that maps high-dimensional sparse features to low-dimensional dense space through mathematical modeling.

[0050] With the rapid development of computer technology, especially emerging technologies such as the metaverse, virtual reality, and augmented reality, the demand for high-quality three-dimensional assets has exploded and has become the core infrastructure of the digital content industry. Driven by this technology and the demand for the development of three-dimensional assets, the use of scene images to construct three-dimensional scenes has become a trend. However, in related technologies, the use of scene images to construct three-dimensional scenes mostly uses point clouds of the entire scene, which has limited aesthetics and accuracy, the quality of individual three-dimensional assets, and poor scene effects. Moreover, constructing the entire scene in the form of a point cloud is prone to incorrect placement of objects, such as interpenetration and floating, further affecting the effect of the three-dimensional scene.

[0051] Based on this, an embodiment of the present application provides a three-dimensional scene generation method. First, the target objects and the target point cloud data corresponding to each target object are determined from the scene image. Then, the scene image and the target point cloud data corresponding to each target object are input into a three-dimensional scene generation model. The three-dimensional assets of each target object are generated using the three-dimensional scene generation model, thereby forming an initial three-dimensional scene corresponding to the scene image. In this way, the three-dimensional assets of each target object can be generated based on the target point cloud data of each target object in the scene image, thereby constructing the initial three-dimensional scene corresponding to the scene image, thereby improving the quality of the three-dimensional assets of the individual target objects in the scene. After obtaining the initial three-dimensional scene, the physical relationship of the three-dimensional assets corresponding to each target object in the initial three-dimensional scene can be corrected based on the target physical relationship between the target objects in the scene image, thereby preventing incorrect placement positions such as interpenetration and floating between the target objects, thereby further improving the quality of the generated three-dimensional scene. In this way, multiple high-quality three-dimensional assets can be generated using the scene image, and these three-dimensional assets are correctly combined into the scene in the original scene image. Among them, the three-dimensional scene generation model can be a neural network model obtained by training the initial three-dimensional scene generation model of the neural network architecture using sample scene images.

[0052] The 3D scene generation method provided in the embodiments of the present application can be applied to an electronic device or to a 3D scene generation device provided in the electronic device. The 3D scene generation device can be implemented by software, hardware, or a combination of both. The electronic device can include at least one of a server, a mobile phone, a computer, a tablet computer, and a wearable device. The server can include a standalone server, a virtual server, a cluster server, and the like.

[0053] The following combination Figures 1 to 3 , taking the execution subject as an electronic device as an example, the three-dimensional scene generation method provided in the embodiment of the present application is described in detail.

[0054] Figure 1 The flow chart of the three-dimensional scene generation method provided by the embodiment of the present application is shown. Figure 1 As shown, the three-dimensional scene generation method may include the following steps 110 to 130.

[0055] Step 110: Determine the target objects in the scene image and the target point cloud data corresponding to each target object.

[0056] The electronic device can obtain a scene image of the scene to be generated into a 3D scene, use an object recognition algorithm to identify objects in the scene image, and then determine the objects for which the 3D asset needs to be generated from the identified objects to obtain target objects. For each identified target object, the electronic device can extract point cloud data for each target object in the scene image to obtain target point cloud data corresponding to each target object.

[0057] In one embodiment, step 110 of determining the target objects in the scene image and the target point cloud data corresponding to each target object can be implemented through the following steps 111 to 113.

[0058] Step 111: Determine the target objects in the scene image and mark the edge of each target object.

[0059] For example, an electronic device can use the Florence-2 visual model to identify objects in a scene image and obtain recognition results, which may include text descriptions and bounding boxes for each identified object. The electronic device can then use the Large Language Model (LLM) to label meaningful components of the identified objects, i.e., objects that require 3D asset construction, and identify these labeled objects as target objects in the scene image.

[0060] Among them, the visual model Florence-2 is a model architecture that uses sequence-to-sequence methods, using text prompts as task instructions and generating desired results in text form. An example of a large language model is the multimodal language model GPT-4v with visual capabilities. Text descriptions can include object categories and attribute information. For example, if the object category is "table," the corresponding attribute information may include shape, size, etc. A bounding box is a rectangular box that can be considered the target detection box for the object. Meaningful component objects are objects that require three-dimensional asset construction. For example, in an office scene, important objects such as tables, printers, cabinets, trash cans, and computers in an office scene image are meaningful component objects. However, small objects such as floor patterns and wall patterns that appear in the scene image can be considered meaningless and will not be constructed into three-dimensional assets.

[0061] For example, the electronic device can call the open-source model Grounded SAM-v2 to mark the edges of each target object. Grounded SAM-v2 is an open-source model based on deep learning that can implement functions such as target localization, segmentation, and tracking.

[0062] Step 112: extracting scene point cloud data of the scene image, and determining point cloud data in the scene point cloud data that is aligned with the pixels of each target object.

[0063] For example, the electronic device can use a neural network model that recovers the three-dimensional geometric structure from a monocular image to extract the scene point cloud data of the scene picture, and obtain the global camera parameters corresponding to the scene picture, and then align the scene point cloud data with the pixels of each target object in the scene picture according to the global camera parameters to obtain point cloud data aligned with the pixels of each target object. For example, the electronic device can use a multi-object generative (MoGe) model to extract the scene point cloud data of the scene picture and obtain the global camera parameters corresponding to the scene picture. The MoGe model includes a computer vision model (Vision Transformer, ViT) encoder based on the Transformer architecture and a convolutional decoder, which can directly predict the affine invariant point map and the mask that excludes undefined geometric areas, so as to further derive the camera displacement, camera focal length and depth map.

[0064] In this way, through pixel alignment, the point cloud data in the scene point cloud data can be matched with each target object to obtain the point cloud data belonging to each target object.

[0065] Step 113: Match the point cloud data of each target object with the edge to obtain target point cloud data corresponding to each target object.

[0066] After the electronic device determines the point cloud data corresponding to each target object, it can match the point cloud data of each target object with the edge of the target object. In this way, the point cloud data within the edge of each target object is the target point cloud data of the target object.

[0067] Step 120: Input the scene image and the target point cloud data into the 3D scene generation model to obtain an initial 3D scene corresponding to the scene image output by the 3D scene generation model.

[0068] The 3D scene generation model can be a neural network model obtained by training an initial 3D scene generation model of a neural network architecture using sample scene images. The 3D scene generation model is used to generate each target object in the scene image into a 3D asset based on the target point cloud data corresponding to each target object, and construct the 3D assets into an initial 3D scene.

[0069] For example, a sample scene image can be used as input, and the sample 3D scene corresponding to the sample scene image can be used as a label. The initial 3D scene generation model can be trained based on the model loss function until the model loss function converges, thereby obtaining a trained 3D scene generation model. The point cloud of the sample 3D model of the sample object in the sample 3D scene is the complete point cloud of the sample object in the real scene corresponding to the sample scene image.

[0070] In one embodiment, Figure 2 One of the schematic diagrams showing the structure and working principle of the three-dimensional scene generation model provided in the embodiment of the present application is shown. Figure 2 As shown, the 3D scene generation model may include a feature extraction module 20 and a 3D asset generation module 30. The feature extraction module 20 is used to extract information describing the 3D assets of each target object in the scene image; the 3D asset generation module 30 is used to generate a 3D asset corresponding to each target object and generate each 3D asset into a 3D scene.

[0071] Specifically, feature extraction module 20 is a neural network model capable of extracting information from images. It is capable of extracting 3D asset information from all target objects in a scene image, obtaining latent vectors for 3D reconstruction of the target objects. This module is capable of extracting 3D asset information not only from unoccluded objects but also from occluded objects. Feature extraction module 20 can be obtained by self-supervised training of an initial feature extraction model based on sample scene images. The initial feature extraction model is a neural network model with visual feature extraction capabilities. 3D asset information is information that can describe 3D assets, and this information can represent the corresponding target objects.

[0072] For example, the initial feature extraction model can be a computer vision self-supervised model DINOv2, which can learn visual features from a large number of images without using any labels or annotations. When the initial feature extraction model is self-supervised trained, the sample scene image can be provided as input to the DINOv2 model, and the DINOv2 model generates a digital vector, usually called an embedding or visual feature, which contains a deep understanding of the input sample scene image. Random occlusions can be added to the sample scene image. In the process of training the initial feature extraction model, the initial feature extraction model is required to reconstruct the input sample scene image to optimize the network parameters in the model until the training is completed, and the feature extraction module 20 can be obtained. In this way, for target objects that are partially occluded in the scene image, the feature extraction module 20 can also extract information for reconstructing three-dimensional assets.

[0073] The 3D asset generation module 30 can be a neural network model capable of generating a 3D scene corresponding to an image based on the 3D reconstructed latent vectors and corresponding point cloud data of objects in the image. The 3D asset generation module 30 can use the 3D sample reconstructed latent vectors and sample point cloud data corresponding to the sample scene image as model inputs, and the sample 3D scene corresponding to the sample scene image as labels. It can train an initial 3D asset generation model based on a training loss function to obtain a trained initial 3D asset generation model, i.e., the 3D asset generation module 30. The point cloud of the sample object in the sample 3D scene is the complete point cloud of the sample object in the real scene corresponding to the sample scene image. For example, if the sample scene image includes a table, the point cloud of the table in the sample 3D scene used as a label during training is the complete point cloud of the entire table in the real scene.

[0074] based on Figure 2 The three-dimensional scene generation model shown in step 120 inputs the scene image and the target point cloud data into the three-dimensional scene generation model to obtain the initial three-dimensional scene corresponding to the scene image output by the three-dimensional scene generation model, which can be achieved through the following steps 121 to 122.

[0075] Step 121: Input the scene image into the feature extraction module 20 of the 3D scene generation model to obtain the 3D reconstruction latent vector corresponding to each target object output by the feature extraction module 20.

[0076] Among them, the 3D reconstruction latent vector is used to represent the 3D asset information of the target object.

[0077] Step 122: For each target object, the corresponding 3D reconstructed latent vector and the corresponding target point cloud data of each target object are input into the 3D asset generation module 30 of the 3D scene generation model to obtain the initial 3D scene output by the 3D asset generation module 30.

[0078] In one embodiment, Figure 3 The second schematic diagram shows the structure and working principle of the three-dimensional scene generation model provided in the embodiment of the present application, referring to Figure 3 As shown, the 3D asset generation module 30 of the 3D scene generation model may include a point cloud image generation layer 31, a point cloud alignment layer 32, and a scene generation layer 33. The point cloud image generation layer 31 is used to construct a 3D asset based on the 3D reconstruction latent vectors and the target point cloud data; the point cloud alignment layer 32 is used to align the target point cloud data with the complete point cloud of the target object in the real scene corresponding to the scene image based on the 3D asset latent vectors; and the scene generation layer 33 is used to generate a 3D scene from the input point cloud.

[0079] Specifically, the point cloud image generation layer 31 can normalize the target point cloud data of the target object to the canonical space of [-1, 1] and generate a 3D asset latent vector for the target object based on a 3D asset generation algorithm. For example, the point cloud image generation layer 31 can be a trained 3D model generation algorithm model, which can be trained using the initial point cloud image generation layer based on the sample 3D reconstruction latent vector corresponding to the sample object image, the sample object point cloud corresponding to the sample object image, and the sample 3D asset latent vector as a label. Specifically, the sample three-dimensional reconstruction latent vector corresponding to the sample object image and the sample object point cloud corresponding to the sample object image can be used as input, and the sample three-dimensional asset latent vector of the sample object in the sample object image can be used as a label to train the initial point cloud image generation layer. During the training process, the sample object point cloud and the sample three-dimensional reconstruction latent vector are used as guidance, and the initial point cloud image generation layer is used to deduce the sample three-dimensional reconstruction latent vector of the sample object in the sample object image as the goal. The network parameters of the initial point cloud image generation layer are optimized until the initial point cloud image generation layer converges to obtain a trained initial point cloud image generation layer, which can be used as the point cloud image generation layer 31.

[0080] The sample object image can be an image of a single sample object in a sample scene image, and the sample object point cloud can be a partially occluded and noisy point cloud. The initial point cloud image generation layer can be a neural network model based on a 3D model generation algorithm. For example, it can be a 3D generative model CLAY, which includes a multi-resolution variational autoencoder (VAE) and a diffusion transformer (DiT). The VAE is responsible for encoding 3D geometric shapes at different levels of detail into the latent space, and the DiT is responsible for generating these geometric shapes. At the same time, precise control of the generated results can be achieved through the specified point cloud or bounding box.

[0081] The point cloud alignment layer 32 can generate point cloud data that is more aligned with the complete point cloud of the target object in the real scene, based on the input 3D asset latent vector of a single target object and its corresponding target point cloud data. This point cloud alignment layer 32 can be trained by the initial point cloud alignment layer based on the sample point cloud data of each sample object in the sample scene image, the sample 3D asset latent vector of each sample object, and label data. The label data is the complete point cloud of the sample object in the real scene. Specifically, during the training process, the sample point cloud data and the sample three-dimensional asset latent vector of each sample object in the sample scene image are input into the initial point cloud alignment layer to obtain the result point cloud output by the initial point cloud alignment layer, and then the result point cloud is compared with the complete point cloud of the sample object in the real scene as a label, and the network parameters of the initial point cloud alignment layer are optimized according to the comparison result. This cycle is repeated until the initial point cloud alignment layer converges. At this time, the result point cloud output by the initial point cloud alignment layer is optimally aligned with the complete point cloud of the sample object in the real scene, and the training is completed to obtain the trained initial point cloud alignment layer, which is the point cloud alignment layer 32.

[0082] based on Figure 3 In the three-dimensional scene generation model shown, step 122 inputs the three-dimensional reconstructed latent vector and the corresponding target point cloud data corresponding to each target object into the three-dimensional asset generation module of the three-dimensional scene generation model to obtain the initial three-dimensional scene output by the three-dimensional asset generation module, which can be achieved through the following steps 1221 to 1223.

[0083] Step 1221: For each target object, the 3D reconstructed latent vector and target point cloud data corresponding to each target object are input into the point cloud image generation layer 31 to obtain the 3D asset latent vector corresponding to each target object output by the point cloud image generation layer 31.

[0084] Among them, the three-dimensional asset latent vector is used to represent the three-dimensional assets of the target object, that is, the three-dimensional model of the target object.

[0085] Step 1222: Input the three-dimensional asset latent vector and target point cloud data corresponding to each target object into the point cloud alignment layer 32 to obtain the aligned point cloud corresponding to each target object output by the point cloud alignment layer 32.

[0086] Step 1223: Generate an initial three-dimensional scene based on the scene generation layer 33 and the aligned point cloud corresponding to each target object.

[0087] For example, the electronic device may call the scene generation layer 33 , which combines the aligned point clouds of the target objects into a three-dimensional scene based on the target physical relationship of the target objects in the scene image to obtain an initial three-dimensional scene.

[0088] For example, the electronic device can use the aligned point cloud corresponding to each target object as target point cloud data, and re-input it into the point cloud image generation layer 31 together with the three-dimensional reconstructed latent vector corresponding to the target object, and repeat the three-dimensional asset construction based on the point cloud image generation layer 31 and the alignment operation based on the point cloud alignment layer 32 until the similarity between the aligned point cloud output by the point cloud alignment layer 32 and the complete point cloud of the target object in the real scene is greater than or equal to a preset similarity threshold; then, the aligned point cloud finally output by the point cloud alignment layer 32 is input into the scene generation layer 33 to obtain the initial three-dimensional scene output by the scene generation layer 33.

[0089] The scene generation layer 33 may combine the aligned point clouds of the target objects ultimately output by the point cloud alignment layer 32 into a three-dimensional scene based on the target physical relationship of the target objects in the scene image, thereby obtaining an initial three-dimensional scene.

[0090] In this way, through the continuous alignment operation of the point cloud alignment layer 32, the alignment effect of the aligned point cloud can be improved, so that the aligned point cloud used to construct the initial three-dimensional scene is closer to the complete point cloud of the target object in the real scene, thereby improving the quality of the three-dimensional assets of the target object in the initial three-dimensional scene, and thus improving the effect of the initial three-dimensional scene.

[0091] Step 130: Obtain the target physical relationship between each target object in the scene image, and perform physical relationship correction on the three-dimensional assets corresponding to each target object in the initial three-dimensional scene based on the target physical relationship to obtain the target three-dimensional scene corresponding to the scene image.

[0092] The target physical relationship reflects the positional relationship between target objects. Electronic devices can use the large language model to identify the physical relationship between target objects in a scene image and obtain the target physical relationship. For example, an electronic device can use the Florence-2 visual model to identify objects in a scene image, obtain recognition results, and then input the recognition results into the large language model to obtain the target objects output by the large language model and the target physical relationships between the target objects.

[0093] After the electronic device obtains the target physical relationship, it can further correct the placement of the three-dimensional assets corresponding to each target object in the initial three-dimensional scene based on the target physical relationship to prevent incorrect placement such as penetration and floating between target objects, thereby improving the effect of the generated three-dimensional scene.

[0094] Specifically, step 130 corrects the physical relationship of the three-dimensional assets corresponding to each target object in the initial three-dimensional scene based on the target physical relationship to obtain the target three-dimensional scene corresponding to the scene image, which can be achieved through the following steps 131 to 134.

[0095] Step 131: Determine the object to be corrected among all target objects in the initial three-dimensional scene according to the target physical relationship.

[0096] The target physical relationship reflects the positional relationship between target objects. The electronic device can determine the target objects with contact relationship in the initial three-dimensional scene as objects to be corrected based on the target physical relationship.

[0097] Step 132: for two target objects to be corrected among the objects to be corrected, determine target moving objects and corresponding target correction loss functions among the two target objects to be corrected based on the target physical relationship between the two target objects to be corrected.

[0098] For two target objects to be corrected among the objects to be corrected, such as a table and a teacup placed on the table, the target moving object that can be moved among the two target objects to be corrected can be determined based on the target physical relationship between the two target objects to be corrected, and then the corresponding target correction loss function can be determined based on the target moving object.

[0099] The target correction loss function is used to characterize the distances of all three-dimensional points of one target moving object relative to the other target moving object among the two target moving objects to be corrected.

[0100] Specifically, in one embodiment, determining a target moving object in two target objects to be corrected and a corresponding target correction loss function based on the target physical relationship between the two target objects to be corrected may include:

[0101] When the target physical relationship between two target objects to be corrected is a placement relationship, the placed object among the two target objects to be corrected is determined as the target moving object, and the correction loss function corresponding to the placed object is determined as the target correction loss function; when the target physical relationship between two target objects to be corrected is a non-placement contact relationship, both target objects to be corrected are determined as target moving objects, and the sum of the correction loss functions corresponding to each target object to be corrected is determined as the target correction loss function.

[0102] For example, assuming that the two target objects to be corrected are target object A and target object B, the correction loss function of target object A and target object B can be established respectively. The correction loss function can be defined as: traversing all three-dimensional points of the object itself, when the distance between each three-dimensional point and the other object is greater than 0, the total distance is the minimum.

[0103] If the target physical relationship between target object A and target object B is a placement relationship, for example, target object A is placed on target object B, which means that the two are in contact, but target object B is like a base and cannot be moved, then target object A can be determined as a target moving object. At this time, the correction loss function corresponding to target object A is determined as the target correction loss function. In this way, the positional relationship between objects A and B can be corrected only by modifying the position of object A.

[0104] If the target physical relationship between target object A and target object B is a non-placement contact relationship, for example, target object A is leaning next to target object B, which means that target object A is leaning against target object B and can move relative to each other, then target object A and target object B can both be determined as target moving objects. At this time, the corrected loss function corresponding to target object A and the corrected loss function corresponding to target object B can be added to obtain the target corrected loss function.

[0105] Step 133: By moving the three-dimensional assets corresponding to the target moving object, the loss value of the target correction loss function is iteratively optimized until the target correction loss function converges, thereby obtaining the initial three-dimensional scene after the physical relationship is corrected.

[0106] For example, referring to the example of target objects A and B in step 132, if target object A is placed on target object B, the modified loss function corresponding to target object A can be iteratively optimized by moving the 3D asset corresponding to target object A until convergence. If the target physical relationship between target objects A and B is a non-placement contact relationship, the sum of the modified loss functions corresponding to target object A and target object B can be iteratively optimized by moving the 3D assets corresponding to target objects A and B, respectively, until convergence.

[0107] When the target correction loss function converges, it means that the distances of all three-dimensional points of one target moving object among the two target objects to be corrected and the other target correction object are greater than 0, and the distance relative to the other target object to be corrected is the smallest.

[0108] Step 134: Determine the initial three-dimensional scene after the physical relationship correction as the target three-dimensional scene.

[0109] After the physical relationship of the initial three-dimensional scene is corrected, the initial three-dimensional scene with the corrected physical relationship can be determined as the target three-dimensional scene. At this time, the target objects in the target three-dimensional scene can be placed with the correct physical relationship, and incorrect placement relationships such as penetration and floating between target objects can be avoided, thereby improving the quality of the generated target three-dimensional scene.

[0110] The three-dimensional scene generation method provided in the embodiment of the present application determines the target objects in the scene image and the target point cloud data corresponding to each target object, inputs the scene image and the target point cloud data into the three-dimensional scene generation model, and obtains the initial three-dimensional scene corresponding to the scene image output by the three-dimensional scene generation model, wherein the three-dimensional scene generation model is used to generate each target object in the scene image as a three-dimensional asset based on the target point cloud data corresponding to each target object, and constructs the three-dimensional assets into an initial three-dimensional scene. In this way, the three-dimensional assets of each target object can be generated for the target point cloud data of each target object in the scene image, and then the initial three-dimensional scene corresponding to the scene image is constructed, thereby improving the quality of the three-dimensional assets of a single target object in the scene; further, by obtaining the target physical relationship between each target object in the scene image, and correcting the physical relationship of the three-dimensional assets corresponding to each target object in the initial three-dimensional scene based on the target physical relationship, the target three-dimensional scene corresponding to the scene image is obtained, so that each target object in the target three-dimensional scene can be placed in the correct physical relationship, further improving the quality of the generated target three-dimensional scene.

[0111] The present application also provides a three-dimensional scene generation device. Figure 4 The schematic diagram of the structure of the three-dimensional scene generation device provided in the embodiment of the present application is shown. Figure 4 As shown, the three-dimensional scene generating device may include:

[0112] A determination unit 410 is configured to determine target objects in a scene image and target point cloud data corresponding to each target object;

[0113] The 3D scene generation unit 420 is configured to input a scene image and target point cloud data into a 3D scene generation model to obtain an initial 3D scene corresponding to the scene image output by the 3D scene generation model. The 3D scene generation model is configured to generate each target object in the scene image into a 3D asset based on the target point cloud data corresponding to each target object, and to construct the 3D assets into an initial 3D scene.

[0114] The correction unit 430 is used to obtain the target physical relationship between each target object in the scene image, and perform physical relationship correction on the three-dimensional assets corresponding to each target object in the initial three-dimensional scene based on the target physical relationship to obtain the target three-dimensional scene corresponding to the scene image.

[0115] In one embodiment, the three-dimensional scene generation unit 420 may include: a feature extraction subunit, configured to input a scene image into a feature extraction module of a three-dimensional scene generation model, and obtain a three-dimensional reconstruction latent vector corresponding to each target object output by the feature extraction module; wherein the feature extraction module is configured to extract information for describing three-dimensional assets of each target object in the scene image, and the three-dimensional reconstruction latent vector is configured to characterize the three-dimensional asset information of the target object; a three-dimensional asset generation subunit, configured to input, for each target object, the three-dimensional reconstruction latent vector corresponding to each target object and the corresponding target point cloud data into the three-dimensional asset generation module of the three-dimensional scene generation model, and obtain an initial three-dimensional scene output by the three-dimensional asset generation module; wherein the three-dimensional asset generation module is configured to generate a three-dimensional asset corresponding to each target object, and generate each three-dimensional asset into a three-dimensional scene.

[0116] In one embodiment, the 3D asset generation module of the 3D scene generation model includes a point cloud image generation layer, a point cloud alignment layer, and a scene generation layer. Accordingly, the 3D asset generation subunit is specifically configured to: for each target object, input the 3D reconstruction latent vector and target point cloud data corresponding to each target object into the point cloud image generation layer, thereby obtaining a 3D asset latent vector corresponding to each target object output by the point cloud image generation layer; input the 3D asset latent vector and target point cloud data corresponding to each target object into the point cloud alignment layer, thereby obtaining an aligned point cloud corresponding to each target object output by the point cloud alignment layer; and generate an initial 3D scene based on the scene generation layer and the aligned point cloud corresponding to each target object. The 3D asset latent vector is used to represent the 3D assets of the target object; the point cloud image generation layer is used to construct the 3D asset based on the 3D reconstruction latent vector and target point cloud data; the point cloud alignment layer is used to align the target point cloud data with the complete point cloud of the target object in the real scene corresponding to the scene image based on the 3D asset latent vector; and the scene generation layer is used to generate a 3D scene from the input point cloud.

[0117] In one embodiment, when the three-dimensional asset generation subunit generates an initial three-dimensional scene based on the scene generation layer and the aligned point cloud corresponding to each target object, it is specifically used to: use the aligned point cloud corresponding to each target object as target point cloud data, and re-input it into the point cloud image generation layer together with the three-dimensional reconstructed latent vector, repeat the three-dimensional asset construction based on the point cloud image generation layer and the alignment operation based on the point cloud alignment layer until the similarity between the aligned point cloud output by the point cloud alignment layer and the complete point cloud is greater than or equal to a preset similarity threshold; input the aligned point cloud finally output by the point cloud alignment layer into the scene generation layer to obtain the initial three-dimensional scene output by the scene generation layer.

[0118] In one embodiment, the correction unit 430 may include: a first determination subunit, used to determine the object to be corrected among all target objects in the initial three-dimensional scene based on the target physical relationship; a second determination subunit, used to determine the target moving object and the corresponding target correction loss function among the two target objects to be corrected based on the target physical relationship between the two target objects to be corrected; an optimization subunit, used to iteratively optimize the loss value of the target correction loss function by moving the three-dimensional assets corresponding to the target moving objects until the target correction loss function converges to obtain the initial three-dimensional scene after the physical relationship correction; a third determination subunit, used to determine the initial three-dimensional scene after the physical relationship correction as the target three-dimensional scene.

[0119] In one embodiment, the second determination subunit is specifically used to: when the target physical relationship between two target objects to be corrected is a placement relationship, determine the placed object among the two target objects to be corrected as a target moving object, and determine the correction loss function corresponding to the placed object as the target correction loss function; when the target physical relationship between the two target objects to be corrected is a non-placement contact relationship, determine both target objects to be corrected as target moving objects, and determine the sum of the correction loss functions corresponding to each target object to be corrected as the target correction loss function; wherein, the target correction loss function is used to characterize the distance of all three-dimensional points of one of the two target moving objects to be corrected relative to the other target object to be corrected.

[0120] In one embodiment, the determination unit 410 may include: a marking subunit, used to determine the target objects in the scene image and mark the edges of each target object; a point cloud determination subunit, used to extract scene point cloud data of the scene image and determine the point cloud data in the scene point cloud data that is aligned with the pixels of each target object; and a matching subunit, used to match the point cloud data of each target object with the edges to obtain target point cloud data corresponding to each target object.

[0121] The implementation principles and beneficial effects of the three-dimensional scene generation device provided in the embodiment of the present application are similar to those of the three-dimensional scene generation method provided in the above embodiment, and will not be repeated here.

[0122] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, i.e., they may be located in one place or distributed across multiple network units. Some or all of the modules may be selected based on actual needs to achieve the objectives of the present embodiment. Those skilled in the art will be able to understand and implement the present invention without inventive effort.

[0123] An embodiment of the present application also provides an electronic device, including a memory and a processor, wherein the memory stores a computer program that can be run on the processor. When the processor executes the program, the steps of the three-dimensional scene generation method described in any of the above method embodiments are implemented, which will not be repeated here.

[0124] Based on the 3D scene generation method described in any of the above embodiments, embodiments of the present application further provide a computer-readable storage medium. For example, the non-transitory computer-readable storage medium may be a read-only memory (ROM), a random access memory (RAM), a CD-ROM, a magnetic tape, a floppy disk, or an optical data storage device. The storage medium stores computer instructions for executing the 3D scene generation method described in any of the above embodiments, which will not be further described herein.

[0125] Those skilled in the art will understand that all or part of the steps to implement the above embodiments may be accomplished by hardware, or may be accomplished by a program to instruct the relevant hardware, and the program may be stored in a computer-readable storage medium, and the above-mentioned storage medium may be a read-only memory, a disk or an optical disk, etc.

[0126] Those skilled in the art will readily appreciate other embodiments of the present application after considering the specification and practicing the contents disclosed herein. This application is intended to cover any variations, uses, or adaptations herein that follow the general principles of this application and include common knowledge or customary techniques in the art not disclosed herein. The description and examples are to be considered merely exemplary, and the true scope and spirit of this application are indicated by the claims.

Claims

1. A three-dimensional scene generation method, characterized in that: include: Determine target objects in the scene image and target point cloud data corresponding to each target object; Inputting the scene image and the target point cloud data into a three-dimensional scene generation model to obtain an initial three-dimensional scene corresponding to the scene image output by the three-dimensional scene generation model; The three-dimensional scene generation model is used to generate each target object in the scene image into a three-dimensional asset based on the target point cloud data corresponding to each target object, and construct the three-dimensional assets into the initial three-dimensional scene; A target physical relationship between each target object in the scene image is obtained, and based on the target physical relationship, a physical relationship correction is performed on the three-dimensional assets corresponding to each target object in the initial three-dimensional scene to obtain a target three-dimensional scene corresponding to the scene image.

2. The three-dimensional scene generation method according to claim 1, characterized in that: The step of inputting the scene image and the target point cloud data into a three-dimensional scene generation model to obtain an initial three-dimensional scene corresponding to the scene image output by the three-dimensional scene generation model includes: Inputting the scene image into a feature extraction module of the 3D scene generation model, and obtaining a 3D reconstruction latent vector corresponding to each target object output by the feature extraction module; wherein the feature extraction module is used to extract information describing the 3D asset of each target object in the scene image; and the 3D reconstruction latent vector is used to represent the 3D asset information of the target object; For each target object, the three-dimensional reconstructed latent vector and the corresponding target point cloud data corresponding to each target object are input into the three-dimensional asset generation module of the three-dimensional scene generation model to obtain the initial three-dimensional scene output by the three-dimensional asset generation module; wherein the three-dimensional asset generation module is used to generate a three-dimensional asset corresponding to each target object, and generate each three-dimensional asset into a three-dimensional scene.

3. The three-dimensional scene generation method according to claim 2, characterized in that: The 3D asset generation module of the 3D scene generation model includes a point cloud image generation layer, a point cloud alignment layer, and a scene generation layer; inputting the 3D reconstructed latent vector corresponding to each target object and the corresponding target point cloud data into the 3D asset generation module of the 3D scene generation model to obtain the initial 3D scene output by the 3D asset generation module includes: For each target object, the 3D reconstruction latent vector and the target point cloud data corresponding to each target object are input into the point cloud image generation layer, and a 3D asset latent vector corresponding to each target object is output by the point cloud image generation layer; wherein the 3D asset latent vector is used to represent the 3D asset of the target object; and the point cloud image generation layer is used to construct the 3D asset based on the 3D reconstruction latent vector and the target point cloud data. Inputting the three-dimensional asset latent vector and the target point cloud data corresponding to each target object into the point cloud alignment layer, obtaining an aligned point cloud corresponding to each target object output by the point cloud alignment layer; the point cloud alignment layer is used to align the target point cloud data with the complete point cloud of the target object in the real scene corresponding to the scene image based on the three-dimensional asset latent vector; The initial three-dimensional scene is generated based on the scene generation layer and the aligned point cloud corresponding to each target object; the scene generation layer is used to generate the input point cloud into a three-dimensional scene.

4. The three-dimensional scene generation method according to claim 3, characterized in that: Generating the initial three-dimensional scene based on the scene generation layer and the aligned point cloud corresponding to each target object includes: The aligned point cloud corresponding to each target object is used as the target point cloud data and re-inputted into the point cloud image generation layer together with the 3D reconstruction latent vector. The 3D asset construction based on the point cloud image generation layer and the alignment operation based on the point cloud alignment layer are repeated until the similarity between the aligned point cloud output by the point cloud alignment layer and the complete point cloud is greater than or equal to a preset similarity threshold. The aligned point cloud finally output by the point cloud alignment layer is input into the scene generation layer to obtain the initial three-dimensional scene output by the scene generation layer.

5. The three-dimensional scene generation method according to any one of claims 1 to 4, characterized in that: The performing physical relationship correction on the three-dimensional assets corresponding to the target objects in the initial three-dimensional scene based on the target physical relationship to obtain the target three-dimensional scene corresponding to the scene image includes: Determining an object to be corrected among all the target objects in the initial three-dimensional scene according to the target physical relationship; For two target objects to be corrected among the objects to be corrected, determining a target moving object and a corresponding target correction loss function among the two target objects to be corrected based on the target physical relationship between the two target objects to be corrected; By moving the three-dimensional asset corresponding to the target moving object, iteratively optimizing the loss value of the target correction loss function until the target correction loss function converges, thereby obtaining the initial three-dimensional scene after physical relationship correction; The initial three-dimensional scene after the physical relationship correction is determined as the target three-dimensional scene.

6. The three-dimensional scene generation method according to claim 5, characterized in that: The determining of a target moving object and a corresponding target correction loss function in the two target objects to be corrected based on the target physical relationship between the two target objects to be corrected includes: In a case where the target physical relationship between the two target objects to be corrected is a placement relationship, determining the placed object among the two target objects to be corrected as the target moving object, and determining the correction loss function corresponding to the placed object as the target correction loss function; When the target physical relationship between the two target objects to be corrected is a non-placement contact relationship, both the target objects to be corrected are determined as the target moving objects, and the sum of the correction loss functions corresponding to each target object to be corrected is determined as the target correction loss function; The target correction loss function is used to characterize the distances of all three-dimensional points of one of the two target moving objects to be corrected relative to the other target moving object to be corrected.

7. The three-dimensional scene generation method according to any one of claims 1 to 4, characterized in that: The determining of the target objects in the scene image and the target point cloud data corresponding to each of the target objects includes: Identify target objects in the scene image and mark the edges of each target object; Extracting scene point cloud data of the scene image, and determining point cloud data in the scene point cloud data that is aligned with pixels of each of the target objects; The point cloud data of each target object is matched with the edge to obtain the target point cloud data corresponding to each target object.

8. A three-dimensional scene generation device, characterized in that: include: a determination unit, configured to determine target objects in a scene image and target point cloud data corresponding to each target object; a 3D scene generation unit, configured to input the scene image and the target point cloud data into a 3D scene generation model, and obtain an initial 3D scene corresponding to the scene image output by the 3D scene generation model; The three-dimensional scene generation model is used to generate each target object in the scene image into a three-dimensional asset based on the target point cloud data corresponding to each target object, and construct the three-dimensional assets into the initial three-dimensional scene; A correction unit is used to obtain the target physical relationship between each target object in the scene image, and based on the target physical relationship, perform physical relationship correction on the three-dimensional assets corresponding to each target object in the initial three-dimensional scene to obtain the target three-dimensional scene corresponding to the scene image.

9. An electronic device comprising a memory and a processor, wherein the memory stores a computer program that can be run on the processor, wherein: When the processor executes the program, the steps of the three-dimensional scene generation method according to any one of claims 1 to 7 are implemented.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the three-dimensional scene generation method according to any one of claims 1 to 7 are implemented.

Citation Information

Patent Citations

  • Three-dimensional scene virtual generation method and system based on GAN network

    CN110706328A

  • 3D scene graph generation method, system and device based on point cloud and readable storage medium

    CN114266863A

  • Three-dimensional scene generation method and device, program product and medium

    CN119579800A

  • Digital twin modeling method, device and equipment and storage medium

    CN119625213A

  • Generation optimization method and device of intelligent body three-dimensional virtual scene, equipment and medium

    CN120014209A