3D scene generation method and related device

By acquiring image and labeling information, the target neural network is used to generate 3D scenes that meet the preset layout, solving the problem of difficult adjustment of object layout and collision occlusion in the prior art, and achieving efficient 3D scene design and development.

CN120259600APending Publication Date: 2025-07-04ORIENTAL POWER HOLDINGS LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202410003858.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-01-02
Publication Date
2025-07-04

AI Technical Summary

Technical Problem

The object types and layouts in 3D scenes generated by existing AI technologies are difficult to flexibly adjust according to user needs, and the objects in the generated scenes are prone to collision and occlusion relationships.

Method used

By obtaining the image and labeling information containing the object frame, the target neural network is used to generate a 3D scene that satisfies the preset layout, and training is combined with the latent vector and signed distance function SDF to ensure the correct position and independence of the object in the 3D scene.

Benefits of technology

It realizes the generation of 3D scenes that meet the preset layout requirements, improves design and development efficiency, and ensures the independence and correct position of objects in the scene, reducing collision and occlusion problems.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120259600A_ABST
    Figure CN120259600A_ABST
Patent Text Reader

Abstract

The invention provides a 3D scene generation method and a related device. The embodiment of the invention can be applied to the field of artificial intelligence. The method comprises the steps that a first image and annotation information are acquired, the first image comprises at least one object frame, and the annotation information indicates the category of a target object corresponding to the object frame; and processing the first image based on the annotation information through the target neural network to obtain a three-dimensional (3D) scene comprising the target object, the projection position of the target object on the first plane of the 3D scene corresponding to the position of the object frame on the first image. According to the method, the function of generating the 3D scene through the two-dimensional planar graph and semantics is realized, the generated 3D scene meets the preset layout requirement, and the design and development efficiency of the 3D scene is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of artificial intelligence technology, and in particular, to a 3D scene generation method and related devices. Background Art

[0002] Artificial Intelligence (AI) is the theory, method, technology, and application system that uses digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use knowledge to obtain the best results. Among them, the application of AI technology in three-dimensional (3D) model generation has been relatively extensive. Compared with the traditional method of manually operating 3D scene modeling, AI technology can automatically or semi-automatically generate 3D models through machine learning and deep learning methods.

[0003] Based on AI technology, a corresponding 3D scene can be generated through given text. However, this 3D scene is usually just a random scene that conforms to the semantic content of the text. The types of objects in the scene and the layout of each object in the scene are both random, and it is difficult to flexibly adjust according to user needs. Summary of the Invention

[0004] Embodiments of this application provide a 3D scene generation method and related devices for generating a 3D scene that meets a preset layout.

[0005] The first aspect of this application provides a 3D scene generation method, including:

[0006] Obtain a first image and annotation information, where the first image includes at least one object box, and the annotation information indicates the category of the target object corresponding to the object box;

[0007] Process the first image based on the annotation information through a target neural network to obtain a three-dimensional (3D) scene including the target object, and the projection position of the target object on the first plane of the 3D scene corresponds to the position of the object box on the first image.

[0008] In a possible implementation method, processing the first image based on the annotation information through a target neural network to obtain a 3D scene including the target object includes:

[0009] Extract a latent vector of the first image based on the annotation information through a target neural network;

[0010] Process the latent vector through a target neural network to obtain a 3D scene.

[0011] In a possible implementation method, the number of object bounding boxes is at least two, including a first object bounding box and a second object bounding box. The first object bounding box corresponds to a first target object, and the second object bounding box corresponds to a second target object;

[0012] When the first object bounding box occludes the second object bounding box, in the 3D scene, the first target object is located on the other side of the second target object relative to the first plane.

[0013] In a possible implementation method, after processing the first image based on the annotation information through a target neural network to obtain a 3D scene including target objects, it further includes:

[0014] In response to a movement operation on the target object, change the position of the target object in the 3D scene.

[0015] In a possible implementation method, the object bounding box is a rectangular box, and the projection range of the target object on the first plane corresponds to the size range of the rectangular box.

[0016] The second aspect of this application provides a training method for a neural network, including:

[0017] Obtain a training 3D scene and annotation information. The training 3D scene includes at least one target object, and the annotation information indicates the category of the target object;

[0018] Obtain training images, where the training images include object bounding boxes, and the position of the object bounding boxes on the training images corresponds to the projection position of the target object on the first plane of the training 3D scene;

[0019] Process the training images based on the annotation information through a first neural network to obtain a predicted 3D scene including predicted objects;

[0020] Train the first neural network according to a first loss function. The first loss function indicates the similarity between the predicted layout and the training layout. The predicted layout is the layout of the predicted objects in the predicted 3D scene, and the training layout is the layout of the target objects in the training 3D scene.

[0021] In a possible implementation method, after obtaining the training 3D scene and annotation information, it further includes:

[0022] Extract a second latent vector of the training 3D scene according to the annotation information;

[0023] Processing the training images based on the annotation information through a first neural network to obtain a predicted 3D scene including predicted objects includes:

[0024] Extract a first latent vector of the training images according to the annotation information through a first neural network;

[0025] Process the first latent vector through a first neural network to obtain a predicted 3D scene;

[0026] The first loss function indicates the similarity between the first latent vector and the second latent vector.

[0027] In a possible implementation method, before extracting the second latent vector of the training 3D scene according to the annotation information, it further includes:

[0028] Encode the training 3D scene into first parameters through a second neural network, where the first parameters are triplane parameters;

[0029] Process the first parameters based on the annotation information through a second neural network to obtain the target signed distance function (SDF) corresponding to each target object;

[0030] Train the second neural network according to the second loss function, where the second loss function indicates the similarity between the target SDF and the training SDF, and the training SDF is obtained by measuring the target objects in the training 3D scene;

[0031] Extracting the second latent vector of the training 3D scene according to the annotation information includes:

[0032] Encode the training 3D scene into second parameters according to the target SDF, where the second parameters are triplane parameters;

[0033] Extract the second latent vector of the second parameters.

[0034] In a possible implementation method, the number of target objects in the training 3D scene is at least two;

[0035] Obtaining the training 3D scene and annotation information includes:

[0036] Obtain a sample 3D scene and annotation information, where the sample 3D scene includes at least two target objects;

[0037] Perform collision detection on the target objects in the sample 3D scene;

[0038] Adjust the positions of the target objects in the sample 3D scene according to the detection results to obtain a training 3D scene, where at least two target objects in the training 3D scene do not penetrate each other.

[0039] In a possible implementation method, performing collision detection on the target objects in the sample 3D scene includes:

[0040] Based on the annotation information, perform watertight processing on the target objects in the sample 3D scene to obtain corresponding object meshes;

[0041] Perform collision detection on the object grid to obtain a detection result, where the detection result includes the conflict area between the object grids.

[0042] In a possible implementation method, after processing the training image based on the annotation information through the first neural network to obtain a predicted 3D scene including predicted objects, it further includes:

[0043] Establish corresponding 3D modules for the predicted objects in the predicted 3D scene through the first neural network;

[0044] In response to a movement operation on the predicted object, change the position of the predicted object in the predicted 3D scene.

[0045] The third aspect of this application provides a 3D scene generation device, including:

[0046] A first acquisition unit for acquiring a first image and annotation information, where the first image includes at least one object box, and the annotation information indicates the category of the target object corresponding to the object box;

[0047] A first modeling unit for processing the first image based on the annotation information through a target neural network to obtain a three-dimensional 3D scene including the target object, and the projection position of the target object on the first plane of the 3D scene corresponds to the position of the object box on the first image.

[0048] In a possible implementation method,

[0049] The first modeling unit is specifically configured to extract the latent vector of the first image based on the annotation information through the target neural network; process the latent vector through the target neural network to obtain a 3D scene.

[0050] In a possible implementation method, the number of object boxes is at least two, including a first object box and a second object box, the first object box corresponds to a first target object, and the second object box corresponds to a second target object;

[0051] When the first object box occludes the second object box, in the 3D scene, the first target object is located on the other side of the second target object relative to the first plane.

[0052] In a possible implementation method, it further includes:

[0053] A first movement unit for changing the position of the target object in the 3D scene in response to a movement operation on the target object.

[0054] In a possible implementation method, the object box is a rectangular box, and the projection range of the target object on the first plane corresponds to the size range of the rectangular box.

[0055] The fourth aspect of the present application provides a training device for a neural network, including:

[0056] A second acquisition unit, configured to acquire a training 3D scene and annotation information, where the training 3D scene includes at least one target object, and the annotation information indicates the category of the target object;

[0057] The second acquisition unit is further configured to acquire a training image, where the training image includes an object box, and the position of the object box on the training image corresponds to the projection position of the target object on the first plane of the training 3D scene;

[0058] A second modeling unit, configured to process the training image based on the annotation information through a first neural network to obtain a predicted 3D scene including predicted objects;

[0059] A training unit, configured to train the first neural network according to a first loss function, where the first loss function indicates the similarity between the predicted layout and the training layout, the predicted layout is the layout of the predicted objects in the predicted 3D scene, and the training layout is the layout of the target objects in the training 3D scene.

[0060] In a possible implementation method,

[0061] A first extraction unit, configured to extract a second latent vector of the training 3D scene according to the annotation information;

[0062] The second modeling unit is specifically configured to extract a first latent vector of the training image according to the annotation information through a first neural network; process the first latent vector through the first neural network to obtain a predicted 3D scene; the first loss function indicates the similarity between the first latent vector and the second latent vector.

[0063] In a possible implementation method, it further includes:

[0064] A second parameter adjustment unit, configured to encode the training 3D scene into a first parameter through a second neural network, where the first parameter is a three-plane parameter; process the first parameter based on the annotation information through the second neural network to obtain a target signed distance function SDF corresponding to each target object;

[0065] The training unit is further configured to train the second neural network according to a second loss function, where the second loss function indicates the similarity between the target SDF and the training SDF, and the training SDF is obtained by measuring the target objects in the training 3D scene;

[0066] The first extraction unit is specifically configured to encode the training 3D scene into a second parameter according to the target SDF, where the second parameter is a three-plane parameter; extract the second latent vector of the second parameter.

[0067] In a possible implementation method, the number of target objects in the training 3D scene is at least two;

[0068] The second acquisition unit is specifically configured to acquire a sample 3D scene and annotation information, where the sample 3D scene includes at least two target objects; perform collision detection on the target objects in the sample 3D scene; adjust the positions of the target objects in the sample 3D scene according to the detection results to obtain a training 3D scene, and there is no penetration between at least two target objects in the training 3D scene.

[0069] In a possible implementation method, performing collision detection on the target objects in the sample 3D scene includes:

[0070] Based on the annotation information, perform watertight processing on the target objects in the sample 3D scene to obtain corresponding object meshes;

[0071] Perform collision detection on the object meshes to obtain detection results, where the detection results include the conflict regions between the object meshes.

[0072] In a possible implementation method,

[0073] The second modeling unit is further configured to establish a corresponding 3D module for the predicted object in the predicted 3D scene through a first neural network;

[0074] It further includes: a second moving unit, configured to change the position of the predicted object in the predicted 3D scene in response to a moving operation on the predicted object.

[0075] The fifth aspect of the present application provides a computer device, including:

[0076] A memory, a transceiver, a processor, and a bus system;

[0077] Wherein, the memory is used to store programs;

[0078] The processor is configured to execute the programs in the memory, including executing the methods in the above aspects;

[0079] The bus system is used to connect the memory and the processor to enable communication between the memory and the processor.

[0080] The sixth aspect of the present application provides a computer-readable storage medium, where instructions are stored in the computer-readable storage medium, and when the instructions are run on a computer, the computer is caused to execute the methods in the above aspects.

[0081] A seventh aspect of the present application provides a computer program product or a computer program. The computer program product or the computer program includes computer instructions, and the computer instructions are stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device executes the methods provided in the above aspects.

[0082] As can be seen from the above technical solutions, the embodiments of the present application have the following advantages:

[0083] The present application provides a 3D scene generation method and a related device. By obtaining a first image and annotation information, the first image includes at least one object box, and the annotation information indicates the category of the target object corresponding to the object box; through a target neural network, the first image is processed based on the annotation information to obtain a three-dimensional (3D) scene including the target object, and the projection position of the target object on the first plane of the 3D scene corresponds to the position of the object box on the first image, realizing the function of generating a 3D scene through a two-dimensional plan view and semantics. The generated 3D scene meets the preset layout requirements, which is beneficial to improving the design and development efficiency of the 3D scene. BRIEF DESCRIPTION OF THE DRAWINGS

[0084] Figure 1 It is an application environment diagram of the 3D scene generation method in the embodiments of the present application;

[0085] Figure 2 It is a method flow chart of the 3D scene generation method provided by an embodiment of the present application;

[0086] Figure 3 It is a schematic diagram of the 3D scene generation method provided by the embodiments of the present application;

[0087] Figure 4 It is a method flow chart of the training method of the neural network provided by the embodiments of the present application;

[0088] Figure 5 It is a method flow chart of the training method of the neural network provided by the embodiments of the present application;

[0089] Figure 6a It is a flow chart for preprocessing the training 3D scene;

[0090] Figure 6b It is a three-dimensional representation flow chart of the three-plane parameters;

[0091] Figure 7 It is a schematic flow diagram of the training method of the neural network provided by the embodiments of the present application;

[0092] Figure 8 It is a schematic diagram of an embodiment of the 3D scene generation device in the embodiments of the present application;

[0093] Figure 9 This is a schematic diagram of an embodiment for training a neural network in an embodiment of the present application;

[0094] Figure 10 This is a schematic diagram of a server structure provided by an embodiment of the present application. Detailed implementation manners

[0095] The terms "first", "second", "third", "fourth", etc. (if any) in the specification, claims and drawings of the present application are used to distinguish similar objects and do not necessarily describe a specific order or sequence. It should be understood that such data can be interchanged under appropriate circumstances so that the embodiments of the present application described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "comprising" and "corresponding to" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device that includes a series of steps or units does not necessarily limit to those clearly listed steps or units, but may include other steps or units not clearly listed or inherent to these process, method, product or device.

[0096] Artificial Intelligence (AI) is to use a digital computer or a machine controlled by a digital computer to simulate, extend and expand human intelligence, a theory, method, technology and application system that can perceive the environment, acquire knowledge and use knowledge to obtain the best results. In other words, artificial intelligence is a comprehensive technology in computer science. It attempts to understand the essence of intelligence and produce a new intelligent machine that can react in a way similar to human intelligence. Artificial intelligence also studies the design principles and implementation methods of various intelligent machines to enable the machines to have the functions of perception, reasoning and decision-making.

[0097] Artificial intelligence technology is an interdisciplinary subject, involving a wide range of fields, including both hardware-level technologies and software-level technologies. Artificial intelligence basic technologies generally include technologies such as sensors, dedicated artificial intelligence chips, cloud computing, distributed storage, big data processing technology, operation / interaction systems, and mechatronics. Artificial intelligence software technologies mainly include several major directions such as computer vision technology, speech processing technology, natural language processing technology, and machine learning / deep learning.

[0098] Computer Vision Technology (CV) Computer vision is a science that studies how to enable machines to "see". More specifically, it refers to machine vision that uses cameras and computers to replace human eyes for target recognition and measurement, and further performs image processing to make the computer-processed images more suitable for human eyes to observe or be transmitted to instruments for detection. As a scientific discipline, computer vision research related theories and technologies, and attempts to establish artificial intelligence systems that can obtain information from images or multi-dimensional data. Computer vision technology usually includes image processing, image recognition, image semantic understanding, image retrieval, OCR, video processing, video semantic understanding, video content / behavior recognition, 3D object reconstruction, 3D technology, virtual reality, augmented reality, simultaneous localization and mapping, and other technologies.

[0099] Natural Language Processing (NLP) is an important direction in the fields of computer science and artificial intelligence. It studies various theories and methods that can achieve effective communication between humans and computers in natural language. Natural language processing is a science that integrates linguistics, computer science, and mathematics. Therefore, the research in this field will involve natural language, that is, the language people use in daily life, so it has a close connection with the research of linguistics. Natural language processing technology usually includes text processing, semantic understanding, machine translation, robot question answering, knowledge graph, and other technologies.

[0100] Machine Learning (ML) is an interdisciplinary subject that involves multiple disciplines such as probability theory, statistics, approximation theory, convex analysis, and algorithm complexity theory. It specifically studies how computers simulate or implement human learning behaviors to acquire new knowledge or skills, and reorganize the existing knowledge structure to continuously improve their own performance. Machine learning is the core of artificial intelligence and the fundamental way to make computers intelligent, and its applications cover all fields of artificial intelligence. Machine learning and deep learning usually include technologies such as artificial neural networks, belief networks, reinforcement learning, transfer learning, inductive learning, and formal learning.

[0101] With the research and progress of artificial intelligence technology, artificial intelligence technology has been studied and applied in multiple fields. For example, common ones include smart homes, smart wearable devices, virtual assistants, smart speakers, smart marketing, driverless, autonomous driving, drones, robots, smart healthcare, smart customer service, etc. It is believed that with the development of technology, artificial intelligence technology will be applied in more fields and play an increasingly important role.

[0102] The solution provided by the embodiments of this application designs artificial intelligence technologies such as image recognition, semantic analysis, and machine learning, and is used to analyze images by combining annotation semantics to generate a 3D model that meets the layout requirements, which will be specifically described through the embodiments in the following text.

[0103] Text2Room is an existing 3D scene generation technology that is used to generate a textured room-scale textured three-dimensional mesh model from a given text prompt. It uses a pre-trained two-dimensional (2D) diffusion model to synthesize a series of 2D images from different poses, and then uses a pre-trained depth estimation model to obtain a three-channel red-green-blue depth (RGB-D) image. After obtaining the RGB-D image, Text2Room slightly moves the camera position from the previous frame. Since the camera position has changed, the RGB-D image of the previous frame cannot be directly used for the current frame. Therefore, the RGB-D image of the previous frame is extended or "painted" to the current frame through out-painting; by continuously moving the camera position, a series of consecutive RGB-D images are obtained; and a 3D scene model is obtained by fusing multiple frames of RGB-D images.

[0104] However, Text2Room has some drawbacks: it is overly dependent on the depth estimation model, resulting in distorted 3D geometry and low accuracy. It is difficult to avoid collisions between objects in the scene and coordinate the occlusion relationships between objects in the generated multiple frames of RGB-D images, which may lead to incorrect or incomplete 3D geometries; in addition, the indoor scenes generated by Text2Room are usually a single piece of mesh, and the objects in the scene are adhered to each other, making it difficult to adjust.

[0105] To solve the above problems, the present application provides a 3D scene generation method and related devices. It should be understood that the 3D scene generation method provided by the present application is used to construct 3D scenes. For example, the present application can be widely applied to the following application scenarios to improve the design and development efficiency for such scenarios: Virtual Reality (VR): By constructing 3D scenes, users or virtual objects are brought into virtual scenes for interaction. Virtual scenes include virtual game scenes, VR scenes, etc.; Augmented Reality (AR): By adding virtual objects to the real environment, the combination of 3D scenes and the real world is achieved; Extended Reality (XR): Through 3D scene generation technology, realistic virtual scenes and objects are created and integrated with the real scene; Architecture and Design: 3D scene technology is used to create and visualize designs for better communication and improvement of designs. This technology can also be applied to the fields of interior design and decoration; Game Development: 3D scene generation technology is often used in game development to create realistic game environments and characters. In addition, 3D digital content generation tools for ordinary users are also key elements in realizing the metaverse.

[0106] For ease of understanding, please refer to Figure 1 , Figure 1 which is the application environment diagram of the 3D scene generation method in the embodiments of the present application. As Figure 1 shown, the 3D scene generation method in the embodiments of the present application is applied to a 3D scene generation system. The 3D scene generation system includes: a server and a terminal device; wherein, the server can be an independent physical server, or a server cluster or distributed system composed of multiple physical servers, or a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, Content Delivery Network (CDN), and big data and artificial intelligence platforms. The terminal can be a smart phone, a tablet computer, a laptop computer, a desktop computer, a smart speaker, a smart watch, etc., but is not limited thereto. The terminal and the server can be directly or indirectly connected through wired or wireless communication methods, and the embodiments of the present application do not limit this.

[0107] The method provided by this application is mainly applied to a server, which can be a Central Processing Unit (CPU) or a Graphics Processing Unit (GPU). The terminal device sends the collected training data to the server, and the training data includes a training 3D scene, annotation information, and training images. The server trains based on the training data to obtain a target neural network. When the terminal device sends a separate target image to the server, the server can input the target image into the target neural network, and output a target 3D scene corresponding to the target image through the target neural network.

[0108] Next, from the perspective of the server, the 3D scene generation method in this application will be introduced. Please refer to Figure 2 , the 3D scene generation method provided by the embodiment of this application includes:

[0109] 201, Obtain a first image and annotation information. The first image includes at least one object box, and the annotation information indicates the category of the target object corresponding to the object box.

[0110] In this embodiment, first, a first image and corresponding annotation information need to be obtained. The first image is a plan view and includes at least one object box. An object box refers to a way of annotating a target object in the first image, which frames the position of the target object in the image. Annotation information usually refers to the information for marking the target object in the image, indicating the category of the target object.

[0111] Please refer to Figure 3 , Figure 3 is a schematic diagram of the 3D scene generation method provided by the embodiment of this application. Figure 3 On the left side in

[0112] is the first image 300, and there are multiple object boxes on the first image 300, which are object boxes 301 to 306 respectively. Each object box has corresponding annotation information, indicating the category of the target object corresponding to the object box. For example, the annotation information of object box 301 is ceiling lamp, the annotation information of object box 302 is dining table, the annotation information of object box 303 is ceiling lamp, the annotation information of object box 304 is sofa, the annotation information of object box 305 is clothes hanger, and the annotation information of object box 306 is TV cabinet.

[0113] As Figure 3As shown, the object frame is represented by a rectangular frame, which can be used to indicate the size relationship between corresponding furniture. For example, if the size of the rectangular frame 306 is smaller than that of the rectangular frame 304, then in the 3D scene generated according to the first image 300, the floor area occupied by the TV cabinet is smaller than that of the sofa. Of course, the object frame can also be other graphics, which can correspond to the actual projected area of the furniture on the ground. For example, if the object frame 302 is a dining table and a round tabletop is desired for the generated dining table, the object frame 302 can be drawn as a circle. The embodiments of the present application do not limit the shape of the object frame. In short, the positional relationship and size relationship between the object frames are used to reflect the placement positions and relative sizes of the corresponding target objects.

[0114] Furthermore, the color of the object frame can also be set. This color can correspond to the annotation information or can be used to distinguish different target objects, such as Figure 3 As shown, the first image 300 includes 6 object frames, and these 6 object frames correspond to 5 object types, namely: the object frames 301 and 303 both represent ceiling lights, the object frame 302 represents a dining table, the object frame 304 represents a sofa, the object frame 305 represents a clothes hanger, and the object frame 306 represents a TV cabinet.

[0115] Then when setting the color of the object frame, the 6 object frames can either be set to different colors to distinguish each individual; or the object frames corresponding to the same type of target objects can be set to the same color. For example, the annotation information of the object frames 301 and 303 is both ceiling lights, so the colors of the object frames 301 and 303 are the same, and the colors of the other object frames are all different.

[0116] 202, through the target neural network, based on the annotation information, the first image is processed to obtain a 3D scene including the target objects, and the projection position of the target objects on the first plane of the 3D scene corresponds to the position of the object frames on the first image.

[0117] In this embodiment, using the annotation information, the first image is processed through the target neural network, and the result of the processing is to generate a 3D scene including the target objects. In this 3D scene, the projection position of the target objects on the first plane corresponds to the position of the object frames on the first image. It can be understood that the target neural network is a pre-trained neural network, and machine learning is used to let the neural network learn the mapping relationship from the first image to the 3D scene.

[0118] Such as Figure 3 As shown, Figure 3On the right side is a 3D scene 310 generated based on the first image. The 3D scene 310 includes multiple target objects, namely target objects 311 to 316. It can be understood that target objects 311 to 316 correspond one-to-one with object frames 301 to 306 on the first image 300. It can be understood that the first plane in step 202 corresponds to the bottom surface in the 3D scene 310, and the projection of the target objects in the 3D scene 310 on the first plane corresponds to the first image 300.

[0119] In a possible implementation method, the object frame in the first image is a rectangular frame, and the projection range of the target object on the first plane corresponds to the coverage range of the rectangular frame.

[0120] As Figure 3 shown, the object frame in the first image is represented by a rectangular frame. Taking object frame 304 as an example, in the 3D scene 310, it corresponds to target object 314 (sofa). The size range of object frame 304 in the first image 300 corresponds to the projection range of the sofa on the ground.

[0121] It can be understood that the size range in this embodiment corresponds to the projection range, which is used to represent the size of the target object in the 3D scene, rather than limiting the size of the target object. The corresponding relationship between the above size range and the projection range can mean that the projection range and the size range meet preset requirements. For example, if the similarity between the projection range and the size range is higher than a preset threshold, it means that the projection range and the size range correspond.

[0122] The 3D scene generation method provided by the embodiments of the present application obtains a first image and annotation information. The first image includes at least one object frame, and the annotation information indicates the category of the target object corresponding to the object frame; through a target neural network, the first image is processed based on the annotation information to obtain a three-dimensional 3D scene including target objects. The projection position of the target object on the first plane of the 3D scene corresponds to the position of the object frame on the first image, realizing the function of generating a 3D scene through a two-dimensional plan view and semantics. The generated 3D scene meets the preset layout requirements, which is beneficial to improving the design and development efficiency of the 3D scene.

[0123] In a possible implementation method, step 202 specifically includes:

[0124] 2021, through a target neural network, extract the latent vector of the first image based on the annotation information;

[0125] 2022, process the latent vector through a target neural network to obtain a 3D scene.

[0126] It can be understood that a latent vector refers to a low-dimensional representation hidden in data, capable of capturing the main features and structure of the input data. By processing the first image through a pre-trained target neural network based on the annotation information, the latent vector of the first image can be extracted, thereby capturing the latent features in the first image. In this embodiment, the target neural network includes an autoencoder or a variational autoencoder: An autoencoder is an unsupervised deep learning model whose goal is to learn the compression and encoding method of the input data and then reconstruct the input data as much as possible. During the training process of the autoencoder, the autoencoder will learn a latent representation; A variational autoencoder is an extension of the autoencoder that learns the distribution of latent vectors by introducing variational inference. In a variational autoencoder, the latent vector is regarded as a random variable and described using a probability distribution. By maximizing the expected value of the latent vector likelihood function, the variational autoencoder can learn a more accurate latent representation.

[0127] After obtaining the latent vector, it is also necessary to generate a 3D scene through the target neural network using the latent vector. In this embodiment, the target neural network includes a generator network: The generator network is used to generate a 3D scene according to the learned distribution.

[0128] In this embodiment, considering that image data belongs to two-dimensional data, when the image data carries annotation information, the combined data is more complex and difficult for the neural network model to understand. Therefore, the latent vector is introduced in this application embodiment. The latent vector can reduce the high-dimensional data to a low-dimensional vector, thereby reducing the complexity and dimension of the data and facilitating further data processing and analysis; In addition, since the latent vector belongs to a data compression technology, after extracting the latent vector from the data, not only can the data storage space and transmission time be reduced, but also computing resources can be saved during the subsequent data processing; At the same time, the latent vector can extract the main features and structure of the data, thus facilitating machine learning tasks such as classification, clustering, and recognition.

[0129] In a possible implementation method, the number of object boxes in the first image is at least two, including a first object box and a second object box, where the first object box corresponds to a first target object and the second object box corresponds to a second target object;

[0130] When the first object box occludes the second object box, in the 3D scene, the first target object is located on the other side of the second target object relative to the first plane.

[0131] In this embodiment, the first image includes object frames that occlude each other. At this time, the position depths of the target objects corresponding to the two object frames can be determined according to the occlusion relationship. That is, when the first object frame occludes the second object frame, it means that in the first image, the layer of the first object is above the layer of the second object. Then, in the generated 3D image, when the viewing angle of the 3D image corresponds to the first image, the first target object is closer to the viewing angle direction than the second target object, that is, the first target object is located on the other side of the second target object relative to the first plane.

[0132] As Figure 3 shown, the first plane corresponds to the ground of the 3D image 310. In the first image 300, the object frame 301 occludes the object frame 302. Therefore, in the 3D image 310, the target object 311 is above the target object 312.

[0133] In a possible implementation method, it further includes:

[0134] 203. In response to a movement operation on the target object, change the position of the target object in the 3D scene.

[0135] In this embodiment, the target objects in the generated 3D scene can be flexibly moved. It can be understood that since the 3D scene generation method in the embodiments of the present application generates a 3D scene based on annotation information and a layout image (the first image), in this 3D scene, each target object also carries its own annotation information, that is, the target objects can be independent of each other. Therefore, in response to a movement operation on the target object, the position of the target object in the 3D scene can be changed accordingly.

[0136] The following introduces a training method for a neural network provided by the embodiments of the present application, which is used to train the above target neural network. For ease of understanding, please refer to Figure 4 , Figure 4 which is the method flowchart of the training method for the neural network provided by the embodiments of the present application, and includes:

[0137] 401. Obtain a training 3D scene and annotation information. The training 3D scene includes at least one target object, and the annotation information indicates the category of the target object.

[0138] In this embodiment, the server first obtains training data, which includes a training 3D scene and annotation information. The training data can be used to train the target neural network or to make predictions on the target neural network.

[0139] Training 3D scenes are usually created by professional modeling software or scanning devices, which include various types of target objects. In the case of modeling an indoor scene, the types of target objects include furniture, household appliances, pets, daily necessities, etc.; in the case of modeling a virtual environment, the types of target objects include animals, plants, buildings, vehicles, signs, etc. According to the size, specifications, and types of the required 3D scenes, the target objects included in the 3D scenes also have corresponding differences.

[0140] The annotation information is only the category of the target object, which can be added manually or automatically generated by machine learning algorithms. This application does not limit the method for obtaining the annotation information.

[0141] In the training 3D scene, machine learning algorithms can be used to learn how to identify and classify target objects based on this annotation information. Through repeated training and learning of these scenes, the machine learning algorithms can gradually improve the accuracy of their identification and classification.

[0142] 402. Obtain training images, where the training images include object frames, and the positions of the object frames on the training images correspond to the projection positions of the target objects on the first plane of the training 3D scene.

[0143] In this embodiment, the training images include object frames, and there is a one-to-one correspondence between the object frames and the target objects. The positions of the object frames on the training images correspond to the projection positions of the target objects in the training 3D scene.

[0144] It can be understood that the training images should correspond to a certain plane view or a certain view of the training 3D scene. Therefore, the training images can be manually drawn based on the training 3D scene or generated according to the training 3D scene. For example: after obtaining the training 3D scene and the annotation information, obtain the projection image of the training 3D scene on the first plane; determine the range corresponding to each target object on the projection image according to the annotation information; delimit the object frames according to the range corresponding to each target object, thereby obtaining the training images. It can be understood that there may be an overlap or mutual occlusion of the projections of two target objects on the first plane, then it can be considered that the depths of the two objects on the first plane are different. Therefore, when determining the training images, the object frame corresponding to the target object farther from the first plane will occlude the object frame corresponding to the other target object.

[0145] 403. Process the training images based on the annotation information through the first neural network to obtain a predicted 3D scene including predicted objects.

[0146] It can be understood that this step refers to the process of generating a predicted 3D scene using the first neural network. During the training process, the first neural network is trained to recognize and classify target objects from the training images and transform them into models in the predicted 3D scene, and this process is based on the annotation information. The annotation information is used to indicate the categories of the target objects, and the training images are used to indicate information such as the positions and sizes of the target objects, and this information is used to supervise the training process of the neural network.

[0147] 404, the first neural network is trained according to the first loss function, and the first loss function indicates the similarity between the predicted layout and the training layout. The predicted layout is the layout of the predicted objects in the predicted 3D scene, and the training layout is the layout of the target objects in the training 3D scene.

[0148] In this embodiment, the predicted 3D scene is generated based on the training images, and the expected scene corresponding to the training images is the training 3D scene. It can be understood that the purpose of the first neural network is to generate a 3D scene with the same or similar layout according to the training images, that is, the categories of the corresponding target objects in the predicted 3D scene and the training 3D scene should be the same, but they can have different appearances. For example: the dining table in the training 3D scene is a blue oval dining table, then in the corresponding training image, the dining table can be represented by an object frame of a rectangle, a circle, an ellipse or an irregular figure with the same or approximate planar size, and the object frame is labeled as a dining table; in the predicted 3D scene obtained according to the training image, the corresponding predicted object only needs to be a dining table, and the position where the dining table is located is the same as or similar to that in the training 3D scene.

[0149] The first loss function indicates the similarity between the predicted layout and the training layout. The predicted layout is the layout of the predicted objects in the predicted 3D scene, which is generated from the output result of the first neural network. The training layout is the layout of the target objects in the training 3D scene. It is a known true result. By calculating the similarity between the predicted layout and the training layout, the performance and accuracy of the first neural network can be evaluated. During the training process, the first neural network gradually improves its recognition and classification accuracy of the target objects by learning the annotation information. At the same time, the first loss function adjusts the parameters of the neural network according to the similarity between the predicted layout and the training layout to minimize the value of the loss function. Through repeated training and learning, the first neural network can finally achieve high accuracy and performance.

[0150] In the Figure 4 corresponding method, in an optional embodiment, please refer to Figure 5 , Figure 5 which is the method flow chart of the training method of the neural network provided by the embodiment of the present application,

[0151] 501. Obtain a sample 3D scene and annotation information, where the sample 3D scene includes at least two target objects;

[0152] 502. Perform collision detection on the target objects in the sample 3D scene;

[0153] 503. Adjust the positions of the target objects in the sample 3D scene according to the detection results to obtain a training 3D scene, where at least two target objects in the training 3D scene do not penetrate each other.

[0154] It can be understood that steps 501 to 503 correspond to step 401 in the Figure 4 corresponding embodiment. In this embodiment, data preprocessing is further included, specifically converting the sample 3D scene into a training 3D scene, where any two target objects in the training 3D scene do not penetrate each other, that is, eliminating the penetration phenomenon in the training 3D scene.

[0155] First, obtain a sample 3D scene, which includes at least two target objects. It can be understood that only when there are at least two objects in the scene will there be a penetration phenomenon. Among them, the target objects are not limited to movable objects, but also include immovable objects such as walls and floors. According to the annotation information, the category of each target object in the sample 3D scene can be determined.

[0156] Then perform collision detection on these target objects. Collision detection is an important technology in computer graphics, which is used to detect the overlap or collision between 3D models. By performing collision detection, it can be determined whether there is any overlap or collision between the target objects in the sample 3D scene.

[0157] If there is any collision, adjust the positions of the target objects in the sample 3D scene according to the detection results to eliminate the collision. This process can be achieved by adjusting parameters such as the position and rotation angle of each target object. By adjusting the positions of the target objects, it can be ensured that at least two target objects in the training 3D scene do not penetrate each other, that is, there is no overlap or collision.

[0158] The training 3D scene obtained through the above steps can be used to train the first neural network.

[0159] In a possible implementation method, step 502 includes:

[0160] 5021. Based on the annotation information, perform watertight processing on the target objects in the sample 3D scene to obtain corresponding object meshes;

[0161] 5022. Perform collision detection on the object meshes to obtain detection results, where the detection results include the conflict areas between the object meshes.

[0162] In this embodiment, by performing watertight processing on the target object in the sample 3D scene, the corresponding watertight object mesh is obtained. Watertight processing refers to converting a complex 3D model into a simple mesh model to facilitate collision detection. In this process, the surface of the target object is discretized into small triangles or tetrahedrons, and each triangle or tetrahedron represents a part of the surface of the target object. The watertight object mesh has a strictly defined interior and exterior, and can more accurately represent the shape and boundary of the object. In collision detection, an accurate shape representation can reduce the possibility of false detection and missed detection, and improve the accuracy of detection.

[0163] Then, collision detection is performed based on the watertight object mesh. In this process, each object mesh can be traversed and compared with other object meshes to detect whether there are any conflict regions. A conflict region refers to the overlapping region or contact point between two object meshes. The detection result can be obtained through collision detection, which contains the conflict regions between the object meshes.

[0164] It can be understood that by performing watertight processing on the target object in the sample 3D scene, there are the following advantages: accurately representing the object shape: Watertight processing can represent the 3D model as an accurate mesh model, including the boundaries and details of the object. In contrast, non-watertight models may have shape distortions or inaccuracies, which may lead to errors in collision detection. Through watertight processing, the object shape can be more accurately represented, thereby reducing the possibility of false detection and missed detection; reducing calculation errors: In the process of watertight processing, the complex 3D model is simplified into a simple triangle or tetrahedron mesh. This simplification can reduce the calculation errors generated during collision detection. Since calculation errors may cause small position deviations, these deviations may be eliminated or reduced after watertight processing, thereby improving the accuracy of collision detection.

[0165] As Figure 6a shown, Figure 6a is a flowchart for preprocessing the training 3D scene, including: obtaining the sample 3D scene; classifying the target objects in the sample 3D scene; performing watertight processing on the target objects; performing collision detection to obtain the training 3D scene (as shown in the 3D diagram on the right in Figure 6a ).

[0166] 504, according to the annotation information, extract the second latent vector of the training 3D scene;

[0167] In this embodiment, considering that training a 3D scene belongs to three-dimensional data and the training 3D scene needs to be combined with annotation information, the data type is relatively complex and it is difficult for a neural network model to understand. Therefore, a latent vector is introduced in the embodiments of the present application. The latent vector can reduce high-dimensional data to low-dimensional vectors, thereby reducing the complexity and dimension of the data and facilitating further data processing and analysis. At the same time, the latent vector can extract the main features and structures of the data, thus facilitating machine learning tasks such as classification, clustering, and recognition. For the relevant description of the latent vector, please refer to Figure 2 the relevant description corresponding to step 202 in the corresponding embodiment, which will not be elaborated here.

[0168] It can be understood that after extracting the second latent vector of the training 3D scene according to the annotation information, and Figure 4 in the steps corresponding to step 404 in the corresponding embodiment, it is also necessary to extract the first latent vector of the training image, and train the neural network model through the correlation between the latent vectors. For the specific description, please refer to the relevant descriptions of steps 509 and 510 in the following text.

[0169] In a possible implementation method, before step 504, it further includes:

[0170] 505, encoding the training 3D scene into a first parameter through a second neural network, where the first parameter is a tri-planar parameter;

[0171] 506, processing the first parameter based on the annotation information through a second neural network to obtain the target SDF corresponding to each target object;

[0172] 507, training the second neural network according to a second loss function, where the second loss function indicates the similarity between the target SDF and the training SDF, and the training SDF is obtained by measuring the target object in the training 3D scene;

[0173] At this time, step 504 specifically includes:

[0174] 5041, encoding the training 3D scene into a second parameter according to the target SDF, where the second parameter is a tri-planar parameter;

[0175] 5042, extracting the second latent vector of the second parameter.

[0176] In this embodiment, three-plane parameters are introduced to obtain the second latent vector. The three-plane parameters are a type of parameter that describes the position, shape, and orientation of a plane. It consists of three parametric equations, which respectively describe the relationships between the x, y, and z coordinates on the plane and the three parameters. The parametric equations can effectively describe the position, shape, and orientation of the plane. At the same time, the 3D geometry is represented by a signed distance function (SDF), and the three-plane parameters of the training 3D scene are trained with this, so that the three-plane parameters carry the category features of the target object.

[0177] First, the training 3D scene is encoded into the first parameter by the second neural network, where the first parameter is the three-plane parameter, which compresses and reduces the dimension of the high-dimensional training 3D scene, facilitating further data processing and analysis. For each object category in the training 3D scene, the features of the first parameter can be converted into multiple signed distance functions (SDFs) through a multilayer perceptron (MLP), where each SDF corresponds to an object of a category in the training 3D scene.

[0178] The SDF represents determining the distance from a point to the boundary of a finite region in space and simultaneously defining the sign of the distance: the point is positive inside the region boundary, negative outside, and 0 when located on the boundary. That is, in the training 3D scene, for a point on the edge of the target object, its SDF should be 0. Then, according to the characteristic that the SDF of the points on the boundary of the target object in the training 3D scene is 0, the first parameter is trained to obtain the second parameter, which also is the three-plane parameter and includes the category of the target object. Finally, the second latent vector is extracted through the second parameter.

[0179] As Figure 6b shown, Figure 6b is the three-dimensional representation flowchart of the three-plane parameter.

[0180] Figure 6b On the left side of [figure], the training 3D scene is encoded to obtain the first parameter (three-plane parameter), and then the features of the first parameter are converted into multiple target SDFs through the MLP. Each target SDF corresponds to an object of a category in the training 3D scene (including sofas, chairs, cabinets, beds, walls, etc.). After training the second neural network through the second loss function, the target SDFs are optimized, and thus the second parameter (three-plane parameter) is generated according to the target SDFs. The SDFs in the 3D scene corresponding to the second parameter correspond to the training SDFs in the training 3D scene.

[0181] More specifically, the tri-plane uses three 2D feature planes fxy, fxz, fyz ∈ RN×N×C, where the spatial resolution of each channel is N×N and the dimension of the feature channels is C. An MLP is used to decode the features sampled from the planes to output the SDF values for each category. The 3D coordinates p are queried by projecting them onto each axis-aligned plane (i.e., the x-y, x-z, and y-z planes). For the queried features (fxy(p), fxz(p), fyz(p)), they are added together, and the result is decoded using a lightweight MLP to obtain the value S(p) ∈ RM of point p. Here, M is the number of object categories. The specific formula is:

[0182] S(p) = MLP(fxy(p) + fxz(p) + fyz(p))

[0183] The tri-plane and the MLP can be jointly optimized. The optimization objective function includes the following losses:

[0184] 1) Object surface loss:

[0185] According to the definition of SDF, the SDF value of the points sampled on the surface should be 0. Here, the L1 norm is used:

[0186] Lsurf = λsurfΣi∈{1,…,M}Σp∈Ω||S(p)i||1

[0187] where S(p)i is the numerical value of the SDF corresponding to the i-th object category.

[0188] 2) Minimum SDF loss: Let min(S(p)) represent the distance from the points in space to the nearest object. For this value, the L1 loss is used:

[0189] Lsdf = λsdfΣp∈Θ||min(S(p)) - min(S(p)GT)||1

[0190] 3) Surface normal loss: The L2 norm of the surface normal of the points sampled on the object surface:

[0191] Lnormal = λnormalΣp∈Ω||Δ(p) - Δ(p)GT||2

[0192] 4) Distance field derivative loss: According to the definition of the signed distance field (SDF), the derivative of the SDF with respect to point P should be 1 everywhere. Here, the L2 norm is used:

[0193] Lgradient = λgradientΣp∈{Ω,Θ}||Δ(p) - 1||2

[0194] Where Ω is the set of points randomly sampled on the surface of the geometry. Θ is the set of points randomly sampled in space. λgradient, λnormal, λsdf, and λsurf are the weights for each loss. Δ(p) is the derivative of the SDF value of point p with respect to that point, i.e., the normal vector of that point. The GT flag is the ground truth value of the corresponding numerical value. The ground truth data can be directly calculated from the preprocessed object mesh. For all 3D scenes in the training data, its tri-plane parameters (the second parameter) can be optimized in the above manner.

[0195] 508. Obtain a training image, where the training image includes an object bounding box, and the position of the object bounding box on the training image corresponds to the projection position of the target object on the first plane of the training 3D scene;

[0196] It can be understood that step 508 is similar to step 402 in the Figure 4 corresponding embodiment, and details will not be elaborated here.

[0197] 509. Through a first neural network, extract a first latent vector of the training image according to the annotation information;

[0198] 510. Through the first neural network, process the first latent vector to obtain a predicted 3D scene.

[0199] It can be understood that step 509 and step 510 correspond to Figure 4 step 403 in the corresponding embodiment. In this embodiment, considering that the training image belongs to two-dimensional data, when the image data carries annotation information, the combined data is more complex and difficult for the neural network model to understand. Therefore, the present application embodiment introduces a latent vector. The latent vector can reduce the high-dimensional data to a low-dimensional vector, thereby reducing the complexity and dimension of the data and facilitating further data processing and analysis; at the same time, the latent vector can extract the main features and structures of the data, thereby facilitating machine learning tasks such as classification, clustering, and recognition.

[0200] At this time, corresponding to Figure 4 step 404 in the corresponding embodiment, it specifically includes:

[0201] 511. Train the first neural network according to a first loss function, where the first loss function indicates the similarity between the first latent vector and the second latent vector.

[0202] It can be understood that after converting both the 3D scene and the 2D image into latent vectors, the first loss function specifically indicates the similarity between the first latent vector and the second latent vector. The higher the similarity between the first latent vector and the second latent vector, the higher the similarity between the predicted layout and the training layout.

[0203] Please refer to Figure 7 , Figure 7 which is a schematic flowchart of the training method of the neural network provided by the embodiment of the present application, including:

[0204] Step 1, use a variational autoencoder (VAE) to perform auto-encoding on the tri-plane (equivalent to steps 505 to 507) to obtain a second latent vector (equivalent to step 504) of the tri-plane for predicting the 3D scene (the left side of step 1 in Figure 7 ): After the training of step 1 is completed, convert the tri-plane in the training data into a latent vector, which is used as the ground truth generated by the diffusion model for generating the tri-plane of the predicted 3D model based on the training images (the right side of step 1 in Figure 7 ).

[0205] Step 2, use the diffusion model to generate a first latent vector from random noise: At the same time as step 1, the 2D layout encoder can also be used to encode the 2D layout map (training image or test image) input by the user, and then use the Transformer to input the encoded data into the diffusion process of the diffusion model, so as to achieve the effect of using the 2D layout to control 3D geometry generation.

[0206] In a possible implementation method, it further includes:

[0207] 512. Through the first neural network, establish a corresponding 3D module for the predicted object in the predicted 3D scene;

[0208] 513. In response to the movement operation of the predicted object, change the position of the predicted object in the predicted 3D scene.

[0209] In this embodiment, after generating the predicted 3D scene, it further includes establishing a corresponding 3D module for the predicted object in the predicted 3D scene, so that each predicted object is independent and can be flexibly moved in the predicted 3D scene. It can be understood that since the 3D scene generation method in the embodiment of the present application generates a 3D scene based on the annotation information and the layout image (the first image), each target object in this 3D scene also carries its own annotation information, that is, each target object can be independent, so it can respond to the movement operation for the target object and correspondingly change the position of the target object in the 3D scene.

[0210] The training method of the neural network provided by the embodiments of the present application realizes the function of generating a three-dimensional indoor semantic scene by the user through a two-dimensional floor plan of the scene. Compared with the existing solutions, the geometry output by this solution has built-in semantic information and the module meshes (Meshes) of objects of different categories are independent, which is more convenient for user interaction. The method provided by the embodiments of the present application can give birth to new and diverse extended reality (XR) applications. XR covers interactive experiences including real and virtual, such as augmented reality (AR), virtual reality (VR), mixed reality (MR), etc., and can be used to create large-scale user-defined 3D scene content, with broad application prospects, such as being applied to the construction of the metaverse. In addition, this method can also be used to generate training scenes required for AI robot navigation to help the robot better understand and adapt to the environment. At the same time, the present application can also be widely applied in the game industry. For example, it can be used as a user-generated content (UGC) module in games to assist users in generating content efficiently; it can be used as a map editor in open-world games to help players or game developers create game maps; it can also be used as a development assistance tool for games to help developers quickly generate game prototypes, test scenes, characters, etc., and improve the efficiency of game design and development.

[0211] The 3D scene generation device in the present application will be described in detail below. Please refer to Figure 8 。 Figure 8 FIG. 800 is a schematic diagram of an embodiment of a 3D scene generation device 800 in an embodiment of the present application. The 3D scene generation device 800 includes:

[0212] A first acquisition unit 801, configured to acquire a first image and annotation information. The first image includes at least one object box, and the annotation information indicates the category of the target object corresponding to the object box.

[0213] A first modeling unit 802, configured to process the first image based on the annotation information through a target neural network to obtain a three-dimensional (3D) scene including the target object, and the projection position of the target object on the first plane of the 3D scene corresponds to the position of the object box on the first image.

[0214] The 3D scene generation device provided by the embodiments of the present application obtains a first image and annotation information. The first image includes at least one object box, and the annotation information indicates the category of the target object corresponding to the object box. Through a target neural network, the first image is processed based on the annotation information to obtain a three-dimensional (3D) scene including the target object. The projection position of the target object on the first plane of the 3D scene corresponds to the position of the object box on the first image, realizing the function of generating a 3D scene through a two-dimensional plan view and semantics. The generated 3D scene meets the preset layout requirements, which is beneficial to improving the design and development efficiency of the 3D scene.

[0215] In a possible implementation method,

[0216] The first modeling unit 802 is specifically configured to extract a latent vector of the first image based on the annotation information through a target neural network, and process the latent vector through the target neural network to obtain a 3D scene.

[0217] In this embodiment, considering that image data belongs to two-dimensional data, when the image data carries annotation information, the combined data is more complex and difficult for the neural network model to understand. Therefore, the latent vector is introduced in the embodiments of the present application. The latent vector can reduce the high-dimensional data to a low-dimensional vector, thereby reducing the complexity and dimension of the data and facilitating further data processing and analysis. At the same time, the latent vector can extract the main features and structures of the data, thus facilitating machine learning tasks such as classification, clustering, and recognition.

[0218] In a possible implementation method, the number of object boxes is at least two, including a first object box and a second object box. The first object box corresponds to a first target object, and the second object box corresponds to a second target object.

[0219] When the first object box occludes the second object box, in the 3D scene, the first target object is located on the other side of the second target object relative to the first plane.

[0220] In this embodiment, the first image includes object boxes that occlude each other. At this time, the position depth of the target objects corresponding to the two object boxes can be determined according to the occlusion relationship. That is, when the first object box occludes the second object box, it means that in the first image, the layer of the first object is above the layer of the second object. Then, in the corresponding generated 3D image, when the perspective of the 3D image corresponds to the first image, the first target object is closer to the perspective direction than the second target object, that is, the first target object is located on the other side of the second target object relative to the first plane.

[0221] In a possible implementation method, it further includes:

[0222] The first moving unit 803 is configured to change the position of the target object in the 3D scene in response to a moving operation on the target object.

[0223] In this embodiment, the target object in the generated 3D scene can be flexibly moved. It can be understood that since the 3D scene generation method in the embodiments of the present application generates the 3D scene based on the annotation information and the layout image (the first image), in this 3D scene, each target object also carries its own annotation information, that is, the target objects can be independent of each other. Therefore, in response to a moving operation on the target object, the position of the target object in the 3D scene can be changed accordingly.

[0224] In a possible implementation, the object box is a rectangular box, and the projection range of the target object on the first plane corresponds to the size range of the rectangular box.

[0225] In this embodiment, the size range corresponds to the projection range, which is used to represent the size of the target object in the 3D scene, rather than limiting the size of the target object. The corresponding relationship between the above size range and the projection range may mean that the projection range and the size range meet the preset requirements. For example, if the similarity between the projection range and the size range is higher than the preset threshold, it means that the projection range and the size range correspond.

[0226] The training device of the neural network in the present application will be described in detail below. Please refer to Figure 9 . Figure 9 FIG. is a schematic diagram of an embodiment of the training device 900 of the neural network in the embodiments of the present application. The training device 900 of the neural network includes:

[0227] The second obtaining unit 901 is configured to obtain a training 3D scene and annotation information. The training 3D scene includes at least one target object, and the annotation information indicates the category of the target object;

[0228] The second obtaining unit 901 is further configured to obtain a training image, where the training image includes an object box, and the position of the object box on the training image corresponds to the projection position of the target object on the first plane of the training 3D scene;

[0229] The second modeling unit 902 is configured to process the training image based on the annotation information through a first neural network to obtain a predicted 3D scene including predicted objects;

[0230] The training unit 903 is configured to train the first neural network according to a first loss function. The first loss function indicates the similarity between the predicted layout and the training layout. The predicted layout is the layout of the predicted objects in the predicted 3D scene, and the training layout is the layout of the target objects in the training 3D scene.

[0231] In a possible implementation method,

[0232] A first extraction unit, configured to extract a second latent vector for training a 3D scene according to the annotation information;

[0233] A second modeling unit 902, specifically configured to extract a first latent vector of a training image according to the annotation information through a first neural network; process the first latent vector through the first neural network to obtain a predicted 3D scene; a first loss function indicates the similarity between the first latent vector and the second latent vector.

[0234] In a possible implementation method, it further includes:

[0235] A second parameter tuning unit, configured to encode a training 3D scene into a first parameter through a second neural network, the first parameter being a trilinear parameter; process the first parameter based on the annotation information through the second neural network to obtain a target signed distance function SDF corresponding to each target object;

[0236] A training unit 903 is further configured to train the second neural network according to a second loss function, the second loss function indicating the similarity between the target SDF and the training SDF, and the training SDF is obtained by measuring the target objects in the training 3D scene;

[0237] The first extraction unit is specifically configured to encode the training 3D scene into a second parameter according to the target SDF, the second parameter being a trilinear parameter; extract a second latent vector of the second parameter.

[0238] In a possible implementation method, the number of target objects in the training 3D scene is at least two;

[0239] A second acquisition unit 901 is specifically configured to acquire a sample 3D scene and annotation information, where the sample 3D scene includes at least two target objects; perform collision detection on the target objects in the sample 3D scene; adjust the positions of the target objects in the sample 3D scene according to the detection results to obtain a training 3D scene, and the at least two target objects in the training 3D scene do not penetrate each other.

[0240] In a possible implementation method, performing collision detection on the target objects in the sample 3D scene includes:

[0241] Based on the annotation information, perform watertight processing on the target objects in the sample 3D scene to obtain corresponding object meshes;

[0242] Perform collision detection on the object meshes to obtain a detection result, and the detection result includes a conflict area between the object meshes.

[0243] In a possible implementation method,

[0244] The second modeling unit 902 is further configured to establish a corresponding 3D module for the predicted object in the predicted 3D scene through the first neural network;

[0245] It further includes: a second moving unit 904, configured to change the position of the predicted object in the predicted 3D scene in response to a moving operation on the predicted object.

[0246] The training device of the neural network provided by the embodiments of the present application corresponds to Figure 4 and Figure 5 the training method of the neural network in the corresponding embodiment. For related descriptions, please refer to the above text and will not be elaborated here.

[0247] Figure 10 FIG. is a schematic structural diagram of a server provided by an embodiment of the present application. The server 300 may vary greatly due to different configurations or performances, and may include one or more central processing units (CPUs) 322 (for example, one or more processors) and a memory 332, and one or more storage media 330 (for example, one or more mass storage devices) for storing application programs 342 or data 344. Among them, the memory 332 and the storage media 330 may be transient storage or persistent storage. The program stored in the storage media 330 may include one or more modules (not shown in the figure), and each module may include a series of instruction operations on the server. Further, the central processing unit 322 may be configured to communicate with the storage media 330 and execute a series of instruction operations in the storage media 330 on the server 300.

[0248] The server 300 may further include one or more power supplies 326, one or more wired or wireless network interfaces 350, one or more input / output interfaces 358, and / or one or more operating systems 341, such as Windows ServerTM, Mac OS XTM, UnixTM, LinuxTM, FreeBSDTM, etc.

[0249] The steps performed by the server in the above embodiments may be based on the Figure 10 server structure shown.

[0250] Those skilled in the art can clearly understand that for the convenience and simplicity of description, the specific working processes of the systems, devices, and units described above may refer to the corresponding processes in the foregoing method embodiments and will not be elaborated here.

[0251] In the embodiments of the present application, the term "module" or "unit" refers to a computer program with a predetermined function or a part of a computer program, which works together with other related parts to achieve a predetermined goal, and can be fully or partially implemented by using software, hardware (such as a processing circuit or a memory), or a combination thereof. Similarly, one processor (or multiple processors or memories) can be used to implement one or more modules or units. In addition, each module or unit can be a part of an overall module or unit that includes the function of that module or unit.

[0252] In several embodiments provided by the present application, it should be understood that the disclosed systems, devices, and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of units is only a logical function division, and there may be other division methods in actual implementation. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed coupling or direct coupling or communication connection between each other can be through some interfaces, and the indirect coupling or communication connection of devices or units can be in electrical, mechanical, or other forms.

[0253] The units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they can be located in one place, or can be distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0254] In addition, the functional units in each embodiment of the present application can be integrated in one processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit. The above-mentioned integrated units can be implemented in the form of hardware or in the form of software functional units.

[0255] When the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods in various embodiments of this application. The foregoing storage medium includes: various media that can store program codes, such as USB flash drives, mobile hard disks, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical discs.

[0256] In the above, the above embodiments are only used to illustrate the technical solution of this application, rather than to limit it; although this application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements on some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of various embodiments of this application.

Claims

1. A 3D scene generation method, characterized in that, Including: Obtain a first image and annotation information, where the first image includes at least one object bounding box, and the annotation information indicates the category of the target object corresponding to the object bounding box; Process the first image based on the annotation information through a target neural network to obtain a three-dimensional (3D) scene including the target object, and the projection position of the target object on a first plane of the 3D scene corresponds to the position of the object bounding box on the first image.

2. The method according to claim 1, wherein The process of processing the first image based on the annotation information through the target neural network to obtain a 3D scene including the target object includes: Extract a latent vector of the first image based on the annotation information through the target neural network; Process the latent vector through the target neural network to obtain the 3D scene.

3. The method according to claim 1, characterized in that The number of the object bounding boxes is at least two, including a first object bounding box and a second object bounding box, the first object bounding box corresponds to a first target object, and the second object bounding box corresponds to a second target object; When the first object bounding box occludes the second object bounding box, in the 3D scene, the first target object is located on the other side of the second target object relative to the first plane.

4. The method according to claim 1, wherein After processing the first image based on the annotation information through the target neural network to obtain a 3D scene including the target object, it further includes: In response to a movement operation on the target object, change the position of the target object in the 3D scene.

5. The method according to claim 1, wherein The object bounding box is a rectangular box, and the projection range of the target object on the first plane corresponds to the size range of the rectangular box.

6. A training method for a neural network, characterized in that, Including: Obtain a training 3D scene and annotation information, where the training 3D scene includes at least one target object, and the annotation information indicates the category of the target object; Obtain a training image, where the training image includes an object bounding box, and the position of the object bounding box on the training image corresponds to the projection position of the target object on a first plane of the training 3D scene; Process the training image based on the annotation information through a first neural network to obtain a predicted 3D scene including predicted objects; Train the first neural network according to a first loss function, where the first loss function indicates the similarity between a predicted layout and a training layout, the predicted layout is the layout of the predicted objects in the predicted 3D scene, and the training layout is the layout of the target objects in the training 3D scene.

7. The method according to claim 6, characterized in that, After obtaining the training 3D scene and annotation information, it further includes: Extract a second latent vector of the training 3D scene according to the annotation information; The process of processing the training image based on the annotation information through the first neural network to obtain a predicted 3D scene including predicted objects includes: Extract a first latent vector of the training image according to the annotation information through the first neural network; Process the first latent vector through the first neural network to obtain the predicted 3D scene; The first loss function indicates the similarity between the first latent vector and the second latent vector.

8. The method according to claim 7, characterized in that, Before extracting the second latent vector of the training 3D scene according to the annotation information, the following steps are further included: Encoding the training 3D scene into a first parameter through a second neural network, where the first parameter is a tri-planar parameter; Processing the first parameter based on the annotation information through a second neural network to obtain a target signed distance function (SDF) corresponding to each target object; Training the second neural network according to a second loss function, where the second loss function indicates the similarity between the target SDF and the training SDF, and the training SDF is obtained by measuring the target objects in the training 3D scene; The step of extracting the second latent vector of the training 3D scene according to the annotation information includes: Encoding the training 3D scene into a second parameter according to the target SDF, where the second parameter is a tri-planar parameter; Extracting the second latent vector of the second parameter.

9. The method according to claim 6, wherein The number of target objects in the training 3D scene is at least two; The step of obtaining the training 3D scene and the annotation information includes: Obtaining a sample 3D scene and the annotation information, where the sample 3D scene includes the at least two target objects; Performing collision detection on the target objects in the sample 3D scene; Adjusting the positions of the target objects in the sample 3D scene according to the detection results to obtain the training 3D scene, where the at least two target objects in the training 3D scene do not penetrate each other.

10. The method according to claim 9, wherein The step of performing collision detection on the target objects in the sample 3D scene includes: Based on the annotation information, performing watertight processing on the target objects in the sample 3D scene to obtain corresponding object meshes; Performing collision detection on the object meshes to obtain the detection results, where the detection results include the conflict regions between the object meshes.

11. The method according to claim 6, wherein After processing the training image based on the annotation information through a first neural network to obtain a predicted 3D scene including predicted objects, the following steps are further included: Establishing corresponding 3D modules for the predicted objects in the predicted 3D scene through a first neural network; In response to a movement operation on the predicted objects, changing the positions of the predicted objects in the predicted 3D scene.

12. A 3D scene generation device, characterized in that, It includes: A first acquisition unit for acquiring a first image and annotation information, where the first image includes at least one object bounding box, and the annotation information indicates the category of the target object corresponding to the object bounding box; A first modeling unit for processing the first image based on the annotation information through a target neural network to obtain a three-dimensional (3D) scene including the target objects, where the projection position of the target objects on the first plane of the 3D scene corresponds to the position of the object bounding box on the first image.

13. A training device for a neural network, characterized in that, It includes: A second acquisition unit for acquiring a training 3D scene and annotation information, where the training 3D scene includes at least one target object, and the annotation information indicates the category of the target object; The second acquisition unit is further configured to acquire a training image, where the training image includes an object box, and the position of the object box on the training image corresponds to the projection position of the target object on the first plane of the training 3D scene; The second modeling unit is configured to process the training image based on the annotation information through a first neural network to obtain a predicted 3D scene including predicted objects; The training unit is configured to train the first neural network according to a first loss function, where the first loss function indicates the similarity between the predicted layout and the training layout, the predicted layout is the layout of the predicted objects in the predicted 3D scene, and the training layout is the layout of the target objects in the training 3D scene.

14. A computer device, characterized in that, Comprising: A memory, a transceiver, a processor, and a bus system; Wherein, the memory is used to store programs; The processor is configured to execute the programs in the memory, including executing the 3D scene generation method according to any one of claims 1 to 5, or executing the training method of the neural network according to any one of claims 6 to 11; The bus system is configured to connect the memory and the processor to enable communication between the memory and the processor.

15. A computer-readable storage medium, including instructions, which when running on a computer, cause the computer to execute the 3D scene generation method according to any one of claims 1 to 5, or execute the training method of the neural network according to any one of claims 6 to 11.