Three-dimensional scene construction method, device, equipment and storage medium

By combining image data and point cloud data, the object category and bounding box are determined, and the 3D model is generated and adjusted, which solves the problems of 3D scene construction speed and accuracy in existing technologies and realizes efficient and accurate 3D scene construction and editing.

CN120431296BActive Publication Date: 2025-09-09SHENZHEN XGRIDS-INNOVATION CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510938965.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-07-08
Publication Date
2025-09-09
Estimated Expiration
2045-07-08

AI Technical Summary

Technical Problem

Existing technologies make it difficult to quickly and accurately construct three-dimensional scenes, and are unable to edit three-dimensional scenes based on object category information.

Method used

By acquiring image data and point cloud data of real 3D scenes, target detection is performed, object categories and 3D bounding boxes are determined, a 3D model is generated, and the model size and position are adjusted according to the bounding box information to construct a virtual 3D scene.

Benefits of technology

It improves the efficiency and accuracy of 3D scene construction, ensures the consistency of object categories, and supports subsequent scene editing.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120431296B_ABST
    Figure CN120431296B_ABST
Patent Text Reader

Abstract

This application relates to the field of computer vision technology and discloses a method, apparatus, device, and storage medium for constructing a three-dimensional scene. The method comprises: acquiring data of a real three-dimensional scene to obtain image data and point cloud data; performing target detection based on the image data and point cloud data to determine the object category and three-dimensional bounding box of each object in the real three-dimensional scene; for each object, determining target image data including the object from the image data; generating a first three-dimensional model of the object based on the object category and target image data of the object; adjusting the size of the first three-dimensional model of the object based on the size information carried by the three-dimensional bounding box of each object to obtain a second three-dimensional model of the object; and constructing a virtual three-dimensional scene corresponding to the real three-dimensional scene based on the position information and orientation angle information carried by the three-dimensional bounding box of each object and the second three-dimensional model of the object. This application enables the rapid and accurate construction of three-dimensional scenes.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The embodiments of the present application relate to the field of computer vision technology, and specifically to a three-dimensional scene construction method, apparatus, device, and storage medium. Background Art

[0002] Currently, the main methods for building virtual three-dimensional scenes include manual modeling and image sensor-based reconstruction technology, such as reconstructing three-dimensional scenes through photogrammetry and laser scanning. Although manual modeling can highly customize three-dimensional scenes, it is time-consuming and costly, and it is difficult to quickly generate large-scale and complex three-dimensional scenes. Although image sensor-based reconstruction technology can capture the geometry and texture information of objects from the real world and then construct a three-dimensional scene, this method is difficult to ensure that the categories of each object in the constructed virtual three-dimensional scene are consistent with the categories of objects in the real three-dimensional scene when constructing a complex three-dimensional scene, resulting in the constructed virtual three-dimensional scene not matching the real three-dimensional scene.

[0003] Furthermore, since the virtual 3D scene constructed in the above manner does not include object category information, the object category of each object in the virtual 3D scene cannot be determined, resulting in the inability to subsequently edit the constructed virtual 3D scene. Summary of the Invention

[0004] In view of the above problems, the embodiments of the present application provide a three-dimensional scene construction method, device, equipment and storage medium, which are used to solve the problems in the prior art that a three-dimensional scene cannot be constructed quickly and accurately, and that a three-dimensional scene cannot be edited according to the object category information of objects in the three-dimensional scene.

[0005] According to one aspect of an embodiment of the present application, a three-dimensional scene construction method is provided, the method comprising: acquiring image data and point cloud data of a real three-dimensional scene; performing target detection based on the image data and the point cloud data, and determining the object category and three-dimensional bounding box of each object in the real three-dimensional scene, wherein the three-dimensional bounding box carries position information, size information, and orientation angle information of the bounding box; for each object, determining target image data including the object from the image data; generating a first three-dimensional model of the object based on the object category and the target image data of the each object; adjusting the size of the first three-dimensional model of the object based on the size information carried by the three-dimensional bounding box of the each object to obtain a second three-dimensional model of the object; and constructing a virtual three-dimensional scene corresponding to the real three-dimensional scene based on the position information, the orientation angle information, and the second three-dimensional model of the object carried by the three-dimensional bounding box of the each object.

[0006] In an optional manner, the target detection is performed based on the image data and the point cloud data to determine the object category and the three-dimensional bounding box of each object in the real three-dimensional scene, including: inputting the image data and the point cloud data into a dual-stream feature extraction module to obtain a semantic feature vector and a three-dimensional geometric feature vector output by the dual-stream feature extraction module, wherein the dual-stream feature extraction module includes a two-dimensional semantic feature extraction branch and a three-dimensional geometric feature extraction branch, the two-dimensional semantic feature extraction branch is used to receive the image data and perform feature extraction on the image data to obtain the semantic feature vector, and the three-dimensional geometric feature extraction branch is used to extract the feature of the image data to obtain the semantic feature vector. The branch is used to receive the point cloud data and perform feature extraction on the point cloud data to obtain the three-dimensional geometric feature vector; input the semantic feature vector and the three-dimensional geometric feature vector into a feature fusion module to obtain the fusion feature output by the feature fusion module; input the fusion feature into a depth fusion module to obtain the depth fusion feature output by the depth fusion module; input the depth fusion feature into a prediction head module including a parallel object classification head and a three-dimensional bounding box regression head to obtain the object category of each object output by the object classification head, and the three-dimensional bounding box of each object output by the three-dimensional bounding box regression head.

[0007] In an optional manner, for each object, determining the target image data including the object from the image data includes: for each object, selecting the first image data including the object from the image data; for the object, if there is only one first image data, determining the first image data as the target image data; if there are multiple first image data, scoring the multiple first image data, and determining the target image data according to the score of each first image data.

[0008] In an optional manner, determining the target image data based on the score of each piece of the first image data includes: if there is second image data with a score greater than or equal to a preset score threshold among the multiple pieces of the first image data, determining the second image data as the target image data; if the scores of the multiple pieces of the first image data are all less than the preset score threshold, determining the first image data with the highest score among the multiple pieces of the first image data as the target image data.

[0009] In an optional manner, generating a first three-dimensional model of the object based on the object category and the target image data of each object includes: for each object, if there is only one target image data, generating the first three-dimensional model based on the object category and the target image data of the object through a deep learning model; if there are multiple target image data, generating the first three-dimensional model based on the object category and the multiple target image data of the object through a multi-viewpoint stereo vision method.

[0010] In an optional manner, constructing a virtual three-dimensional scene corresponding to the real three-dimensional scene based on the position information, the orientation angle information and the second three-dimensional model of the object carried by the three-dimensional bounding box of each object includes: initially aligning the second three-dimensional model according to the position information and the orientation angle information carried by the three-dimensional bounding box to obtain a first three-dimensional scene constructed by the second three-dimensional model, wherein the position and orientation angle of the second three-dimensional model in the first three-dimensional scene respectively match the position information and the orientation angle information; accurately aligning the second three-dimensional model in the first three-dimensional scene through an iterative nearest point algorithm to obtain a second three-dimensional scene; and performing scene optimization processing on the second three-dimensional scene to obtain the virtual three-dimensional scene.

[0011] In an optional embodiment, after constructing a virtual three-dimensional scene corresponding to the real three-dimensional scene based on the position information, the orientation angle information and the second three-dimensional model of the object carried by the three-dimensional bounding box of each object, the method further includes: determining the mold-penetrating area between different second three-dimensional models in the virtual three-dimensional scene; determining the target second three-dimensional model containing the mold-penetrating area based on the mold-penetrating area; processing the target second three-dimensional model by at least one of adjusting transparency, moving, rotating and adjusting size, and / or receiving an operation performed by a user on the target second three-dimensional model, and processing the target second three-dimensional model based on the operation until there is no mold-penetrating area between different second three-dimensional models in the virtual three-dimensional scene.

[0012] According to another aspect of an embodiment of the present application, a three-dimensional scene construction device is provided, the device comprising: a data acquisition module for acquiring image data and point cloud data of a real three-dimensional scene; a target detection module for performing target detection based on the image data and the point cloud data, and determining the object category and three-dimensional bounding box of each object in the real three-dimensional scene, wherein the three-dimensional bounding box carries position information, size information and orientation angle information of the bounding box; a determination module for determining, for each object, target image data including the object from the image data; a three-dimensional model generation module for generating a first three-dimensional model of the object based on the object category and the target image data of the object; a three-dimensional model adjustment module for adjusting the size of the first three-dimensional model of the object based on the size information carried by the three-dimensional bounding box of the object to obtain a second three-dimensional model of the object; and a three-dimensional scene construction module for constructing a virtual three-dimensional scene corresponding to the real three-dimensional scene based on the position information, the orientation angle information and the second three-dimensional model of the object carried by the three-dimensional bounding box of the object.

[0013] According to another aspect of an embodiment of the present application, a three-dimensional scene construction device is provided, including a memory, a processor, and a computer program stored in the memory, wherein the processor executes the computer program to implement the three-dimensional scene construction method as described above.

[0014] According to another aspect of the embodiments of the present application, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the three-dimensional scene construction method as described above is implemented.

[0015] In an embodiment of the present application, the category and number of objects in the real three-dimensional scene are accurately determined through image data of the real three-dimensional scene, and the position and size of each object are determined through point cloud data. Then, a first three-dimensional model that is relatively similar to the shape of the object is generated based on the object category of each object and the target image data including the object, and a virtual three-dimensional scene is constructed based on the size, position information and the first three-dimensional model of the object, thereby improving the efficiency and accuracy of constructing the three-dimensional scene.

[0016] Moreover, since each first three-dimensional model is determined based on the object category of each object, and the second three-dimensional model is obtained after adjusting the size of the first three-dimensional model, for the virtual three-dimensional scene constructed in the embodiment of the present application, the object category of the object corresponding to each second three-dimensional model in the virtual three-dimensional scene can be determined, so that the virtual three-dimensional scene can be edited based on the object category of the object corresponding to each second three-dimensional model.

[0017] The above description is only an overview of the technical solutions of the embodiments of the present application. In order to more clearly understand the technical means of the embodiments of the present application, they can be implemented in accordance with the contents of the specification. In order to make the above and other purposes, features and advantages of the embodiments of the present application more obvious and easy to understand, the specific implementation methods of the present application are listed below. BRIEF DESCRIPTION OF THE DRAWINGS

[0018] The accompanying drawings are only used to illustrate the embodiments and are not to be considered as limiting the present application. In addition, the same reference symbols are used to represent the same components throughout the drawings. In the drawings:

[0019] Figure 1 A schematic diagram of a process for constructing a three-dimensional scene according to an embodiment of the present invention is shown;

[0020] Figure 2 Shown Figure 1 Schematic diagram of the sub-step flow of step 120;

[0021] Figure 3 A schematic diagram illustrating a method for determining the object category and three-dimensional bounding box of each object in a real three-dimensional scene according to an embodiment of the present application is shown;

[0022] Figure 4 Shown Figure 1 Schematic diagram of the sub-step flow chart of step 130;

[0023] Figure 5 A schematic diagram of the structure of a three-dimensional scene construction device provided in an embodiment of the present application is shown;

[0024] Figure 6 A schematic structural diagram of a three-dimensional scene construction device provided in an embodiment of the present application is shown. DETAILED DESCRIPTION

[0025] The exemplary embodiments of the present application will be described in more detail below with reference to the accompanying drawings. Although exemplary embodiments of the present application are shown in the drawings, it should be understood that the present application can be implemented in various forms and should not be limited to the embodiments set forth herein.

[0026] Virtual 3D scenes are playing an increasingly important role in fields such as Virtual Reality (VR), Augmented Reality (AR), and Mixed Reality (MR). Their application has expanded to industries such as game development, urban planning, industrial design, education and training, and healthcare. By providing an immersive experience, virtual environments significantly enhance the user's interactivity with digital content and the depth of their perception. Virtual 3D scene construction technology based on real-world perception aims to leverage real-world perception and data capture to create more realistic and immersive virtual experiences, thereby bridging the gap between the digital and physical worlds.

[0027] Currently, the main methods for constructing virtual 3D scenes include manual modeling and image sensor-based reconstruction techniques, such as photogrammetry and laser scanning. While manual modeling allows for precise control over every detail of a scene and allows for the construction of a virtual 3D scene that perfectly mirrors the real 3D scene, it is time-consuming and costly, making it difficult to rapidly generate large and complex 3D scenes (e.g., large-scale virtual worlds, game development, etc.). Image sensor-based reconstruction techniques can capture the geometry and texture information of real-world objects to construct 3D scenes, but when constructing complex 3D scenes, this approach may not be able to fully and accurately identify and distinguish the categories of all objects in the real 3D scene due to issues such as occlusion, lighting variations, and sensor accuracy limitations. For example, laser scanning may misidentify multiple similar objects as a single entity or fail to distinguish subtle differences in object type. If subsequent applications (such as simulation and analysis) require the differentiation of specific object categories, this inconsistency in category information can lead to semantic discrepancies between the virtual 3D scene and the real 3D scene, not just visual differences.

[0028] Furthermore, since the virtual 3D scene constructed in the above manner does not include object category information, the object category of each object in the virtual 3D scene cannot be determined, resulting in the inability to subsequently edit the constructed virtual 3D scene.

[0029] In order to quickly construct a three-dimensional scene while ensuring that the object categories in the constructed virtual three-dimensional scene are consistent with the object categories in the real three-dimensional scene, and the virtual three-dimensional scene can be subsequently edited based on the object category information of the objects in the virtual three-dimensional scene, the present application proposes a three-dimensional scene construction method, which obtains image data and point cloud data of the real three-dimensional scene, performs target detection based on the image data and point cloud data, and determines the object category and three-dimensional bounding box of each object in the real three-dimensional scene, wherein the three-dimensional bounding box refers to the smallest rectangular block used to enclose an object in three-dimensional space, which can be used to represent the position, size and direction of the object, and then generates a three-dimensional model of each object according to the object category and image data of each object, and adjusts the size of the three-dimensional model according to the size information of the three-dimensional bounding box, and finally arranges the three-dimensional model of each object according to the position information of the three-dimensional bounding box, so as to obtain a virtual three-dimensional scene corresponding to the real three-dimensional scene.

[0030] Figure 1 A flow chart of a three-dimensional scene construction method provided in an embodiment of the present application is shown, and the method is executed by a terminal device, which may be a terminal device including one or more processors, such as a computer, a server, a touch-screen phone, a smart phone, a tablet computer, a portable electronic device, a three-dimensional scanning device or other electronic device, wherein the three-dimensional scanning device includes a camera and a lidar. The processor may be a central processing unit (CPU), or an application-specific integrated circuit (ASIC), or one or more integrated circuits configured to implement the embodiments of the present application, which are not limited here. The one or more processors included in the terminal device may be processors of the same type, such as one or more CPUs; or they may be processors of different types, such as one or more CPUs and one or more ASICs, which are not limited here. Figure 1 As shown, the method includes the following steps:

[0031] Step 110: Acquire image data and point cloud data of a real three-dimensional scene.

[0032] Image data refers to images captured of a real 3D scene, such as RGB images. Point cloud data refers to data obtained by scanning a real 3D scene using a LiDAR (LiDar). 3D scanning equipment is typically used to collect both image data and point cloud data. 3D scanning equipment integrates a camera and LiDAR. The camera captures image data, while the LiDAR captures point cloud data.

[0033] To obtain image data, in an embodiment of the present application, a camera device in a 3D scanning device can be used to capture a real 3D scene to obtain image data, which can then be transmitted to a terminal device executing an embodiment of the present application. The 3D scanning device is connected to the terminal device via one or more of a local area network (LAN), a metropolitan area network (MAN), a wide area network (WAN), a 4G / 5G network, Wi-Fi, Bluetooth, and a point-to-point (P2P) communication network to transmit the collected image data to the terminal device. In an embodiment of the present application, if the terminal device is integrated with a camera device, the camera device in the terminal device can also be used to directly capture the real 3D scene to obtain image data.

[0034] It is worth noting that in this step, the image data obtained is used to subsequently identify various objects existing in the real three-dimensional scene from the image data. Therefore, in order to be able to identify various objects from the image data obtained in this step, preferably, when using a three-dimensional scanning device to photograph the real three-dimensional scene, it should be ensured that every object and every area in the real three-dimensional scene is photographed.

[0035] When the scale of the real three-dimensional scene is large and includes many objects, if each object is photographed one by one using a three-dimensional scanning device, it will take a lot of time and reduce the efficiency of building a virtual three-dimensional model. Therefore, in the example of this application, preferably, when using a three-dimensional scanning device to photograph the real three-dimensional scene, the three-dimensional scanning device is controlled to move along a path passing through all objects and continue photographing, so as to ensure that each object is photographed and improve the efficiency of obtaining image data.

[0036] Since the position information of each object is needed when constructing a virtual three-dimensional scene to ensure that the position of each object in the constructed virtual three-dimensional scene matches the position of the object in the real three-dimensional scene, and image data can only determine the category, shape, etc. of the objects included in the real three-dimensional scene. Therefore, in order to determine the position, size, and other information of each object in the real three-dimensional scene, in an embodiment of the present application, a laser radar in a three-dimensional scanning device can be used to collect point cloud data, and then the position, size, and other information of each object can be determined based on the point cloud data. Specifically, laser pulses can be emitted into the real three-dimensional scene by the laser radar. These pulses are reflected back by objects in the real three-dimensional scene. The laser radar captures the reflected signals and calculates the flight time of the laser pulses based on the emitted laser pulses and the captured reflected signals, thereby determining the distance between the object and the laser radar. At the same time, combined with the horizontal rotation and vertical angle information of the laser radar, the position of each point of the object in three-dimensional space is determined, thereby generating point cloud data.

[0037] In the embodiment of the present application, the point cloud data can be collected by the laser radar in the 3D scanning device and then transmitted to the terminal device. Alternatively, if the terminal device is integrated with a laser radar, the point cloud data can also be directly collected by the laser radar in the terminal device.

[0038] Step 120: Perform target detection based on the image data and point cloud data to determine the object category and three-dimensional bounding box of each object in the real three-dimensional scene.

[0039] In this step, the object category of each object in the three-dimensional scene can be determined based on the image content in the image data, and a known target detection algorithm can be used to perform target detection.

[0040] As mentioned above, a 3D bounding box refers to the smallest cuboid used to enclose an object in 3D space. Point cloud data is composed of a large number of discrete 3D points, each of which carries its specific coordinate values ​​on the X, Y, and Z axes. These coordinate values ​​directly describe the distribution of the object surface in space, that is, the collection of these points reflects the geometric shape and spatial occupancy of the object surface. Therefore, in this step, the 3D bounding box of each object in the real 3D scene can be directly determined through the point cloud data. It can be understood that if there are multiple objects in the real 3D scene, multiple 3D bounding boxes will be determined accordingly in this step, and the multiple 3D bounding boxes correspond one-to-one to the multiple objects.

[0041] In the embodiment of the present application, the three-dimensional bounding box carries the position information (eg, the position information of the center point), size information, and orientation angle information of the bounding box.

[0042] Step 130: For each object, determine target image data including the object from the image data.

[0043] In this step, for each object in the real 3D scene, target image data is selected from the image data acquired in step 110, where the target image data is the image data that includes the object. It should be noted that if a piece of image data includes multiple objects, then that image data can serve as the target image data for each of the multiple objects.

[0044] Step 140 : Generate a first three-dimensional model of each object according to the object category and target image data of the object.

[0045] For each object in the real 3D scene, since the target image data includes the object, the shape of the object can be determined from the target image data. In this step, by combining the object category and shape of the object, a first 3D model that approximates the shape of the object can be generated.

[0046] For each object in a real three-dimensional scene, if a first three-dimensional model of the object is directly generated after only identifying the object's shape from the target image data, the shape of the generated first three-dimensional model may be significantly different from the shape of the object, resulting in a low similarity between the virtual three-dimensional scene subsequently constructed based on the first three-dimensional model and the real three-dimensional scene, that is, the accuracy of the constructed virtual three-dimensional scene is low. For example, when the target image data of the object only includes images obtained by photographing the object from a certain perspective, and the shape of the object from various perspectives cannot be determined through the target image data, or when part of the object in the target image data is obscured by other objects, resulting in the inability to determine the overall shape of the object through the target image data, if the first three-dimensional model is generated only based on the target image data, the shape of the first three-dimensional model will be significantly different from the actual shape of the object.

[0047] In an embodiment of the present application, in addition to determining the shape of each object based on the target image data, the first three-dimensional model is generated in combination with the object category of each object. This allows the first three-dimensional model to restore the shape of each object to the greatest extent possible, thereby improving the accuracy of the virtual three-dimensional scene constructed based on the first three-dimensional model. For example, if the category of an object is a sofa chair, its shape has a certain similarity to the shape of a sofa and the shape of a chair. If the first three-dimensional model is generated directly based on the target image data of the object, the shape of the generated first three-dimensional model may be biased towards the shape of a sofa or the shape of a chair. In the embodiment of the present application, the first three-dimensional model is generated in combination with its object category. Therefore, the shape of the generated first three-dimensional model is closer to the shape of the object, and will not be biased towards the shape of a sofa or the shape of a chair, thereby improving the accuracy of the virtual three-dimensional scene subsequently constructed based on the first three-dimensional model.

[0048] Step 150: Adjust the size of the first three-dimensional model of the object according to the size information carried by the three-dimensional bounding box of each object to obtain a second three-dimensional model of the object.

[0049] Since a 3D bounding box is the smallest rectangular parallelepiped that encloses an object in 3D space, the size information of the 3D bounding box (the length, width, and height of the 3D bounding box) can be used to adjust the size of the first 3D model so that the size of the resulting second 3D model matches the size of the object (i.e., the size is the same or substantially the same).

[0050] Step 160: Construct a virtual three-dimensional scene corresponding to the real three-dimensional scene based on the position information and orientation angle information carried by the three-dimensional bounding box of each object and the second three-dimensional model of the object.

[0051] Since the position information and orientation angle carried by the three-dimensional bounding box are the position and orientation angle of the object in the real three-dimensional scene, in this step, a virtual three-dimensional scene corresponding to the real three-dimensional scene can be constructed by placing the second three-dimensional model according to the position information and orientation angle carried by the three-dimensional bounding box.

[0052] In an embodiment of the present application, the category and number of objects in the real three-dimensional scene are accurately determined through image data of the real three-dimensional scene, and the position and size of each object are determined through point cloud data. Then, a first three-dimensional model that is relatively similar to the shape of the object is generated based on the object category of each object and the target image data including the object, and a virtual three-dimensional scene is constructed based on the size, position information and the first three-dimensional model of the object, thereby improving the efficiency and accuracy of constructing the three-dimensional scene.

[0053] Moreover, since each first three-dimensional model is determined based on the object category of each object, and the second three-dimensional model is obtained after adjusting the size of the first three-dimensional model, for the virtual three-dimensional scene constructed in the embodiment of the present application, the object category of the object corresponding to each second three-dimensional model in the virtual three-dimensional scene can be determined, so that the virtual three-dimensional scene can be edited based on the object category of the object corresponding to each second three-dimensional model.

[0054] In order to accurately determine the object category and 3D bounding box of each object in the real 3D scene and improve the accuracy of the constructed virtual 3D scene, Figure 2 Shown Figure 1 The sub-step flow diagram of step 120 is as follows: Figure 2 As shown, step 120 includes steps 121 to 124.

[0055] In this embodiment, a Hierarchical Cross-Modal Fusion Network for 3D Scene Layout Perception (HCF-Net) is used to determine object categories and 3D bounding boxes. The HCF-Net deep learning network builds on traditional models like PointNet++ by adding a multimodal fusion module and introducing open-lexicon semantic modeling capabilities, enabling it to recognize and classify any object category, including those not covered during training. Figure 3 The schematic diagram of determining the object category and 3D bounding box of each object in a real 3D scene provided by an embodiment of the present application is shown. Figure 2 and Figure 3 .

[0056] Step 121: input the image data and point cloud data into the dual-stream feature extraction module to obtain the semantic feature vector and the three-dimensional geometric feature vector output by the dual-stream feature extraction module.

[0057] like Figure 3 As shown, the dual-stream feature extraction module includes a two-dimensional semantic feature extraction branch and a three-dimensional geometric feature extraction branch. The two-dimensional semantic feature extraction branch receives image data and performs feature extraction on the image data to obtain a semantic feature vector; the three-dimensional geometric feature extraction branch receives point cloud data and performs feature extraction on the point cloud data to obtain a three-dimensional geometric feature vector.

[0058] The 2D semantic feature extraction branch uses the Visual Transformer as its backbone network. This branch is responsible for extracting multi-scale deep semantic features from image data, capturing the color, texture, and component information of objects. The 3D geometric feature extraction branch utilizes point cloud processing networks such as the Point Transformer. This branch focuses on learning the local geometric details, topological relationships between points, and overall spatial distribution characteristics of point clouds.

[0059] Step 122: Input the semantic feature vector and the three-dimensional geometric feature vector into a feature fusion module to obtain a fusion feature output by the feature fusion module.

[0060] The feature fusion module dynamically fuses and semantically expands semantic feature vectors across modalities with 3D geometric feature vectors. Specifically, it constructs a multi-level cross-attention map between the 2D semantic feature vectors and the 3D geometric feature vectors, dynamically assigns fusion weights for texture and geometric information, and introduces contrastive learning loss to enhance semantic discrimination, particularly for object categories not covered during training.

[0061] This feature fusion module enables each 3D point to "focus on" and fuse the object category semantic information of the most relevant objects in the image, thereby enhancing its own feature representation ability and contextual perception of scene understanding. In other words, the fused features output by this module carry both object category semantic information and geometric features.

[0062] Step 123: Input the fusion features into the depth fusion module to obtain the depth fusion features output by the depth fusion module.

[0063] The deep fusion module is a Deep Fusion Transformer Module. The aligned and enhanced 3D geometric feature vectors (enriched with preliminary semantics) and the associated 2D semantic feature vectors (region-selective image features) are jointly transmitted to the Deep Fusion Module. In this module, the 3D geometric feature vectors and 2D semantic feature vectors are treated as distinct "tokens" in a unified sequence. The powerful self-attention mechanism within the Transformer captures the complex dependencies and deep interactions between these cross-modal tokens, generating a highly fused and information-rich joint feature representation.

[0064] Step 124: Input the deep fusion features into a prediction head module including a parallel object classification head and a 3D bounding box regression head to obtain the object category of each object output by the object classification head and the 3D bounding box of each object output by the 3D bounding box regression head.

[0065] The object classification head consists of a small multi-layer perceptron (MLP), which is responsible for predicting the category of each perceived object. The 3D bounding box regression head also uses an MLP structure to regress the 3D center point, size (length, width, and height), and orientation of the object.

[0066] It is worth noting that in this application, the object classification head is changed from a fixed category softmax to a similarity matching structure, and the open-domain text encoder will be called, and the features of the output of the visual feature encoder will be projected into the language space. Moreover, the last layer of the traditional classifier is Linear(C,d), and its weight matrix can be regarded as the prototype of each category. In the open-domain text encoder in this application, the fixed weight matrix is ​​replaced by a dynamic text embedding. It is worth noting that the text encoder Text Encoder is located at the input end and is used to process descriptive text (such as object category names or noun phrases) and generate text embedding vectors, which are responsible for converting text information into semantic representations aligned with visual features, providing a basis for subsequent cross-modal interaction. At the same time, the object classification head uses text embedding as a comparison target to achieve recognition of any object category.

[0067] When using LiDAR to collect point cloud data of a real 3D scene, if there are two objects that are close to each other in the real 3D scene, the point cloud data of these two objects may be mistakenly determined as the point cloud data of a single object. This can lead to the inability to accurately determine the 3D bounding boxes of each object based on the point cloud data, resulting in low accuracy in the constructed virtual 3D scene. In the embodiment of the present application, by fusing image data and point cloud data, point cloud data belonging to different objects in the point cloud data can be better distinguished by combining the image data, thereby more accurately identifying the boundaries of each independent object and accurately determining the 3D bounding box of each object, thereby improving the accuracy of the constructed virtual 3D scene.

[0068] In order to accurately determine the target image data and accurately generate the three-dimensional model of each object, Figure 4 Shown Figure 1 Schematic diagram of the sub-step flow chart of step 130. Figure 4 As shown, step 130 includes the following steps 131 to 137.

[0069] Step 131: For each object, select first image data including the object from the image data.

[0070] As previously described, when using a camera to capture image data of a real three-dimensional scene, the camera is controlled to move along a path passing through all objects and continuously capture images. Therefore, the camera may capture the same object from different angles. Therefore, for a given object, multiple images may include the object in the image data acquired in step 110. Based on this, in this step, the first image data including each object is selected from the image data acquired in step 110.

[0071] Step 132: For the object, determine whether there is only one first image data. If so, go to step 133; if not, go to step 134.

[0072] Step 133: Determine the first image data as target image data.

[0073] Wherein, for a certain object in a real three-dimensional scene, if there is only one first image data including the object, then in this step, the first image data is determined as the target image data.

[0074] Step 134: Score the plurality of first image data.

[0075] It is understood that when a real three-dimensional scene contains many objects, the object in the first image data may be partially obscured by other objects, and the degree of obscuration of the object in different first image data may vary. For example, in some first image data, the area of ​​the object obscured by other objects may be larger, while in other first image data, the area of ​​the object obscured by other objects may be smaller. However, it is understood that the smaller the area of ​​the object obscured by other objects in the first image data, the more accurately the overall shape of the object can be determined using the first image data.

[0076] Therefore, for each object in a real three-dimensional scene, when there are multiple first image data that include the object, in this step, each of the multiple first image data is scored, so that the target image data of the object can be subsequently determined based on the score of each first image data. Specifically, the score of each photo can be obtained by taking a weighted sum of the size of the object's occluded area in the first image data, the image clarity and resolution of the first image data, etc. The score of the first image data is negatively correlated with the size of the object's occluded area in the first image data, and positively correlated with the image clarity and resolution of the first image data.

[0077] Step 135: Determine whether there is any second image data with a score greater than or equal to a preset score threshold among the plurality of first image data. If yes, go to step 136; if not, go to step 137.

[0078] The preset score threshold can be determined as needed.

[0079] Step 136: Determine the second image data as the target image data.

[0080] If the score of the second image data is greater than or equal to the preset score threshold, it means that the second image data meets the requirements, and the second image data is used as the target image data.

[0081] Step 137 : Determine the first image data with the highest score among the plurality of first image data as the target image data.

[0082] Among them, if the scores of multiple first image data are all less than the preset score threshold, in order to subsequently determine the first three-dimensional model of the object based on the target image data, in this step, the first image data with the highest score is determined as the target image data.

[0083] In an embodiment of the present application, for each object in a real three-dimensional scene, if there are multiple first image data that include the object, the target image data of the object is determined by the above method. Compared with the method of randomly determining one first image data as the target image data from multiple first image data, the accuracy of the first three-dimensional model determined based on the target image data can be improved, and the shape of the generated first three-dimensional model can be closer to the shape of the actual object, thereby improving the accuracy of the constructed three-dimensional model.

[0084] In order to further improve the accuracy of the generated first 3D model and make its shape closer to the shape of the real object, Figure 4 Based on the provided embodiments, in the embodiments of the present application, for each object in a real three-dimensional scene, if there is only one target image data, a first three-dimensional model is generated based on the object category and target image data of the object through a deep learning model; if there are multiple target image data, a first three-dimensional model is generated based on the object category and multiple target image data of the object through a multi-view stereo vision method (Multi-ViewStereo, MVS).

[0085] Specifically, for a certain object, if there is only one target image data, a deep learning model with a Transformer architecture is used (for example, through a large reconstruction model (LRM) or CatFree3D, etc.) to directly predict the three-dimensional mesh, and at the same time use large-scale data sets for training and combine adversarial loss, perceptual loss, etc. to optimize the three-dimensional scene. For a certain object, if there are multiple target image data, an MVS method is used (such as a depth map estimation-based or voxel-based method, referring to the ideas of tool chains such as 360MVSNet and COLMAP) to estimate the depth information, fuse it to generate a point cloud, and then perform meshing and post-processing (smoothing, hole filling) to construct a three-dimensional scene.

[0086] In order to obtain a more accurate virtual three-dimensional scene, in the embodiment of the present application, step 160 includes:

[0087] Step a1: Initially align the second 3D model based on the position information and orientation angle information carried by the 3D bounding box to obtain a first 3D scene constructed by the second 3D model.

[0088] In this step, the second three-dimensional model is preliminarily placed using the information carried by the three-dimensional bounding box to construct the first three-dimensional scene, and the position and orientation angle of the second three-dimensional model in the first three-dimensional scene respectively match the position information and orientation angle information carried by the three-dimensional bounding box (for example, the two are approximately the same).

[0089] Step a2: Accurately align the second 3D model in the first 3D scene using an iterative closest point (ICP) algorithm to obtain a second 3D scene.

[0090] In this step, the key points of the model are extracted based on geometric features (such as curvature, normal vector), texture features or shape descriptors (such as FPFH, SHOT), and the initial correspondence is calculated by matching these features. Then, the ICP algorithm is used to iteratively optimize the rigid transformation (rotation, translation) to minimize the Euclidean distance error between point clouds or meshes, and finally achieve high-precision alignment of the second 3D model.

[0091] Step a3: performing scene optimization processing on the second three-dimensional scene to obtain a virtual three-dimensional scene.

[0092] After constructing the virtual three-dimensional scene, in order to improve rendering efficiency and operating performance, the second three-dimensional scene is optimized in this step, including reducing the number of polygons and merging materials. Reducing the number of polygons refers to simplifying the geometric structure of the second three-dimensional model through an algorithm to reduce the number of its vertices and faces. Moreover, the reduction is usually small, such as 30%, so it will not cause the shape of the second three-dimensional model to be inconsistent with the shape of the real object. Material merging refers to merging multiple adjacent models or mesh objects using the same material into a whole, and mainly merging those "attachments". For example, the second three-dimensional scene includes a table and a laptop attached to the table, then the contact surfaces of the two in the real three-dimensional scene can be fused together, thereby avoiding the need to re-do the fusion process caused by the connection when the three-dimensional scene is constructed next time, and it can also simplify the complexity of the scene model.

[0093] In the embodiments of the present application, by first performing an initial alignment on the second 3D model and then performing a precise alignment, the positions of the respective second 3D models in the second 3D scene constructed by the second 3D model match the positions of the real objects, thereby improving the accuracy of the constructed virtual 3D scene. Furthermore, by performing scene optimization processing on the second 3D scene, rendering efficiency can be improved.

[0094] Given the presence of multiple objects in a real 3D scene, and the corresponding second 3D scene also including multiple second 3D models, in order to eliminate noticeable seams between the models, the present application also includes performing scene fusion processing on the second 3D scene before step a3. Specifically, advanced mesh fusion algorithms (such as those based on the Poisson equation) are used to smooth the seams of the second 3D scene and optimize the relationships between the second 3D models based on physical constraints. Poisson-based methods can be used to generate transition surfaces or adjust local geometric structures to ensure that the gradient (i.e., the geometric change trend) of the spliced ​​area is consistent with that of the surrounding area, thereby eliminating cracks or sudden changes.

[0095] In order to further improve the accuracy of the constructed virtual 3D scene, Figure 1 Based on the embodiment provided, in the embodiment of the present application, after step 160, the following steps are further included:

[0096] Step b1: determining the interpenetration regions between different second three-dimensional models in the virtual three-dimensional scene.

[0097] Intersection refers to the abnormal overlap or penetration of two or more secondary 3D models. In this step, you can combine object categories and 3D bounding boxes to use precise collision detection (such as based on the GJK algorithm) for key objects and axis-aligned bounding boxes (AABB) for the background to quickly detect intersection areas.

[0098] Step b2: determining a target second three-dimensional model including the mold-piercing area according to the mold-piercing area.

[0099] If penetration occurs in some of the second three-dimensional models, these second three-dimensional models are determined as target second three-dimensional models.

[0100] Step b3: Processing the target second three-dimensional model by at least one of adjusting transparency, moving, rotating, and resizing until no intervening regions exist between different second three-dimensional models in the virtual three-dimensional scene.

[0101] Among them, if some second three-dimensional models have penetration phenomenon, the position or size of one or more second three-dimensional models can be adjusted so that these second three-dimensional models no longer have penetration phenomenon. Specifically, the position of the second three-dimensional model can be adjusted by moving and rotating the second three-dimensional model. In addition to the above methods, the depth buffer (Z-buffer) or depth sorting similar to the depth sorting algorithm (Painter's Algorithm) can also be used to process occlusion and penetration areas, sort and render transparent objects, or use advanced technologies such as Order-Independent Transparency (OIT) to adjust the transparency of the second three-dimensional model to process occlusion and penetration areas. For areas with occlusion in the virtual three-dimensional scene, frustum culling, backface culling and occlusion query-based technologies can be used to process the occluded areas, thereby reducing rendering overhead.

[0102] In the embodiment of the present application, the virtual three-dimensional scene is adjusted in the above manner to avoid situations that violate reality in the virtual three-dimensional scene, thereby improving the accuracy of the virtual three-dimensional scene.

[0103] In some embodiments, to improve the accuracy of the virtual three-dimensional scene, a user operation performed on a target second three-dimensional model is received and the target second three-dimensional model is processed according to the operation until no cross-over areas exist between different second three-dimensional models in the virtual three-dimensional scene. For example, the terminal device displays the virtual three-dimensional scene to the user, and after the user performs an operation (such as rotating or resizing) on ​​a second three-dimensional model, the terminal device receives the user operation performed on the second three-dimensional model and then performs the operation on the second three-dimensional model accordingly, thereby updating the virtual three-dimensional scene until no cross-over areas exist between different second three-dimensional models in the virtual three-dimensional scene.

[0104] Figure 5 FIG. 1 shows a schematic diagram of the structure of a three-dimensional scene construction device provided in an embodiment of the present application. Figure 5 As shown, the three-dimensional scene construction device 200 includes: a data acquisition module 201, a target detection module 202, a determination module 203, a three-dimensional model generation module 204, a three-dimensional model adjustment module 205, and a three-dimensional scene construction module 206.

[0105] The data acquisition module 201 acquires image data and point cloud data of a real-world 3D scene. The object detection module 202 performs object detection based on the image data and point cloud data, determining the object category and 3D bounding box of each object in the real-world 3D scene. The 3D bounding box carries information about the bounding box's position, size, and orientation angle. The determination module 203 determines, for each object, the target image data containing that object from the image data. The 3D model generation module 204 generates a first 3D model of each object based on the object category and target image data. The 3D model adjustment module 205 adjusts the size of the first 3D model of each object based on the size information carried by the 3D bounding box of each object, thereby obtaining a second 3D model of the object. The 3D scene construction module 206 constructs a virtual 3D scene corresponding to the real-world 3D scene based on the position information, orientation angle information carried by the 3D bounding box of each object and the second 3D model of the object.

[0106] The 3D scene construction device 200 provided in this embodiment is used to execute the technical solution of the 3D scene construction method in the aforementioned method embodiment. Its implementation principle and technical effects are similar and will not be described in detail here.

[0107] It is worth noting that the 3D scene construction device 200 provided in this embodiment further includes other modules for executing the steps of the above-mentioned 3D scene construction method embodiment, which will not be described in detail here.

[0108] Figure 6A schematic structural diagram of a three-dimensional scene construction device provided in an embodiment of the present application is shown. The specific embodiment of the present application does not limit the specific implementation of the three-dimensional scene construction device.

[0109] like Figure 6 As shown, the three-dimensional scene construction device 300 may include a processor 302 and a memory 304 .

[0110] The memory 304 is used to store a computer program 306. The memory 304 may include a high-speed RAM memory, or may also include a non-volatile memory (non-volatile memory), such as at least one disk memory. The computer program 306 may include computer-executable instructions.

[0111] The processor 302 is configured to execute the computer program 306 to implement the above-mentioned embodiment of the three-dimensional scene construction method.

[0112] Processor 302 may be a central processing unit (CPU), an application-specific integrated circuit (ASIC), or one or more integrated circuits configured to implement the embodiments of the present application. The one or more processors included in the 3D scene construction device 300 may be processors of the same type, such as one or more CPUs, or may be processors of different types, such as one or more CPUs and one or more ASICs.

[0113] An embodiment of the present application provides a computer-readable storage medium, which stores a computer program. When the computer program is executed by a processor, the above-mentioned three-dimensional scene construction method embodiment is implemented.

[0114] An embodiment of the present application provides a computer program that can be executed by a processor to implement the above-mentioned three-dimensional scene construction method embodiment.

[0115] An embodiment of the present application provides a computer program product, which includes a computer program. When the computer program is executed by a processor, it implements the above-mentioned embodiment of the three-dimensional scene construction method.

[0116] In the several embodiments provided in this application, if any function is implemented in the form of a software function module / unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, all or part of the technical solution of this application can be embodied in the form of a software product, and the computer software product is stored in a storage medium, including a number of instructions for enabling a computer device (which can be a personal computer, server or other electronic device) to execute all or part of the steps of the method described in each embodiment of this application. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (ROM), random access memory (RAM), disk or optical disk, and other media that can store computer program code.

[0117] The algorithm or demonstration provided here are not inherently relevant to any particular computer, virtual system or other equipment. Various general purpose systems can also be used together with the teachings based on this. According to the above description, it is obvious that the structure required for constructing this type of system. In addition, the present application embodiment is not directed to any specific programming language yet. It should be understood that various programming languages ​​can be utilized to realize the content of the present application described here, and the above description of specific languages ​​is for the purpose of disclosing the best mode of implementation of the present application.

[0118] It should be noted that the above embodiments illustrate rather than limit the present application, and that a person skilled in the art may devise alternative embodiments without departing from the scope of the appended claims. In the claims, any reference signs placed between brackets should not be construed as limiting the claims. The word "comprising" does not exclude the presence of elements or steps not listed in the claims. The word "a" or "an" preceding an element does not exclude the presence of a plurality of such elements. The present application may be implemented by means of hardware comprising several different elements and by means of a suitably programmed computer. In claims that list several means, several units or modules of these means may be embodied by the same item of hardware. The use of the words first, second, and third etc. does not indicate any order. These words may be interpreted as names. The steps in the above embodiments should not be understood as limiting the order of execution unless otherwise specified.

[0119] The above-described embodiments merely represent several implementation methods of the present application. While the descriptions are relatively specific and detailed, they should not be construed as limiting the scope of the present application. It should be noted that a person of ordinary skill in the art may make various modifications and improvements without departing from the spirit of the present application, and these modifications and improvements fall within the scope of protection of the present application. Therefore, the scope of protection of the present application shall be determined by the appended claims.

Claims

1. A three-dimensional scene construction method, characterized in that: The method comprises: Acquire image data and point cloud data of real three-dimensional scenes; performing target detection based on the image data and the point cloud data to determine an object category and a three-dimensional bounding box of each object in the real three-dimensional scene, wherein the three-dimensional bounding box carries position information, size information, and orientation angle information of the bounding box; For each of the objects, determining target image data including the object from the image data; generating a first three-dimensional model of each object according to the object category and the target image data of the object; adjusting the size of the first three-dimensional model of the object according to the size information carried by the three-dimensional bounding box of each object to obtain a second three-dimensional model of the object; constructing a virtual three-dimensional scene corresponding to the real three-dimensional scene based on the position information and the orientation angle information carried by the three-dimensional bounding box of each object and the second three-dimensional model of the object; The step of generating a first three-dimensional model of each object according to the object category and the target image data includes: For each of the objects, if there is only one target image data, generating the first three-dimensional model based on the object category and the target image data of the object through a deep learning model; If there are multiple target image data, the first three-dimensional model is generated based on the object category of the object and the multiple target image data using a multi-viewpoint stereo vision method.

2. The method according to claim 1, characterized in that The performing target detection based on the image data and the point cloud data to determine the object category and three-dimensional bounding box of each object in the real three-dimensional scene includes: Inputting the image data and the point cloud data into a dual-stream feature extraction module to obtain a semantic feature vector and a three-dimensional geometric feature vector output by the dual-stream feature extraction module, wherein the dual-stream feature extraction module includes a two-dimensional semantic feature extraction branch and a three-dimensional geometric feature extraction branch, the two-dimensional semantic feature extraction branch is used to receive the image data and perform feature extraction on the image data to obtain the semantic feature vector, and the three-dimensional geometric feature extraction branch is used to receive the point cloud data and perform feature extraction on the point cloud data to obtain the three-dimensional geometric feature vector; Inputting the semantic feature vector and the three-dimensional geometric feature vector into a feature fusion module to obtain a fusion feature output by the feature fusion module; Inputting the fusion features into a depth fusion module to obtain the depth fusion features output by the depth fusion module; The deep fusion features are input into a prediction head module including a parallel object classification head and a three-dimensional bounding box regression head to obtain the object category of each object output by the object classification head and the three-dimensional bounding box of each object output by the three-dimensional bounding box regression head.

3. The method according to claim 1, characterized in that For each of the objects, determining target image data including the object from the image data comprises: For each of the objects, selecting first image data including the object from the image data; For the object, if there is only one piece of the first image data, the first image data is determined as the target image data; if there are multiple pieces of the first image data, the multiple pieces of the first image data are scored, and the target image data is determined based on the score of each piece of the first image data.

4. The method according to claim 3, characterized in that The determining the target image data according to the score of each piece of the first image data includes: If there is second image data with a score greater than or equal to a preset score threshold among the plurality of first image data, determining the second image data as the target image data; If the scores of the plurality of first image data are all smaller than the preset score threshold, the first image data with the highest score among the plurality of first image data is determined as the target image data.

5. The method according to claim 1, wherein The constructing a virtual three-dimensional scene corresponding to the real three-dimensional scene based on the position information, the orientation angle information and the second three-dimensional model of the object carried by the three-dimensional bounding box of each object includes: Initially aligning the second three-dimensional model according to the position information and the orientation angle information carried by the three-dimensional bounding box to obtain a first three-dimensional scene constructed by the second three-dimensional model, wherein the position and orientation angle of the second three-dimensional model in the first three-dimensional scene match the position information and the orientation angle information, respectively; Accurately aligning the second three-dimensional model in the first three-dimensional scene using an iterative closest point algorithm to obtain a second three-dimensional scene; The second three-dimensional scene is subjected to scene optimization processing to obtain the virtual three-dimensional scene.

6. The method according to claim 1, characterized in that After constructing a virtual three-dimensional scene corresponding to the real three-dimensional scene based on the position information, the orientation angle information carried by the three-dimensional bounding box of each object and the second three-dimensional model of the object, the method further includes: determining a penetration region between different second three-dimensional models in the virtual three-dimensional scene; Determine a target second three-dimensional model including the mold-piercing area according to the mold-piercing area; The target second three-dimensional model is processed by at least one of adjusting transparency, moving, rotating and resizing, and / or receiving an operation performed by a user on the target second three-dimensional model, and the target second three-dimensional model is processed according to the operation until no penetration area exists between different second three-dimensional models in the virtual three-dimensional scene.

7. A three-dimensional scene construction device, characterized in that: The device comprises: A data acquisition module is used to acquire image data and point cloud data of real three-dimensional scenes; an object detection module, configured to perform object detection based on the image data and the point cloud data, and determine an object category and a three-dimensional bounding box of each object in the real three-dimensional scene, wherein the three-dimensional bounding box carries position information, size information, and orientation angle information of the bounding box; a determination module, configured to determine, for each of the objects, target image data including the object from the image data; a three-dimensional model generation module, configured to generate a first three-dimensional model of each object based on the object category and the target image data of each object, and further configured to, for each object, if only one target image data exists, generate the first three-dimensional model based on the object category and the target image data of the object using a deep learning model; and if multiple target image data exist, generate the first three-dimensional model based on the object category and the multiple target image data of the object using a multi-view stereo vision method; a three-dimensional model adjustment module, configured to adjust the size of the first three-dimensional model of each object according to the size information carried by the three-dimensional bounding box of each object, to obtain a second three-dimensional model of the object; A three-dimensional scene construction module is used to construct a virtual three-dimensional scene corresponding to the real three-dimensional scene based on the position information, the orientation angle information and the second three-dimensional model of the object carried by the three-dimensional bounding box of each object.

8. A three-dimensional scene construction device, comprising a memory, a processor, and a computer program stored in the memory, characterized in that: The processor executes the computer program to implement the three-dimensional scene construction method according to any one of claims 1 to 6.

9. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the three-dimensional scene construction method according to any one of claims 1 to 6 is implemented.

Citation Information

Patent Citations

  • Multi-view outdoor three-dimensional scene reconstruction method and system

    CN117036590A

  • Real-Time Dynamic Three-Dimensional Adaptive Object Recognition and Model Reconstruction

    US20160071318A1