Three-dimensional visualization method and device for target scene
By constructing and optimizing the three-dimensional Gaussian body collection, the problems of high labeling costs, low efficiency and insufficient details in the three-dimensional reconstruction of large-scale urban market scenes are solved, and high-quality three-dimensional modeling and efficiency improvement are achieved.
Patent Information
- Application Number
- CN202411958808.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-27
- Publication Date
- 2025-05-30
AI Technical Summary
When facing large-scale urban market scenarios, existing three-dimensional reconstruction technology has problems such as high labeling costs, limited model capacity and insufficient detail fidelity, and is low in efficiency.
By obtaining images from multiple shooting angles, a three-dimensional Gaussian body set of the target scene is constructed, and divided into multiple subsets for optimization processing, and the three-dimensional visual model of the target scene is finally determined.
High-quality three-dimensional modeling of large-scale complex scenarios is realized, which reduces data preparation costs, improves the efficiency and detail fidelity of three-dimensional reconstruction, and avoids the problem of excessive memory overhead of graphics processors.
Smart Images

Figure CN120070695A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of 3D modeling technology, and more specifically, to a 3D visualization method and device for a target scene. Background Art
[0002] 3D reconstruction is a technology that transforms the scenes of the real world into digital 3D models. This technology is widely used in fields such as urban planning, virtual reality, and augmented reality.
[0003] The methods for urban scene reconstruction include traditional methods and optimized methods. Among them, the traditional method converts the real city into a digital 3D model through steps such as image data acquisition, feature extraction and matching, sparse and dense reconstruction, mesh reconstruction, and texture mapping. Although 3D models of relatively high quality can be generated in this way, when faced with large-scale urban scenes, there are problems such as high annotation costs and limited model capacity. The optimized methods are mainly dominated by methods based on Neural Radiance Fields (NeRF), and representative works include Block-NeRF, BungeeNeRF, and ScaNeRF. NeRF directly models the density and color information of the scene in terms of spatial position and viewing direction through a neural network, thereby achieving high-quality 3D reconstruction. This method has great deficiencies in the detail fidelity of the scene, and moreover, shows slow performance in the efficiency of 3D reconstruction. Summary of the Invention
[0004] An object of an embodiment of the present invention is to provide a new technical solution for 3D visualization of a target scene to improve the quality and efficiency of 3D reconstruction of the target scene.
[0005] According to a first aspect of the present invention, there is provided a 3D visualization method for a target scene, which includes:
[0006] Obtaining a plurality of first images captured for the target scene; wherein, the plurality of first images are obtained by photographing the target scene from a plurality of shooting perspectives;
[0007] According to the plurality of first images, constructing a 3D Gaussian volume set of the target scene; wherein, the 3D Gaussian volume set includes a plurality of 3D Gaussian volumes, and each 3D Gaussian volume in the plurality of 3D Gaussian volumes corresponds to a set of Gaussian volume parameters, and the Gaussian volume parameters include the center position coordinates of the Gaussian volume;
[0008] According to the center position coordinates corresponding to each 3D Gaussian volume in the 3D Gaussian volume set, dividing the 3D Gaussian volume set into a plurality of first subsets;
[0009] For each of the multiple first subsets, perform optimization processing on the first subset according to at least one of the first images associated with the first subset to obtain multiple optimized first subsets corresponding to the multiple first subsets;
[0010] Determine a three-dimensional visualization model of the target scene according to the multiple optimized first subsets.
[0011] Optionally, the constructing a three-dimensional Gaussian volume set of the target scene according to the multiple first images includes:
[0012] Construct a sparse point cloud of the target scene according to the camera pose information matched by each of the multiple first images;
[0013] Align the sparse point cloud according to the edge features of the entity objects included in the multiple first images to obtain an aligned sparse point cloud;
[0014] Construct a three-dimensional Gaussian volume set of the target scene according to the aligned sparse point cloud.
[0015] Optionally, the aligning the sparse point cloud according to the edge features of the entity objects included in the multiple first images to obtain an aligned sparse point cloud includes:
[0016] Extract the edge features of the entity objects included in the multiple first images;
[0017] When the edge features of the entity objects include parallel edge features and vertical edge features, determine a horizontal plane corresponding to the parallel edge features in the sparse point cloud according to the parallel edge features, and determine a vertical plane corresponding to the vertical edge features in the sparse point cloud according to the vertical edge features;
[0018] Perform parallel alignment processing on the horizontal plane points of the horizontal plane and vertical alignment processing on the vertical plane points of the vertical plane to obtain the aligned sparse point cloud.
[0019] Optionally, the dividing the three-dimensional Gaussian volume set into multiple first subsets according to the central position coordinates corresponding to each three-dimensional Gaussian volume in the three-dimensional Gaussian volume set includes:
[0020] Project the central position coordinates corresponding to each three-dimensional Gaussian volume onto a horizontal plane to obtain the central position projection coordinates of each three-dimensional Gaussian volume;
[0021] Divide the three-dimensional Gaussian volume set into multiple first subsets according to the central position projection coordinates of each three-dimensional Gaussian volume.
[0022] Optionally, dividing the set of three-dimensional Gaussian bodies into a plurality of first subsets according to the central position projection coordinates of each three-dimensional Gaussian body includes:
[0023] Dividing the projection area of the set of three-dimensional Gaussian bodies on the horizontal plane into a plurality of first sub-areas of equal size at a first set area interval; wherein, there is a set proportion of boundary overlapping areas between adjacent first sub-areas among the plurality of first sub-areas;
[0024] Allocating each three-dimensional Gaussian body to a corresponding first sub-area according to the central position projection coordinates of each three-dimensional Gaussian body;
[0025] For each first sub-area among the plurality of first sub-areas, determining a first subset corresponding to the first sub-area according to the plurality of three-dimensional Gaussian bodies allocated to the first sub-area, and obtaining a plurality of first subsets corresponding to the plurality of first sub-areas.
[0026] Optionally, the step of determining at least one of the first images associated with the first subset includes:
[0027] Projecting the first subset onto the image plane of each of the plurality of shooting perspectives according to the plurality of shooting perspectives to obtain projection information of each shooting perspective;
[0028] For each of the shooting perspectives, determining whether the shooting perspective and the first subset are strongly associated according to the projection information corresponding to the shooting perspective;
[0029] In the case where the shooting perspective and the first subset are strongly associated, associating the first image corresponding to the shooting perspective with the first subset to obtain at least one first image associated with the first subset.
[0030] Optionally, the projection information includes the number of two-dimensional Gaussian distributions. Determining whether the shooting perspective and the first subset are strongly associated according to the projection information corresponding to the shooting perspective includes:
[0031] Determining a projection ratio of the shooting perspective according to the number of two-dimensional Gaussian distributions in the projection information corresponding to the shooting perspective and the number of three-dimensional Gaussian bodies included in the first subset;
[0032] In the case where the projection ratio of the shooting perspective is greater than or equal to a projection ratio threshold, determining that the shooting perspective and the first subset are strongly associated.
[0033] Optionally, the determining the three-dimensional visualization model of the target scene according to the plurality of optimized first subsets includes:
[0034] Divide the projected area of the three-dimensional Gaussian body set on the horizontal plane into a plurality of second sub-regions of equal size at a second set region interval; wherein, there is no boundary overlapping region between adjacent second sub-regions among the plurality of second sub-regions, and the second set region interval is greater than the first set region interval;
[0035] According to the plurality of second sub-regions, re-divide the plurality of optimized first sub-sets to obtain a second sub-set corresponding to each second sub-region among the plurality of second sub-regions;
[0036] Determine a three-dimensional visualization model of the target scene according to the second sub-set corresponding to each second sub-region.
[0037] Optionally, after determining the three-dimensional visualization model of the target scene, the method further includes:
[0038] Obtain and output a second image of the three-dimensional visualization model from the target rendering perspective according to the target rendering perspective;
[0039] Perform super-resolution processing on the second image to obtain and output a third image from the target rendering perspective.
[0040] According to a second aspect of the present invention, there is also provided a three-dimensional visualization device for a target scene, which includes a memory and a processor, the memory is used to store executable instructions; the processor is used to operate according to the control of the instructions to execute the method as described in the first aspect of the present invention.
[0041] One beneficial effect of the present invention is that by constructing a three-dimensional Gaussian body set of the target scene, and each three-dimensional Gaussian body in the three-dimensional Gaussian body set contains a detailed description of the target scene in a certain area, various details in the complex target scene can be represented more precisely, and the dependence on a large amount of manually labeled data in the traditional method can be reduced, the cost of data preparation is reduced, and the efficiency of three-dimensional reconstruction is higher. By dividing the three-dimensional Gaussian body set into a plurality of first sub-sets, each of the plurality of first sub-sets is respectively optimized, and after obtaining the optimized first sub-sets, they are spliced to obtain a three-dimensional visualization model of the target scene, which can realize three-dimensional modeling of large-scale complex scenes, and the optimization of the plurality of first sub-sets can be processed in parallel independently, thus avoiding the problem of excessive graphics processing unit (GPU) memory overhead during large-scale scene modeling, improving the efficiency of three-dimensional modeling. In addition, by optimizing each first sub-set through at least one first image associated with it, the detail fidelity in large-scale scenes can also be improved. BRIEF DESCRIPTION OF THE DRAWINGS
[0042] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate embodiments of the invention and, together with the description, serve to explain the principles of the invention.
[0043] Figure 1 is a schematic diagram of the hardware structure of a three-dimensional visualization device for a target scenario according to an embodiment of the present invention;
[0044] Figure 2 is a schematic flowchart of a three-dimensional visualization method for a target scenario according to an embodiment of the present invention;
[0045] Figure 3 is a schematic diagram of the structure of a three-dimensional visualization device for a target scenario according to an embodiment of the present invention. Detailed Embodiments
[0046] Various exemplary embodiments of the present invention will now be described in detail with reference to the accompanying drawings. It should be noted that: unless otherwise specifically stated, the relative arrangements of components and steps, numerical expressions, and numerical values set forth in these embodiments do not limit the scope of the present invention.
[0047] The following description of at least one exemplary embodiment is merely illustrative in nature and is in no way intended to limit the present invention, its application, or its use.
[0048] Technologies, methods, and devices known to those of ordinary skill in the relevant art may not be discussed in detail, but where appropriate, such technologies, methods, and devices should be considered as part of the specification.
[0049] In all examples shown and discussed herein, any specific values should be construed as merely exemplary and not as limitations. Thus, other examples of exemplary embodiments may have different values.
[0050] It should be noted that: like reference numerals and letters denote like items in the following drawings, and thus, once an item is defined in one drawing, it need not be further discussed in subsequent drawings.
[0051] <Hardware Configuration>
[0052] Figure 1 is a schematic diagram of the structure of a three-dimensional visualization device 100 for a target scenario according to an embodiment of the present invention.
[0053] As Figure 1 shown, the three-dimensional visualization device 100 for a target scenario can be any electronic device, such as a PC, a laptop, a server, etc.
[0054] In this embodiment, with reference to Figure 1As shown, the three-dimensional visualization device 100 of the target scenario may include a processor 1100, a memory 1200, an interface device 1300, a communication device 1400, a display device 1500, an input device 1600, a speaker 1700, a microphone 1800, and so on.
[0055] The processor 1100 may be a mobile version processor. The memory 1200 includes, for example, ROM (Read Only Memory), RAM (Random Access Memory), non-volatile memory such as a hard disk, etc. The interface device 1300 includes, for example, a USB interface, a headphone interface, etc. The communication device 1400 can perform wired or wireless communication, for example. The communication device 1400 may include a short-distance communication device, for example, any device that performs short-distance wireless communication based on short-distance wireless communication protocols such as the Hilink protocol, WiFi (IEEE 802.11 protocol), Mesh, Bluetooth, ZigBee, Thread, Z-Wave, NFC, UWB, LiFi, etc. The communication device 1400 may also include a remote communication device, for example, any device that performs WLAN, GPRS, 2G / 3G / 4G / 5G remote communication. The display device 1500 is, for example, a liquid crystal display screen, a touch display screen, etc. The display device 1500 is used to display the collected remote sensing images. The input device 1600 may include, for example, a touch screen, a keyboard, etc. The user can input / output voice information through the speaker 1700 and the microphone 2800.
[0056] In this embodiment, the memory 1200 of the three-dimensional visualization device 100 of the target scenario is used to store instructions for controlling the processor 1100 to operate to at least execute the three-dimensional visualization method of the target scenario according to any embodiment of the present invention. Those skilled in the art can design the instructions according to the solutions disclosed in the present invention. How the instructions control the processor to operate is well known in the art and will not be described in detail here.
[0057] Although Figure 1 multiple devices of the three-dimensional visualization device 100 of the target scenario are shown, the present invention may only relate to some of the devices. For example, the three-dimensional visualization device 100 of the target scenario only relates to the memory 1200, the processor 1100, and the display device 1500.
[0058] In this embodiment, the three-dimensional visualization device 100 of the target scenario implements the method according to any embodiment of the present invention based on multiple first images of multiple shooting perspectives in the target scenario to obtain a three-dimensional visualization model of the target scenario.
[0059] <Method Embodiment>
[0060] Figure 2It is a schematic flowchart of a three-dimensional visualization method for a target scene according to an embodiment of the present invention, and this method can be implemented by a three-dimensional visualization device 2000 for the target scene.
[0061] According to Figure 2 As shown, the three-dimensional visualization method for the target scene in this embodiment may include the following steps S2100 to S2500:
[0062] Step S2100, obtain a plurality of first images captured for the target scene.
[0063] In this embodiment, the target scene may be an urban scene, a natural landscape, an indoor environment, a historical site, etc., which is not limited here.
[0064] The plurality of first images are obtained by capturing the target scene from a plurality of shooting perspectives, and each shooting perspective in the plurality of shooting perspectives may correspond to at least one first image. Among them, the plurality of first images may be images captured for the target scene in advance, or images obtained through a third-party platform, which is not limited here.
[0065] The shooting perspective may be the perspective information of the camera when shooting the target scene.
[0066] In some examples, the perspective information includes camera pose information.
[0067] In other examples, the perspective information may include information such as camera pose information, field of view angle, focal length, aperture, exposure parameters, etc.
[0068] Those skilled in the art should understand that the perspective information is not specifically limited here.
[0069] Step S2200, construct a three-dimensional Gaussian volume set for the target scene according to the plurality of first images.
[0070] In this embodiment, the target scene is described by a three-dimensional Gaussian volume set obtained from a plurality of first images. Among them, the three-dimensional Gaussian volume set includes a plurality of three-dimensional Gaussian volumes, and each three-dimensional Gaussian volume in the plurality of three-dimensional Gaussian volumes corresponds to a set of Gaussian volume parameters. The set of Gaussian volume parameters includes: the central position coordinates of the Gaussian volume. The central position coordinates of the Gaussian volume are used to characterize the position of the Gaussian volume in space.
[0071] In some examples, the set of Gaussian volume parameters includes: the central position coordinates of the Gaussian volume, the scale parameter of the Gaussian volume, the color information of the Gaussian volume, and the rotation matrix of the Gaussian volume. Among them, the scale parameter of the Gaussian volume is used to characterize the width, height, and depth of the Gaussian volume, that is, the "size" of the Gaussian volume in three-dimensional space. The rotation matrix of the Gaussian volume is used to characterize the direction of the Gaussian volume relative to the coordinate system. The color information of the Gaussian volume is used to characterize the color of the Gaussian volume.
[0072] In some embodiments, constructing the three-dimensional Gaussian volume set of the target scene according to the multiple first images in step S2200 includes: steps S2200.1 to S2200.3.
[0073] Step S2200.1: Construct a sparse point cloud of the target scene according to the camera pose information matched with each first image in the multiple first images.
[0074] In this embodiment, the camera pose information may be the position and orientation of the camera in three-dimensional space, which is represented by a position vector (representing the position of the camera in the world coordinate system) and a rotation matrix (or Euler angles, quaternions, etc., representing the direction of the camera relative to the world coordinate system). The camera pose information is crucial for determining how the camera observes and records the target scene.
[0075] The sparse point cloud of the target scene is a set of key points extracted from multiple first images for characterizing the target scene.
[0076] This step can be implemented by feature point matching and Structure from Motion (SfM) algorithms. That is, the SfM algorithm can extract feature points from each of the multiple first images, calculate the camera pose information corresponding to the first image, and then construct a three-dimensional sparse point cloud of the target scene according to the camera pose information matched with each first image in the multiple first images.
[0077] Step S2200.2: Align the sparse point cloud according to the edge features of the entity objects included in the multiple first images to obtain an aligned sparse point cloud.
[0078] In this embodiment, the target scene includes multiple entity objects.
[0079] In an example where the target scene is a city, the multiple entity objects included in the target scene may be: buildings, transportation vehicles, roads and paths, artificial landscapes, people, animals, plants, etc.
[0080] The multiple entity objects of the target scene may include the same edge features, different edge features, or some entity objects among the multiple entities may have the same edge features while the other part of the entity objects have different edge features, which is not limited here.
[0081] The edge features of the multiple entity objects in the target scene may be parallel edge features, vertical edge features, curved edge features, broken line edge features, etc., which is not limited here.
[0082] When performing this step, the edge features of the entity objects included in each of the multiple first images can be extracted, and then, based on certain specific edge features among the extracted edge features, the sparse point cloud can be aligned to obtain a processed sparse point cloud.
[0083] In some embodiments, in step S2200.2, aligning the sparse point cloud according to the edge features of the entity objects included in the multiple first images to obtain an aligned sparse point cloud includes: steps S2200.21 to S2200.23.
[0084] Step S2200.21, extracting the edge features of the entity objects included in the multiple first images.
[0085] In this embodiment, the extracted edge features of the entity objects can be one or more of parallel edge features, vertical edge features, curved edge features, polyline edge features, etc., and are not limited herein.
[0086] In this step, extracting the edge features of the entity objects in the first image can be implemented by an edge detection algorithm, such as the Canny edge detector, or by using more advanced computer vision techniques, such as a deep learning model, to identify edges, and is not limited herein.
[0087] Step S2200.22, in the case where the edge features of the entity objects include parallel edge features and vertical edge features, determining a horizontal plane corresponding to the parallel edge features in the sparse point cloud according to the parallel edge features, and determining a vertical plane corresponding to the vertical edge features in the sparse point cloud according to the vertical edge features.
[0088] In this embodiment, the parallel edge feature can be a feature indicating that the entity object in the target scene is parallel to the ground. In an example where the corresponding target scene is an urban scene, the entity objects having parallel edge features can be, for example, the walls of buildings, columns, lamp posts, etc.
[0089] The vertical edge feature can be a feature indicating that the entity object in the target scene is perpendicular to the ground. In an example where the corresponding target scene is an urban scene, the entity objects having vertical edge features can be, for example, the roofs of buildings, roads, bridge decks, etc.
[0090] When the edge features of the entity object include parallel edge features and vertical edge features, the sparse point cloud can be aligned by the Manhattan alignment technique. The alignment process of the Manhattan alignment technique for the sparse point cloud can be as follows: First, according to the parallel edge features, local plane region detection is performed in the sparse point cloud to identify the plane parallel to the ground as the horizontal plane corresponding to the parallel edge features, and according to the vertical edge features, local plane region detection is performed in the sparse point cloud to identify the plane perpendicular to the ground as the vertical plane corresponding to the vertical edge features. Then, step S2200.23 is executed to obtain the aligned sparse point cloud.
[0091] Among them, the local plane region detection can be implemented by a clustering algorithm, a RANSAC (Random Sample Consensus) algorithm, or other statistical methods, which are not limited here.
[0092] Those skilled in the art should understand that the local plane region detection is common general knowledge in the art and will not be elaborated here.
[0093] In step S2200.23, parallel alignment processing is performed on the horizontal plane points of the horizontal plane, and vertical alignment processing is performed on the vertical plane points of the vertical plane to obtain the aligned sparse point cloud.
[0094] In this embodiment, the horizontal plane points are the points in the sparse point cloud that belong to the horizontal plane. The vertical plane points are the points in the sparse point cloud that belong to the vertical plane.
[0095] That is to say, the parallel alignment processing in this step is for the points (i.e., horizontal plane points) in the sparse point cloud that are identified as belonging to the horizontal plane (such as the ground, the desktop, etc.). If the horizontal plane points are not exactly located on the horizontal plane, these horizontal plane points will be adjusted by moving them up and down to make them completely aligned with the horizontal plane (i.e., make them completely perpendicular to the Z-axis vertical axis).
[0096] The vertical alignment processing in this step is for the points (i.e., vertical plane points) in the sparse point cloud that are identified as belonging to the vertical plane (such as the wall surface, the column surface, etc.). If these vertical plane points are not exactly located on the vertical plane, these points will be adjusted by moving them left and right to make them completely aligned with the vertical plane, and the positions of the vertical plane points will be adjusted to make them located in the positive direction of the Z-axis vertical axis.
[0097] According to the embodiments of the present application, when the edge features of the entity object include parallel edge features and vertical edge features, according to the parallel edge features, a horizontal plane corresponding to the parallel edge features is determined in the sparse point cloud, and according to the vertical edge features, a vertical plane corresponding to the vertical edge features is determined in the sparse point cloud. The horizontal plane points of the horizontal plane are subjected to parallel alignment processing, and the vertical plane points of the vertical plane are subjected to vertical alignment processing to obtain an aligned sparse point cloud. In this way, the aligned sparse point cloud can accurately reflect the spatial relationship of the entity object in the target scene, facilitating more accurate three-dimensional modeling of the target scene.
[0098] Step S2200.3, construct a three-dimensional Gaussian volume set of the target scene according to the aligned sparse point cloud.
[0099] In this embodiment, the center position of each point in the aligned sparse point cloud can be used as the center position of the three-dimensional Gaussian volume to initialize a three-dimensional Gaussian volume. Moreover, during the process of initializing a three-dimensional Gaussian volume at the position of each point in the aligned sparse point cloud, information such as the scale parameter of the Gaussian volume, the color information of the Gaussian volume, the selection attribute of the Gaussian volume, and the opacity of the Gaussian volume also needs to be considered.
[0100] The color information of the Gaussian volume can be obtained according to the color information of the point corresponding to the center position of the Gaussian volume in the aligned sparse point cloud, or according to the color information extracted from the first image, or calculated according to a certain color estimation algorithm.
[0101] The scale parameter of the Gaussian volume is used to characterize the width, height, and depth of the Gaussian volume, that is, the "size" of the Gaussian volume in three-dimensional space. The scale parameter of the Gaussian volume can be fixed or dynamically adjusted according to the local density of the points in the aligned sparse point cloud. For example, if the points in the aligned sparse point cloud are relatively dense, a smaller-scale Gaussian volume can be used to represent the local structure more precisely, and if the points in the aligned sparse point cloud are relatively sparse, a larger-scale Gaussian volume can be used to cover a wider area.
[0102] The selection attribute of the Gaussian volume can be whether the user or the algorithm selects to include this Gaussian volume in visualization or further processing. In some applications, it may be necessary to decide whether to retain the Gaussian volume according to specific criteria (such as the scale parameter of the Gaussian volume, the center position of the Gaussian volume, or the degree of overlap with other Gaussian volumes).
[0103] The opacity of the Gaussian volume can be the opposite of the transparency of the Gaussian volume in visualization, which determines the visible degree of the Gaussian volume during rendering. The opacity can be used to represent the uncertainty or importance of the Gaussian volume. For example, a Gaussian volume with a lower opacity may represent a lower confidence level or importance.
[0104] After initializing a three-dimensional Gaussian volume at each point position of the aligned sparse point cloud, a set of three-dimensional Gaussian volumes of the target scene can be obtained.
[0105] Through the above steps S2100 and S2200, a set of three-dimensional Gaussian volumes of the target scene can be constructed. Moreover, each three-dimensional Gaussian volume in this set contains a detailed description of the target scene in a certain area, so that various details in the complex target scene can be represented more precisely, making it perform well in the reconstruction and rendering of complex large scenes (such as urban scenes). And by constructing a set of three-dimensional Gaussian volumes, the dependence on a large amount of manually labeled data in traditional methods can be reduced, the cost of data preparation can be lowered, and the efficiency of three-dimensional reconstruction is higher.
[0106] After obtaining the set of three-dimensional Gaussian volumes, in the related art, images from different viewpoints are often generated through the three-dimensional Gaussian splatting (3DGauSianSplatting, 3DGS) technology, that is, the set of three-dimensional Gaussian volumes is directly processed through view transformation and volume rendering methods to generate images from different viewpoints. However, this 3DGS technology mainly focuses on single objects or small-scale scenes. If 3DGS is directly applied to large-scale scenes, such as urban scenes, natural landscape scenes, etc., the construction of its three-dimensional visualization model will cause excessive memory overhead of the graphics processing unit (GPU), the efficiency of three-dimensional reconstruction is low, and there are also serious deficiencies in the detail fidelity of the constructed three-dimensional visualization model.
[0107] For the above reasons, the embodiments of the present application provide a method for distributed reconstruction of a three-dimensional visualization model, that is, the three-dimensional visualization model of the target scene is constructed through the following steps S2300 to S2500. The method for distributed reconstruction of the three-dimensional visualization model will be described in detail below in combination with steps S2300 to S2500.
[0108] Step S2300, divide the set of three-dimensional Gaussian volumes into multiple first subsets according to the central position coordinates corresponding to each three-dimensional Gaussian volume in the set of three-dimensional Gaussian volumes.
[0109] In this embodiment, the set of three-dimensional Gaussian volumes can be divided into multiple first subsets through various three-dimensional space partitioning methods.
[0110] These partitioning methods can be: grid partitioning, K-Means clustering, density-based clustering methods, octree, etc., which are not limited here.
[0111] In one example, multiple first subsets can be obtained through meshing. Meshing can be dividing the space occupied by the three-dimensional Gaussian body set into a regular grid, with each grid cell corresponding to a subspace. According to the coordinates of the centers of the Gaussian bodies, multiple three-dimensional Gaussian bodies in the three-dimensional Gaussian body set are assigned to the corresponding grid cells, thereby obtaining a first subset corresponding to each grid cell.
[0112] In another example, multiple first subsets can be obtained by partitioning based on a density-based clustering method. This method can be dividing the three-dimensional Gaussian body set into multiple first subsets with the same density according to the distribution density of the three-dimensional Gaussian bodies.
[0113] Since there will be a problem of low three-dimensional reconstruction efficiency due to the limitations of the calculation and storage of the graphics processor when directly partitioning the three-dimensional Gaussian body set in space, therefore, to avoid the problem of low three-dimensional reconstruction efficiency caused by the storage and calculation limitations of the graphics processor, the three-dimensional Gaussian body set can be partitioned according to the projection of the three-dimensional Gaussian body set on the horizontal plane.
[0114] Based on this, in some embodiments, in step S2300, dividing the three-dimensional Gaussian body set into multiple first subsets according to the central position coordinates corresponding to each three-dimensional Gaussian body in the three-dimensional Gaussian body set includes: step S2300.1 and step S2300.2.
[0115] Step S2300.1, projecting the central position coordinates corresponding to each three-dimensional Gaussian body onto the horizontal plane to obtain the central position projection coordinates of each three-dimensional Gaussian body.
[0116] In this embodiment, since the central position coordinates of the three-dimensional Gaussian body can reflect the position of the three-dimensional Gaussian body in space, the central position coordinates corresponding to each three-dimensional Gaussian body in the three-dimensional Gaussian body set can be projected onto the horizontal plane to obtain the central position projection coordinates of each three-dimensional Gaussian body.
[0117] The horizontal plane is the plane formed by the X-axis and the Y-axis, where the X-axis, Y-axis, and Z-axis are perpendicular to each other in pairs.
[0118] Step S2300.2, dividing the three-dimensional Gaussian body set into multiple first subsets according to the central position projection coordinates of each three-dimensional Gaussian body.
[0119] In this embodiment, multiple three-dimensional Gaussian body central position projection coordinates can be divided through various plane partitioning methods, thereby realizing the partitioning of the three-dimensional Gaussian body set.
[0120] The various plane partitioning methods can be, for example, checkerboard partitioning, adaptive meshing, quadtrees, etc., which are not limited here.
[0121] According to the embodiments of the present application, the central position coordinates corresponding to each three-dimensional Gaussian body can be projected onto the horizontal plane to obtain the central position projection coordinates of each three-dimensional Gaussian body. According to the central position projection coordinates of each three-dimensional Gaussian body, the three-dimensional Gaussian body set is divided into multiple first subsets. This horizontal plane projection division method can adapt to target scenarios of different scales and complexities. Moreover, for large-scale target scenarios, dividing the Gaussian body set in three-dimensional space may encounter computational and storage limitations, while this two-dimensional projection division of horizontal plane projection can effectively process large-scale data. And the horizontal plane projection division is more convenient for realizing parallel and independent processing of each sub-region on the horizontal plane, improving the efficiency of three-dimensional model reconstruction.
[0122] In some embodiments, in step S2300.2, dividing the three-dimensional Gaussian body set into multiple first subsets according to the central position projection coordinates of each three-dimensional Gaussian body includes: steps S2300.21 to S2300.23.
[0123] In step S2300.21, the projection region of the three-dimensional Gaussian body set on the horizontal plane is divided into multiple first sub-regions of equal size at a first set region interval.
[0124] In this embodiment, the first set region interval can be the size of each first sub-region (i.e., the spatial range size of each first sub-region) or the distance between adjacent first sub-regions.
[0125] The first set region interval can be determined according to the size of the projection region of the three-dimensional Gaussian body set on the horizontal plane.
[0126] For example, if the projection region of the three-dimensional Gaussian body set on the horizontal plane is 10 square meters and the first set region interval can be 1 square meter, then 10 first sub-regions can be obtained.
[0127] There is a set proportion of boundary overlap regions between adjacent first sub-regions among the multiple first sub-regions. Among them, the boundary overlap region is the region where the boundaries of two adjacent first sub-regions overlap.
[0128] The set proportion can refer to the proportion of the boundary overlap region to the first sub-region.
[0129] In one example, the set proportion can be 10%.
[0130] Those skilled in the art should understand that the specific size of the set proportion and the size of the first set region interval are not limited here.
[0131] In one example, the projection area of the three-dimensional Gaussian body on the horizontal plane can be divided into 36 first sub-regions of 6×6 by the checkerboard division method. Moreover, there is a 10% overlapping area between adjacent first sub-regions. That is to say, if a first sub-region is region A and its adjacent first sub-region on the right is region B, then 10% of the area between region A and region B is shared by both.
[0132] Step S2300.22: According to the central position projection coordinates of each three-dimensional Gaussian body, allocate each three-dimensional Gaussian body to the corresponding first sub-region.
[0133] In this embodiment, each first sub-region can be represented by a region coordinate interval, that is, a first sub-region can be described by its minimum and maximum coordinate values on the X-axis (i.e., the X coordinate interval) and its minimum and maximum coordinate values on the Y-axis (i.e., the Y coordinate interval).
[0134] At this time, the central position projection coordinates of each three-dimensional Gaussian body can be matched with the region coordinate interval corresponding to each first sub-region, so as to allocate the central position projection coordinates of each three-dimensional Gaussian body to the corresponding first sub-region.
[0135] Step S2300.23: For each first sub-region among the multiple first sub-regions, determine the first subset corresponding to the first sub-region according to the multiple three-dimensional Gaussian bodies allocated to the first sub-region, and obtain the first subset corresponding to each first sub-region among the multiple first sub-regions.
[0136] Exemplarily, the multiple first sub-regions are 16 regions. After performing step S2300.22 to allocate each three-dimensional Gaussian body in the three-dimensional Gaussian body set to the corresponding first sub-region, for any one of the 16 sub-regions, according to the multiple three-dimensional Gaussian bodies allocated to the sub-region, obtain the first subset corresponding to the sub-region, so as to obtain 16 first subsets corresponding to the 16 first sub-regions.
[0137] By evenly dividing the projection area of the three-dimensional Gaussian body on the horizontal plane into multiple first sub-regions, the number of Gaussian bodies in each first sub-region can be balanced to a certain extent, so that the subsequent data processing for each first subset is more balanced. Moreover, each first subset obtained by dividing the Gaussian body set into multiple first subsets can be processed independently, such as local refinement, feature extraction or visualization, and the multiple first subsets can be allocated to different computing units for parallel processing, thereby greatly improving the three-dimensional reconstruction efficiency.
[0138] Step S2400: For each of the multiple first subsets, optimize the first subset according to at least one of the first images associated with the first subset, to obtain multiple optimized first subsets corresponding to the multiple first subsets.
[0139] In this embodiment, the multiple first subsets can be assigned to different computing units for optimization processing, to obtain the optimized first subsets corresponding to each of the multiple first subsets. In this way, independent optimization processing of each first subset can be achieved, and moreover, the multiple first subsets can be optimized in parallel.
[0140] In some embodiments, in step S2400, optimizing the first subset according to at least one of the first images associated with the first subset includes:
[0141] Optimize the first subset by using the three-dimensional Gaussian splashing technique according to at least one of the first images associated with the first subset.
[0142] In this embodiment, at least one first image associated with the first subset can be used to optimize parameters such as the position, size parameters, orientation information, and color information of the Gaussian bodies in the first subset, so as to better fit the three-dimensional visualization model constructed from the multiple first images.
[0143] Those skilled in the art should understand that the process of optimizing the first subset by using the three-dimensional Gaussian splashing technique according to at least one of the first images associated with the first subset is well-known in the art, and the specific process will not be elaborated here.
[0144] Through the 3DGS technique, the first images from multiple shooting perspectives can be fully utilized to comprehensively optimize the Gaussian bodies in the first subset, thereby generating a more accurate and detailed three-dimensional model.
[0145] In some embodiments, the step of determining at least one of the first images associated with the first subset includes: step S3100 to step S3300.
[0146] Step S3100: According to the multiple shooting perspectives, project the first subset onto the image plane of each of the multiple shooting perspectives, to obtain the projection information of each shooting perspective.
[0147] In this embodiment, the process of projecting the first subset onto the image plane of each of the multiple shooting perspectives actually refers to converting the three-dimensional spatial coordinates of the multiple three-dimensional Gaussian bodies included in the first subset into corresponding two-dimensional pixel coordinates.
[0148] Here, the process of converting the three-dimensional spatial coordinates of a three-dimensional Gaussian body into the two-dimensional pixel coordinates of the three-dimensional Gaussian body is as follows: First, through the camera pose information (i.e., the external camera parameters), the three-dimensional space of the three-dimensional Gaussian body is transformed from the world coordinate system to the camera coordinate system, and the coordinates of the three-dimensional Gaussian body in the camera coordinate system are obtained. Then, through the internal parameters of the camera, the coordinates of the three-dimensional Gaussian body in the camera coordinate system are projected onto the image plane to obtain the two-dimensional pixel coordinates of the three-dimensional Gaussian body.
[0149] After a first subset is back-projected onto the image plane of a capture view, a set of two-dimensional Gaussian distributions can be obtained. Among them, one three-dimensional Gaussian body corresponds to one two-dimensional Gaussian distribution. This set of two-dimensional Gaussian distributions can be characterized by projection information.
[0150] In one example, the projection information includes the number of two-dimensional Gaussian distributions.
[0151] In this example, the number of two-dimensional Gaussian distributions can represent the number of two-dimensional Gaussian distributions projected by the first subset at a certain capture view.
[0152] In another example, the projection information includes: projection points, the number of two-dimensional Gaussian distributions, the colors of the two-dimensional Gaussian distributions, etc.
[0153] In this example, the projection points are the projection positions of the centers or key points of the three-dimensional Gaussian bodies on the two-dimensional image plane. The colors of the two-dimensional Gaussian distributions are the color information of the three-dimensional Gaussian bodies projected onto the image plane to form a color distribution map.
[0154] Step S3200, for each of the multiple capture views, determine whether the capture view and the first subset are strongly correlated according to the projection information corresponding to the capture view.
[0155] In some embodiments, the projection information includes the number of two-dimensional Gaussian distributions. In step S3200, determining whether the capture view and the first subset are strongly correlated according to the projection information corresponding to the capture view includes: step S3200.1 and step S3200.2.
[0156] Step S3200.1, determine the projection ratio of the capture view according to the number of two-dimensional Gaussian distributions in the projection information corresponding to the capture view and the number of three-dimensional Gaussian bodies included in the first subset.
[0157] Exemplarily, the shooting perspectives may include Perspective A, Perspective B, and Perspective C. For any first subset, the projection information of the first subset at these three shooting perspectives can be calculated. Among them, the projection information of the first subset at these three shooting perspectives includes the A projection information corresponding to Perspective A, the B projection information corresponding to Perspective B, and the C projection information corresponding to Perspective C. According to the number of two-dimensional Gaussian distributions in the A projection information corresponding to Perspective A and the number of Gaussian bodies included in the first subset, determine the projection ratio of Perspective A for the first subset. According to the number of two-dimensional Gaussian distributions in the B projection information corresponding to Perspective B and the number of Gaussian bodies included in the first subset, determine the projection ratio of Perspective B for the first subset. According to the number of two-dimensional Gaussian distributions in the C projection information corresponding to Perspective C and the number of Gaussian bodies included in the first subset, determine the projection ratio of Perspective C for the first subset.
[0158] In this embodiment, the projection ratio is the ratio of the number of two-dimensional Gaussian distributions obtained by projecting the first subset at the shooting perspective to the number of three-dimensional Gaussian bodies included in the first subset.
[0159] Step S3200.2, in the case where the projection ratio of the shooting perspective is greater than or equal to the projection ratio threshold, determine that the shooting perspective and the first subset are strongly correlated.
[0160] In this embodiment, the projection ratio threshold can be, for example, 1%, or other projection ratio thresholds can be set, which are not limited here.
[0161] Continuing with the above example, if the projection ratio of Perspective A for the first subset is greater than the projection ratio threshold of 1%, it is considered that Perspective A is important for the first subset, that is, Perspective A and the first subset are strongly correlated. If the projection ratio of Perspective B for the first subset is less than the projection ratio threshold of 1%, it is determined that Perspective B and the first subset are not strongly correlated. If the projection ratio of Perspective C for the first subset is greater than the projection ratio threshold of 1%, it is determined that Perspective C and the first subset are strongly correlated.
[0162] Step S3300, in the case where the shooting perspective and the first subset are strongly correlated, associate the first image corresponding to the shooting perspective with the first subset to obtain at least one first image associated with the first subset.
[0163] Continuing with the above example, Perspective A and Perspective C are strongly correlated with the first subset. At this time, the first image corresponding to Perspective A and the first image corresponding to Perspective C are associated with the first subset to obtain two first images associated with the first subset.
[0164] According to the embodiments of the present application, by setting at least one first image associated with the first subset as an image corresponding to the shooting perspective strongly associated with the first subset, this way of associating the first image with the first subset can, when optimizing the first subset, only optimize the first subset through the first images of the shooting perspectives strongly associated with it, which can achieve improving the optimization efficiency while ensuring the optimization effect.
[0165] Step S2500, determine the three-dimensional visualization model of the target scene according to the multiple optimized first subsets.
[0166] In this embodiment, the multiple optimized first subsets can be spliced according to the way of dividing the multiple first subsets in step S2300 to obtain the three-dimensional visualization model of the target scene.
[0167] In the example of obtaining multiple first subsets by dividing the horizontal plane into 36 first sub-regions of 6×6 by the checkerboard division method, and there is a 10% overlapping region between adjacent first sub-regions, the splicing of the 36 optimized first subsets in this step can be: according to the arrangement of the 36 first sub-regions corresponding to the 36 optimized first subsets, splice the 36 optimized first subsets corresponding to the 36 first sub-regions to obtain the three-dimensional visualization model of the target scene.
[0168] Since when dividing the three-dimensional Gaussian body set by the first set region interval, there may be some Gaussian bodies near the boundaries of the first sub-regions, and these Gaussian bodies located at the boundaries of the first sub-regions may be repeatedly calculated in multiple first sub-regions. To reduce the redundancy of repeated calculations, after obtaining multiple optimized first subsets, the horizontal plane can be re-divided into multiple second sub-regions at the second set region interval, and according to the central position projection coordinates of the multiple three-dimensional Gaussian bodies included in the multiple optimized first subsets, assign the corresponding three-dimensional Gaussian bodies to each second sub-region to obtain the second subset corresponding to each second sub-region. Then splice the multiple second subsets corresponding to the multiple second sub-regions to obtain the three-dimensional visualization model of the target scene.
[0169] In some embodiments, determining the three-dimensional visualization model of the target scene according to the multiple optimized first subsets in step S2500 includes: steps S2500.1 to S2500.3.
[0170] Step S2500.1, divide the projection region of the three-dimensional Gaussian body set on the horizontal plane into multiple second sub-regions of equal size at the second set region interval.
[0171] In this embodiment, there is no boundary overlapping area between adjacent second sub-regions among the multiple second sub-regions, and the second set region interval is greater than the first set region interval.
[0172] The method for dividing the multiple second sub-regions in this step is basically the same as that in step S2300.21, and will not be elaborated here.
[0173] Step S2500.2: According to the multiple second sub-regions, re-divide the multiple optimized first subsets to obtain a second subset corresponding to each second sub-region among the multiple second sub-regions.
[0174] In this embodiment, based on the projection coordinates of the centers of all three-dimensional Gaussian bodies included in the multiple optimized first subsets on the horizontal plane, each three-dimensional Gaussian body can be assigned to the corresponding second sub-region. Then, according to the multiple three-dimensional Gaussian bodies assigned to any second sub-region, determine the second subset corresponding to this second sub-region, and obtain the second subset corresponding to each second sub-region among the multiple second sub-regions.
[0175] Step S2500.3: According to the second subset corresponding to each second sub-region, determine the three-dimensional visualization model of the target scene.
[0176] In this embodiment, for the second subset corresponding to each second sub-region, according to the regional coordinate interval of this second sub-region, splice the multiple second subsets to obtain the three-dimensional visualization model of the target scene.
[0177] According to the embodiments of the present application, by constructing a set of three-dimensional Gaussian bodies of the target scene, and each three-dimensional Gaussian body in this set of three-dimensional Gaussian bodies contains a detailed description of the target scene in a certain area, various details in the complex target scene can be represented more precisely, and the dependence on a large amount of manually labeled data in the traditional method can be reduced, the cost of data preparation is reduced, and the efficiency of three-dimensional reconstruction is higher. By dividing the set of three-dimensional Gaussian bodies into multiple first subsets, and respectively performing optimization processing on each first subset among the multiple first subsets, after obtaining the optimized first subsets and then splicing them, the three-dimensional visualization model of the target scene can be realized. Moreover, the optimization of the multiple first subsets can be processed in parallel independently, so as to avoid the problem of excessive memory overhead of the graphics processing unit (GPU) when modeling large-scale scenes, improve the efficiency of three-dimensional modeling. In addition, by optimizing each first subset through at least one first image associated with it, the detail fidelity in large-scale scenes can also be improved.
[0178] For large-scale complex scenes, the main factor affecting the rendering speed of 3D visualization models is depth sorting. Among them, depth sorting refers to the process of arranging all 3D Gaussian bodies in the order from near to far during the rendering of 3D visualization models to ensure the correct visual effect. For large-scale complex scenes, the complexity of depth sorting will directly lead to a very slow rasterization process (i.e., the process of converting a 3D visualization model into a 2D image). That is to say, as the number of Gaussian bodies increases, the rendering performance will decrease significantly. When the number of Gaussian bodies increases to the millions level, the rendering process will become extremely slow. For this, the present application provides a method for improving rendering efficiency, which is as follows:
[0179] In some embodiments, after determining the 3D visualization model of the target scene in step S2500, the method further includes: step S4100 and step S4200.
[0180] Step S4100, obtaining and outputting a second image of the 3D visualization model from the target rendering perspective.
[0181] In this embodiment, the target rendering perspective can be at least one of the multiple shooting perspectives in step S2100, or a perspective completely different from the multiple shooting perspectives in step S2100, which is not limited here.
[0182] The second image is a low-resolution image.
[0183] Exemplarily, since the 3D visualization model is constructed based on 3D Gaussian bodies, when outputting a low-resolution second image from the target rendering perspective, a low-resolution second image can be obtained through 3DGS rendering.
[0184] Those skilled in the art should understand that the method of obtaining a low-resolution second image through 3DGS rendering is well-known in the art and will not be elaborated here.
[0185] Step S4200, performing super-resolution processing on the second image to obtain and output a third image from the target rendering perspective.
[0186] In this embodiment, the resolution of the third image is greater than that of the second image. The second image can be subjected to super-resolution processing through a super-resolution model (ESRGAN model), or the second image can be subjected to super-resolution processing through other means, which is not limited here.
[0187] ESRGAN (Enhanced Super-Resolution Generative Adversarial Networks) is a super-resolution model based on generative adversarial networks (GAN). It improves the performance and image quality of super-resolution through adversarial training. By inputting a low-resolution second image into the ESRGAN model, the generator attempts to generate a high-resolution third image. The discriminator compares the generated high-resolution third image with the real high-resolution image and provides feedback to the generator to improve its performance. Through iterative training, the generator learns how to generate higher-quality high-resolution images. The ESRGAN model can generate visually more natural high-resolution images, reducing the artifacts and blurring common in the super-resolution process. The model can better preserve the detail and texture information of the image, improving the clarity of the image.
[0188] According to the embodiments of the present application, by obtaining and outputting a low-resolution second image of the three-dimensional visualization model from the target rendering perspective, and performing super-resolution processing on the low-resolution second image to obtain and output a high-resolution third image, the efficiency and image quality of rendering large-scale complex scenes can be improved.
[0189] In some embodiments, the target rendering perspective can be multiple shooting perspectives in step S2100. Thus, after performing step S4100 and step S4200, a third image corresponding to each shooting perspective in the multiple shooting perspectives can be obtained. At this time, the three-dimensional visualization model of the target scene can also be evaluated according to the pixel error value between the third image corresponding to each shooting perspective and the first image.
[0190] In this embodiment, the PSNR metric can be used to represent the pixel error value between the third image and the first image of each shooting perspective.
[0191] PSNR is calculated based on MSE (Mean Squared Error). MSE measures the average squared error between the first image and the third image. Among them, the unit of PSNR is decibel (dB). The higher the value, the smaller the distortion of the third image compared to the first image, and thus the three-dimensional visualization model of the target scene is closer to the real target scene.
[0192] In order to verify the accuracy of the method in the above embodiments of the present application when constructing large-scale scenes, the inventor also tested through the above method in the Small City scene in MatrixCity. There are 5620 training images of different perspectives in this scene (i.e., multiple first images in step S2100). The following steps are performed on these 5620 images to obtain the three-dimensional visualization model corresponding to this scene:
[0193] Step S1, obtain 5,620 first images of the Small City scene.
[0194] Step S2, match camera pose information for each first image.
[0195] Step S3, construct a sparse point cloud of the SmallCity scene based on the camera pose information matched for each of the 5,620 first images.
[0196] Step S4, extract the edge features of the entity objects included in the 5,620 first images.
[0197] Step S5, in the case where the edge features of the entity objects include parallel edge features and vertical edge features, determine a horizontal plane corresponding to the parallel edge features in the sparse point cloud according to the parallel edge features, and determine a vertical plane corresponding to the vertical edge features in the sparse point cloud according to the vertical edge features.
[0198] Step S6, perform parallel alignment processing on the horizontal plane points of the horizontal plane and vertical alignment processing on the vertical plane points of the vertical plane to obtain an aligned sparse point cloud.
[0199] Step S7, construct a three-dimensional Gaussian volume set of the Small City scene based on the aligned sparse point cloud.
[0200] Step S8, project the central position coordinates corresponding to each three-dimensional Gaussian volume in the three-dimensional Gaussian volume set onto the horizontal plane to obtain the central position projection coordinates of each three-dimensional Gaussian volume.
[0201] In this example, the horizontal plane is the xy plane.
[0202] Step S9, divide the projection area of the three-dimensional Gaussian volume set on the horizontal plane into 36 equal-sized first sub-regions by the checkerboard division method at a first set region interval.
[0203] In this embodiment, there is a set proportion of boundary overlap regions between adjacent first sub-regions among the 36 first sub-regions.
[0204] Step S10, assign each three-dimensional Gaussian volume to the corresponding first sub-region according to the central position projection coordinates of each three-dimensional Gaussian volume.
[0205] Step S11, determine a first subset corresponding to the first sub-region according to the multiple three-dimensional Gaussian volumes assigned to the first sub-region, and obtain a first subset corresponding to each first sub-region among the 36 first sub-regions.
[0206] In this example, the 36 first sub-regions correspond to 36 first subsets.
[0207] Step S12: Project the first subset onto the image plane of each of the multiple shooting perspectives according to the multiple shooting perspectives, and obtain the projection information of each shooting perspective.
[0208] Step S13: For any one of the multiple shooting perspectives, determine the projection ratio of the shooting perspective according to the number of two-dimensional Gaussian distributions in the projection information of the shooting perspective and the number of three-dimensional Gaussian bodies included in the first subset.
[0209] Step S14: When the projection ratio of the shooting perspective is greater than or equal to the projection ratio threshold, determine that the shooting perspective and the first subset are strongly correlated.
[0210] Step S15: When the shooting perspective and the first subset are strongly correlated, associate the first image of the shooting perspective with the first subset to obtain at least one first image associated with the first subset.
[0211] Step S16: For each of the multiple first subsets, perform optimization processing on the first subset according to at least one first image associated with the first subset by using the three-dimensional Gaussian splashing technique to obtain multiple optimized first subsets corresponding to the multiple first subsets.
[0212] Step S17: Divide the projection area of the three-dimensional Gaussian body set on the horizontal plane into multiple second sub-regions of equal size at a second set region interval; wherein, there is no boundary overlapping region between adjacent second sub-regions among the multiple second sub-regions, and the second set region interval is greater than the first set region interval.
[0213] Step S18: Re-divide the multiple optimized first subsets according to the multiple second sub-regions to obtain a second subset corresponding to each second sub-region among the multiple second sub-regions.
[0214] Step S19: Determine the three-dimensional visualization model of the Small City scene according to the second subset corresponding to each second sub-region.
[0215] After obtaining the three-dimensional visualization model of the Small City scene, test it with 740 test images from different perspectives. For the convenience of description, the 740 test images can be the real images of the Small City scene under 740 test perspectives. For comparison, it is also necessary to obtain the rendered images under these 740 test perspectives. Specifically, the rendered images of the 740 test perspectives can be obtained through the following steps.
[0216] Step S20: Obtain and output the second image of the three-dimensional visualization model under each of the 740 test perspectives according to the 740 test perspectives.
[0217] In step S21, super-resolution processing is performed on 740 second images, and 740 rendered images are obtained and output.
[0218] After that, the PSNR value between the rendered image and the real image corresponding to each test view among the 740 test views is calculated, and the obtained PSNR reaches 27.84, indicating that the three-dimensional visualization model of the Small City scene has high accuracy.
[0219] <Embodiment of storage medium>
[0220] An embodiment of the present application provides a readable storage medium, on which a program or instruction is stored, and when the program or instruction is executed by a processor, the steps of the method described in any of the above method embodiments are implemented.
[0221] <Embodiment of device>
[0222] Figure 3 It is a structural block diagram of a three-dimensional visualization device 3000 for a target scene according to an embodiment of the present invention.
[0223] In this embodiment, as Figure 3 shown, the three-dimensional visualization device 3000 for the target scene includes a memory 3001 and a processor 3002. The memory 3001 is used to store executable instructions, and the processor 3002 is used to operate according to the control of the instructions to execute the method described in any of the above embodiments.
[0224] The present invention may be a system, a method, and / or a computer program product. The computer program product may include a computer-readable storage medium having thereon computer-readable program instructions for causing a processor to implement various aspects of the present invention.
[0225] A computer-readable storage medium can be a tangible device that can hold and store instructions for use by an instruction execution device. A computer-readable storage medium may be, for example—but not limited to—an electrical storage device, a magnetic storage device, an optical storage device, an electromagnetic storage device, a semiconductor storage device, or any suitable combination of the foregoing. More specific examples (a non-exhaustive list) of the computer-readable storage medium include: a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), a static random access memory (SRAM), a portable compact disc read-only memory (CD-ROM), a digital versatile disc (DVD), a memory stick, a floppy disk, a mechanically encoded device such as a punched card or raised structures in grooves having instructions stored thereon, and any suitable combination of the foregoing. The computer-readable storage medium as used herein is not construed to be an instantaneous signal per se, such as a radio wave or other freely propagating electromagnetic wave, an electromagnetic wave propagated through a waveguide or other transmission medium (e.g., an optical pulse through an optical fiber cable), or an electrical signal transmitted through a wire.
[0226] The computer-readable program instructions described herein can be downloaded from a computer-readable storage medium to various computing / processing devices, or downloaded to an external computer or external storage device via a network, such as the Internet, a local area network, a wide area network, and / or a wireless network. The network may include a copper transmission cable, an optical fiber transmission, a wireless transmission, a router, a firewall, a switch, a gateway computer, and / or an edge server. A network adapter card or network interface in each computing / processing device receives the computer-readable program instructions from the network and forwards the computer-readable program instructions for storage in a computer-readable storage medium in each computing / processing device.
[0227] The computer program instructions for performing the operations of the present invention may be assembly instructions, instruction set architecture (ISA) instructions, machine instructions, machine - related instructions, microcode, firmware instructions, state - setting data, or source code or object code written in any combination of one or more programming languages, including object - oriented programming languages such as Smalltalk, C++, etc., and conventional procedural programming languages such as the "C" language or similar programming languages. The computer - readable program instructions may be executed entirely on the user's computer, partially on the user's computer, executed as a stand - alone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the case of a remote computer, the remote computer may be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., through the Internet using an Internet service provider). In some embodiments, by using the state information of the computer - readable program instructions to customize an electronic circuit, such as a programmable logic circuit, a field - programmable gate array (FPGA), or a programmable logic array (PLA), the electronic circuit can execute the computer - readable program instructions to implement various aspects of the present invention.
[0228] Aspects of the present invention are described herein with reference to the flowcharts and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the present invention. It should be understood that each block of the flowcharts and / or block diagrams, and combinations of blocks in the flowcharts and / or block diagrams, can be implemented by computer - readable program instructions.
[0229] These computer - readable program instructions can be provided to a processor of a general - purpose computer, a special - purpose computer, or other programmable data - processing apparatus to produce a machine such that the instructions, when executed by the processor of the computer or other programmable data - processing apparatus, create a means for implementing the functions / actions specified in one or more blocks of the flowchart and / or block diagram. These computer - readable program instructions can also be stored in a computer - readable storage medium, which causes a computer, a programmable data - processing apparatus, and / or other devices to operate in a particular manner, so that the computer - readable medium storing the instructions includes a manufacture, which includes instructions for implementing various aspects of the functions / actions specified in one or more blocks of the flowchart and / or block diagram.
[0230] Computer-readable program instructions may also be loaded onto a computer, other programmable data processing apparatus, or other device, causing a series of operational steps to be performed on the computer, other programmable data processing apparatus, or other device to produce a computer-implemented process such that the instructions executed on the computer, other programmable data processing apparatus, or other device implement the functions / acts specified in one or more blocks of the flowchart and / or block diagram.
[0231] The flowcharts and block diagrams in the figures illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of the present invention. In this regard, each block in the flowchart or block diagram may represent a module, a segment of a program, or a portion of an instruction, which contains one or more executable instructions for implementing the specified logical function. In some alternative implementations, the functions noted in the blocks may occur out of the order noted in the figures. For example, two consecutive blocks may in fact be executed substantially in parallel, or they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block of the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented by special purpose hardware-based systems that perform the specified functions or acts, or by combinations of special purpose hardware and computer instructions. As is well known to those skilled in the art, implementations by hardware, by software, and by the combination of software and hardware are equivalent.
[0232] The embodiments of the present invention have been described above. The above description is exemplary, not exhaustive, and is not limited to the disclosed embodiments. Many modifications and variations will be apparent to those of ordinary skill in the art without departing from the scope and spirit of the described embodiments. The choice of terms used herein is intended to best explain the principles of the embodiments, the practical application, or improvements made to the technology in the market, or to enable other ordinary skilled persons in the art to understand the embodiments disclosed herein. The scope of the present invention is defined by the appended claims.
Claims
1. A three-dimensional visualization method for a target scene, characterized in that: The method comprises: Acquire a plurality of first images obtained by photographing a target scene; wherein the plurality of first images are obtained by photographing the target scene at a plurality of shooting angles; Constructing a three-dimensional Gaussian volume set of the target scene according to the multiple first images; wherein the three-dimensional Gaussian volume set includes multiple three-dimensional Gaussian volumes, each of the multiple three-dimensional Gaussian volumes corresponds to a set of Gaussian volume parameters, and the Gaussian volume parameters include the coordinates of the center position of the Gaussian volume; Dividing the three-dimensional Gaussian body set into a plurality of first subsets according to the center position coordinates corresponding to each three-dimensional Gaussian body in the three-dimensional Gaussian body set; For each first subset of the multiple first subsets, optimizing the first subset according to at least one first image associated with the first subset, to obtain multiple optimized first subsets corresponding to the multiple first subsets; A three-dimensional visualization model of the target scene is determined according to the multiple optimized first subsets.
2. The method according to claim 1, characterized in that The step of constructing a set of three-dimensional Gaussian volumes of the target scene according to the plurality of first images comprises: Constructing a sparse point cloud of the target scene according to camera pose information matched to each first image in the plurality of first images; Performing alignment processing on the sparse point cloud according to edge features of the entity objects included in the plurality of first images to obtain an aligned sparse point cloud; A three-dimensional Gaussian volume set of the target scene is constructed according to the aligned sparse point cloud.
3. The method according to claim 2, characterized in that The aligning process is performed on the sparse point cloud according to the edge features of the entity objects included in the plurality of first images to obtain the aligned sparse point cloud, including: Extracting edge features of the entity objects included in the plurality of first images; In the case where the edge features of the physical object include parallel edge features and vertical edge features, determining a horizontal plane corresponding to the parallel edge features in the sparse point cloud according to the parallel edge features, and determining a vertical plane corresponding to the vertical edge features in the sparse point cloud according to the vertical edge features; The horizontal plane points of the horizontal plane are parallelly aligned, and the vertical plane points of the vertical plane are vertically aligned to obtain the aligned sparse point cloud.
4. The method according to claim 1, characterized in that: The step of dividing the three-dimensional Gaussian body set into a plurality of first subsets according to the center position coordinates corresponding to each three-dimensional Gaussian body in the three-dimensional Gaussian body set includes: Projecting the center position coordinates corresponding to each three-dimensional Gaussian body onto a horizontal plane to obtain the center position projection coordinates of each three-dimensional Gaussian body; The three-dimensional Gaussian volume set is divided into a plurality of first subsets according to the projection coordinates of the center position of each three-dimensional Gaussian volume.
5. The method according to claim 4, characterized in that The step of dividing the three-dimensional Gaussian body set into a plurality of first subsets according to the projection coordinates of the center position of each three-dimensional Gaussian body comprises: The projection area of the three-dimensional Gaussian volume set on the horizontal plane is divided into a plurality of first sub-areas of equal size at first set area intervals; wherein, there are boundary overlap areas of a set ratio between adjacent first sub-areas in the plurality of first sub-areas; Allocating each of the three-dimensional Gaussian bodies to a corresponding first sub-region according to the projection coordinates of the center position of each of the three-dimensional Gaussian bodies; For each first sub-region among the multiple first sub-regions, a first subset corresponding to the first sub-region is determined according to multiple three-dimensional Gaussian volumes allocated to the first sub-region, so as to obtain multiple first subsets corresponding to the multiple first sub-regions.
6. The method according to claim 1, characterized in that The step of determining at least one of the first images associated with the first subset comprises: According to the multiple shooting angles, projecting the first subset onto an image plane of each shooting angle among the multiple shooting angles to obtain projection information of each shooting angle; For each of the shooting angles, determining, according to projection information corresponding to the shooting angle, whether the shooting angle is strongly associated with the first subset; In the case where the shooting angle of view is strongly associated with the first subset, the first image corresponding to the shooting angle of view is associated with the first subset to obtain at least one first image associated with the first subset.
7. The method according to claim 6, characterized in that The projection information includes a two-dimensional Gaussian distribution quantity, and determining whether the shooting angle is strongly associated with the first subset according to the projection information corresponding to the shooting angle includes: Determining a projection ratio of the shooting angle of view according to the number of two-dimensional Gaussian distributions in the projection information corresponding to the shooting angle of view and the number of three-dimensional Gaussian volumes included in the first subset; When the projection ratio of the shooting angle of view is greater than or equal to a projection ratio threshold, it is determined that the shooting angle of view is strongly associated with the first subset.
8. The method according to claim 1, characterized in that The step of determining the three-dimensional visualization model of the target scene according to the plurality of optimized first subsets includes: Divide the projection area of the three-dimensional Gaussian volume set on the horizontal plane into a plurality of second sub-areas of equal size at a second set area interval; wherein there is no border overlap between adjacent second sub-areas in the plurality of second sub-areas, and the second set area interval is greater than the first set area interval; Re-dividing the plurality of optimized first subsets according to the plurality of second sub-regions to obtain a second subset corresponding to each second sub-region in the plurality of second sub-regions; A three-dimensional visualization model of the target scene is determined according to the second subset corresponding to each second sub-area.
9. The method according to claim 1, characterized in that: After determining the three-dimensional visualization model of the target scene, the method further includes: According to the target rendering perspective, obtaining and outputting a second image of the three-dimensional visualization model at the target rendering perspective; The second image is subjected to super-resolution processing to obtain and output a third image under the target rendering perspective.
10. A three-dimensional visualization device for a target scene, comprising a memory and a processor, wherein the memory is used to store executable instructions; and the processor is used to operate according to the control of the instructions to execute the method as claimed in any one of claims 1 to 9.