A method for high-fidelity three-dimensional modeling of large-scale scenes with unbounded expansion

Through a large-scene three-dimensional modeling method that can be expanded without boundaries, using drone data and 3D Gaussian Splatting technology, the problems of limited model capacity and long rendering time in the existing technology are solved, and borderless city-level high-fidelity three-dimensional modeling and fast rendering effects are achieved.

CN118608706BActive Publication Date: 2025-06-13ZHONGKE XINGTU DIGITAL EARTH HEFEI CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202410653406.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-05-24
Publication Date
2025-06-13
Estimated Expiration
2044-05-24

AI Technical Summary

Technical Problem

When the existing technology uses drones for urban-level three-dimensional reconstruction, the model capacity is limited and large-scale urban-level scenarios cannot be trained. The training and rendering time is too long, which can easily lead to memory overflow.

Method used

A three-dimensional modeling method for high-resistance real scenes that can be expanded without boundaries is adopted. By collecting multi-view image data of the drone, using colmap to generate camera pose maps and scene point cloud data, it is processed in chunks and passed in 3D Gaussian Splating for parallel training, and combined with fast α-blending rendering technology, it achieves fast and visually realistic rendering effects.

Benefits of technology

It supports the construction of borderless city-level large scenarios, and can be trained in parallel between scenes to save computing resources and time. After blocking, it is integrated on the browser and rendered in real time with a boundary high-remaining real scene. The algorithm is characterized by high stability and visualization.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118608706B_ABST
    Figure CN118608706B_ABST
Patent Text Reader

Abstract

The present invention discloses a method for high-fidelity three-dimensional modeling of large-scale scenes with unbounded expansion, comprising the following steps: S1, collecting multi-view image data of unmanned aerial vehicles; S2, inputting the collected data into colmap to generate the pose graph of cameras and the scene point cloud data; S3, dividing the entire scene into multiple small blocks, calculating the centroid of each block, and obtaining the point cloud and camera poses of different blocks after scene division; S4, transmitting the processed data information into 3DGS for parallel training to obtain the models of each different block; S5, performing fast rendering; S6, rendering and fusing the trained blocks on the front-end browser to obtain a three-dimensional scene from any perspective; S7, adjusting the height of the front-end camera to obtain three-dimensional scenes with different precisions. The present invention supports the construction of unbounded city-level large-scale scenes, and parallel training can be performed between scenes, saving computing resources and time. The algorithm has the characteristics of high stability and being visualizable.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the fields of computer vision and computer graphics, and particularly relates to a method for high-fidelity three-dimensional modeling of large scenes that can be infinitely extended. Background Art

[0002] City-level three-dimensional modeling is a technology for digitally representing the geographical information of a city or urban area and its surrounding environment. It can simulate the appearance, structure, and terrain of a city by creating accurate three-dimensional models, providing support for fields such as urban planning, land management, traffic planning, architectural design, and visual simulation. The use of three-dimensional modeling technology can provide technical support for the development of digital twins and the metaverse, promoting the construction of smart cities and the digital economy.

[0003] Currently, technologies for three-dimensional reconstruction using drones include SLAM, SFM, NeRF, etc. However, they all have some problems.

[0004] (1) The model capacity is limited and cannot train large-scale city-level scenes.

[0005] (2) The training and rendering time is too long and may cause memory overflow. Summary of the Invention

[0006] To solve the existing problems, the present invention provides a method for high-fidelity three-dimensional modeling of large scenes that can be infinitely extended. The specific solution is as follows:

[0007] A method for high-fidelity three-dimensional modeling of large scenes that can be infinitely extended, comprising the following steps:

[0008] S1, collecting multi-view image data of drones;

[0009] S2, inputting the collected image data into colmap, and using colmap to generate the pose graph of the camera and the scene point cloud data;

[0010] S3, dividing the entire scene into multiple small blocks, using the coordinate information stored in the point cloud to find the centroid of each block, partitioning the sparse point cloud of the scene according to the centroid of the sub-region, and simultaneously retrieving the camera pose corresponding to each sub-region using the correspondence between the point cloud and the camera, obtaining the point cloud and camera pose of different blocks after scene partitioning;

[0011] S4, inputting the processed data information into 3D Gaussian Splatting (i.e., 3DGS) for parallel training to obtain the model of each different block;

[0012] S5, achieving a fast and visually realistic rendering effect through fast α-blending rendering;

[0013] S6. Render and fuse the trained blocks on the front-end browser to obtain a 3D scene from any perspective;

[0014] S7. Adjust the height of the front-end camera to obtain 3D scenes with different precisions.

[0015] Preferably, step S1 is specifically to select an area and use a drone to perform oblique photography on the scene. The drone flies along a dot matrix grid route according to the planned path; an automatic flight plan is generated by the drone software, and the drone is adjusted to the automatic flight mode to perform the flight task according to the planned route. At the same time, the height of the drone is changed to collect drone image data at different height levels.

[0016] Preferably, the calculation method of the centroid in step S3 is as follows: Set all the point clouds as a set N. According to the preset number of blocks, calculate the centroid position of each block according to the formula where n represents the position of each point cloud and x represents the centroid of a certain block. At the same time, the method of dividing blocks in step S3 is: divide the point clouds and camera poses of the entire scene into multiple corresponding blocks according to the number of blocks and the centroid position.

[0017] Preferably, step S4 is specifically: input the processed data into 3D Gaussian Splatting for parallel training, and calculate the loss function of the model using the formula where The total loss function is the weighted sum of two different loss terms; λ is a weight coefficient between 0 and 1, used to balance the contributions of the two losses; represents the loss, measuring the absolute difference between the predicted value and the true value; The structural similarity loss D-SSIM is a variant of SSIM (Structural Similarity Index Measure), used to measure the similarity in visual quality between two images.

[0018] Preferably, step S6 is specifically: load a certain block, and then dynamically swap the data of different blocks from the memory according to the visual position of the camera to prevent memory overflow.

[0019] The present invention also discloses a computer-readable storage medium, on which a computer program is stored. After the computer program runs, it executes the method described in any one of the above.

[0020] The present invention also discloses a computer system, including a processor and a storage medium. A computer program is stored on the storage medium, and the processor reads and runs the computer program to execute the method described in any one of the above.

[0021] The beneficial effects of the present invention are as follows:

[0022] The present invention supports the construction of a borderless urban-scale large scene, where parallel training can be carried out between scenes, saving computing resources and time. It supports the compression and optimization of point clouds, enabling seamless integration and real-time rendering of a borderless high-fidelity real scene large scene on a browser after segmentation. The algorithm features high stability and visualization. BRIEF DESCRIPTION OF THE DRAWINGS

[0023] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.

[0024] Figure 1 It is a flowchart of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0025] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts fall within the scope of protection of the present invention.

[0026] As Figure 1 , a method for high-fidelity three-dimensional modeling of a large scene with unbounded expansion includes the following steps:

[0027] S1. Collect multi-view image data of unmanned aerial vehicles.

[0028] Specifically, in step S1, an area is selected, and the scene is obliquely photographed using an unmanned aerial vehicle. The unmanned aerial vehicle forms a dot matrix grid flight path according to the planned path; an automatic flight path plan is generated by the unmanned aerial vehicle software, and the unmanned aerial vehicle is adjusted to the automatic flight mode to perform the flight task according to the planned flight path. At the same time, the height of the unmanned aerial vehicle is changed, and the scene is photographed again at the preset levels and heights to collect unmanned aerial vehicle image data at different height levels.

[0029] S2. Input the collected picture data into colmap, and use colmap to generate the pose map of the camera and the scene point cloud data.

[0030] S3 divides the entire scene into multiple small blocks, uses the coordinate information stored in the point cloud to find the centroid of each block, divides the scene sparse point cloud into blocks according to the centroid of the sub-area, and uses the correspondence between the point cloud and the camera to retrieve the camera pose corresponding to each sub-area, and obtains the point cloud and camera pose of different blocks after the scene is divided.

[0031] The method for calculating the centroid in step S3 is: set all point clouds as a set N, and divide them into blocks according to the pre-set number according to the formula The centroid position of each block is calculated, where n represents the position of each point cloud and x represents the centroid of a block. Meanwhile, the block division method in step S3 is: according to the number of blocks and the position of the centroid, the point cloud and camera pose of the entire scene are divided into a plurality of corresponding blocks.

[0032] S4, the processed data, i.e., the point cloud and camera pose data information of different blocks after the scene division obtained in step S3, is passed into 3D Gaussian Splatting (3DGS) for parallel training to obtain the model of each different block.

[0033] Specifically, the processed data is passed into the 3DGS for initializing the 3DGS set.

[0034]

[0035] Defined using a 3D Gaussian function, where G(x) is the Gaussian function value for a given x, representing the probability density at position x, x is the input vector, and Σ is the covariance matrix. This formula represents a 3D Gaussian distribution with a mean of 0, which is used to describe the spatial distribution characteristics of each point.

[0036]

[0037] Then the weight calculation is performed. i Represents the weight of each Gaussian point, which is used to combine different Gaussian points to generate the final rendered image. i is the transparency of each Gaussian point, and the denominator of the formula ensures that the sum of all weights is 1. i Represents the position vector or variable associated with the i-th component. Σ i : The covariance matrix of the i-th component, which describes the distribution shape and range of the component in space. Σ j The sum symbol indicates that all components are summed to normalize the weight w i After that, the 3D Gaussian image needs to be projected to 2D for rendering.

[0038] Σ′-JW∑WTT

[0039] This formula is used to transform the covariance matrix Σ into the camera coordinate system, where J is the Jacobian matrix of the projection transformation of the affine approximation, and W is the transformation matrix from the world to the camera coordinates, usually including rotation and translation. This formula ensures the correct representation of the Gaussian shape in the camera coordinate system and is applicable to the optimization process.

[0040] 2 = RRSSTRT

[0041] In order to maintain the validity of the covariance matrix during the training process. This formula represents the covariance matrix Σ as a combination of a scaling matrix S and a rotation matrix R. This allows for independent optimization of the scaling and rotation factors.

[0042]

[0043] After that, the model is trained using the loss function. Among them The total loss function is the weighted sum of two different loss terms; λ is a weight coefficient between 0 and 1, used to balance the contributions of the two losses; represents the loss, measuring the absolute difference between the predicted value and the true value; The structural similarity loss (D-SSIM) is a variant of SSIM (Structural Similarity Index Measure), used to measure the similarity in visual quality between two images. The loss function combines the L1 loss and the D-SSIM term to measure the difference between the rendered image and the training view. This loss function design helps to better capture the view-related appearance and thus improve the rendering quality.

[0044]

[0045] The gradient of the loss function L is calculated through gradients for optimizing the positions, colors, and other parameters of the Gaussian points. Among them The gradient of the loss function L represents the derivative of the loss function with respect to each parameter. is the partial derivative, representing the derivative of the loss function L with respect to the weight w i of. This indicates the degree of influence of changing the i-th weight on the overall loss function. is the weight w iThe gradient vector. Through the chain rule, the gradient of the loss function with respect to the weights is transformed into the gradient with respect to the Gaussian point parameters, thereby guiding the optimization process. At the same time, adaptive density control is added. During the optimization process, the density of the Gaussian is adaptively controlled to better represent the scene, balancing the need to capture details and reducing the computational load in less critical areas. For small Gaussian functions in under-reconstructed regions, new geometry that must be created needs to be covered. For this case, it is best to clone the Gaussian function by simply creating an identical-sized copy and moving it along the direction of the position gradient. On the other hand, large Gaussian distributions in high-variance regions need to be split into small Gaussian distributions. Such a Gaussian function is replaced by two new Gaussian functions.

[0046] Finally, Gaussian colors are obtained through color synthesis for rendering the picture.

[0047]

[0048] This formula calculates the final color C(x) through weighted averaging, where c i is the color of each Gaussian point. Each color value is weighted according to its corresponding weight w i for weighting.

[0049] S5, rendered through fast α-blending to achieve fast and visually realistic rendering effects. In computer graphics, fast α-blending is a rendering technique used to synthesize multiple image layers or image elements so that they can be displayed transparently or semi-transparently together. In this technique, the α value refers to the transparency level of each image element or layer.

[0050] S6, rendering and fusing the trained blocks on the front-end browser to obtain a three-dimensional scene from any perspective: First, load the scene of a block at a predetermined level, and according to the movement of the camera, dynamically swap in new scene blocks and swap out old scene blocks from memory to achieve seamless connection and browsing of the scene.

[0051] S7, adjusting the height of the front-end camera to obtain three-dimensional scenes with different precisions.

[0052] The present invention supports the construction of borderless city-level large scenes, and the scenes can be trained in parallel, saving computing resources and time. It supports compressing and optimizing point clouds, and after partitioning, it can be fused on the browser and real-time render a borderless high-fidelity real scene large scene. The algorithm has the characteristics of high stability and being visualizable.

[0053] The present invention also discloses a computer-readable storage medium with a computer program stored thereon. After the computer program runs, it executes the method described in any one of the above.

[0054] The present invention also discloses a computer system, including a processor and a storage medium. A computer program is stored on the storage medium. The processor reads and runs the computer program from the storage medium to execute the method described in any one of the above.

[0055] Those skilled in the art will further appreciate that the various illustrative logical blocks, modules, circuits, and algorithmic steps described in connection with the embodiments disclosed herein can be implemented as electronic hardware, computer software, or a combination of both. To clearly illustrate this interchangeability of hardware and software, the various illustrative components, blocks, modules, circuits, and steps are described above in terms of their functionality. Whether such functionality is implemented as hardware or software depends upon the particular application and the design constraints imposed on the overall system. Skilled artisans may implement the described functionality in varying ways for each particular application, but such implementation decisions should not be interpreted as causing a departure from the scope of the present invention.

[0056] The various illustrative logical blocks, modules, and circuits described in connection with the embodiments disclosed herein can be implemented using a general-purpose processor, a digital signal processor (DSP), an application specific integrated circuit (ASIC), a field programmable gate array (FPGA) or other programmable logic device, discrete gate or transistor logic, discrete hardware components, or any combination thereof designed to perform the functions described herein. A general-purpose processor may be a microprocessor, but in the alternative, the processor may be any conventional processor, controller, microcontroller, or state machine. The processor may also be implemented as a combination of computing devices, such as a combination of a DSP and a microprocessor, multiple microprocessors, one or more microprocessors cooperating with a DSP core, or any other such configuration.

[0057] The steps of a method or algorithm described in connection with the embodiments disclosed herein can be embodied directly in hardware, in a software module executed by a processor, or in a combination of both. The software module may reside in RAM memory, flash memory, ROM memory, EPROM memory, EEPROM memory, registers, a hard disk, a removable disk, a CD-ROM, or any other form of storage medium known in the art. An exemplary storage medium is coupled to the processor such that the processor can read from, and write to, the storage medium. In the alternative, the storage medium may be integrated into the processor. The processor and the storage medium may reside in an ASIC. The ASIC may reside in a user terminal. In the alternative, the processor and the storage medium may reside as discrete components in a user terminal.

[0058] In one or more exemplary embodiments, the described functionality may be implemented in hardware, software, firmware, or any combination thereof. If implemented in software as a computer program product, the functions may be stored on or transmitted via a computer-readable medium as one or more instructions or code. The computer-readable medium includes both computer storage media and communication media including any medium that facilitates transfer of a computer program from one place to another. The storage media may be any available media that can be accessed by a computer. By way of example and not limitation, such computer-readable media can comprise RAM, ROM, EEPROM, CD-ROM or other optical disk storage, magnetic disk storage or other magnetic storage devices, or any other medium that can be used to carry or store desired program code in the form of instructions or data structures and that can be accessed by a computer. Any connection is properly termed a computer-readable medium. For example, if the software is transmitted from a web site, server, or other remote source using a coaxial cable, fiber optic cable, twisted pair, DSL, or wireless technologies such as infrared, radio, and microwave, then the coaxial cable, fiber optic cable, twisted pair, DSL, or wireless technologies such as infrared, radio, and microwave are included in the definition of medium. As used herein, disk and disc include compact disc (CD), laser disc, optical disc, digital versatile disc (DVD), floppy disk, and Blu-ray disc where disks typically reproduce data magnetically, while discs reproduce data optically with lasers. The above combination should also be included within the scope of computer-readable media.

[0059] The foregoing description of the disclosure is provided to enable any person skilled in the art to make or use the disclosure. Various modifications to the disclosure will be readily apparent to those skilled in the art, and the generic principles defined herein may be applied to other variations without departing from the spirit or scope of the disclosure. Thus, the disclosure is not intended to be limited to the examples and designs described herein but is to be accorded the widest scope consistent with the principles and novel features disclosed herein.

[0060] Although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that: they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements on some of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. A large-scale high-fidelity 3D modeling method that can be expanded without boundaries, characterized in that: The following steps are involved: S1, collects multi-view image data from drones; S2, input the collected image data into colmap, and use colmap to generate the camera pose graph and scene point cloud data; S3, divide the entire scene into multiple small blocks, use the coordinate information stored in the point cloud to find the centroid of each block, divide the scene sparse point cloud into blocks according to the centroid of the sub-region, and use the correspondence between the point cloud and the camera to retrieve the camera pose corresponding to each sub-region, and obtain the point cloud and camera pose of different blocks after the scene is divided; S4, the processed data, i.e., the point cloud and camera pose data of different blocks after the scene division obtained in step S3, is passed into 3D Gaussian Splatting (3DGS) for parallel training to obtain the model of each different block; S5, rendering with fast alpha-blending for fast and visually realistic rendering; S6, rendering and fusing the trained blocks on the front-end browser to obtain a three-dimensional scene at any perspective; S7, adjust the height of the front-end camera to obtain three-dimensional scenes with different precisions; In S4: Perform weight calculation, where w i Represents the weight of each Gaussian point, which is used to combine different Gaussian points to generate the final rendered image, α i is the transparency of each Gaussian point. The denominator of the formula ensures that the sum of all weights is 1. i represents the position vector or variable associated with the ith component, Σ i : The covariance matrix of the i-th component, which describes the distribution shape and range of the component in space, Σ j The sum symbol indicates that all components are summed to normalize the weight w i ,After that, the 3D Gaussian image needs to be projected into 2D for rendering, Σ′=JWΣW T J T This formula is used to transform the covariance matrix Σ into the camera coordinate system, where J is the Jacobian matrix of the projection transformation of the affine approximation, and W is the transformation matrix from the world to the camera coordinate system, which usually includes rotation and translation. This formula ensures that the shape of the Gaussian is correctly represented in the camera coordinate system and is suitable for the optimization process. S = RSS T R T In order to maintain the validity of the covariance matrix during training, this formula represents the covariance matrix Σ as a combination of a scaling matrix S and a rotation matrix R, and independently optimizes the scaling and rotation factors. The model is then trained using the loss function, where The total loss function is the weighted sum of two different loss terms; λ is a weight coefficient between 0 and 1, which is used to balance the contribution of the two losses; represents the loss, which measures the absolute difference between the predicted value and the true value; Structural Similarity Loss (D-SSIM) is a variant of SSIM (Structural Similarity Index Measure) and is used to measure the similarity of two images in visual quality. The loss function combines L1 loss and D-SSIM terms to measure the difference between the rendered image and the training view. This loss function design helps to better capture the view-related appearance, thereby improving the rendering quality. The gradient of the loss function L is optimized by gradient calculation to optimize the position, color and other parameters of the Gaussian points. The gradient of the loss function L represents the derivative of the loss function with respect to each parameter. is the partial derivative, which represents the loss function L with respect to the weight w i The derivative of , which shows the impact of changing the i-th weight on the overall loss function, is the weight w i The gradient vector of .

2. The method according to claim 1, characterized in that: Step S1 specifically includes selecting an area, using a drone to take oblique photographs of the scene, and the drone flies along a dot grid route according to the planned path; the drone software automatically generates a route plan, and the drone is switched to automatic flight mode to perform the flight mission according to the planned route. At the same time, the drone's altitude is changed to collect drone image data at different altitude levels.

3. The method according to claim 1, characterized in that The method for calculating the centroid in step S3 is as follows: set all point clouds to a set N, and divide them into blocks according to the pre-set number of blocks according to the formula The centroid position of each block is calculated, where n represents the position of each point cloud, and x represents the centroid of a block. At the same time, the block division method in step S3 is: according to the number of blocks and the position of the centroid, the point cloud and camera pose of the entire scene are divided into multiple corresponding blocks.

4. The method according to claim 1, characterized in that: Step S6 specifically includes: loading a certain block, and then dynamically exchanging data of different blocks from the memory according to the visual position of the camera to prevent memory overflow.

5. A computer-readable storage medium, characterized in that: A computer program is stored on the medium, and after the computer program is run, the method according to any one of claims 1 to 4 is executed.

6. A computer system, characterized in that: The method comprises a processor and a storage medium, wherein a computer program is stored in the storage medium, and the processor reads and runs the computer program from the storage medium to execute the method as claimed in any one of claims 1 to 4.

Citation Information

Patent Citations

  • Object surface three-dimensional reconstruction method and device based on point cloud data

    CN116597113A

  • Fast large-scale three-dimensional reconstruction method based on neural radiation field

    CN117726758A