Outdoor scene modeling method and device based on depth image and electronic equipment
By collecting multi-perspective color images by drones and combining them with the SFM and PatchMatch algorithms to generate 3D sparse point clouds, the problems of high equipment cost, low efficiency and difficulty in handling dynamic interference in existing technologies are solved, and high-precision, real-time outdoor scene modeling is achieved.
Patent Information
- Application Number
- CN202510791894.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-13
- Publication Date
- 2025-09-12
AI Technical Summary
Existing outdoor scene modeling technologies have problems such as high equipment cost, low reconstruction efficiency, difficulty in handling dynamic interference, high reconstruction failure rate in weak texture areas, and large impact of lighting changes, making it difficult to achieve efficient, low-cost and robust three-dimensional model reconstruction.
The method uses drone-based multi-view color images to extract sparse point clouds, combines SFM technology and PatchMatch algorithm, generates 3D sparse point clouds through color and geometric constraints, and performs 3D Gaussian lattice rendering reconstruction to reduce equipment costs and improve reconstruction efficiency.
It achieves high-precision, real-time 3D model reconstruction, overcomes the influence of dynamic interference and illumination changes, reduces costs, and improves reconstruction efficiency and model integrity.
Smart Images

Figure CN120635322A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the technical field of computer vision and three-dimensional reconstruction, and in particular to a method, device and electronic device for outdoor scene modeling based on depth images. Background Art
[0002] Outdoor scene modeling refers to the process of creating outdoor scenes using 3D modeling software. It is a technology that converts outdoor objects and environments in the real world into virtual 3D models and is widely used in many fields.
[0003] Outdoor scene modeling mainly faces three major challenges: multiple dynamic interferences, large data scale, and complex geometric details.
[0004] Current mainstream technical solutions have the following limitations:
[0005] 1. Laser point cloud reconstruction solution: Use vehicle-mounted / airborne LiDAR to obtain high-precision 3D point clouds, and combine GPS / IMU positioning data to build scene models.
[0006] Disadvantages: The equipment is expensive (millions of yuan), making it difficult to deploy on a large scale; point cloud density is unevenly distributed (sparse point clouds at long distances), resulting in loss of details on building facades and vegetation; dynamic objects (such as moving vehicles) produce motion artifacts that require manual correction later.
[0007] 2. Multi-view geometric reconstruction solution: Sparse point clouds are extracted from multi-view RGB images, and dense point cloud meshes are generated through dense matching.
[0008] Defects: Failure in weak texture areas: Low-texture areas such as the sky and road surface have reconstruction holes due to insufficient feature points (for example, the hole rate in urban scenes is >15%); Efficiency bottleneck: Reconstruction of thousands of images takes >24 hours (GPU cluster acceleration is required); Scale drift problem: Accumulated errors during large-scale reconstruction lead to model distortion (scale error of 0.5-1m per kilometer).
[0009] 3. Neural Radiance Field (NeRF) method: uses implicit neural fields to represent scene geometry and materials, and outputs new perspective images through volume rendering.
[0010] Disadvantages: Long training time: Single-scenario training takes 6-12 hours, which cannot meet real-time requirements;
[0011] Poor support for dynamic scenes: unable to effectively model moving objects;
[0012] Low outdoor applicability: Strong lighting changes lead to loss of details in highlight areas (PSNR < 25dB).
[0013] In summary, existing approaches to outdoor scene modeling suffer from a precision-efficiency trade-off: Traditional point cloud / mesh methods rely on expensive hardware and struggle to handle dynamic interference; visual geometry methods suffer from a high reconstruction failure rate in weakly textured areas; and neural rendering methods are computationally expensive and unable to output complete models in real time. Therefore, a highly efficient, cost-effective, and robust outdoor scene modeling technology is urgently needed. Summary of the Invention
[0014] The embodiments of the present application provide a method, device, and electronic device for outdoor scene modeling based on depth images to solve the technical problems existing in the prior art.
[0015] Other features and advantages of the present application will become apparent from the following detailed description, or may be learned in part by practice of the present application.
[0016] According to a first aspect of an embodiment of the present application, a method for outdoor scene modeling based on a depth image is provided, comprising:
[0017] Collect multi-view color images of objects in outdoor scenes based on drones;
[0018] Extracting an initial sparse point cloud and a depth image of the object based on the multi-view color image;
[0019] Performing color constraints on the initial sparse point cloud based on the multi-view color image and performing geometric constraints on the initial sparse point cloud based on the depth image to obtain a 3D sparse point cloud;
[0020] 3D Gaussian lattice rendering and reconstruction are performed based on the 3D sparse point cloud.
[0021] In some embodiments of the present application, based on the aforementioned solution, the multi-view color image of the object collected by the drone includes:
[0022] Set the drone's collection direction and flight trajectory based on the object's appearance and shape;
[0023] The drone is flown along the flight trajectory, and multi-view color images of the object are collected according to the collection direction during the flight.
[0024] In some embodiments of the present application, based on the aforementioned solution, extracting an initial sparse point cloud of an object based on the multi-view color image includes:
[0025] Using SFM technology to perform feature extraction and matching on the multi-view color image;
[0026] Based on the extraction and matching results, the linear triangulation method is used for initial reconstruction;
[0027] Based on the initial reconstruction, the extended view is incrementally reconstructed;
[0028] Perform global optimization and output the initial sparse point cloud.
[0029] In some embodiments of the present application, based on the aforementioned solution, extracting a depth image of an object based on the multi-view color image includes:
[0030] performing distortion correction, illumination normalization, and feature consistency processing on the multi-view color image;
[0031] Use the PatchMatch stereo matching algorithm to calculate the depth image corresponding to each perspective color image;
[0032] Each depth image is subjected to outlier filtering, weighted median filtering, and hole filling processing.
[0033] In some embodiments of the present application, based on the aforementioned solution, the color constraint on the initial sparse point cloud based on the multi-view color image includes:
[0034] Calibrate the coordinate systems of the multi-view color image and the initial sparse point cloud;
[0035] Establishing a point-pixel mapping relationship between the initial sparse point cloud and the multi-view color image;
[0036] Colors are assigned to the initial sparse point cloud based on the point-pixel mapping relationship.
[0037] In some embodiments of the present application, based on the aforementioned scheme, during the color assignment process, for the multi-perspective overlapping area, the minimum and maximum points of the color difference are calculated, and the colors are fused using linear interpolation based on the minimum and maximum points.
[0038] In some embodiments of the present application, based on the aforementioned solution, performing geometric constraints on the initial sparse point cloud based on the depth image includes:
[0039] generating a depth point cloud based on the depth image;
[0040] Inputting the depth point cloud into an occupancy grid to generate an occupancy probability grid;
[0041] Setting a threshold, deleting low-probability grids below the threshold, and obtaining a sparse voxel structure;
[0042] The sparse voxel structure is used as a constraint framework to constrain the initial sparse point cloud.
[0043] In some embodiments of the present application, based on the aforementioned solution, performing 3D Gaussian lattice rendering reconstruction based on the 3D sparse point cloud includes:
[0044] Initialize a 3D Gaussian lattice based on the 3D sparse point cloud;
[0045] Performing differentiable rendering processing on the 3D Gaussian lattice and optimizing Gaussian parameters;
[0046] During the optimization process, the density of Gaussian points is dynamically adjusted according to the complexity of the scene where the object is located;
[0047] Iteratively training the 3D Gaussian lattice;
[0048] Render the trained 3D Gaussian lattice to complete the reconstruction.
[0049] According to a second aspect of an embodiment of the present application, there is provided an outdoor scene modeling device based on a depth image, comprising:
[0050] An acquisition unit, used to acquire multi-view color images of objects in outdoor scenes based on a drone;
[0051] An extraction unit, configured to extract an initial sparse point cloud and a depth image of an object based on the multi-view color image;
[0052] a constraint unit, configured to perform color constraints on the initial sparse point cloud based on the multi-view color image and perform geometric constraints on the initial sparse point cloud based on the depth image, to obtain a 3D sparse point cloud;
[0053] A reconstruction unit is used to perform 3D Gaussian lattice rendering reconstruction based on the 3D sparse point cloud.
[0054] According to a third aspect of an embodiment of the present application, there is provided an electronic device, including: a memory and a processor;
[0055] The memory is used to store computer instructions;
[0056] The processor is configured to call the computer instructions stored in the memory so that the electronic device executes the method according to the first aspect.
[0057] The technical solution of the present application realizes high-precision, real-time, and fully automatic three-dimensional model reconstruction of outdoor scenes based on depth images, effectively overcomes the influence of dynamic interference and illumination changes, and reduces costs.
[0058] It should be understood that the foregoing general description and the following detailed description are exemplary and explanatory only and are not restrictive of the present application. BRIEF DESCRIPTION OF THE DRAWINGS
[0059] The accompanying drawings are incorporated into and constitute a part of the specification, illustrating embodiments consistent with the present application and, together with the specification, explaining the principles of the present application. Obviously, the drawings described below are only some embodiments of the present application, and those skilled in the art can derive other drawings based on these drawings without inventive effort. In the drawings:
[0060] Figure 1 A schematic diagram of a process for modeling an outdoor scene based on a depth image according to an embodiment of the present application is shown;
[0061] Figure 2 A block diagram of an outdoor scene modeling device based on a depth image according to an embodiment of the present application is shown;
[0062] Figure 3 A block diagram of an electronic device according to an embodiment of the present application is shown;
[0063] Figure 4 A schematic diagram of the structure of a computer system suitable for implementing an electronic device according to an embodiment of the present application is shown. DETAILED DESCRIPTION
[0064] Example embodiments will now be described more fully with reference to the accompanying drawings. However, example embodiments can be implemented in many forms and should not be construed as limited to the examples set forth herein; rather, these embodiments are provided so that this application will be thorough and complete and will fully convey the concepts of the example embodiments to those skilled in the art.
[0065] In addition, described feature, structure or characteristic can be combined in one or more embodiments in any suitable manner.In the following description, many specific details are provided so as to provide a full understanding of the embodiments of the present application. However, it will be appreciated by those skilled in the art that the technical scheme of the present application can be put into practice without one or more of the specific details, or other methods, components, devices, steps etc. can be adopted. In other cases, known methods, devices, implementations or operations are not shown or described in detail to avoid blurring the various aspects of the application.
[0066] The block diagrams shown in the accompanying drawings are merely functional entities and do not necessarily correspond to physically separate entities. That is, these functional entities may be implemented in software, in one or more hardware modules or integrated circuits, or in different networks and / or processor devices and / or microcontroller devices.
[0067] The flowcharts shown in the accompanying drawings are for illustrative purposes only and do not necessarily include all contents and operations / steps, nor must they be executed in the order described. For example, some operations / steps may be decomposed, while others may be combined or partially combined. Therefore, the actual execution order may vary depending on the actual situation.
[0068] To make the objectives, technical solutions, and advantages of the present invention more clear, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts shall fall within the scope of protection of the present invention.
[0069] The following will describe some embodiments of the present application in detail with reference to the accompanying drawings. In the absence of conflict, the following embodiments and features therein may be combined with each other.
[0070] See also Figure 1 , shows a flow chart of a method for outdoor scene modeling based on depth images according to an embodiment of the present application.
[0071] like Figure 1 As shown, a method for outdoor scene modeling based on depth images is presented, including steps S100 to S400.
[0072] refer to Figure 1 ,Step S100, collecting multi-view color pictures of objects in outdoor scenes based on the drone.
[0073] It is understandable that drones can easily capture pictures of outdoor objects, reducing shooting costs and improving shooting efficiency.
[0074] In some feasible embodiments, based on the above solution, step S100 includes:
[0075] Set the drone's collection direction and flight trajectory based on the object's appearance and shape;
[0076] The drone is flown along the flight trajectory, and multi-view color images of the object are collected according to the collection direction during the flight.
[0077] It can be understood that the set collection direction and flight trajectory can ensure that the drone can collect complete object images from multiple perspectives.
[0078] Continue to refer Figure 1 , step S200, extracting the initial sparse point cloud and depth image of the object based on the multi-view color picture.
[0079] It can be understood that the initial sparse point cloud can construct the initial macrostructure of the object, while the depth image can directly reflect the geometry of the visible surface of the object.
[0080] In some feasible embodiments, based on the aforementioned solution, extracting an initial sparse point cloud of an object based on the multi-view color image includes:
[0081] Using SFM technology to perform feature extraction and matching on the multi-view color image;
[0082] Based on the extraction and matching results, the linear triangulation method is used for initial reconstruction;
[0083] Based on the initial reconstruction, the extended view is incrementally reconstructed;
[0084] Perform global optimization and output the initial sparse point cloud.
[0085] It should be noted that SFM technology refers to Structure from Motion, which is a technology in the field of computer vision that is used to recover three-dimensional scene structures from image sequences.
[0086] It can be understood that this step converts the 2D multi-view color image into an initial sparse point cloud through the SFM technology and incremental optimization mechanism, providing an initial structure for subsequent model reconstruction.
[0087] In some feasible embodiments, based on the aforementioned solution, extracting a depth image of an object based on the multi-view color image includes:
[0088] performing distortion correction, illumination normalization, and feature consistency processing on the multi-view color image;
[0089] Use the PatchMatch stereo matching algorithm to calculate the depth image corresponding to each perspective color image;
[0090] Each depth image is subjected to outlier filtering, weighted median filtering, and hole filling processing.
[0091] It should be noted that, in this embodiment, preprocessing such as distortion correction, illumination normalization, and feature consistency is performed on the multi-view color image, which reduces the impact of illumination differences and preserves edges while suppressing noise.
[0092] It should be noted that, in this embodiment, the outlier filtering process specifically includes:
[0093] Consistency check: Compare the depth consistency of the reference view and the source view, and remove points with large differences.
[0094] Depth Range Constraint: Removes depth values that are outside of physical possibility.
[0095] Weighted median filtering is a process that smoothes the depth map while preserving edges.
[0096] Hole filling processing refers to filling the missing areas using plane fitting or neighborhood interpolation.
[0097] Continue to refer Figure 1 In step S300 , color constraints are performed on the initial sparse point cloud based on the multi-view color image and geometric constraints are performed on the initial sparse point cloud based on the depth image to obtain a 3D sparse point cloud.
[0098] In some feasible embodiments, based on the above solution, the performing color constraints on the initial sparse point cloud based on the multi-view color image includes:
[0099] Calibrate the coordinate systems of the multi-view color image and the initial sparse point cloud;
[0100] Establishing a point-pixel mapping relationship between the initial sparse point cloud and the multi-view color image;
[0101] Colors are assigned to the initial sparse point cloud based on the point-pixel mapping relationship.
[0102] It can be understood that this step improves the color authenticity while maintaining the sparse point cloud structure, making it more consistent with the real scene.
[0103] In some feasible embodiments, based on the aforementioned solution, during the color assignment process, for the multi-view overlapping area, the minimum and maximum points of the color difference are calculated, and the colors are fused using linear interpolation based on the minimum and maximum points.
[0104] It can be understood that this step solves the problem of color difference between overlapping areas.
[0105] In some feasible embodiments, based on the above solution, performing geometric constraints on the initial sparse point cloud based on the depth image includes:
[0106] generating a depth point cloud based on the depth image;
[0107] Inputting the depth point cloud into an occupancy grid to generate an occupancy probability grid;
[0108] Setting a threshold, deleting low-probability grids below the threshold, and obtaining a sparse voxel structure;
[0109] The sparse voxel structure is used as a constraint framework to constrain the initial sparse point cloud.
[0110] It can be understood that this step can significantly improve the geometric accuracy and structural integrity of the sparse point cloud.
[0111] Continue to refer Figure 1 , step S400, performing 3D Gaussian lattice rendering reconstruction based on the 3D sparse point cloud.
[0112] It should be noted that 3D Gaussian lattice rendering is an advanced explicit 3D reconstruction technology that can convert sparse point clouds into renderable high-fidelity models.
[0113] In some feasible embodiments, based on the above solution, performing 3D Gaussian lattice rendering reconstruction based on the 3D sparse point cloud includes:
[0114] Initialize a 3D Gaussian lattice based on the 3D sparse point cloud;
[0115] Performing differentiable rendering processing on the 3D Gaussian lattice and optimizing Gaussian parameters;
[0116] During the optimization process, the density of Gaussian points is dynamically adjusted according to the complexity of the scene where the object is located;
[0117] Iteratively training the 3D Gaussian lattice;
[0118] Render the trained 3D Gaussian lattice to complete the reconstruction.
[0119] It can be understood that through this step, the 3D sparse point cloud can be converted into a high-quality renderable model, which can achieve real-time lighting dynamic effects while maintaining geometric accuracy.
[0120] The following describes an embodiment of the device of the present application, which can be used to implement a depth image-based outdoor scene modeling method in the above embodiment of the present application. For details not disclosed in the embodiment of the device of the present application, please refer to the embodiment of the method of the above application.
[0121] Reference Figure 2 As shown, according to an embodiment of the present application, a depth image-based outdoor scene modeling device 200 includes:
[0122] The acquisition unit 201 is configured to acquire multi-view color images of objects in an outdoor scene using a drone;
[0123] An extraction unit 202 is configured to extract an initial sparse point cloud and a depth image of an object based on the multi-view color image;
[0124] A constraint unit 203 is configured to perform color constraints on the initial sparse point cloud based on the multi-view color image and to perform geometric constraints on the initial sparse point cloud based on the depth image to obtain a 3D sparse point cloud;
[0125] The reconstruction unit 204 is configured to perform 3D Gaussian lattice rendering reconstruction based on the 3D sparse point cloud.
[0126] like Figure 3 As shown, an embodiment of the present application also provides an electronic device 300, including a memory 310, a processor 320, and a computer program 311 stored in the memory 310 and executable on the processor. When the processor 320 executes the computer program 311, the steps of the above-mentioned outdoor scene modeling method based on depth images are implemented.
[0127] Since the electronic device introduced in this embodiment is a device used to implement an outdoor scene modeling device based on depth images in the embodiment of the present application, based on the method introduced in the embodiment of the present application, technical personnel in this field can understand the specific implementation of the electronic device of this embodiment and its various variations. Therefore, how the electronic device implements the method in the embodiment of the present application is no longer introduced in detail here. As long as the equipment used by technical personnel in this field to implement the method in the embodiment of the present application falls within the scope of protection to be protected by this application.
[0128] During the specific implementation process, when the computer program 311 is executed by the processor, any implementation method of the embodiments corresponding to the first aspect can be implemented.
[0129] Figure 4 A schematic diagram of the structure of a computer system suitable for implementing an electronic device according to an embodiment of the present application is shown.
[0130] It should be noted that Figure 4 The computer system 400 of the electronic device shown is only an example and should not bring any limitation to the functions and scope of use of the embodiments of the present application.
[0131] like Figure 4 As shown, the computer system 400 includes a central processing unit (CPU) 401, which can perform various appropriate actions and processes according to the program stored in the read-only memory (ROM) 402 or the program loaded from the storage part 408 into the random access memory (RAM) 403, such as executing the method described in the above embodiment. Various programs and data required for system operation are also stored in the RAM 403. The CPU 401, ROM 402 and RAM 403 are connected to each other via a bus 404. An input / output (I / O) interface 405 is also connected to the bus 404.
[0132] The following components are connected to the I / O interface 405: an input section 406 including a keyboard, a mouse, and the like; an output section 407 including devices such as a cathode ray tube (CRT), a liquid crystal display (LCD), and a speaker; a storage section 408 including a hard disk and the like; and a communication section 409 including a network interface card such as a LAN (Local Area Network) card or a modem. The communication section 409 performs communication processing via a network such as the Internet. A drive 410 is also connected to the I / O interface 405 as needed. Removable media 411, such as a magnetic disk, an optical disk, a magneto-optical disk, or a semiconductor memory, is installed in the drive 410 as needed, so that computer programs read therefrom can be installed into the storage section 408 as needed.
[0133] In particular, according to an embodiment of the present application, the process described above with reference to the flowchart can be implemented as a computer software program. For example, an embodiment of the present application includes a computer program product, which includes a computer program carried on a computer-readable medium, and the computer program includes program code for executing the method shown in the flowchart. In such an embodiment, the computer program can be downloaded and installed from a network via the communication section 409, and / or installed from a removable medium 411. When the computer program is executed by the central processing unit (CPU) 401, the various functions defined in the system of the present application are executed.
[0134] It should be noted that the computer-readable medium shown in the embodiments of the present application can be a computer-readable signal medium or a computer-readable storage medium or any combination of the above two. The computer-readable storage medium can be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, device or device, or any combination of the above. More specific examples of computer-readable storage media can include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM), a flash memory, an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present application, a computer-readable storage medium can be any tangible medium containing or storing a program that can be used by or in combination with an instruction execution system, device or device. In the present application, a computer-readable signal medium can include a data signal propagated in baseband or as part of a carrier wave, which carries a computer-readable program code. Such propagated data signals may take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. A computer-readable signal medium may also be any computer-readable medium other than a computer-readable storage medium that can transmit, propagate, or transport a program for use by or in connection with an instruction execution system, apparatus, or device. Program code embodied on a computer-readable medium may be transmitted using any suitable medium, including but not limited to wireless, wired, or any suitable combination thereof.
[0135] The flowcharts and block diagrams in the accompanying drawings illustrate the possible implementation architecture, functions and operations of the systems, methods and computer program products according to various embodiments of the present application. Among them, each box in the flowchart or block diagram can represent a module, program segment, or a part of the code, and the above-mentioned module, program segment, or a part of the code contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the box can also occur in an order different from that marked in the accompanying drawings. For example, two boxes represented in succession can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram or flowchart, and the combination of boxes in the block diagram or flowchart, can be implemented with a dedicated hardware-based system that performs the specified function or operation, or can be implemented with a combination of dedicated hardware and computer instructions.
[0136] The units involved in the embodiments described in this application may be implemented by software or hardware, and the units described may also be set in a processor. In some cases, the names of these units do not constitute limitations on the units themselves.
[0137] As another aspect, the present application further provides a computer program product or computer program, which includes computer instructions stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium and executes the computer instructions, causing the computer device to perform the outdoor scene modeling method based on a depth image described in the above embodiment.
[0138] As another aspect, the present application also provides a computer-readable medium, which may be included in the electronic device described in the above embodiment; or may exist independently without being incorporated into the electronic device. The computer-readable medium carries one or more programs, and when the one or more programs are executed by the electronic device, the electronic device implements the outdoor scene modeling method based on depth images described in the above embodiment.
[0139] It should be noted that, although several modules or units of the device for action execution are mentioned in the above detailed description, this division is not mandatory. In fact, according to the embodiment of the application, the features and functions of two or more modules or units described above can be concretized in one module or unit. On the contrary, the features and functions of one module or unit described above can be further divided into multiple modules or units to be concretized.
[0140] Through the description of the above embodiments, it is easy for those skilled in the art to understand that the example embodiments described herein can be implemented by software or by combining software with necessary hardware. Therefore, the technical solution according to the embodiments of the present application can be embodied in the form of a software product, which can be stored in a non-volatile storage medium (which can be a CD-ROM, a USB flash drive, a mobile hard disk, etc.) or on a network, and includes several instructions to enable a computing device (which can be a personal computer, a server, a touch terminal, or a network device, etc.) to execute the method according to the embodiments of the present application.
[0141] Those skilled in the art will readily appreciate other embodiments of the present application after considering the specification and practicing the embodiments disclosed herein. This application is intended to cover any variations, uses, or adaptations of the present application that follow the general principles of the present application and include common knowledge or customary techniques in the art that are not disclosed in this application. It should be understood that this application is not limited to the precise structures described above and shown in the accompanying drawings, and various modifications and changes can be made without departing from the scope thereof. The scope of this application is limited only by the appended claims.
Claims
1. A method for outdoor scene modeling based on depth images, characterized in that: include: Collect multi-view color images of objects in outdoor scenes based on drones; Extracting an initial sparse point cloud and a depth image of the object based on the multi-view color image; Performing color constraints on the initial sparse point cloud based on the multi-view color image and performing geometric constraints on the initial sparse point cloud based on the depth image to obtain a 3D sparse point cloud; 3D Gaussian lattice rendering and reconstruction are performed based on the 3D sparse point cloud.
2. The method according to claim 1, characterized in that The multi-view color image of the object collected by the drone includes: Set the drone's collection direction and flight trajectory based on the object's appearance and shape; The drone is flown along the flight trajectory, and multi-view color images of the object are collected according to the collection direction during the flight.
3. The method according to claim 1, characterized in that Extracting an initial sparse point cloud of an object based on the multi-view color image includes: Using SFM technology to perform feature extraction and matching on the multi-view color image; Based on the extraction and matching results, the linear triangulation method is used for initial reconstruction; Based on the initial reconstruction, the extended view is incrementally reconstructed; Perform global optimization and output the initial sparse point cloud.
4. The method according to claim 1, wherein Extracting a depth image of an object based on the multi-view color image includes: performing distortion correction, illumination normalization, and feature consistency processing on the multi-view color image; Use the PatchMatch stereo matching algorithm to calculate the depth image corresponding to each perspective color image; Each depth image is subjected to outlier filtering, weighted median filtering, and hole filling processing.
5. The method according to claim 1, wherein The performing color constraints on the initial sparse point cloud based on the multi-view color image includes: Calibrate the coordinate systems of the multi-view color image and the initial sparse point cloud; Establishing a point-pixel mapping relationship between the initial sparse point cloud and the multi-view color image; Colors are assigned to the initial sparse point cloud based on the point-pixel mapping relationship.
6. The method according to claim 5, characterized in that In the process of color assignment, for the multi-view overlapping area, the minimum point and the maximum point of the color difference are calculated, and the colors are fused using a linear interpolation method based on the minimum point and the maximum point.
7. The method according to claim 1, characterized in that The performing geometric constraints on the initial sparse point cloud based on the depth image includes: generating a depth point cloud based on the depth image; Inputting the depth point cloud into an occupancy grid to generate an occupancy probability grid; Setting a threshold, deleting low-probability grids below the threshold, and obtaining a sparse voxel structure; The sparse voxel structure is used as a constraint framework to constrain the initial sparse point cloud.
8. The method according to claim 1, characterized in that The performing 3D Gaussian dot matrix rendering and reconstruction based on the 3D sparse point cloud includes: Initialize a 3D Gaussian lattice based on the 3D sparse point cloud; Performing differentiable rendering processing on the 3D Gaussian lattice and optimizing Gaussian parameters; During the optimization process, the density of Gaussian points is dynamically adjusted according to the complexity of the scene where the object is located; Iteratively training the 3D Gaussian lattice; Render the trained 3D Gaussian lattice to complete the reconstruction.
9. A device for modeling outdoor scenes based on depth images, characterized in that: include: An acquisition unit, used to acquire multi-view color images of objects in outdoor scenes based on a drone; An extraction unit, configured to extract an initial sparse point cloud and a depth image of an object based on the multi-view color image; a constraint unit, configured to perform color constraints on the initial sparse point cloud based on the multi-view color image and perform geometric constraints on the initial sparse point cloud based on the depth image, to obtain a 3D sparse point cloud; A reconstruction unit is used to perform 3D Gaussian lattice rendering reconstruction based on the 3D sparse point cloud.
10. An electronic device, characterized in that: include: memory and processor; The memory is used to store computer instructions; The processor is configured to call the computer instructions stored in the memory so that the electronic device executes the method according to any one of claims 1 to 8.