Underground mine scene reconstruction method based on three-dimensional gaussian sputtering non-uniform illumination
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2026-03-31
- Publication Date
- 2026-08-11
AI Technical Summary
[0005]基于此,本发明的目的是提供一种基于三维高斯溅射的非均匀光照的地下矿井场景重建方法,旨在解决现有技术中缺少一种兼顾鲁棒性、高精度与实时性的基于三维高斯溅射的非均匀光照的地下矿井场景重建方法的问题
[0054] This invention acquires visual data from underground mine scenes and processes it into multi-view images. Figure 2This method uses 3D image data to generate sparse point clouds, and then initializes a 3D Gaussian ellipsoid attribute set containing dynamic structural features to obtain an initial 3D Gaussian volume set for the mine scene. Simultaneously, it extracts a physical prior structure map adapted to the mine environment as a supervision signal for Gaussian ellipsoid geometric growth. This effectively solves the problems of existing motion structure recovery algorithms generating sparse and noisy initial point clouds that cannot support accurate optimization, and the lack of effective structural supervision during Gaussian volume geometric growth that easily leads to structural breakage in the reconstruction. Furthermore, by decoupling the initial 3D Gaussian volume set through rendering to generate vector chromaticity maps, physical radiation maps, and structural feature maps, it removes the interference of lighting on the intrinsic colors of objects, overcoming the limitations of existing techniques that use RGB. Color spaces result in strong coupling between color and lighting, making it easy for models to incorporate lighting effects into geometric structures, leading to color flickering and distortion in reconstructed images. This paper addresses the problem of inconsistent exposure in multi-view images of mines, which easily produces floating artifacts and geometric collapse. Furthermore, sensor noise and dust scattering noise are amplified during rendering, making it difficult for traditional denoising methods to balance detail preservation and noise suppression. The paper further describes the process of reconstructing clean physical intensity maps and vector chromaticity maps into HSV components and inversely transforming them to obtain RGB images. Multiple loss constraint functions are constructed based on the RGB images, structural feature maps, physical prior structure maps, and real images of the mine scene, employing Adam... The optimization algorithm performs end-to-end training and updates of the learnable parameters of the 3D Gaussian ellipsoid until the loss converges, overcoming the shortcomings of existing technologies such as the lack of effective multi-dimensional constraints in model training and insufficient reconstruction accuracy and robustness to extreme mine environments. By performing viewpoint adaptive rendering based on the converged Gaussian volume set parameters, it outputs the 3D reconstruction result of the underground mine scene from the desired viewpoint, solving the problems of slow training speed and insufficient real-time rendering capabilities of NeRF and its variants in existing technologies, which cannot meet the needs of rapid reconstruction and real-time navigation in mine sites. Therefore, this invention solves the problem of the lack of a robust, high-accuracy, and real-time underground mine scene reconstruction method based on non-uniform illumination using 3D Gaussian sputtering.
Smart Images

Figure CN121962472B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of underground mine scene reconstruction technology, and in particular to a method for underground mine scene reconstruction based on non-uniform illumination using three-dimensional Gaussian sputtering. Background Technology
[0002] Safe production and intelligent mining in underground mines are the core directions for high-quality development in the mining industry. 3D reconstruction of underground mine scenes, as a key technology for smart mine construction, is a crucial foundation for realizing digital twins of mines, navigation and planning for unmanned inspection robots, emergency drills for underground disasters, and intelligent equipment operation and maintenance. Accurate 3D mine scene models can intuitively reconstruct key information such as roadway topology, equipment layout, and spatial structure, providing data support for risk warning, path planning, and remote control in the mine production process. This plays an irreplaceable role in improving the level of intelligent mine production and reducing the safety risks of manual operations. With the development of computer vision and 3D reconstruction technology, visual acquisition-based 3D reconstruction methods have become the mainstream technical route for underground mine scene reconstruction due to their non-contact and high convenience advantages. Their technical performance directly determines the effectiveness of intelligent mine applications.
[0003] In the field of 3D scene reconstruction technology, 3D Gaussian Splatting (3DGS) has gradually replaced traditional reconstruction methods as the mainstream solution for 3D reconstruction of industrial scenes due to its advantages of fast real-time rendering and high reconstruction accuracy. Compared with reconstruction techniques such as Neural Radiance Field (NeRF) and its variants, 3D Gaussian Splatting does not require a complex voxel rendering process, and can meet the real-time requirements of industrial sites while ensuring reconstruction accuracy. However, it has many significant drawbacks when directly applied to non-uniform lighting scenes in mines.
[0004] While existing 3D Gaussian sputtering technology is the mainstream solution for 3D reconstruction, its direct application to non-uniform lighting scenes in underground mines lacks adaptation to extreme environmental characteristics such as low light and weak textures, strong noise, dust and water vapor interference, and photometric consistency failure caused by light source movement. This results in numerous key technical defects. The initial point cloud generated by the motion recovery structure algorithm is sparse and contains a large amount of noise, failing to provide a reliable foundation for subsequent accurate optimization. Furthermore, the use of the RGB color space leads to strong coupling between color and lighting, making it easy for the model to incorrectly integrate lighting effects into the geometric structure, causing color flickering distortion. The inconsistent exposure of multi-view images of the mine can also directly cause floating artifacts and geometric collapse in the reconstruction. In addition, the ability to suppress noise from mine sensors and dust scattering is insufficient, and this noise is amplified during the rendering process, ultimately leading to structural fractures and frequent artifacts in the reconstructed model. The reconstruction accuracy and environmental robustness are difficult to meet the requirements of the mine. While NeRF and its variants can adapt to complex lighting conditions to some extent, they suffer from slow training speed and insufficient real-time rendering capabilities, making them unsuitable for real-time applications such as rapid reconstruction of mine sites and unmanned inspection navigation. Summary of the Invention
[0005] Therefore, the purpose of this invention is to provide a method for reconstructing underground mine scenes based on non-uniform illumination using three-dimensional Gaussian sputtering, aiming to solve the problem that there is a lack of a method for reconstructing underground mine scenes based on non-uniform illumination using three-dimensional Gaussian sputtering that takes into account robustness, high accuracy and real-time performance in the prior art.
[0006] A method for reconstructing an underground mine scene based on non-uniform illumination using three-dimensional Gaussian sputtering, according to an embodiment of the present invention, the method comprising:
[0007] Visual data of an underground mine scene is acquired, and the visual data is processed to obtain multi-view results. Figure 2 3D image data, to be based on the multi-view Figure 2 A sparse point cloud is generated from the 3D image data, and a set of 3D Gaussian ellipsoidal attributes with dynamic structural features is initialized based on the sparse point cloud to obtain an initial set of 3D Gaussian volumes for the mine scene.
[0008] According to the multi-view Figure 2 Physical prior structure diagrams adapted to the mine environment are extracted from 3D image data and used as a supervision signal for Gaussian ellipsoid geometric growth;
[0009] The initial three-dimensional Gaussian volume set is decoupled and rendered to generate a vector chromaticity map, a physical radiation map, and a structural feature map;
[0010] Based on the camera parameters when the visual data was acquired, the physical radiometric map was exposed and compensated, and dark noise was suppressed by the intensity collapse function to obtain a clean physical intensity map.
[0011] The clean physical intensity map and the vector chromaticity map are reconstructed into HSV components, and an inverse transform is performed to obtain an RGB image.
[0012] Based on the RGB image, the structural feature map, the physical prior structure map, and the real image of the mine scene, a multi-loss constraint function is constructed. The Adam optimization algorithm is used to perform end-to-end training and update of the learnable parameters of the three-dimensional Gaussian ellipsoid until the loss converges.
[0013] Based on the parameters of the Gaussian body set after training convergence, the 3D reconstruction result of the underground mine scene under the desired viewpoint is output through viewpoint adaptive rendering.
[0014] In addition, the underground mine scene reconstruction method based on non-uniform illumination of three-dimensional Gaussian sputtering according to the above embodiments of the present invention may also have the following additional technical features:
[0015] Furthermore, the step of initializing a three-dimensional Gaussian ellipsoidal attribute set containing dynamic structural features based on the sparse point cloud to obtain an initial three-dimensional Gaussian volume set for the mine scene includes:
[0016] Visual data of an underground mine scene is acquired using a camera under sparse perspective, and the visual data is processed to obtain multi-view results. Figure 2 3D image data, to extract from the multi-view image using a motion recovery structure algorithm. Figure 2 Generate sparse point clouds from 3D image data;
[0017] Three-dimensional scene reconstruction is performed based on the sparse point cloud using three-dimensional Gaussian sputtering technology;
[0018] In the initial stage of the reconstruction process, each point in the sparse point cloud is initialized with the geometric properties of a three-dimensional Gaussian ellipsoid, and the coverage of the Gaussian ellipsoid is expanded according to a preset ratio to adapt to the sparse point cloud scene.
[0019] Based on the RGB information of underground mine images, the vector chromaticity attribute is initialized through inverse Gamma correction, and the physical emissivity attribute is initialized by combining the color features of sparse point clouds and the low light characteristics of the mine.
[0020] The structural feature dimensions are determined based on the sparse point cloud density and scene type, and the initial values of the structural features are assigned.
[0021] Furthermore, according to the aforementioned multi-view Figure 2 The steps for extracting a physical prior structure map adapted to the mine environment from 3D image data, and using it as a supervision signal for Gaussian ellipsoidal geometric growth, include:
[0022] The physical prior theory based on photometric gradient is used to calculate the pixel brightness gradient of each pixel in the image, and the physical prior structure diagram is constructed by combining spatial differential operators.
[0023] For the completely black pixel area in the mine, the pseudo-structure signal is filled based on the neighborhood gradient to avoid invalid areas in the physical prior structure map;
[0024] For areas with strong reflection, a noise suppression factor is introduced to smooth the gradient, reducing the impact of reflection interference on the physical prior structure map.
[0025] Furthermore, the step of decoupling and rendering the initial three-dimensional Gaussian volume set to generate a vector chromaticity map, a physical radiation map, and a structural feature map includes:
[0026] Perform α-blending rasterization on a 3D Gaussian volume to generate a vector chromaticity map, a physical radiance map, and a structural feature map, respectively.
[0027] For high-noise regions in the vector chromaticity map whose physical radiance is below a preset threshold, their rendering weights are weighted and smoothed.
[0028] For structural feature maps, if multi-dimensional structural features are used, their channel comprehensive value is taken as the supervision signal; if single-dimensional structural features are used, a single-channel map is directly output to balance the structural supervision effect and computational efficiency.
[0029] Among them, the vector chromaticity map is used to characterize the intrinsic color of an object after the illumination is removed, the physical radiation map is used to characterize the true light field distribution including the influence of the light source, and the structural feature map is used to maintain geometric continuity in the dark region.
[0030] Furthermore, the steps of performing exposure compensation on the physical radiometer based on the camera parameters used when acquiring the visual data, and suppressing dark noise using an intensity collapse function to obtain a clean physical intensity map include:
[0031] Determine whether the camera parameters provide EXIF metadata;
[0032] If so, the exposure gain is calculated directly based on the exposure-related parameters;
[0033] If not, exposure parameters are calculated based on image statistical features, and outliers are removed using multi-view statistical information.
[0034] For the physical radiometric map after exposure compensation, the adjusted intensity collapse function is used for denoising. The pixel value truncation range in the function is adapted to the strong light source area of the mine. At the same time, parameters are set to adapt to the noise suppression of dark areas, resulting in a clean physical intensity map.
[0035] Furthermore, the step of reconstructing the clean physical intensity map and the vector chromaticity map into HSV components and inversely transforming them to obtain an RGB image includes:
[0036] The hue and saturation components are recovered from the vector chromaticity diagram, and the saturation is enhanced by a preset amplitude to improve the recognizability of mine equipment markings;
[0037] The clean physical intensity map is used as the luminance component and combined to obtain the HSV image;
[0038] The HSV image is inversely transformed into an sRGB image using a differentiable color space conversion function, and the sRGB image is optimized using a lightweight tone mapping module to ensure that the image color and detail meet the observation requirements of the mine scene.
[0039] Furthermore, the steps of constructing multiple loss constraint functions based on the RGB image, the structural feature map, the physical prior structure map, and the real image of the mine scene, and using the Adam optimization algorithm to perform end-to-end training and updating of the learnable parameters of the three-dimensional Gaussian volume until the model converges include:
[0040] Construct a total loss function consisting of photometric loss, structural consistency loss, and buoyancy suppression loss;
[0041] For both large-scale tunnel scenarios and close-range equipment scenarios, corresponding optimization parameters and iteration stages are configured respectively;
[0042] Configure the hyperparameters based on a combination of general and scenario-specific parameters, use the Adam optimization algorithm to update the learnable parameters of the Gaussian ellipsoid end-to-end, and perform Gaussian densification pruning and fine optimization in stages until the loss converges.
[0043] Specifically, the photometric loss is based on the difference between the RGB image and the real image of the mine scene, using a weighted combination of multiple loss types; the structural consistency loss is based on the difference constraint between the structural feature map and the physical prior structure map; and the buoyancy suppression loss introduces a gradient scaling strategy.
[0044] Another objective of this invention is to provide a system for reconstructing underground mine scenes based on non-uniform illumination using three-dimensional Gaussian sputtering, for implementing the aforementioned method for reconstructing underground mine scenes based on non-uniform illumination using three-dimensional Gaussian sputtering. The system includes:
[0045] The data acquisition and processing module is used to acquire visual data of the underground mine scene and process the visual data to obtain multi-view results. Figure 2 3D image data, to be based on the multi-view Figure 2 A sparse point cloud is generated from the 3D image data, and a set of 3D Gaussian ellipsoidal attributes with dynamic structural features is initialized based on the sparse point cloud to obtain an initial set of 3D Gaussian volumes for the mine scene.
[0046] The physical prior structure diagram determination module is used to determine the physical prior structure diagram based on the multi-view... Figure 2Physical prior structure diagrams adapted to the mine environment are extracted from 3D image data and used as a supervision signal for Gaussian ellipsoid geometric growth;
[0047] The decoupled rendering module is used to decouple and render the initial three-dimensional Gaussian volume set to generate a vector chromaticity map, a physical radiation map, and a structural feature map;
[0048] The clean physical intensity map determination module is used to perform exposure compensation on the physical radiometric map based on the camera parameters when the visual data is acquired, and to suppress dark noise through an intensity collapse function to obtain a clean physical intensity map.
[0049] An RGB image determination module is used to reconstruct the clean physical intensity map and the vector chromaticity map into HSV components and perform an inverse transformation to obtain an RGB image.
[0050] The model training module is used to construct multiple loss constraint functions based on the RGB image, the structural feature map, the physical prior structure map and the real image of the mine scene, and to use the Adam optimization algorithm to perform end-to-end training and update of the learnable parameters of the three-dimensional Gaussian ellipsoid until the loss converges.
[0051] The 3D reconstruction module is used to output the 3D reconstruction results of the underground mine scene from the desired viewpoint through viewpoint adaptive rendering based on the parameters of the Gaussian body set after training convergence.
[0052] Another objective of this invention is to provide a storage medium storing a computer program that, when executed by a processor, implements the steps of the above-described method for reconstructing an underground mine scene based on non-uniform illumination using three-dimensional Gaussian sputtering.
[0053] Another objective of this invention is to provide an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the program to implement the steps of the above-described method for reconstructing an underground mine scene based on non-uniform illumination using three-dimensional Gaussian sputtering.
[0054] This invention acquires visual data from underground mine scenes and processes it into multi-view images. Figure 2This method uses 3D image data to generate sparse point clouds, and then initializes a 3D Gaussian ellipsoid attribute set containing dynamic structural features to obtain an initial 3D Gaussian volume set for the mine scene. Simultaneously, it extracts a physical prior structure map adapted to the mine environment as a supervision signal for Gaussian ellipsoid geometric growth. This effectively solves the problems of existing motion structure recovery algorithms generating sparse and noisy initial point clouds that cannot support accurate optimization, and the lack of effective structural supervision during Gaussian volume geometric growth that easily leads to structural breakage in the reconstruction. Furthermore, by decoupling the initial 3D Gaussian volume set through rendering to generate vector chromaticity maps, physical radiation maps, and structural feature maps, it removes the interference of lighting on the intrinsic colors of objects, overcoming the limitations of existing techniques that use RGB. Color spaces result in strong coupling between color and lighting, making it easy for models to incorporate lighting effects into geometric structures, leading to color flickering and distortion in reconstructed images. This paper addresses the problem of inconsistent exposure in multi-view images of mines, which easily produces floating artifacts and geometric collapse. Furthermore, sensor noise and dust scattering noise are amplified during rendering, making it difficult for traditional denoising methods to balance detail preservation and noise suppression. The paper further describes the process of reconstructing clean physical intensity maps and vector chromaticity maps into HSV components and inversely transforming them to obtain RGB images. Multiple loss constraint functions are constructed based on the RGB images, structural feature maps, physical prior structure maps, and real images of the mine scene, employing Adam... The optimization algorithm performs end-to-end training and updates of the learnable parameters of the 3D Gaussian ellipsoid until the loss converges, overcoming the shortcomings of existing technologies such as the lack of effective multi-dimensional constraints in model training and insufficient reconstruction accuracy and robustness to extreme mine environments. By performing viewpoint adaptive rendering based on the converged Gaussian volume set parameters, it outputs the 3D reconstruction result of the underground mine scene from the desired viewpoint, solving the problems of slow training speed and insufficient real-time rendering capabilities of NeRF and its variants in existing technologies, which cannot meet the needs of rapid reconstruction and real-time navigation in mine sites. Therefore, this invention solves the problem of the lack of a robust, high-accuracy, and real-time underground mine scene reconstruction method based on non-uniform illumination using 3D Gaussian sputtering. Attached Figure Description
[0055] Figure 1 This is a flowchart of the underground mine scene reconstruction method based on non-uniform illumination of three-dimensional Gaussian sputtering in the first embodiment of the present invention;
[0056] Figure 2 This is a schematic diagram of the underground mine scene reconstruction system based on non-uniform illumination of three-dimensional Gaussian sputtering according to the second embodiment of the present invention.
[0057] Figure 3 This is a schematic diagram of the structure of the electronic device in the third embodiment of the present invention;
[0058] The following detailed description, in conjunction with the accompanying drawings, will further illustrate the present invention. Detailed Implementation
[0059] To facilitate understanding of the present invention, a more complete description will be given below with reference to the accompanying drawings. Several embodiments of the invention are illustrated in the drawings. However, the invention can be implemented in many different forms and is not limited to the embodiments described herein. Rather, these embodiments are provided so that this disclosure will be thorough and complete.
[0060] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this invention pertains. The terminology used herein in the description of the invention is for the purpose of describing particular embodiments only and is not intended to be limiting of the invention. The term "and / or" as used herein includes any and all combinations of one or more of the associated listed items.
[0061] Example 1
[0062] Please see Figure 1 The figure shows a method for reconstructing an underground mine scene based on non-uniform illumination using three-dimensional Gaussian sputtering in the first embodiment of the present invention. The method specifically includes steps S1-S7.
[0063] S1, acquire visual data of the underground mine scene, and process the visual data to obtain multi-view... Figure 2 3D image data, to be based on the multi-view Figure 2 A sparse point cloud is generated from the 3D image data, and a set of 3D Gaussian ellipsoidal attributes containing dynamic structural features is initialized based on the sparse point cloud to obtain an initial 3D Gaussian volume set for the mine scene.
[0064] Specifically, visual data of an underground mine scene is acquired using a camera under sparse perspective, and the visual data is processed to obtain multi-view... Figure 2 3D image data, to extract from the multi-view image using a motion recovery structure algorithm. Figure 2 A sparse point cloud is generated from the 3D image data; a 3D scene is reconstructed based on the sparse point cloud using 3D Gaussian sputtering technology.
[0065] In the initial stage of the reconstruction process, each point in the sparse point cloud is initialized with the geometric properties of a three-dimensional Gaussian ellipsoid, and the coverage of the Gaussian ellipsoid is expanded according to a preset ratio to adapt to the sparse point cloud scene. Based on the RGB information of the underground mine image, the vector chromaticity attribute is initialized through inverse Gamma correction, and the physical emissivity attribute is initialized by combining the color features of the sparse point cloud and the low light characteristics of the mine. Based on the density of the sparse point cloud and the scene type, the dimension of the structural features is determined and the initial value of the structural features is completed.
[0066] In specific implementation, S1.1, mining scene data is collected by using an intrinsically safe smart handheld terminal for mining (such as an explosion-proof mobile phone equipped with a high-sensitivity module) or a camera of an inspection robot.
[0067] Specifically, the acquisition method supports two modes: one is to directly capture still photos; the other is to record a video stream and then perform frame extraction at a fixed frame rate (e.g., 5 frames / second). During the acquisition process, the device moves along the tunnel or around the device to acquire multiple views. Figure 2 A sequence of 3D images is used to ensure a preset overlap rate between adjacent views, acquiring a total of 50-200 images. The image resolution is normalized to 1920×1080 or 2048×1536.
[0068] S1.2 The above image sequence is processed using the Structure from Motion (SfM) algorithm.
[0069] Specifically, feature points are extracted from each image using feature point extraction algorithms (such as ORB, SIFT, etc.), then mismatched feature points are removed using the RANSAC algorithm, and the camera pose (including position vector) is calculated based on the correspondence of the remaining feature points. and rotation matrix , (for view indexing) and generating sparse point clouds ,in , For the first The spatial coordinates of the points The total number of point clouds is 2-15 points / m². Although the point cloud generated by SfM is sparse and noisy, the present invention can overcome the problem of poor initialization quality of SfM by using a density enhancement strategy guided by structural features (S6.2) and a physical prior structure graph (S2).
[0070] S1.3. Based on the sparse point cloud generated in S1.2, each point... Expanded into a three-dimensional Gaussian body Initialize the set of three-dimensional Gaussian volumes Among them, the set of three-dimensional Gaussian volumes is initialized. Notation:
[0071]
[0072] Among them, the superscript " " indicates the initial stage (i.e., 0 iterations), subscript Indicates the first A Gaussian body; This represents the geometric properties of the Gaussian body in its initial stage; This represents the vector chromaticity attribute in the initial stage; This indicates the physical emissivity attribute in the initial stage; Indicates the structural characteristic attributes of the initial stage;
[0073] (1) Geometric properties ,
[0074] Geometric properties of Gaussian solids, defining the nth Gaussian solid. The probability distribution at any point x in three-dimensional space:
[0075]
[0076] Among them, parameters Gaussian body The initial values of the coordinates of the center point in three-dimensional space. for The coordinates are then iteratively updated; where the parameters... Representing a three-dimensional Gaussian body The learnable covariance matrix is derived from the scaling vector. and rotation quaternions Construct and describe the ellipsoidal shape, size, and orientation of the Gaussian body in the three-dimensional space of an underground mine tunnel. The formula for calculating the covariance matrix is: ;parameter The upper limit of opacity is represented by a learnable parameter. Obtained by mapping using the Sigmoid function. in, Use the Sigmoid activation function;
[0077] (2) Vector chromaticity attribute ,
[0078] Extracting corresponding points from the acquired image in S1.1 RGB pixel values at the projection position Illumination coupling is removed by inverse Gamma correction, and the correction formula is:
[0079] ,
[0080] Convert linear RGB to HSV space and extract hue. and saturation The initial vector chromaticity coordinates are obtained through vector transformation: ,make sure , To avoid numerical breaks in the red area.
[0081] (3) Physical emissivity attribute ,
[0082] Take the maximum channel value of linear RGB in S1.3(2). Adaptability coefficient for low-light scenes in mines ,in The initialization formula is: ,ensure This enhances signal strength in low-light areas.
[0083] (4) Structural characteristics and attributes ,
[0084] Based on the point cloud density ( S The area to be reconstructed for the mine scene, unit Adaptive selection of dimensions based on scene type Specifically:
[0085] when Or when the scene is a large-scale tunnel scene, .
[0086] .
[0087] The set of indices of all Gaussian solids intersecting point x After sorting by depth from near to far, and combining the vector chromaticity attributes With physical emissivity properties The color at point x is obtained using forward blending:
[0088] ,
[0089] in, The transparency of the preceding Gaussian volume is multiplied to ensure the correct occlusion relationship between equipment and tunnels in the mine scene.
[0090] S1.4 After initializing the 3D Gaussian volume set, it is necessary to use differentiable rendering and parameter optimization to gradually match the real geometry and appearance features of the mine scene.
[0091] Specifically, a 3D Gaussian volume is projected onto the 2D image plane of each viewpoint acquired in step S1 using differentiable rendering. The difference between the reconstructed image and the input real mine image is used as the reconstruction loss. The Adam optimization algorithm is used to iteratively optimize the geometric and attribute parameters of each 3D Gaussian volume until the loss converges, thus obtaining an initial set of 3D Gaussian volumes adapted to the mine scene. Further, in the t-th iteration:
[0092] (1) Gaussian projection: using the camera projection function of the k-th view The center point of the three-dimensional Gaussian solid Projecting onto a two-dimensional pixel plane yields the projection position. :
[0093] ,
[0094] in, This is the camera intrinsic parameter matrix (obtained from the calibration of the industrial camera). Let k be the extrinsic parameter matrix of view. It is the rotation matrix of the camera coordinate system relative to the world coordinate system. , It is a translation vector. Both describe the camera's shooting posture in the mine scene; Represents the homogeneous coordinates of the center point of the Gaussian body, used for linear operations adapted to projection transformation; the superscript "(t)" indicates t iterations;
[0095] (2) Calculate the two-dimensional covariance matrix after projection: calculate the projection function. right Jacobian matrix Combined with the transformation matrix from the world coordinate system to the camera coordinate system , three-dimensional covariance Convert to two-dimensional covariance :
[0096] ,
[0097] Take its top-left 2×2 submatrix as the two-dimensional covariance after projection:
[0098] ,
[0099] Wherein, projection function It is a linear transformation, and its Jacobian matrix is... Equivalent to camera intrinsic parameter matrix The first two lines describe the effect of minute changes in three-dimensional coordinates on the two-dimensional projection position. It is a 3×3 matrix after three-dimensional covariance attitude transformation and projection transformation; the symbol “[ ]” “1:2,1:2” indicates that the first and second rows and the first and second columns of the matrix are extracted;
[0100] (3) Calculate pixel-level opacity:
[0101] Opacity The calculation is based on a two-dimensional Gaussian kernel, and the formula is:
[0102]
[0103] in, Let k be the two-dimensional pixel coordinates of view k; For a two-dimensional Gaussian kernel term, it represents the three-dimensional Gaussian volume at pixel coordinates. The degree of coverage at a given location indicates that the closer the value of the two-dimensional Gaussian kernel term is to 1, the greater the contribution of the Gaussian volume to that pixel. It is the upper limit of the opacity of the Gaussian body, which controls its maximum contribution intensity.
[0104] (4) Calculate the pixel coordinates in the two-dimensional image plane Color of the location:
[0105] color The calculation is performed by combining vector chromaticity with physical radiance, followed by α-blending, as shown in the formula:
[0106]
[0107] in, To cover pixel coordinates The Gaussian volume index set; It is the fusion of vector chromaticity (color) and physical emissivity (brightness) to obtain the complete color information corresponding to the Gaussian body; It is the cumulative product of the opacity of the preceding Gaussian volume—for example, if the opacity of the preceding Gaussian volume is 0.5, then the current Gaussian volume can only transmit 50% of the brightness, thus realizing the realistic spatial occlusion relationship of "near equipment occluding the far tunnel wall" in the mine scene.
[0108] (5) Reconstruction Loss Calculation: Reconstruct the scene of the specified view in the two-dimensional image plane and calculate the reconstruction loss. For the view... Corresponding original image Let its corresponding reconstructed image be ,definition The loss function measures the difference between the reconstructed image and the original image. In one embodiment of the present invention, the formula for the loss function is:
[0109] ,
[0110] in, For the total number of views, Let k be the set of pixels. The total number of pixels across all views;
[0111] (6) Parameter update: Calculate the parameters using the backpropagation algorithm. about The gradients are calculated, and these parameters are updated using the Adam optimization algorithm:
[0112] ,
[0113] in, For a set of learnable parameters, For learning rate, Used to adjust the position of the Gaussian volume to match the sparse point cloud. Used to optimize color and prevent red device logos from breaking. Used to enhance brightness in low-light areas Used to enhance brightness in low-light areas Used to adjust the shape and size of a Gaussian body to fit the surface of a real object; Used to adjust the opacity of a Gaussian volume to reflect the spatial occlusion relationships of objects; Used to fit local structural features to maintain geometric continuity.
[0114] (7) Iteration Termination: Let Then, by repeating steps (1)-(6) until the loss converges, assuming the iteration terminates after a certain number of iterations. The initialized set of three-dimensional Gaussian volumes is obtained. :
[0115]
[0116] S2, according to the multi-view Figure 2 Physical prior structure diagrams adapted to the mine environment are extracted from 3D image data and used as a supervision signal for the geometric growth of Gaussian ellipsoids.
[0117] Specifically, the pixel brightness gradient of each pixel in the image is calculated based on the physical prior theory of photometric gradient, and a physical prior structure map is constructed by combining the spatial differential operator; for the all-black pixel area of the mine, the pseudo-structure signal is filled based on the neighborhood gradient to avoid invalid areas in the physical prior structure map; for the strongly reflective pixel area, a noise suppression factor is introduced to smooth the gradient and reduce the impact of reflection interference on the physical prior structure map.
[0118] In practical implementation, for each two-dimensional image in the multi-view set described in step S1, an illumination-invariant physical prior structure map is extracted using the physical prior theory based on photometric gradients and the Sobel spatial differential operator. This provides precise structural monitoring signals for subsequent Gaussian body geometric growth, ensuring that the Gaussian body distribution conforms to the real-world structures such as tunnel walls and equipment. The processing flow is as follows:
[0119] S2.1 Calculate pixel brightness values and gradients: For each mine image By pixel Calculate pixel brightness value , The calculation formula is:
[0120] ,
[0121] in, For pixels The RGB channel values are calculated using a weighted formula that aligns with the human eye's sensitivity to the brightness of the three channels, avoiding fluctuations in the values of a single channel from affecting the calculation of pixel brightness values, thus meeting the brightness extraction needs under low-light conditions in mines.
[0122] Then, three types of gradients are further calculated to fully capture scene structure (device edges, tunnel textures, etc.):
[0123] First gradient Along pixel space Direction calculation is performed using the Sobel operator;
[0124] Second gradient :along The pixel step size is used to calculate the second-order gradient in space, focusing on the finer structural textures in the scene (such as cracks in tunnel walls and the texture of device screws).
[0125] Third gradient :along The pixel step size is used to calculate the third-order gradient in space, which helps to enhance the features of deep structures and compensate for the loss of features of small components in low-light environments.
[0126] S2.2 Generate a Physical Prior Structure Map (PGT) and calculate pixels according to the formula. Physical prior structure diagram value Integrate the structural information of the three types of gradients:
[0127]
[0128] in, It is a pixel brightness value estimation. It is a discrete gradient operator (such as Sobel); Set the gradient weight coefficients. The first-order gradient weight is set to 1.0 by default, which is used to extract the edges of major structures such as equipment and tunnels. The second and third-order gradient weights are lower to avoid excessive amplification of subtle noise and to balance the contribution ratio of different orders of gradients.
[0129] S2.3 Targeted optimizations were made to address the unique characteristics of underground mine scenarios, such as completely dark areas and strong reflection interference:
[0130] Processing completely black pixels: This is prone to occur deep within well tunnels. For all-black pixels, fill with the average gradient of a 3×3 neighborhood, using the following formula:
[0131]
[0132] in, For pixels of 3×3 Neighborhood sets, by using neighborhood information to complete the structure of the entirely black region, ensure the integrity of the structure graph; if 3×3 If the neighborhood is invalid, expand the neighborhood range or use the global mean.
[0133] Processing of highly reflective pixels: These pixels are prone to causing issues on the metal surfaces of mining equipment. Highly reflective pixels can introduce false gradient noise. This can be addressed by introducing a noise suppression factor. The gradient is smoothed using Gaussian blur, and the formula is:
[0134]
[0135] in, Standard deviation Gaussian kernel, satisfying Gaussian blur can soften abrupt gradients in regions of strong reflection. Suppress noise while preserving the structure.
[0136] The final set of physical prior structure diagrams for all views is as follows: , will this As a Ground Truth, it provides reliable input for the supervision of Gaussian structure.
[0137] S3, decouple and render the initial three-dimensional Gaussian volume set to generate a vector chromaticity map, a physical radiation map, and a structural feature map.
[0138] Specifically, an α-blending rasterization operation is performed on the 3D Gaussian volume to generate a vector chromaticity map, a physical radiometric map, and a structural feature map. For high-noise regions in the vector chromaticity map where the physical radiometric rate is below a preset threshold, the rendering weights are weighted and smoothed. For the structural feature map, if multi-dimensional structural features are used, their channel composite value is taken as the supervision signal; if single-dimensional structural features are used, a single-channel map is directly output to balance the structural supervision effect and computational efficiency. Among these, the vector chromaticity map is used to characterize the intrinsic color of the object after the illumination is stripped, the physical radiometric map is used to characterize the true light field distribution, including the influence of the light source, and the structural feature map is used to maintain geometric continuity in dark regions.
[0139] In practical implementation, to achieve independent optimization of the color, brightness, and structural features of the mine scene, the initial three-dimensional Gaussian volume set obtained in S1.4 is... The process involves performing α-blending rasterization, rendering in three branches to obtain a vector chromaticity map, a physical radiosity map, and a structural feature map. This ensures the purity of each feature while maintaining scene hierarchy through a uniform opacity contribution. The processing flow and calculation method are as follows:
[0140] S3.1, Definition of the first A Gaussian solid for the view Pixels Opacity contribution The effective intensity of the Gaussian body after considering preceding occlusion is used to calculate its opacity contribution.
[0141] ,
[0142] in, The final opacity calculated in S1.4(3); product term Multiply the transparency of the preceding Gaussian volume to ensure that the spatial logic of "near equipment occluding the far tunnel wall" in the mine scene remains consistent during rendering.
[0143] S3.2 Based on the calculated weights The rendering of the following three branch feature maps is performed in parallel.
[0144] (1) Vector chromaticity map (HV-Map) rendering: In order to remove the interference of brightness on color, a 2-channel pure chromaticity feature map containing hue and saturation information is rendered. 、 The formula is:
[0145]
[0146]
[0147] in, To cover pixels A set of Gaussian volume indices; in low-light conditions in underground mines. The region is prone to chromatic noise. Multiply by 0.8 for weighted smoothing to prevent noise from being amplified during rendering and to ensure the stability of key colors such as device identification colors.
[0148] (2) Rendering of physical radiance map, rendering a single-channel physical intensity map. Focusing on real-world information, the formula is:
[0149]
[0150] This image directly reflects the physical light intensity distribution of the mine scene, providing a clean brightness basis for subsequent exposure correction and noise reduction, and avoiding correction deviations caused by interference between color and brightness.
[0151] (3) Rendering of Structure-Map: Rendering feature maps used for structure supervision. According to structural feature dimensions Adaptable to different mining scenarios:
[0152] when When adapting to complex equipment structures (such as motor and pipe combinations), the average of the three channels is used as the monitoring signal to improve the richness of structural representation. The formula is as follows:
[0153]
[0154] when Time: Adapts to simple tunnels and large-scale scenarios, directly outputs a single-channel image, reducing computational complexity. The formula is:
[0155]
[0156] in, When the iteration terminates (i.e., the 1st iteration) t * The optimization of the (nth iteration) n The structural characteristic attribute values of a Gaussian body.
[0157] S4. Based on the camera parameters when acquiring the visual data, the physical radiometric map is exposed and compensated, and dark noise is suppressed by the intensity collapse function to obtain a clean physical intensity map.
[0158] Specifically, it is determined whether the camera parameters provide EXIF metadata; if so, the exposure gain is directly calculated based on the exposure-related parameters; if not, the exposure parameters are calculated based on image statistical features, and outliers are removed using multi-view statistical information; for the physical radiometric map after exposure compensation, the adjusted intensity collapse function is used for denoising, the pixel value truncation range in the function is adapted to the strong light source area of the mine, and parameters adapted to suppress noise in dark areas are set to obtain a clean physical intensity map.
[0159] The physical radiation map obtained in step S3 represents the true light intensity of the scene, but the acquired training images are affected by camera exposure parameters and sensitivity noise, resulting in domain differences between the two. This step unifies the brightness benchmark through exposure gain correction and suppresses dark noise through intensity collapse technology, thereby providing high-quality brightness data for subsequent image reconstruction.
[0160] In practical implementation, S4.1, considering the diversity of mine image acquisition equipment, the exposure gain is calculated in two cases: whether or not EXIF metadata exists.
[0161] 1) When EXIF metadata exists: Read the EXIF parameters (exposure time) of view k. ,aperture ISO ), calculate the exposure gain according to the formula :
[0162] ,
[0163] in, This is a preset adjustment coefficient for the mine scene; in the low-light environment of a mine, the camera often uses long exposure and high ISO settings; in this embodiment, it is set to... Gain can be appropriately reduced to avoid noise amplification caused by high sensitivity.
[0164] 2) When EXIF metadata is missing, the exposure gain is estimated using the image's own brightness features. The steps are as follows:
[0165] 1. Calculate the global mean brightness of the image. Dark pixel ratio Peak value of strong light source ;
[0166] 2. Obtain the estimated exposure gain value through the fitting formula: The higher the proportion of dark areas and the lower the average brightness relative to the peak value, the greater the gain, which is suitable for the low light characteristics of mines.
[0167] 3. Outlier Correction: For Outliers are corrected using linear interpolation:
[0168] ,
[0169] in, For all views of the median.
[0170] S4.2 Based on the calculated gain, the physical radiation map generated in step S3 is first converted into an intensity map simulating the image received by the camera sensor. .
[0171] ,
[0172] Then, in order to remove sensor thermal noise in low-light environments, an intensity collapse function is introduced. ,right Nonlinear filtering is performed to obtain a pure brightness image. :
[0173] ,
[0174] in, Will Truncate to the [0,1] range to prevent overexposure from overflowing. hour, This enables noise collapse in dark areas, while the sine function can enhance details in low-brightness areas, making it suitable for effective information extraction in low-light environments like mines.
[0175] S5, the clean physical intensity map and the vector chromaticity map are recombined into HSV components, and the RGB image is obtained by inverse transformation.
[0176] Specifically, the hue and saturation components are recovered from the vector chromaticity map, and the saturation is enhanced by a preset amplitude to improve the recognizability of mine equipment markings; the clean physical intensity map is used as the luminance component and combined to obtain an HSV image; the HSV image is inversely transformed into an sRGB image through a differentiable color space conversion function, and a lightweight tone mapping module is used to optimize the sRGB image to ensure that the image color and details meet the observation requirements of the mine scene.
[0177] After obtaining the pure chromaticity features (from S3) and the denoised luminance features (from S4), this step uses HSV color space reconstruction and tone mapping networks to restore the true colors and visual effects of the mine scene. In specific implementation, step S5 includes:
[0178] S5.1. Recombine the decoupled features into the three components of the HSV color space and perform HSV component recovery:
[0179] 1) Hue: Deduced from the vector chromaticity diagram to avoid breaks in the red marking area of mining equipment. The formula is:
[0180] ,
[0181] 2) Saturation: Calculated using a formula to enhance the visibility of equipment identification colors, addressing the issue of low saturation in low-light conditions in mines. The formula is:
[0182] ,
[0183] Among them, 1.1 is the saturation enhancement coefficient, which improves the distinction between the equipment identification color and the background without amplifying noise.
[0184] 3) Brightness (Value): Directly use the denoised brightness map output in step S4. :
[0185] ,
[0186] After obtaining the HSV components in S5.2, a differentiable 3D LUT (i.e., a three-dimensional lookup table) trilinear interpolation function is used. The formula for inverse transforming the HSV components into an sRGB image is as follows:
[0187] ,
[0188] This process preserves gradient transitivity, enabling the entire rendering pipeline to support end-to-end training.
[0189] Finally, to further optimize the visual effects, a network structure of " "Lightweight tone mapping network" Process and optimize images The aim is to fine-tune local tonal differences and normalize pixel values to the [0,1] display range. The final output is the predicted image. It takes into account the accurate reproduction of colors, the uniform distribution of brightness, and the clear presentation of structure, providing high-quality predictive input for the loss calculation of S6.
[0190] S6. Based on the RGB image, the structural feature map, the physical prior structure map, and the real image of the mine scene, construct a multi-loss constraint function, and use the Adam optimization algorithm to perform end-to-end training and update of the learnable parameters of the three-dimensional Gaussian ellipsoid until the loss converges.
[0191] Specifically, a total loss function is constructed, consisting of photometric loss, structural consistency loss, and buoyancy suppression loss. For both large-scale tunnel scenarios and close-range equipment scenarios, corresponding optimization parameters and iteration stages are configured. Based on a combination of general and scenario-specific hyperparameters, the Adam optimization algorithm is used to update the learnable parameters of the Gaussian ellipsoid end-to-end. Gaussian pruning and fine-tuning are performed in stages until the loss converges. Specifically, the photometric loss uses a weighted combination of multiple loss classes based on the difference between the RGB image and the real image of the mine scene; the structural consistency loss is constrained by the difference between the structural feature map and the physical prior structure map; and the buoyancy suppression loss incorporates a gradient scaling strategy.
[0192] To drive accurate updates of the aforementioned model parameters (including Gaussian volume properties, tone mapping network parameters, etc.), this step constructs a multi-loss function encompassing photometric, structural, and regularized values, and employs a phased optimization strategy to ensure the model's convergence and robustness in complex mining scenarios. In specific implementation, S6.1 first defines the total loss function. It consists of photometric loss, structural consistency loss, and regularization loss, and the formula is:
[0193] ,
[0194] Among them, settings These are the weighting coefficients; It can enhance structural monitoring and adapt to the structural accuracy requirements of mining scenarios. A lower value can smooth parameters while avoiding excessive suppression of details. The calculation of each loss term is as follows:
[0195] 1) Photometric loss Combining L1 loss and D-SSIM loss, the effect of noise in mine images is suppressed by the following formula:
[0196] ,
[0197] in, , The images used are real images of the mine scene (or high-fidelity images collected); L1 loss is robust to noise, and D-SSIM loss focuses on structural similarity. The combination of the two ensures pixel-level accuracy while avoiding structural distortion.
[0198] 2) Structural consistency loss The difference between the constrained rendering structure graph and the physical prior structure graph is used to ensure that the Gaussian volume distribution fits the real scene structure. The formula is:
[0199] ,
[0200] Where H and W are the image height and width, respectively, the rendering structure is forced to align with the physical prior by using the global pixel average error.
[0201] 3) Regularization loss To prevent buoyancy artifacts caused by sudden parameter changes during training, L2 smoothing constraints are applied to the physical radiance and vector chromaticity. The formula is as follows:
[0202] ,
[0203] in, The total number of Gaussian bodies is determined by penalizing excessively large parameter values to ensure a smooth transition of Gaussian body properties.
[0204] S6.2 This embodiment designs a two-stage optimization strategy from coarse to fine, utilizing the physical prior structure diagram generated in step S2 ( As a guiding signal throughout:
[0205] In the first phase (rounds 1-5000), geometric topology construction and structure-oriented densification are performed.
[0206] 1) Weighted structure-guided densification: changing the traditional reliance solely on location gradients Encryption logic, introducing As a weight for structural confidence, the weighted gradient norm is calculated. :
[0207] ,
[0208] in, For Gaussian solid projection coordinates, For enhancement coefficient;
[0209] Splitting strategy: When the Gaussian body satisfies When the gradient threshold (set to 0.0002 in the example) is greater than 1% of the point cloud spacing, it is split into two Gaussian bodies.
[0210] 2) Aggressive pruning based on opacity:
[0211] To remove "floating artifacts" caused by mine dust and sparse viewpoints, a double pruning process was performed:
[0212] Threshold pruning: Every 100 iterations, remove the upper limit of opacity. A Gaussian body (set to 0.005).
[0213] Periodic reset: Every 3000 iterations, the opacity of all Gaussian volumes is reset to 0.01. This forces the model to "compete" for survival again through the geometric constraints of step S1, effectively clearing accumulated translucent smoke-like artifacts.
[0214] In the second phase (rounds 5001-15000), the geometry locking and physical appearance decoupling are refined:
[0215] 1) Geometric parameter locking mechanism: Stop all densification and pruning operations, and freeze all spatial geometric parameters of the 3D Gaussian body (including position coordinates). Rotation Quaternions With scaling vector This prevents it from accepting gradient updates.
[0216] Only optimized vector chromaticity is available Physical emissivity Structural features and tone mapping network The weight.
[0217] 2) Based on Spatial adaptive gradient scaling:
[0218] Pixel-level gradient scaling mask is constructed using the physical prior structure map generated by S2. Differentiated learning is implemented for different frequency domain regions:
[0219] ,
[0220] in, for High quantile threshold, This is the lower bound of the basic learning rate.
[0221] Device / Edge Area ( high): It allows for full-intensity gradient backhaul, ensuring clear restoration of high-frequency details such as dashboard readings and warning signs.
[0222] Tunnel wall / ground area ( Low): It attenuates the gradient amplitude, thus playing a regularization role and smoothing out the sensor's chromaticity noise under low light conditions.
[0223] 3) Cosine annealing scheduling in the time dimension:
[0224] During the iteration process, the learning rate for the open parameters... Cosine annealing attenuation:
[0225] ,
[0226] in, This is the initial refinement learning rate. for The percentage is 1%. This strategy ensures that the model can quickly adapt to the adjustment of lighting features in the early stage of the phase, and converge slightly at the end of the phase, avoiding parameter oscillation and obtaining delicate rendering quality.
[0227] Iteration terminates: when The fluctuation is less than 500 consecutive iterations Training should be stopped when the time comes.
[0228] The final optimized set of three-dimensional Gaussian volumes is obtained. Its definition is:
[0229] ,
[0230] in, This represents the total number of Gaussian bodies retained after phased optimization. The optimized geometric properties are designed to fit the actual spatial structure of mine tunnels and equipment. The color features reproduce the equipment identification colors and the natural colors of the environment. The structural properties are highly aligned with the physical priors, thus meeting the high-precision 3D reconstruction requirements of mine scenarios.
[0231] S7, based on the Gaussian body set parameters after training convergence, outputs the 3D reconstruction result of the underground mine scene from the desired viewpoint through viewpoint adaptive rendering.
[0232] Specifically, after training, based on step S6, the final optimized set of three-dimensional Gaussian volumes includes four decoupled attributes: geometry, chromaticity, physics, and structure. and the 3D LUT interpolation function trained in step S5. and tone mapping network It performs inverse color space transformation and viewpoint adaptive rendering to output high-precision, highly robust 3D data of the mine scene. The specific process includes:
[0233] S7.1 Rendering attribute loading and view space mapping: For any target viewpoint specified by the user, obtain its camera extrinsic matrix. and intrinsic parameter matrix Traversing the set Each Gaussian body Using the Gaussian projection formula in step S1, calculate its two-dimensional projection center from the target's perspective. Two-dimensional covariance matrix and depth-based sorted indexes .
[0234] S7.2 Parallel rasterization of decoupled features, based on sorted Gaussian volume index, utilizing... - The blending technique renders two core decoupled branches in parallel to obtain intermediate feature maps from the target's perspective:
[0235] 1) Vector Chromaticity Branch: Utilizing optimized vector chromaticity attributes The hue feature map from the target viewpoint is obtained by rendering. and saturation feature map ;
[0236] 2) Physical radiation branch: Utilizing the optimized physical emissivity attribute Rendering yields a pure physical intensity map from the target's perspective. The formula is expressed as:
[0237]
[0238] in, The opacity of the Gaussian body under the target viewpoint is assigned a weight, and the weight is calculated in the same way as the opacity contribution formula in S3.1 to ensure consistency across steps.
[0239] S7.3 HSV spatial reconstruction reassembles the three independent channel features obtained in S7.2 into a complete HSV image. At this stage, the system can enhance specific areas of the mine according to actual needs (e.g., artificially improve...). The gain is used to simulate the effect of turning on a searchlight, or to enhance... (To highlight warning signs)
[0240] S7.4 Inverse Color Transformation and Final Imaging: To convert the HSV image in physical space into an sRGB image visible to the human eye and eliminate the non-uniform lighting artifacts unique to the mining environment, The colors are passed through the trained color mapping module sequentially.
[0241] 1) Inverse color space transformation: Trilinear interpolation is performed using the 3D LUT function optimized in step S5.2.
[0242] 2) Tone mapping and detail enhancement: Input lightweight tone mapping network Output the final high-fidelity reconstructed image. : .
[0243] Based on the above embodiments, this invention proposes a high-fidelity 3D reconstruction and multi-mode rendering method for non-uniform lighting and weak texture environments in underground mines. This method uses 3D Gaussian sputtering as the explicit scene representation basis, combined with color space decoupling and lighting physical properties, and sequentially executes the following steps: First, preprocessing the mine multi-view data initializes a four-dimensional decoupled Gaussian volume set containing geometry, vector chromaticity, physical emissivity, and structural features; second, extracting a lighting-independent physical prior structure diagram based on the physical prior theory of photometric gradients to provide strong constraints for dark geometry; third, constructing a physically-based decoupled rendering pipeline, explicitly simulating camera exposure and lighting response, and combining... Guided Gaussian geometric growth and pruning are performed to complete chroma-luminance reorganization and inverse transformation. Scene parameters are jointly optimized end-to-end through photometric loss, structural consistency loss, and HSV space regularization. Finally, based on the optimized decoupling properties, inverse HSV transformation and intensity collapse are performed to achieve illumination-independent pure scene extraction (lighting-free mode) and physically driven virtual relighting (simulation mode), and cross-view... Figure 1Consistent restoration of low-confidence textures. Through the above process, this invention can achieve geometrically accurate and texture-clear 3D reconstruction in complex mining environments with interference from moving miner lamps and frequent dynamic shadows; effectively solves the problem of traditional 3DGS generating a large number of floating artifacts and texture noise in low-light environments; successfully achieves visual enhancement and detail restoration of overexposed or extremely dark areas by decoupling physical properties; and further supports low-cost real-time relighting editing of reconstructed scenes.
[0244] This technology improves reconstruction accuracy in extreme underground mine environments. By decoupling Gaussian sphere properties and optimizing physical prior extraction, it addresses the low reconstruction accuracy issues in non-uniform lighting and sparse point cloud scenarios in mines, achieving scene reconstruction with centimeter- to millimeter-level accuracy. It also ensures real-time application requirements by balancing reconstruction accuracy and computational efficiency through dynamic structural feature dimensions and lightweight processing, meeting the real-time rendering needs of scenarios such as mine inspection and navigation. Furthermore, it enhances environmental robustness by employing exposure compensation adaptation and intensity collapse denoising schemes to improve the model's adaptability to extreme conditions such as low light, strong reflection, and noise in mines, reducing reconstruction artifacts and structural fractures. This technology meets the urgent application needs in fields such as smart mine digital twins, underground disaster emergency drills, and navigation planning for unmanned inspection robots.
[0245] Example 2
[0246] Please see Figure 2 The diagram shows a structural block diagram of the underground mine scene reconstruction system based on non-uniform illumination using 3D Gaussian sputtering, as proposed in the second embodiment of the present invention. This underground mine scene reconstruction system 200 based on non-uniform illumination using 3D Gaussian sputtering includes: a data acquisition and processing module 21, a physical prior structure diagram determination module 22, a decoupled rendering module 23, a clean physical intensity map determination module 24, an RGB image determination module 25, a model training module 26, and a 3D reconstruction module 27, wherein:
[0247] The data acquisition and processing module 21 is used to acquire visual data of the underground mine scene and process the visual data to obtain multi-view data. Figure 2 3D image data, to be based on the multi-view Figure 2 A sparse point cloud is generated from the 3D image data, and a set of 3D Gaussian ellipsoidal attributes with dynamic structural features is initialized based on the sparse point cloud to obtain an initial set of 3D Gaussian volumes for the mine scene.
[0248] The physical prior structure diagram determination module 22 is used to determine the physical prior structure diagram based on the multi-view Figure 2 Physical prior structure diagrams adapted to the mine environment are extracted from 3D image data and used as a supervision signal for Gaussian ellipsoid geometric growth;
[0249] Decoupled rendering module 23 is used to decouple and render the initial three-dimensional Gaussian volume set to generate a vector chromaticity map, a physical radiation map and a structural feature map.
[0250] Clean Physical Intensity Map Determination Module 24 is used to perform exposure compensation on the physical radiometric map based on the camera parameters when the visual data is acquired, and to suppress dark noise through an intensity collapse function to obtain a clean physical intensity map.
[0251] The RGB image determination module 25 is used to reconstruct the clean physical intensity map and the vector chromaticity map into HSV components and inversely transform them to obtain an RGB image.
[0252] Model training module 26 is used to construct multiple loss constraint functions based on the RGB image, the structural feature map, the physical prior structure map and the real image of the mine scene, and to use the Adam optimization algorithm to perform end-to-end training and update of the learnable parameters of the three-dimensional Gaussian ellipsoid until the loss converges.
[0253] The 3D reconstruction module 27 is used to output the 3D reconstruction result of the underground mine scene from the desired viewpoint through viewpoint adaptive rendering based on the parameters of the Gaussian body set after training convergence.
[0254] Example 3
[0255] In another aspect, the present invention also proposes an electronic device, please refer to [link to relevant documentation]. Figure 3 The diagram shows an electronic device according to the third embodiment of the present invention, including a memory 20, a processor 10, and a computer program 30 stored in the memory and executable on the processor. When the processor 10 executes the computer program 30, it implements the underground mine scene reconstruction method based on non-uniform illumination of three-dimensional Gaussian sputtering as described above.
[0256] In some embodiments, the processor 10 may be a central processing unit (CPU), controller, microcontroller, microprocessor or other data processing chip, used to run program code stored in memory 20 or process data, such as executing access restriction programs.
[0257] The memory 20 includes at least one type of readable storage medium, such as flash memory, hard disk, multimedia card, card-type memory (e.g., SD or DX memory), magnetic memory, magnetic disk, optical disk, etc. In some embodiments, the memory 20 can be an internal storage unit of an electronic device, such as the hard disk of the electronic device. In other embodiments, the memory 20 can also be an external storage device of the electronic device, such as a plug-in hard disk, smart media card (SMC), secure digital (SD) card, flash card, etc. Furthermore, the memory 20 can include both internal and external storage units of the electronic device. The memory 20 can be used not only to store application software and various types of data of the electronic device, but also to temporarily store data that has been output or will be output.
[0258] It should be pointed out that, Figure 3 The structure shown does not constitute a limitation on the electronic device. In other embodiments, the electronic device may include fewer or more components than shown, or combine certain components, or have different component arrangements.
[0259] This invention also proposes a computer-readable storage medium storing a computer program that, when executed by a processor, implements the above-described method for reconstructing underground mine scenes based on non-uniform illumination using three-dimensional Gaussian sputtering.
[0260] Those skilled in the art will understand that the logic and / or steps represented in the flowcharts or otherwise described herein, for example, can be considered as a ordered list of executable instructions for implementing logical functions, and can be embodied in any computer-readable medium for use by, or in conjunction with, an instruction execution system, apparatus, or device (such as a computer-based system, a processor-included system, or other system that can fetch and execute instructions from, an instruction execution system, apparatus, or device). For the purposes of this specification, "computer-readable medium" can mean any means that can contain, store, communicate, propagate, or transmit programs for use by, or in conjunction with, an instruction execution system, apparatus, or device.
[0261] More specific examples of computer-readable media (a non-exhaustive list) include: electrical connections (electronic devices) having one or more wires, portable computer disk drives (magnetic devices), random access memory (RAM), read-only memory (ROM), erasable and editable read-only memory (EPROM or flash memory), fiber optic devices, and portable optical disc read-only memory (CDROM). Furthermore, computer-readable media can even be paper or other suitable media on which the program can be printed, because the program can be obtained electronically, for example, by optically scanning the paper or other medium, followed by editing, interpreting, or otherwise processing as necessary, and then stored in computer memory.
[0262] It should be understood that various parts of the present invention can be implemented in hardware, software, firmware, or a combination thereof. In the above embodiments, multiple steps or methods can be implemented in software or firmware stored in memory and executed by a suitable instruction execution system. For example, if implemented in hardware, as in another embodiment, it can be implemented using any one or a combination of the following techniques known in the art: discrete logic circuits having logic gates for implementing logical functions on data signals, application-specific integrated circuits (ASICs) having suitable combinational logic gates, programmable gate arrays (PGAs), field-programmable gate arrays (FPGAs), etc.
[0263] In the description of this specification, references to terms such as "one embodiment," "some embodiments," "example," "specific example," or "some examples," etc., indicate that a specific feature, structure, material, or characteristic described in connection with that embodiment or example is included in at least one embodiment or example of the invention. In this specification, the illustrative expressions of the above terms do not necessarily refer to the same embodiment or example. Furthermore, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in one or more embodiments or examples.
[0264] The above embodiments merely illustrate several implementation methods of the present invention, and their descriptions are relatively specific and detailed, but they should not be construed as limiting the scope of the present invention. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of the present invention, and these all fall within the protection scope of the present invention. Therefore, the protection scope of this patent should be determined by the appended claims.
Claims
1. A method for reconstructing underground mine scenes based on non-uniform illumination using 3D Gaussian sputtering, characterized in that, The method includes: Visual data of an underground mine scene is acquired and processed to obtain multi-view two-dimensional image data. Sparse point cloud is generated based on the multi-view two-dimensional image data, and a three-dimensional Gaussian ellipsoid attribute set containing dynamic structural features is initialized based on the sparse point cloud to obtain an initial three-dimensional Gaussian volume set of the mine scene. Based on the multi-view two-dimensional image data, a physical prior structure diagram adapted to the mine environment is extracted as a supervision signal for the geometric growth of the Gaussian ellipsoid. The initial three-dimensional Gaussian volume set is decoupled and rendered to generate a vector chromaticity map, a physical radiation map, and a structural feature map; Based on the camera parameters when the visual data was acquired, the physical radiometric map was exposed and compensated, and dark noise was suppressed by the intensity collapse function to obtain a clean physical intensity map. The clean physical intensity map and the vector chromaticity map are reconstructed into HSV components, and an inverse transform is performed to obtain an RGB image. Based on the RGB image, the structural feature map, the physical prior structure map, and the real image of the mine scene, a multi-loss constraint function is constructed. The Adam optimization algorithm is used to perform end-to-end training and update of the learnable parameters of the three-dimensional Gaussian ellipsoid until the loss converges. Based on the parameters of the Gaussian body set after training convergence, the 3D reconstruction result of the underground mine scene under the desired viewpoint is output through viewpoint adaptive rendering. Based on the RGB image, the structural feature map, the physical prior structure map, and the real image of the mine scene, a multi-loss constraint function is constructed. The Adam optimization algorithm is used to perform end-to-end training and updating of the learnable parameters of the three-dimensional Gaussian ellipsoid until the loss converges. The steps include: Construct a total loss function consisting of photometric loss, structural consistency loss, and buoyancy suppression loss; For both large-scale tunnel scenarios and close-range equipment scenarios, corresponding optimization parameters and iteration stages are configured respectively; Configure the hyperparameters based on a combination of general and scenario-specific parameters, use the Adam optimization algorithm to update the learnable parameters of the Gaussian ellipsoid end-to-end, and perform Gaussian densification pruning and fine optimization in stages until the loss converges. The photometric loss is based on the difference between the RGB image and the real image of the mine scene, and adopts a weighted combination of multiple types of losses. The structural consistency loss is based on the difference constraint between the structural feature map and the physical prior structure map. The buoyancy suppression loss introduces a gradient scaling strategy.
2. The method for reconstructing an underground mine scene based on non-uniform illumination using three-dimensional Gaussian sputtering according to claim 1, characterized in that, The steps for initializing a set of three-dimensional Gaussian ellipsoidal attributes containing dynamic structural features based on the sparse point cloud to obtain an initial set of three-dimensional Gaussian volumes for the mine scene include: Visual data of an underground mine scene is acquired by a camera from a sparse perspective, and the visual data is processed to obtain multi-view two-dimensional image data. Sparse point clouds are then generated from the multi-view two-dimensional image data based on the structure-recovery-motion algorithm. Three-dimensional scene reconstruction is performed based on the sparse point cloud using three-dimensional Gaussian sputtering technology; In the initial stage of the reconstruction process, each point in the sparse point cloud is initialized with the geometric properties of a three-dimensional Gaussian ellipsoid, and the coverage of the Gaussian ellipsoid is expanded according to a preset ratio to adapt to the sparse point cloud scene. Based on the RGB information of underground mine images, the vector chromaticity attribute is initialized through inverse Gamma correction, and the physical emissivity attribute is initialized by combining the color features of sparse point clouds and the low light characteristics of the mine. The structural feature dimensions are determined based on the sparse point cloud density and scene type, and the initial values of the structural features are assigned.
3. The method for reconstructing underground mine scenes based on non-uniform illumination using three-dimensional Gaussian sputtering according to claim 1, characterized in that, The steps of extracting a physical prior structure diagram adapted to the mine environment from the multi-view two-dimensional image data, and using it as a supervision signal for Gaussian ellipsoid geometric growth, include: The physical prior theory based on photometric gradient is used to calculate the pixel brightness gradient of each pixel in the image, and the physical prior structure diagram is constructed by combining spatial differential operators. For the completely black pixel area in the mine, the pseudo-structure signal is filled based on the neighborhood gradient to avoid invalid areas in the physical prior structure map; For areas with strong reflection, a noise suppression factor is introduced to smooth the gradient, reducing the impact of reflection interference on the physical prior structure map.
4. The method for reconstructing underground mine scenes based on non-uniform illumination using three-dimensional Gaussian sputtering according to claim 1, characterized in that, The steps of decoupling and rendering the initial three-dimensional Gaussian volume set to generate a vector chromaticity map, a physical radiation map, and a structural feature map include: Perform α-blending rasterization on a 3D Gaussian volume to generate a vector chromaticity map, a physical radiance map, and a structural feature map, respectively. For high-noise regions in the vector chromaticity map whose physical radiance is below a preset threshold, their rendering weights are weighted and smoothed. For structural feature maps, if multi-dimensional structural features are used, their channel comprehensive value is taken as the supervision signal; if single-dimensional structural features are used, a single-channel map is directly output to balance the structural supervision effect and computational efficiency. Among them, the vector chromaticity map is used to characterize the intrinsic color of an object after the illumination is removed, the physical radiation map is used to characterize the true light field distribution including the influence of the light source, and the structural feature map is used to maintain geometric continuity in the dark region.
5. The method for reconstructing an underground mine scene based on non-uniform illumination using three-dimensional Gaussian sputtering according to claim 1, characterized in that, The steps of performing exposure compensation on the physical radiometric map based on the camera parameters used when acquiring the visual data, and suppressing dark noise using an intensity collapse function to obtain a clean physical intensity map include: Determine whether the camera parameters provide EXIF metadata; If so, the exposure gain is calculated directly based on the exposure-related parameters; If not, exposure parameters are calculated based on image statistical features, and outliers are removed using multi-view statistical information. For the physical radiometric map after exposure compensation, the adjusted intensity collapse function is used for denoising. The pixel value truncation range in the function is adapted to the strong light source area of the mine. At the same time, parameters are set to adapt to the noise suppression of dark areas, resulting in a clean physical intensity map.
6. The method for reconstructing an underground mine scene based on non-uniform illumination using three-dimensional Gaussian sputtering according to claim 1, characterized in that, The steps of reconstructing the clean physical intensity map and the vector chromaticity map into HSV components and then inversely transforming them to obtain an RGB image include: The hue and saturation components are recovered from the vector chromaticity diagram, and the saturation is enhanced by a preset amplitude to improve the recognizability of mine equipment markings; The clean physical intensity map is used as the luminance component and combined to obtain the HSV image; The HSV image is inversely transformed into an sRGB image using a differentiable color space conversion function, and the sRGB image is optimized using a lightweight tone mapping module to ensure that the image color and detail meet the observation requirements of the mine scene.
7. A system for reconstructing underground mine scenes based on non-uniform illumination using three-dimensional Gaussian sputtering, characterized in that, The system for implementing the underground mine scene reconstruction method based on non-uniform illumination using three-dimensional Gaussian sputtering as described in any one of claims 1 to 6 comprises: The data acquisition and processing module is used to acquire visual data of the underground mine scene, process the visual data to obtain multi-view two-dimensional image data, generate sparse point cloud based on the multi-view two-dimensional image data, and initialize a three-dimensional Gaussian ellipsoid attribute set containing dynamic structural features based on the sparse point cloud to obtain an initial three-dimensional Gaussian volume set of the mine scene. The physical prior structure diagram determination module is used to extract a physical prior structure diagram adapted to the mine environment based on the multi-view two-dimensional image data, as a supervision signal for Gaussian ellipsoid geometric growth; The decoupled rendering module is used to decouple and render the initial three-dimensional Gaussian volume set to generate a vector chromaticity map, a physical radiation map, and a structural feature map; The clean physical intensity map determination module is used to perform exposure compensation on the physical radiometric map based on the camera parameters when the visual data is acquired, and to suppress dark noise through an intensity collapse function to obtain a clean physical intensity map. An RGB image determination module is used to reconstruct the clean physical intensity map and the vector chromaticity map into HSV components and perform an inverse transformation to obtain an RGB image. The model training module is used to construct multiple loss constraint functions based on the RGB image, the structural feature map, the physical prior structure map and the real image of the mine scene, and to use the Adam optimization algorithm to perform end-to-end training and update of the learnable parameters of the three-dimensional Gaussian ellipsoid until the loss converges. The 3D reconstruction module is used to output the 3D reconstruction results of the underground mine scene from the desired viewpoint through viewpoint adaptive rendering based on the parameters of the Gaussian body set after training convergence.
8. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the program is executed by the processor, it implements the steps of the underground mine scene reconstruction method based on non-uniform illumination of three-dimensional Gaussian sputtering as described in any one of claims 1 to 6.
9. An electronic device, characterized in that, The system includes a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the program, it implements the method for reconstructing underground mine scenes based on non-uniform illumination using three-dimensional Gaussian sputtering as described in any one of claims 1-6.
Citation Information
Patent Citations
Historical block scene three-dimensional reconstruction method and system based on Gaussian sputtering
CN120318431A
Scene reconstruction method based on delayed rendering and three-dimensional Gaussian
CN120635288A