Efficient multi-view three-dimensional reconstruction method based on low-illumination image enhancement

By using the image enhancement method of encoder-decoder structure and coarse to fine MVS reconstruction model in low light environments, the problem of low three-dimensional reconstruction quality under low light is solved, and high-quality image enhancement and three-dimensional reconstruction effects are achieved.

CN120219614APending Publication Date: 2025-06-27GUANGDONG UNIV OF TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510233627.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-28
Publication Date
2025-06-27

AI Technical Summary

Technical Problem

In low-light environments, traditional multi-view three-dimensional reconstruction methods cannot provide sufficient image quality and depth information, resulting in a lack of detail and accuracy in reconstruction results.

Method used

Using a low-light image enhancement method based on the encoder-decoder structure, high-quality enhanced images are generated by processing the original RAW image, and combined with coarse to fine MVS reconstruction model, extract depth features, and perform depth map generation and three-dimensional reconstruction.

Benefits of technology

It significantly improves the quality of low-light images, restores more image details and retains critical depth information, improves the accuracy and effectiveness of 3D reconstruction, and reduces the dependence of computing resources.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120219614A_ABST
    Figure CN120219614A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of image processing, in particular to an efficient multi-view three-dimensional reconstruction method based on low-illumination image enhancement, which comprises the following steps of: 1, acquiring an original RAW image as network input data, and inputting the original RAW image into a low-illumination image enhancement module for processing to obtain an enhanced image; 2, inputting the enhanced image into a low-illumination image enhancement network and a coarse-to-fine MVS reconstruction model which are trained in advance in an MVS reconstruction module, and extracting depth features; according to the method, features are extracted from a multi-view RAW image, image enhancement is carried out, details and quality of a low-illumination image are remarkably improved, a depth map is refined step by step through multi-stage depth estimation, different resolutions and depth plane assumptions are used in each stage, the calculation complexity is reduced in combination with a prediction result of the previous stage, and the prediction efficiency is improved. The depth precision in the low-illumination image is improved, the calculation cost is effectively reduced through the multi-stage optimization process, and meanwhile the reconstruction quality is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of image processing, and particularly to an efficient multi-view three-dimensional reconstruction method based on low-light image enhancement. Background Art

[0002] In the fields of computer vision and three-dimensional reconstruction, the multi-view stereo (MVS) technology is to recover the three-dimensional structure of a scene from images taken from multiple angles. However, in low-light environments, traditional MVS methods often fail to provide sufficient image quality and depth information, resulting in reconstructed results lacking details and accuracy. In low-light conditions, the brightness and contrast of images are low, often leading to increased image noise, loss of details, and signal degradation and quantization errors are prone to occur in the non-RAW data of image sensors, which makes it particularly difficult to perform MVS reconstruction based on these images.

[0003] In the prior art, many low-light image enhancement methods focus on improving the visual effect of a single image, but in multi-view stereo reconstruction, these methods often fail to effectively combine multi-view information, resulting in the reconstruction quality and accuracy not meeting the requirements for generating high-quality three-dimensional models. In addition, traditional MVS reconstruction methods mainly rely on textures and feature points in images, and in low light, there are fewer textures and feature points in images, further exacerbating the challenges of reconstruction.

[0004] Therefore, an efficient multi-view three-dimensional reconstruction method based on low-light image enhancement is proposed to solve the above-mentioned problems. Summary of the Invention

[0005] Technical Problems to be Solved

[0006] In view of the above-mentioned drawbacks of the prior art, the present invention provides an efficient multi-view three-dimensional reconstruction method based on low-light image enhancement, which can effectively solve the problem in the prior art that multi-view stereo three-dimensional reconstruction cannot be effectively performed in low-light environments and high-quality three-dimensional models cannot be generated.

[0007] Technical Solutions

[0008] To achieve the above object, the present invention is realized through the following technical solutions:

[0009] The present invention provides an efficient multi-view three-dimensional reconstruction method based on low-light image enhancement, including the following steps:

[0010] Step 1: Collect the original RAW images as network input data and input them into the low-light image enhancement module for processing to obtain enhanced images;

[0011] Step 2: Input the enhanced image into the pre-trained low-light image enhancement network and the coarse-to-fine MVS reconstruction model in the MVS reconstruction module, extract the depth features, and return the depth features to the low-light image enhancement network and the coarse-to-fine MVS reconstruction model to obtain the depth map;

[0012] Step 3: Perform back-projection operation on the depth map through the camera parameters, map the pixel points in the two-dimensional image to the three-dimensional space to form a point cloud model, and complete the three-dimensional reconstruction.

[0013] Furthermore, the method for collecting the original RAW image in Step 1 includes:

[0014] Collect 8-bit RAW data in Bayer format, separate and package it into four-channel image data, and at the same time reduce the resolution of the four-channel image data to 50% of the original image, and adjust the four-channel image data through the brightness amplification factor.

[0015] Furthermore, the method for processing the network input data in Step 1 includes:

[0016] Extract the feature information from the original RAW image through the encoder, and input the feature information into the decoder to restore the enhanced image.

[0017] Furthermore, the construction method of the low-light image enhancement network and the coarse-to-fine MVS reconstruction model in Step 2 includes:

[0018] Define the encoder-decoder structure as the basic structure of the low-light image enhancement network and the coarse-to-fine MVS reconstruction model, and define that the low-light image enhancement module and the MVS reconstruction module share the same encoder structure;

[0019] Input the enhanced image I i into the encoder to obtain the depth feature map F i , where i represents the number of the viewing angle;

[0020] Define a depth hypothesis d, and construct a cost volume C(x, y, d) in the reference viewing angle based on the depth range;

[0021] Input the depth feature map F i and the cost volume C(x, y, d) into the decoder to obtain the depth map D k (x, y), where (x, y) represents the pixel position and k represents the number of the stage K.

[0022] Furthermore, the training method of the low-light image enhancement network and the coarse-to-fine MVS reconstruction model includes:

[0023] Collect the original RAW image data of several low-light perspectives, and use the low-light image enhancement module to process the original RAW image data to obtain enhanced images of several perspectives as the training data for the low-light image enhancement network and the coarse-to-fine MVS reconstruction model;

[0024] Initialize the network parameters of the low-light image enhancement network and the coarse-to-fine MVS reconstruction model, including: depth hypothesis d, cost volume C(x, y, d), and stage K;

[0025] Define the loss function L of the low-light image enhancement network and the coarse-to-fine MVS reconstruction model:

[0026]

[0027] where d(p) is the true depth of pixel p; is the final depth estimate; is the set of valid pixels; λ k is the loss weight for stage K; ξ = 1.2;

[0028] Input the enhanced image into the encoder structure, and after passing through multiple convolutional, pooling operations, feature warping, and cascading cost volume operations in sequence, obtain the depth map of each perspective, calculate the corresponding loss function value, calculate the error gradients of the low-light image enhancement network and the coarse-to-fine MVS reconstruction model in reverse, and backpropagate the error gradients along the cascading cost volume; According to the error gradients, use the optimization algorithm to update the network parameters to reduce the value of the loss function; Repeat the training of the enhanced images of several perspectives until the low-light image enhancement network and the coarse-to-fine MVS reconstruction model converge or reach the pre-set number of iterations in advance; That is, complete the training of the low-light image enhancement network and the coarse-to-fine MVS reconstruction model.

[0029] Furthermore, the decoder includes a feature processing module and a depth estimation sub-module;

[0030] The feature processing module maps the depth feature map F i to the coordinate system of the reference perspective based on the feature warping technique, and the mapping method is:

[0031] Generate multiple depth planes using the depth hypothesis d, and warp the depth feature map F i according to the depth hypothesis d to form a feature volume;

[0032] Generate discrete depth plane hypotheses based on the plane sweep stereo algorithm. Under the depth plane hypothesis, the feature mapping between perspectives is achieved by calculating the homography matrix H i (d) of the reference perspective, and the calculation formula is: where K i 、R i 、ti and K r 、R r 、t r are the internal and external parameters of the cameras for the i-th perspective and the reference perspective respectively; n r is the normal vector of the reference plane.

[0033] Furthermore, the depth estimation sub-module gradually optimizes the cost volume C(x, y, d) at several resolutions based on the cascaded cost volume. The size of the cost volume C(x, y, d) is F×D×W×H; where F is the number of depth feature maps F i , D is the number of depth plane hypotheses, and W and H are the spatial dimensions of the depth feature map F i ;

[0034] At each stage, according to the prediction of the previous stage, the depth range of the hypothesis is gradually narrowed and the resolution is increased;

[0035] The number of depth plane hypotheses at stage K is given by , where G k and V k are the depth range and the depth interval respectively; the resolution at each stage is gradually increased to and where W is the width of the image and H is the height of the image;

[0036] Then, a 3D convolutional neural network CNN is used to perform regularization and regression operations on the cost volume C(x, y, d).

[0037] Advantageous Effects

[0038] The technical solution provided by the present invention has the following advantageous effects compared with the prior art:

[0039] By combining the encoder-decoder structure with RAW image processing technology, the present invention significantly improves the quality of low-light images, and solves the problems of detail loss and noise amplification faced by traditional image enhancement methods in low-light environments. By enhancing low-light images, more image details can be restored and key depth information can be retained, thereby effectively improving the visibility and realism of images. This enhancement method can not only make low-light images clearer and more colorful, but also ensure the accuracy of depth information during the enhancement process, providing accurate input for subsequent 3D reconstruction. Compared with traditional low-light image processing methods, the present invention can effectively reduce image noise while ensuring image details, improving image quality. This innovative enhancement method lays a more solid foundation for multi-view stereo reconstruction in low-light environments, significantly improving the accuracy and effect of subsequent 3D reconstruction;

[0040] In the process of multi-view stereo reconstruction, the present invention adopts the combination of a shared encoder and a coarse-to-fine decoder structure, which optimizes the calculation process and reduces the dependence on computing resources. The design of the shared encoder highly integrates the feature extraction process of multi-view input images, significantly reducing redundant calculations and improving computing efficiency. The coarse-to-fine decoder structure adopts a step-by-step optimization method. First, a rough depth estimation is performed, and then the depth map is gradually refined at a higher resolution, thereby improving the accuracy of depth estimation and effectively reducing memory consumption. In this process, the step-by-step refined stagewise decoding can optimize the depth information of the image at different resolutions at multiple levels, making the finally obtained 3D model more accurate and delicate. In addition, the coarse-to-fine design can also accelerate the convergence speed of depth estimation, avoid the problem of memory overload in traditional reconstruction methods, and reduce the demand for GPU resources. Therefore, the present invention can maintain a high computing efficiency when processing large-scale data, and improve the depth estimation accuracy and 3D reconstruction quality;

[0041] To address the special challenges of 3D reconstruction in low-light environments, the present invention constructs a dedicated low-light dataset, including multi-view captured low-light and normal-light images, accompanying 3D models and depth maps, and provides high-quality reference data for network training through fine annotation. These low-light datasets not only contain typical low-light scenarios, such as night-time shooting, indoor low-light environments, and high-dynamic range (HDR) scenes, but also cover various noise and uneven illumination conditions, capable of simulating various complex low-light environments in the real world. By using these datasets, the present invention can perform high-quality 3D reconstruction under low-light conditions, avoiding the impact of low-light environments on traditional multi-view reconstruction methods, and further enhancing the practicality and accuracy of low-light 3D reconstruction. In addition, the datasets of the present invention can also provide valuable experimental materials for other research in the fields of low-light image enhancement and 3D reconstruction, promoting the development of low-light computer vision technology and providing more reliable technical support for related industries;

[0042] The technical solution of the present invention has strong robustness and can cope with various challenges in low-light environments. Whether under extreme lighting conditions or in the presence of a large amount of noise and uneven lighting, the present invention can maintain high performance. The low-light image enhancement module effectively removes the noise brought by the low-light environment through deep learning methods, restores details, and improves the image quality; while the multi-view stereo reconstruction module can stably perform depth estimation in low-light environments and generate high-quality 3D models. This robustness makes the present invention outstanding in practical applications and is applicable to various actual scenarios, such as security monitoring, autonomous driving, virtual reality, architectural modeling, archaeological excavation and other fields. Especially in nighttime or insufficient light environments, it can complete accurate 3D reconstruction tasks. In addition, the robustness of the present invention also ensures the wide adaptability of the system, can process a variety of different types of low-light scenarios, and has strong versatility and reliability.

[0043] Through innovative low-light image enhancement methods, efficient multi-view stereo reconstruction designs, specially constructed low-light datasets, and strong robustness, the present invention significantly improves the 3D reconstruction effect under low-light conditions. These advantages enable the present invention to have a significant technological breakthrough in the field of 3D reconstruction under low-light environments and can provide reliable technical solutions for image enhancement and 3D modeling applications in various low-light environments. Brief Description of the Drawings

[0044] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.

[0045] Figure 1 It is a schematic flow chart of the efficient multi-view 3D reconstruction method in the embodiment of the present invention. Detailed Embodiments

[0046] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the drawings in the embodiments of the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of the present invention.

[0047] The following further describes the present invention with reference to the embodiments.

[0048] Embodiment:

[0049] See the appendix Figure 1 In this case, an efficient multi-view 3D reconstruction method based on low-light image enhancement is proposed, including the following steps:

[0050] Step 1: Collect the original RAW images from multiple perspectives as the network input data, process the network input data, and input the processed network input data into the low-light image enhancement module for processing to obtain enhanced images;

[0051] Step 2: Input the enhanced images into the pre-constructed low-light image enhancement network and the coarse-to-fine MVS reconstruction model in the MVS reconstruction module, extract the depth features of the enhanced images, and return the depth features to the low-light image enhancement network and the coarse-to-fine MVS reconstruction model to obtain depth maps;

[0052] Step 3: Perform back-projection operations on the depth maps through camera parameters, map the pixel points in the two-dimensional images to the three-dimensional space, form a high-precision point cloud model, and complete the 3D reconstruction.

[0053] The present invention constructs a deep learning network model based on the Encoder-Decoder structure for enhancing low-light images. This structure is widely used in image processing tasks due to its excellent feature extraction ability, simple network design, and efficient computational performance.

[0054] Specifically, in this case, the original RAW images use 8-bit RAW data as the network input data. Compared with non-RAW image data, 8-bit RAW image data can retain more original detail information, especially in low-light conditions, which helps to reduce the impact of noise and signal loss; it overcomes the problem that non-RAW image data generated by traditional image sensors often has signal damage, resulting in a large number of pixel values being quantized to zero, thus significantly affecting the effect of the reconstruction process; while 8-bit RAW image data can effectively reduce the impact of such problems, provide a richer data basis for subsequent processing, improve the brightness, contrast, and detail quality of low-light images through the low-light image enhancement module, and provide better input for the MVS reconstruction module.

[0055] The present invention uses the original RAW images from multiple perspectives as the network input data. Considering the need for multi-view reconstruction, the original RAW images from multiple perspectives can provide more image information, effectively enhance the quality of low-light images, and thus improve the overall effect of multi-view 3D reconstruction.

[0056] The ways to process the network input data include:

[0057] Using Bayer image separation and four-channel packing technology, the Bayer format (BGGR) of the original RAW image is separated and packed into four-channel image data. This helps to process the image data more efficiently and reduces the image resolution to 50% of the original image to reduce image noise and improve processing speed, creating favorable conditions for subsequent calculations and feature extraction;

[0058] To ensure the brightness consistency of the enhanced image, all four-channel image data are adjusted through a brightness amplification factor; this step can effectively improve the overall dark condition of the low-light image, make the brightness of the enhanced image more suitable for subsequent processing and analysis, and enhance the visibility and information richness of the image; the value of the brightness amplification factor is set to 250, and experiments have verified that the brightness amplification factor performs well in various low-light scenes.

[0059] More specifically, the processed network input data is processed in the following ways:

[0060] The encoder extracts features from the processed original RAW image to obtain feature information. With its powerful feature extraction capability, the encoder can mine deep feature information from the original RAW image, which is crucial for subsequent processing.

[0061] The extracted feature information is input into the decoder, and the decoder restores the enhanced image based on the feature information, that is, the enhanced image; in order to improve the computational efficiency, the grayscale image is used as input in the decoding stage to maintain the original resolution of the image;

[0062] After the above operations, the output enhanced image is significantly improved in terms of brightness, contrast and detail quality, providing better quality input data for the subsequent MVS reconstruction module.

[0063] Furthermore, in this case, the low-light image enhancement module can restore the original RAW image in low light to an effect close to that in normal light. However, since there is still a large difference between the original RAW image in low light and the image in normal light, the feature information extracted when directly used for MVS reconstruction is not comprehensive enough. To solve this problem, the present invention adopts a coarse-to-fine MVS reconstruction method based on enhanced image features. Compared with the traditional method, the MVS reconstruction method of the present invention has made effective improvements in image feature extraction and depth estimation.

[0064] The low-light image enhancement module shares the same encoder structure with the MVS reconstruction module. By sharing the encoder structure between the low-light image enhancement module and the MVS reconstruction module, efficient multi-view 3D reconstruction can be achieved. The design of sharing the encoder structure highly integrates the feature extraction process of multi-view input images. When processing multi-view images, only one feature extraction operation is required, significantly reducing redundant calculations and thus improving the computational efficiency. At the same time, this design makes the feature connection between image enhancement and MVS reconstruction closer. The depth features extracted from the enhanced image through the shared encoder structure can better serve subsequent depth estimation, thereby improving the effect of low-light images in multi-view stereo reconstruction.

[0065] The ways to extract the depth features of the enhanced image include:

[0066] Define the encoder-decoder structure as the basic structure of the low-light image enhancement network and the coarse-to-fine MVS reconstruction model.

[0067] Define the input of the encoder: the enhanced images I of multiple views i , where i represents the view number; the encoder structure includes multiple layers of convolution and pooling operations to gradually extract multi-resolution feature representations, obtaining the output of the encoder: the depth features of the enhanced image, that is, the depth feature map F of each view i ;

[0068] Define the input of the decoder: the depth feature map F of each view i and the cost volume C(x, y, d) in the reference view, where d is the depth hypothesis and (x, y) is the pixel position;

[0069] The decoder includes a multi-resolution feature processing module and a specific depth estimation sub-module. The feature processing module is used to process the depth feature map F extracted by the encoder i , and the depth estimation sub-module is used to generate multi-stage depth estimation results, that is, the output of the decoder: the multi-stage refined depth map D k (x, y), where k represents the coarse-to-fine stage number; using depth features for depth estimation greatly improves the effect of low-light images in multi-view stereo reconstruction.

[0070] This design of sharing the encoder structure not only improves the computational efficiency but also makes the features between image enhancement and MVS reconstruction more closely combined.

[0071] Furthermore, in this case, the feature processing module processes the depth feature map F based on the feature warping technique i ; feature warping is a differentiable operation used to warp features from different views (i.e., the depth feature map F i)The coordinate system mapped to the reference view, and the mapping method is as follows:

[0072] Define a depth hypothesis d, generate multiple depth planes using the depth hypothesis d, and warp the depth feature map F of each view according to the depth hypothesis d i to form a feature volume; generate discrete depth plane hypotheses based on the plane sweep stereo algorithm; under each depth plane hypothesis, calculate the homography matrix H i (d) between each view and the reference view to achieve feature mapping between views, and the calculation formula is: In the formula, K i , R i , t i and K r , R r , t r are the internal and external camera parameters of the i-th view and the reference view respectively; n r is the normal vector of the reference plane; through the homography matrix H i (d), the depth feature maps F of other views i can be warped to the reference view to form a feature volume for each depth plane hypothesis, laying the foundation for subsequent depth estimation;

[0073] The depth estimation sub-module gradually optimizes the cost volume C(x, y, d) in the reference view at multiple resolutions (from coarse to fine). The cost volume C(x, y, d) is a four-dimensional array, which is constructed based on a certain depth range hypothesis at a coarser resolution, and its size is F×D×W×H; where F is the number of depth feature maps F i , D is the number of depth plane hypotheses, and W and H are the spatial dimensions of the depth feature map F i ; then as the processing stage progresses, gradually narrow the depth range and increase the resolution; for example, the cascade stage gradually refines the depth estimation result through stage K = 3:

[0074] Stage 1: Wide range of depth hypothesis d, low resolution;

[0075] Stage 2: Narrower depth range, higher resolution;

[0076] Stage 3: Further refine the range of depth hypothesis d and higher resolution;

[0077] In each stage, gradually narrow the assumed depth range according to the prediction of the previous stage; the number of depth plane hypotheses in stage K is given by , in the formula, G k and V k are the depth range and depth interval respectively; the resolution in each stage is gradually increased to and to improve the GPU memory usage efficiency, where W is the width of the image and H is the height of the image; in this way, the depth estimation result is continuously refined, and finally a high-resolution depth map is generated to more accurately reflect the three-dimensional structure of the scene.

[0078] Use a 3D convolutional neural network CNN to perform regularization and regression operations on the cost volume C(x, y, d). The 3D convolutional neural network CNN can effectively learn and process the features in the cost volume C(x, y, d), remove the influence of uncertain factors brought by noise, occlusion, weak texture and other regions, so as to obtain a more accurate depth estimation result. After this process, a high-quality depth map is finally obtained, providing key data support for 3D reconstruction;

[0079] The training methods of the low-light image enhancement network and the coarse-to-fine MVS reconstruction model include:

[0080] Collect the original RAW image data of low light from multiple viewpoints, and use the low-light image enhancement module to process the original RAW image data to obtain enhanced images from multiple viewpoints as the training data of the low-light image enhancement network and the coarse-to-fine MVS reconstruction model; initialize the network parameters of the low-light image enhancement network and the coarse-to-fine MVS reconstruction model, including: depth hypothesis d, cost volume C(x, y, d) and stage K; define the loss function L of the low-light image enhancement network and the coarse-to-fine MVS reconstruction model:

[0081]

[0082] where d(p) is the true depth of pixel p; is the final depth estimation; is the set of valid pixels; λ k is the loss weight of stage K; ξ = 1.2;

[0083] Input the enhanced image into the encoder structure. After passing through multiple layers of convolution, pooling operations, feature warping and cascading cost volume operations in sequence, obtain the depth map of each viewpoint, calculate the corresponding loss function value, calculate the error gradient of the low-light image enhancement network and the coarse-to-fine MVS reconstruction model in reverse, and backpropagate the error gradient along the cascading cost volume; according to the error gradient, use the optimization algorithm to update the network parameters to reduce the value of the loss function; repeat the training of the enhanced images from multiple viewpoints until the low-light image enhancement network and the coarse-to-fine MVS reconstruction model converge or reach the pre-set number of iterations in advance; that is, complete the training of the low-light image enhancement network and the coarse-to-fine MVS reconstruction model.

[0084] It should be noted that, in order to ensure the efficient training of the low-light image enhancement network and the coarse-to-fine MVS reconstruction model, the present invention has carried out refined data preparation, including multiple steps, to ensure the diversity, accuracy, and effectiveness of the training data; specifically including:

[0085] Viewpoint group selection: In order to make full use of multi-view image information, the present invention first selects viewpoint groups based on the camera parameters obtained from the Structure from Motion (SfM) algorithm. Each group of viewpoints contains multiple different shooting angles to ensure the lighting consistency and geometric alignment between multi-view images. In each group of viewpoints, we select a reference viewpoint and form the viewpoint group by calculating the similarity with other viewpoints and using the geometric relationship and overlapping area between viewpoints. This strategy ensures that each viewpoint group contains representative and high-quality images, suitable for subsequent training tasks;

[0086] Iterative Closest Point (ICP) algorithm alignment: In order to obtain accurate depth maps for comparison with the results after low-light image enhancement, we use the ICP algorithm to precisely align the 3D model. The ICP algorithm finds the optimal rigid transformation parameters (including rotation, translation, and scale transformation) by minimizing the Euclidean distance between two point sets. This step ensures that the generated depth maps can be aligned with the actual scene captured by multi-view, providing high-precision reference data for training and evaluation.

[0087] Dataset generation and model training: The dataset of the present invention includes multi-view low-light images, normal-light images, as well as corresponding 3D models and depth maps; these data are used to train and validate the low-light image enhancement network and the coarse-to-fine MVS reconstruction model. The dataset covers images under different shooting conditions, including both low-light and normal-light images, ensuring the richness and diversity of the training data; in addition, the 3D models and depth maps in the dataset are processed by ICP alignment to ensure that the real-world data used in the training process can be precisely compared with the enhanced images, improving the training effect and generalization ability of the model.

[0088] The above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements for some of the technical features; and these modifications or replacements will not cause the essence of the corresponding technical solutions to deviate from the protection scope of the technical solutions of the various embodiments of the present invention.

Claims

1. An efficient multi-view 3D reconstruction method based on low-light image enhancement, characterized in that: The following steps are involved: Step 1: Collect the original RAW image as the network input data and input it into the low-light image enhancement module for processing to obtain an enhanced image; Step 2: Input the enhanced image into the pre-trained low-light image enhancement network and coarse-to-fine MVS reconstruction model in the MVS reconstruction module to extract the depth features, and return the depth features to the low-light image enhancement network and coarse-to-fine MVS reconstruction model to obtain the depth map; Step 3: Back-project the depth map through the camera parameters, map the pixels in the two-dimensional image into three-dimensional space, form a point cloud model, and complete three-dimensional reconstruction.

2. The efficient multi-view 3D reconstruction method based on low-light image enhancement according to claim 1, characterized in that: The method of acquiring the original RAW image in step 1 includes: The 8-bit RAW data in Bayer format is collected, separated and packaged into four-channel image data. At the same time, the resolution of the four-channel image data is reduced to 50% of the original image, and the four-channel image data is adjusted by the brightness magnification factor.

3. The efficient multi-view 3D reconstruction method based on low-light image enhancement according to claim 1, characterized in that: The method of processing the network input data in step 1 includes: The encoder extracts feature information from the original RAW image, inputs the feature information into the decoder, and restores the enhanced image.

4. The efficient multi-view 3D reconstruction method based on low-light image enhancement according to claim 1, characterized in that: The construction method of the low-light image enhancement network and the coarse-to-fine MVS reconstruction model in step 2 includes: Define the encoder-decoder structure as the basic structure of the low-light image enhancement network and the coarse-to-fine MVS reconstruction model, and define the low-light image enhancement module and the MVS reconstruction module to share the same encoder structure; The enhanced image I i Input into the encoder to get the deep feature map F i , where i represents the number of the viewing angle; Define a depth hypothesis d and construct a cost volume C(x, y, d) under the reference perspective based on the depth range; The deep feature map F i And the cost body C(x, y, d) is input into the decoder to obtain the depth map D k (x, y), where (x, y) represents the pixel position and k represents the number of stage K.

5. The efficient multi-view 3D reconstruction method based on low-light image enhancement according to claim 4, characterized in that: The training method of the low-light image enhancement network and the coarse-to-fine MVS reconstruction model includes: Collect the original RAW image data of low light from several viewing angles, process the original RAW image data using the low light image enhancement module, and obtain enhanced images of several viewing angles as training data for the low light image enhancement network and the coarse-to-fine MVS reconstruction model; Initialize the network parameters of the low-light image enhancement network and the coarse-to-fine MVS reconstruction model, including: depth hypothesis d, cost volume C(x, y, d), and stage K; Define the loss function L of the low-light image enhancement network and the coarse-to-fine MVS reconstruction model: Where d(p) is the true depth of pixel p; is the final depth estimate; is the effective pixel set; k is the loss weight of stage K; ξ=1.2; The enhanced image is input into the encoder structure, and after multiple layers of convolution, pooling operation, feature warping and cascade cost body operation, the depth map of each perspective is obtained, and the corresponding loss function value is calculated. The error gradient of the low-light image enhancement network and the coarse-to-fine MVS reconstruction model is reversely calculated, and the error gradient is back-propagated along the cascade cost body; according to the error gradient, the network parameters are updated using the optimization algorithm to reduce the value of the loss function; the enhanced images of several perspectives are repeatedly trained until the low-light image enhancement network and the coarse-to-fine MVS reconstruction model converge or reach the preset number of iterations; that is, the training of the low-light image enhancement network and the coarse-to-fine MVS reconstruction model is completed.

6. The efficient multi-view 3D reconstruction method based on low-light image enhancement according to claim 5, characterized in that: The decoder includes a feature processing module and a depth estimation submodule; The feature processing module transforms the deep feature map F i Mapped to the coordinate system of the reference perspective, the mapping method is: Use the depth hypothesis d to generate multiple depth planes, and according to the depth hypothesis d, the depth feature map F i Distort to form a characteristic volume; Based on the plane scanning stereo algorithm, a discrete depth plane hypothesis is generated. Under the depth plane hypothesis, the homography matrix H of the reference view is calculated. i (d) to realize the feature mapping between perspectives. The calculation formula is: In the formula, K i , R i ,t i and K r , R r ,t r are the camera internal and external parameters of the i-th viewing angle and the reference viewing angle respectively; n r is the normal vector of the reference plane.

7. The efficient multi-view 3D reconstruction method based on low-light image enhancement according to claim 6, characterized in that: The depth estimation submodule gradually optimizes the cost volume C(x, y, d) at several resolutions based on the cascaded cost volume, and the size of the cost volume C(x, y, d) is F×D×W×H; where F is the depth feature map F i , D is the number of depth plane hypotheses, W and H are the depth feature maps F i The size of the space; Each stage progressively narrows the assumed depth range and improves the resolution based on the predictions of the previous stage; The number of depth plane hypotheses at stage K is given by Given, where G k and V k are depth range and depth interval respectively; the resolution of each stage is gradually improved to and Where W is the width of the image and H is the height of the image; The 3D convolutional neural network CNN is then used to regularize and regress the cost volume C(x, y, d).