Low-light image unsupervised training method based on binocular camera
Through the unsupervised training method of the low-light image of binocular camera, the problem of poor image enhancement effect of monocular cameras under low-light conditions is solved, and high-quality image enhancement and reconstruction in low-illumination scenes are achieved, maintaining the accuracy of the three-dimensional geometric structure and the realism of the image.
Patent Information
- Application Number
- CN202510874483.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-27
- Publication Date
- 2025-07-25
- Estimated Expiration
- 2045-06-27
AI Technical Summary
In the prior art, monocular cameras are difficult to obtain high-quality paired data sets under low light conditions, resulting in poor image enhancement effect, and monocular images are difficult to accurately restore geometric structures in complex scenes, and image noise and texture loss are serious.
The low-light image unsupervised training method based on binocular camera is adopted, and image pairs are captured simultaneously through binocular cameras, and lightweight optimization networks and dual-branch networks are built. Combined with feature consistency constraint strategies, the total loss function is constructed, image decomposition and feature alignment are performed, image brightness and detail clarity are improved, while maintaining the three-dimensional geometric structure.
No need to rely on paired data sets, which significantly improves the adaptability and practicality of image enhancement, maintains the three-dimensional geometric consistency of the original scene, avoids structural distortion problems, and significantly improves the realism and visual quality of the image.
Smart Images

Figure CN120374429A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of computer vision and image processing, and particularly relates to an unsupervised training method for low-light images based on a binocular camera. Background Art
[0002] The low-light image enhancement technology has wide application value in scenarios such as video surveillance, autonomous driving, medical imaging, and robot perception. In recent years, significant progress has been made in image enhancement methods based on deep learning, and related methods are generally divided into two categories: supervised and unsupervised. Among them, the supervised method relies on a large number of paired low-light images and normal-light images as training data. Although it can achieve good results under ideal conditions, it is significantly difficult to obtain a real and rich paired data set in practical applications.
[0003] To reduce the dependence on labeled data, unsupervised image enhancement methods have gradually received attention. Among them, the Retinex theory is a classic image decomposition model, which believes that an image can be represented as the pixel-by-pixel product of illumination and reflectance. In the low-light image enhancement task, this theory enhances the illumination component to increase the image brightness while maintaining the reflectance component to preserve the image structure.
[0004] Although most current methods adopt a structure modeling based on the Retinex theory, they still generally rely on monocular images. Due to the limited information obtained by a monocular camera, it is difficult to accurately restore the geometric structure in complex scenes. Especially under low illumination conditions, problems such as image noise, contrast reduction, and texture loss will further reduce the accuracy of depth estimation and image restoration. Summary of the Invention
[0005] Aiming at the deficiencies of the above technologies, the purpose of the present invention is to provide an unsupervised training method for low-light images based on a binocular camera to solve the problems that it is difficult to obtain a high-quality paired data set and a monocular camera can only obtain image information from a single perspective. This method jointly optimizes the Retinex decomposition model and the feature consistency constraint strategy, while enhancing the image brightness and detail clarity, maintaining the original three-dimensional geometric structure information, and is applicable to image enhancement and reconstruction tasks under low illumination scenarios.
[0006] To achieve the above purpose, the technical solution adopted by the present invention is as follows: An unsupervised training method for low-light images based on a binocular camera, comprising the following steps: Step S1, synchronously capture a pair of low-light images of the real world through a binocular camera, remove the blurred images therein, and construct a data set; Step S2, build and use a lightweight optimization network to remove the sensor noise and compression artifacts in the dataset in Step S1, and generate an optimized image pair; Step S3, build a dual-branch network based on the retina theory. The left-branch network decomposes the left image of the optimized image pair to obtain a left reflection image and a left illumination image, and the right-branch network decomposes the right image of the optimized image pair to obtain a right reflection image and a right illumination image; Step S4, build a feature extraction network TB-FEN, extract the features of the left image of the optimized image pair and the left reflection image in Step S3, and the features of the right image of the optimized image pair and the right reflection image in Step S3, respectively, and perform feature consistency constraints; Step S5, construct a total loss function according to the feature consistency constraints in Steps S2, S3, and S4 to guide network training.
[0007] Furthermore, in Step S2: The built lightweight optimization network is specifically: The built lightweight optimization network consists of 4 convolutional activation layers and 1 convolutional layer with a kernel size of 3x3; each convolutional activation layer contains a convolutional layer with a kernel size of 3x3 and a ReLU activation function; the output channel number of the convolutional layer with a kernel size of 3x3 at the end is 3, which is used to generate the optimized image; The left image and the right image of the low-light image pair pass through the lightweight optimization network to obtain the optimized image, as shown in the expression; ; The projection consistency loss function is: ; where, I L is the left image of the low-light image pair, I R is the right image of the low-light image pair; i L is the left image of the optimized image pair, i R is the right image of the optimized image pair, is the optimization network, represents the square of the Euclidean norm, is the value of the projection consistency loss function.
[0008] Furthermore, in Step S3, build a dual-branch network based on the retina theory. The left-branch network decomposes the left image of the optimized image pair to obtain a left reflection image and a left illumination image, and the right-branch network decomposes the right image of the optimized image pair to obtain a right reflection image and a right illumination image; specifically: Step S31, the dual-branch network processes the left image of the optimized image pair and the right image of the optimized image pair respectively. The left branch and the right branch in the dual-branch network are symmetrically set and have the same network structure, which consists of a reflection decomposition network and an illumination decomposition network; Step S32: The reflection decomposition network extracts the reflection image, which consists of 5 convolutional activation layers. The first 4 convolutional activation layers each contain a convolutional layer with a kernel size of 3x3 and a ReLU activation function, and the last convolutional activation layer contains a convolutional layer with a kernel size of 3x3 and a Sigmoid activation function, with the number of output channels being 3; The illumination decomposition network extracts the illumination image, which consists of 5 convolutional activation layers. The first 4 convolutional activation layers each contain a convolutional layer with a kernel size of 3x3 and a ReLU activation function, and the last convolutional activation layer contains a convolutional layer with a kernel size of 3x3 and a Sigmoid activation function, with the number of output channels being 1; Step S33: The image decomposition process is expressed as: ; where are the left and right images of the optimized image pair, and respectively represent the left and right images of the reflection image and the left and right images of the illumination image, represents the reflection decomposition network, represents the illumination decomposition network; Step S34: Introduce an illumination smoothing loss function based on total variation regularization, and the expression is; ; where , represents the gradient of pixel (i,j) on the x-axis in the horizontal direction, , represents the gradient of pixel (i,j) on the y-axis in the vertical direction, represents summation, represents absolute value, represents the value of the illumination smoothing loss function; Step S35: Introduce a reflection reconstruction error loss, and the expression is; ; where represents the value of the reflection reconstruction error loss function, represents the square of the L2 norm, represents dot product.
[0009] Furthermore, in Step S4: Build a feature extraction network TB-FEN to extract the features of the left image of the optimized image pair and the left reflection image in Step S3, and the features of the right image of the optimized image pair and the right reflection image in Step S3, and perform feature consistency constraints; specifically: Step S41: The feature extraction network TB-FEN respectively extracts from the left and right images of the optimized image pair The left and right images of the reflected image Extract image semantic information at different levels from them and achieve feature alignment; Step S42, the feature extraction network TB-FEN processes the left and right images of the input optimized image pair The left and right images of the reflected image Use a weight encoder with parameter sharing for feature extraction to ensure the consistency of the extracted features; Step S43, the weight encoder with parameter sharing includes three levels of feature extraction modules, corresponding to the shallow feature extraction module, the middle feature extraction module, and the deep feature extraction module respectively. Each level of feature extraction module consists of a residual connection module; Inside the residual connection module, it consists of a convolutional layer with a convolution kernel size of 3x3, a batch normalization layer, and a ReLU activation function, and adds the input and output to form a residual connection. Each residual connection module uses the same network structure configuration; Step S44, shallow feature extraction module: The shallow feature consistency loss is defined as: ; Among them, represents the L1 norm, which is used to measure the absolute difference between features, represents the shallow feature extraction module, represents the shallow feature consistency loss; Middle feature extraction module, measuring the similarity between the optimized image and the reflected image in the local structure distribution: The middle feature consistency loss is calculated based on the structural similarity index and is defined as: ; Among them, represents the structural similarity index, represents the middle feature extraction module, represents the middle feature consistency loss; Deep feature extraction module: The loss function of the deep feature extraction module is based on the cosine similarity and is defined as: ; Among them, represents the deep feature extraction module, represents the deep feature consistency loss; Step S45, use a weighted total loss function to optimize the feature alignment effect at multiple levels. The total feature consistency loss function is defined as follows: ; Among them, are the weight hyperparameters of the shallow, middle, and deep loss terms, which are used to control the contribution ratio of features at different levels in the total loss, Represents the total feature consistency loss function.
[0010] Furthermore, in step S5, a total loss function is constructed to guide the network training; specifically: The total loss function is: 。
[0011] The present invention adopts another technical solution: a low-light image enhancement method based on a binocular camera, applying the unsupervised training method for low-light images based on a binocular camera, including the following steps: Step T1, synchronously capture the left image I of the low-light image pair through a binocular camera L and the right image I of the low-light image pair R ; Step T2, load the parameters of the lightweight optimization network and the dual-branch network trained through steps S2 and S3; Step T3, input the left image I of the image pair L and the right image I R into the lightweight optimization network to obtain the left optimized image i of the optimized image pair L and the right optimized image i of the optimized image pair R . The left branch network decomposes the left optimized image i of the optimized image pair through the reflection decomposition network and the illumination decomposition network L into the left reflection image R L and the left illumination image L L , The right branch network decomposes the right optimized image i of the optimized image pair through the reflection decomposition network and the illumination decomposition network R into the right reflection image R R and the right illumination image L R ; Step T4, enhance the left and right images L of the illumination image x and fuse the enhanced illumination image with the reflection image to obtain the image pair after illumination enhancement .
[0012] Furthermore, in step T4, enhance the left and right images L of the illumination image x and fuse the enhanced illumination image with the reflection image to obtain the image pair after illumination enhancement ; specifically: Step T41, for the left and right images L of the illumination image x , obtain the enhanced illumination image L Hx through gamma transformation, formula: ; where is the gamma transformation for the left and right images L of the illumination image x ; Step T42, combine the enhanced illumination image L Hx with the reflection image to obtain an image pair with enhanced illumination. The formula is: L Hx ; where H x is the image pair with enhanced illumination, represents the left and right images of the illumination image.
[0013] Compared with the prior art, the advantages of the present invention are as follows: (1) The image enhancement method proposed by the present invention does not depend on paired low-light and normal illumination image data during the training process, solves the problem that it is difficult to obtain high-quality paired data in supervised methods, and significantly improves the adaptability and practicability of the method; (2) The present invention uses a binocular camera as the input, and utilizes the structural redundant information between the left and right views for image enhancement and feature alignment, effectively maintaining the three-dimensional geometric consistency of the original scene and avoiding the common structure distortion problem in the monocular image enhancement process; (3) In the image decomposition stage, the illumination and reflection modeling based on the Retinex theory of the retina is introduced, and at the same time, the consistency constraint functions of shallow, middle, and deep features are designed to align the optimized image and the reflection image from the detail, structure to semantic levels, significantly improving the realism and visual quality of the enhanced image. Description of the Drawings
[0014] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following drawings are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.
[0015] Figure 1 is the flowchart of the unsupervised training method for low-light images based on a binocular camera according to the present invention; Figure 2 is the flowchart of the low-light image enhancement method based on a binocular camera according to the present invention. Detailed Embodiments
[0016] To make the objectives, technical solutions, and advantages of the embodiments of the present application clearer, the following will clearly and completely describe the technical solutions in the embodiments of the present application with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are part of the embodiments of the present application, rather than all of them. All other embodiments obtained by those of ordinary skill in the art without creative efforts based on the embodiments of the present application belong to the scope of protection of the present application.
[0017] The embodiments of the present invention will be further described in detail below in conjunction with the accompanying drawings.
[0018] As Figure 1 shown, a method for unsupervised training of low-light images based on a binocular camera includes the following steps: Step S1, synchronously capture a pair of low-light images of the real world through a binocular camera, remove the blurred images therein, and construct a data set; Step S2, build and use a lightweight optimization network to remove the sensor noise and compression artifacts of the data set in Step S1, and generate an optimized image pair; Step S3, build a two-branch network based on the retina theory. The left-branch network decomposes the left image of the optimized image pair to obtain a left reflection image and a left illumination image, and the right-branch network decomposes the right image of the optimized image pair to obtain a right reflection image and a right illumination image; Step S4, build a feature extraction network TB-FEN, extract the features of the left image of the optimized image pair and the left reflection image in Step S3, and the features of the right image of the optimized image pair and the right reflection image in Step S3, respectively, and perform feature consistency constraints; Step S5, construct a total loss function according to the feature consistency constraints in Steps S2, S3, and S4 to guide network training.
[0019] Further, in the above Step S2: The built lightweight optimization network is specifically: The built lightweight optimization network consists of 4 convolutional activation layers and 1 convolutional layer with a convolutional kernel size of 3x3; each convolutional activation layer contains a convolutional layer with a convolutional kernel size of 3x3 and a ReLU activation function; the output channel number of the convolutional layer with a convolutional kernel size of 3x3 at the end is 3, which is used to generate an optimized image; The left image and the right image of the low-light image pair are passed through the lightweight optimization network to obtain an optimized image, as shown in the expression; ; The projection consistency loss function is: ; where, I L is the left image of the low-light image pair, I R is the right image of the low-light image pair; i L is the left image of the optimized image pair, i R is the right image of the optimized image pair, is the optimization network, represents the square of the Euclidean norm, is the value of the projection consistency loss function.
[0020] Further, in step S3, a dual-branch network is constructed based on the retina theory. The left-branch network decomposes the left image of the optimized image pair into a left reflection image and a left illumination image, and the right-branch network decomposes the right image of the optimized image pair into a right reflection image and a right illumination image. Specifically: In step S31, the dual-branch network processes the left image and the right image of the optimized image pair respectively. The left branch and the right branch in the dual-branch network are symmetrically arranged and have the same network structure, which consists of a reflection decomposition network and an illumination decomposition network. In step S32, the reflection decomposition network extracts the reflection image, which consists of 5 convolutional activation layers. The first 4 convolutional activation layers contain a convolutional layer with a kernel size of 3x3 and a ReLU activation function, and the last convolutional activation layer contains a convolutional layer with a kernel size of 3x3 and a Sigmoid activation function, and the number of output channels is 3. The illumination decomposition network extracts the illumination image, which consists of 5 convolutional activation layers. The first 4 convolutional activation layers contain a convolutional layer with a kernel size of 3x3 and a ReLU activation function, and the last convolutional activation layer contains a convolutional layer with a kernel size of 3x3 and a Sigmoid activation function, and the number of output channels is 1. In step S33, the image decomposition process is expressed as: ; Among them, are the left and right images of the optimized image pair, and respectively represent the left and right images of the reflection image and the left and right images of the illumination image, represents the reflection decomposition network, represents the illumination decomposition network; In step S34, an illumination smoothing loss function based on total variation regularization is introduced, and the expression is; ; Among them, , represents the gradient of pixel (i,j) on the x-axis in the horizontal direction, , represents the gradient of pixel (i,j) on the y-axis in the vertical direction, represents summation, represents absolute value, represents the value of the illumination smoothing loss function; In step S35, a reflection reconstruction error loss is introduced, and the expression is; ; Among them represents the value of the reflection reconstruction error loss function, represents the square of the L2 norm. represents the dot product.
[0021] Furthermore, in step S4: Build the feature extraction network TB-FEN, extract the features of the left image of the optimized image pair and the left reflected image in step S3, and the features of the right image of the optimized image pair and the right reflected image in step S3 respectively, and perform feature consistency constraints; specifically: Step S41, the feature extraction network TB-FEN respectively extracts the image semantic information of different levels from the left and right images of the optimized image pair and the left and right images of the reflected image and realizes feature alignment; Step S42, the feature extraction network TB-FEN uses a weight encoder with shared parameters to extract features from the left and right images of the input optimized image pair and the left and right images of the reflected image to ensure the consistency of the extracted features; Step S43, the weight encoder with shared parameters includes three levels of feature extraction modules, corresponding to the shallow feature extraction module, the middle layer feature extraction module, and the deep layer feature extraction module respectively. Each level of feature extraction module is composed of a residual connection module; The internal of the residual connection module consists of a convolutional layer with a convolutional kernel size of 3x3, a batch normalization layer, and a ReLU activation function, and adds the input and output to form a residual connection. Each residual connection module uses the same network structure configuration; Step S44, shallow feature extraction module: The shallow feature consistency loss is defined as: ; where represents the L1 norm, which is used to measure the absolute difference between features, represents the shallow feature extraction module, represents the shallow feature consistency loss; Middle layer feature extraction module, measuring the similarity of the optimized image and the reflected image in the local structure distribution: The middle layer feature consistency loss is calculated based on the structural similarity index and is defined as: ; where represents the structural similarity index, represents the middle layer feature extraction module, represents the middle layer feature consistency loss; Deep layer feature extraction module: The loss function of the deep layer feature extraction module is based on the cosine similarity and is defined as: ; where Denote the deep feature extraction module, Denote the deep feature consistency loss; Step S45, adopt a weighted total loss function to optimize the feature alignment effects at multiple levels. The total feature consistency loss function is defined as follows: ; Wherein, Are the weight hyperparameters of the shallow, middle, and deep loss terms, used to control the contribution ratio of features at different levels in the total loss, Denote the total feature consistency loss function.
[0022] Furthermore, in step S5, construct the total loss function to guide network training; specifically: The total loss function is: .
[0023] As Figure 2 shown, the present invention adopts another technical solution: a low-light image enhancement method based on a binocular camera, applying the unsupervised training method for low-light images based on a binocular camera, including the following steps: Step T1, synchronously capture the left image I L of the low-light image pair and the right image I R of the low-light image pair through a binocular camera; Step T2, load the parameters of the lightweight optimization network and the dual-branch network trained through step S2 and step S3; Step T3, input the left image I L and the right image I R of the image pair into the lightweight optimization network to obtain the left optimized image i L of the optimized image pair and the right optimized image i R of the optimized image pair. The left branch network decomposes the left optimized image i L of the optimized image pair into a left reflection image R L and a left illumination image L L , The right branch network decomposes the right optimized image i R of the optimized image pair into a right reflection image R R and a right illumination image L R ; Step T4, enhance the left and right images L x of the illumination image, and fuse the enhanced illumination image with the reflection image to obtain an image pair with enhanced illumination .
[0024] Furthermore, in step T4, the left and right images Lx Enhance it, and fuse the enhanced illumination image with the reflection image to obtain an image pair with enhanced illumination Specifically: Step T41: For the left and right images L of the illumination image x , obtain the enhanced illumination image L through gamma transformation Hx , formula: ; Wherein, is to perform gamma transformation on the left and right images L of the illumination image x ; Step T42: Combine the enhanced illumination image L Hx with the reflection image to obtain an image pair with enhanced illumination, formula: L Hx ; Wherein, H x is the image pair with enhanced illumination, represents the left and right images of the illumination image
[0025] The above are only specific embodiments of the present application, enabling those skilled in the art to understand or implement the present application. Various modifications to these embodiments will be obvious to those skilled in the art. The general principles defined herein can be implemented in other embodiments without departing from the spirit or scope of the present application. Therefore, the present application will not be limited to these embodiments shown herein, but rather to the broadest scope consistent with the principles and novel features claimed herein.
Claims
1. An unsupervised training method for low-light images based on a binocular camera, characterized in that It includes the following steps: Step S1: Synchronously capture low-light image pairs of the real world through a binocular camera, remove the blurred images among them, and construct a dataset; Step S2: Build and use a lightweight optimization network to remove sensor noise and compression artifacts in the dataset in Step S1, and generate optimized image pairs; Step S3: Build a two-branch network based on the retina theory. The left-branch network decomposes the left image of the optimized image pair to obtain a left reflection image and a left illumination image, and the right-branch network decomposes the right image of the optimized image pair to obtain a right reflection image and a right illumination image; Step S4: Build a feature extraction network TB-FEN, extract the features of the left image of the optimized image pair and the left reflection image in Step S3, and the features of the right image of the optimized image pair and the right reflection image in Step S3, and perform feature consistency constraints; Step S5: According to the feature consistency constraints in Steps S2, S3, and S4, construct a total loss function to guide network training.
2. The unsupervised training method for low-light images based on a binocular camera according to claim 1, wherein In Step S2: The built lightweight optimization network is specifically: The built lightweight optimization network consists of 4 convolutional activation layers and 1 convolutional layer with a kernel size of 3x3; each convolutional activation layer contains a convolutional layer with a kernel size of 3x3 and a ReLU activation function; the output channel number of the convolutional layer with a kernel size of 3x3 at the end is 3, which is used to generate optimized images; The left image and the right image of the low-light image pair pass through the lightweight optimization network to obtain optimized images, as shown in the expression; ; The projection consistency loss function is as follows: ; Among them, I L is the left image of the low-light image pair, and I R is the right image of the low-light image pair; i L is the left image of the optimized image pair, and i R is the right image of the optimized image pair, is the optimization network, represents the square of the Euclidean norm, is the value of the projection consistency loss function.
3. A method for unsupervised training of low-light images based on a binocular camera according to claim 2, characterized in that, Step S3: Build a two-branch network based on the retina theory. The left-branch network decomposes the left image of the optimized image pair to obtain a left reflection image and a left illumination image, and the right-branch network decomposes the right image of the optimized image pair to obtain a right reflection image and a right illumination image; specifically: Step S31: The two-branch network processes the left image of the optimized image pair and the right image of the optimized image pair respectively. The left branch and the right branch in the two-branch network are symmetrically set and have the same network structure, which consists of a reflection decomposition network and an illumination decomposition network; Step S32: The reflection decomposition network extracts the reflection image, which consists of 5 convolutional activation layers. The first 4 convolutional activation layers contain a convolutional layer with a kernel size of 3x3 and a ReLU activation function, and the last convolutional activation layer contains a convolutional layer with a kernel size of 3x3 and a Sigmoid activation function, and the output channel number is 3; The illumination decomposition network extracts the illumination image, which consists of 5 convolutional activation layers. The first 4 convolutional activation layers contain a convolutional layer with a kernel size of 3x3 and a ReLU activation function, and the last convolutional activation layer contains a convolutional layer with a kernel size of 3x3 and a Sigmoid activation function, and the output channel number is 1; Step S33: The image decomposition process is expressed as: ; Among them, are the left and right images of the optimized image pair, and respectively represent the left and right images of the reflected image and the left and right images of the illumination image, represents the reflection decomposition network, represents the illumination decomposition network; Step S34: Introduce an illumination smoothing loss function based on total variation regularization, and the expression is; ; Among them, , represents the gradient of pixel (i,j) on the x-axis in the horizontal direction, , represents the gradient of pixel (i,j) on the y-axis in the vertical direction, represents summation, represents absolute value, represents the value of the illumination smoothing loss function; Step S35: Introduce a reflection reconstruction error loss, and the expression is; ; where represents the value of the reflection reconstruction error loss function, represents the square of the L2 norm, represents the dot product.
4. A method for unsupervised training of low-light images based on a binocular camera according to claim 3, characterized in that, In step S4: Build the feature extraction network TB-FEN, extract the features of the left image of the optimized image pair and the left reflected image in step S3, and the features of the right image of the optimized image pair and the right reflected image in step S3 respectively, and perform feature consistency constraints; Specifically: Step S41, the feature extraction network TB-FEN extracts the image semantic information of different levels from the left and right images of the optimized image pair and the left and right images of the reflection image respectively, and realizes feature alignment; Step S42, the left and right images of the input optimized image pair and the left and right images of the reflected image are subjected to feature extraction by the feature extraction network TB-FEN using a weight encoder with parameter sharing to ensure the consistency of the extracted features; and the left and right images of the reflected image are subjected to feature extraction by a weight encoder with parameter sharing to ensure the consistency of the extracted features; In step S43, the weight encoder with parameter sharing contains three levels of feature extraction modules, corresponding to the shallow feature extraction module, the middle layer feature extraction module, and the deep layer feature extraction module respectively. Each level of feature extraction module is composed of a residual connection module; Inside the residual connection module, it consists of a convolutional layer with a convolutional kernel size of 3x3, a batch normalization layer, and a ReLU activation function, and the input and output are added to form a residual connection. Each residual connection module uses the same network structure configuration; In step S44, the shallow feature extraction module: The shallow feature consistency loss is defined as: ; Among them, represents the L1 norm, which is used to measure the absolute difference between features, represents the shallow feature extraction module, represents the shallow feature consistency loss; The middle layer feature extraction module measures the similarity between the optimized image and the reflected image in the local structure distribution: The middle layer feature consistency loss is calculated based on the structural similarity index and is defined as: ; Among them, represents the structural similarity index, represents the middle-level feature extraction module, represents the middle-level feature consistency loss; The deep layer feature extraction module: The loss function of the deep layer feature extraction module is based on the cosine similarity and is defined as: ; Among them, represents the deep feature extraction module, represents the deep feature consistency loss; In step S45, use a weighted total loss function to optimize the feature alignment effects of multiple levels. The total feature consistency loss function is defined as follows: ; Among them, is the weight hyperparameter of the shallow, middle, and deep loss terms, which is used to control the contribution ratio of features at different levels to the total loss. represents the total feature consistency loss function.
5. The unsupervised training method for low-light images based on a binocular camera according to claim 4, wherein, In step S5, construct a total loss function to guide network training; specifically: The total loss function is: .
6. A low-light image enhancement method based on a binocular camera, applying the unsupervised training method for low-light images based on a binocular camera described in claim 5, characterized in that, It includes the following steps: Step T1, synchronously capture the left image I of the low-light image pair and the right image I of the low-light image pair through a binocular camera L using a binocular camera R ; In step T2, load the parameters of the lightweight optimization network and the double-branch network trained through step S2 and step S3; Step T3, input the left image I of the image pair L and the right image I R into the lightweight optimization network to obtain the left optimized image i of the optimized image pair L and the right optimized image i of the optimized image pair R . The left branch network decomposes the left optimized image i of the optimized image pair L into the left reflection image R L and the left illumination image L L ; The right-branch network decomposes the right-optimized image i of the optimized image pair into a right-reflection image R R and a right-illumination image L R through a reflection decomposition network and an illumination decomposition network R ; Step T4, enhance the left and right images L of the illumination image x and fuse the enhanced illumination image with the reflection image to obtain an image pair with enhanced illumination .
7. A low-light image enhancement method based on a binocular camera according to claim 6, characterized in that: Step T4, enhance the left and right images L of the illumination image x and fuse the enhanced illumination image with the reflection image to obtain an image pair with enhanced illumination ; specifically: Step T41, for the left and right images L of the illumination image x , obtain the enhanced illumination image L through gamma transformation Hx , formula: ; Among them, perform gamma transformation on the left and right images L of the illumination image x ; Step T42, combine the enhanced illumination image L Hx with the reflection image to obtain an image pair with enhanced illumination. The formula is: L Hx ; Among them, H x is the image pair after enhanced illumination, representing the left and right images of the illumination image.
Citation Information
Patent Citations
Binocular image super-resolution method and device based on graph neural network
CN114170078A
Visual SLAM (Simultaneous Localization and Mapping) method and system for image enhancement under low-light condition
CN116894791A
Multi-scale low-illumination binocular stereo image enhancement method integrated with low-frequency information
CN118096561A
Two-way Transform image super-resolution method and system
CN118918004A
Low-light image enhancement method based on zero-reference Retinex decomposition network
CN119168895A