A waveband compensating imaging method for laser glare suppression
By using a three-channel photoelectric imaging system with temporal/spatial spectral separation and an image restoration algorithm based on generative adversarial networks, the problem of image information loss caused by laser glare was solved, and high-quality imaging under laser interference was achieved.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- NAT UNIV OF DEFENSE TECH
- Filing Date
- 2024-07-25
- Publication Date
- 2026-05-29
AI Technical Summary
When existing optoelectronic imaging systems are subjected to laser glare interference, the image sensor is prone to saturation areas, leading to information loss and affecting image quality and system performance.
A three-channel photoelectric imaging system with temporal/spatial spectral separation and an image restoration algorithm based on generative adversarial networks (GANs) are employed to recover lost image information through joint hardware and software optimization. The system uses filters with interconnected but non-overlapping transmission bands and combines a Swin-Unet network structure and a discriminant network to achieve image restoration and fusion.
In the event of laser glare, it can effectively recover lost image information, improve the performance of the imaging system in complex optoelectronic environments, and ensure imaging quality.
Smart Images

Figure CN118741328B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of band compensation technology, and in particular to a band compensation imaging method for suppressing laser glare. Background Technology
[0002] A photoelectric imaging system mainly consists of an optical system, an image sensor, and an image processing system. The image sensor is typically placed near the focal plane of the optical system and perpendicular to its optical axis. During operation, focusing is used to image the target scene onto the detector plane to obtain a clear image output. When this type of photoelectric imaging system is irradiated by low-energy interfering lasers, it can cause a large saturation area in the image sensor, leading to the loss of image information, i.e., laser glare interference. Therefore, while pursuing excellent imaging capabilities, photoelectric imaging systems also need to enhance their laser glare suppression capabilities. To ensure the normal operation of the photoelectric imaging system, reasonable and feasible technical means need to be adopted to suppress laser glare interference, reduce information loss, and thus improve the performance of photoelectric imaging equipment in complex photoelectric environments.
[0003] Depending on the methods employed, glare suppression techniques in photoelectric imaging systems can be divided into two categories: one involves processing at the image end, and the other involves processing through front-end optical system design combined with back-end image processing. Image processing-based laser glare suppression typically requires that the interfering light spot does not completely obscure the target information. First, the area and location of the interfering light spot are determined, and then techniques such as digital corona discharge are used to recover the image within the diffused light spot. Laser glare suppression using optical system design has no requirements regarding the size and location of the interfering light spot, making it more applicable. However, it requires changes to the imaging optical path, increasing system complexity. Its core idea is to utilize the difference in coherence between the imaging light field and the laser light field. By frequency modulation of the incident light field, spatiotemporally separated spectral band images are obtained. Correlation fusion of these spatiotemporally separated spectral band image data is then used to achieve laser glare interference suppression and image information recovery, thereby solving the problem of recovering visible light target image information from laser interference. Summary of the Invention
[0004] To address the aforementioned technical problems, this invention proposes a band-compensation imaging method for laser glare suppression. This invention focuses on a visible light color imaging system, conducting optical design and algorithm research. A three-channel photoelectric imaging system with time-spectral separation is designed, and an image restoration algorithm based on generative adversarial networks is developed to recover the lost image information caused by glare, thus achieving glare suppression. The feasibility of the proposed band-compensation technique for suppressing laser glare is analyzed, and a three-channel photoelectric imaging experimental system with time-spectral separation is built to verify the system's color imaging capabilities. Furthermore, training and test set images are acquired using the imaging system to test the performance of the compensation processing algorithm, thereby experimentally verifying the glare suppression capability of the band-compensation technique.
[0005] This invention discloses a band-compensatory imaging method for laser glare suppression. The method utilizes a three-channel photoelectric imaging system with temporal / spatial spectral separation and an image restoration algorithm. Through joint hardware and software optimization, it recovers the lost image information caused by glare, thereby achieving laser glare suppression; wherein:
[0006] For the aforementioned three-channel photoelectric imaging system with time / space spectral separation, the three-color RGB filter or four-color CMYK filter used in the color camera with overlapping transmission bands is replaced with a filter whose transmission bands are connected but do not overlap, thereby limiting the glare effect of interfering laser on the color imaging system to one band.
[0007] The image restoration algorithm uses the color statistical attributes of the reference image of the target scene or area scene before glare interference, combined with the image information of the un-glare-interference band, to predict the image information of the glare-interference band through machine learning, and then restores the color image under the condition of no glare interference through image fusion.
[0008] When there is no interfering laser, the three-channel photoelectric imaging system transmits the received monochrome image to the image processing system for de-mosaicing, white balance and color correction to obtain a color image that meets the clarity requirements and outputs it for display.
[0009] In the presence of interfering lasers, one channel of the three-channel photoelectric imaging system is affected by laser glare interference, causing a saturation region in the image sensor and resulting in the loss of image information. The lost image information caused by the glare interference is recovered by using image information from the unaffected band and image recovery algorithms, thereby suppressing laser glare interference.
[0010] In a preferred embodiment, the time / space spectral separation three-channel photoelectric imaging system in the method employs the time / space spectral separation principle; wherein:
[0011] Spatial spectral separation refers to the structure that uses optical field modulation elements to spatially separate the mixed optical field of interfering laser and signal light and image it onto multiple image sensors. The optical field modulation of spatial spectral separation is achieved by using prisms, dichroic mirrors and filters to simultaneously acquire spectral images of all channels and perform real-time image processing and glare suppression.
[0012] Temporal spectral separation refers to the structure that uses optical field manipulation elements to temporally separate the mixed optical field of interfering laser and signal light and image it onto an image sensor. The optical field manipulation of temporal spectral separation is achieved by using a filter wheel mounting bracket, a sliding filter mounting bracket, and a filter.
[0013] In a preferred embodiment, the three-channel photoelectric imaging system in the method includes: a collimating and expanding telescope C, dichroic mirrors DBS1 and DBS2, bandpass filter groups BP1, BP2, and BP3, imaging lenses L1, L2, and L3, and image sensors S1, S2, and S3; wherein the bandpass bands of the bandpass filter groups BP1, BP2, and BP3 are selected such that the corresponding imaging spectral ranges of the image sensors S1, S2, and S3 are independent and do not overlap; wherein:
[0014] Bandpass filter groups BP1, BP2 and BP3 are respectively placed in a mounting slot of filter wheel FW. By rotating the filters, the first channel Ch.1, the second channel Ch.2 and the third channel Ch.3 are imaged step by step. The spectral images of the three channels are obtained by time-spectral separation.
[0015] The center wavelengths of the bandpass filter groups BP1, BP2 and BP3 are 450±5nm, 550±5nm and 650±5nm, respectively, the full width at half maximum (FWHM) of the spectrum is 50±5nm, and the stopband optical density OD≥4.
[0016] In a preferred embodiment, in the method:
[0017] When there is a single wavelength interfering laser in the external environment, the interfering laser will cause glare and interference to one channel of the image sensor, while the remaining two channels of the image sensor will work normally.
[0018] When there is dual-wavelength interfering laser in the external environment, the interfering laser glare interferes with the image sensors of two channels, while the image sensor of the other channel works normally.
[0019] When there are three or more interfering lasers in the external environment, no practical interfering laser source is found, and the image sensor of at least one channel is working normally.
[0020] In a preferred embodiment, the image restoration algorithm employs a generative adversarial network (GAN), comprising a generator network and a discriminator network. The generator network uses a Swin-Unet network structure, which is an image segmentation network that integrates the SwinTransformer model and the classic U-Net architecture. The basic module structure of the discriminator network adopts the discriminator structure of a super-resolution GAN, replacing the output layer with a convolutional layer. Wherein:
[0021] The Swin-Unet network employs an encoder structure that initially processes the input image, converting it into a sequential embedding representation. Specifically, the input image is first segmented into several non-overlapping 4×4 pixel blocks. This segmentation process transforms the feature dimension of each block into 4×4×3, and a linear embedding layer maps the feature dimension to the model-defined dimension C. The transformed patch markers then undergo feature extraction through several Swin Transformer blocks and a patch merging layer. The patch merging layer performs downsampling and increases the feature dimension, generating hierarchical feature representations, integrating local features, and reducing sequence length. The Swin Transformer layer then extracts feature information.
[0022] For the encoder, the C-dimensional input image is fed into two consecutive Swing Transformer blocks for feature representation learning, during which the feature dimension and resolution remain unchanged; at the same time, the patch merging layer reduces the number of labels by 2x downsampling and expands the feature dimension to twice the original size. This operation is performed three times inside the encoder.
[0023] For the patch merging layer, the input block is divided into 4 sub-blocks, which are then merged through the patching process, downsampling the feature resolution to half of the original. At the same time, since the merging operation increases the feature dimension by four times, a linear layer is applied to the merged features to readjust the feature dimension back to twice the original size.
[0024] For the decoder, it is built on the basis of the Swing Transformer block. Corresponding to the encoder, the decoder uses a patch extension layer to perform feature map upsampling. The patch extension layer restores the higher resolution feature map by rearranging adjacent feature map blocks, achieving 2x resolution upsampling. At the same time, the number of feature channels is halved to maintain feature dimension consistency.
[0025] In a preferred embodiment, the method includes a discriminant network used to determine whether each image is original data or a synthesized image generated by a generator network. The structure of the discriminant network includes:
[0026] The input layer is used to receive image input of size (C, H, W), where C represents the number of channels of the image, and H and W represent the height and width of the image, respectively.
[0027] Several basic convolutional modules are provided, each of which includes a convolutional layer, a batch normalization layer, and a LeakyReLU activation function. In the convolutional layer, the number of input channels is converted into the specified number of output channels. The convolutional layer uses a 3×3 kernel with a stride of 1 and edge padding of 1.
[0028] Batch standardization layer;
[0029] LeakyReLU activation function;
[0030] Several convolutional layers with the same number of channels are used, with a 3×3 kernel, a stride of 2, and edge padding of 1. Each convolutional layer is followed by a batch normalization layer and a LeakyReLU activation function.
[0031] The stacked convolutional blocks are configured as follows:
[0032] The first convolutional block does not contain a batch normalization layer, has C input channels (C=3 for RGB images), and 64 output channels.
[0033] The second convolutional block contains two basic convolutional modules with 64 input channels and 128 output channels.
[0034] The third convolutional block contains two basic convolutional modules with 128 input channels and 256 output channels.
[0035] The fourth convolutional block contains two basic convolutional modules with 256 input channels and 512 output channels.
[0036] Finally, the output layer is a convolutional layer with 512 input channels and 1 output channel, using a 3×3 convolutional kernel with a stride of 1 and padding of 1. The output size generated by this layer is (1, patch_h, patch_w), where patch_h = H / 2^4 and patch_w = W / 2^4.
[0037] In a preferred embodiment, the method uses a generative adversarial network to recover the input image, with the loss function being:
[0038] L=αL gen-MSE +L VGG / i,j +L msssim
[0039] Where α = 0.01, L gen-MSE L VGG / i,j and L msssim These are the loss functions for generative networks, VGG-based perceptual networks, and MSSSIM, respectively.
[0040] The loss function for the generator network is:
[0041]
[0042] Where G is the generator network, D is the discriminator network, MSE is the mean squared error, n is the output size of the discriminator, and x is the input image;
[0043] The discriminant network loss function is:
[0044]
[0045] Where y is the reference image;
[0046] The perceptual loss function based on VGG is:
[0047]
[0048] Among them, W i,j H i,j The value θ represents the feature map size, φ() represents the features extracted from the image by the VGG network, and the subscript θ represents the feature map size. G This represents the generated network weight parameters;
[0049] The MSSSIM loss function is:
[0050]
[0051] Where M represents different scale levels, μ p and μ g σ represents the average value of the predicted image and the reference image, respectively. p and σ g σ represents the standard deviation of the predicted image and the reference image, respectively. pg β represents the covariance between the predicted image and the reference image. m and γ m To indicate the relative importance between the two terms, constant terms c1 and c2 are introduced as smoothing terms.
[0052] In summary, this invention designs a three-channel imaging system based on the concept of band compensation, employing spatial and temporal spectral separation for the verification of the technology's principle. An image restoration algorithm based on generative adversarial networks is developed, trained using a public dataset, and tested to evaluate the average peak signal-to-noise ratio (PSNR) and multi-scale structural similarity (MSSSIM) of the restored images when a single channel is interfered with by laser glare, thus verifying the information recovery capability of the band compensation technology. A three-channel imaging system based on temporal spectral separation is constructed, and preliminary color image acquisition experiments are conducted to test the glare performance of the imaging system under different laser energies. Furthermore, the imaging system is used to acquire training and test set images to test the performance of the image restoration algorithm. The results show that the average PSNR and MSSSIM of the restored images reach 29.358 dB and 0.958, respectively, experimentally verifying the glare suppression capability of the band compensation technology.
[0053] This invention can ensure imaging quality while having laser glare suppression function, suppress the interference of laser glare on the imaging system, and improve the working performance of the imaging system in complex optoelectronic environments. Attached Figure Description
[0054] To more clearly illustrate the specific embodiments of the present invention or the technical solutions in the prior art, the drawings used in the description of the specific embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of the present invention. For those skilled in the art, other drawings can be obtained from these drawings without creative effort.
[0055] Figure 1 This is a schematic diagram of the band compensation technology of the present invention;
[0056] Figure 2 This is a schematic diagram of the spatial spectral separation three-channel photoelectric imaging system of the present invention;
[0057] Figure 3 This is a schematic diagram of the time-spectral separation three-channel photoelectric imaging system of the present invention;
[0058] Figure 4 This is a spectral transmittance curve of the time-spectral separation three-channel photoelectric imaging system of the present invention;
[0059] Figure 5 This is a diagram of the generative network structure of the image restoration algorithm of this invention;
[0060] Figure 6 This is a diagram of the discriminant network structure of the image restoration algorithm of this invention;
[0061] Figure 7 This is a loss function curve of the image restoration algorithm of the present invention;
[0062] Figure 8 This is a graph showing the peak signal-to-noise ratio (PSNR) and multi-scale structural similarity (MSSSIM) curves of the image restoration algorithm of this invention.
[0063] Figure 9 This is a simulation result diagram of the image restoration algorithm of the present invention;
[0064] Figure 10 This is a diagram showing the de-mosaicing effect of the time-spectral separation three-channel photoelectric imaging system of the present invention;
[0065] Figure 11 This is a diagram showing the white balance adjustment effect of the time-spectral separation three-channel photoelectric imaging system of the present invention;
[0066] Figure 12 The diagram shows the color correction processing effect of the time-spectral separation three-channel photoelectric imaging system of the present invention, a 24-color test chart (a), and experimental processing results (b).
[0067] Figure 13 This is a diagram illustrating the dazzling effect of the time-spectral separation three-channel photoelectric imaging system of the present invention;
[0068] Figure 14 It is a dazzling effect image from a color camera;
[0069] Figure 15 These are typical experimental images acquired by the time-spectral separation three-channel photoelectric imaging system of the present invention;
[0070] Figure 16 This is a visual representation of the experimental results of the image restoration algorithm of this invention. Detailed Implementation
[0071] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0072] First Embodiment
[0073] This embodiment provides an imaging technology that suppresses laser glare while maintaining image quality. The solution is: a photoelectric imaging system with laser glare suppression function. This system is a band compensation imaging system, which includes: a three-channel photoelectric imaging system with temporal / spatial spectral separation and an image restoration algorithm. The three-channel photoelectric imaging system with temporal / spatial spectral separation replaces the three-color (RGB) or four-color (CMYK) filters used by the color camera, which have overlapping transmission bands, with filters whose transmission bands are interconnected but do not overlap, so that the glare effect of interfering lasers on the color imaging system is limited to one band. The image restoration algorithm, based on the color statistical attributes of the reference image of the target or area scene before glare interference, combined with the image information of the bands not affected by glare, infers the image information of the band affected by glare through machine learning, and restores a color image that is similar to that under the condition of no interference through image fusion.
[0074] When there is no interfering laser, the three-channel photoelectric imaging system sends the received monochrome image to the image processing system for de-mosaicing, white balance, and color correction to obtain a clear color image, which is then output and displayed. When there is interfering laser, one channel of the three-channel photoelectric imaging system will be affected by laser glare, causing saturation areas in the image sensor and resulting in the loss of image information. The lost image information caused by glare is recovered by using information that is not affected by glare and image recovery algorithms, thereby suppressing the effect of glare interference and achieving the purpose of laser glare suppression.
[0075] Furthermore, the aforementioned three-channel photoelectric imaging system with temporal / spatial spectral separation employs the principle of temporal / spatial spectral separation. Spatial spectral separation refers to the structure that uses optical field manipulation elements to spatially separate the mixed optical field of interfering laser and signal light and image it onto multiple image sensors. Optical field manipulation for spatial spectral separation can be achieved using elements such as prisms, dichroic mirrors, and filters. Temporal spectral separation refers to the structure that uses optical field manipulation elements to temporally separate the mixed optical field of interfering laser and signal light and image it onto a single image sensor. Optical field manipulation for temporal spectral separation can be achieved using elements such as filter wheel mounts, sliding filter mounts, and filters. The advantages of the spatial spectral separation optical system are that it can simultaneously acquire spectral images of all channels, perform real-time image processing, and has real-time glare suppression capabilities. Its disadvantage is its complex system structure. The advantages of the temporal spectral separation optical system are its simple system structure. Its disadvantage is that it requires acquiring spectral images of each channel individually and then performing image processing to remove glare interference, resulting in poor real-time performance.
[0076] Furthermore, the image restoration algorithm is implemented using a generative adversarial network, which consists of two parts: a generator network and a discriminator network. They compete with each other during training, hence the term "adversarial".
[0077] Furthermore, the generative network adopts the Swin-Unet network structure, which is an image segmentation network that integrates the Swin Transformer model and the classic U-Net architecture.
[0078] Furthermore, the basic module structure of the discriminant network is based on the discriminant structure design of super-resolution generative adversarial networks, but the output layer is replaced with a convolutional layer.
[0079] Furthermore, the loss function is defined as:
[0080] L=αL gen-MSE +L VGG / i,j =L msssim
[0081] Where α = 0.01, L gen-MSE L VGG / i,j and L msssim These are the loss functions for generative networks, VGG-based perceptual networks, and MSSSIM, respectively.
[0082] Compared with the prior art, the significant advantage of this invention is that it recovers the lost image information caused by glare through joint optimization of hardware and software, thereby achieving the function of laser glare suppression.
[0083] Second Embodiment
[0084] Figure 1 A schematic diagram of the band compensation technology of the present invention is shown. For example... Figure 1 As shown, the band compensation technology with laser glare suppression function of the present invention is an imaging technology that combines hardware and software optimization, including: a three-channel photoelectric imaging system with temporal / spatial spectral separation and an image restoration algorithm.
[0085] A band compensation technique is used to modulate the optical field frequency, thereby suppressing laser glare in the imaging system. A three-channel photoelectric imaging system with spatial spectral separation is designed, and the schematic diagram of the system is shown below. Figure 2 As shown.
[0086] Figure 2 In the diagram, C represents a collimating and expanding telescope, DBS1 and DBS2 are dichroic mirrors, BP1, BP2, and BP3 are bandpass filter groups, L1, L2, and L3 are imaging lenses, and S1, S2, and S3 are identical image sensors. The selection of the bandpass bands for the bandpass filter groups BP1, BP2, and BP3 in the optical system must ensure that the corresponding imaging spectral ranges of the image sensors S1, S2, and S3 are independent and do not overlap.
[0087] Design a time-spectral separation three-channel photoelectric imaging system. The schematic diagram of the system is shown below. Figure 3 As shown in (a), FW is the filter wheel, and bandpass filter groups BP1, BP2, and BP3 are placed in a mounting slot of the wheel. By rotating the filters, images of Ch.1, Ch.2, and Ch.3 can be formed step by step. Three-channel spectral band images are acquired through time-spectral separation. In the figure, L1, L2, and L3 are imaging lenses, and S1, S2, and S3 are identical image sensors. The experimental setup is as follows... Figure 3 As shown in (b), the bandpass filter group is located in the mounting slot of the rotary wheel. The rotary wheel (FWM6C, LBTEK) is connected to the camera (CS165CU / M, Thorlabs) via a 30mm coaxial system. The imaging lens (MVL35M23, Navitar) is fixed to the camera via a C-interface.
[0088] The design schemes of the bandpass filter groups BP1, BP2 and BP3 in the system are shown in Table 1. The center wavelengths of the bandpass filter groups BP1, BP2 and BP3 are 450±5nm, 550±5nm and 650±5nm, respectively, the full width at half maximum (FWHM) of the spectrum is 50±5nm, and the stopband optical density OD≥4.
[0089] Table 1: Filter parameters for time-spectral separation imaging systems
[0090] Component model center wavelength Spectral half width at half maximum (FWHM) bandstop optics MBF10-450-50 450±5 50±5 OD≥4 MBF10-550-50 550±5 50±5 OD≥4 MBF10-650-50 650±5 50±5 OD≥4
[0091] The visible light spectral transmittance curves of bandpass filter groups BP1, BP2, and BP3 are shown below. Figure 4As shown, the transmission bands of the filters are independent and do not overlap. This design ensures that the optical system can achieve normal target detection and identification while effectively suppressing laser glare interference.
[0092] Typical application scenarios are as follows: When there is a single-wavelength interfering laser in the external environment, the laser can only dazzle and interfere with the image sensor of one channel, while the image sensor of the remaining two channels can work normally; when there is a dual-wavelength interfering laser in the external environment, the interfering laser can dazzle the image sensor of two channels, while the image sensor of the other channel can still work normally; when there are three or more interfering lasers in the external environment, literature research shows that single-frequency blue laser technology or supercontinuum laser technology with spectral coverage below 500nm, high intensity, and high stability is not mature enough, and no practical interfering laser source has been found. Therefore, at least one of the channels can still work normally.
[0093] When capturing images using a three-channel photoelectric imaging experimental system with time-spectral separation, the image acquired in the channel with high-transmittance interference laser wavelengths exhibits local saturation, ultimately resulting in a single-channel damaged RGB image and causing loss of image information. Therefore, a compensatory processing algorithm based on generative adversarial networks (GANs) is proposed for image restoration of single-channel damaged images. The generative network adopts the Swin-Unet network structure, while the basic module structure of the discriminant network is borrowed from the discriminant network of super-resolution GANs. Simultaneously, convolutional layers are used instead of convolutional layers for the output layer, transforming the discriminant network output into a matrix corresponding to different image patches, which helps the discriminant network capture more local details in the image.
[0094] The generative network employs the Swin-Unet network architecture, which is an image segmentation network that integrates the Swin Transformer model and the classic U-Net architecture. The Swin Transformer is a deep neural network based on a self-attention mechanism. This model effectively addresses the long-range dependency problem in visual tasks by introducing a hierarchical Transformer structure. U-Net, on the other hand, is an encoder-decoder network specifically designed for image segmentation, characterized by the use of skip connections to preserve spatial information during the downsampling process. The Swin-Unet network architecture combines the efficient feature extraction capabilities of the Swin Transformer with the spatial information preservation advantages of U-Net.
[0095] Figure 5As shown, the Swin-Unet network employs an innovative encoder structure that initially processes the input image, transforming it into a sequential embedded representation. Specifically, the input image is first segmented into several non-overlapping 4×4 pixel blocks. This segmentation process reduces the feature dimension of each block to 4×4×3 = 48. Then, a linear embedding layer is applied to map the feature dimensions to the model-defined dimension C. The transformed patch tokens undergo feature extraction through several Swin Transformer blocks and a patch merging layer. The patch merging layer performs downsampling and increases the feature dimension, resulting in a hierarchical feature representation. The main function of the patch merging layer is to integrate local features and reduce sequence length, while the task of the Swin Transformer layers is to extract rich feature information.
[0096] In the encoder section, a C-dimensional input image with a resolution of is fed into two consecutive Swing Transformer blocks for feature representation learning. During the learning process, the feature dimension and resolution remain unchanged. Simultaneously, the patch merging layer reduces the number of labels by a factor of 2 and expands the feature dimension to twice its original size; this operation is performed three times within the encoder.
[0097] Patch Merging: The input block is divided into four sub-blocks, which are then merged using the Patch Merging process. This operation downsamples the feature resolution to half of its original value. Simultaneously, since the merging operation quadruples the feature dimension, a linear layer needs to be applied to the merged features to readjust the feature dimension back to twice its original size.
[0098] The decoder is built upon the Swing Transformer block and is designed to correspond to the encoder. Unlike the patch merging layer used in the encoder, the decoder employs a patch expanding layer to perform feature map upsampling. The patch expanding layer restores a higher-resolution feature map by rearranging adjacent feature map patches, achieving a 2x resolution upsampling while halving the number of feature channels to ensure feature dimension consistency.
[0099] The basic module structure of the discriminative network is based on the discriminative structure design of super-resolution generative adversarial networks
[119] , but the output layer is replaced with a convolutional layer. In this way, the discriminative network can determine whether each image is the original data or a synthetic image generated by the generative network. Precise guidance may be more conducive to the learning of the discriminative network, thereby enabling the generative network to generate higher quality images in the adversarial process. The purpose of this design is to provide more accurate gradient information to the discriminative network to promote its adversarial training, thereby driving the generative network to produce higher quality image output.
[0100] The structure of the discriminant network is as follows Figure 6 As shown, the specific configuration is described below:
[0101] The input layer receives an image input of size (C, H, W), where C represents the number of channels in the image, and H and W represent the height and width of the image, respectively.
[0102] This discriminative network comprises several basic convolutional modules, each consisting of a convolutional layer, a batch normalization layer, and a LeakyReLU activation function. In the convolutional layer, the number of input channels is converted to a specified number of output channels. The convolutional layer uses a 3×3 kernel with a stride of 1 and padding of 1. Following the convolution is a batch normalization layer, immediately followed by the LeakyReLU activation function. Subsequent convolutional layers maintain the same number of channels and use a 3×3 kernel, but with a stride of 2 and padding of 1. Each convolutional layer is followed by a batch normalization layer and a LeakyReLU activation function.
[0103] The stacked convolutional blocks are configured as follows:
[0104] The first convolutional block does not contain a batch normalization layer, has C input channels (C=3 for RGB images), and 64 output channels.
[0105] The second convolutional block contains two basic convolutional modules with 64 input channels and 128 output channels.
[0106] The third convolutional block contains two basic convolutional modules with 128 input channels and 256 output channels.
[0107] The fourth convolutional block contains two basic convolutional modules with 256 input channels and 512 output channels.
[0108] Finally, the output layer is a convolutional layer with 512 input channels and 1 output channel, using a 3×3 convolutional kernel with a stride of 1 and padding of 1. The output size generated by this layer is (1, patch_h, patch_w), where patch_h = H / 2^4 and patch_w = W / 2^4.
[0109] In tasks involving input image reconstruction using generative adversarial networks (GANs), the selection of the loss function is crucial to the effectiveness of model training. To accurately assess the quality of the reconstructed image, a combination of multiple loss functions is typically employed to measure its quality. Specifically, the design of the loss function must consider not only the pixel-level reconstruction accuracy of the image but also the reconstruction quality at the structural and perceptual levels, thereby guiding the network to generate visually satisfactory results.
[0110] In the task of restoring saturated images using generative adversarial networks, the loss function is defined as:
[0111] L=αL gen-MSE +L VGG / i,j +L msssim (1)
[0112] Where α = 0.01, L gen-MSE L VGG / i,j and L msssim These are the loss functions for generative networks, VGG-based perceptual networks, and MSSSIM, respectively.
[0113] Adversarial Loss: A traditional adversarial loss function is used to measure the realism of the generated image. The similarity between the generated image and the real image in the feature space is calculated, encouraging the generative network to generate a more realistic reconstruction.
[0114] The loss function of the generator network is defined as:
[0115]
[0116] Where G is the generator network, D is the discriminator network, MSE is the mean squared error, n is the output size of the discriminator (n=14), and x is the input image.
[0117] The discriminant network loss function is defined as:
[0118]
[0119] Where y is the reference image, i.e., Ground Truth.
[0120] Perceptual loss: measures the structural and textural similarity between the generated and original images. Mean squared error loss is used to calculate the distance between the original and generated images in the feature space, and a pre-trained VGG model is used for feature extraction. By training the model using this objective function, the reconstructed image can continuously approach the original image in the feature space.
[0121] The perceptual loss function based on VGG16 is defined as:
[0122]
[0123] Among them, W i,j H i,j The value φ() represents the feature map size, and θ represents the features extracted from the image by the VGG network. G This indicates the generated network weight parameters.
[0124] Content loss: Since the perceptual loss function is based on feature extraction, it may lead to the loss of low-level information such as texture and color in the image. Therefore, the Multi-Scale Structural Similarity (MSSSIM)
[120] loss function is introduced to compensate for this defect. The MSSSIM loss function can evaluate the visual quality of the generated image from multiple scales, thereby driving the generative network to produce restoration results that are closer to the original image in terms of color and texture details. The design aims to ensure the naturalness and authenticity of the restored image in visual perception, while maintaining the richness and integrity of image details.
[0125] The MSSSIM loss function is defined as follows:
[0126]
[0127] Where M represents different scale levels, μ p and μ g σ represents the average value of the predicted image and the reference image, respectively. p and σ g σ represents the standard deviation of the predicted image and the reference image, respectively. pg β represents the covariance between the predicted image and the reference image. m and γ m Used to indicate the relative importance between two terms. Constant terms c1 and c2 are introduced as smoothing terms to avoid the denominator approaching zero.
[0128] Third Embodiment
[0129] The compensation processing algorithm was trained and tested using real-world scene photos from the vangogh2photo dataset to evaluate its performance. To simulate the saturation phenomenon caused by glare, the training data needed to be synthesized. Specifically, the images in the dataset were first processed using Python, and their dimensions were uniformly adjusted to 224×224 pixels through downsampling. Next, a channel was randomly selected from the RGB color channels, and a circular region with a grayscale value of 255 (8-bit image bit depth) was randomly generated within that channel to simulate the saturation phenomenon in a localized area of that channel caused by glare.
[0130] This synthetic processing allows for the simulation of glare images obtained by a time-spectral separated three-channel photoelectric imaging experimental system under laser glare interference, enabling the model to learn and identify local saturation conditions that may occur in real-world scenarios. The introduction of this synthetic processing aims to enhance the model's understanding of saturation phenomena, thereby improving its performance in practical applications.
[0131] In this study, code was written using the PyTorch deep learning framework, and model training was performed on a computer equipped with an NVIDIA GeForce RTX 3090 graphics card. For model training, the Adam algorithm was chosen as the optimizer, with a learning rate of 5e-4 and a batch size of 4. The model was trained for 10 epochs using 6184 images, taking approximately 26 hours to complete.
[0132] During model training, the loss function serves to evaluate the difference between the model's predictions and the actual data. A decrease in the loss function value indicates an improvement in model performance. Figure 7 The chart shows how the loss function changes with training batches during model training. The loss function of the generator network decreases rapidly in the early stages of training and gradually stabilizes as the number of training batches increases. This indicates that the generator network gradually learns how to generate samples that closely resemble real data during training, and its generation ability continuously improves. The loss function of the discriminator network decreases rapidly in the early stages of training and begins to fluctuate after reaching a low value. This fluctuation may be because the samples generated by the generator network increasingly approximate real samples, making it difficult for the discriminator network to consistently distinguish between real and generated samples. Nevertheless, the overall trend of the loss function is still downward, meaning that the performance of the discriminator network is continuously improving.
[0133] During model training, the peak signal-to-noise ratio (PSNR) and multi-scale structural similarity (MSSSIM) curves of the generated samples output by the generative network are as follows: Figure 8 As shown, the results indicate that as the number of training batches increases, PSNR and MSSSIM rapidly increase to a stable value and fluctuate around the stable value, indicating that the generated samples output by the generator network gradually approach the real samples, and the ability of the generator network continues to improve.
[0134] To evaluate the algorithm's performance, after training, 100 images not used in the training were used for restoration. The Peak Signal-to-Noise Ratio (PSNR) and Multi-Scale Structural Similarity (MSSSIM) of the restored images were calculated. Experimental results show that the average values of the two metrics for the restored images are: PSNR = 29.243 dB and MSSSIM = 0.985. The intuitive results are as follows: Figure 9As shown, the first row represents the original image, the second row represents the glare image, and the third row represents the restored image. The results show that the algorithm can recover the information lost due to glare, and the restored image has high clarity, fine texture, good color fidelity, and a visual effect very close to the original image.
[0135] Further experimental research will be conducted to verify the glare suppression performance of this technology. After an image sensor captures an image, it needs to undergo a digital image processing (ISP) process to convert the raw sensor data into a high-quality image before it can be stored or displayed. The differences in image data captured by color cameras and monochrome cameras lead to different ISP processes. The purpose of the ISP process in a color camera is to reconstruct a color image from the image data captured by an image sensor equipped with a color filter array. This process typically includes the following key steps:
[0136] 1. De-mosaic: This step uses an interpolation algorithm to process the incomplete color image captured by the color filter array into a complete three-channel color image.
[0137] 2. White Balance Adjustment: The purpose of this step is to correct the color temperature of the image to ensure the naturalness of color rendering, especially for the accurate reproduction of white and neutral tones.
[0138] 3. Color Correction: This step corrects and optimizes the colors of the image to compensate for color deviations in the image sensor and optical system, ensuring color accuracy and consistency.
[0139] 4. Tone Mapping and Gamma Correction: This step adjusts the contrast and brightness of the image, optimizes the dynamic range, and uses gamma correction to adapt to the non-linear response of the human eye to brightness.
[0140] 5. Sharpening and noise suppression: Sharpening algorithms enhance the clarity of details, while noise suppression algorithms aim to reduce visual noise in images and improve image quality.
[0141] 6. Compression: Compress images using data compression technology to save storage space and optimize transmission efficiency.
[0142] Monochrome cameras do not require color information processing, so their ISP process is relatively simple, with the main steps including tone mapping and gamma correction, sharpening, noise suppression, and compression.
[0143] A three-channel photoelectric imaging system uses a monochrome camera to capture color images, requiring processing such as depigmentation, white balance adjustment, and color correction. First, the monochrome camera captures images of a white card in three channels (RGB) over time. Depigmentation is achieved by directly overlaying the three channel images, as shown in the image. Figure 10 As shown, the white card exhibits an overall color cast.
[0144] Next, use a white card to set the correct white balance adjustment. Figure 10 It can be seen that the green channel has the highest brightness, therefore, taking the green channel as the reference, its gain coefficient is set to 1. The specific method for white balance adjustment is to calculate the gain coefficients for the red and blue channels of the image, and then multiply the coefficients by the corresponding channel. The calculation formula is as follows:
[0145]
[0146] In the formula, R′, G′, and B′ are the three-channel pixel values after white balance adjustment, and k r and k b These are the white balance gain coefficients for the red and blue channels, respectively, where R, G, and B are the pixel values before white balance adjustment. The white card region in the image is located and identified. The mean values of the pixels in this region for the red, green, and blue channels are calculated. The white balance gain coefficients for the red and blue channels are calculated by dividing the mean values of the red and green channels by the mean values of the blue channel, resulting in values of 3.77 and 1.37, respectively. The white balance adjustment result is as follows: Figure 11 As shown, the color shift of the white card is suppressed, ensuring the authenticity of the white card's colors and the consistency of the visual effect.
[0147] Then, color correction is achieved by adjusting colors using a color correction card. The position of the color correction card in the image is detected, and the spatial and color information of the color patches within it is measured. Next, based on the RGB color reference values of the color patches, a color correction matrix is calculated using the least squares method. This matrix can optimally map the RGB color measurements of the color patches to their reference values through affine transformation. Simultaneously, this matrix can be used to correct the colors in images acquired by the imaging system, making the colors of images captured under different lighting conditions closer to the standard or desired colors. The color correction process of the imaging system can be represented by a color correction matrix as follows:
[0148]
[0149] In the formula, R′, G′, and B′ are the three-channel pixel values after color correction processing, and R, G, and B are the pixel values before color correction processing. Figure 12 As a result of the color correction process, the color accuracy, contrast, and saturation of the color chart are significantly improved, which greatly enhances the overall tone of the image.
[0150] Finally, by calculating the gain coefficient and color correction matrix, color image capture of a three-channel photoelectric imaging system using a monochrome camera is achieved.
[0151] The preceding discussion focused on the imaging performance of the three-channel photoelectric imaging system without the influence of glare lasers. Next, we will investigate the imaging performance of the system under the influence of glare lasers and further test the glare suppression performance of the band compensation technique.
[0152] The three-channel photoelectric imaging system was illuminated with a 532nm laser. Images of the red, green, and blue channels were obtained at a laser power of 36.16μW, as shown below. Figure 13 As shown in (a), (b), and (c), it can be seen that the images captured by the red and blue channels are darker, while the images captured by the green channel are brighter. Figure 13 (d) is the composite color image obtained after white balance adjustment and color correction. The results show that the saturation area of the green channel is larger, resulting in severe glare, while the saturation areas of the red and blue channels are smaller, resulting in less glare.
[0153] The glare effect of a color camera was tested under the same conditions, serving as a control group for the glare effect of the three-channel photoelectric imaging system. The glare effect of the color camera is as follows: Figure 14 As shown in (a). Figure 14 (b), (c), and (d) are grayscale images of the corresponding red, green, and blue channels, respectively. The results show that when the glare laser is green, the red, green, and blue channels of the color camera are simultaneously severely saturated. This phenomenon is caused by crosstalk from the image sensor. Comparing the glare effects of the three-channel photoelectric imaging system and the color camera, it can be seen that the advantage of the three-channel photoelectric imaging system lies in its ability to suppress glare interference from two channels and maintain better imaging performance.
[0154] Further tests were conducted on the glare characteristics of the three-channel photoelectric imaging system and the color camera under different laser energies. The number of saturated pixels in the red, green, and blue channels was quantitatively calculated. The experimental images were 8-bit deep, and the saturation threshold was set to 254. The statistical results of the number of saturated pixels are listed in Table 2. The results show that the red and blue channels of the three-channel photoelectric imaging system can effectively suppress glare interference. Meanwhile, the number of saturated pixels in the green channel of the three-channel photoelectric imaging system is greater than that of the color camera. This is because the monochromatic camera used in the three-channel photoelectric imaging system has a higher quantum efficiency at 532 nm wavelength than the color camera.
[0155] Table 2 Number of saturated pixels in red, green and blue channels at different laser energies
[0156]
[0157] In the research on compensatory processing algorithms, an experimental scene including a toy remote-controlled car and a lawn was constructed. A three-channel photoelectric imaging system was used for image acquisition, and color images were obtained after white balance adjustment and color correction. Typical experimental images acquired by the imaging system are shown below. Figure 15 As shown, (a), (b), and (c) are experimental images captured by the R, G, and B channels, respectively, and (d) is the color image obtained after white balance adjustment and color correction.
[0158] Initially, 300 grayscale images were captured. After white balance adjustment and color correction, 100 color images were obtained. Fourteen images with obvious quality problems due to drastic lighting changes were removed, resulting in 86 color images for the experiment. These color images were segmented to form a training set of 320 images and a test set of 24 images. The glare spots measured experimentally were used as a mask to preprocess the training set images to simulate the glare effect. The preprocessed training set was then used to train a compensation processing algorithm, with a batch size of 4 and a training epoch of 25.
[0159] To evaluate the algorithm's performance, a test set was used for image restoration after training, and the Peak Signal-to-Noise Ratio (PSNR) and Multi-Scale Structural Similarity (MSSSIM) of the restored image were calculated. Experimental results show that the average values of the two metrics for the restored image are: PSNR = 29.358 dB and MSSSIM = 0.958. Figure 16 The results show the algorithm's visual output, including the original image (first row), the glare image (second row), and the restored image (third row). The results demonstrate that the algorithm can provide reasonable restoration of the glare area, with the R and B channels showing better restoration performance than the G channel. This is likely because the grass background is predominantly green, leading to significant information loss in the G channel under glare, thus affecting the restoration results.
[0160] In summary, the band compensation technology of this invention has laser glare suppression performance far superior to that of conventional optoelectronic imaging systems, and can greatly improve the adaptability of the imaging system in complex optoelectronic environments.
[0161] Please note that the technical features of the above embodiments can be combined arbitrarily. For the sake of brevity, not all possible combinations of the technical features in the above embodiments have been described. However, as long as the combination of these technical features does not contradict each other, it should be considered within the scope of this specification. The above embodiments only illustrate several implementation methods of this application, and their descriptions are relatively specific and detailed, but they should not be construed as limiting the scope of the invention patent. It should be pointed out that for those skilled in the art, several modifications and improvements can be made without departing from the concept of this application, and these all fall within the protection scope of this application. Therefore, the protection scope of this patent application should be determined by the appended claims.
Claims
1. A band-compensation imaging method for laser glare suppression, characterized in that, The method utilizes a three-channel photoelectric imaging system with time-spectral separation and an image restoration algorithm. Through joint hardware and software optimization, it recovers the lost image information caused by glare, thereby achieving laser glare suppression; wherein: For the aforementioned time-spectral separation three-channel photoelectric imaging system, the three-color RGB filter or four-color CMYK filter used in the color camera with overlapping transmission bands is replaced with a bandpass filter, thereby limiting the glare effect of interfering laser on the color imaging system to one band. The image restoration algorithm uses the color statistical attributes of the reference image of the target scene or area scene before glare interference, combined with the image information of the un-glare-interference band, to predict the image information of the glare-interference band through machine learning, and then restores the color image under the condition of no glare interference through image fusion. When there is no interfering laser, the three-channel photoelectric imaging system transmits the received monochrome image to the image processing system for de-mosaicing, white balance and color correction to obtain a color image that meets the clarity requirements and outputs it for display. In the presence of interfering lasers, one channel of the three-channel photoelectric imaging system is affected by laser glare interference, causing a saturation area in the image sensor and resulting in the loss of image information. The lost image information caused by the glare interference is recovered by using the image information of the unaffected band and the image recovery algorithm, thereby suppressing laser glare interference. In the method, the time-spectral separation three-channel photoelectric imaging system employs the time-spectral separation principle; wherein: Temporal spectral separation refers to the structure that uses optical field modulation elements to separate the mixed optical field of interfering laser and signal light in time and image it onto an image sensor. The optical field modulation of temporal spectral separation is achieved by a filter wheel mounting bracket, a sliding filter mounting bracket, and a filter. In the method described, the three-channel photoelectric imaging system includes: a collimating and expanding telescope C, dichroic mirrors DBS1 and DBS2, bandpass filter groups BP1, BP2, and BP3, imaging lenses L1, L2, and L3, and image sensors S1, S2, and S3; wherein, the bandpass bands of the bandpass filter groups BP1, BP2, and BP3 are selected such that the corresponding imaging spectral ranges of the image sensors S1, S2, and S3 are independent and do not overlap; wherein: Bandpass filter groups BP1, BP2 and BP3 are respectively placed in a mounting slot of filter wheel FW. By rotating the filters, the first channel Ch.1, the second channel Ch.2 and the third channel Ch.3 are imaged step by step. The spectral images of the three channels are obtained by time-spectral separation. The center wavelengths of the bandpass filter groups BP1, BP2 and BP3 are 450±5nm, 550±5nm and 650±5nm, respectively, the full width at half maximum (FWHM) of the spectrum is 50±5nm, and the stopband optical density OD≥4.
2. The band-compensation imaging method for laser glare suppression according to claim 1, characterized in that, In the method: When there is a single wavelength interfering laser in the external environment, the interfering laser will cause glare and interference to one channel of the image sensor, while the remaining two channels of the image sensor will work normally. When there is dual-wavelength interfering laser in the external environment, the interfering laser glare interferes with the image sensors of two channels, while the image sensor of the other channel works normally. When there are three or more interfering lasers in the external environment, no practical interfering laser source is found, and the image sensor of at least one channel is working normally.
3. The band-compensation imaging method for laser glare suppression according to claim 2, characterized in that, In the method, the image restoration algorithm employs a generative adversarial network, including a generator network and a discriminator network. The generator network adopts the Swin-Unet network structure, which is an image segmentation network that integrates the Swin Transformer model and the classic U-Net architecture. The basic module structure of the discriminator network adopts the discriminator structure of a super-resolution generative adversarial network, and replaces the output layer with a convolutional layer. Wherein: The Swin-Unet network employs an encoder structure that initially processes the input image, converting it into a sequential embedding representation. Specifically, the input image is first segmented into several non-overlapping 4×4 pixel blocks. This segmentation process transforms the feature dimension of each block into 4×4×3, and a linear embedding layer maps the feature dimension to the model-defined dimension C. The transformed patch markers then undergo feature extraction through several Swin Transformer blocks and a patch merging layer. The patch merging layer performs downsampling and increases the feature dimension, generating hierarchical feature representations, integrating local features, and reducing sequence length. The Swin Transformer layer then extracts feature information. For the encoder, the C-dimensional input image is fed into two consecutive Swing Transformer blocks for feature representation learning, during which the feature dimension and resolution remain unchanged; at the same time, the patch merging layer reduces the number of labels by 2x downsampling and expands the feature dimension to twice the original size. This operation is performed three times inside the encoder. For the patch merging layer, the input block is divided into 4 sub-blocks, which are then merged through the patching process, downsampling the feature resolution to half of the original. At the same time, since the merging operation increases the feature dimension by four times, a linear layer is applied to the merged features to readjust the feature dimension back to twice the original size. For the decoder, it is built on the basis of the Swing Transformer block. Corresponding to the encoder, the decoder uses a patch extension layer to perform feature map upsampling. The patch extension layer restores the higher resolution feature map by rearranging adjacent feature map blocks, achieving 2x resolution upsampling. At the same time, the number of feature channels is halved to maintain feature dimension consistency.
4. A band-compensation imaging method for laser glare suppression according to claim 3, characterized in that, In the method, a discriminant network is used to determine whether each image is original data or a synthetic image generated by a generative network. The structure of the discriminant network includes: The input layer is used to receive image input of size (C, H, W), where C represents the number of channels of the image, and H and W represent the height and width of the image, respectively. Several basic convolutional modules are provided, each of which includes a convolutional layer, a batch normalization layer, and a LeakyReLU activation function. In the convolutional layer, the number of input channels is converted into the specified number of output channels. The convolutional layer uses a 3×3 kernel with a stride of 1 and edge padding of 1. Batch standardization layer; LeakyReLU activation function; Several convolutional layers with the number of channels remaining constant, using 3×3 convolutional kernels with a stride of 2 and edge padding of 1, followed by batch normalization layers and LeakyReLU activation function after each convolutional layer; The stacked convolutional blocks are configured as follows: The first convolutional block: does not contain a batch normalization layer, has C input channels (C=3 for RGB images), and 64 output channels; The second convolutional block contains two basic convolutional modules with 64 input channels and 128 output channels. The third convolutional block contains two basic convolutional modules with 128 input channels and 256 output channels. The fourth convolutional block contains two basic convolutional modules with 256 input channels and 512 output channels. Finally, the output layer is a convolutional layer with 512 input channels and 1 output channel. It uses a 3×3 convolutional kernel with a stride of 1 and padding of 1. The output size generated by this layer is (1, patch_h, patch_w), where patch_h = H / 2^4 and patch_w = W / 2^4.
5. A band-compensation imaging method for laser glare suppression according to claim 4, characterized in that, In the method described above, a generative adversarial network is used to recover the input image, and the loss function is: in, , , and These are the loss functions for generative networks, VGG-based perceptual networks, and MSSSIM, respectively. The loss function for the generator network is: in, It is a generative network. It is to identify the network. It is the mean square error. It is the output size of the discriminator. Input image; The discriminant network loss function is: in, For reference image; The perceptual loss function based on VGG is: Where WH represents the feature map size. Indicates the features extracted from the image by the VGG network, subscript This represents the generated network weight parameters; The MSSSIM loss function is: in, Indicates different scale levels, and These represent the average values of the predicted image and the reference image, respectively. and These represent the standard deviations of the predicted image and the reference image, respectively. This represents the covariance between the predicted image and the reference image. and The constant term indicates the relative importance between two terms. and It was introduced as a smoothing term.