Active single-pixel imaging method for strongly scattering media based on conditional generative adversarial networks

By designing a single-pixel optical imaging system based on Hadamard transform and an AMSPI-LSCGAN model of the least squares condition generation adversarial network, the imaging quality problems in low sampling rate and strong scattering environments are solved, and clear imaging under high turbidity conditions is achieved, which is suitable for single-pixel scattering imaging systems.

CN115601621BActive Publication Date: 2025-09-02HUBEI UNIV OF TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202211264651.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-10-17
Publication Date
2025-09-02
Estimated Expiration
2042-10-17

AI Technical Summary

Technical Problem

The existing single-pixel imaging methods have poor imaging quality in low sampling rates and strong scattering environments, especially when the turbidity is greater than 50 NTU, the target reconstruction takes a long time and the imaging quality is greatly affected by the scattering medium.

Method used

A single pixel optical imaging system based on Hadamard transformation is designed, and an AMSPI-LSCGAN model based on the least squares condition generated adversarial network is constructed. Combining the least squares loss, content loss and average structural similarity loss as joint loss functions, image reconstruction is carried out through a deep convolutional network, enhancing the imaging robustness under strong scattering conditions.

Benefits of technology

The image reconstruction quality is significantly improved at low sampling rate, and the object profile can be distinguished when the water turbidity reaches 144 NTU. It has high robustness and versatility. It is suitable for single-pixel scattering imaging systems and improves the target imaging quality in a strong scattering environment.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115601621B_ABST
    Figure CN115601621B_ABST
Patent Text Reader

Abstract

The present invention discloses an active single-pixel imaging method for strongly scattering media based on a conditional generative adversarial network. The method comprises the following steps: designing a single-pixel optical imaging system based on a Hadamard transform to obtain one-dimensional detection signal values ​​of a test target under a series of different scattering conditions and sampling rates; shaping the sampled one-dimensional detection signal values ​​under different scattering conditions and sampling rates into two-dimensional signal images through data averaging preprocessing, which are used as test set images; constructing a deep convolutional conditional generative adversarial network (AMSPI-LSCGAN) based on a least squares loss function model with a compression-excitation block and a residual block; combining the least squares loss with a content loss and an average structural similarity loss as the loss function of the AMSPI-LSCGAN network for training, thereby avoiding training collapse and obtaining reconstructed images with higher fidelity; and inputting the test set images into the trained AMSPI-LSCGAN network to reconstruct the test target.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a single-pixel imaging method for a strongly scattering medium in an active mode based on a least squares condition-generated adversarial network at a low sampling rate, and belongs to the technical field of optical imaging analysis. Background Art

[0002] It is well known that when light passes through scattering media such as turbid liquids, haze, smoke, and biological tissue, due to the uneven refractive index distribution within the medium, it undergoes multiple, unknown scatterings and is dispersed in all directions. This prevents the complete collection and utilization of the ballistic light containing object information. Ultimately, only a speckle pattern is detected on the observation surface, making it difficult to image the target. This makes imaging through strongly scattering media a significant challenge. In recent years, the application of computational optics imaging to scattering and the emergence of high-precision wavefront modulation devices such as spatial light modulators (SLMs) and digital micromirror devices (DMDs) have further promoted research in scattering imaging. For example, methods such as wavefront shaping, transfer matrices, and speckle autocorrelation have all opened up possibilities for imaging through scattering media. However, these methods are sensitive to changes in the scattering medium, and imaging quality is easily affected by the environment. Therefore, they place very high demands on the target object selection and system stability. In recent years, single-pixel imaging (SPI) has attracted widespread attention as a new measurement technique. It does not require pixelated photodetectors to detect the light signal. Instead, it uses a single point detector to measure the integrated intensity of the object after illumination, thus simplifying the complexity of the experimental system. This new imaging mechanism provides a new solution for low-light imaging and imaging through scattering media. However, target reconstruction in the compressed sensing algorithm based on single-pixel imaging is limited by the selected sparse basis, and target reconstruction is time-consuming and the imaging quality is affected by the scattering medium.

[0003] In recent years, deep learning (DL) algorithms, also known as deep neural networks (DNNs), have been widely used to solve inverse problems. By learning the internal patterns of sample data, neural networks can learn and adapt to these mapping relationships through known mapping examples, thus possessing powerful computational and fitting capabilities. Therefore, deep learning has also been used to solve scattering imaging problems to improve the signal-to-noise ratio of objects reconstructed in scattering media. However, for existing DL-based scattering media computational imaging reconstruction schemes, the target imaging quality in strong scattering media (turbidity greater than 50 NTU) and at low sampling (0 ≤ sampling rate ≤ 30%) remains less than ideal. Summary of the Invention

[0004] In order to solve the problem of low imaging quality of existing single-pixel imaging methods at low sampling rates and in strong scattering environments, the present invention provides an active mode single-pixel imaging method for strongly scattering media based on a least squares conditional generative adversarial network.

[0005] The technical solution employed by the present invention is to first design and implement a single-pixel optical imaging system (HSPI) based on the Hadamard transform. The system comprises a He-Ne laser, a digital micromirror device (DMD), a data acquisition card, a single-pixel detector, a beam terminator, various lenses, and supporting components. The imaging principle is as follows: After laser light is collimated and expanded, it is modulated on the DMD, which has been pre-loaded with a Hadamard speckle pattern. When a pixel in the speckle pattern is 1, the DMD's micromirror is set to the "on" state; when a pixel in the speckle pattern is 0, the DMD's micromirror is set to the "off" state. The reflected light then passes through a test target in a standard turbid solution. Subsequently, a single-pixel detector synchronously acquires and records the total light intensity. Furthermore, the DMD projection rate and the DAQ system of the digital acquisition card are synchronized by a trigger signal. Finally, the resulting light intensity signal is shaped into a two-dimensional image and input into a trained network to reconstruct images of the test target after passing through different scattering media at different sampling rates.

[0006] In order to achieve the above object, the technical solution provided by the present invention is: a single-pixel imaging method based on a least squares conditional generative adversarial network, comprising the following steps:

[0007] Step 1: Design a single-pixel optical imaging system based on Hadamard transform to obtain the one-dimensional detection signal value of the test target under a series of different scattering conditions and sampling rates;

[0008] Step 2: The one-dimensional detection signal values ​​obtained under different scattering conditions and sampling rates are pre-processed by data averaging and then reshaped into a two-dimensional signal image as the test set image;

[0009] Step 3: Construct a deep convolutional conditional generative adversarial network AMSPI-LSCGAN based on a least squares loss function model with a compression-excitation block and a residual block;

[0010] Step 4: Combine the least squares loss with the content loss and the average structural similarity loss as the loss function of the AMSPI-LSCGAN network for training to avoid training collapse and obtain higher-fidelity reconstructed images.

[0011] In step 5, the test set image is input into the AMSPI-LSCGAN network trained in step 4 to reconstruct the test target.

[0012] Furthermore, the single-pixel optical imaging system in step 1 includes a He-Ne laser, a digital micromirror device (DMD), a data acquisition card, a single-pixel detector, a beam terminator, various lenses, and supporting parts. The imaging principle is as follows: after the light emitted by the laser is collimated and expanded, it is modulated on the DMD pre-loaded with Hadamard speckle. When a pixel in the speckle pattern is 1, the micromirror of the DMD is set to the "on" state; when a pixel in the speckle pattern is 0, the micromirror of the DMD is set to the "off" state. The reflected light then passes through a test target in a standard turbid liquid, and the single-pixel detector then synchronously collects and records the total light intensity.

[0013] Furthermore, in step 3, the datasets used for network training come from the MNIST and EMNIST handwriting datasets. First, the image size is adjusted. Then, each image is multiplied by the Hadamard measurement matrix and all pixels are normalized into a column of one-dimensional detection values ​​to simulate the measurement values ​​collected by a single-pixel detector. Finally, the network is trained on image pairs consisting of signal images formed by preprocessing and normalizing the one-dimensional detection values ​​and real images to obtain optimized network models at different sampling rates.

[0014] Furthermore, the Hadamard measurement matrix is ​​calculated as follows;

[0015] When the matrix is ​​composed of +1 and -1 elements and satisfies HH T = NZ, such a matrix is ​​called a Hadamard matrix, H is an N-order square matrix, H T is the transpose of H, Z is the N-order unit matrix; for any matrix with row and column dimensions of 2 m The Hadamard matrix of is used as the measurement matrix, which can be recursively obtained by the following formula:

[0016]

[0017] Among them, 2 m-1 represents the matrix dimension. The Hadamard matrix is ​​generated by random Hadamard transformation. The matrix H is passed through the digital micromirror device (DMD) to obtain the illumination speckle P(x, y). (x, y) represents the image coordinates. Since the matrix H is composed of +1 and -1, and the DMD can only be used to debug the measurement matrix composed of 0 and 1, an H is designed in the complementary mode, which can be obtained by the following formula:

[0018]

[0019] Where E is a matrix whose elements are all 1; H + The +1 elements in H are retained, and the -1 elements in H are converted to 0; H -Indicates that the elements originally 1 in H are converted to 0, and the elements -1 are replaced by +1; in this way, H can be modulated on the DMD device. + With H - matrix to obtain the matrix H.

[0020] Furthermore, the specific implementation of step 3 is as follows:

[0021] AMSPI-LSCGAN consists of a generator network G and a discriminator network D. The generator network G is a deep convolutional U-shaped network structure that restores the original image through an encoder consisting of four downsampling convolution modules, a decoder consisting of four upsampling convolution modules, and three residual blocks added between the encoder and decoder.

[0022] The encoder stage is a downsampling process. The input is an image of a certain size directly reshaped by the measured signal value of the barrel detector. The convolution layer is used to extract image features, and the maximum pooling layer is used to reduce the spatial dimension of the image. The information in the input image is encoded to better understand the edge and texture structure of the image. A squeeze-excitation SE block is added after each convolution layer. The SE block consists of two parts: compression and excitation. W and H are set to represent the width and height of the feature map respectively, C represents the number of channels, and the size of the input feature map is W×H×C. The first step is the compression operation. The input feature map is compressed into a 1×1×C vector through global average pooling. This vector has a global receptive field to some extent, and the output dimension matches the number of input feature channels; the second step is the excitation operation, which consists of two fully connected layers. The output is a 1×1×C vector. Finally, the weight value of each channel calculated by the SE block is multiplied by the two-dimensional matrix of the channel corresponding to the original feature map to obtain the final output result.

[0023] Furthermore, the calculation formula of the residual block is as follows;

[0024]

[0025] In Equation 8, p is the input of the residual block, q is the output of the residual block, and W i Represents the parameters of the i-th layer, obtained through training, f(p,W i ) represents the residual mapping, as well as are the partial derivatives of p and q respectively. For the network structure of the residual block, the input features are first passed through two 3×3 convolutional layers to obtain the residual map, and then the input is added to the output through the shortcut connection to complete the feature fusion.

[0026] Furthermore, the specific loss function in step 4 is as follows;

[0027] Least squares loss function LLSGAN as follows:

[0028]

[0029] Among them, x is the real sample, P x is the actual sample distribution; z is the signal value obtained by the single pixel system, P z is the generated sample distribution defined by the input generation network G(z), E represents the mathematical expectation, and D represents the discriminant network; b is set to 1 to represent real data; a is set to 0 to represent forged data; c is set to 0 to deceive the discriminant network D; the use of the least squares loss function will prevent the gradient of D from decreasing to 0, and the data at the boundary will also receive a penalty proportional to the distance, thereby ensuring that the network obeys more gradient information, thereby improving the stability of training;

[0030] In order to make the reconstructed image closer to the true value, the loss function based on the mean absolute error (MAE) is selected to minimize the pixel-level difference between the real image and the generated image. The objective function is:

[0031] L L1 =E[||xG(z)||1] (11)

[0032] ||·||1 represents the L1 norm;

[0033] Then, the average structural similarity loss is used to evaluate the quality of the entire image, which can be expressed as:

[0034]

[0035] where u x and u G(z) are the average values ​​of the total pixels in the real image and the reconstructed image, σ x 2 and σ G(z) 2 is the variance between the real image and the reconstructed image, σ xG(z) is the covariance of the real image and the reconstructed image; C1, C2 and C3 are constants used to avoid division by 0; u x , σ x and σ xG(z) By using a circularly symmetric Gaussian weighted matrix window to calculate, and then using MSSIM to evaluate the quality of the entire image, it can be expressed as:

[0036]

[0037] Among them, U is the real image block, R is the predicted image block, x v represents the real image content of the vth window, G v(z) represents the image content generated at the vth window, and K represents the total number of windows; therefore, the loss of MSSIM can be expressed as:

[0038]

[0039] Therefore, the final joint loss function is as follows:

[0040] L G =λ1L LSGAN +[λ2L MSSIM +(1-λ2)L L1 ]λ3 (15)

[0041] Among them, λ1, λ2, and λ3 are all constants.

[0042] Compared with the prior art, the advantages and beneficial effects of the present invention are:

[0043] (1) The present invention discloses an active mode single-pixel imaging method for strongly scattering media based on a least squares conditional generative adversarial network. An active mode single-pixel imaging (AMSPI) based on a least squares conditional generative adversarial network algorithm (AMSPI-LSCGAN) is designed and proposed. In this model, the generator adopts a U-shaped structure and adds a jump connection with an attention gate to enhance the salient features of the target under strong scattering conditions. In the encoder and decoder structures, a compression-excitation (SE) block is added to better remove noise and redundant feature information and improve the reconstruction effect of the network. Adding a residual block between the encoder and decoder structures can not only reuse feature information but also avoid the problem of gradient vanishing leading to network training collapse. At the same time, in order to further improve the image reconstruction quality of AMSPI under low sampling rate strong scattering conditions, it is proposed for the first time to combine the least squares loss, content loss and mean structural similarity loss (MSSIM) into a joint loss function. Adding the MSSIM loss can significantly improve the image perception quality and effectively eliminate image artifacts and redundant noise. In addition, the method has certain robustness to scattering media and has practical application value for denoising and enhancing scattering medium imaging.

[0044] (2) The present invention discloses an active mode single-pixel imaging method for strongly scattering media based on a least squares conditional generative adversarial network. An active mode single-pixel imaging system based on a Hadamard pattern is designed and constructed, achieving strong scattering medium imaging at a low sampling rate. When the water turbidity reaches 144 NTU and the sampling rate is 19.14%, the traditional CGI method can still distinguish the outline of the object. Therefore, the imaging system constructed by the present invention exhibits excellent performance in imaging strongly scattering media.

[0045] (3) The present invention discloses an active mode single-pixel imaging method for strongly scattering media based on a least squares conditional generative adversarial network. In practical applications, it is only necessary to input a set of one-dimensional bucket detector signals in physical experiments into a trained network to reconstruct the target image without adding additional processing and devices. It is suitable for current single-pixel scattering imaging systems and has strong versatility and practicality.

[0046] (4) The present invention discloses an active mode single-pixel imaging method for strongly scattering media based on a least squares conditional generative adversarial network, which uses deep learning as a powerful tool to further develop application scenarios of single-pixel imaging, that is, combining the active mode single-pixel imaging system model based on the Hadamard pattern with the deep learning data model to greatly improve the target imaging quality in a strong scattering environment and a low sampling rate.

[0047] (5) The present invention discloses an active mode single-pixel imaging method for strongly scattering media based on a least squares conditional generative adversarial network. Experiments have shown that good measurement noise characterization can improve the peak signal-to-noise ratio and structural similarity of the reconstructed image under strong scattering conditions, effectively promoting the combination of single-pixel scattering medium imaging and deep learning methods. BRIEF DESCRIPTION OF THE DRAWINGS

[0048] Figure 1 This is a detailed flow chart of the active mode single-pixel scattering imaging system and experimental method based on conditional generative adversarial networks. Figure 1 (a) is a schematic diagram of the measurement optical path. Figure 1 (b) represents data preprocessing, Figure 1 (c) represents network training, Figure 1 (d) represents network reconstruction.

[0049] Figure 2 It is the network framework of AMSPI-LSCGAN. Figure 2 (a) represents the generation network, Figure 2 (b) represents the identification network.

[0050] Figure 3 The network module structure. (a) Compression and excitation block structure. (b) Residual block structure.

[0051] Figure 4 are the reconstruction results of different recovery methods under low scattering and strong scattering conditions at different sampling rates. Figure 4 (a) shows the reconstruction results of the CGI, CSCGI, Pix2Pix and AMSPI-LSCGAN methods under low scattering conditions (SNR=25dB) when the number "0" is sampled at rates of 19.14%, 14.06%, 9.77% and 3.52% respectively. Figure 4(b) shows the reconstruction results of the digit 0 using the CGI, CSCGI, Pix2Pix and AMSPI-LSCGAN methods under strong scattering conditions (SNR = 10dB) at sampling rates of 19.14%, 14.06%, 9.77% and 3.52% respectively.

[0052] Figure 5 are the reconstruction results of different methods under strong scattering and low sampling rate conditions. Figure 5 (a) and (b) show the reconstruction results of the letter "F" using the CGI, CSCGI, Pix2Pix, and AMSPI-LSCGAN methods under strong scattering conditions (SNR = 13dB and 10dB) at sampling rates of 9.77% and 3.52%. Figure 5 (c) and (d) show the reconstruction results of the letter “H” using the CGI, CSCGI, Pix2Pix, and AMSPI-LSCGAN methods under strong scattering conditions (SNR = 13 dB and 10 dB) at sampling rates of 9.77% and 3.52%.

[0053] Figure 6 These are the network generalization simulation test results with sampling rates of 14.06% and 9.77% under scattering conditions of SNR = 20dB and 15dB.

[0054] Figure 7 This is a diagram of the experimental setup.

[0055] Figure 8 In the physical experiment, under different turbidity conditions (57, 86, 115 and 144 NTU), the number "7" and the letter "F" were reconstructed using different methods under different sampling rates. Figure 8 (a) and (c) show the reconstruction results of CGI, CSCGI, Pix2Pix and AMSPI-LSCGAN methods when the sampling rate of the test target is 19.14%. Figure 8 (b) and (d) show the reconstruction results of CGI, CSCGI, Pix2Pix and AMSPI-LSCGAN methods when the sampling rate of the test target is 3.52%.

[0056] Figure 9 Here are the PSNR and SSIM of the reconstructed images of the number "7" and the letter "F" at different sampling rates under different turbidity conditions (57, 86, 115 and 144 NTU). Figure 9 (a), (c) and (b), (d) show the PSNR and SSIM of the reconstructed image of the number “2” at sampling rates of 19.14% and 3.52%, respectively. Figure 9(e), (g) and (f), (h) show the PSNR and SSIM of the reconstructed images of the letter “F” at sampling rates of 19.14% and 3.52%, respectively.

[0057] Figure 10 These are the network generalization physical experiment test results under the scattering conditions of 57NTU, 71NTU, 86NTU, and 100NTU, with sampling rates of 19.14% and 3.52% respectively.

[0058] Figure 11 Figure 2 shows the reconstruction results of the previous training without noise addition and the new training with noise addition. (a) Reconstruction results of the word "F" under scattering conditions of 57NTU, 86NTU, 115NTU, and 144NTU at a sampling rate of 19.14%. (b) Quantitative comparison of PSNR. (c) Quantitative comparison of SSIM. DETAILED DESCRIPTION

[0059] The technical solution of the present invention will be further described below with reference to the accompanying drawings and embodiments.

[0060] The active mode single-pixel imaging method for strongly scattering media based on a least squares conditional generative adversarial network includes the following steps:

[0061] Step 1: Build a model of active-mode single-pixel scattering imaging based on Hadamard patterns. The specific implementation includes the following sub-steps:

[0062] Step 1.1: Hadamard measurement matrix: When the matrix is ​​composed of +1 and -1 elements and satisfies HH T = NZ, such a matrix is ​​called a Hadamard matrix (H is an N-order square matrix, H T (where H is the transpose of H and Z is the N-order unit matrix). We chose this matrix as the measurement matrix because the Hadamard matrix is ​​orthogonal and all values ​​are binary, so it has the characteristics of fast transformation. In addition, differential measurement is used in actual measurement, which is more resistant to external noise and more suitable for SPI based on DMD modulation. More importantly, the Hadamard measurement matrix has higher reconstruction efficiency due to its low computational complexity. For any matrix with row and column dimensions of 2 m The Hadamard matrix of is used as the measurement matrix, which can be recursively obtained by the following formula:

[0063]

[0064] Among them, 2 m-1Denotes the matrix dimension. The Hadamard matrix is ​​generated by a randomized Hadamard transform. Matrix H is passed through a digital micromirror device (DMD) to obtain the illumination pattern P(x,y), where (x,y) represents the image coordinates. Since matrix H consists of +1s and -1s, and the DMD can only be used to debug measurement matrices consisting of combinations of 0s and 1s, we design an H in complementary mode, which can be obtained from the following equation:

[0065]

[0066] Where E is a matrix whose elements are all 1; H + The +1 elements in H are retained, and the -1 elements in H are converted to 0; H - Indicates that the elements originally 1 in H are converted to 0, and the elements -1 are replaced by +1. In this way, H can be modulated on the DMD device. + With H - matrix to obtain the matrix H.

[0067] Step 1.2: Assume that the measurement matrix is ​​H i + , where i = 1, 2, 3, ..., M represents the i-th measurement, and M is the number of measurements. After being modulated by the DMD, it passes through the scattering medium to reach the target and is finally captured by the single-pixel detector to obtain the measurement value S i + It can be expressed as:

[0068]

[0069] In formula 3: I(x,y) represents the target information; P i +′ (x, y) represents the illumination speckle distortion caused by the influence of the scattering medium in the path, which is represented by the measurement matrix H + Similarly, we can get i - The matrix is ​​used as the measurement matrix for the illumination speckle distortion P i -′ (x,y) and the measured value S i - , then the intensity value S can be calculated i :

[0070] S i =S i + -S i - =∑ x,y [P i +′ (x,y)-P i -′(x,y)]I(x,y) (4)

[0071] After M measurements, the reconstructed object information I*(x,y) can be obtained by formula (4):

[0072] I * (x,y)= i P i ′(x,y)>- i > <P i ′(x,y)> (5)

[0073] In formula 5, <·> represents the average value of M measurements, P i ′ represents the illumination speckle intensity distribution after differential measurement. However, thousands of measurements are usually required to reconstruct a higher-quality image. Here, a compressed sensing (CS) algorithm can be introduced to reduce the number of measurements. In compressed sensing computational ghost imaging (CSCGI), the entire speckle pattern distribution is often represented by a matrix. Therefore, the pattern can be converted into a row vector to form a measurement matrix P, where P is an M×N matrix and N is the number of elements in each speckle pattern. The transmittance distribution of the target can be converted into a column vector matrix X. Therefore, the column vector matrix S formed by the measurement signal can be expressed as:

[0074] S=PX (6)

[0075] Since the solution of X is not unique, the transformation process can be expressed as:

[0076]

[0077] Where X′ represents the target reconstruction result of the CSCGI method, Ψ is the sparse transformation operator, l is the regularization parameter, and ||·||2 is the L2 norm. However, in the above method, the quality of the reconstructed image of the target will be severely degraded due to the presence of scattering media. Therefore, it is very important to design a novel LSRCGAN method to improve the imaging quality after penetrating the scattering medium.

[0078] Step 2: We designed and proposed a deep convolutional conditional generative adversarial network (AMSPI-LSCGAN) learning framework based on a least squares loss model with a compression-excitation block and a residual block. This network can automatically learn feature information from unlabeled data and can pre-process one-dimensional sampled data before inputting it into the network, making the resulting reconstructed image more specific.

[0079] ​​Step 2.1: Generative Adversarial Network (GAN) model: It consists of two competing networks, namely the generator network (G) and the discriminator network (D). The role of G is to generate fake images that are very close to the original image to deceive D; while D attempts to distinguish between real samples and generated samples. Based on the characteristics of the GAN network, the specific network framework structure of the AMSPI-LSCGAN we designed is as follows: Figure 2 shown.

[0080] The generator network (G) is a deep convolutional U-shaped network structure that restores the original image through an encoder consisting of four downsampling convolution modules, a decoder consisting of four upsampling convolution modules, and three residual blocks added between the encoder and decoder. Figure 2 (a) shows the encoder stage. The encoder stage is a downsampling process, with the input being a 32×32 pixel image directly reshaped from the measured signal values ​​of the barrel detector. Convolutional layers extract image features, while max pooling layers reduce the spatial dimensionality of the image. This encodes information in the input image, enabling a better understanding of the image's edges and texture structure. A squeeze-excitation (SE) block is added after each convolutional layer. This SE module reduces redundant information and improves the quality of feature maps, making it suitable for image reconstruction under strong scattering conditions. A residual block is applied between the encoder and decoder to shorten the distance between previous and next layers and reuse target feature information. The decoder stage is an upsampling process, with the output being a 32×32 pixel image. Each upsampling convolutional module undergoes 2×2 bilinear interpolation to restore the spatial resolution and increase the size of the output image. Furthermore, SE blocks are added to the decoder's convolutional modules. In addition, attention gates (AGs) are added to traditional skip connections, enhancing the salient features of skip connections, removing more redundant information, and avoiding the loss of large amounts of spatially precise detail during decoding. Therefore, our method can restore more details and reduce the adverse effects of strong scattering media on image quality degradation. The first four parts of the identification network (D) are mainly composed of convolutional layers, normalization layers, and activation function layers alternately forming convolutional modules to form a convolutional neural network, and finally the output is mapped through a fully connected layer, such as Figure 2As shown in (b). The design purpose of D is to make the image generated by G as close to the original true value as possible, and to continuously update the relevant parameters of G, so that the ability to generate values ​​close to the true value is enhanced. In addition, due to the use of least squares loss, the last layer of D is not activated by the sigmoid function. In each convolution module, the convolution layer is first used to realize feature extraction at different scales; secondly, the normalization layer is added to speed up the feature mapping ability of the network and can also serve as a regularizer; finally, the leaky rectified linear unit (Leaky ReLU) is used instead of the traditional rectified linear unit (ReLU) layer to prevent the problem of gradient disappearance during training. The input of D is the real image and the image generated by G, and the output is a one-dimensional feature vector to realize the discriminant function.

[0081] Step 2.2: Both the encoder and decoder of G use SE blocks, which can automatically obtain the importance of each feature channel, thereby improving useful features and suppressing useless features. It is suitable for image reconstruction problems under scattering conditions. The specific structure is as follows Figure 3 As shown. The SE block mainly consists of two parts: compression and excitation. W and H represent the width and height of the feature map respectively. C represents the number of channels, and the size of the input feature map is W×H×C. The first step is the compression operation, which compresses the input feature map into a 1×1×C vector through global average pooling. This vector has a global receptive field to a certain extent, and the dimension of the output matches the number of input feature channels. The second step is the excitation operation, which consists of two fully connected layers, and the output is a 1×1×C vector. Finally, the weight value of each channel calculated by the SE block is multiplied by the two-dimensional matrix of the channel corresponding to the original feature map to obtain the final output result.

[0082] Step 2.3: In order to extract more detailed information from the image, a residual block is added between the encoder and decoder to replace the traditional convolutional layer in the generated network G. The basic calculation formula is as follows:

[0083]

[0084] In Equation 8, p is the input of the residual block, q is the output of the residual block, and W i represents the parameters of the i-th layer, f(p,W i ) represents the residual mapping, as well as are the partial derivatives of p and q respectively. For the network structure of the residual block, the input features are first passed through two 3×3 convolutional layers to obtain the residual map, and then the input is added to the output through a shortcut connection to complete the feature fusion. The residual block not only reduces the difficulty of training deep networks, but also avoids the network training crash caused by gradient disappearance. The specific structure is as follows Figure 3 shown.

[0085] Step 3: In order to avoid training collapse and obtain higher fidelity reconstructed images, the original loss function is improved and the least square loss is combined with the content loss and the average structural similarity loss. The traditional GAN ​​network loss function is:

[0086]

[0087] Among them, x is the real sample, P x is the actual sample distribution; z is the signal value obtained by the single pixel system, P z It is the generated sample distribution defined by the input generation network G(z), E represents the mathematical expectation, and D represents the discriminant network. But here we improve the traditional loss function, the least squares loss function L LSGAN as follows:

[0088]

[0089] Among them, b is set to 1 to represent real data; a is set to 0 to represent forged data; and c is set to 0 to represent deceiving D. Using the least squares loss function will prevent the gradient of D from decreasing to 0, and the data at the boundary will also receive a penalty proportional to the distance, thereby ensuring that the network obeys more gradient information and improving the stability of training.

[0090] In order to make the reconstructed image closer to the true value, a loss function based on the mean absolute error (MAE) is selected to minimize the pixel-level difference between the real image and the generated image. The objective function is:

[0091] L L1 =E[||xG(z)||1] (11)

[0092] ||·||1 represents the L1 norm. Compared with the L2 norm, the L1 norm can reduce the blur and have better image reconstruction quality.

[0093] Structural Similarity Index (SSIM) is one of the quality assessment methods. SSIM evaluates the similarity between two images by comparing brightness, contrast, and structure. It can be expressed as:

[0094]

[0095] where u x and u G(z) are the average values ​​of the total pixels in the real image and the reconstructed image respectively. x 2 and σ G(z) 2 is the variance between the real image and the reconstructed image. xG(z)is the covariance between the real image and the reconstructed image. C1, C2 and C3 are small positive numbers used to avoid division by 0. It is worth noting that in network training, since some areas of the real image contain a lot of target information, while some areas are blank, it is best to apply SSIM locally. Here, the local statistic u x , σ x and σ xG(z) It can be calculated by using a circularly symmetric Gaussian weighted matrix window of 11×11 units with a standard deviation of 1.5. Then, MSSIM is used to evaluate the quality of the entire image, which can be expressed as.

[0096]

[0097] Among them, U is the real image block, R is the predicted image block, x v represents the real image content of the vth window, G v (z) represents the image content generated at the vth window, and K represents the total number of sliding windows. Therefore, the loss of MSSIM can be expressed as:

[0098]

[0099] Therefore, the proposed joint loss function is as follows:

[0100] L G =λ1L LSGAN +[λ2L MSSIM +(1-λ2)L L1 ]λ3 (15)

[0101] In the proposed network model, λ1 = 1, λ2 = 0.84, and λ3 = 200.

[0102] The simulation experiment of the method of the present invention is given below.

[0103] Step 4: Use the designed network to train the model with sampling rates of 19.14%, 14.06%, 9.77%, and 3.52% under different datasets;

[0104] Step 4.1: The training dataset is from the MNIST (6,000 training images and 100 test images) and EMNIST (5,200 training images and 100 test images) handwriting datasets. The images are resized from 28×28 pixels to 32×32 pixels. Each image is then multiplied by the Hadamard measurement matrix and all pixels are normalized into a column of one-dimensional detection values ​​to simulate the measurements collected by a single-pixel detector. Finally, the network is trained on pairs of signal images, formed by preprocessing and normalizing the one-dimensional detection values, and ground truth images. The network model is optimized at sampling rates of 19.14%, 14.06%, 9.77%, and 3.52%. Note that the reconstructed target image size is 32×32, and the ratio of the number of Hadamard patterns required for measurement to the image size is defined as the sampling rate "V." For example, at a sampling rate of 1, 1024 Hadamard patterns need to be measured. Therefore, the above-mentioned sampling rates of 9.14%, 14.06%, 9.77% and 3.52% require 194, 144, 100 and 36 Hadamard modes, respectively.

[0105] The program is written in Python 3.6, and Pytorch is used to implement our LSRCGAN model on an NVIDIA RTX 3080Ti GPU. The learning rate is set to 0.0002, and the Adam optimizer is used to optimize and update the convolution kernel parameters. The training steps are 600 times.

[0106] Step 4.2: In particular, in terms of test data set preparation, the Hadamard pattern may change due to differences in the concentration of the scattering medium. The increase in the concentration of the scattering medium also leads to stronger absorption and scattering, which further reduces the signal-to-noise ratio (SNR). Assuming that we only consider the intensity change of the Hadamard pattern, different Gaussian white noises of 5dB, 10dB, 15dB and 20dB are added to the original Hadamard pattern. Then, the noisy Hadamard pattern is multiplied by the target image, and all pixels are added as one data to simulate the measurement signal value of the bucket detector to obtain the final test signal image. In this way, we simulate the measurement results under different scattering conditions with different signal-to-noise ratios, and the target image can be reconstructed by inputting the above simulated test signal image into the optimized network.

[0107] Step 5: Input the simulation test set data into the network model that has been trained at the corresponding sampling rate, and a high-quality reconstructed target image can be output. Compared with other methods, it is found that even under strong scattering conditions SNR = 5dB, the LSRCGAN network can reconstruct the target image with better fidelity and perceptual quality at a low sampling rate of 3.52%, such as Figure 3 and 4 As shown;

[0108] Step 6: In addition, to verify the generalization of the AMSPI-LSCGAN method, only the MNIST dataset was used as the training set, and some other patterns that were not in the training set were simulated and reconstructed at sampling rates of 14.06% and 9.77% under scattering conditions of SNR = 20dB and 15dB. These patterns consisted of the two-character pattern "02", the English letters "T" and "b", and the special symbol "three equidistant slits" as test targets. The test results are shown in Figure 6. Figure 8 This demonstrates that even if these test objects are not present in the training dataset, the LSRCGAN method can still effectively learn the correspondence between the compressed sampled data and the original image, and can even reconstruct part of the target image. Therefore, these experiments demonstrate the good generalization of our method, and further expansion and optimization of the network dataset will lead to even better reconstruction results.

[0109] The physical experiment of the method of the present invention is given below.

[0110] Step 1: Design and build a scattering AMSPI system, including He-Ne laser, DMD, data acquisition card, single pixel detector, beam terminator, various lenses and support parts, such as Figure 7 As shown;

[0111] Step 2: Select the number 7 and the letter F on a transmissive card as test targets. After collimation and beam expansion, the laser illuminates a computer-controlled DMD with a preloaded Hadamard measurement matrix to achieve compressed sampling and information modulation of the target. The light intensity signals passing through turbidity concentrations of 57, 86, 115, and 144 NTU are measured to obtain a series of one-dimensional detection signal values ​​at different sampling rates.

[0112] Step 3: Experimental data preprocessing: The one-dimensional detection signal values ​​obtained by the experimental system under different scattering conditions and sampling rates are preprocessed by data averaging and then reshaped into a two-dimensional signal image, which is used as the test set image of the input network;

[0113] Step 4: Network reconstruction: Input the test set image into the AMSPI-LSCGAN network trained in step 4 to reconstruct the test target.

[0114] Step 5: Result analysis: In addition to using the network of the present invention for testing, the test target reconstruction is also performed under different scattering conditions using computational ghost imaging (CGI), compressed sensing computational ghost imaging (CSCGI) and conditional generative adversarial network (Pix2Pix) methods at sampling rates of 19.14% and 3.52%. The comparative imaging results are displayed as an intuitive evaluation and analysis. Furthermore, the statistical average of the two evaluation indicators, structural similarity (SSIM) and peak signal-to-noise ratio (PSNR), are selected as objective indicators to evaluate the reconstruction results. Through the intuitive display of subjective imaging results and objective quantitative evaluation indicators, the reconstruction capability and model robustness of the method of the present invention are discussed and analyzed in detail.

[0115] like Figure 8 As shown in Figure 1, this is the subjective imaging result of this method. Figure 8 It can be clearly seen that the quality of the reconstructed images of the four methods has decreased due to the decrease in sampling rate and the increase in turbidity concentration. Figure 8 As shown in (a) and (c), the restoration effect of the CGI method always has blurring artifacts when the sampling rate V = 19.14%, and it is even difficult to restore the image under strong scattering conditions (115NTU and 144NTU). The CSCGI method is also affected by strong scattering media, and the imaging quality is seriously degraded. However, Pix2Pix and AMSPI-LSCGAN can show better restoration effects than the above methods. Figure 8 As shown in (b) and (d), the differences between the images reconstructed by different methods are more obvious at a lower sampling rate V = 3.52%. The CGI and CSCGI methods have difficulty in obtaining contour information from the images of the number "7" and the letter "F" under different turbidity conditions. The images reconstructed by the Pix2Pix method also have artifacts and even local deformations. When the turbidity is 144NTU, the reconstructed images of the lower half of the number "7" and the upper half of the letter "F" also show obvious distortion, which seriously reduces the image quality. In contrast, our AMSPI-LSCGAN method can effectively resist the interference of scattering media and remove the ringing artifacts caused by undersampled images at low sampling rates, thereby obtaining target images with higher fidelity and perceptual quality.

[0116] like Figure 9 As shown in , it is the objective evaluation index of this method. As the sampling rate decreases and the turbidity concentration increases, the PSNR and SSIM of the reconstructed image will decrease. Figure 9 When the turbidity concentration is 57NTU, the PSNR and SSIM of the reconstructed image number 7 by the AMSPI-LSCGAN method at a sampling rate of 19.14% exceed 16dB and 0.59 respectively, and the PSNR and SSIM of the reconstructed image letter F exceed 27dB and 0.56 respectively. Figure 9(a), (b) and (e), (f). When the turbidity concentration increases further, the PSNR and SSIM of the reconstructed image by AMSPI-LSCGAN still show significant improvement compared with other methods. When the sampling rate is 3.52%, the SSIM of CGI and CSCGI drops rapidly, and the PSNR and SSIM of Pix2Pix also drop significantly with the increase of turbidity concentration. However, the AMSPI-LSCGAN network can better restore the target image in the sampling rate of 3.52% and 144NTU environment, as shown in Figure 2. Figure 9 As shown in (c), (g) and (d), (h). For example, the PSNR and SSIM of the reconstructed image of the number 7 are 12% and 33% higher than those of the Pix2Pix method, respectively; the PSNR and SSIM of the reconstructed image of the letter F are 8% and 41% higher than those of the Pix2Pix method, respectively.

[0117] We conduct generalization physical experiments to reconstruct two-character patterns using only the optimized model trained on the MNIST dataset to demonstrate the generalization of our AMSPI-LSCGAN method. Figure 10 As shown in the figure, when the turbidity concentration is 100 NTU, AMSPI-LSCGAN can still reconstruct the target image with a sampling rate of 19.14%, even if the test target is not present in the training set. As the scattering concentration increases, when the sampling rate is 3.52%, the reconstructed image becomes deformed and the edge contour is lost. New training and optimization datasets can produce higher quality reconstructed images.

[0118] In simulations and physical experiments, we also found that the quality of reconstructed images at high sampling rates decreases rapidly with the deepening of the scattering medium concentration. It is well known that higher sampling rates help reconstruct more object details for single-pixel imaging, but sometimes increasing the sampling rate does not significantly improve the quality of the reconstructed image. Here, different signal-to-noise ratios of previous simulations are introduced into the training process as prior information, so that the neural network can learn to optimize reconstruction at high sampling rates under strong scattering conditions. First, the EMNIST dataset is selected as the real image, and the method of step 4.2 is used to obtain a test signal image with a sampling rate of 19.04% and SNR=20dB and SNR=5dB. Then, the signal image with added noise is paired with the real image. Finally, the new paired image and the original paired image without added noise are combined into 15,600 images for training. The experimental reconstruction results are shown in the figure below. Figure 11As shown in Figure 3, the results of training with noisy images are superior to those of previous training, improving reconstruction quality under strong scattering conditions. Adding different levels of Gaussian white noise to the speckle pattern to simulate different scattering conditions as prior information does not fully reveal the scattering characteristics, but experiments show that a good representation of measurement noise can improve the quality of target reconstruction in strong scattering conditions with AMSPI-LSCGAN.

[0119] The specific embodiments described herein are merely illustrative of the spirit of the present invention. Persons skilled in the art may make various modifications, additions, or substitutions to the described specific embodiments without departing from the spirit of the present invention or exceeding the scope of the appended claims.

Claims

1. A single-pixel imaging method based on least squares conditional generative adversarial network, characterized in that: The steps include: Step 1: Design a single-pixel optical imaging system based on Hadamard transform to obtain the one-dimensional detection signal value of the test target under a series of different scattering conditions and sampling rates; Step 2: The one-dimensional detection signal values ​​obtained under different scattering conditions and sampling rates are pre-processed by data averaging and then reshaped into a two-dimensional signal image as the test set image; Step 3: Construct a deep convolutional conditional generative adversarial network AMSPI-LSCGAN based on a least squares loss function model with a compression-excitation block and a residual block; The specific implementation of step 3 is as follows: AMSPI-LSCGAN includes a generation network G and identification network D , where the network is generated G It is a deep convolutional U-shaped network structure that restores the original image through an encoder consisting of 4 downsampling convolution modules, a decoder consisting of 4 upsampling convolution modules, and 3 residual blocks added between the encoder and decoder; The encoder stage is a downsampling process. The input is an image of a certain size directly reshaped by the measured signal value of the barrel detector. The convolution layer is used to extract image features, and the maximum pooling layer is used to reduce the spatial dimension of the image. The information in the input image is encoded to better understand the edge and texture structure of the image. A squeeze-excitation SE block is added after each convolution layer. The SE block consists of two parts: compression and excitation. W and H are set to represent the width and height of the feature map respectively, C represents the number of channels, and the size of the input feature map is W×H×C. The first step is the compression operation. The input feature map is compressed into a 1×1×C vector through global average pooling. This vector has a global receptive field to some extent, and the output dimension matches the number of input feature channels; the second step is the excitation operation, which consists of two fully connected layers. The output is a 1×1×C vector. Finally, the weight value of each channel calculated by the SE block is multiplied by the two-dimensional matrix of the channel corresponding to the original feature map to obtain the final output result; Step 4: Combine the least squares loss with the content loss and the average structural similarity loss as the loss function of the AMSPI-LSCGAN network for training to avoid training collapse and obtain higher-fidelity reconstructed images. In step 5, the test set image is input into the AMSPI-LSCGAN network trained in step 4 to reconstruct the test target.

2. The single-pixel imaging method based on least squares conditional generative adversarial network according to claim 1, characterized in that: The single-pixel optical imaging system in step 1 includes a He-Ne laser, a digital micromirror device (DMD), a data acquisition card, a single-pixel detector, a beam terminator, various lenses, and supporting components. Its imaging principle is as follows: After the laser light is collimated and expanded, it is modulated on the DMD, which has been pre-loaded with a Hadamard speckle pattern. When a pixel in the speckle pattern is 1, the DMD's micromirror is set to the "on" state. When a pixel in the speckle pattern is 0, the DMD's micromirror is set to the "off" state; the reflected light then passes through the test target in the standard turbid liquid, and then the single-pixel detector synchronously collects and records the total light intensity.

3. The single-pixel imaging method based on least squares conditional generative adversarial network according to claim 1, characterized in that: In step 3, the network training data sets used are from the MNIST and EMNIST handwriting data sets. First, the image size is resized. Then, each image is multiplied by the Hadamard measurement matrix and all pixels are normalized into a column of one-dimensional detection values ​​to simulate the measurement values ​​collected by a single-pixel detector. Finally, the network is trained on image pairs consisting of signal images formed by preprocessing and normalizing the one-dimensional detection values ​​and the real images to obtain optimized network models at different sampling rates.

4. The single-pixel imaging method based on least squares conditional generative adversarial network according to claim 3, characterized in that: The Hadamard measurement matrix is ​​calculated as follows; When the matrix is ​​composed of +1 and -1 elements and satisfies HH T = NZ When , such a matrix is ​​called a Hadamard matrix, H is a N Order square, H T yes H The transpose of Z yes N The unit matrix of order; For any matrix with row and column dimensions of 2 m The Hadamard matrix of is used as the measurement matrix, which can be recursively obtained by the following formula: (1) Among them, 2 m-1 Represents the matrix dimension. The Hadamard matrix is ​​generated by random Hadamard transformation; the matrix H Illumination speckle is obtained by digital micromirror device (DMD) P ( x , y ),( x , y ) represents the image coordinates, since the matrix H It is composed of +1 and -1, and DMD can only be used to debug the measurement matrix of 0 and 1. Therefore, a complementary mode is designed. H , can be obtained by the following formula: (2) in, E represents a matrix whose elements are all 1; H + Retained H The element with +1 in H All elements with -1 in the string are converted to 0; H - Indicates that H The elements that were originally 1 are converted to 0, and the elements that were -1 are replaced by +1; in this way, the DMD device can be modulated using H + and H - Matrix to get the matrix H .

5. The single-pixel imaging method based on least squares conditional generative adversarial network according to claim 1, characterized in that: The calculation formula of the residual block is as follows; (8) In formula 8, p is the input of the residual block, q is the output of the residual block, W i Indicates the i The parameters of the layer are obtained through training, f ( p , W i ) represents the residual mapping, as well as They are p and q For the network structure of the residual block, the input features are first passed through two 3×3 convolutional layers to obtain the residual map, and then the input is added to the output through the shortcut connection to complete the feature fusion.

6. The single-pixel imaging method based on least squares conditional generative adversarial network according to claim 1, characterized in that: The specific loss function in step 4 is as follows; Least squares loss function L LSGAN as follows: (10) in, x It is a real sample. P x is the true sample distribution; z is the signal value obtained by the single pixel system, P z is the input generation network G ( z ) defines the generated sample distribution, E represents mathematical expectation, D represents the discriminant network; b is set to 1, representing real data; a Set to 0, it means forging data; set c to 0, it means deceiving the discriminant network D ; Using the least squares loss function will make D The gradient will not decrease to 0, and the data at the boundary will also receive a penalty proportional to the distance, which ensures that the network obeys more gradient information, thereby improving the stability of training; In order to make the reconstructed image closer to the true value, the loss function based on the mean absolute error (MAE) is selected to minimize the pixel-level difference between the real image and the generated image. The objective function is: (11) || · ||1 represents the L1 norm; Then, the average structural similarity loss is used to evaluate the quality of the entire image, which can be expressed as: (12) in u x and u G(z) are the average values ​​of the total pixels in the real image and the reconstructed image, respectively, σ x 2 and σ G(z) 2 is the variance between the real image and the reconstructed image, σ xG(z) is the covariance of the real image and the reconstructed image; C1, C2 and C3 are constants used to avoid division by 0; u x 、 σ x and σ xG(z) By using a circularly symmetric Gaussian weighted matrix window to calculate, and then using MSSIM to evaluate the quality of the entire image, it can be expressed as: (13) Among them, U is the real image block, R is the predicted image block, x v Indicates the v The actual image content of the window, G v ( z ) indicates the v The image content generated at each window, K represents the total number of windows; therefore, the loss of MSSIM can be expressed as: (14) Therefore, the final joint loss function is as follows: (15) Among them, λ1, λ 2 , λ3 are all constants.

Citation Information

Patent Citations

  • Method for improving single-pixel color imaging performance in scattering environment based on deep learning

    CN112950507A

  • High-sensitivity anti-scattering imaging method based on coding and network

    CN114187375A