A passive non-line-of-sight imaging method and device based on deep learning

Through the deep learning passive non-line-of-sight imaging method, the scattered images projected by self-luminous objects and the imaging neural network with symmetrical structure are utilized to solve the imaging problems caused by obstructions in the imaging scene, achieve high-quality and highly generalized non-line-of-sight imaging, and significantly improve the imaging distance and effect.

CN119155563BActive Publication Date: 2025-09-26TSINGHUA UNIVERSITY
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411284340.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-09-13
Publication Date
2025-09-26
Estimated Expiration
2044-09-13

AI Technical Summary

Technical Problem

Existing passive non-line-of-sight imaging technology has problems with poor imaging quality and generalization. In particular, when there are obstructions in the imaging scene, the imaging problem is pathological and it is difficult to achieve high quality and high generalization.

Method used

A passive non-line-of-sight imaging method based on deep learning is adopted. The scattered image generated by the projected image of the self-luminous object on the intermediate wall is used for imaging through an imaging neural network composed of a pair of symmetrical encoders and decoders, including 9 layers of downsampling and 9 layers of upsampling units, combined with convolutional layers, batch normalization layers and activation functions to realize an end-to-end imaging process.

Benefits of technology

It achieves high-quality imaging without specific scene restrictions, and the imaging distance is increased from the traditional 1-2m to 10m. The imaging effect is excellent, and implicit information can be effectively utilized in complex environments, which improves the practicality and versatility of passive non-line-of-sight imaging.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119155563B_ABST
    Figure CN119155563B_ABST
Patent Text Reader

Abstract

The present invention proposes a passive non-line-of-sight imaging method and device based on deep learning, which belongs to the field of optical non-line-of-sight imaging technology. The method includes: in a passive non-line-of-sight imaging scene, using a camera to collect a scattered image generated by a projection image projected by a self-luminous object on an intermediate wall; inputting the scattered image into a preset imaging neural network, and the imaging neural network outputs an imaging result of the projection image; wherein the imaging neural network is composed of a pair of encoders and decoders with a symmetrical structure; the encoder is composed of 9 layers of downsampling units connected in sequence, and the decoder is composed of 9 layers of upsampling units connected in sequence. The present invention uses deep learning to achieve passive non-line-of-sight imaging, has no special restrictions on imaging scenes, and can increase the imaging distance, thereby greatly improving the versatility of passive non-line-of-sight imaging.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of optical non-line-of-sight imaging, and in particular proposes a passive non-line-of-sight imaging method and device based on deep learning. Background Art

[0002] Non-line-of-sight imaging technology enables optical detection and imaging of areas outside the linear field of view. This technology utilizes relay reflective surfaces in the scene to obtain information, which is then solved using a specific algorithm to achieve imaging around obstacles. This technology expands the optical imaging range from the full field of view to beyond the field of view, and has enormous application potential in areas such as navigation and autonomous driving, disaster relief, and endoscopic medical diagnosis.

[0003] Non-line-of-sight imaging technologies are categorized as active and passive, depending on whether they use active light sources to controllably illuminate hidden scenes. Active methods have drawbacks such as high hardware costs, complex system setup, long imaging acquisition times, high risk, and poor concealment. Passive methods can effectively circumvent these drawbacks, but due to the limited information acquired during the detection phase, the imaging calculation process is highly ill-posed, resulting in poor image quality and generalizability, and often requiring or restricting specific imaging scenarios. Summary of the Invention

[0004] The present invention aims to overcome the shortcomings of existing technologies by proposing a method and apparatus for passive non-line-of-sight imaging based on deep learning. This invention utilizes deep learning to implement passive non-line-of-sight imaging, without any specific restrictions on the imaging scene. It can also increase the imaging distance, thereby significantly enhancing the versatility of passive non-line-of-sight imaging.

[0005] A first embodiment of the present invention provides a passive non-line-of-sight imaging method based on deep learning, comprising:

[0006] In a passive non-line-of-sight imaging scenario, a camera is used to capture the scattered image generated by the projection image of the self-luminous object on the intermediate wall.

[0007] The scattered image is input into a preset imaging neural network, and the imaging neural network outputs the imaging result of the projection image; wherein the imaging neural network is composed of a pair of encoders and decoders with a symmetrical structure; the encoder is composed of 9 layers of downsampling units connected in sequence, and the decoder is composed of 9 layers of upsampling units connected in sequence.

[0008] In a specific embodiment of the present invention, the projected image is a square binary image with a pixel number less than or equal to the maximum pixel number of the self-luminous object.

[0009] In a specific embodiment of the present invention, before inputting the scattering image into a preset imaging neural network, the method further includes:

[0010] The scatter image is sampled to an input size set by the imaging neural network.

[0011] In a specific embodiment of the present invention, each downsampling unit of the encoder includes a two-dimensional convolution layer for achieving downsampling, a batch normalization layer for controlling the stability of internal parameters of the model, and an activation function for introducing nonlinearity; each upsampling unit of the decoder includes a two-dimensional deconvolution layer for achieving upsampling, a batch normalization layer for controlling the stability of internal parameters of the model, a jump connection for enhancing feature fusion, and an activation function for introducing nonlinearity.

[0012] In a specific embodiment of the present invention, it also includes:

[0013] The imaging neural network initially inputs a scattering image of size 512×512, and after passing through the first downsampling unit, outputs an image of size 256×256 and number of features 4, which enters the second downsampling unit; after passing through the second downsampling unit, outputs an image of size 128×128 and number of features 4, which enters the third downsampling unit; after passing through the third downsampling unit, outputs an image of size 64×64 and number of features 8, which enters the fourth downsampling unit; after passing through the fourth downsampling unit, outputs an image of size 32×32 and number of features 16, which enters the fifth downsampling unit; after passing through the fifth downsampling unit, outputs an image of size 16×16 and number of features 32, which enters the sixth downsampling unit; after passing through the sixth downsampling unit, outputs an image of size 8 ×8, the feature number of 32 enters the seventh downsampling unit; after the seventh downsampling unit, the output image of size 4×4, the feature number of 32 enters the eighth downsampling unit; after the eighth downsampling unit, the output image of size 2×2, the feature number of 32 enters the ninth downsampling unit; after the ninth downsampling unit, the output image of size 1×1, the feature number of 32 enters the first upsampling unit; after the first upsampling unit, the output image of size 2×2, the feature number of 32 is jumped connected with the image output by the eighth downsampling unit, and is spliced ​​along the feature dimension into an image of size 2×2, the feature number of 64 enters the second upsampling unit; after the second upsampling unit, the output image of size 4×4, the feature number of 32 The image is skipped and connected with the output image of the seventh down-sampling unit, and is spliced ​​along the feature dimension into an image of size 4×4 and feature number 64 to enter the third up-sampling unit; after the third up-sampling unit, an image of size 8×8 and feature number 32 is output, and the image is skipped and connected with the output image of the sixth down-sampling unit, and is spliced ​​along the feature dimension into an image of size 8×8 and feature number 64 to enter the fourth up-sampling unit; after the fourth up-sampling unit, an image of size 16×16 and feature number 32 is output, and the image is skipped and connected with the output image of the fifth down-sampling unit, and is spliced ​​along the feature dimension into an image of size 16×16 and feature number 64 to enter the fifth up-sampling unit; after the fifth up-sampling unit, an image of size 32 is output. ×32, an image with a feature number of 16, this image is skip-connected with the output image of the fourth down-sampling unit, and is spliced ​​along the feature dimension into an image with a size of 32×32 and a feature number of 32 to enter the sixth up-sampling unit; after the sixth up-sampling unit, an image with a size of 64×64 and a feature number of 8 is output, this image is skip-connected with the output image of the third down-sampling unit, and is spliced ​​along the feature dimension into an image with a size of 64×64 and a feature number of 16 to enter the seventh up-sampling unit; after the seventh up-sampling unit, an image with a size of 128×128 and a feature number of 4 is output, this image is skip-connected with the output image of the second down-sampling unit, and is spliced ​​along the feature dimension into an image with a size of 128×128 and a feature number of 8 to enter the eighth up-sampling unit;After passing through the eighth upsampling unit, the output image is a 256×256 image with 4 features. This image is jump-connected to the output image of the first downsampling unit, and spliced ​​along the feature dimension to form an image of 256×256 with 8 features. This image then enters the ninth upsampling unit. After passing through the ninth upsampling unit, the output image is a 512×512 image with 4 features. This image enters the readout layer of the imaging neural network. The readout layer outputs an image of 512×512 with 1 feature, which is the final imaging result.

[0014] In a specific embodiment of the present invention, it also includes:

[0015] Each downsampling unit includes a convolution layer with a convolution kernel size of 4×4 and a sliding step size of 2 and a LeakyRELU activation function; wherein the second downsampling unit to the eighth downsampling unit also include a batch normalization layer;

[0016] Each upsampling unit layer includes a deconvolution layer with a convolution kernel size of 4×4 and a sliding step size of 2, a batch normalization layer, and a RELU activation function; wherein, the first upsampling unit to the third upsampling unit also use dropout with a probability of 0.5;

[0017] The readout layer includes a convolution layer with a convolution kernel size of 1×1 and a sigmoid function.

[0018] In a specific embodiment of the present invention, before inputting the scattering image into a preset imaging neural network, the method further includes:

[0019] training the imaging neural network;

[0020] The training of the imaging neural network comprises:

[0021] 1) In the passive non-line-of-sight imaging scenario, the self-luminous object is used to project a plurality of true-value images, and a scattered image generated by each true-value image on the intermediate wall is acquired by the camera; wherein the plurality of true-value images have the same characteristics;

[0022] 2) Sampling each ground truth image and corresponding scatter image to a size set by the imaging neural network to form a training sample, wherein the input of the training sample is the resized scatter image and the ground truth label is the corresponding resized ground truth image; all training samples constitute a training set;

[0023] 3) constructing the imaging neural network;

[0024] 4) Using the training set obtained in step 2) to train the imaging neural network to obtain the trained imaging neural network.

[0025] A second embodiment of the present invention provides a passive non-line-of-sight imaging device based on deep learning, comprising:

[0026] A scattered image acquisition module is used to use a camera to acquire a scattered image generated by a projection image projected by a self-luminous object on an intermediate wall in a passive non-line-of-sight imaging scenario;

[0027] An imaging module is configured to input the scattered image into a preset imaging neural network, which outputs an imaging result of the projection image; wherein the imaging neural network is composed of a pair of encoders and decoders with a symmetrical structure; the encoder is composed of nine layers of downsampling units connected in sequence, and the decoder is composed of nine layers of upsampling units connected in sequence.

[0028] A third embodiment of the present invention provides an electronic device, including:

[0029] at least one processor; and a memory communicatively coupled to the at least one processor;

[0030] The memory stores instructions that can be executed by the at least one processor, and the instructions are configured to execute the above-mentioned passive non-line-of-sight imaging method based on deep learning.

[0031] A fourth aspect of the present invention provides a computer-readable storage medium storing computer instructions for enabling the computer to execute the above-mentioned deep learning-based passive non-line-of-sight imaging method.

[0032] Features and beneficial effects of the present invention:

[0033] The present invention does not impose any special restrictions on the inaccessible hidden space (from the hidden target to the relay wall) in the non-line-of-sight imaging task. It only uses a data-driven method to reduce the ill-posedness of the problem, greatly improving the practicality and versatility of passive non-line-of-sight imaging. In addition, the present invention increases the distance from the camera to the relay wall from 1-2m in the traditional method to 10m, further expanding the application scenarios of non-line-of-sight imaging. Under high-precision timing control, the present invention can realize high-speed automated data acquisition, greatly facilitating the acquisition of a large amount of experimental data required for neural network training; at the same time, after the network training is completed, real-time high-speed imaging can be achieved at the same speed, and the imaging effect is excellent.

[0034] The main advantages of the present invention are as follows:

[0035] Traditional passive non-line-of-sight imaging methods introduce occlusions between the hidden target and the intervening wall to reduce the pathological nature of the imaging problem. However, high-quality imaging requires occlusions to be as complex as possible, while high generalization requires occlusions to be as simple, arbitrary, or even non-existent as possible. This creates an unavoidable trade-off between the two. In contrast, the present invention utilizes deep learning methods to effectively circumvent this problem.

[0036] The imaging method using deep learning has no additional restrictions on the non-field of view area, and has no specific requirements for the placement distance and angle of the screen (hidden target), whether there is occlusion or scattering during the transmission process, whether the fixed background light is uniform, whether the observation relay wall is flat, etc. The present invention can fully extract and utilize these implicit environmental information. In fact, the more complex the environment, the more information can be used, the less pathological the problem is, and the more conducive to imaging. Since the hidden area behind the relay wall is often unknown or difficult to enter in the actual application of non-field of view detection, it is not reasonable to limit the area to meet specific requirements. The present invention does not impose restrictions on it, which is more in line with the real scene and can effectively improve the application scope and practical value of the technology.

[0037] The present invention adopts an end-to-end neural network that includes additional scene prior information, and there is no need to explicitly solve and express this information. The prior effect can be directly utilized without specific knowledge of the prior. Compared with traditional methods, it is more direct and efficient, and can avoid errors caused by approximate modeling. BRIEF DESCRIPTION OF THE DRAWINGS

[0038] Figure 1 This is an overall flow chart of a deep learning-based passive non-line-of-sight imaging method according to an embodiment of the present invention;

[0039] Figure 2 is a schematic diagram of a passive non-line-of-sight imaging scenario according to a specific embodiment of the present invention;

[0040] Figure 3 is a schematic diagram of the imaging neural network structure of a specific embodiment of the present invention;

[0041] Figure 4 This is a diagram showing the training effect of an imaging neural network according to a specific embodiment of the present invention;

[0042] Figure 5 This is a graph of imaging results of an imaging neural network test according to a specific embodiment of the present invention. DETAILED DESCRIPTION

[0043] The present invention proposes a passive non-line-of-sight imaging method and device based on deep learning, which is further described in detail below with reference to the accompanying drawings and specific embodiments.

[0044] A first embodiment of the present invention provides a passive non-line-of-sight imaging method based on deep learning, comprising:

[0045] In a passive non-line-of-sight imaging scenario, a camera is used to capture the scattered image generated by the projection image of the self-luminous object on the intermediate wall.

[0046] The scattered image is input into a preset imaging neural network, and the imaging neural network outputs the imaging result of the projection image; wherein the imaging neural network is composed of a pair of encoders and decoders with a symmetrical structure; the encoder is composed of 9 layers of downsampling units connected in sequence, and the decoder is composed of 9 layers of upsampling units connected in sequence.

[0047] In a specific implementation of the present invention, the passive non-line-of-sight imaging method based on deep learning has the following overall process: Figure 1 As shown, it is divided into training phase and testing phase, including the following steps:

[0048] 1) Training phase.

[0049] 1-1) Build a passive non-line-of-sight imaging scene.

[0050] In a specific embodiment of the present invention, the passive non-line-of-sight imaging scene schematic diagram is as follows: Figure 2As shown. In this scenario, the imaging target is an image displayed by a self-luminous object. The self-luminous object is approximately two-dimensional and can be any object that meets the brightness requirements. A screen is generally used. In this embodiment, an LED screen is selected, but an LCD screen that meets the brightness requirements can also be used. In this embodiment, the self-luminous object is a monochrome, high-brightness LED screen with an average brightness greater than 2000 nits and a spacing of more than 5 mm between the LED beads on the screen. The camera uses a CMOS industrial camera with a pixel count greater than 300,000 and an output accuracy of 16 bits. The parameters of the imaging lens used with the camera are determined according to the specific imaging scenario, so that the scattered image to be collected on the intervening wall can completely fill the camera frame. The reflectivity of the monochromatic light emitted by the intervening wall to the LED screen is greater than 20%. In this imaging scenario, the distance between the LED screen and the intervening wall is less than or equal to 60 cm, and the distance between the intervening wall and the camera is less than or equal to 10 m. There is obstruction between the camera and the LED screen, preventing the camera from directly observing the image displayed on the LED screen, but it can observe the scattered image transmitted to the intervening wall. Specifically, in this embodiment, the LED screen measures 488mm × 244mm, has 64 × 32 pixels, and a single pixel size of 7.62mm × 7.62mm. Its average brightness is approximately 2500 nits. The screen emits monochromatic green light with a central wavelength of 516nm and a spectral linewidth of approximately 33nm. The camera used is a PCO panda4.2bi mono sCMOS scientific camera with a maximum resolution of 2048 × 2048, a pixel size of 6.5μm × 6.5μm, a photosensitive surface size of 13.3mm × 13.3mm, a dynamic range of 87dB, and a maximum quantum efficiency greater than 80%. The imaging lens is a 100mm fixed-focus lens from Zhonglian Science and Technology. The intermediate wall is a whiteboard with rough-textured wallpaper to simulate a common brick wall. The distance between the LED screen and the intermediate wall is approximately 60cm, and the distance between the intermediate wall and the camera is approximately 10m.

[0051] 1-2) Use the scene built in step 1-1) to build a training set.

[0052] In this embodiment, different true value images are displayed on the self-luminous object, and then the camera is used to capture the corresponding scattered image on the intermediate wall. Among them, the true value image should be a square binary image with a pixel number less than or equal to the maximum number of pixels of the self-luminous object. In order to cooperate with the imaging network structure, the input and output of the network are both 512×512 size images. Therefore, when constructing the training set, it is necessary to re-interpolate and sample the true value image displayed on the self-luminous object and the scattered image captured by the camera to a size of 512×512, and then use the interpolated scattered image as input and the corresponding interpolated true value image as the true value label to form a set of training samples. This embodiment uses no less than 5000 groups of training samples to form a training set. The images displayed by the self-luminous objects in the training set have common features (for example, they are all handwritten numbers, handwritten letters, human body movements, etc.). In the process of constructing the training set, the imaging scene is kept fixed and free of interference.

[0053] In one specific embodiment of the present invention, a training set was constructed using the publicly available MNIST dataset of handwritten digits, whose original images were sized 28×28 pixels, as images displayed on an LED screen. In this embodiment, the original images were preprocessed and converted into 512×512 binary images before being input to the LED screen for display. A corresponding 800×800 scattered image was then captured using a camera in a constructed passive non-line-of-sight imaging scene. The image was then preprocessed and converted into 512×512 pixels. This embodiment used a total of 5,500 different images of handwritten digits, resulting in a training set consisting of 5,500 training samples.

[0054] 1-3) Construct an imaging neural network.

[0055] In this embodiment, the imaging neural network is based on the classic convolutional neural network Unet structure and consists of a pair of symmetrical encoders and decoders; wherein the encoder and decoder are respectively a contraction path and an expansion path, and the encoder consists of 9 layers of downsampling units connected in sequence, each of which contains a two-dimensional convolution layer for downsampling, a batch normalization layer for controlling the stability of the model's internal parameters, and an activation function for introducing nonlinearity; the encoder and decoder are directly connected, and the encoder output is the decoder input; the decoder consists of 9 layers of upsampling units connected in sequence, each of which contains a two-dimensional deconvolution layer for upsampling, a batch normalization layer for controlling the stability of the model's internal parameters, a jump connection for enhancing feature fusion, and an activation function for introducing nonlinearity. The encoder can gradually reduce the spatial dimension of the input image to achieve feature extraction, and the decoder gradually restores the features to the original image dimension output.

[0056] In this embodiment, the overall structure of the imaging neural network is as follows: Figure 3As shown in Figure 2, the imaging neural network initially inputs a scattering image of size 512×512. After passing through downsampling unit 1 (including a convolution layer with a convolution kernel size of 4×4, a sliding step of 2, and a LeakyRELU activation function), the output image of size 256×256 and feature number 4 enters downsampling unit 2; after passing through downsampling unit 2 (including a convolution layer with a convolution kernel size of 4×4, a sliding step of 2, a batch normalization layer, and a LeakyRELU activation function), the output image of size 128×128 and feature number 4 enters downsampling unit 3; after passing through downsampling unit 3 (including a convolution layer with a convolution kernel size of 4×4, a sliding step of 2, a batch normalization layer, and a LeakyRELU activation function), the output image of size 128×128 and feature number 4 enters downsampling unit 3. ELU activation function), the output size is 64×64 and the number of features is 8, which enters the downsampling unit 4; after the downsampling unit 4 (including the convolution layer with a convolution kernel size of 4×4 and a sliding step of 2, the batch normalization layer and the LeakyRELU activation function), the output size is 32×32 and the number of features is 16, which enters the downsampling unit 5; after the downsampling unit 5 (including the convolution layer with a convolution kernel size of 4×4 and a sliding step of 2, the batch normalization layer and the LeakyRELU activation function), the output size is 16×16 and the number of features is 32, which enters the downsampling unit 6; after the downsampling unit 6 (including the convolution layer with a convolution kernel size of 4×4 and a sliding step of 2, the batch normalization layer and the LeakyRELU activation function), the output size is 16×16 and the number of features is 32, which enters the downsampling unit 6. Normalization layer and LeakyRELU activation function), the output size is 8×8 and the number of features is 32, which enters the downsampling unit 7; after the downsampling unit 7 (including the convolution layer with a convolution kernel size of 4×4 and a sliding step of 2, the batch normalization layer and the LeakyRELU activation function), the output size is 4×4 and the number of features is 32, which enters the downsampling unit 8; after the downsampling unit 8 (including the convolution layer with a convolution kernel size of 4×4 and a sliding step of 2, the batch normalization layer and the LeakyRELU activation function), the output size is 2×2 and the number of features is 32, which enters the downsampling unit 9; after the downsampling unit 9 (including the convolution layer with a convolution kernel size of 4×4 and a sliding step of 2 After the convolution layer and RELU activation function, the output image of size 1×1 and number of features is sent to upsampling unit 1; after the upsampling unit 1 (including the deconvolution layer with convolution kernel size of 4×4 and sliding step size of 2, batch normalization layer, RELU activation function, and dropout with probability 0.5), the output image of size 2×2 and number of features is sent to upsampling unit 2; after the upsampling unit 2 (including the deconvolution layer with convolution kernel size of 4×4 and sliding step size of 2, batch normalization layer, RELU activation function, and dropout with probability 0.5), the output image of size 2×2 and number of features is sent to upsampling unit 2.5 dropout) after the output size of the image is 4 × 4, the number of features is 32, the image and the down-sampling unit 7 output image jump connection, along the feature dimension splicing into a size of 4 × 4, the number of features of 64 images into the up-sampling unit 3; after the up-sampling unit 3 (including the convolution kernel size of 4 × 4, the sliding step size of 2 deconvolution layer, batch normalization layer, RELU activation function, and the use of probability 0.5 dropout), the output size of the image is 8 × 8, the number of features is 32, the image and the down-sampling unit 6 output image jump connection, along the feature dimension splicing into a size of 8 × 8, the number of features of 64 images into the up-sampling unit 4; after the up-sampling unit 4 (including the convolution kernel size of 4 × 4, the sliding step size of 2 deconvolution layer, batch normalization layer, RELU activation function, and the use of probability 0.5 dropout), the output size of the image is 8 × 8, the number of features is 32, the image and the down-sampling unit 6 output image jump connection, along the feature dimension splicing into a size of 8 × 8, the number of features of 64 images into the up-sampling unit 4; After passing through the upsampling unit 5 (including the deconvolution layer with a convolution kernel size of 4×4 and a sliding step size of 2, the batch normalization layer and the RELU activation function), the output image is 16×16 in size and 32 in feature number. This image is skip-connected with the output image of the downsampling unit 5, and is spliced ​​along the feature dimension into an image of 16×16 in size and 64 in feature number, and enters the upsampling unit 5; after passing through the upsampling unit 5 (including the deconvolution layer with a convolution kernel size of 4×4 and a sliding step size of 2, the batch normalization layer and the RELU activation function), the output image is 32×32 in size and 16 in feature number. This image is skip-connected with the output image of the downsampling unit 4, and is spliced ​​along the feature dimension into an image of 32×32 in size and 32 in feature number, and enters the upsampling unit 6; after passing through the upsampling unit 6 (including the deconvolution layer with a convolution kernel size of 4×4 and a sliding step size of After passing through the upsampling unit 7 (including a deconvolution layer with a convolution kernel size of 4×4 and a sliding step size of 2, a batch normalization layer and a RELU activation function), the output image is 64×64 in size and 8 in feature number. This image is skip-connected with the output image of the downsampling unit 3, and is spliced ​​along the feature dimension into an image with a size of 64×64 and 16 in feature number. The image then enters the upsampling unit 7. After passing through the upsampling unit 7 (including a deconvolution layer with a convolution kernel size of 4×4 and a sliding step size of 2, a batch normalization layer and a RELU activation function), the output image is 128×128 in size and 4 in feature number. This image is skip-connected with the output image of the downsampling unit 2, and is spliced ​​along the feature dimension into an image with a size of 128×128 and 8 in feature number. The image then enters the upsampling unit 8. After passing through the upsampling unit 8 (including a deconvolution layer with a convolution kernel size of 4×4 and a sliding step size of 2 After a deconvolution layer, batch normalization layer, and RELU activation function, the output is a 256×256 image with 4 features. This image is skip-connected to the output image of downsampling unit 1, and concatenated along the feature dimension to form a 256×256 image with 8 features, which enters upsampling unit 9. After upsampling unit 9 (which includes a deconvolution layer with a convolution kernel size of 4×4 and a sliding step of 2, a batch normalization layer, and a RELU activation function), the output is a 512×512 image with 4 features, which enters the readout layer. The readout layer (which includes a convolution layer with a convolution kernel size of 1×1 and a sigmoid function) ultimately outputs a 512×512 image with 1 feature, which is the final imaging result output by the imaging neural network.

[0057] It should be noted that the imaging neural network of this embodiment has been optimized in several ways to address the characteristics of the passive non-line-of-sight imaging problem. The main innovations are as follows:

[0058] a) The depth of the traditional U-net network is increased, using nine layers of downsampling units in the encoder and nine layers of upsampling units in the decoder, increasing model complexity and parameter space. Based on the physical characteristics of the imaging problem, the input and output dimensions and number of features of each layer in the neural network are adjusted, while ensuring high imaging quality and fast training speed.

[0059] b) Unlike the traditional U-net network that uses a maximum pooling layer for downsampling, the imaging neural network in the present invention directly achieves downsampling through a convolutional layer with a sliding step size of 2, using a convolutional layer with learnable parameters for downsampling to achieve simultaneous feature extraction and translation.

[0060] c) LeakyRELU is used as the activation function in the encoder instead of the RELU function to improve gradient flow and make it suitable for training deeper networks while maintaining its advantage of simple gradient operations. The decoder still uses the RELU activation function, intentionally introducing a certain degree of neuron death to keep the expansion path network sparse, thereby preventing overfitting and achieving faster computation speed and generalization.

[0061] d) Before each convolutional network layer's output passes through the activation function, batch normalization is performed. This operation first translates and scales the network output to a normalized form with a mean of 0 and a variance of 1. Two learnable parameters are then introduced to perform a small translation and scaling on the normalized output. This operation improves the stability of the neural network training process, helps reduce the network's sensitivity to initialization, and increases the permissible range of the learning rate, thereby accelerating model convergence and improving training speed.

[0062] e) Use Dropout regularization in some upsampling units. Dropout is a regularization method used to prevent overfitting in neural networks. By randomly dropping neurons with a certain probability during training, these neurons are excluded from the current calculation and parameter update, thereby reducing the network's dependence on specific neurons. This can improve the generalization ability of the network in that layer and make the training process more robust.

[0063] 1-4) Using the training set obtained in step 1-2), the imaging neural network established in step 1-3) is trained.

[0064] In this embodiment, the Adam optimizer is used to train the imaging neural network. The parameters of the Adam optimizer are β1 = 0.9, β2 = 0.99, and the weight decay rate is 2×10-5 , the mean square error MSE is selected as the loss function in training, and the initial learning rate is 2.5×10 -4 , the training cycle is not less than 100, and the imaging quality improves with the increase of training cycle. After the training is completed, the final imaging neural network is obtained.

[0065] Figure 4 This is a diagram showing the effect of imaging neural network training according to a specific embodiment of the present invention. Figure 4 The peak signal-to-noise ratio (PSNR) between the output of the imaging neural network and the target true value varies with the number of training cycles. After training, the highest PSNR of the imaging neural network test set reaches over 18dB, indicating that the imaging result in this embodiment is very close to the true value image.

[0066] 2)Testing phase.

[0067] 2-1) Obtain a test image.

[0068] In this embodiment, the test image and the true value image of the training set have the same features (handwritten numbers, handwritten letters, human body movements, etc.).

[0069] In a specific embodiment of the present invention, the test image is an MNIST handwritten digit image with a size of 28×28. 2-2) In the imaging scene constructed in step 1), the LED screen displays the test image of step 2-1), and light is transmitted to the intermediate wall to form a corresponding original scattered image, which is captured by the camera.

[0070] The collected original scattering image is sampled to the input size set by the imaging neural network, which is 512×512 in this embodiment, to obtain the final scattering image.

[0071] 2-3) The final scattered image obtained in step 2-2) is input into the trained imaging neural network, and the network outputs an imaging result corresponding to the scattered image, thereby realizing the passive non-line-of-sight imaging process.

[0072] In a specific embodiment of the present invention, the output of the imaging neural network is compared and analyzed with the true value image displayed on the LED screen. The imaging result of this embodiment is as follows: Figure 5 As shown, Figure 5 (a) and Figure 5In (b), the left column shows scattered images directly captured by the camera from the intermediate wall, the middle column shows the image output after the scattered images are input into the imaging neural network, and the right column shows the true image displayed on the LED screen. This shows that the method described in this embodiment can restore scattered images to true images with high quality, achieving the goal of passive non-line-of-sight imaging. Specific tests have shown that the two-dimensional images produced using the method described in this embodiment have a resolution of less than 1 cm, an imaging speed greater than 5 frames per second, a peak signal-to-noise ratio greater than 15 dB, and a recognition accuracy exceeding 90%, resulting in excellent imaging results.

[0073] To implement the above embodiment, a second embodiment of the present invention provides a deep learning-based passive non-line-of-sight imaging device, comprising:

[0074] A scattered image acquisition module is used to use a camera to acquire a scattered image generated by a projection image projected by a self-luminous object on an intermediate wall in a passive non-line-of-sight imaging scenario;

[0075] An imaging module is configured to input the scattered image into a preset imaging neural network, which outputs an imaging result of the projection image; wherein the imaging neural network is composed of a pair of encoders and decoders with a symmetrical structure; the encoder is composed of nine layers of downsampling units connected in sequence, and the decoder is composed of nine layers of upsampling units connected in sequence.

[0076] It should be noted that the aforementioned explanation of an embodiment of a passive non-line-of-sight imaging method based on deep learning is also applicable to a passive non-line-of-sight imaging device based on deep learning in this embodiment, and will not be repeated here. According to an embodiment of the present invention, a passive non-line-of-sight imaging device based on deep learning is proposed. In a passive non-line-of-sight imaging scene, a camera is used to capture a scattered image generated by a projection image projected by a self-luminous object on an intermediate wall surface; the scattered image is input into a preset imaging neural network, and the imaging neural network outputs an imaging result of the projection image; wherein the imaging neural network is composed of a pair of encoders and decoders with a symmetrical structure; the encoder is composed of 9 layers of downsampling units connected in sequence, and the decoder is composed of 9 layers of upsampling units connected in sequence.

[0077] This makes it possible to use deep learning to achieve passive non-line-of-sight imaging, without any special restrictions on the imaging scene, and to increase the imaging distance, thereby greatly improving the versatility of passive non-line-of-sight imaging.

[0078] To implement the above embodiment, a third aspect of the present invention provides an electronic device, including:

[0079] at least one processor; and a memory communicatively coupled to the at least one processor;

[0080] The memory stores instructions that can be executed by the at least one processor, and the instructions are configured to execute the above-mentioned passive non-line-of-sight imaging method based on deep learning.

[0081] To implement the above-mentioned embodiment, the fourth aspect of the present invention proposes a computer-readable storage medium, which stores computer instructions, and the computer instructions are used to enable the computer to execute the above-mentioned passive non-line-of-sight imaging method based on deep learning.

[0082] It should be noted that the computer-readable medium mentioned above in the present disclosure may be a computer-readable signal medium or a computer-readable storage medium, or any combination of the two. A computer-readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, device, or component, or any combination of the above. More specific examples of computer-readable storage media may include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present disclosure, a computer-readable storage medium may be any tangible medium that contains or stores a program that can be used by or in conjunction with an instruction execution system, device, or component. In the present disclosure, a computer-readable signal medium may include a data signal propagated in baseband or as part of a carrier wave, which carries computer-readable program code. Such a propagated data signal may take a variety of forms, including but not limited to an electromagnetic signal, an optical signal, or any suitable combination of the above. A computer-readable signal medium may also be any computer-readable medium other than a computer-readable storage medium that can transmit, propagate, or transport a program for use by or in conjunction with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium may be transmitted using any suitable medium, including but not limited to wires, optical cables, RF (radio frequency), etc., or any suitable combination thereof.

[0083] The computer-readable medium may be included in the electronic device or may exist independently without being incorporated into the electronic device. The computer-readable medium carries one or more programs that, when executed by the electronic device, cause the electronic device to perform the deep learning-based passive non-line-of-sight imaging method described in the above embodiment.

[0084] Computer program code for performing the operations of the present disclosure may be written in one or more programming languages, or a combination thereof, including object-oriented programming languages ​​such as Java, Smalltalk, C++, and conventional procedural programming languages ​​such as "C" or similar programming languages. The program code may be executed entirely on the user's computer, partially on the user's computer, as a stand-alone software package, partially on the user's computer and partially on a remote computer, or entirely on the remote computer or server. In cases involving a remote computer, the remote computer may be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., through the Internet using an Internet service provider).

[0085] In the description of this specification, the description with reference to the terms "one embodiment", "some embodiments", "example", "specific example", or "some examples" means that the specific features, structures, materials or characteristics described in conjunction with the embodiment or example are included in at least one embodiment or example of the present application. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials or characteristics described can be combined in any one or more embodiments or examples in a suitable manner. In addition, those skilled in the art can combine and combine different embodiments or examples described in this specification and features of different embodiments or examples without contradiction.

[0086] Furthermore, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of the technical features being referred to. Thus, a feature defined as "first" or "second" may explicitly or implicitly include at least one of such features. Throughout the description of this application, "plurality" means at least two, for example, two, three, etc., unless otherwise specifically defined.

[0087] Any process or method description in a flowchart or otherwise described herein may be understood to represent a module, segment or portion of code comprising one or more executable instructions for implementing the steps of a specific logical function or process, and the scope of the preferred embodiments of the present application includes alternative implementations in which functions may be performed out of the order shown or discussed, including performing functions in a substantially simultaneous manner or in the reverse order depending on the functions involved, which should be understood by those skilled in the art to which the embodiments of the present application belong.

[0088] The logic and / or steps represented in the flowcharts or otherwise described herein, for example, can be considered as a sequenced list of executable instructions for implementing the logical functions, and can be embodied in any computer-readable medium for use by, or in conjunction with, an instruction execution system, apparatus, or device (e.g., a computer-based system, a system including a processor, or other system that can fetch and execute instructions from an instruction execution system, apparatus, or device). For purposes of this specification, a "computer-readable medium" can be any device that can contain, store, communicate, propagate, or transport a program for use by, or in conjunction with, an instruction execution system, apparatus, or device. More specific examples (a non-exhaustive list) of computer-readable media include the following: an electrical connection with one or more wires (electronic devices), a portable computer disk cartridge (magnetic device), random access memory (RAM), read-only memory (ROM), erasable and programmable read-only memory (EPROM or flash memory), fiber optic devices, and a portable compact disc read-only memory (CDROM). Furthermore, the computer-readable medium may even be paper or other suitable medium on which the program is printed, since the program may be obtained electronically, for example, by optically scanning the paper or other medium and then editing, interpreting or otherwise processing it in a suitable manner if necessary, and then storing it in a computer memory.

[0089] It should be understood that various parts of the present application can be implemented using hardware, software, firmware, or a combination thereof. In the above embodiments, multiple steps or methods can be implemented using software or firmware stored in a memory and executed by a suitable instruction execution system. For example, if implemented using hardware, as in another embodiment, any one of the following technologies known in the art or a combination thereof can be used to implement: a discrete logic circuit having a logic gate circuit for implementing a logic function on a data signal, an application-specific integrated circuit having a suitable combination of logic gate circuits, a programmable gate array (PGA), a field programmable gate array (FPGA), etc.

[0090] Those skilled in the art will understand that all or part of the steps in the method of the above embodiment can be completed by instructing related hardware through a program, and the program can be stored in a computer-readable storage medium. When the program is executed, it includes one or a combination of the steps of the method embodiment.

[0091] In addition, the functional units in the various embodiments of the present application may be integrated into a processing module, or each unit may exist physically separately, or two or more units may be integrated into a module. The above-mentioned integrated module may be implemented in the form of hardware or in the form of a software functional module. If the integrated module is implemented in the form of a software functional module and sold or used as an independent product, it may also be stored in a computer-readable storage medium.

[0092] The storage medium mentioned above may be a read-only memory, a magnetic disk, or an optical disk, etc. Although the embodiments of the present application have been shown and described above, it is understood that the above embodiments are exemplary and should not be construed as limiting the present application. Persons skilled in the art may make changes, modifications, substitutions, and variations to the above embodiments within the scope of the present application.

Claims

1. A passive non-line-of-sight imaging method based on deep learning, characterized in that: include: In a passive non-line-of-sight imaging scenario, a camera is used to capture the scattered image generated by the projection image of the self-luminous object on the intermediate wall. The scattering image is input into a preset imaging neural network, and the imaging neural network outputs the imaging result of the projection image; wherein the imaging neural network is composed of a pair of encoders and decoders with a symmetrical structure; the encoder is composed of 9 layers of downsampling units connected in sequence, and the decoder is composed of 9 layers of upsampling units connected in sequence; the imaging neural network initially inputs a scattering image of size 512×512, and after passing through the first downsampling unit, the output image of size 256×256 and number of features 4 enters the second downsampling unit; after passing through the second downsampling unit, the output image of size 128×128 and number of features 4 enters the third downsampling unit; after passing through the third downsampling unit, the output image of size 64×64 and number of features 8 enters the fourth downsampling unit; after passing through the fourth downsampling unit, the output image of size 64×64 and number of features enters the fourth downsampling unit; After the unit, the output image with a size of 32×32 and a feature number of 16 enters the fifth downsampling unit; after passing through the fifth downsampling unit, the output image with a size of 16×16 and a feature number of 32 enters the sixth downsampling unit; after passing through the sixth downsampling unit, the output image with a size of 8×8 and a feature number of 32 enters the seventh downsampling unit; after passing through the seventh downsampling unit, the output image with a size of 4×4 and a feature number of 32 enters the eighth downsampling unit; after passing through the eighth downsampling unit, the output image with a size of 2×2 and a feature number of 32 enters the ninth downsampling unit; after passing through the ninth downsampling unit, the output image with a size of 1×1 and a feature number of 32 enters the first upsampling unit; each downsampling unit contains a LeakyRELU activation function; each upsampling unit contains a RELU activation function.

2. The method according to claim 1, characterized in that The projected image is a square binary image with a pixel number less than or equal to the maximum pixel number of the self-luminous object.

3. The method according to claim 1, characterized in that Before inputting the scattering image into a preset imaging neural network, the method further includes: The scatter image is sampled to an input size set by the imaging neural network.

4. The method according to claim 1, wherein Each downsampling unit of the encoder includes a two-dimensional convolution layer for achieving downsampling, a batch normalization layer for controlling the stability of internal parameters of the model, and an activation function for introducing nonlinearity; each upsampling unit of the decoder includes a two-dimensional deconvolution layer for achieving upsampling, a batch normalization layer for controlling the stability of internal parameters of the model, a jump connection for enhancing feature fusion, and an activation function for introducing nonlinearity.

5. The method according to claim 4, characterized in that Also includes: After passing through the first upsampling unit, the imaging neural network outputs an image with a size of 2×2 and a feature number of 32. This image is jump-connected with the image output by the eighth downsampling unit, and is spliced ​​along the feature dimension into an image with a size of 2×2 and a feature number of 64, which enters the second upsampling unit; after passing through the second upsampling unit, the output is an image with a size of 4×4 and a feature number of 32. This image is jump-connected with the image output by the seventh downsampling unit, and is spliced ​​along the feature dimension into an image with a size of 4×4 and a feature number of 64, which enters the third upsampling unit; after passing through the third upsampling unit, the output is an image with a size of 8×8 and a feature number of 32. This image is jump-connected with the image output by the sixth downsampling unit, and is spliced ​​along the feature dimension into an image with a size of 8×8 and a feature number of 64, which enters the fourth upsampling unit; after passing through the fourth upsampling unit, the output is an image with a size of 16×16 and a feature number of 32. This image is jump-connected with the image output by the fifth downsampling unit, and is spliced ​​along the feature dimension into an image with a size of 16×16 and a feature number of 64, which enters the fifth upsampling unit; After the fifth upsampling unit, an image with a size of 32×32 and a feature number of 16 is output. This image is jump-connected with the output image of the fourth downsampling unit, and is spliced ​​along the feature dimension into an image with a size of 32×32 and a feature number of 32 to enter the sixth upsampling unit; after the sixth upsampling unit, an image with a size of 64×64 and a feature number of 8 is output. This image is jump-connected with the output image of the third downsampling unit, and is spliced ​​along the feature dimension into an image with a size of 64×64 and a feature number of 16 to enter the seventh upsampling unit; after the seventh upsampling unit, an image with a size of 128×128 and a feature number of 4 is output. This image is jump-connected with the output image of the second downsampling unit, and is spliced ​​along the feature dimension into an image with a size of 64×64 and a feature number of 16 to enter the seventh upsampling unit. The output image is jump-connected and spliced ​​along the feature dimension into an image with a size of 128×128 and a feature number of 8, which enters the eighth upsampling unit; after passing through the eighth upsampling unit, an image with a size of 256×256 and a feature number of 4 is output, and this image is jump-connected with the image output from the first downsampling unit, and spliced ​​along the feature dimension into an image with a size of 256×256 and a feature number of 8, which enters the ninth upsampling unit; after passing through the ninth upsampling unit, an image with a size of 512×512 and a feature number of 4 is output and enters the readout layer of the imaging neural network; the readout layer outputs an image with a size of 512×512 and a feature number of 1, which is the final imaging result.

6. The method according to claim 5, characterized in that Also includes: Each downsampling unit includes a convolution layer with a convolution kernel size of 4×4 and a sliding step size of 2; wherein the second downsampling unit to the eighth downsampling unit also include a batch normalization layer; Each upsampling unit includes a deconvolution layer with a convolution kernel size of 4×4 and a sliding step size of 2, and a batch normalization layer; wherein, the first upsampling unit to the third upsampling unit also use dropout with a probability of 0.5; The readout layer includes a convolution layer with a convolution kernel size of 1×1 and a sigmoid function.

7. The method according to claim 6, characterized in that Before inputting the scattering image into a preset imaging neural network, the method further includes: training the imaging neural network; The training of the imaging neural network comprises: 1) In the passive non-line-of-sight imaging scenario, the self-luminous object is used to project a plurality of true-value images, and a scattered image generated by each true-value image on the intermediate wall is obtained by the camera; wherein the plurality of true-value images have the same characteristics; 2) Sampling each ground-truth image and corresponding scattering image to a size set by the imaging neural network to form a training sample, wherein the input of the training sample is the resized scattering image and the ground-truth label is the corresponding resized ground-truth image; all training samples constitute a training set; 3) constructing the imaging neural network; 4) Using the training set obtained in step 2) to train the imaging neural network, to obtain the trained imaging neural network.

8. A passive non-line-of-sight imaging device based on deep learning, characterized in that: include: A scattered image acquisition module is used to use a camera to acquire a scattered image generated by a projection image projected by a self-luminous object on an intermediate wall in a passive non-line-of-sight imaging scenario; An imaging module is provided, which is used to input the scattered image into a preset imaging neural network, and the imaging neural network outputs the imaging result of the projection image; wherein the imaging neural network is composed of a pair of encoders and decoders with a symmetrical structure; the encoder is composed of 9 layers of downsampling units connected in sequence, and the decoder is composed of 9 layers of upsampling units connected in sequence; the scattered image is input into a preset imaging neural network, and the imaging neural network outputs the imaging result of the projection image; wherein the imaging neural network is composed of a pair of encoders and decoders with a symmetrical structure; the encoder is composed of 9 layers of downsampling units connected in sequence, and the decoder is composed of 9 layers of upsampling units connected in sequence; the imaging neural network initially inputs a scattered image of size 512×512, and after passing through the first downsampling unit, it outputs an image of size 256×256 and feature number 4 and enters the second downsampling unit; after passing through the second downsampling unit, it outputs an image of size 128×128 and feature number The image with a size of 4 enters the third downsampling unit; after passing through the third downsampling unit, the image with a size of 64×64 and a feature number of 8 enters the fourth downsampling unit; after passing through the fourth downsampling unit, the image with a size of 32×32 and a feature number of 16 enters the fifth downsampling unit; after passing through the fifth downsampling unit, the image with a size of 16×16 and a feature number of 32 enters the sixth downsampling unit; after passing through the sixth downsampling unit, the image with a size of 8×8 and a feature number of 32 enters the seventh downsampling unit; after passing through the seventh downsampling unit, the image with a size of 4×4 and a feature number of 32 enters the eighth downsampling unit; after passing through the eighth downsampling unit, the image with a size of 2×2 and a feature number of 32 enters the ninth downsampling unit; after passing through the ninth downsampling unit, the image with a size of 1×1 and a feature number of 32 enters the first upsampling unit; each downsampling unit contains a LeakyRELU activation function; each upsampling unit contains a RELU activation function.

9. An electronic device, characterized in that: include: at least one processor; and, a memory communicatively coupled to the at least one processor; The memory stores instructions that can be executed by the at least one processor, and the instructions are configured to execute the method according to any one of claims 1 to 7.

10. A computer-readable storage medium, characterized in that The computer-readable storage medium stores computer instructions, and the computer instructions are used to enable the computer to execute the method according to any one of claims 1 to 7.

Citation Information

Patent Citations

  • Passive non-vision field image recognition method and recognition device based on deep learning

    CN113837217A

  • Non-vision field imaging method based on untrained deep decoding neural network

    CN114494480A