Rapid double-reflection non-vision-field imaging method based on important light screening
By screening important rays and sampling difference rays and building a neural implicit model, the problems of large reconstruction errors and low efficiency in non-field of view imaging are solved, and fast and accurate reconstruction of hidden targets is achieved, which is suitable for static and dynamic scenes.
Patent Information
- Application Number
- CN202511129614.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-13
- Publication Date
- 2025-09-12
- Estimated Expiration
- 2045-08-13
AI Technical Summary
When reconstructing hidden targets, existing non-line-of-sight imaging technology is greatly affected by the accuracy of projection point positions and the clarity of shadow boundaries, resulting in large reconstruction errors. In addition, the number of rays increases exponentially with resolution, affecting efficiency. Implicit neural modeling makes it difficult to extract feature points and shadow trajectories.
By adjusting the projection position of the laser point to form a multi-perspective shadow image, screening important light, building a neural implicit model, combining multi-resolution sampling and difference light sampling, optimizing the NeRF model, and achieving fast and accurate hidden target reconstruction.
Fast 3D reconstruction of static and dynamically occluded spaces is achieved, with a relative depth deviation of 2% and an absolute deviation of 0.2 meters, which improves reconstruction accuracy and efficiency and reduces computational costs.
Smart Images

Figure CN120635335A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of non-line-of-sight imaging, and in particular to a fast double-reflection non-line-of-sight imaging method based on important light screening. Background Art
[0002] Traditional detection methods adhere to the imaging principles of geometric optics, limiting photon capture to the line of sight. Non-line-of-sight (NLoS) imaging, which involves sensing targets beyond the line of sight, has become a significant research area. The basic concept of NLoS imaging involves modeling and analyzing the optical transmission process and target structure using indirectly transmitted photons in scenarios where the optical path between the target and the detector is blocked.
[0003] The current mainstream solutions include the coherent information method based on coherent light sources, the photon time-of-flight method using ultrafast photon detectors, and the two-dimensional intensity information method using shadows. The coherent information method is affected by the scattering effect of the correlated light on the intermediate surface. The effective coherent information is insufficient to reconstruct the hidden scene, making the implementation of high-quality surround imaging more complicated. The photon time-of-flight method uses the captured photon flight time information to model the three-reflection process of the photon and is currently a relatively mature surround imaging solution. However, due to the high cost of light sources and photon detectors, many researchers are limited to simulation environments and public datasets, which hinders the expansion of these methods in a wider range of research and application fields. The two-dimensional intensity information method uses traditional cameras to record the light intensity information in the scene. Its effectiveness depends on the lighting quality of the target environment and is difficult to adapt to daily tasks in complex environments.
[0004] In 2020, Connor Henley et al. proposed a method using two intermediary surfaces (Henley, C., et al. Imaging behind occluders using two-bounce light. in Computer Vision–ECCV 2020: 16th European Conference, Glasgow, UK, August 23–28, 2020, Proceedings, Part XXIX 16. 2020. Springer. ), defining it as a distributed scheme: one surface serves as the light receiving surface, and the other serves as the shadow imaging surface. By continuously adjusting the projection angle of the laser point light source, the shadow and light source position are used to carve the structure of the hidden scene. Schematic diagrams of integrated and distributed non-line-of-sight schemes are shown below. Figure 1As shown in Figure 2 , in this configuration, the active light source and detector only need to modulate and respond to light intensity. Compared to integrated approaches that use scattered coherent light information to image a scene, this design significantly reduces reliance on light source and sensor performance. In practical applications, distributed solutions allow for flexible selection of reflective and imaging surfaces based on different tasks, offering greater potential in terms of observation field of view, device size, and cost.
[0005] However, the quality of the engraving depends on the accuracy of the projected point positions and the clarity of the shadow boundaries. Deviations from these two factors can lead to reconstruction errors, and the laser spot projected onto the intermediate surface inevitably produces local scattering, introducing errors into the engraving process. Secondly, the accuracy of the reconstruction depends on the number and fineness of the rays, but the number of rays increases exponentially with the resolution of the shadow binary image and the number of projections, which significantly affects the efficiency of the reconstruction of the hidden scene.
[0006] In 2021, the paper (Mildenhall, B., et al., Nerf: Representing scenes as neural radiance fields for view synthesis. Communications of the ACM, 2021.65(1): p. 99-106.) proposed Neural Radiance Fields (NeRF), a deep learning model for implicit 3D spatial modeling. Compared to previous algorithms, NeRF only requires multi-view images, corresponding intrinsic and extrinsic parameters, and pose, without requiring an explicit 3D reconstruction process. It can construct an implicit model of the current scene for rendering and 3D reconstruction from new perspectives. The idea of implicit modeling has also been applied to non-line-of-sight imaging. In 2021, the paper (Shen, S., et al., Non-line-of-sight imaging via neural transient fields. IEEE Transactions on Pattern Analysis and Machine Intelligence, 2021. 43(7): p. 2257-2268.) proposed the concept of neural transient fields (NeTF). In a transient surround vision system based on photon time-of-flight, the captured photon transient signals were used to construct a neural implicit representation model, significantly improving reconstruction performance. In 2022, the paper (Mu, F., et al., Physics to the rescue: Deep non-line-of-sight reconstruction for high-speed imaging. IEEE Transactions on Pattern Analysis and Machine Intelligence, 2022.) used conditional radiance fields in transient reconstruction, introducing light transport and volume rendering physical priors into neural networks to reconstruct fine details. Currently, there has been no application of implicit neural modeling to surround shadow reconstruction. The key issues are the difficulty in extracting feature points for matching in the shadows and the difficulty in accurately obtaining the shadow trajectory. Summary of the Invention
[0007] In order to overcome the shortcomings of the existing technology, the present invention provides a fast double-reflection non-line-of-sight imaging method based on important light screening. By adjusting the projection position of the laser point to form a multi-perspective shadow image, the important light that forms the shadow of the hidden target is screened out from the shadow image, and a neural implicit model is constructed to achieve fast and accurate reconstruction of the hidden target.
[0008] The technical solution for achieving the purpose of the present invention is: a fast double-reflection non-line-of-sight imaging method based on important light screening, comprising:
[0009] Step 1: Calibrate the sampling system based on the NLOS scenario to determine the size of the space to be reconstructed;
[0010] Step 2: Use the galvanometer system to change the laser light path to illuminate the hidden space, use the camera to obtain the shadow image produced by the active illumination, and process the shadow image based on the calibration information to obtain the pose and shadow binary image in 3D coordinates;
[0011] Step 3: Perform light modeling based on the double-reflection NLOS scene and build a 3D reconstruction network for the shadow scene;
[0012] Step 4: Sample the shadow binary image obtained in step 2 using a multi-resolution sampling method based on edge intensity;
[0013] Step 5: Input the sampled shadow binary image and the pose in step 2 into the reconstruction network in step 3 for training;
[0014] Step 6: Repeat step 2, use the difference ray sampling method to dynamically sample the shadow binary image obtained in step 2, and repeat step 5;
[0015] Step 7: Repeat step 6 to achieve 3D reconstruction of the NLOS dynamic scene.
[0016] Furthermore, step 1 is specifically as follows:
[0017] The sampling system is calibrated based on the NLOS scenario; the center point of the relay wall is , the center point of the projection wall is ; to connect and The midpoint of the line segment is the origin , establish a three-dimensional space coordinate system; the distance between the two walls is , the width of the relay wall is W and the height is .
[0018] Furthermore, step 2 is specifically as follows:
[0019] In order to realize the reconstruction of hidden objects, a double-reflection non-line-of-sight imaging system is constructed, which consists of a camera and a laser scanning galvanometer system. The incident direction of the laser is changed by the laser scanning galvanometer system, and the positions on the relay wall are respectively 、 ,..., Use the camera to capture the changing shadows on the projection wall, thereby obtaining a set of shadow images that can reflect the characteristics of the hidden object;
[0020] The data processing process is as follows:
[0021] (1) Perform perspective transformation on the light spot under oblique viewing angle to obtain the light spot under normal viewing angle; let the coordinates of the four corner points on the original image be ,in ;
[0022] The coordinates of the four corresponding points on the target plane are ,in:
[0023] ;
[0024] and are the pixel width and height of the sampled image respectively;
[0025] The homogeneous coordinate representation of the perspective transformation is:
[0026] ;
[0027] For each pair of corresponding points , ,in is the coordinate of the light spot under the oblique angle before perspective transformation, is the coordinate of the light spot under the positive viewing angle after perspective transformation; solve the various value: 、 ,..., ;By performing perspective transformation on all pixels in the captured image, the corrected spot position is obtained;
[0028] (2) Traverse the grayscale shadow image For all pixels in The corresponding pixel coordinates ( , ); Set the pixel coordinates ( , ) is converted to three-dimensional coordinates ( ), the conversion formula is as follows:
[0029] ;
[0030] Assuming that there are N light spots corresponding to N shadow images, the corresponding light spot coordinates are expressed as ;
[0031] (3) Perform the same perspective transformation operation on the shadow image to obtain the corrected shadow image; perform histogram equalization on the shadow image and use an adaptive threshold to perform a binarization operation.
[0032] Furthermore, step 3 is specifically as follows:
[0033] Based on the principle of neural radiation fields, a multi-layer perceptron is used to implicitly model light and hidden objects. Light rays are constructed based on the light spot position and shadow image, and 3D coordinates are sampled along each ray. An MLP is used to predict the probability of the existence of an object at a 3D coordinate. The 3D structure is reconstructed by accumulating the existence probabilities, forming a self-supervised learning framework for implicitly modeling the hidden space.
[0034] The probability of the existence of hidden objects is calculated using the cumulative transmittance It describes the light in the interval The probability of light propagating without being blocked; The expected cumulative transmittance is expressed as ,in Indicates the distance from the surface of the object to the relay wall, Indicates the spot position, t indicates the light propagation time; near the boundary and far boundaries In the range between To estimate and calculate the cumulative transmittance :
[0035] ;
[0036] In the imaging task, enter the spot position , which can query the three-dimensional coordinates in the hidden space The cumulative transmittance at ; Continuous Implicit Neural Shadow Field Expressed as:
[0037] ;
[0038] From the spot position Construct the light corresponding to the shadow image and trace the light to the projection wall; use the layered sampling method of discrete samples to estimate the cumulative transmittance The integral result of:
[0039] ;
[0040] in, Indicates adjacent moments The distance between samples, represents the cumulative transmittance, Indicates opacity;
[0041] Perform binary segmentation on the shadow image to obtain the distribution of light passing through the hidden object; mark the pixels in the bright area as "1", indicating that the light has passed through the hidden scene; mark the pixels in the dark area as "0", indicating that the light is blocked; the black background indicates that the light is not blocked when passing through the unknown space, and the white foreground indicates that the light is blocked by the object, thus forming a shadow on the wall; supervise the optimization process of the network by calculating the cumulative transmittance; the loss function of the cumulative transmittance Expressed as:
[0042] ;
[0043] in, is the set of rays in each batch during training, indicating the rays involved in training;
[0044] Finally, the depth of the hidden object is calculated from multiple perspectives by numerically integrating N rays. :
[0045] .
[0046] Furthermore, step 4 is specifically as follows:
[0047] There are N light spots on the relay wall ( , , ,..., ), the light is emitted from these N light spots to the projection wall; the resolution of the collected shadow binary image is ; Emit light from the light spot on the relay wall to the projection wall ,in , , so that we get a ray of light;
[0048] The rays containing structural information are defined as important rays. A multi-resolution sampling method based on edge intensity is proposed, which uses different downsampling factors for different regions and divides the regions into sparse sampling regions and dense sampling regions according to edge intensity.
[0049] For a size of The shadow binary image is divided into rectangular regions, each with a height of , with a width of ; Detect the edges within each rectangular area and calculate its edge strength; use the Sobel operator to calculate the average edge strength of all pixels in each area , and set a threshold , in order to decide whether to perform focused sampling in this area;
[0050] when When , the region is classified as a densely sampled region, using the first downsampling factor To preserve more important light, ;when When , the region is classified as a sparsely sampled region and the second downsampling factor is used Downsample the sparsely sampled area, ;
[0051] In this way, the sampling density is adjusted in different regions; the light from each region that has undergone different downsampling processes is merged and input into the MLP for training; the latent space is reconstructed using light of different resolutions, and the original resolution of the structure is preserved.
[0052] Furthermore, step 6 is specifically as follows:
[0053] In a dual-reflection non-line-of-sight system, the positions of the light spot and shadow are changed by a galvanometer to acquire shadow images and position data. The relay wall is scanned repeatedly to obtain temporally continuous dynamic data. In the dynamic reconstruction task, galvanometer scanning and data acquisition operations are continuously performed, and the newly sampled shadows and coordinates are simultaneously input into the MLP to achieve rapid reconstruction.
[0054] A differential sampling method is used to reconstruct dynamic scenes. The changing light is extracted from consecutive frames captured at the same spot position to model the dynamics of hidden objects. Based on the complete reconstruction of the first frame, subsequent frames only reconstruct the dynamic part.
[0055] Adopt image pyramid strategy to supplement global feature information;
[0056] The shadow binary image is downsampled at different scales. The downsampled images are stacked and fused to create a multi-scale shadow map and a multi-scale ray set. At the same time, the light rays that change between consecutive frames are extracted and stored in a differential ray set. Finally, the resulting multi-scale ray set and differential ray set are mixed and fed into the MLP for training.
[0057] By training the network with multi-scale information including changing and overlapping light in old and new frames, the network can quickly reconstruct dynamic scenes.
[0058] An electronic device comprises a memory, a processor and a computer program stored in the memory and executable on the processor, wherein the steps of the above method are implemented when the processor executes the program.
[0059] A computer-readable storage medium stores a computer program, which implements the steps of the above method when executed by a processor.
[0060] A computer program product comprises a computer program, which implements the steps of the above method when executed by a processor.
[0061] Compared with the prior art, the present invention has the following significant advantages:
[0062] (1) The present invention performs spatial and temporal decomposition on multi-view shadow images to extract important rays that mainly shape the hidden scene, and optimizes and reduces the NeRF model with a large number of rays with repeated structures to achieve accurate and efficient reconstruction of hidden objects; the present invention demonstrates the ability to quickly reconstruct three-dimensional space in dynamic and static occlusion, achieving a scene relative depth deviation of 2%, and The absolute deviation at scene scale is 0.2 meters.
[0063] (2) The present invention constructs a new dual-reflection non-line-of-sight imaging system, which uses a laser galvanometer acquisition system to continuously scan the scene and update the sampled scene data. The reconstruction system uses the latest light sampling method to obtain the difference light between two consecutive samples, and uses the difference light to update the existing reconstruction model, thereby realizing the function of reconstructing NLOS dynamic scenes. BRIEF DESCRIPTION OF THE DRAWINGS
[0064] Figure 1 It is an integrated and distributed non-line-of-sight solution.
[0065] Figure 2 It is a dual-reflection non-line-of-sight (NLOS) imaging system, where (a) is the system composition, (b) is the light sampling and reconstruction, and (c) is the reconstruction framework.
[0066] Figure 3 A diagram showing the visualization of scene parameters.
[0067] Figure 4 Schematic diagram of data processing.
[0068] Figure 5 A diagram visualizing light sampling.
[0069] Figure 6 Schematic diagram of the multi-resolution sampling method based on feature intensity partitioning.
[0070] Figure 7 Schematic diagram of the training method combining multi-resolution and frame difference.
[0071] Figure 8 Comparison of full sampling and multi-resolution sampling methods at different shadow resolutions. Gray text annotation: Root mean square error (RMSE) value of overall depth.
[0072] Figure 9Comparison of full sampling and multi-resolution sampling methods under different scene complexities. Gray text annotation: Root mean square error (RMSE) value of overall depth.
[0073] Figure 10 Comparison of full sampling method and differential sampling method in dynamic scenes. Gray text annotation: root mean square error (RMSE) value of overall depth.
[0074] Figure 11 Comparison of initialization between full sampling and multi-resolution sampling methods.
[0075] Figure 12 Comparison of the efficiency of full sampling and multi-resolution sampling methods at different resolutions.
[0076] Figure 13 Comparison of full sampling and differential sampling methods in multi-target and background scenes.
[0077] Figure 14 Comparison of reconstruction accuracy between full sampling and multi-resolution sampling methods in real static scenes.
[0078] Figure 15 Comparison of full sampling and differential sampling methods in real static scenes. DETAILED DESCRIPTION
[0079] The technical solution of the present invention is described in detail below with reference to the accompanying drawings and embodiments.
[0080] Non-line-of-sight imaging uses a wall to reflect laser light and a wall to receive shadows, and constructs a ray carving model to achieve three-dimensional reconstruction of the hidden scene. However, the level of detail in depicting the hidden target relies on a large number of multi-view rays, which causes the computational delay to increase exponentially. The present invention proposes a fast implicit ray carving method that uses a hierarchical neural radiation field (NeRF) model to reconstruct hidden objects. Specifically, the multi-view shadow image is spatially and temporally decomposed to extract important rays that mainly shape the hidden scene, and the NeRF model that optimizes and reduces a large number of rays with repeated structures is optimized to achieve accurate and efficient reconstruction of hidden objects. The method of the present invention demonstrates the ability to quickly reconstruct dynamic and static occluded spaces, achieving a scene relative depth deviation of 2% (in The absolute deviation at scene scale is 0.2 meters).
[0081] First, we will introduce the dual-reflection non-line-of-sight (NLOS) imaging technology. When an unknown target is obscured by obstacles in its line of sight, the detector cannot directly obtain its structural information. The two sides of the hidden object (the relay wall and the projection wall) are Lambertian walls. Assuming a fixed illumination angle, the surface brightness of these two walls is the same when viewed from all directions. A laser beam is emitted from outside the space and reflects off the relay wall, producing a light spot and scattered light. The scattered light travels through the hidden space, reaches the projection wall, and casts a shadow. To simplify the process, higher-order reflected light is ignored. The shadow reflects the structure of the hidden space. By varying the position of the light spot, the scattered light can be tracked to reconstruct the structure of the hidden object.
[0082] The present invention performs system calibration based on scenarios. Figure 3 As shown, the center point of the relay wall is , the center point of the projection wall is . To connect and The midpoint of the line segment is the origin , establish a three-dimensional space coordinate system. In the system of the present invention, the distance between the two walls is , the width of the relay wall is W and the height is .
[0083] In order to realize the reconstruction of hidden objects, the present invention constructs a new double-reflection non-line-of-sight imaging system, such as Figure 2 As shown in Figure (a) in the figure. The system mainly consists of a camera and a laser scanning galvanometer system. The galvanometer changes the incident direction of the laser, and the positions on the relay wall are 、 ,..., The camera is used to capture the changing shadows on the projection wall, thereby obtaining a set of shadow images that can reflect the characteristics of the hidden object.
[0084] What makes the new system different from the previous dual-reflection non-line-of-sight imaging system is that it has added a dynamic reconstruction function. The previous imaging system only scanned once and statically modeled the scene. The new system has a dynamic reconstruction function. Specifically, the data acquisition system repeatedly scans the scene and inputs the data sampled each time into the reconstruction system. The reconstruction system uses the latest light sampling method to obtain the difference light between two consecutive samples, and uses the difference light to update the existing reconstructed scene, thereby achieving the function of reconstructing a dynamic scene.
[0085] Data processing process such as Figure 4 As shown:
[0086] (1) Perform perspective transformation on the light spot under oblique viewing angle to obtain the light spot under normal viewing angle. Let the coordinates of the four corner points on the original image be ,in .
[0087] The coordinates of the four corresponding points on the target plane are ,in:
[0088] ;
[0089] and are the pixel width and height of the sampled image respectively.
[0090] The homogeneous coordinate representation of the perspective transformation is:
[0091] ;
[0092] For each pair of corresponding points ( , ), the various elements in the perspective transformation matrix can be solved according to the above transformation formula By performing perspective transformation on all pixels in the captured image, the corrected spot position can be obtained.
[0093] (2) Traverse the grayscale shadow image For all pixels in , find the maximum grayscale value (i.e. the center of the light spot) The corresponding pixel coordinates ( , Then, the pixel coordinates ( , ) is converted to three-dimensional coordinates ( ), the conversion formula is as follows:
[0094] ;
[0095] Assuming that there are N light spots corresponding to N shadow images, the corresponding light spot coordinates can be expressed as .
[0096] (3) Perform the same perspective transformation operation on the shadow image to obtain the corrected shadow image. Perform histogram equalization on the shadow image and use an adaptive threshold to perform a binarization operation.
[0097] The following is a detailed description of the light modeling process. Based on the principle of Neural Radiance Field (NeRF), a multi-layer perceptron (MLP) is used to implicitly model light and hidden objects. Figure 2 This is shown in Figure (c) of [1]. Rays are constructed based on the spot position and shadow image, and 3D coordinates are sampled along each ray. The MLP is used to predict the probability of an object at that 3D coordinate (i.e., the probability of the object existing in the latent space). The shadow image is then reconstructed by accumulating these probabilities, forming a self-supervised learning framework that implicitly models the latent space.
[0098] The probability of the existence of hidden objects is expressed as cumulative transmittance It describes the light in the interval The probability that light will propagate without being blocked. The expected cumulative transmittance is expressed as ,in Indicates the distance from the surface of the object to the relay wall, Indicates the spot position, t indicates the light propagation time. and far boundaries In the range between To estimate and calculate the cumulative transmittance :
[0099] ;
[0100] In the imaging task, enter the spot position , you can query the three-dimensional coordinates in the hidden space The cumulative transmittance at ,like Figure 2 This continuous implicit neural shadow field It can be expressed as:
[0101] ;
[0102] From the spot position Construct the rays corresponding to the shadow image and trace the rays that reach the projection wall. Use a stratified sampling method with discrete samples to estimate the cumulative transmittance. The integral result of:
[0103] ;
[0104] in, Indicates adjacent moments The distance between samples, represents the cumulative transmittance, Indicates opacity.
[0105] Perform binary segmentation on the shadow image to obtain the distribution of light passing through the hidden object. Mark the pixels in the bright area as "1", indicating that the light has passed through the hidden scene; mark the pixels in the dark area as "0", indicating that the light is blocked. The black background indicates that the light is not blocked when passing through the unknown space, and the white foreground indicates that the light is blocked by the object, thus forming a shadow on the wall. The optimization process of the network is supervised by calculating the cumulative transmittance. The loss function of the cumulative transmittance is: It can be expressed as:
[0106]
[0107] in, is the set of rays in each batch during training, representing the rays participating in training.
[0108] Finally, the depth of the hidden object is calculated from multiple perspectives by numerically integrating N rays. :
[0109]
[0110] The following describes the optimized light sampling scheme of the present invention. The reconstruction method of the present invention is based on light modeling. The light generation process is as follows: Figure 5 As shown. There are N light spots on the relay wall ( , , ,..., ). The light is emitted from these N spots to the projection wall. The resolution of the collected shadow binary image is . Emit light from the light spot on the relay wall to the projection wall ,in , , so that we get A ray of light.
[0111] The full sampling method is to obtain The order of the rays is randomly shuffled and then fed into the network in batches for training. During training, each ray is treated as equally important. After multiple ray shuffling and training cycles, the structure of the hidden scene can be recovered. However, this method relies on a large number of repeated training iterations, resulting in low reconstruction efficiency.
[0112] At the beginning of training, due to the lack of effective model initialization prior information, the randomly input rays cannot guarantee the uniformity of structural information, which can easily cause the model to fall into local extremes and lead to model collapse. In addition, the edges of shadow binary images usually contain critical structural information, while high-resolution shadow binary images contain a large number of flat pixels (i.e., invalid rays). The full sampling method uniformly samples and models each ray, causing the network to focus too much on invalid rays during training, resulting in limited accuracy when reconstructing the target (especially the details). In dynamic scenes, sampling all rays frame by frame within a scanning cycle and training from scratch is very inefficient. The differences between consecutive frames should be considered so that the model can prioritize rays with dynamic structural information.
[0113] Rays containing structural information (such as edges and outlines) are defined as important rays. To increase the proportion of important rays in the overall training set, this paper proposes a multi-resolution sampling method based on edge strength. This method not only improves training efficiency but also solves the problem of initialization corruption. The multi-resolution sampling method applies different downsampling factors to different regions and divides regions into sparsely sampled and densely sampled regions based on edge strength.
[0114] Blank bright areas in a shadow binary image indicate that light is not blocked by an object, and the light from these areas provides less structural information for reconstruction. Therefore, a larger downsampling factor is used in these areas. Inner shadow areas indicate that light is completely blocked by an object and similarly provide less structural information, so a larger downsampling factor is also used. For edge regions, focused sampling or full sampling is used to improve the model's reconstruction accuracy for edges and details. Since blank areas contribute only limitedly to the reconstruction of hidden objects, a larger downsampling factor can reduce invalid light and redundant computation.
[0115] For a size of The shadow binary image is divided into rectangular regions, each with a height of , with a width of ,like Figure 6 As shown. Detect the edges in each rectangular area and calculate its edge strength. Use the Sobel operator to calculate the average edge strength of all pixels in each area. , and set a threshold , in order to decide whether to perform focus sampling in this area.
[0116] when When , the region is classified as a densely sampled region, using the first downsampling factor To preserve more important light, ;when When , the area is classified as a sparsely sampled area, and only a small number of rays are needed to indicate whether the area is occupied. The second downsampling factor is used Downsample the sparsely sampled area, , reducing the number of rays, thus achieving efficient optimization.
[0117] This approach allows for flexible adjustment of sampling density across different regions, ensuring sufficient capture of important rays while effectively reducing redundant computation. Finally, the rays from each region, having undergone different downsampling processes, are combined and fed into the MLP for training. Reconstructing the latent space using rays of varying resolution preserves the original resolution of the structure, achieving high-resolution reconstruction at a low computational cost.
[0118] In a dual-reflection non-line-of-sight system, the position of the light spot and shadow is changed by a galvanometer to obtain shadow images and position data. The relay wall is scanned repeatedly to obtain temporally continuous dynamic data. Due to the high-speed scanning of the galvanometer, the data acquisition speed is much higher than the movement speed of the hidden object, so the dynamic changes of the object within one scanning cycle can be ignored. In the dynamic reconstruction task, the galvanometer scanning and data acquisition operations are continuously performed, and the newly sampled shadows and coordinates are input into the MLP to achieve fast reconstruction. The full sampling method requires a large number of training iterations to reconstruct a frame of image in one scanning cycle, which is too slow for dynamic reconstruction.
[0119] This paper designs a differential sampling method to reconstruct dynamic scenes. Taking into account the small differences between consecutive frames, the method extracts the changing light from consecutive frames captured at the same spot position to model the dynamics of hidden objects. Based on the complete reconstruction of the first frame, subsequent frames reconstruct only the dynamic portion, achieving rapid reconstruction.
[0120] Because neural networks suffer from catastrophic forgetting, they often significantly forget previously learned information when encountering new tasks during continuous learning. To prevent the neural network from forgetting the static parts of the latent space when encountering new frames, an image pyramid strategy is used to supplement global feature information.
[0121] like Figure 7 As shown in the figure, from the perspective of the same light spot, the frame difference in a dynamic scene is mainly reflected in the displacement of the person's arm, while the rest of the scene remains unchanged. Based on this characteristic, the neural network should mainly update the changed light to achieve efficient reconstruction in dynamic scenes.
[0122] The shadow binary image is downsampled at different scales. The low-resolution downsampled images are stacked and fused to create a multi-scale shadow map and a multi-scale ray set. Light rays that change between consecutive frames are extracted and stored in a differential ray set. Finally, the resulting multi-scale ray set and differential ray set are combined and fed into the MLP for training.
[0123] By training the network with multi-scale information including changing and overlapping light in old and new frames, the network can quickly reconstruct dynamic scenes.
[0124] We create both simulated and real datasets to evaluate the proposed method. These datasets contain static and dynamic scenes under various conditions.
[0125] Using Blender, we created a simulated dataset based on the scale of a real scene. We reproduced the ambient lighting and other interfering factors found in real scenes in the simulated scene. Based on the surface reflection characteristics of objects, we simulated the light propagation process in real scenes, making the simulated dataset more realistic as a reference. We recorded the spatial position of the light spot and the shadow binary image pairs. Each frame contained 25 pairs of light spot and shadow binary image pairs, and the resolution of the shadow binary image was 2.5. .
[0126] In order to further verify the light sampling method in practice, the present invention constructed a typical real scene to ensure the performance of the method in the real world. The experimental verification platform is a standard office space ( To simulate real-world occlusion, a 4m planar barrier was used for geometric alignment to create controllable line-of-sight occlusion. In the real-world dataset, the number of binary image pairs of light spots and shadows per frame was also 25. The specifications of the laser scanning galvanometer are as follows: 500mW continuous wave output at 520nm spectral emission; bidirectional angular resolution in the axial plane ; The reflective coating maintains 99% reflectivity across the entire operating spectrum.
[0127] The proposed dual-reflection non-viewing area reconstruction network was implemented using PyTorch and an NVIDIA GeForce RTX 3090 GPU. For each pixel in the binary segmentation map, ray sampling is performed using 64 hierarchical coarse samples, followed by 64 importance-weighted fine samples. This sampling process is supported by a hierarchical neural architecture consisting of a 4-layer coarse network for initial estimation and a 6-layer fine network for detail refinement. During training, to balance memory efficiency and convergence stability, the ray batch size is maintained at 1024 for each optimization iteration.
[0128] Two static scenes are selected to evaluate the accuracy of the dual-reflection non-line-of-sight imaging framework and are analyzed at different resolutions ( and ) and different scene complexities (single object and multi-object), the reconstruction accuracy of multi-resolution sampling and full sampling methods are benchmarked.
[0129] For the two sampling methods, the scene reconstruction accuracy is compared and analyzed at different training iterations. Figure 8 As shown, a residential interior scene containing a chandelier, sofa, coffee table and chairs is taken as an example. In the low-resolution scene ( ), multi-resolution sampling reduces the number of rays by 79.5% compared to full sampling while maintaining comparable reconstruction accuracy.
[0130] The low-resolution scene test shows the excellent efficiency of the multi-resolution method. The full sampling method requires 1-2 times more training iterations to achieve comparable accuracy. ), the multi-resolution sampling method used only 2.6% of the number of rays required for full sampling. The full sampling method was unable to reconstruct the structure of the chandelier, while the multi-resolution method achieved a detailed reconstruction of the chair after 1000 training iterations and a high-precision recovery of the chandelier at 2000 iterations. At this stage, the accuracy of the full sampling method degraded significantly. The root mean square error (RMSE) was calculated for the entire scene, which quantifies the depth deviation in the reconstructed 3D geometry. The full sampling method had a depth deviation of 4.2% in the low-resolution scene and 4.1% in the high-resolution scene, while the multi-resolution method achieved a depth deviation of only 2.6% at both resolutions.
[0131] Experimental results show that the multi-resolution method achieves high reconstruction accuracy at all tested resolutions, and its ability to enhance detail is particularly significant in high-resolution scenes. At a simulated scene size of 10m, the multi-resolution method achieves an absolute depth deviation of 0.26m, significantly better than the 0.42m deviation of the full-sampling method.
[0132] In order to evaluate the sampling performance under different scene complexities, the present invention sets up single-target and multi-target scenes, such as Figure 9 Quantitative validation shows that in single-target scenes, the multi-resolution approach achieves: (1) a 98.3% reduction in the number of rays (79.1% of which are important rays); and (2) enhanced structure preservation in the reconstruction of slender objects, which cannot be recovered by the full sampling approach.
[0133] Under complex multi-object conditions, the optimized method achieved 68.5% important ray utilization while maintaining a 98.5% ray reduction. After 500 training iterations, multi-resolution sampling demonstrated complete scene reconstruction and achieved high-accuracy output at 1,000 iterations. In comparison, the full sampling method required four times the computational effort (2,000 iterations) to achieve similar visual results.
[0134] In single-object scenes, the full-sampling method achieves a final depth deviation of 4.5%, while the multi-resolution method reduces this deviation to 4.2%. For multi-object scenes, the full-sampling method achieves a higher final depth deviation of 6.8%, while the multi-resolution method achieves a lower deviation of 3.4%.
[0135] The multi-resolution sampling method increases the proportion of important rays in the total number of rays through partitioned downsampling, which helps the network recover detailed features and achieve accurate reconstruction using fewer rays, greatly reducing the amount of computation.
[0136] Set up a moving person to test the reconstruction efficiency in dynamic scenes. Figure 10 As shown, in a moving sequence, frames are sampled at intervals of one and five frames to simulate slow and fast motion, respectively. In slow motion, the smaller inter-frame differences enable efficient reconstruction via differential ray sampling. While the full sampling approach exhibits a depth bias of 2.7%, the differential sampling approach reduces this bias to 1.9%. At a scene scale of 10 meters, the absolute depth bias is 0.19 meters. The differential sampling approach achieves accurate reconstruction of moving objects within 100 training iterations, outperforming the full sampling approach, which requires 500 iterations to achieve the same intermediate result. In fast motion, there is no spatial overlap between the sampled frames. The differential sampling approach completes reconstruction within 150 training iterations, while the full sampling approach still produces blurry reconstructions after 500 iterations. Compared to the full sampling approach's depth bias of 3.7%, the differential sampling approach reduces the bias to 3.3%, resulting in a relative improvement of 10.8% in reconstruction accuracy.
[0137] Within a limited number of iterations, full sampling methods cannot complete the training of the current ray batch in time, resulting in information loss. Differential sampling methods extract dynamic difference rays, preserving the reconstructed structure while updating the dynamic region. This strategy achieves robust dynamic reconstruction and prevents network forgetfulness by fully preserving features.
[0138] This result shows that the sampling method of the present invention has a significant advantage in slow motion with large correlation between consecutive frames. At the same time, for fast motion scenes, it also improves the reconstruction accuracy of dynamic scenes by reducing the number of training iterations required.
[0139] Initialization success is crucial to scene reconstruction efficiency. When processing a large number of rays, it becomes difficult for the network to build an initial model from the chaotic first batch of inputs. The initialization phase, which includes parameter assignment, data preparation, and model configuration, requires significant time and resources. Initialization failures require re-initialization, significantly reducing reconstruction efficiency. In dynamic reconstruction, an already reconstructed model may become corrupted during data updates.
[0140] Successful initialization depends heavily on the characteristics of the first batch of rays. The random initial rays of full sampling methods often lead to insufficient information and can also cause model collapse. Multi-resolution sampling methods effectively alleviate this limitation through two mechanisms: (1) ray number optimization: targeted reduction of the total number of rays while increasing the proportion of important rays; (2) multi-scale information supplementation: hierarchical spatial information integration replaces random ray selection.
[0141] like Figure 11As shown in Figure 3, the initialization success rate of full sampling in high-resolution scenarios is extremely low, which seriously affects the reconstruction efficiency. In contrast, the multi-resolution method maintains a 100% initialization success rate at all resolution scales.
[0142] Using simulated data of a lamp, the reconstruction efficiency of the sampling method is tested. Figure 12 As shown, in At the resolution of 100,000, the multi-resolution method reduces the number of rays from 6.56 million to 156,000, a reduction of 97.6%. It is worth noting that 96.3% (151,000) of these sampled rays are important rays, which significantly improves the sampling efficiency. It is worth mentioning that the reconstruction quality achieved by multi-resolution sampling in 15 seconds exceeds the result after 60 seconds of full sampling training, achieving a 4x speedup in static scene reconstruction. Similar improvements are observed in scenarios with high resolution. The multi-resolution approach effectively reduces the total number of rays while increasing the information density.
[0143] Multi-resolution sampling methods exhibit two distinct operating modes: at low resolutions, where the total number of rays limits training throughput, multi-resolution methods prioritize rays to improve sampling efficiency; as resolution increases, the computational advantages of reducing the number of rays gradually become dominant.
[0144] The multi-resolution method can reduce the total number of rays by one to two orders of magnitude. By properly screening rays and enhancing information density, it reduces unnecessary repeated training iterations, significantly lowers computational costs, and improves reconstruction efficiency.
[0145] To evaluate the sampling performance in dynamic environments, two different scene configurations are proposed: (1) a multi-target dynamic scene containing a fast-walking pedestrian and a stationary person with arm movements; and (2) a static background dynamic scene that highlights isolated arm movements in a fixed environment. Figure 13 Three advantages are demonstrated: (1) Reduction of temporal artifacts: The full sampling method cannot capture walking pedestrians in time after the data update is completed; in contrast, the differential sampling method achieves complete motion trajectory reconstruction without frame loss; (2) Accelerated convergence: The differential sampling method completes detailed scene reconstruction in 10 seconds, which is 6 times faster than the 60 seconds required by the full sampling method, reducing training time by 83% without sacrificing spatiotemporal accuracy; (3) Enhanced motion sensitivity: Through differential ray priority, the differential method allocates 78% of computing resources to the changing area, making the deformation parameters converge 92% faster than the full sampling method.
[0146] The differential ray mechanism achieves accurate detection of scene spatial disturbances through spatiotemporal gradient analysis, thereby optimizing the allocation of computing resources and accelerating scene reconstruction after dynamic changes.
[0147] To test the sampling method's performance in typical static and dynamic scenarios, we constructed a real-world experimental environment. The experimental setting was an indoor space with a distance of 5.22 meters between two walls. A carefully arranged office scene, containing hidden objects of varying sizes and shapes, was constructed within the hidden area. Subsequent dynamic and static scenarios were constructed based on this model.
[0148] The experimental results of real static scenes are as follows Figure 14 As shown. The resolution of the shadow binary image used for reconstruction is , with 64 images. In this experiment, the full sampling approach generated 11 million rays, while the multiresolution approach reduced the number of rays to 4% and increased the density of important rays to 97%. The multiresolution sampling framework achieved reconstruction accuracy comparable to the full sampling 2000 iteration baseline in just 500 training cycles. The multiresolution approach successfully resolved key structural components, including the geometry of the table legs and the outline of the ergonomic chair.
[0149] The experimental results of real dynamic scenes are as follows Figure 15 This dynamic scene serves as a controlled experimental framework for evaluating two key capabilities: (1) the ability to integrate the addition of scene dynamics in real time; and (2) the ability to maintain the stability of the baseline geometry during continuous updates. To establish consistency across the experiments, both sampling methods are initialized using the same pre-trained scene representation to ensure standardized initialization across the data update sequence.
[0150] When new data is input, the differential sampling method quickly updates the model based on the subsequent data and immediately reconstructs the outline of the newly added person, quickly completing a detailed reconstruction of the body outline. In contrast, the full sampling method has difficulty in instantly reconstructing new people in the same scene.
[0151] The multi-resolution sampling method has the ability to continuously refine previously incompletely reconstructed static backgrounds. By performing detailed modeling, it accurately reconstructs missing structural components, such as table legs, while progressively enhancing other static background regions, significantly improving reconstruction accuracy.
[0152] This paper proposes an efficient dual-reflection non-line-of-sight imaging framework that combines neural implicit representation with an optimized light sampling strategy. By integrating Neural Radiance Field (NeRF) and shadow binary images, it achieves high-precision reconstruction of hidden scenes using cumulative transmittance modeling.
[0153] The present invention has three key improvements: First, a low-cost galvanometer-based laser scanning system is developed for capturing shadow images. Compared with time-of-flight (TOF) transient imaging methods, this system significantly reduces hardware complexity while retaining the ease of calibration. At one-fifth the cost of traditional non-line-of-sight devices, the system achieves a portable and deployable scene configuration. Second, the present invention's sampling optimization method improves computational efficiency through ray selection, reducing the amount of data required by 1-2 orders of magnitude compared to full sampling methods. This innovation enables large-scale environments ( ) with a depth deviation of only 2%. Third, the framework achieves rapid dynamic scene reconstruction through high-frequency data acquisition, achieving low latency in reflecting dynamic scene changes and completing the reconstruction of moving objects.
Claims
1. A fast double-reflection non-line-of-sight imaging method based on important light screening, characterized in that: include: Step 1: Calibrate the sampling system based on the NLOS scenario to determine the size of the space to be reconstructed; Step 2: Use the galvanometer system to change the laser light path to illuminate the hidden space, use the camera to obtain the shadow image produced by the active illumination, and process the shadow image based on the calibration information to obtain the pose and shadow binary image in 3D coordinates; Step 3: Perform light modeling based on the double-reflection NLOS scene and build a 3D reconstruction network for the shadow scene; Step 4: Sample the shadow binary image obtained in step 2 using a multi-resolution sampling method based on edge intensity; Step 5: Input the sampled shadow binary image and the pose in step 2 into the 3D reconstruction network in step 3 for training; Step 6: Repeat step 2, use the difference ray sampling method to dynamically sample the shadow binary image obtained in step 2, and repeat step 5; Step 7: Repeat step 6 to achieve 3D reconstruction of the NLOS dynamic scene.
2. The fast double-reflection non-line-of-sight imaging method based on important light screening according to claim 1, characterized in that: The step 1 is specifically as follows: The sampling system is calibrated based on the NLOS scenario; the center point of the relay wall is , the center point of the projection wall is ; to connect and The midpoint of the line segment is the origin , establish a three-dimensional space coordinate system; the distance between the two walls is , the width of the relay wall is W and the height is .
3. The fast double-reflection non-line-of-sight imaging method based on important light screening according to claim 2, characterized in that: The step 2 is specifically as follows: In order to realize the reconstruction of hidden objects, a double-reflection non-line-of-sight imaging system is constructed, which consists of a camera and a laser scanning galvanometer system. The incident direction of the laser is changed by the laser scanning galvanometer system, and the positions on the relay wall are respectively 、 ,..., Use the camera to capture the changing shadows on the projection wall, thereby obtaining a set of shadow images that can reflect the characteristics of the hidden object; The data processing process is as follows: (1) Perform perspective transformation on the light spot under oblique viewing angle to obtain the light spot under normal viewing angle; let the coordinates of the four corner points on the original image be ,in ; The coordinates of the four corresponding points on the target plane are ,in: ; and are the pixel width and height of the sampled image respectively; The homogeneous coordinate representation of the perspective transformation is: ; For each pair of corresponding points , ,in is the coordinate of the light spot under the oblique angle before perspective transformation, is the coordinate of the light spot under the positive viewing angle after perspective transformation; solve the various value: 、 ,..., ; By performing perspective transformation on all pixels in the captured image, the corrected spot position is obtained; (2) Traverse the grayscale shadow image For all pixels in The corresponding pixel coordinates ( , ); Set the pixel coordinates ( , ) is converted to three-dimensional coordinates ( ), the conversion formula is as follows: ; Assuming that there are N light spots corresponding to N shadow images, the corresponding light spot coordinates are expressed as ; (3) Perform the same perspective transformation operation on the shadow image to obtain the corrected shadow image; perform histogram equalization on the shadow image and use an adaptive threshold to perform a binarization operation.
4. The fast double-reflection non-line-of-sight imaging method based on important light screening according to claim 3, characterized in that: The step 3 is specifically as follows: Based on the principle of neural radiation fields, a multi-layer perceptron is used to implicitly model light and hidden objects. Light rays are constructed based on the light spot position and shadow image, and 3D coordinates are sampled along each ray. An MLP is used to predict the probability of the existence of an object at a 3D coordinate. The 3D structure is reconstructed by accumulating the existence probabilities, forming a self-supervised learning framework for implicitly modeling the hidden space. The probability of the existence of hidden objects is expressed as cumulative transmittance It describes the light in the interval The probability of light propagating without being blocked; The expected cumulative transmittance is expressed as ,in Indicates the distance from the surface of the object to the relay wall, represents the light spot position, and t represents the light propagation time; Near the border and far boundaries In the range between To estimate and calculate the cumulative transmittance : ; In the imaging task, enter the spot position , which can query the three-dimensional coordinates in the hidden space The cumulative transmittance at ; Continuous Implicit Neural Shadow Field Expressed as: ; From the spot position Construct the light corresponding to the shadow image and trace the light to the projection wall; use the layered sampling method of discrete samples to estimate the cumulative transmittance The integral result of: ; in, Indicates adjacent moments The distance between samples, represents the cumulative transmittance, Indicates opacity; Perform binary segmentation on the shadow image to obtain the distribution of light passing through the hidden object; mark the pixels in the bright area as "1", indicating that the light has passed through the hidden scene; mark the pixels in the dark area as "0", indicating that the light is blocked; a black background indicates that the light is not blocked when passing through the unknown space, and a white foreground indicates that the light is blocked by the object, thus forming a shadow on the wall; the optimization process of the network is supervised by calculating the cumulative transmittance; the loss function of the cumulative transmittance is Expressed as: ; in, is the set of rays in each batch during training, indicating the rays involved in training; Finally, the depth of the hidden object is calculated from multiple perspectives by numerically integrating N rays. : 。 5. The fast double-reflection non-line-of-sight imaging method based on important light screening according to claim 4, characterized in that: The step 4 is specifically as follows: There are N light spots on the relay wall ( , , ,..., ), the light is emitted from these N light spots to the projection wall; the resolution of the collected shadow binary image is ; Emit light from the light spot on the relay wall to the projection wall ,in , , so that we get a ray of light; The rays containing structural information are defined as important rays. A multi-resolution sampling method based on edge intensity is proposed, which uses different downsampling factors for different regions and divides the regions into sparse sampling regions and dense sampling regions according to edge intensity. For a size of The shadow binary image is divided into rectangular regions, each with a height of , with a width of ; Detect the edges within each rectangular area and calculate its edge strength; Use the Sobel operator to calculate the average edge strength of all pixels in each region , and set a threshold , in order to decide whether to perform focused sampling in this area; when When , the region is classified as a densely sampled region, using the first downsampling factor To preserve more important light, ;when When , the region is classified as a sparsely sampled region and the second downsampling factor is used Downsample the sparsely sampled area, ; In this way, the sampling density is adjusted in different regions; the light from each region that has undergone different downsampling processes is merged and input into the MLP for training; the latent space is reconstructed using light of different resolutions, and the original resolution of the structure is preserved.
6. The fast double-reflection non-line-of-sight imaging method based on important light screening according to claim 5, characterized in that: The step 6 is specifically as follows: In a dual-reflection non-line-of-sight system, the positions of the light spot and shadow are changed by a galvanometer to obtain shadow images and position data; the relay wall is scanned repeatedly to obtain temporally continuous dynamic data; In the dynamic reconstruction task, the galvanometer scanning and data acquisition operations are continuously performed, and the newly sampled shadows and coordinates are input into the MLP; A differential sampling method is used to reconstruct dynamic scenes. The changing light is extracted from consecutive frames captured at the same spot position to model the dynamics of hidden objects. Based on the complete reconstruction of the first frame, subsequent frames only reconstruct the dynamic part. Adopt image pyramid strategy to supplement global feature information; The shadow binary image is downsampled at different scales. The downsampled images are stacked and fused to create a multi-scale shadow map and a multi-scale ray set. At the same time, the light rays that change between consecutive frames are extracted and stored in a differential ray set. Finally, the resulting multi-scale ray set and differential ray set are mixed and input into the MLP for training. By training the network with multi-scale information including changing and overlapping light in old and new frames, the network can quickly reconstruct dynamic scenes.
7. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the program, the steps of the method according to any one of claims 1 to 6 are implemented.
8. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the program is executed by a processor, the steps of the method according to any one of claims 1 to 6 are implemented.
9. A computer program product comprising a computer program, characterized in that When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 6 are implemented.
Citation Information
Patent Citations
Digital twin modeling method and system based on SAM large model and NeRF
CN117671138A
View-around scene three-dimensional reconstruction method based on implicit representation
CN118864734A
Building group multi-scale high-efficiency high-precision digital twinning multi-resolution neural radiation field method
CN119783536A
Method and device for optical system online designing based on intelligent light computing, and storage medium
US20250155702A1
Cited By
A physically guided non-line-of-sight three-dimensional imaging method and system
CN122435162A