Fast bi-directional non-visibility imaging method based on important ray screening

By selecting important rays and optimizing the ray training process, and utilizing a laser scanning galvanometer system and a neural implicit model, the problem of low reconstruction efficiency and accuracy in non-visual field imaging was solved, achieving rapid and accurate 3D reconstruction in complex environments.

CN120635335BActive Publication Date: 2025-10-24NANJING UNIV OF SCI & TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511129614.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-08-13
Publication Date
2025-10-24
Estimated Expiration
2045-08-13

AI Technical Summary

Technical Problem

Existing non-view-of-sight imaging technologies are greatly affected by the accuracy of the projection point position and the clarity of the shadow boundary when reconstructing hidden targets. Furthermore, the number of light rays increases exponentially with resolution, resulting in low reconstruction efficiency and difficulty in adapting to and achieving high-precision reconstruction in complex environments.

Method used

By employing a fast dual-reflection non-viewpoint imaging method based on important ray selection, a laser scanning galvanometer system is used to adjust the projection position, construct a neural implicit model, select important rays for 3D reconstruction, and optimize the ray training process by combining multi-resolution sampling and differential ray sampling.

Benefits of technology

It achieves rapid and accurate 3D reconstruction in complex environments, reduces computational latency, improves reconstruction accuracy and efficiency, and significantly enhances dynamic scene reconstruction capabilities.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120635335B_ABST
    Figure CN120635335B_ABST
Patent Text Reader

Abstract

The application discloses a kind of fast double reflection non-visual area imaging methods based on important light screening, which comprises: based on NLOS scene, calibrating sampling system, determine the size of the space to be reconstructed;Change laser light path illumination hidden space, obtain the shadow image generated by active illumination, process the shadow image based on calibration information, get the pose and shadow binary image under three-dimensional coordinates;Based on double reflection NLOS scene, light modeling is carried out, and the three-dimensional reconstruction network of shadow scene is built;The shadow binary image is sampled using a multi-resolution sampling method based on edge intensity;The shadow binary image after sampling is input into the reconstruction network for training with the pose;The shadow binary image obtained is dynamically sampled using a difference light sampling method;The three-dimensional reconstruction of NLOS dynamic scene is realized.The application realizes low delay in reflecting the dynamic change of the scene and completes the reconstruction of the moving object.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the field of non-line-of-sight imaging, and in particular to a fast double-bounce non-line-of-sight imaging method based on important ray screening. BACKGROUND

[0002] Traditional detection methods follow the imaging criteria of geometric optics, limiting the capture of photons within the line-of-sight range. Non-line-of-sight (NLoS) imaging, i.e., the perception of targets outside the line of sight, has become an important research field. The basic configuration of non-line-of-sight imaging is in a scene where the light propagation path between the target and the detector is blocked. Through the indirectly propagated photons, the light transmission process and the target structure are modeled and analyzed.

[0003] Current mainstream solutions include coherent information methods based on coherent light sources, photon time-of-flight methods using ultrafast photon detectors, and two-dimensional intensity information methods using shadows. The coherent information method is affected by the intermediate surface scattering effect of the relevant light. The effective coherent information is not enough to reconstruct the hidden scene, making it more complex to achieve high-quality around-the-view imaging. The photon time-of-flight method uses the time-of-flight information of the captured photons to model the three-reflecting process of the photons, which is a relatively mature around-the-view solution. However, due to the high cost of light sources and photon detectors, many researchers can only be limited to simulation environments and public data sets, which hinders the expansion of these methods in more extensive research and application fields. The two-dimensional intensity information method uses a traditional camera to record the light intensity information in the scene, and its effectiveness depends on the quality of the target environment lighting, making it difficult to adapt to daily tasks in complex environments.

[0004] In 2020, Connor Henley et al. proposed a method using two intermediate surfaces (Henley, C., et al. Imaging behind occluders using two-bounce light. in Computer Vision–ECCV 2020: 16th European Conference, Glasgow, UK, August 23–28, 2020, Proceedings, Part XXIX 16. 2020. Springer.) defined as a distributed scheme: one surface as a light source receiving surface, and the other as a shadow imaging surface. By constantly adjusting the projection angle of the laser point light source, the structure of the hidden scene is carved using shadows and light source positions. The schematic diagram of the integrated and distributed non-line-of-sight scheme is as follows Figure 1The configuration is shown in FIG. 1. In this configuration, the active light source and detector only need to modulate and respond to the light intensity. Compared with the integrated method of imaging the scene using scattered coherent light information, this design significantly reduces the dependence on the performance of the light source and sensor. In practical applications, the distributed scheme can flexibly select the reflecting surface and imaging surface according to different tasks, and has greater potential in terms of observation field of view, device size and cost.

[0005] However, the effect of engraving depends on the accuracy of the projection point position and the clarity of the shadow boundary. Deviation of these two factors will cause reconstruction error, and the laser points projected to the intermediate surface will inevitably produce local scattering, thereby introducing errors in the engraving process. Secondly, the accuracy of the reconstruction depends on the number and fineness of the light rays, but the number of light rays increases exponentially with the resolution of the shadow binary image and the number of projections, which greatly affects the reconstruction efficiency of the hidden scene.

[0006] In 2021, the paper (Mildenhall, B., et al., Nerf: Representing scenes as neural radiance fields for view synthesis. Communications of the ACM, 2021. 65(1): p. 99-106.) proposed a neural radiance field (NeRF), which is a deep learning model for three-dimensional implicit space modeling. Compared with previous algorithms, it only needs multi-view images, corresponding internal and external parameters and poses, without explicit three-dimensional reconstruction process, to construct the implicit model of the current scene for new view rendering and three-dimensional reconstruction. The idea of implicit modeling has also been applied to non-line-of-sight imaging. In 2021, the paper (Shen, S., et al., Non-line-of-sight imaging via neural transient fields. IEEE Transactions on Pattern Analysis and Machine Intelligence, 2021. 43(7): p. 2257-2268.) proposed the concept of neural transient field (NeTF), which uses the captured photon transient signal to construct a neural implicit representation model in a transient around-view vision system based on photon time of flight, significantly improving the reconstruction performance. In 2022, the paper (Mu, F., et al., Physics to the rescue: Deep non-line-of-sight reconstruction for high-speed imaging. IEEE Transactions on Pattern Analysis and Machine Intelligence, 2022.) used conditional radiance field in transient reconstruction, introducing light transmission and volume rendering physics prior into the neural network to reconstruct fine details. So far, there is no case of implicit neural modeling applied to around-view shadow reconstruction, and the key problem is that it is difficult to extract feature points for matching in the shadow and to accurately obtain the shadow trajectory. SUMMARY

[0007] In order to overcome the shortcomings of the prior art, the present application provides a fast double-reflection non-line-of-sight imaging method based on important light screening, which forms multi-view shadow images by adjusting the projection position of the laser point, screens important light from the shadow image which forms the shadow of the hidden target, constructs a neural implicit model, and realizes fast and accurate reconstruction of the hidden target.

[0008] The technical scheme for achieving the object of the present application is as follows: a fast double-reflection non-visual field imaging method based on important light screening, comprising:

[0009] Step 1: calibrate the sampling system based on the NLOS scene, and determine the size of the space to be reconstructed;

[0010] Step 2: change the laser light path to illuminate the hidden space through the galvanometer system, use the camera to obtain the shadow image generated by the active illumination, process the shadow image based on the calibration information, and obtain the pose and shadow binary image under the three-dimensional coordinates;

[0011] Step 3: model the light based on the double-reflection NLOS scene, and build a three-dimensional reconstruction network for the shadow scene;

[0012] Step 4: use a multi-resolution sampling method based on edge intensity to sample the shadow binary image obtained in step 2;

[0013] Step 5: input the sampled shadow binary image and the pose in step 2 into the reconstruction network in step 3 for training;

[0014] Step 6: repeat step 2, use a difference light sampling method to dynamically sample the shadow binary image obtained in step 2, and repeat step 5;

[0015] Step 7: repeat step 6 to realize three-dimensional reconstruction of the NLOS dynamic scene.

[0016] Further, step 1 is specifically as follows:

[0017] The sampling system is calibrated based on the NLOS scene; the center point of the relay wall is , the center point of the projection wall is ; the midpoint of the line segment connecting and is taken as the origin , and a three-dimensional space coordinate system is established; the distance between the two walls is , the width of the relay wall is W, and the height is .

[0018] Further, step 2 is specifically as follows:

[0019] In order to realize the reconstruction of hidden objects, a double-reflection non-visual field imaging system is constructed, which is composed of a camera and a laser scanning galvanometer system; the incident direction of the laser is changed through the laser scanning galvanometer system, and light spots with positions , , are generated on the relay wall; the changing shadows on the projection wall are captured by the camera, thereby obtaining a group of shadow images that can reflect the characteristics of the hidden objects;

[0020] The data processing process is as follows:

[0021] (1) The perspective transformation is performed on the light spot under the oblique angle to obtain the light spot under the normal angle; the coordinates of the four corner points on the original image are set as , wherein ;

[0022] The coordinates of the four corresponding points on the target plane are , wherein:

[0023] ;

[0024] and are the pixel width and height of the sampling picture, respectively;

[0025] The homogeneous coordinate representation of the perspective transformation is:

[0026] ;

[0027] For each pair of corresponding points , , wherein is the light spot coordinate under the oblique angle before the perspective transformation, and is the light spot coordinate under the normal angle after the perspective transformation; the values of each in the perspective transformation matrix are solved according to the above transformation formula: , ,..., ; the corrected light spot position is obtained by performing perspective transformation on all pixel points in the collected image;

[0028] (2) All pixels in the gray-scale shadow image are traversed to find the pixel coordinates corresponding to the maximum gray value ; the pixel coordinates are converted into three-dimensional coordinates , and the conversion formula is as follows: ;

[0029] Assuming that there are N light spots corresponding to N shadow images, the corresponding light spot coordinates are represented as ;

[0030] (3) The same perspective transformation operation is performed on the shadow image to obtain the corrected shadow image; the histogram equalization processing is performed on the shadow image, and the adaptive threshold is used for the binarization operation.

[0031]

[0032] Further, step 3 is specifically:

[0033] Based on the principle of neural radiance field, a multi-layer perceptron is used to implicitly model the light and hidden objects; the light is constructed according to the spot position and the shadow image, and the three-dimensional coordinates are sampled along each light; the MLP is used to predict the possibility of the existence of objects on the three-dimensional coordinates; the three-dimensional structure is reconstructed by accumulating the existence possibility, forming a self-supervised learning framework for implicitly modeling the hidden space;

[0034] The possibility of the existence of the hidden object is represented by the cumulative transmittance , which describes the probability of the propagation of the light without being blocked in the interval ; the expected cumulative transmittance of the light is represented as , where represents the distance from the surface of the object to the relay wall, represents the spot position, and t represents the light propagation time; in the range between the near boundary and the far boundary , the cumulative transmittance is estimated and calculated using :

[0035] ;

[0036] In the imaging task, the input spot position can query the cumulative transmittance of the three-dimensional coordinates in the hidden space; the continuous implicit neural shadow field is represented as:

[0037] ;

[0038] The light corresponding to the shadow image is constructed from the spot position , and the light reaching the projection wall is tracked; the integral result of the cumulative transmittance is estimated using the hierarchical sampling method of discrete samples:

[0039] ;

[0040] wherein represents the distance between the samples between adjacent time points , represents the cumulative transmittance, represents the opacity;

[0041] ​​The shadow image is binarized to obtain the light distribution through the hidden object; the pixels in the bright area are marked as "1", indicating that the light passes through the hidden scene; the pixels in the dark area are marked as "0", indicating that the light is blocked; the black background indicates that the light is not blocked when passing through the unknown space, and the white foreground indicates that the light is blocked by the object, thereby forming a shadow on the wall; the optimization process of the network is supervised by calculating the cumulative transmittance; the loss function of the cumulative transmittance is represented as:

[0042] ;

[0043] wherein, is the set of light rays for each batch during training, and represents the light rays participating in training;

[0044] Finally, the depth of the hidden target is calculated from multiple perspectives by numerically integrating N light rays :

[0045] .

[0046] Further, step 4 is specifically:

[0047] There are N light spots on the relay wall , , ,..., , and light rays are emitted from these N light spots to the projection wall; the resolution of the collected shadow binary image is ; light rays are emitted from the light spots on the relay wall to the projection wall , wherein , , so a total of light rays are obtained;

[0048] The light rays containing structural information are defined as important light rays; a multi-resolution sampling method based on edge intensity is proposed, which uses different down-sampling factors for different regions, and divides the regions into sparse sampling regions and dense sampling regions according to the edge intensity;

[0049] For a shadow binary image with a size of , it is divided into rectangular regions, each with a height of and a width of ; the edges in each rectangular region are detected, and the edge intensity is calculated; the average edge intensity of all pixels in each region is calculated using the Sobel operator , and a threshold is set to determine whether the region is focused on sampling;

[0050] When When the area is classified as a dense sampling area, a first down-sampling factor is used to retain more important rays, When the area is classified as a sparse sampling area, a second down-sampling factor is used to down-sample the sparse sampling area, ;

[0051] In this way, the sampling density is adjusted in different areas; the rays in each area after different down-sampling processing are combined and input into the MLP for training; the hidden space is reconstructed using rays of different resolutions, and the original resolution of the structure is retained.

[0052] Further, step 6 is specifically:

[0053] In the double-reflection non-visual area system, the position of the light spot and the shadow is changed by the galvanometer to obtain the shadow image and position data; the relay wall is repeatedly scanned to obtain time-continuous dynamic data; in the dynamic reconstruction task, the galvanometer scanning and data acquisition operation are continuously performed, and the newly sampled shadow and coordinates are input into the MLP to realize fast reconstruction;

[0054] A differential sampling method is used to reconstruct the dynamic scene; the changed rays are extracted from the continuous frames captured from the same light spot position to model the dynamics of the hidden object; on the basis of the complete reconstruction of the first frame, only the dynamic part is reconstructed in the subsequent frames;

[0055] An image pyramid strategy is used to supplement global feature information;

[0056] Different scale down-sampling processing is performed on the shadow binary image; the down-sampled images are stacked and fused to create multi-scale shadow maps and construct multi-scale ray sets; at the same time, the rays changed between continuous frames are extracted and stored in the differential ray set; finally, the obtained multi-scale ray set and differential ray set are mixed and input into the MLP for training;

[0057] By using multi-scale information containing changed rays and overlapping rays in new and old frames to train the network, the network can quickly reconstruct the dynamic scene.

[0058] An electronic device includes a memory, a processor, and a computer program stored on the memory and executable on the processor, and the processor implements the steps of the above method when executing the program.

[0059] A computer readable storage medium has a computer program stored thereon, and the program is executed by a processor to implement the steps of the above method.

[0060] ​A computer program product comprising a computer program which, when executed by a processor, implements the steps of the above method.

[0061] Compared with the prior art, the present application has the following advantages:

[0062] (1) The present application performs spatial and temporal decomposition on multi-view shadow images to extract important rays that mainly shape the hidden scene, optimizes the NeRF model to reduce a large number of rays with repetitive structures, and realizes accurate and efficient reconstruction of hidden objects; the present application demonstrates the rapid three-dimensional reconstruction capability of dynamic and static occlusion space, realizes a scene relative depth deviation of 2%, and an absolute deviation of 0.2 meters under the scene scale.

[0063] (2) The present application constructs a brand-new double-reflection non-visual area imaging system, continuously scans the scene by using a laser galvanometer acquisition system, updates the sampled scene data, uses the latest ray sampling method to obtain the difference rays in two continuous samplings, updates the existing reconstruction model by using the difference rays, and realizes the function of reconstructing the NLOS dynamic scene. BRIEF DESCRIPTION OF DRAWINGS

[0064] Figure 1 It is an integrated and distributed non-visual area scheme.

[0065] Figure 2 It is a double-reflection non-visual area (NLOS) imaging system, wherein (a) is the system composition, (b) is the ray sampling and reconstruction, and (c) is the reconstruction framework.

[0066] Figure 3 It is a scene parameter visualization schematic diagram.

[0067] Figure 4 It is a data processing schematic diagram.

[0068] Figure 5 It is a ray sampling visualization schematic diagram.

[0069] Figure 6 It is a multi-resolution sampling method schematic diagram based on feature intensity partitioning.

[0070] Figure 7 It is a training method schematic diagram combining multi-resolution and frame difference.

[0071] Figure 8 It is a comparison of full sampling and multi-resolution sampling methods under different shadow resolutions, and the gray text annotation is the root mean square error (RMSE) value of the overall depth.

[0072] Figure 9 ​For comparison of full sampling and multi-resolution sampling methods under different scene complexities, the root mean square error (RMSE) of the overall depth is annotated in gray text.

[0073] Figure 10 For comparison of full sampling and differential sampling methods in dynamic scenes, the root mean square error (RMSE) of the overall depth is annotated in gray text.

[0074] Figure 11 For initialization comparison of full sampling and multi-resolution sampling methods.

[0075] Figure 12 For efficiency comparison of full sampling and multi-resolution sampling methods under different resolutions.

[0076] Figure 13 For comparison of full sampling and differential sampling methods in multi-target and background scenes.

[0077] Figure 14 For comparison of full sampling and multi-resolution sampling methods in real static scenes.

[0078] Figure 15 For comparison of full sampling and differential sampling methods in real static scenes. DETAILED DESCRIPTION

[0079] The technical solutions of the present application will be described in detail below in combination with the drawings and examples.

[0080] Non-line-of-sight imaging uses a wall to reflect laser light and a wall to receive shadows, and constructs a ray sculpting model to realize three-dimensional reconstruction of hidden scenes. However, the degree of detail of hidden object delineation depends on a large number of multi-view rays, which leads to a fold increase in computational latency. The present application proposes a fast implicit ray sculpting method, which uses a hierarchical neural radiance field (NeRF) model to reconstruct hidden objects. Specifically, the multi-view shadow images are spatially and temporally decomposed to extract important rays that mainly shape the hidden scene, and the NeRF model is optimized to reduce a large number of rays with repetitive structures, to realize accurate and efficient reconstruction of hidden objects. The method of the present application demonstrates the ability of fast three-dimensional reconstruction of dynamic and static occlusion spaces, and realizes a scene relative depth deviation of 2% (absolute deviation of 0.2 meters under the scene scale).

[0081] ​Firstly, the double reflection non-line-of-sight (NLOS) imaging technology is introduced. An unknown object is blocked by an obstacle in the line of sight, and the detector cannot directly obtain its structural information. The two sides of the hidden object (relay wall and projection wall) are Lambertian walls, and it is assumed that the surface brightness of the two walls is the same when observed from various directions under a fixed illumination angle. A laser beam is emitted from outside the space, and after being reflected by the relay wall, it produces a light spot and scattered light. The scattered light passes through the hidden space, reaches the projection wall and projects a shadow. To simplify the process, high-order reflected light is ignored. The shadow reflects the structure of the hidden space, and by changing the position of the light spot, the scattered light can be tracked, and the structure of the hidden object can be reconstructed.

[0082] The present application calibrates the system based on the scene. As shown in Figure 3 , the center point of the relay wall is , and the center point of the projection wall is . A three-dimensional coordinate system is established with the midpoint of the line segment connecting and as the origin . In the system of the present application, the distance between the two walls is , the width of the relay wall is W, and the height is .

[0083] In order to realize the reconstruction of the hidden object, the present application constructs a brand new double reflection non-line-of-sight imaging system, as shown in Figure 2 (a). The system mainly consists of a camera and a laser scanning galvanometer system. By changing the incidence direction of the laser through the galvanometer, light spots with positions , , are generated on the relay wall. The changing shadow on the projection wall is captured by the camera, and a set of shadow images reflecting the characteristics of the hidden object is obtained.

[0084] Unlike the previous double reflection non-line-of-sight imaging system, the new system adds a dynamic reconstruction function. The previous imaging system only scans once and models the scene statically. The new system has a dynamic reconstruction function, which is specifically reflected in the data acquisition system repeatedly scanning the scene, inputting the data sampled each time into the reconstruction system, and using the latest light sampling method to obtain the difference light between two consecutive samplings, and using the difference light to update the existing reconstructed scene, achieving the function of reconstructing dynamic scenes.

[0085] The data processing process is shown in Figure 4 :

[0086] (1) Perspective transformation is performed on the light spot under the oblique angle to obtain the light spot under the orthoview angle. Let the coordinates of the four corner points on the original image be , where .

[0087] The coordinates of the four corresponding points on the target plane are wherein:

[0088] ;

[0089] and are the pixel width and height of the sampling image, respectively.

[0090] The homogeneous coordinate representation of the perspective transformation is:

[0091] ;

[0092] For each pair of corresponding points , , the values of each in the perspective transformation matrix can be solved according to the above transformation formula. By performing perspective transformation on all pixel points in the collected image, the corrected spot position can be obtained.

[0093] (2) Traverse all pixels in the gray-scale shadow image to find the maximum gray value (i.e., the center of the spot) corresponding pixel coordinates , . Then, convert the pixel coordinates , to three-dimensional coordinates , and the conversion formula is as follows:

[0094] ;

[0095] Assuming that there are N spots corresponding to N shadow images, the corresponding spot coordinates can be represented as .

[0096] (3) Perform the same perspective transformation operation on the shadow image to obtain the corrected shadow image. Perform histogram equalization processing on the shadow image and use adaptive thresholding for binaryzation operation.

[0097] The light modeling process is described in detail below. Based on the principle of neural radiance field (NeRF), a multi-layer perceptron (MLP) is used to implicitly model the light and hidden objects, as shown in figure (c) in Figure 2 . According to the spot position and shadow image, construct the light and sample the three-dimensional coordinates along each light. The MLP is used to predict the possibility of the existence of objects at the three-dimensional coordinates (i.e., the possibility of the existence of objects in the hidden space). Then, by accumulating the existence possibility, the shadow image is reconstructed, forming a self-supervised learning framework for implicit modeling of the hidden space.

[0098] The possibility of hidden object existence is represented by the cumulative transmittance , which describes the probability of a light ray propagating without being blocked in the interval . The expected cumulative transmittance of a light ray is denoted as , where is the distance from the object surface to the relay wall, is the light spot position, and t is the light propagation time. In the range between the near boundary and the far boundary , the cumulative transmittance is estimated and calculated using :

[0099]

[0100] In the imaging task, the input light spot position can be used to query the cumulative transmittance at the three-dimensional coordinates in the hidden space, as shown in Fig. (b) in . This continuous implicit neural shadow field can be represented as: Figure 2

[0101]

[0102] From the light spot position , the light rays corresponding to the shadow image are constructed and traced to the projection wall. The integral result of the cumulative transmittance is estimated using the hierarchical sampling method of discrete samples:

[0103]

[0104] where is the distance between samples between adjacent time instants, is the cumulative transmittance, is the opacity.

[0105] The shadow image is binarized to obtain the light distribution through the hidden object. The pixels in the bright area are marked as “1”, indicating that the light ray has passed through the hidden scene; the pixels in the dark area are marked as “0”, indicating that the light ray has been blocked. The black background represents the light ray that has not been blocked while passing through the unknown space, and the white foreground represents the light ray that has been blocked by the object, thereby forming a shadow on the wall. The optimization process of the network is supervised by calculating the cumulative transmittance. The loss function of the cumulative transmittance can be represented as:

[0106] ​​​​​​​​

[0107] wherein, is the set of light rays for each batch during training, representing the light rays participating in training.

[0108] Finally, the depth of the hidden target is calculated from multiple perspectives by numerically integrating N light rays :

[0109]

[0110] The optimized light ray sampling scheme of the present application is introduced below. The reconstruction method of the present application is based on light ray modeling, and the light ray generation process is as shown in Figure 5 There are N light spots on the relay wall , , ,..., . Light rays are emitted from these N light spots to the projection wall. The resolution of the collected shadow binary image is . Light rays are emitted from the light spots on the relay wall to the projection wall , where , , so a total of light rays are obtained.

[0111] The full sampling method is to randomly shuffle the order of the obtained light rays, and then input them into the network in batches for training. During training, each light ray is considered to have the same importance. After multiple light ray shuffling and training, the structure of the hidden scene can be recovered. However, this method relies on a large number of repeated training iterations, resulting in low reconstruction efficiency.

[0112] At the beginning of training, due to the lack of effective model initialization prior information, the randomly input light rays cannot guarantee the uniformity of the structural information, which is easy to make the model fall into local extreme value, and further lead to the collapse of the model. In addition, the edges of the shadow binary image usually contain key structural information, while there are a large number of flat pixels (i.e. invalid light rays) in the high-resolution shadow binary image. The full sampling method uniformly samples and models each light ray, making the network pay too much attention to invalid light rays during training, resulting in limited accuracy of the network in reconstructing the target (especially the detailed part). In dynamic scenes, it is very inefficient to sample all light rays frame by frame and train from scratch within a scanning period. The difference between consecutive frames should be considered to enable the model to prioritize processing light rays with dynamic structural information.

[0113] Light rays containing structural information such as edges and contours are defined as important light rays. To improve the proportion of important light rays in the overall training set, the present application proposes a multi-resolution sampling method based on edge intensity. This method not only improves training efficiency, but also solves the problem of initialization collapse. The multi-resolution sampling method uses different down-sampling factors for different regions, and divides the regions into sparse sampling regions and dense sampling regions according to edge intensity.

[0114] The white bright area in the shadow binary image means that the light ray is not blocked by the object, and the light ray in these areas provides less structural information for reconstruction. Therefore, a larger down-sampling factor is used in these areas. The internal shadow area indicates that the light ray is completely blocked by the object, which also provides less structural information, so a larger down-sampling factor is also used. For edge regions, focused sampling or full sampling is used to improve the reconstruction accuracy of the model in terms of edges and details. Since the blank area has limited contribution to the reconstruction of hidden objects, a larger down-sampling factor can reduce invalid light rays and redundant calculations.

[0115] For a shadow binary image with a size of , it is divided into rectangular regions, each with a height of and a width of , as shown in Figure 6 . Detect the edges in each rectangular region and calculate the edge intensity. Use the Sobel operator to calculate the average edge intensity of all pixels in each region , and set a threshold to determine whether the region needs to be focused sampled.

[0116] When , the region is classified as a dense sampling region, and a first down-sampling factor is used to retain more important light rays, ; when , the region is classified as a sparse sampling region, and only a small number of light rays are needed to indicate whether the region is occupied, and a second down-sampling factor is used to down-sample the sparse sampling region, , reducing the number of light rays and achieving efficient optimization.

[0117] In this way, the sampling density can be flexibly adjusted in different regions, ensuring that important light rays are captured and redundant calculations are effectively reduced. Finally, the light rays from different down-sampling regions are combined and input into the MLP for training. Different resolution light rays are used to reconstruct the hidden space and preserve the original resolution of the structure, achieving high-resolution reconstruction at a lower computational cost.

[0118] In the double-reflection non-visual range system, the positions of the light spot and the shadow are changed by the galvanometer, and then the shadow image and position data are obtained. The relay wall is repeatedly scanned to obtain time-continuous dynamic data. Due to the high-speed scanning of the galvanometer, the data acquisition speed is much higher than the movement speed of the hidden object, so the dynamic change of the object can be ignored in a scanning period. In the dynamic reconstruction task, the galvanometer scanning and data acquisition operations are continuously performed, and the newly sampled shadow and coordinates are input into the MLP to realize fast reconstruction. The full sampling method needs a large number of training iterations for image reconstruction in one frame of a scanning period, which is too slow for dynamic reconstruction.

[0119] The present application designs a differential sampling method to reconstruct dynamic scenes. In view of the small difference between consecutive frames, the changed light is extracted from the consecutive frames captured from the same light spot position to model the dynamics of the hidden object. On the basis of the complete reconstruction of the first frame, only the dynamic part is reconstructed in the subsequent frames, thereby realizing fast reconstruction.

[0120] Due to the catastrophic forgetting mechanism of the neural network, when encountering a new task in the continuous learning process, the neural network often significantly forgets the previously learned information. In order to prevent the neural network from forgetting the static part in the hidden space when encountering a new frame, an image pyramid strategy is used to supplement the global feature information.

[0121] As shown in Figure 7 From the perspective of the same light spot, the frame difference in the dynamic scene mainly reflects on the displacement of the person's arm, and the rest part does not change significantly. Based on this feature, the neural network should mainly update the changed light to realize efficient reconstruction in the dynamic scene.

[0122] The binary shadow image is down-sampled at different scales. The low-resolution down-sampled images are stacked and fused to create multi-scale shadow maps and construct multi-scale light sets. At the same time, the changed light between consecutive frames is extracted and stored in the differential light set. Finally, the obtained multi-scale light set and differential light set are mixed and input into the MLP for training.

[0123] By using the multi-scale information containing the changed light and the overlapping light in the new and old frames to train the network, the network can quickly reconstruct the dynamic scene.

[0124] The present application creates simulation data sets and real data sets to evaluate the proposed method. These data sets contain static and dynamic scenes under various conditions.

[0125] Using Blender, we created a simulated dataset based on the scale of a real scene. We reproduced the ambient lighting and other interfering factors found in real scenes in the simulated scene. Based on the surface reflection characteristics of objects, we simulated the light propagation process in real scenes, making the simulated dataset more realistic as a reference. We recorded the spatial position of the light spot and the shadow binary image pairs. Each frame contained 25 pairs of light spot and shadow binary image pairs, and the resolution of the shadow binary image was 2.5. .

[0126] In order to further verify the light sampling method in practice, the present invention constructed a typical real scene to ensure the performance of the method in the real world. The experimental verification platform is a standard office space ( To simulate real-world occlusion, a 4m planar barrier was used for geometric alignment to create controllable line-of-sight occlusion. In the real-world dataset, the number of binary image pairs of light spots and shadows per frame was also 25. The specifications of the laser scanning galvanometer are as follows: 500mW continuous wave output at 520nm spectral emission; bidirectional angular resolution in the axial plane ; The reflective coating maintains 99% reflectivity across the entire operating spectrum.

[0127] The proposed dual-reflection non-viewing area reconstruction network was implemented using PyTorch and an NVIDIA GeForce RTX 3090 GPU. For each pixel in the binary segmentation map, ray sampling is performed using 64 hierarchical coarse samples, followed by 64 importance-weighted fine samples. This sampling process is supported by a hierarchical neural architecture consisting of a 4-layer coarse network for initial estimation and a 6-layer fine network for detail refinement. During training, to balance memory efficiency and convergence stability, the ray batch size is maintained at 1024 for each optimization iteration.

[0128] Two static scenes are selected to evaluate the accuracy of the dual-reflection non-line-of-sight imaging framework and are analyzed at different resolutions ( and ) and different scene complexities (single object and multi-object), the reconstruction accuracy of multi-resolution sampling and full sampling methods are benchmarked.

[0129] For the two sampling methods, the scene reconstruction accuracy is compared and analyzed at different training iterations. Figure 8 As shown, a residential interior scene containing a chandelier, sofa, coffee table and chairs is taken as an example. In the low-resolution scene ( ), multi-resolution sampling reduces the number of rays by 79.5% compared to full sampling while maintaining comparable reconstruction accuracy.

[0130] The low-resolution scene test demonstrates the superior efficiency of the multi-resolution method, which requires 1-2 times more training iterations to achieve comparable accuracy to the full-sampling method. In the high-resolution experiment, ), the multi-resolution sampling method only uses 2.6% of the number of rays required by full sampling. The full-sampling method cannot reconstruct the structure of the hanging lamp, while the multi-resolution method achieves a detailed reconstruction of the chair after 1000 training iterations and completes the high-precision recovery of the hanging lamp at 2000 iterations. At this stage, the accuracy of the full-sampling method has decreased significantly. The root mean square error (RMSE) of the entire scene is calculated, which quantifies the depth deviation in the reconstructed three-dimensional geometric structure. The depth deviation of the full-sampling method in the low-resolution scene is 4.2%, and in the high-resolution scene it is 4.1%, while the multi-resolution method achieves a depth deviation of only 2.6% at both resolutions.

[0131] The experimental results show that the multi-resolution method achieves high reconstruction accuracy at all test resolutions, and its ability to enhance details is particularly significant in high-resolution scenes. At a simulated scene size of 10m, the absolute depth deviation corresponding to the multi-resolution method is 0.26m, which is significantly better than the 0.42m deviation of the full-sampling method.

[0132] To evaluate the sampling performance under different scene complexities, the invention sets up single-target and multi-target scenes, as shown in Figure 9 Quantitative verification shows that in single-target scenes, the multi-resolution method achieves: (1) a 98.3% reduction in the number of rays (important rays account for 79.1%); (2) in the reconstruction of slender objects, the structure preservation is enhanced, while the full-sampling method cannot recover such structures.

[0133] Under complex multi-target conditions, the optimized method maintains a 98.5% reduction in rays while achieving an important ray utilization rate of 68.5%. After 500 training iterations, the multi-resolution sampling demonstrates complete scene reconstruction, and achieves high-precision output at 1000 iterations. In contrast, the full-sampling method requires four times the computational effort (2000 iterations) to achieve similar visual effects.

[0134] In single-target scenes, the final depth deviation of the full-sampling method is 4.5%, while the multi-resolution method reduces this deviation to 4.2%. For multi-object scenes, the final depth deviation of the full-sampling method is higher at 6.8%, while the multi-resolution method achieves a lower deviation of 3.4%.

[0135] The multi-resolution sampling method increases the proportion of important rays in the total number of rays through zonal down-sampling, which helps the network recover detailed features and achieve accurate reconstruction using fewer rays, greatly reducing computational effort.

[0136] Set up a moving person to test the reconstruction efficiency in dynamic scenes. Figure 10 As shown, in a moving sequence, frames are sampled at intervals of one and five frames to simulate slow and fast motion, respectively. In slow motion, the smaller inter-frame differences enable efficient reconstruction via differential ray sampling. While the full sampling approach exhibits a depth bias of 2.7%, the differential sampling approach reduces this bias to 1.9%. At a scene scale of 10 meters, the absolute depth bias is 0.19 meters. The differential sampling approach achieves accurate reconstruction of moving objects within 100 training iterations, outperforming the full sampling approach, which requires 500 iterations to achieve the same intermediate result. In fast motion, there is no spatial overlap between the sampled frames. The differential sampling approach completes reconstruction within 150 training iterations, while the full sampling approach still produces blurry reconstructions after 500 iterations. Compared to the full sampling approach's depth bias of 3.7%, the differential sampling approach reduces the bias to 3.3%, resulting in a relative improvement of 10.8% in reconstruction accuracy.

[0137] Full sampling methods cannot complete training for the current ray batch within a limited number of iterations, resulting in information loss. Differential sampling methods extract dynamic difference rays, preserving the reconstructed structure while updating the dynamic region. This strategy achieves robust dynamic reconstruction and prevents network forgetfulness by fully preserving features.

[0138] This result shows that the sampling method of the present invention has a significant advantage in slow motion with large correlation between consecutive frames. At the same time, for fast motion scenes, it also improves the reconstruction accuracy of dynamic scenes by reducing the number of training iterations required.

[0139] Initialization success is crucial to scene reconstruction efficiency. When processing a large number of rays, it becomes difficult for the network to build an initial model from the chaotic first batch of inputs. The initialization phase, which includes parameter assignment, data preparation, and model configuration, requires significant time and resources. Initialization failures require re-initialization, significantly reducing reconstruction efficiency. In dynamic reconstruction, an already reconstructed model may become corrupted during data updates.

[0140] Successful initialization depends heavily on the characteristics of the first batch of rays. The random initial rays of full sampling methods often lead to insufficient information and can also cause model collapse. Multi-resolution sampling methods effectively alleviate this limitation through two mechanisms: (1) ray number optimization: targeted reduction of the total number of rays while increasing the proportion of important rays; (2) multi-scale information supplementation: hierarchical spatial information integration replaces random ray selection.

[0141] like Figure 11As shown in Fig. 6, the initialization success rate of full sampling is extremely low in high-resolution scenarios, which seriously affects the reconstruction efficiency. In contrast, the multi-resolution method maintains a 100% initialization success rate at all resolution scales.

[0142] The reconstruction efficiency of the sampling method is tested using simulated data from a single lamp. As shown in Fig. 7, in the Figure 12 resolution, the multi-resolution method reduces the number of light rays from 6.56 million to 156,000, a reduction of 97.6%. Notably, 96.3% (1.51 million) of these sampled light rays are important light rays, significantly improving sampling efficiency. It is worth mentioning that the reconstruction quality achieved by the multi-resolution sampling within 15 seconds exceeds the result of the full sampling after 60 seconds of training, achieving a 4-fold acceleration in static scene reconstruction. In resolution scenarios, similar improvements are also observed, with the multi-resolution method effectively reducing the total number of light rays while increasing the information density.

[0143] The multi-resolution sampling method exhibits two different operating modes: in low-resolution conditions, the total number of light rays limits the training throughput, and the multi-resolution method prioritizes light rays to improve sampling efficiency; as the resolution increases, the computational advantage brought by reducing the number of light rays gradually dominates.

[0144] The multi-resolution method can reduce the total number of light rays by one to two orders of magnitude. By reasonably selecting light rays and enhancing information density, it reduces unnecessary repeated training iterations, significantly reduces computational cost, and improves reconstruction efficiency.

[0145] To evaluate the sampling performance in dynamic environments, two different scene configurations are proposed: (1) a multi-target dynamic scene containing a fast-walking pedestrian and a stationary person with arm movement; (2) a static background dynamic scene highlighting isolated arm movement in a fixed environment. Figure 13 Three advantages are demonstrated: (1) reducing temporal artifacts: the full sampling method cannot timely capture the walking pedestrian after data update is completed; in contrast, the differential sampling method achieves complete motion trajectory reconstruction without frame loss; (2) accelerating convergence: the differential sampling method completes detailed scene reconstruction within 10 seconds, which is 6 times faster than the full sampling method, reducing training time by 83% without sacrificing spatio-temporal accuracy; (3) enhancing motion sensitivity: through differential light ray priority, the differential method allocates 78% of computational resources to the changing area, making the deformation parameter convergence speed 92% faster than the full sampling method.

[0146] The differential light ray mechanism achieves accurate detection of scene spatial disturbances through spatio-temporal gradient analysis, thereby optimizing the allocation of computational resources and accelerating scene reconstruction after dynamic changes. ​

[0147] To test the performance of the sampling method in typical static and dynamic scenarios, the present application builds an experimental environment in real scenes. The experimental scene is set in an indoor space with a distance of 5.22 meters between the two walls. An office scene is carefully arranged in the hidden area, which includes hidden objects of different volumes and shapes. The subsequent dynamic and static scenes are both based on this.

[0148] The experimental results of the real static scene are shown in Figure 14 . The resolution of the shadow binary image used for reconstruction is , and the number of images is 64. In this experiment, the full sampling method generates 1100 million rays, while the multi-resolution method reduces the number of rays to 4% of the original and increases the density of important rays to 97%. The multi-resolution sampling framework achieves a reconstruction accuracy comparable to the full sampling 2000 iteration benchmark in only 500 training cycles. The multi-resolution method successfully resolves key structural components, including the geometry of the table legs and the contour of the ergonomic chair.

[0149] The experimental results of the real dynamic scene are shown in Figure 15 . This dynamic scene serves as a controlled experimental framework to evaluate two key capabilities: (1) the ability to integrate real-time dynamic information added to the scene; (2) the ability to maintain the stability of the baseline geometry during continuous updates. To establish consistency in the experiment, both sampling methods are initialized using the same pre-trained scene representation to ensure standardization in the initialization of the data update sequence.

[0150] After the input of new data, the differential sampling method quickly updates the model according to the subsequent data and immediately reconstructs the contour of the newly added personnel, quickly completing the fine reconstruction of the body contour. In contrast, the full sampling method has difficulty in immediately reconstructing the newly appearing personnel in the same scene.

[0151] The multi-resolution sampling method has the ability to continuously refine the static background that has not been fully reconstructed before. By performing detailed modeling, it accurately reconstructs missing structural components such as table legs, while gradually enhancing other static background areas, thereby significantly improving the reconstruction accuracy.

[0152] The present application proposes an efficient dual-reflection non-visual domain imaging framework that combines neural implicit representation with an optimized ray sampling strategy. By integrating neural radiance field (NeRF) and shadow binary image, high-precision reconstruction of hidden scenes is achieved using cumulative transmittance modeling.

[0153] The present invention has three key improvements: First, a low-cost galvanometer-based laser scanning system is developed for capturing shadow images. Compared with time-of-flight (TOF) transient imaging methods, this system significantly reduces hardware complexity while retaining the ease of calibration. At one-fifth the cost of traditional non-line-of-sight devices, the system achieves a portable and deployable scene configuration. Second, the present invention's sampling optimization method improves computational efficiency through ray selection, reducing the amount of data required by 1-2 orders of magnitude compared to full sampling methods. This innovation enables large-scale environments ( ) with a depth deviation of only 2%. Third, the framework achieves rapid dynamic scene reconstruction through high-frequency data acquisition, achieving low latency in reflecting dynamic scene changes and completing the reconstruction of moving objects.

Claims

1. A fast two-reflection non-visual imaging method based on important ray screening, characterized in that, The application relates to a method for three-dimensional reconstruction of a hidden space in a non-line-of-sight (NLOS) scene. Step 1: calibrate the sampling system based on the NLOS scene to determine the size of the space to be reconstructed; Step 2: change the laser light path to illuminate the hidden space through a galvanometer system, obtain a shadow image generated by active illumination using a camera, process the shadow image based on the calibration information to obtain the pose and binary shadow image under three-dimensional coordinates; Step 3: model the light rays based on the double-reflection NLOS scene, and build a three-dimensional reconstruction network for the shadow scene, specifically: Based on the principle of neural radiation field, use a multi-layer perceptron to implicitly model the light rays and hidden objects; construct light rays according to the position of the light spot and the shadow image, and sample the three-dimensional coordinates along each light ray; the MLP is used to predict the possibility of the existence of objects at three-dimensional coordinates; the three-dimensional structure is reconstructed by accumulating the possibility of existence, forming a self-supervised learning framework for implicitly modeling the hidden space; The probability of the existence of hidden objects is expressed as cumulative transmittance It describes the light in the interval The probability of light propagating without being blocked; The expected cumulative transmittance is expressed as ,in Indicates the distance from the surface of the object to the relay wall, represents the light spot position, and t represents the light propagation time; In the range between near boundary and far boundary , use to estimate and calculate cumulative transmittance : ; In an imaging task, input spot positions , cumulative transmittance at three-dimensional coordinates in a hidden space can be queried ; continuous implicit neural shadow field is represented as: ; From spot position Construct rays corresponding to the shadow image and trace them to the projection wall; use a hierarchical sampling method of discrete samples to estimate the cumulative transmittance The integral result: ; wherein denotes the distance between samples between adjacent time instants denotes the distance between samples between adjacent time instants denotes the cumulative transmittance, denotes the opacity; The shadow image is binarized to obtain the light distribution through the hidden object; the pixels in the bright area are marked as "1", indicating that the light passes through the hidden scene; the pixels in the dark area are marked as "0", indicating that the light is blocked; the black background indicates that the light is not blocked when passing through the unknown space, and the white foreground indicates that the light is blocked by the object, thereby forming a shadow on the wall; the optimization process of the network is supervised by calculating the cumulative transmittance; the loss function of the cumulative transmittance is represented as: ; wherein, is a set of rays for each batch during training, representing the rays participating in the training; Finally, the depth of the hidden object is calculated from multiple viewpoints by numerically integrating N rays : ; Step 4: sample the binary shadow image obtained in step 2 using a multi-resolution sampling method based on edge intensity; Step 5: input the sampled binary shadow image and the pose in step 2 into the three-dimensional reconstruction network in step 3 for training; Step 6: repeat step 2, dynamically sample the binary shadow image obtained in step 2 using a differential light sampling method, and repeat step 5; Step 7: repeat step 6 to realize three-dimensional reconstruction of the NLOS dynamic scene.

2. The fast two-reflection non-foveal imaging method based on important ray screening according to claim 1, characterized in that, The step 1 is specifically: The sampling system is calibrated based on the NLOS scene; the center point of the relay wall is , the center point of the projection wall is ; the midpoint of the line segment connecting and is taken as the origin , and a three-dimensional space coordinate system is established; the distance between the two walls is , the width of the relay wall is W, and the height is .

3. The fast two-reflection non- visibility imaging method based on important ray screening according to claim 2, characterized in that, The step 2 is specifically: In order to realize the reconstruction of hidden objects, a double reflection non-visual domain imaging system is constructed, which is composed of a camera and a laser scanning galvanometer system; the incident direction of laser is changed by the laser scanning galvanometer system, and the light spots with positions of 、 、 are generated on the relay wall; the changing shadows on the projection wall are captured by the camera, so that a group of shadow images reflecting the characteristics of hidden objects are obtained; The data processing process is as follows: (1) The perspective transformation is performed on the light spot under the angle of squint to obtain the light spot under the angle of orthovision; the coordinates of four corner points on the original image are set as , wherein ; The coordinates of the four corresponding points on the target plane are wherein: ; with respectively the width and height of the pixels of the sample picture; The homogeneous coordinates of the perspective transformation are: ; For each pair of corresponding points , wherein is the spot coordinate in the oblique view before the perspective transformation, is the spot coordinate in the orthographic view after the perspective transformation; each value of the perspective transformation matrix is solved according to the above transformation formula: , , By performing perspective transformation on all pixel points in the collected image, the corrected spot position is obtained;​ (2) Traverse all pixels in the grayed shadow image to find the maximum gray value corresponding to the pixel coordinates , ; convert the pixel coordinates , into three-dimensional coordinates , and the conversion formula is as follows: ; Assuming that N spot corresponds to generate N shadow image, the corresponding spot coordinates are expressed as ; (3) Perform the same perspective transformation operation on the shadow image to obtain the corrected shadow image; perform histogram equalization processing on the shadow image, and use an adaptive threshold for binaryzation operation.

4. The fast two-reflection non- visibility imaging method based on important ray screening according to claim 1, characterized in that, The step 4 is specifically: There are N light spots on the relay wall , , ,..., ), and the light rays are emitted from the N light spots to the projection wall; the resolution of the collected shadow binary image is ; the light rays emitted from the light spots on the relay wall to the projection wall are , wherein , , so that a total of light rays are obtained; The light rays containing structural information are defined as important light rays; a multi-resolution sampling method based on edge intensity is proposed, different down-sampling factors are used in different regions, and the regions are divided into sparse sampling regions and dense sampling regions according to the edge intensity; For a shadow binary image with a size of , it is divided into rectangular regions, each with a height of and a width of ; the edges in each rectangular region are detected and their edge strength is calculated; The average edge intensity of all pixels in each region is calculated using Sobel operators and a threshold is set to determine whether the region is to be focus sampled; When this region is classified as a dense sampling region, using a first down-sampling factor to retain more important rays, ; when this region is classified as a sparse sampling region, using a second down-sampling factor down-sampling the sparse sampling region, ; In this way, the sampling density is adjusted in different regions; the light rays in each region after different down-sampling processing are combined and input into the MLP for training; the hidden space is reconstructed using light rays of different resolutions, and the original resolution of the structure is preserved.

5. The fast two-reflection non- visibility imaging method based on important ray screening according to claim 4, characterized in that, The step 6 is specifically: In the double-reflection non-visual domain system, the position of the light spot and the shadow is changed through the galvanometer to obtain the shadow image and position data; the relay wall is repeatedly scanned to obtain time-continuous dynamic data; In the dynamic reconstruction task, the galvanometer scanning and data acquisition operations are continuously performed, and the newly sampled shadow and coordinates are input into the MLP; A differential sampling method is used to reconstruct the dynamic scene; the changing light rays are extracted from the continuous frames captured from the same light spot position to model the dynamics of the hidden objects; on the basis of the complete reconstruction of the first frame, only the dynamic part is reconstructed in the subsequent frames; An image pyramid strategy is used to supplement the global feature information; The shadow binary image is down-sampled in different scales; the down-sampled images are stacked and fused to create a multi-scale shadow image and build a multi-scale light set; at the same time, the light changes between continuous frames are extracted and stored in a differential light set; finally, the obtained multi-scale light set and the differential light set are mixed and input into an MLP for training; By using multi-scale information containing changing light and overlapping light in new and old frames to train the network, the network can quickly reconstruct a dynamic scene.

6. An electronic device comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, characterized in that, The processor implements the steps of the method of any one of claims 1-5 when executing the program.

7. A computer-readable storage medium having stored thereon a computer program, characterized in that The program is executed by the processor to implement the steps of the method of any one of claims 1-5.

8. A computer program product comprising a computer program, characterized in that, The computer program is executed by the processor to implement the steps of the method of any one of claims 1-5.

Citation Information

Patent Citations

  • Digital twin modeling method and system based on SAM large model and NeRF

    CN117671138A

  • View-around scene three-dimensional reconstruction method based on implicit representation

    CN118864734A