A method and system for reconstruction of depth images of a time-of-flight camera

By decomposing and processing iToF measurement data using neural networks, the problem of insufficient depth measurement accuracy caused by MPI in ToF cameras is solved, and efficient real-time depth image reconstruction is achieved.

CN119672083BActive Publication Date: 2026-04-17TRIGIANTS TECH
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
TRIGIANTS TECH
Filing Date
2024-12-05
Publication Date
2026-04-17

AI Technical Summary

Technical Problem

Existing ToF cameras suffer from insufficient depth measurement accuracy in complex scenes due to multipath interference (MPI). Traditional algorithms are computationally intensive and difficult to meet the needs of real-time applications, while neural networks cannot effectively process high-dimensional data.

Method used

By decomposing the iToF measurement data, and using a neural network-based prediction model and a backscattering model, we can predict low-dimensional vector representations and reconstruct high-dimensional data, thereby reducing computational complexity and removing MPI interference.

Benefits of technology

It improves the accuracy and computational efficiency of depth imaging, is suitable for real-time depth image reconstruction, and reduces the computational resource requirements.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119672083B_ABST
    Figure CN119672083B_ABST
Patent Text Reader

Abstract

This invention belongs to the field of time-of-flight cameras and provides a method and system for reconstructing depth images from a time-of-flight camera. The method includes: acquiring iToF measurement data from an iToF camera; expressing the iToF measurement data as an integral MPI describing equation; converting the MPI describing equation into a discrete form; inputting the iToF measurement data into a prediction model, which predicts the iToF measurement data as a low-dimensional vector of a specified reflected light order; inputting the low-dimensional vector into a backscattering model, which converts the low-dimensional vector into an estimated backscattering vector; calculating new iToF measurement data based on the measurement matrix and the backscattering vector obtained by mapping; and reconstructing the depth image of the iToF camera based on the iToF measurement data. This invention reduces computational complexity by decomposing the ToF data, using a neural network-based prediction model to predict the low-dimensional vector representation of high-dimensional iToF data, and restoring the low-dimensional vector to high-dimensional ToF data using a fixed backscattering model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of time-of-flight cameras, and more particularly to a method and system for reconstructing depth images from a time-of-flight camera. Background Technology

[0002] In the field of Time-of-Flight (ToF) imaging, there are indirect ToF (iToF) cameras and direct ToF (dToF) cameras. iToF cameras suffer from multipath interference (MPI), which has always been a major problem affecting depth measurement accuracy. MPI occurs when the same light signal reaches the camera after multiple reflections, leading to errors in the measurement data. This error makes it difficult for traditional ToF cameras to obtain accurate depth information in complex scenes, especially in situations involving reflection or light refraction. Therefore, effectively eliminating or reducing the impact of MPI on measurement results has been an important research direction in ToF imaging technology.

[0003] To improve the accuracy of depth imaging, many existing technologies attempt to address the MPI problem using various algorithms and techniques. Traditional techniques include using ToF camera data with single or multiple modulation frequencies and compensating for MPI through iterative algorithms, linear programming, or deterministic methods. While these methods are theoretically effective, they are often computationally intensive and slow in practical applications, making them difficult to meet the demands of real-time applications. In particular, ToF cameras acquire high-dimensional data, and traditional neural networks cannot directly predict another high-dimensional data from input high-dimensional data. This is computationally very complex and difficult to implement, limiting the application of neural networks in the field of ToF imaging. Summary of the Invention

[0004] In view of the above technical problems, the present invention provides a method and system for reconstructing depth images from a time-of-flight camera. After decomposing the ToF data, a prediction model based on a neural network is used to predict the low-dimensional vector representation of the high-dimensional iToF data. A fixed backscattering model is used to restore the low-dimensional vector to the high-dimensional ToF data. This method achieves high performance while maintaining low computational complexity.

[0005] Other features and advantages of the invention will become apparent from the following detailed description, or may be learned in part by practice of the invention.

[0006] According to one aspect of the present invention, a method for reconstructing depth images from a time-of-flight camera is disclosed, the method comprising:

[0007] The iToF measurement data of the iToF camera is acquired, and the iToF measurement data is expressed as an integral MPI description equation. The MPI description equation is then converted into a discrete form with N time steps, so that the iToF measurement data can be represented as a measurement matrix and backscatter vector based on a single modulation frequency. The operation is repeated based on multiple different modulation frequencies to obtain the measurement matrix expressed by the iToF measurement data at multiple modulation frequencies and the stack of measurement values ​​in the complex domain.

[0008] The stack of iToF measurement data is input into the prediction model, which is trained based on the nonlinear function of a neural network. The prediction model is used to predict the iToF measurement data as a low-dimensional vector of a specified order of reflected light. The low-dimensional vector represents the amplitude and path length of the reflected light from the iToF measurement data.

[0009] The low-dimensional vector is input into the backscattering model, and the low-dimensional vector is converted into a predicted backscattering vector. The backscattering model is based on a generative adversarial network or a variational autoencoder, and the backscattering model is configured as a mapping from the low-dimensional vector to the backscattering vector.

[0010] Based on the measurement matrix and the backscatter vector obtained by mapping, new iToF measurement data is calculated, and the depth image of the iToF camera is reconstructed based on the iToF measurement data.

[0011] Furthermore, the MPI description equation is a superposition of multiple complex phasors, which are used to represent the reflection signal of each optical path in the iToF measurement data, including the amplitude and phase information of the reflected light.

[0012] Furthermore, the measurement matrix includes modulation frequency, light time step, phase information, scene response, and sensor sensitivity function.

[0013] Furthermore, when training the prediction model, a set of original iToF measurement data at different frequencies is used as input, and the real backscatter vector after removing reflected light interference is used as a label to train the prediction model to obtain a low-dimensional vector that predicts the order of the reflected light.

[0014] Furthermore, the order of the reflected light is specified as 2, including the first-order reflected light and the second-order reflected light, and the label is a backscattering vector containing only the information of the first-order and second-order reflected light.

[0015] Furthermore, during the training of the prediction model, the training parameters are optimized based on the difference between the prediction results and the actual data, wherein:

[0016] The direct difference between the predicted iToF measurement data and the actual iToF measurement data is calculated, and the training parameters are optimized based on the difference.

[0017] The predicted backscatter vector and the true backscatter vector are treated as two different probability mass functions, and the statistical distance between them is measured based on the EMD function to obtain the difference between the predicted backscatter vector and the true backscatter vector. The training parameters are then optimized based on the difference.

[0018] Furthermore, after mapping the backscatter vector or when reconstructing the depth image of the iToF camera based on the iToF measurement data, noise removal is performed on the mapped backscatter vector or the reconstructed depth image based on bilateral filtering.

[0019] Furthermore, the network structure of the prediction model includes:

[0020] An input layer is used to input the stack of iToF measurement data in the complex domain and convert the complex form of the iToF measurement data into a combination of the real and imaginary parts.

[0021] Multiple convolutional layers are used to extract low-level features from high-level features on the input iToF measurement data. The first convolutional layer incorporates a weight kernel, which is used to retrieve information from each pixel.

[0022] A fully connected layer is used to map the features extracted by the convolutional layer to the low-dimensional vector;

[0023] An output layer is used to output the low-dimensional vector.

[0024] According to another aspect of the present invention, a system for reconstructing depth images from a time-of-flight camera is disclosed, the system comprising:

[0025] The acquisition module is used to acquire iToF measurement data from the iToF camera, express the iToF measurement data as an integral MPI description equation, convert the MPI description equation into a discrete form with N time steps, so as to represent the iToF measurement data as a measurement matrix and backscatter vector based on a single modulation frequency. The operation is repeated based on multiple different modulation frequencies to obtain the measurement matrix expressed by the iToF measurement data at multiple modulation frequencies and the stack of measurement values ​​in the complex domain.

[0026] The prediction module is used to input the stack of measurement values ​​of the iToF measurement data into the prediction model. The prediction model is trained based on the nonlinear function of the neural network. The prediction model is used to predict the iToF measurement data as a low-dimensional vector of a specified order of reflected light. The low-dimensional vector represents the amplitude and path length of the reflected light of the iToF measurement data.

[0027] A mapping module is used to input the low-dimensional vector into the backscattering model and convert the low-dimensional vector into an estimated backscattering vector. The backscattering model is based on a generative adversarial network or a variational autoencoder. The backscattering model is configured as a mapping from the low-dimensional vector to the backscattering vector.

[0028] The reconstruction module is used to calculate new iToF measurement data based on the measurement matrix and the backscatter vector obtained by mapping, and to reconstruct the depth image of the iToF camera based on the iToF measurement data.

[0029] The technical solution of the present invention has the following beneficial effects:

[0030] This invention utilizes a deep learning neural network to accurately remove multipath interference caused by reflection paths from iToF measurement data. After generating a low-dimensional vector, it is restored by a backscattering model, which solves the limitation of neural networks, reduces computational complexity, significantly reduces the required computing resources, and improves the training and inference efficiency of the model. It is especially suitable for real-time depth image reconstruction.

[0031] This invention calculates iToF measurement data under multi-frequency modulation, making full use of the MPI feature variations caused by frequency changes, thereby further improving the accuracy and measurement range of depth reconstruction. Attached Figure Description

[0032] Figure 1 This is a flowchart illustrating a method for reconstructing a depth image from a time-of-flight camera, as described in an embodiment of this specification.

[0033] Figure 2 This is a structural block diagram of a depth image reconstruction system for a time-of-flight camera, as described in an embodiment of this specification. Detailed Implementation

[0034] Exemplary embodiments will now be described more fully with reference to the accompanying drawings. However, these exemplary embodiments can be implemented in many forms and should not be construed as limited to the examples set forth herein; rather, these embodiments are provided to make the invention more comprehensive and complete, and to fully convey the concept of the exemplary embodiments to those skilled in the art. The described features, structures, or characteristics may be combined in any suitable manner in one or more embodiments. In the following description, numerous specific details are provided to give a full understanding of embodiments of the invention. However, those skilled in the art will recognize that the technical solutions of the invention may be practiced with one or more of these specific details omitted, or other methods, components, apparatus, steps, etc., may be employed. In other instances, well-known technical solutions are not shown or described in detail to avoid obscuring various aspects of the invention.

[0035] Furthermore, the accompanying drawings are merely illustrative of the invention. The same reference numerals in the drawings denote the same or similar parts, and therefore repeated descriptions of them will be omitted. Some block diagrams shown in the drawings are functional entities and do not necessarily correspond to physically or logically independent entities. These functional entities can be implemented in software, in one or more hardware modules or integrated circuits, or in different network and / or processor devices and / or microcontroller devices.

[0036] In one implementation, such as Figure 1 As shown, this specification provides a method for reconstructing depth images from a time-of-flight camera. The method can be executed by a computer, server, tablet computer, etc. Specifically, the method may include the following steps S101~S104:

[0037] In step S101, iToF measurement data from the iToF camera is acquired, and the iToF measurement data is expressed as an integral MPI description equation. The MPI description equation is then converted into a discrete form with N time steps, so that the iToF measurement data can be represented as a measurement matrix and backscatter vector based on a single modulation frequency. The operation is repeated based on multiple different modulation frequencies to obtain the measurement matrix expressed by the iToF measurement data at multiple modulation frequencies and the stack of measurement values ​​in the complex domain.

[0038] In Time-of-Flight (ToF) imaging, the main idea for acquiring depth information is based on the fact that the speed of light is constant. In dToF, a very short light pulse is emitted towards the scene; the pulse is reflected and eventually received by the sensor, and depth is inferred from the time of flight of the light pulse. iToF uses the same principle but employs a different type of illumination signal, using periodic light modulation and adjusting the depth based on the phase shift between the incident light and an internal reference signal. To obtain depth information, the specific relationships are as follows:

[0039] ;

[0040] in, It is the modulation frequency. It's the speed of light.

[0041] In practice, the transmitter sends a signal with a modulation frequency to the scene. modulation signal At the sensor end, reflected light With sensor sensitivity function As shown in the following equation, the phase of the light measured by ToF will be shifted by a factor. ; camera's raw measurements It is the result of these related operations, assuming If it is a sine wave, then the reflected light can be expressed as an attenuated and delayed version of the original signal:

[0042] ;

[0043] Among them, delay Expressed in terms of phase displacement, if we consider the sensor's sensitivity function as:

[0044] ;

[0045] Assuming the light reflects only once in the scene, then:

[0046] ;

[0047] in, It's the integration time. It is the phase shift applied to sensor sensitivity.

[0048] Due to common sense We can obtain the closed-form solution for the integral:

[0049] ;

[0050] in, A is the strength of the ToF signal, and A is its amplitude. This phase shift is caused by scene depth. It needs to be restored. A Only the sensor sensitivity needs to be adjusted. 4 known phase shifts downsampling We can obtain:

[0051] ;

[0052] ;

[0053] ;

[0054] Since phasor notation is a very convenient way to represent the sinusoidal correlation function, the following formula can be used to express the raw measurement data:

[0055] ;

[0056] in, It is the amplitude. It is the phase of the original sine wave function.

[0057] The above only considers the case where the iToF signal is reflected only once in the scene. However, in real-world scenarios, light reflects multiple times, resulting in multiple reflected beams reaching the same pixel, a phenomenon known as multipath interference (MPI). In this case, the description can be extended by assuming that the final iToF signal is the sum of different interference signals. Each interference signal can be represented by a phasor; that is, the MPI description equation is a superposition of multiple complex phasors. These complex phasors represent the reflected signal of each light path in the iToF measurement data, including the amplitude and phase information of the reflected light. Since sinusoidal signals of the same frequency and their phasor representations are closed under addition, the integral form of the MPI description equation for the iToF measurement data can be expressed as:

[0058] ;

[0059] in, It is the maximum flight time of the interfering light rays under consideration. The backscattering distribution function describes the intensity of interfering rays based on their time of flight. The aforementioned MPI phenomenon is a non-zero mean error in iToF depth measurement, typically leading to overestimation of depth. A key characteristic of MPI distortion is its dependence on the geometry of the scene under consideration, which affects the backscattering distribution function. This has a significant impact. The MPI (Multi-Path Integrator) effect occurs because a single ray of light travels through multiple paths after reflection within a scene. Existing iToF cameras ignore these MPI paths during integration, leading to inaccurate results. If multi-path integration can be avoided, and instead the light intensity arriving at the sensor at each moment can be captured, the effects caused by the varying arrival times of incident light reflection can be isolated.

[0060] Deep neural networks (DNNs) often struggle with high-dimensional data, and the backscattered vectors composed of emitted light are high-dimensional data. Specifically, a backscattered vector can contain thousands of data points in the time direction, while iToF measurement data contains only a small number of data points, making training deep networks extremely difficult. Therefore, to make the data applicable to deep neural networks, the integral form of the MPI description equation can be simplified by sampling the integration time interval into N time steps to represent the discrete form of the MPI description equation:

[0061] ;

[0062] Specifically, the impulse response of the scene can be isolated from the backscatter vector. In, and using matrices The measurement model is typically represented by the modulation frequency, the time step of the light, phase information, scene response, and sensor sensitivity function. Furthermore, considering the frequency-dependent distortion due to interference from light rays, M different modulation frequencies are sampled. Variations in frequency cause changes in the distortion pattern, providing additional information about the MPI distribution. This also helps to achieve a longer, unambiguous measurement range while maintaining depth domain accuracy.

[0063] Extending the discrete form of the MPI describing equation yields:

[0064] ;

[0065] in, It is a stack of iToF measurement data in the complex domain at different modulation frequencies, and It is a measurement matrix. Therefore, given a measurement matrix, we can obtain... From the original iToF measurement data, it is hoped that the backscatter vector can be recovered. .

[0066] In step S102, the stack of measured values ​​of the iToF measurement data is input into the prediction model. The prediction model is trained based on the nonlinear function of the neural network. The prediction model is used to predict the iToF measurement data as a low-dimensional vector of a specified order of reflected light. The low-dimensional vector represents the amplitude and path length of the reflected light of the iToF measurement data.

[0067] The prediction model is the core component of the entire neural network, and its input consists of raw iToF measurement data from different modulation frequencies. The output is the predicted value in the corresponding latent domain Z. Specifically, the prediction model is a highly nonlinear function. It has trainable parameters. Receive iToF measurement data It outputs an estimated latent vector Z. Z consists of a vector of a preset dimension, for example, This represents four Z components, specifically the amplitude and path length of the first and second order interference light paths, i.e., the amplitude and path length of the same light signal when it first directly hits the sensor and the second time it is reflected back to the sensor. The amplitude is the intensity, and the path length is the step size, i.e., time.

[0068] In step S103, the low-dimensional vector is input into the backscattering model to convert the low-dimensional vector into an estimated backscattering vector. The backscattering model is based on a generative adversarial network or a variational autoencoder and is configured as a mapping from the low-dimensional vector to the backscattering vector.

[0069] The objective of the backscattering model is to map a latent variable Z to its corresponding backscattering vector. The backscattering model can be expressed as follows: Then, the mapping process can be expressed as:

[0070] ;

[0071] ;

[0072] in, , These are trainable parameters. It is the domain of all backscatter vectors. This can be based on generative adversarial networks or variational autoencoders, as long as it can achieve the mapping from low-dimensional space to high-dimensional data. In practice, considering that the task is MPI denoising, i.e., multipath interference removal, the backscattering model only needs to accurately locate the direct component (i.e., the first peak of the signal). For the second global component, a more concise expression can be used. Therefore, the case of two reflections can be considered, i.e., the first and second reflections usually contain the largest part of the energy in the backscattering vector. Thus, the backscattering model can be defined as a four-dimensional vector... A mapping model to an approximate backscattering vector is used. In the model, the four Z components represent the amplitude and path length of the first and second interference reflected light paths. The backscattering model converts these four values ​​into an approximate backscattering vector, with all values ​​being zero except for two peaks associated with the first and second interference light paths.

[0073] In step S104, new iToF measurement data is calculated based on the measurement matrix and the backscatter vector obtained by mapping, and the depth image of the iToF camera is reconstructed based on the iToF measurement data.

[0074] Among them, by It can be seen that, in Given the prior knowledge, the reflection vector is obtained by mapping. After removing MPI interference Then it can be obtained and utilized. Then, a depth image can be reconstructed.

[0075] In one embodiment, as a supplement to step S102, when training the prediction model, a set of original iToF measurement data at different frequencies is used as input, and the real backscatter vector after removing reflected light interference is used as a label to train the prediction model to obtain a low-dimensional vector that predicts a specified order of reflected light.

[0076] Specifically, the order of the reflected light is specified as 2, including the first-order reflected light and the second-order reflected light, and the label is a backscattering vector that contains only information about the first-order and second-order reflected light.

[0077] In one embodiment, when training the prediction model, the training parameters are optimized based on the difference between the prediction results and the actual data, wherein: the direct difference between the predicted iToF measurement data and the actual iToF measurement data is calculated, and the training parameters are optimized based on the difference; the predicted backscatter vector and the actual backscatter vector are regarded as two different probability mass functions, and the statistical distance between them is measured based on the EMD function to obtain the difference between the predicted backscatter vector and the actual backscatter vector, and the training parameters are optimized based on the difference.

[0078] The discrepancy calculation is used to ensure consistency between the predicted and actual data. Specifically, the measurement loss ensures that the predicted and mapped backscatter vector corresponds to the original input iToF measurement. This is because the measurement matrix... It is known that for any predicted backscatter vector The corresponding calculation can be performed. , Must be consistent with the input measurement data Consistent. The loss function based on mean squared error can be written in the following form:

[0079] ;

[0080] However, in practical applications, this loss is only a soft constraint. To obtain more accurate results, the loss can be reconstructed to ensure the predicted backscatter vector... With the actual truth value Matching. Since the backscatter vector is a sparse, high-dimensional vector, a suitable distance metric needs to be defined. Existing methods such as MSE (mean squared error) or MAE (mean absolute error) are not suitable, so the EMD (Earth Transporter) function is used. The EMD function amplifies the errors between the predicted and actual peak values ​​in terms of time and intensity components, while giving lower weight to zero-value entries.

[0081] Specifically, the two backscattering vectors are treated as two probability mass functions (PMFs), and the statistical distance between them is calculated and located as follows:

[0082] , ;

[0083] in, To avoid division by zero when the predicted result is all zeros, the true backscatter vector can be normalized. Then, the distance between the two distributions is calculated based on the EMD function:

[0084] ;

[0085] in, and The cumulative distribution function is the original distribution. By weighting the above EMD function, we obtain the reconstruction loss defined between the original backscatter vector and the predicted true backscatter vector:

[0086] ;

[0087] in, and It is the cumulative distribution function of the backscatter vector:

[0088] , ;

[0089] And weight Calculated as follows:

[0090] ;

[0091] Among them, window The weights can be set according to the actual situation. Increasing the weights can be done for data with non-zero samples, thereby balancing the importance of the direct component (first reflection) and the global component (second reflection).

[0092] In one embodiment, after mapping the backscatter vector or when reconstructing the depth image of the iToF camera based on the iToF measurement data, noise removal is performed on the mapped backscatter vector or the reconstructed depth image based on bilateral filtering.

[0093] In one embodiment, the network structure of the prediction model includes:

[0094] An input layer is used to input the stack of iToF measurement data in the complex domain and convert the complex form of the iToF measurement data into a combination of the real and imaginary parts.

[0095] Multiple convolutional layers are used to extract low-level features from high-level features on the input iToF measurement data. The first convolutional layer incorporates a weight kernel, which is used to retrieve information from each pixel.

[0096] A fully connected layer is used to map the features extracted by the convolutional layer to the low-dimensional vector;

[0097] An output layer is used to output the low-dimensional vector.

[0098] Based on the same line of thought, such as Figure 2 As shown, an exemplary embodiment of the present invention also provides a depth image reconstruction system from a time-of-flight camera, the system comprising:

[0099] The acquisition module 201 is used to acquire iToF measurement data from an iToF camera, express the iToF measurement data as an integral MPI description equation, convert the MPI description equation into a discrete form with N time steps, so as to represent the iToF measurement data as a measurement matrix and backscatter vector based on a single modulation frequency, and repeat the operation based on multiple different modulation frequencies to obtain the measurement matrix expressed by the iToF measurement data at multiple modulation frequencies and the stack of measurement values ​​in the complex domain.

[0100] The prediction module 202 is used to input the stack of measurement values ​​of the iToF measurement data into the prediction model. The prediction model is trained based on the nonlinear function of the neural network. The prediction model is used to predict the iToF measurement data as a low-dimensional vector of a specified order of reflected light. The low-dimensional vector represents the amplitude and path length of the reflected light of the iToF measurement data.

[0101] The mapping module 203 is used to input the low-dimensional vector into the backscattering model and convert the low-dimensional vector into a predicted backscattering vector. The backscattering model is based on a generative adversarial network or a variational autoencoder. The backscattering model is configured as a mapping from the low-dimensional vector to the backscattering vector.

[0102] The reconstruction module 204 is used to calculate new iToF measurement data based on the measurement matrix and the backscatter vector obtained by mapping, and to reconstruct the depth image of the iToF camera based on the iToF measurement data.

[0103] In the above embodiments, by utilizing deep learning neural networks, multipath interference caused by reflection paths can be accurately removed from iToF measurement data. After generating low-dimensional vectors, these vectors are restored using a backscattering model, overcoming the limitations of neural networks, reducing computational complexity, significantly lowering the required computational resources, and improving the training and inference efficiency of the model. This is particularly suitable for real-time depth image reconstruction. By calculating iToF measurement data under multi-frequency modulation, the changes in MPI features caused by frequency variations are fully utilized, further improving the accuracy and measurement range of depth reconstruction.

[0104] From the above description of the embodiments, those skilled in the art will readily understand that the exemplary embodiments described herein can be implemented by software or by combining software with necessary hardware. Therefore, the technical solutions of the embodiments of the present invention can be embodied in the form of a software product, which can be stored in a non-volatile storage medium (such as a CD-ROM, USB flash drive, portable hard drive, etc.) or on a network, including several instructions to cause a computing device (such as a personal computer, server, terminal device, or network device, etc.) to execute the method according to the exemplary embodiments of the present invention.

[0105] Furthermore, the above figures are merely illustrative representations of the processes included in the method according to exemplary embodiments of the present invention, and are not intended to be limiting. It is readily understood that the processes shown in the above figures do not indicate or limit the temporal order of these processes. Additionally, it is readily understood that these processes may be executed synchronously or asynchronously, for example, in multiple modules.

[0106] Other embodiments of the invention will readily occur to those skilled in the art upon consideration of the specification and practice of the invention disclosed herein. This application is intended to cover any variations, uses, or adaptations of the invention that follow the general principles of the invention and include common knowledge or customary techniques in the art not disclosed herein. The specification and embodiments are to be considered exemplary only, and the true scope and spirit of the invention are indicated by the claims.

Claims

1. A method of reconstruction of a depth image of a time-of-flight camera, characterized in that, The method includes: The iToF measurement data of the iToF camera is acquired, and the iToF measurement data is expressed as an integral MPI description equation. The MPI description equation is then converted into a discrete form with N time steps, so that the iToF measurement data can be represented as a measurement matrix and backscatter vector based on a single modulation frequency. The operation is repeated based on multiple different modulation frequencies to obtain the measurement matrix expressed by the iToF measurement data at multiple modulation frequencies and the stack of measurement values ​​in the complex domain. The stack of iToF measurement data is input into the prediction model, which is trained based on the nonlinear function of a neural network. The prediction model is used to predict the iToF measurement data as a low-dimensional vector of a specified order of reflected light. The low-dimensional vector represents the amplitude and path length of the reflected light from the iToF measurement data. The low-dimensional vector is input into the backscattering model, and the low-dimensional vector is converted into a predicted backscattering vector. The backscattering model is based on a generative adversarial network or a variational autoencoder, and the backscattering model is configured as a mapping from the low-dimensional vector to the backscattering vector. Based on the measurement matrix and the backscatter vector obtained by mapping, new iToF measurement data is calculated, and the depth image of the iToF camera is reconstructed based on the iToF measurement data.

2. The method of claim 1, wherein, The MPI description equation is a superposition of multiple complex phasors, which are used to represent the reflection signal of each optical path in the iToF measurement data, including the amplitude and phase information of the reflected light.

3. The method of claim 1, wherein, The measurement matrix includes modulation frequency, light time step, phase information, scene response, and sensor sensitivity function.

4. The method for reconstructing depth images from a time-of-flight camera according to claim 1, characterized in that, When training the prediction model, a set of original iToF measurement data at different frequencies is used as input, and the real backscatter vector after removing reflected light interference is used as a label to train the prediction model to obtain a low-dimensional vector that predicts the order of the reflected light.

5. The method for reconstructing depth images from a time-of-flight camera according to claim 4, characterized in that, The specified order of the reflected light is 2, including the first-order reflected light and the second-order reflected light. The label is a backscatter vector that contains only information about the first-order and second-order reflected light.

6. The method for reconstructing depth images from a time-of-flight camera according to claim 4, characterized in that, When training the prediction model, the training parameters are optimized based on the difference between the prediction results and the actual data, wherein: The direct difference between the predicted iToF measurement data and the actual iToF measurement data is calculated, and the training parameters are optimized based on the difference. The predicted backscatter vector and the true backscatter vector are treated as two different probability mass functions, and the statistical distance between them is measured based on the EMD function to obtain the difference between the predicted backscatter vector and the true backscatter vector. The training parameters are then optimized based on the difference.

7. The method for reconstructing depth images from a time-of-flight camera according to claim 1, characterized in that, After mapping the backscatter vector or when reconstructing the depth image of the iToF camera based on the iToF measurement data, noise removal is performed on the mapped backscatter vector or the reconstructed depth image based on bilateral filtering.

8. The method for reconstructing depth images from a time-of-flight camera according to claim 1, characterized in that, The network structure of the prediction model includes: An input layer is used to input the stack of iToF measurement data in the complex domain and convert the complex form of the iToF measurement data into a combination of the real and imaginary parts. Multiple convolutional layers are used to extract low-level features from high-level features on the input iToF measurement data. The first convolutional layer incorporates a weight kernel, which is used to retrieve information from each pixel. A fully connected layer is used to map the features extracted by the convolutional layer to the low-dimensional vector; An output layer is used to output the low-dimensional vector.

9. A system for reconstructing depth images from a time-of-flight camera, characterized in that, The system includes: The acquisition module is used to acquire iToF measurement data from the iToF camera, express the iToF measurement data as an integral MPI description equation, convert the MPI description equation into a discrete form with N time steps, so as to represent the iToF measurement data as a measurement matrix and backscatter vector based on a single modulation frequency. The operation is repeated based on multiple different modulation frequencies to obtain the measurement matrix expressed by the iToF measurement data at multiple modulation frequencies and the stack of measurement values ​​in the complex domain. The prediction module is used to input the stack of measurement values ​​of the iToF measurement data into the prediction model. The prediction model is trained based on the nonlinear function of the neural network. The prediction model is used to predict the iToF measurement data as a low-dimensional vector of a specified order of reflected light. The low-dimensional vector represents the amplitude and path length of the reflected light of the iToF measurement data. A mapping module is used to input the low-dimensional vector into the backscattering model and convert the low-dimensional vector into an estimated backscattering vector. The backscattering model is based on a generative adversarial network or a variational autoencoder. The backscattering model is configured as a mapping from the low-dimensional vector to the backscattering vector. The reconstruction module is used to calculate new iToF measurement data based on the measurement matrix and the backscatter vector obtained by mapping, and to reconstruct the depth image of the iToF camera based on the iToF measurement data.

Citation Information

Patent Citations

  • Crack image compression sampling method based on generative adversarial network

    CN111711820A

  • Low-complexity channel estimation method for IRS-assisted backscatter communication system

    CN118869403A