Camera image data processing method and system based on noise reduction technology
Through the camera image data processing method based on noise reduction technology, the problem of degradation of depth measurement accuracy of multi-frequency signals in complex scenarios is solved, and high-precision depth measurement and noise immunity are achieved in complex scenarios.
Patent Information
- Application Number
- CN202510899293.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-01
- Publication Date
- 2025-08-01
- Estimated Expiration
- 2045-07-01
AI Technical Summary
The prior art lacks dynamic optimization of multi-frequency signal processing in complex scenarios, resulting in a degradation of depth measurement accuracy and insufficient signal-to-noise ratio, especially in long-distance multipath interference scenarios, high-frequency aliasing and low-frequency phase fuzzy problems.
The camera image data processing method based on noise reduction technology is adopted to generate multi-frequency measurement tensors by collecting multi-phase charge amount data, and a convolutional neural network with an encoder-decoder structure is used to perform frequency grouping processing, combining phase compensation and windowing processing, inverse Fourier transform and median filtering are performed, time domain waveform features are extracted, confidence threshold segmentation and multi-peak detection are performed, and the depth map with noise enhancement is output through channel-space dual attention weighting and mixed encoding.
Improves the accuracy and signal-to-noise ratio of depth measurement, reduces noise interference, and ensures depth map accuracy and anti-interference ability in complex scenarios.
Smart Images

Figure CN120416677A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of camera data processing, and particularly to a method and system for processing camera image data based on noise reduction technology. Background Art
[0002] In recent years, due to its real-time depth perception ability, indirect time-of-flight (iToF) sensors have been widely used in the fields of three-dimensional reconstruction, autonomous driving, and robot navigation. The iToF technology calculates the target distance by emitting a modulated optical signal and detecting the phase delay of the reflected signal. Traditional methods usually adopt a single modulation frequency, combine four-phase sampling to solve the depth information, and reduce interference through an ambient light suppression algorithm. With the complexity of application scenarios, researchers have proposed multi-frequency modulation technology to balance the near-distance resolution and the far-distance measurement range. For example, an alternating high-low frequency emission mode is adopted, and the time-delay offset is combined with the charge integration characteristics of different frequencies. In addition, noise modeling technology (such as Gaussian-Poisson mixed noise separation) and non-uniformity correction matrix are introduced to improve the consistency of sensor data.
[0003] However, the existing technologies still have significant deficiencies: firstly, the joint processing of multi-frequency signals lacks a dynamic optimization architecture, resulting in low cross-band feature fusion efficiency. Especially in the long-distance multi-path interference scenario, the problems of high-frequency aliasing and low-frequency phase ambiguity coexist; secondly, due to insufficient compensation for the phase offset of the sensor hardware in the existing frequency-domain to time-domain conversion method, the contrast of the time-domain pulse signal decreases. At the same time, traditional window functions (such as rectangular windows) are prone to introduce spectral leakage, affecting the peak positioning accuracy of weak reflection signals. Summary of the Invention
[0004] In view of the above existing problems, the present invention is proposed.
[0005] Therefore, the present invention provides a method for processing camera image data based on noise reduction technology to solve the problems of decreased depth measurement accuracy and insufficient signal-to-noise ratio caused by multi-frequency aliasing, phase offset, and noise interference in complex scenarios.
[0006] To solve the above technical problems, the present invention provides the following technical solutions: In a first aspect, the present invention provides a method for processing camera image data based on noise reduction technology, which includes collecting multi-phase charge quantity data and converting it into a voltage signal matrix, and generating a multi-frequency measurement tensor after frequency grouping processing; Inputting the multi-frequency measurement tensor into a convolutional neural network with an encoder-decoder structure, and using an adversarial training mechanism to generate an extended frequency-domain tensor covering the full target frequency; Performing phase compensation and windowing processing on the extended frequency-domain tensor, then performing inverse Fourier transform, enhancing the time-domain resolution through interpolation processing, and suppressing transient noise using median filtering to generate a transient time-domain image; Extract the time-domain waveform features from the transient time-domain image, segment the confidence region through a confidence threshold, and combine multi-peak detection and random forest to distinguish valid reflection signals, and output a depth probability map; Perform channel-spatial dual attention weighting on the depth probability map, use graph convolution to model neighborhood correlation and implement hybrid coding, and output a noise-resistant enhanced depth map.
[0007] As a preferred scheme of the camera image data processing method based on the noise reduction technology described in the present invention, wherein: the generation of the multi-frequency measurement tensor includes the following steps, The indirect time-of-flight sensor alternately outputs sinusoidal modulation waveforms of different frequencies; At each frequency, control the integration capacitor to start the charge and discharge cycle at multiple phase points, and collect the charge amount data of each phase point; Arrange the charge amount data of each phase point in the spatial dimension and splice them into an original charge amount matrix, convert it into a voltage signal matrix, and then group and reorganize it by frequency to form a multi-frequency measurement tensor.
[0008] As a preferred scheme of the camera image data processing method based on the noise reduction technology described in the present invention, wherein: inputting the multi-frequency measurement tensor into the convolutional neural network of the encoder-decoder structure includes the following steps, The encoder compresses the spatial dimension and expands the number of channels through multi-level stride convolution, and combines skip connections to save intermediate features of different scales; The decoder gradually restores the spatial resolution through transposed convolution and splices it with the intermediate features of different scales saved by the encoder to generate a final feature map; In the final feature map, generate complex components covering the full target frequency through extended frequency domain extrapolation and adversarial training.
[0009] As a preferred scheme of the camera image data processing method based on the noise reduction technology described in the present invention, wherein: the phase compensation and windowing processing of the extended frequency domain tensor includes the following steps, Load the pre-calibrated sensor frequency response parameter matrix to perform phase compensation on the complex tensor covering the full target frequency; Apply a window function to the phase-compensated complex components based on the Hanning window weight to suppress high-frequency oscillations, and merge the real and imaginary components to generate a frequency domain tensor after windowing processing.
[0010] As a preferred scheme of the camera image data processing method based on the noise reduction technology described in the present invention, wherein: the confidence threshold segmentation and multi-peak detection include the following steps, Divide the high and low confidence regions according to the peak confidence heat map; Perform multi-order derivative analysis on the time-domain waveform of the low confidence region to calculate the curvature value and screen candidate peaks; The candidate peaks are probabilistically classified by a pre-trained random forest model, and the effective peaks are determined by combining the time-first principle.
[0011] As a preferred solution of the camera image data processing method based on the noise reduction technology according to the present invention, wherein: the channel-spatial dual attention weighting includes the following steps, Extract the direct and indirect light paths from the depth probability map, generate a multi-modal tensor, and perform channel-by-channel weighting to generate a spatial feature map; Perform dual weighting on the spatial feature map through the spatial attention weight, combine the graph convolution to aggregate the neighborhood features, and output the optimized depth map.
[0012] As a preferred solution of the camera image data processing method based on the noise reduction technology according to the present invention, wherein: the hybrid coding refers to calculating the uncertainty matrix according to the optimized indirect light path probability map, dividing sub-blocks of a predetermined size, and performing lossless coding on the high-confidence sub-blocks and lossy compression on the low-confidence sub-blocks.
[0013] In a second aspect, the present invention provides a camera image data processing system based on the noise reduction technology, including a multi-frequency acquisition module that acquires multi-phase charge quantity data and converts it into a voltage signal matrix, and generates a multi-frequency measurement tensor after frequency grouping processing; A frequency domain expansion module that inputs the multi-frequency measurement tensor into a convolutional neural network with an encoder-decoder structure, and uses an adversarial training mechanism to generate an extended frequency domain tensor covering the full target frequency; A time domain noise reduction module that performs phase compensation and windowing processing on the extended frequency domain tensor and then performs inverse Fourier transform, enhances the time domain resolution through interpolation processing, and suppresses transient noise using median filtering to generate a transient time domain image; A signal classification module that extracts time domain waveform features from the transient time domain image, divides the confidence region through a confidence threshold, and combines multi-peak detection and a random forest to distinguish effective reflection signals, and outputs a depth probability map; A depth optimization module that performs channel-spatial dual attention weighting on the depth probability map, uses graph convolution to model neighborhood correlation, and implements hybrid coding, and outputs a noise-resistant enhanced depth map.
[0014] In a third aspect, the present invention provides a computer device, including a memory and a processor, where the memory stores a computer program, and wherein: when the computer program is executed by the processor, any step of the camera image data processing method based on the noise reduction technology according to the first aspect of the present invention is implemented.
[0015] Fourthly, the present invention provides a computer-readable storage medium, on which a computer program is stored, wherein: when the computer program is executed by a processor, any step of the camera image data processing method based on the noise reduction technology as described in the first aspect of the present invention is implemented.
[0016] The beneficial effects of the present invention are as follows: Based on the quantum efficiency term and the Box-Muller transform, a Gaussian-Poisson mixed noise model is generated, and the noise ratio is dynamically allocated in combination with the sensor dark field calibration data to effectively simulate the real noise distribution; An extended frequency domain tensor is generated through frequency domain extrapolation adversarial training, and combined with the phase compensation matrix and the Hanning window, the frequency domain resolution is increased to the upper limit of the extrapolated frequency to ensure that the time domain waveform after the inverse Fourier transform has high contrast and low ringing artifacts; A multi-peak detection mechanism guided by a confidence heat map is adopted, and effective reflections and noises are distinguished through curvature analysis and a random forest model, and the earliest arriving peak is preferentially selected; Hybrid coding is implemented based on the uncertainty matrix, retaining lossless details in high-confidence regions and using DCT dynamic quantization to compress noise energy in low-confidence regions, reducing data redundancy while ensuring the accuracy of the depth map. Description of the Drawings
[0017] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings required for the description of the embodiments will be briefly introduced below. Obviously, the drawings in the following description are only some embodiments of the present invention, and those of ordinary skill in the art can obtain other drawings based on these drawings without creative efforts.
[0018] Figure 1 It is a flowchart of a camera image data processing method based on noise reduction technology.
[0019] Figure 2 It is a schematic diagram of multi-frequency measurement tensor generation.
[0020] Figure 3 It is a schematic diagram of the encoder-decoder processing flow.
[0021] Figure 4 It is a schematic diagram of confidence threshold segmentation and multi-peak detection. Detailed Embodiments
[0022] In order to make the above-mentioned objects, features, and advantages of the present invention more obvious and understandable, the following will describe the detailed embodiments of the present invention with reference to the drawings in the specification.
[0023] Many specific details are set forth in the following description to facilitate a thorough understanding of the present invention, but the present invention may be implemented in other ways different from those described herein. Those skilled in the art can make similar extensions without departing from the connotation of the present invention, so the present invention is not limited by the specific embodiments disclosed below.
[0024] Secondly, the so-called "one embodiment" or "embodiment" herein refers to a specific feature, structure, or characteristic that may be included in at least one implementation manner of the present invention. The appearances of "in one embodiment" in different places in this specification do not all refer to the same embodiment, nor are they separate or alternative embodiments that exclude each other from other embodiments.
[0025] Referring to Figures 1 to 4 , which is an embodiment of the present invention. This embodiment provides a method for processing camera image data based on noise reduction technology, including the following steps: S1. Configure the optical signal emission module of the indirect time-of-flight (iToF) sensor to alternately output sine modulation waveforms of 20 MHz and 100 MHz, and set the modulation light duration for each frequency to 5 ms.
[0026] At a modulation frequency of 20 MHz, control the integration capacitor to start the charge and discharge cycle at four phase points of 0°, 90°, 180°, and 270°. The time delay offsets corresponding to each phase point are 0 ns, 12.5 ns, 25 ns, and 37.5 ns, and collect the charge quantity data of each pixel at the four phase points; at a modulation frequency of 100 MHz, repeat the phase point control process, adjust the time delay offset to 0 ns, 2.5 ns, 5 ns, and 7.5 ns, and collect the charge quantity data of the corresponding phase points.
[0027] It should be noted that high frequency (100 MHz) can improve the near-distance resolution, and low frequency (20 MHz) can expand the far-distance measurement range (reduce phase aliasing); 0°, 90°, 180°, and 270° phase sampling can eliminate ambient light interference, extract complete sine cycle information, and ensure the depth calculation accuracy; the time delay offset adapted to the frequency can match the charge integration characteristics of the sensor and maximize the signal utilization rate.
[0028] Arrange the four-phase charge quantity data at 20 MHz and 100 MHz into a sub-tensor of 120×160×4 in the spatial dimension, where the third dimension order is 0°, 90°, 180°, and 270°; splice the sub-tensors of 20 MHz and 100 MHz along the third dimension to generate an original charge quantity matrix with a dimension of 120×160×8; according to the charge-voltage conversion coefficient of the indirect time-of-flight sensor, convert the charge quantity matrix into a voltage signal matrix; Extract the voltage signals by frequency grouping: the 20 MHz frequency group includes channels 1-4, and the 100 MHz frequency group includes channels 5-8, and reorganize them into a 120×160×4×2 tensor; separate the voltage signals of the two frequencies along the fourth dimension to generate an original voltage signal matrix with a final dimension of 120×160×4.
[0029] It should be noted that by separating the 20MHz and 100MHz data through frequency grouping, the independence between frequencies is retained, which facilitates optimization processing for different frequency bands.
[0030] Extract the charge quantity data of four phase angles of each pixel at 20MHz and 100MHz modulation frequencies from the original voltage signal matrix with dimensions of 120×160×4. Perform the same transformation on the charge quantity data of the 20MHz frequency group and the 100MHz frequency group respectively, and calculate the quantum efficiency term respectively; independently generate random noise values that conform to the standard Gaussian distribution for each phase angle through the Box-Muller transformation method, and superimpose the random noise values and the quantum efficiency term according to frequency and phase; reorganize the four-phase data of the 20MHz and 100MHz frequency groups after superposition into an intermediate photon count matrix in the original channel order; compress the third dimension of the intermediate photon count matrix, only retain the valid channel data corresponding to the 0° and 90° phases, and perform zeroing processing on the invalid phase channels, and finally generate an initial photon count matrix with dimensions of 120×160×4.
[0031] It should be noted that performing zeroing processing on the invalid phase channels is to suppress the 180° and 270° phase noises (which are greatly affected by circuit non-ideality), retain the 0° and 90° valid signals, and improve the signal-to-noise ratio.
[0032] Traverse each pixel position and phase channel of the initial photon count matrix, and extract the expected number of photons at the current position; generate a uniformly distributed random number r∈[0,1] for each expected number of photons, and use the linear congruential generator algorithm to generate a random sequence with a fixed seed value to ensure repeatability.
[0033] By calibrating the noise distribution characteristics of the indirect time-of-flight sensor in the dark field (no light input) and under uniform illumination conditions, to balance the contributions of the two types of noises, set the random number threshold to 0.3 to select the noise generation method: When r < 0.3, perform Gaussian noise generation. Calculate the standard deviation, and use the Box-Muller transformation to generate standard normal distribution variables and synthesize Gaussian noise values; When r ≥ 0.3, perform Poisson noise generation. Use the inverse transform method to generate Poisson noise values by comparing the cumulative distribution function; Combine the Gaussian noise values and the Poisson noise values according to the conditional selection logic into the noise photon number of each pixel; perform saturation truncation operation on the noise photon numbers that exceed the maximum photon count capacity of the sensor; reorganize all truncated noise photon values in the dimension order of the initial photon count matrix to generate a noise photon count matrix.
[0034] It should be noted that the noise of the sensor under no-light conditions (dark field) is mainly composed of thermionic emission and circuit noise, and its statistical characteristics follow a Gaussian distribution. By measuring the dark field data, it is found that the noise accounts for about 30% of the total noise energy. Under light conditions, the random arrival of photons results in Poisson noise, whose intensity is proportional to the square root of the signal intensity. When the signal intensity is within the normal working range, the Poisson noise accounts for about 70%. Therefore, the random number threshold is set to 0.3.
[0035] The noise photon count matrix is separated into four sub-matrices in the order of phase channels, corresponding to 0°, 90°, 180°, and 270° phase angles respectively; perform subtraction operations on the 0° and 180° phase sub-matrix data of 20 MHz and 100 MHz frequencies to generate the real part, and perform subtraction operations on the 90° and 270° phase sub-matrix data to generate the imaginary part; combine the real and imaginary parts of 20 MHz and 100 MHz frequencies into complex forms respectively to form complex component matrices containing two channels; splice the complex component matrices of 20 MHz and 100 MHz along the channel dimension to generate an original complex tensor containing four channels.
[0036] It should be noted that the real and imaginary part calculations (0°, 90°, 180°, 270°) can eliminate the DC offset, extract the AC signal components, and enhance the contrast of the effective reflection signal.
[0037] Load the pre-calibrated non-uniformity correction matrix from the non-volatile memory.
[0038] Furthermore, the non-uniformity correction matrix is generated in the uniform light calibration experiment of the indirect time-of-flight sensor, and is used to compensate for pixel-level response differences (such as quantum efficiency deviation, circuit gain fluctuation) to improve the data spatial consistency.
[0039] Traverse each spatial pixel point of the original complex tensor, and extract the complex component data of the four channels in each spatial pixel point; perform a correction operation on the real part channel of 20 MHz frequency, and multiply the complex component data by the correction matrix value at the corresponding position; repeat the same multiplication correction operation for the 20 MHz imaginary part, 100 MHz real part, and imaginary part channels in sequence; recombine the corrected four-channel data into an intermediate tensor in the original order; verify whether the real and imaginary part values of all channels of the intermediate tensor exceed the range of the analog-to-digital converter, and perform saturation truncation on the exceeded part, and output the corrected intermediate tensor.
[0040] Traverse each spatial position and channel index of the corrected intermediate tensor, and extract the real and imaginary complex components corresponding to the 20 MHz and 100 MHz frequencies; calculate the modulus length for the real and imaginary parts of each complex component, use the fast square root algorithm to obtain the square root of the corresponding sum of squares, record all the modulus length calculation results, and determine the global maximum modulus length value through the parallel reduction algorithm; normalize the real and imaginary parts of each complex component respectively based on the global maximum modulus length value; recombine the normalized real and imaginary parts in the original channel order to generate a normalized multi-frequency measurement tensor containing the phase components of 20 MHz and 100 MHz each.
[0041] S2. Extract the four-channel data of each spatial position from the normalized multi-frequency measurement tensor; initialize the three-dimensional convolution kernel of the input convolutional layer and configure the weight parameters; perform a convolution operation on the input four-channel data to generate intermediate eigenvalues of 64 channels; process the intermediate eigenvalues through a non-linear activation function to suppress negative value outputs and obtain the activated feature map; perform channel-level normalization on the activated feature map to stabilize the data distribution; verify the normalized feature map, maintain the original spatial resolution and expand the number of channels to form an initial feature map with a dimension of 120×160×64.
[0042] It should be noted that by fusing the real and imaginary components of multiple frequencies (20 MHz / 100 MHz) through the three-dimensional convolution kernel, cross-band joint features are extracted, enhancing the ability to model the correlation between frequencies; the number of channels is expanded to 64 to improve the feature expression ability, while maintaining the original spatial resolution to avoid early information loss.
[0043] Extract each 3×3 local region from the initial feature map with a dimension of 120×160×64, covering the four-channel input data; initialize 128 independent 3×3 convolution kernels, where the weight matrix of the convolution kernel is initialized by the He normal distribution and the bias term is initialized to zero; Perform a convolution operation with a stride of 2 on each 60×80 output position: sample once every other pixel in the horizontal and vertical directions, and maintain the boundary integrity by padding 1 to output a feature map of 60×80×128; apply the ReLU activation function to the output 60×80×128 feature map to suppress negative values and retain positive activations; calculate the mean and variance of 128 channels along the batch dimension and apply batch normalization for processing; save the processed feature map with a dimension of 60×80×128 as skip connection 1 for use by the first upsampling layer of the decoder.
[0044] Taking the 60×80×128 feature map of skip connection 1 as the input, initialize 256 3×3 convolutional kernels. The parameter initialization method is the same as that of the first downsampling. Perform a convolutional operation with a stride of 2 and a padding of 1, and output a feature map with a dimension of 30×40×256; successively perform ReLU activation and batch normalization on the 30×40×256 feature map, save the result as skip connection 2, and pass it to the second upsampling layer of the decoder.
[0045] Taking the 30×40×256 feature map of skip connection 2 as the input, initialize 512 3×3 convolutional kernels; perform a convolutional operation with the same parameters, and output a feature map with a dimension of 15×20×512; also apply ReLU activation and batch normalization to the feature map with a dimension of 15×20×512, and save the final downsampling result as skip connection 3.
[0046] It should be noted that through three convolutional operations with a stride of 2, the spatial dimension is gradually compressed (4-fold → 16-fold → 64-fold downsampling), the receptive field is expanded, and the global scene structure (such as the contour of large-scale objects) is captured; the number of channels is doubled (64 → 128 → 256 → 512), enhancing the abstraction ability of high-level semantic features (such as material and motion blur); saving skip connections is to retain intermediate features of different scales for fusing multi-resolution information during the upsampling of the decoder to restore details.
[0047] Taking the skip connection 3 output from the third downsampling as the input, extract the feature vectors at each spatial position; initialize 512 3×3 dilated convolutional kernels, and configure the parameters with a dilation rate of 2 and a padding of 2; Perform a dilated convolutional operation on the feature vectors at each spatial position: at each 15×20 position, sample a 3×3 neighborhood at an interval of 2 pixels, actually covering a 7×7 physical area; calculate the feature values at each output position through dot product accumulation, and output a feature map with a dimension of 15×20×512.
[0048] It should be noted that performing dilated convolution is to cover a 7×7 physical area without reducing the resolution (15×20×512), capture long-range dependencies (such as long-distance multi-path interference patterns), avoid the loss of spatial information caused by traditional downsampling, and maintain sensitivity to small targets.
[0049] Initialize 256 2×2 transposed convolutional kernels, initialize the weights through bilinear interpolation, and set the bias term to zero; perform a transposed convolutional operation with a stride of 2 on the feature map with a dimension of 15×20×512: insert zero values at each 15×20 position and perform convolution, increase the spatial resolution to 30×40, and reduce the number of channels to 256, and output a feature map with a dimension of 30×40×256.
[0050] Initialize 128 2×2 transposed convolution kernels, and perform a transposed convolution operation with a stride of 2 on the feature map with dimensions 30×40×256: restore the resolution to 60×80, reduce the number of channels to 128, output a feature map with dimensions 60×80×128, and concatenate it with the encoder skip connection 1 along the channel dimension to generate a fused feature map with dimensions 60×80×256.
[0051] Initialize 64 2×2 transposed convolution kernels, and perform a transposed convolution operation with a stride of 2 on the fused feature map with dimensions 60×80×256 to restore the original resolution of 120×160, reduce the number of channels to 64, output a feature map with dimensions 120×160×64, and concatenate it with the feature map with dimensions 120×160×64 output by the input convolutional layer along the channel dimension to generate a final feature map with dimensions 120×160×128.
[0052] Extract the feature vectors at each spatial position from the final feature map with dimensions 120×160×128; initialize 40 independent 1×1 convolution kernels, and perform a fully connected convolution operation on the feature vectors at each spatial position: perform dot product accumulation on the 128-dimensional feature vectors and the 40 convolution kernels respectively to generate a 40-channel output, and arrange the output channels in ascending order of frequency to generate an extended frequency domain tensor containing multi-frequency real and imaginary components; construct a loss function based on the absolute error between the extended frequency domain tensor containing multi-frequency real and imaginary components and the true Fourier coefficients to guide frequency domain extrapolation; design a discriminator network to analyze the data distribution difference between the extended frequency domain tensor and the true value; optimize the reconstruction accuracy of the high-frequency components of the extended frequency domain tensor through an adversarial training mechanism, and finally output an extended frequency domain tensor covering the full target frequency.
[0053] S3. Extract the real and imaginary part data of the frequency channels at each spatial position from the extended frequency domain tensor covering the full target frequency; load the pre-calibrated sensor frequency response parameter matrix, and the sensor frequency response parameter matrix stores the phase compensation factors at each frequency point; perform a phase compensation operation on the channel indices corresponding to each sensor frequency response parameter matrix.
[0054] Further note that the phase compensation operation refers to calculating the compensated complex component based on the extracted real and imaginary part data of each frequency channel, eliminating the phase error introduced by the hardware, and enhancing the contrast of the pulse signal after frequency domain-time domain conversion.
[0055] It should be noted that loading the frequency response parameter matrix is to compensate for the inherent phase offset of the sensor at different frequency points (such as circuit delay, quantum efficiency difference), ensure the phase consistency of multi-band signals, and improve the full-frequency domain data alignment accuracy.
[0056] Rewrite the real and imaginary part data in the compensated complex components into an extended frequency-domain tensor that covers the entire target frequency in the order of frequency channels to generate a phase-aligned frequency-domain tensor.
[0057] Generate multiple discrete frequency points according to the frequency range; calculate the Hanning window weights for each discrete frequency point; separate the phase-aligned frequency-domain tensor by odd and even indices, where the odd channels are the real part and the even channels are the imaginary part; match the separated frequency-domain tensor with the Hanning window weights corresponding to the frequencies, apply the window function weights to the real and imaginary components of each frequency channel respectively to suppress high-frequency oscillations, merge the windowed real and imaginary data and verify the high-frequency attenuation effect to generate a windowed frequency-domain tensor.
[0058] It should be noted that the calculation of the Hanning window weights is to suppress the spectral leakage caused by the truncation of frequency-domain data, smooth the high-frequency oscillations, and reduce the ringing artifacts in the time-domain waveform; windowing is applied independently to the real and imaginary components to preserve the orthogonality of the frequency-domain signal and avoid the complex signal distortion introduced by the window weights.
[0059] Extract the real and imaginary part data of the frequency channels from each pixel of the windowed frequency-domain tensor and arrange them into a complex sequence; perform high-frequency zero-padding extension before and after the complex sequence to form frequency-domain data of a standard length; perform a fast inverse Fourier transform on the frequency-domain data of the standard length to generate a time-domain waveform; set the total time duration, time interval, and time points in the time domain according to the time-domain waveform, check the corresponding pulse start point and the corresponding waveform attenuation end point to generate the initial time-domain energy.
[0060] It should be noted that the high-frequency zero-padding extension can improve the resolution of the frequency-domain data, increase the interpolation points of the time-domain waveform, and make the generated time-domain pulse more accurate; the inverse Fourier transform is to convert the frequency-domain extended signal into a time-domain waveform.
[0061] According to the extrapolated frequency upper limit, apply the Nyquist sampling theorem to determine the minimum time interval; construct a cubic spline interpolation function for the time-domain waveform of each pixel of the initial time-domain energy, using the initial time point as the node to generate a new time network and output the interpolated time-domain tensor; based on the interpolated time-domain tensor, determine the effective reflection signal interval and the noise-dominated attenuation tail, and intercept the first fixed time point of the time-domain waveform of each pixel to generate a dynamically adjusted time-domain tensor.
[0062] It should be noted that the application of the Nyquist sampling theorem is to determine the minimum time interval based on the extrapolated frequency upper limit to avoid time-domain waveform aliasing distortion; the cubic spline interpolation is to smooth the initial time-domain waveform, improve the time resolution, and enhance the detectability of weak signal peaks.
[0063] Read the dynamically adjusted time-domain tensor and traverse all spatial positions and time points; perform a conditional judgment on each dynamically adjusted time-domain tensor: if the optical signal intensity value at each spatial position and time point is less than zero, the dynamically adjusted time-domain tensor is set to zero, otherwise, the optical signal intensity value at the position before the conditional judgment is retained; count the number of negative value points that are set to zero in the dynamically adjusted time-domain tensor and calculate the negative value ratio; extract the optical signal intensity values of all pixels at the initial moment from the non-negatively processed time-domain tensor, traverse all spatial positions, find the maximum initial optical signal intensity value and perform global scaling correction, and verify the corrected optical signal intensity constraint and energy conservation characteristics, and output a transient time-domain waveform that conforms to physical rationality.
[0064] It should be noted that by excluding the attenuation tail dominated by noise (such as time-domain data of ≥80 ns), retaining the effective reflection interval, reducing computational redundancy and noise interference; forcing negative photon counts to zero to ensure that the time-domain waveform conforms to physical reality (optical and electrical signals cannot be negative) and eliminating abnormal pulses caused by computational errors; global scaling correction avoids excessive differences in light intensity between different pixels (such as strong reflection objects and weak signal backgrounds), prevents numerical overflow and balances the dynamic range.
[0065] Extract the time-domain waveform of each pixel from the transient time-domain waveform and construct continuous sampling points of a sliding window; fill the out-of-bounds positions at the start and end of each pixel time-domain waveform with a mirror image, and collect the continuous intensity values of the sampling points within the window at each filled out-of-bounds position, and take the median value after sorting in ascending order as the intermediate result after median filtering and noise reduction; traverse all pixel positions and time points, replace the intermediate result after noise reduction with the median filtering result to generate a noise-reduced waveform; extend mirror image points at both ends of each time-domain waveform of the noise-reduced waveform, and multiply and accumulate the extended mirror image points with Gaussian kernel points at each position to obtain a transient time-domain image.
[0066] It should be noted that by suppressing instantaneous noise in the time-domain waveform (such as thermal noise spikes), retaining the steep edges of real peaks, and improving peak localization accuracy; by eliminating the truncation effect at the beginning and end of the time-domain waveform, reducing boundary noise, and generating a continuous and smooth transient time-domain image.
[0067] S4. Reshape the time series of each pixel point of the time-domain waveform in the transient time-domain image into a spatio-temporal two-dimensional structure to generate an input tensor; sequentially input the input tensor into three convolutional layers to enhance local feature expression, and output a feature map with a resolution of 30×40×64; apply global average pooling and a fully connected layer to the feature map with a resolution of 30×40×64 to generate a peak confidence heat map.
[0068] Decouple the time series of the time-domain waveform in the transient image from the spatial dimension to form an input tensor, facilitating the convolutional layer to capture the time-domain dynamic change features; perform local feature extraction through three convolutional layers to compress the spatial dimension, enhance the key patterns of the time-domain waveform (such as the steepness of the main peak and the position of the secondary peak), and suppress random noise.
[0069] Extract the confidence value of each pixel from the peak confidence heatmap and set a confidence threshold; traverse all spatial positions, compare the confidence value of each pixel with the confidence threshold. If the confidence value is greater than or equal to the confidence threshold, mark it as a high-confidence region; if the confidence value is less than the confidence threshold, mark it as a low-confidence region.
[0070] Further explanation, the confidence threshold is determined by statistically analyzing the confidence value distribution corresponding to the true peak position during the calibration stage of the uniform reflecting surface. 0.7 is the empirical balance point, which can not only retain high-reliability depth information (such as the main peak of the direct light path), but also avoid over-filtering weak signals (such as low-reflectivity targets at a long distance).
[0071] For the high-confidence region, extract the corresponding time-domain waveform from the transient time-domain image, search for the maximum position along the time axis, and record the peak time and amplitude. For the low-confidence region, extract the corresponding time-domain waveform from the transient time-domain image, perform zero-padding and Gaussian convolution operations to generate smoothed time-domain waveform data; based on the smoothed time-domain waveform data, calculate the first derivative at each time point through the central difference method to describe the waveform slope change; based on the first derivative result, continue to apply the central difference method to calculate the second derivative to reflect the concavity and convexity change of the waveform; substitute the first derivative and the second derivative into the curvature formula to calculate the curvature value at each time point, and the magnitude of the curvature value characterizes the local bending degree of the waveform; scan along the time axis, mark the zero-crossing position where the sign of the curvature changes at the sign change point of the curvature value, statistically analyze the length of the continuous zero-crossing interval, exclude invalid peaks in combination with the amplitude threshold, and correct the time-domain waveform data to finally screen out candidate peaks.
[0072] Further explanation, the amplitude threshold is based on the noise data collected by the sensor during the calibration stage in the dark field (no light input), calculate the mean and standard deviation of the time-domain waveform amplitude, and determine the amplitude threshold of the effective signal through the uniform illumination calibration experiment.
[0073] Load the pre-trained random forest model. The model contains 100 decision trees, with a maximum depth of 10 for each tree, and supports multi-class probability output. Input the candidate peaks generated after multi-peak detection in the low-confidence region into the random forest model, and the random forest outputs probabilities for each candidate peak. Set an effective depth threshold, and filter out the candidate peaks with output probabilities greater than or equal to the effective depth threshold, which are marked as effective peaks. If no candidate peak meets the effective depth threshold, it is marked as an invalid detection. When there are multiple candidate peaks with output probabilities greater than or equal to the effective depth threshold, extract the corresponding times and select the earliest-arriving peak as the effective peak. Integrate all the effective peaks to form a depth probability map. In the depth probability map, record the direct light path depth as channel 1, the indirect light path probability as channel 2, and mark the invalid detection areas.
[0074] Further explanation: The effective depth threshold is determined by comprehensively considering the ROC curve of the calibration experiment and the multi-path suppression requirements, ensuring that under the condition of retaining high-reliability peaks, the detection sensitivity and anti-interference ability are balanced.
[0075] It should be noted that multi-class probability prediction is performed using candidate peak features (amplitude ratio, time interval, area ratio) to distinguish the main peak of the direct light path, multi-path secondary peaks, and noise. The direct light path signal (the earliest-arriving peak) takes precedence over multi-path reflections, which conforms to the physical propagation law and reduces interference in dynamic scenes.
[0076] S5. Extract the direct light path depth value of each pixel from channel 1 of the depth probability map, extract the indirect light path probability value of each pixel from channel 2 of the depth probability map D, traverse each pixel of the transient time-domain image, and extract the maximum amplitude from the time-domain waveform. Stack the direct light path depth values, indirect light path probability values, and maximum amplitude along the channel dimension in sequence to generate a multi-modal feature tensor, where channel 3 is the maximum amplitude.
[0077] It should be noted that the maximum amplitude reflects the signal strength, which helps to distinguish real reflections (high amplitude) from noise (low amplitude) and improve the anti-interference ability.
[0078] Perform global average pooling on each channel of the multi-modal feature tensor, calculate the global mean of each channel to generate a channel importance vector. Convert the channel importance vector into channel attention weights through a fully connected layer and a Sigmoid activation function. Perform channel-wise weighting on the multi-modal feature tensor based on the channel attention weights to generate a channel-weighted feature tensor. Calculate the mean of the channel-weighted feature tensor along the channel dimension to generate a spatial feature map, and input the spatial feature map into a convolutional layer to generate spatial attention weights through convolution and Sigmoid activation. Combine the channel and spatial attention weights to perform double weighting on the channel-weighted feature tensor to generate an enhanced multi-modal feature tensor.
[0079] It should be noted that by means of global average pooling and Sigmoid activation, the importance of each channel (depth, probability, amplitude) is quantified to suppress the interference of redundant or low signal-to-noise ratio channels; based on the channel-weighted features, a spatial feature map is generated to highlight key regions (such as the edges of highly reflective objects and high-incidence areas of multipath), and weaken the contribution of the uniform background.
[0080] Each pixel of the enhanced multi-modal feature tensor is defined as a graph node, a 3×3 neighborhood connection relationship is constructed and the node attributes are bound; through the first-layer graph convolution, the neighborhood features are aggregated and activated to generate high-dimensional intermediate features; the second-layer graph convolution is used to reduce the dimension and reconstruct the high-dimensional intermediate features, restore the original number of channels and output the optimized depth feature map.
[0081] It should be noted that the feature vectors of each pixel are associated with neighboring pixels to model local spatial correlations (such as object continuity and spatial propagation patterns of multipath reflections); through the aggregation and dimensionality reduction of double-layer graph convolution, the spatial consistency of the depth map is improved, and isolated noise points and outliers are reduced.
[0082] Separate the depth channel from the optimized depth feature map: channel 1 is the optimized depth map, and channel 2 is the optimized indirect light path probability map; extract the uncertainty matrix from the optimized indirect light path probability map of channel 2, divide it into 8×8 sub-blocks, and calculate the average uncertainty of each block; in the uniform illumination calibration stage, by analyzing the depth errors corresponding to different uncertainties, set the uncertainty threshold to 0.2; when the average uncertainty is less than or equal to 0.2, it is regarded as a high-confidence sub-block, and the corresponding region is extracted from the optimized depth map of channel 1 and lossless coding is performed; when the average uncertainty is greater than 0.2, it is regarded as a low-confidence sub-block, and the corresponding region of the optimized depth map of channel 1 is lossily compressed, and DCT transform, dynamic quantization and run-length coding are performed; the lossless coding blocks and lossy compression blocks are recombined into a hybrid bitstream according to the spatial position, block type markers and quantization parameters are added, and packaging is performed to output the noise-resistant enhanced depth map.
[0083] It should be noted that the uncertainty is calculated based on the optimized indirect light probability map to quantify the depth reliability of each region; high-confidence blocks retain lossless details to avoid compression distortion, and low-confidence blocks use lossy compression to suppress noise energy.
[0084] This embodiment also provides a camera image data processing system based on noise reduction technology, including: A multi-frequency acquisition module that acquires multi-phase charge quantity data and converts it into a voltage signal matrix, and generates a multi-frequency measurement tensor after frequency grouping processing; A frequency domain expansion module that inputs the multi-frequency measurement tensor into a convolutional neural network with an encoder-decoder structure, and uses an adversarial training mechanism to generate an extended frequency domain tensor covering the full target frequency; The time-domain noise reduction module performs phase compensation and windowing on the extended frequency-domain tensor, then performs inverse Fourier transform, enhances the time-domain resolution through interpolation processing, and suppresses transient noise using median filtering to generate a transient time-domain image; The signal classification module extracts time-domain waveform features from the transient time-domain image, segments the confidence region through a confidence threshold, and combines multi-peak detection and random forest to distinguish valid reflection signals, outputting a depth probability map; The depth optimization module performs channel-spatial dual attention weighting on the depth probability map, models neighborhood correlation using graph convolution, and implements hybrid coding, outputting a noise-resistant enhanced depth map.
[0085] This embodiment also provides a computer device applicable to the case of a camera image data processing method based on noise reduction technology, including: a memory and a processor; the memory is used to store computer-executable instructions, and the processor is used to execute the computer-executable instructions to implement the camera image data processing method based on noise reduction technology proposed in the above embodiment.
[0086] This computer device can be a terminal. The computer device includes a processor, a memory, a communication interface, a display screen, and an input device connected through a system bus. Among them, the processor of this computer device is used to provide computing and control capabilities. The memory of this computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The communication interface of this computer device is used to communicate with an external terminal in a wired or wireless manner. The wireless manner can be achieved through WIFI, a carrier network, NFC (Near Field Communication), or other technologies. The display screen of this computer device can be a liquid crystal display screen or an electronic ink display screen. The input device of this computer device can be a touch layer covering the display screen, or a button, a trackball, or a touchpad provided on the housing of the computer device, or an external keyboard, touchpad, or mouse, etc.
[0087] This embodiment also provides a storage medium, on which a computer program is stored. When the program is executed by a processor, it implements the method for processing camera image data based on noise reduction technology as proposed in the above embodiment; the storage medium can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (Static Random Access Memory, abbreviated as SRAM), electrically erasable programmable read-only memory (Electrically Erasable Programmable Read-Only Memory, abbreviated as EEPROM), erasable programmable read-only memory (Erasable Programmable Read Only Memory, abbreviated as EPROM), programmable read-only memory (Programmable Red-Only Memory, abbreviated as PROM), read-only memory (Read-Only Memory, abbreviated as ROM), magnetic memory, flash memory, magnetic disk or optical disk.
[0088] In summary, the present invention generates a Gaussian-Poisson mixed noise model based on the quantum efficiency term and the Box-Muller transform, dynamically allocates the noise ratio in combination with the sensor dark field calibration data, and effectively simulates the real noise distribution; generates an extended frequency domain tensor through frequency domain extrapolation adversarial training, combines the phase compensation matrix and the Hanning window windowing, and improves the frequency domain resolution to the upper limit of the extrapolation frequency to ensure that the time domain waveform after the inverse Fourier transform has high contrast and low ringing artifacts; adopts a multi-peak detection mechanism guided by a confidence heat map, distinguishes effective reflections and noise through curvature analysis and a random forest model, and preferentially selects the earliest arriving peak; implements hybrid coding based on the uncertainty matrix, retains lossless details in high-confidence regions, and uses DCT dynamic quantization to compress the noise energy in low-confidence regions, reducing data redundancy while ensuring the accuracy of the depth map.
[0089] It should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and not to limit them. Although the present invention has been described in detail with reference to the preferred embodiments, those of ordinary skill in the art should understand that the technical solutions of the present invention can be modified or equivalently replaced without departing from the spirit and scope of the technical solutions of the present invention, and they should all be covered within the scope of the claims of the present invention.
Claims
1. A method for processing camera image data based on noise reduction technology, characterized in that: including Collecting multi-phase charge quantity data and converting it into a voltage signal matrix, and generating a multi-frequency measurement tensor after frequency grouping processing Inputting the multi-frequency measurement tensor into a convolutional neural network with an encoder-decoder structure, and using an adversarial training mechanism to generate an extended frequency domain tensor covering the entire target frequency Performing phase compensation and windowing processing on the extended frequency domain tensor, then performing inverse Fourier transform, enhancing the time domain resolution through interpolation processing and suppressing transient noise using median filtering to generate a transient time domain image Extracting time domain waveform features from the transient time domain image, segmenting the confidence region through a confidence threshold, and combining multi-peak detection and random forest to distinguish effective reflection signals, and outputting a depth probability map Performing channel-spatial dual attention weighting on the depth probability map, using graph convolution to model neighborhood correlation and implementing hybrid coding, and outputting a noise-resistant enhanced depth map 2. The method for processing camera image data based on noise reduction technology according to claim 1, characterized in that: The generating of the multi-frequency measurement tensor includes the following steps Alternately outputting sinusoidal modulation waveforms of different frequencies through an indirect time-of-flight sensor At each frequency, controlling the integration capacitor to start the charge and discharge cycle at multiple phase points, and collecting the charge quantity data of each phase point Arranging the charge quantity data of each phase point in the spatial dimension and splicing them into an original charge quantity matrix, converting it into a voltage signal matrix, and then reorganizing it by frequency grouping into a multi-frequency measurement tensor 3. The method for processing camera image data based on noise reduction technology according to claim 2, wherein: The inputting of the multi-frequency measurement tensor into a convolutional neural network with an encoder-decoder structure includes the following steps The encoder compresses the spatial dimension and expands the number of channels through multi-level stride convolution, and combines skip connections to save intermediate features of different scales The decoder gradually restores the spatial resolution through transposed convolution and splices it with the intermediate features of different scales saved by the encoder to generate a final feature map In the final feature map, through extended frequency domain extrapolation and adversarial training, complex components covering the entire target frequency are generated 4. The method for processing camera image data based on noise reduction technology according to claim 3, characterized in that: The performing of phase compensation and windowing processing on the extended frequency domain tensor includes the following steps Loading a pre-calibrated sensor frequency response parameter matrix to perform phase compensation on the complex tensor covering the entire target frequency Applying a window function based on the Hanning window weight to the phase-compensated complex components to suppress high-frequency oscillations, and merging the real and imaginary components to generate a windowed frequency domain tensor 5. The method for processing camera image data based on noise reduction technology according to claim 4, wherein: The confidence threshold segmentation and multi-peak detection include the following steps Dividing the high and low confidence regions according to the peak confidence heat map Performing multi-order derivative analysis on the time domain waveforms in the low confidence region to calculate the curvature value and screening candidate peaks Performing probability classification on the candidate peaks through a pre-trained random forest model, and determining the effective peaks in combination with the time priority principle 6. The method for processing camera image data based on noise reduction technology according to claim 5, characterized in that: The channel-spatial dual attention weighting includes the following steps Extracting direct and indirect light paths from the depth probability map, generating a multi-modal tensor, and performing channel-wise weighting to generate a spatial feature map Doubly weighting the spatial feature map through spatial attention weights, and aggregating neighborhood features in combination with graph convolution to output an optimized depth map 7. The method for processing camera image data based on noise reduction technology according to claim 6, characterized in that: The hybrid coding refers to calculating an uncertainty matrix according to the optimized indirect light path probability map, dividing sub-blocks of a predetermined size, and performing lossless coding on high-confidence sub-blocks and lossy compression on low-confidence sub-blocks 8. A camera image data processing system based on noise reduction technology, based on the camera image data processing method based on noise reduction technology according to any one of claims 1 to 7, characterized in that: including A multi-frequency acquisition module that acquires multi-phase charge quantity data and converts it into a voltage signal matrix, and generates a multi-frequency measurement tensor after frequency grouping processing; A frequency domain expansion module that inputs the multi-frequency measurement tensor into a convolutional neural network with an encoder-decoder structure, and uses an adversarial training mechanism to generate an expanded frequency domain tensor covering the entire target frequency; A time domain noise reduction module that performs phase compensation and windowing processing on the expanded frequency domain tensor and then performs an inverse Fourier transform, enhances the time domain resolution through interpolation processing, and uses median filtering to suppress transient noise to generate a transient time domain image; A signal classification module that extracts time domain waveform features from the transient time domain image, divides the confidence region through a confidence threshold, and combines multi-peak detection and random forest to distinguish valid reflection signals and outputs a depth probability map; A depth optimization module that performs channel-space dual attention weighting on the depth probability map, models neighborhood correlation using graph convolution, and implements hybrid coding to output a noise-resistant enhanced depth map.
9. A computer device, comprising a memory and a processor, the memory storing a computer program, characterized in that: When the processor executes the computer program, it implements the steps of the camera image data processing method based on noise reduction technology according to any one of claims 1 to 7.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by the processor, it implements the steps of the camera image data processing method based on noise reduction technology according to any one of claims 1 to 7.
Citation Information
Patent Citations
Distance measurement method, terminal and storage medium
CN113219476A
Flight time depth image iterative optimization method based on convolutional neural network
CN113240604A
Method for filtering multipath interference of indirect time-of-flight camera and camera
CN117528263A
Three-dimensional time-of-flight vision system
CN119556301A
Dam safety monitoring system and method based on digital twinning
CN119848786A