A depth image data processing stabilization method, device and storage medium
Patent Information
- Application Number
- CN202611162575.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-08-03
- Publication Date
- 2026-09-18
AI Technical Summary
1.在消除噪点的同时,容易导致物体边缘模糊,影响后续识别精度
通过同步获取TOF原始深度图与辅助影像并进行像素级对齐,利用辅助影像的局部梯度强度构建空间可信度权重,使得加权引导滤波过程在辅助影像存在清晰边缘而深度图数据破碎时,强制深度边缘向真实结构边缘收敛,有效解决了深度图边缘失真问题;基于物体移动特征构建时间模型,依据可信度评分设定深度变化预测范围,当检测到的深度变化超出该范围时即判定为噪声干扰并执行平滑抑制,在抑制时序噪声跳变的同时,保障了深度数据序列的稳定性。
Smart Images

Figure CN122776989A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of image enhancement, and in particular to a method, apparatus, and storage medium for stabilizing depth image data processing. Background Technology
[0002] With the widespread application of active depth sensing technologies such as TOF and structured light, three-dimensional spatial perception has become the core of human-computer interaction. However, in practical applications, factors such as low light, strong reflection, and rapid movement make it difficult to maintain the stable quality of the original depth map output by the sensor.
[0003] Traditional depth data processing often relies on spatial filtering or simple temporal smoothing using a single depth channel, which leads to the following problems: 1. While eliminating noise, it can easily lead to blurred object edges, affecting the accuracy of subsequent recognition.
[0004] 2. Simple time averaging methods can produce significant depth residue or delay when objects are moving rapidly.
[0005] 3. In existing technologies, some techniques attempt to fill in the missing depth through interpolation, but this can easily generate erroneous pseudo-data in complex environments, misleading the control logic.
[0006] 4. It cannot distinguish between vibrations caused by environmental disturbances and actual physical displacement.
[0007] We propose a depth image data processing stabilization method, device, and storage medium to address the aforementioned problems. By introducing structural gradients provided by a heterogeneous auxiliary sensing module as guidance, we utilize auxiliary images to perform dynamic reliability modeling of depth pixels, perform structure-preserving filtering in the spatial domain, and combine a temporal consistency model to adaptively suppress depth jumps. This aims to improve the spatiotemporal stability of the depth map without changing the original ranging principle or generating false pixels; preserve the edge structure of objects; eliminate flickering in low signal-to-noise ratio regions; and provide highly consistent and low-latency basic stable layer data for upper-layer applications. Summary of the Invention
[0008] The purpose of this invention is to provide a method, apparatus, and storage medium for stabilizing depth image data processing, in order to solve the problems mentioned in the background art.
[0009] To achieve the above objectives, the present invention provides the following technical solution: a method for stabilizing depth image data processing, comprising the following steps: S1: Acquire the original depth map collected by the depth sensor and the auxiliary image collected by the auxiliary sensing module, and perform time synchronization and spatial alignment processing on the original depth map and the auxiliary image to obtain the pixel-level corresponding depth map and auxiliary image. S2: Based on the aligned auxiliary image and depth map, construct a cross-modal structural consistency constraint model, including: Calculate the image gradient of the auxiliary image and the depth gradient of the depth map respectively; Based on the difference between image gradient and depth gradient, pixel-level consistency weights are generated to characterize structural consistency. S3: Based on consistency weights and local spatial neighborhood information of the depth map, perform adaptive spatial optimization processing on the original depth map with structural constraints to obtain a spatially optimized depth map. In the case of structurally consistent regions, the original depth value is maintained, and neighborhood smoothing is performed in structurally inconsistent regions. S4: Construct a motion state determination model based on continuous multi-frame depth map data, classify the depth changes of the current pixel, and obtain the corresponding motion state categories. The motion state categories include at least stable state, continuous motion state, and abnormal jump state. S5: Based on the motion state category, determine the corresponding temporal fusion strategy and perform adaptive temporal fusion processing on the spatial optimized depth map and the historical depth map; S6: Jointly fuse the spatial optimization results, temporal fusion results, and structural consistency constraints to output the final stable depth map and update the historical frame data.
[0010] Preferably, step S2 specifically involves: Based on the calibrated external parameters and relative positions between cameras, coordinate transformation and reprojection are performed through the joint calibration parameters between the depth sensor and the auxiliary sensing module to spatially align the original depth map with the auxiliary image. Based on the hardware synchronization signal, the hardware trigger line ensures that the two sensors are exposed at the same physical moment, ensuring that the original depth map and the auxiliary image are captured from the same physical moment for alignment processing. The process of generating consistency weights in step S2 includes: When the gradient direction or magnitude of the auxiliary image meets the preset consistency condition with the gradient direction or magnitude of the depth map, the consistency weight of that pixel position is increased. When the two do not meet the consistency condition, the consistency weight of that pixel position is reduced. Step S2 also includes: The auxiliary images are preprocessed to suppress high-frequency noise, and the image gradient is calculated based on the preprocessed images to avoid the influence of noise on the structural consistency determination.
[0011] Preferably, step S3 specifically includes calculating the gradient magnitude map G of the auxiliary image, calculating the gradient of the aligned auxiliary image in the horizontal and vertical directions, and calculating the comprehensive gradient magnitude based on the gradient.
[0012] Preferably, step S3 specifically includes calculating the spatial confidence weight map based on the gradient magnitude map G. The comprehensive gradient magnitude is input into a preset weight function. The output value of the weight function is close to the first value when the gradient magnitude is large and close to the second value when the gradient magnitude is small, and smoothly transitions in between. The output value constitutes a spatial credibility weight map. Read the weight corresponding to this position ; The closer the value is to 1, the more reliable the depth at that location becomes. The final depth output will almost entirely use the original depth value, thus protecting the edges. As the depth approaches 0, the more it needs to be smoothed, and the final depth output uses more of the average depth of its surrounding pixels to smooth out noise. Use the intermediate value as a reference and mix them proportionally; The space optimization process in step S3 includes: Based on consistency weight, the original depth value and the depth statistics in its local neighborhood are weighted and fused. The weighted fusion retains the original depth when the consistency weight is high and enhances neighborhood smoothing when the consistency weight is low.
[0013] Preferably, step S4 specifically includes: using a spatial confidence weight map to perform spatial guided filtering on the aligned depth map, performing weighted guided filtering on a pixel-by-pixel basis, the algorithm independently processing each pixel in the depth map, and outputting a spatially stable depth map; For any pixel location, its output depth value is a weighted mixture of the original depth value and the average depth of the local neighborhood, with the weights determined by the value of that location in the spatial confidence weight map; The motion state determination model in step S4 includes: Continuity analysis is performed based on the changing trend between the current frame and at least one historical depth map. Local consistency analysis is performed based on the consistency of pixel changes within the spatial neighborhood; The pixel motion state is classified based on the results of continuity analysis and local consistency analysis.
[0014] Preferably, step S5 specifically includes calculating the absolute change T per pixel between the spatially stabilized depth map of the current frame and the final stabilized depth map of the previous frame in the historical frame buffer.
[0015] The preferred step S5 specifically includes: calculating the time fusion weight based on the absolute change amount T, and then making a decision based on the weight. The absolute change amount T is input into a function based on exponential decay to calculate the time fusion weight. The time fusion weight is close to the first weight value when the change amount is small and close to the second weight value when the change amount is large. Step S5 also includes, based on the time fusion weights The current frame and historical frames of the spatially stable depth map are fused together, and the final stable depth map of the current frame is output by fusion using a formula. The time fusion strategy in step S5 includes: Different time fusion methods are used for different motion state categories, including: Increase the weight of historical frames to improve stability in a stable state; Enhance the weight of the current frame during continuous motion to improve response speed; Suppress the impact of outliers in the current frame on the output results during abnormal transition states.
[0016] Preferably, the joint fusion process in step S6 includes: A fusion model that simultaneously considers spatial constraints, temporal constraints, and structural consistency constraints is constructed, and the spatial optimization results and temporal fusion results are weighted and combined.
[0017] A depth image data processing stabilization device, comprising: The data acquisition and alignment module is used to simultaneously acquire the raw depth map from the depth sensor and the auxiliary image from the auxiliary sensing module, and perform spatiotemporal alignment processing on the two to obtain the aligned depth map and the aligned auxiliary image. The weight calculation module is used to calculate the spatial confidence weight of each pixel location based on the aligned auxiliary image. The spatial filtering module is used to perform spatial guided filtering on the aligned depth map using the spatial credibility weight map to generate a spatially stable depth map. The temporal fusion module is used to calculate the temporal fusion weights based on the temporal changes between the current frame and historical frames in the spatially stable depth map, and then fuse the current frame and historical frames according to the temporal fusion weights to output the final stable depth map.
[0018] A readable storage medium storing multiple applications configured to be executed by one or more processors, wherein the multiple applications are configured to implement a depth image data processing stabilization method when executed by the processor.
[0019] The technical effects and advantages of this invention are as follows: By synchronously acquiring the original TOF depth map and auxiliary image and performing pixel-level alignment, spatial credibility weights are constructed using the local gradient intensity of the auxiliary image. This forces the depth edges to converge towards the real structure edges during the weighted guided filtering process when the auxiliary image has clear edges but the depth map data is fragmented, effectively solving the problem of depth map edge distortion. A temporal model is constructed based on object movement features, and the depth change prediction range is set according to the credibility score. When the detected depth change exceeds this range, it is judged as noise interference and smoothing suppression is performed. This suppresses temporal noise jumps while ensuring the stability of the depth data sequence.
[0020] By introducing an auxiliary sensing module to acquire auxiliary images with high texture details, a data-driven spatial credibility weight map is constructed based on the gradient magnitude of the auxiliary images. This weight map utilizes the positive correlation between gradient magnitude and structural saliency to automatically distinguish between high gradient edge regions and low gradient flat regions in the auxiliary images. In high gradient regions, the weights approach 1, and the filter output is dominated by the original depth value, thus avoiding the geometric contour shift and loss of details caused by averaging operations at the edges of the depth map in traditional filtering. In the low gradient region, the weights approach 0, and the filter output is dominated by the local neighborhood mean, which effectively suppresses the random noise and outliers inherent in the TOF depth map in flat regions. By using a gradient adaptive weighting mechanism, edge structure preservation and noise reduction of flat areas are achieved simultaneously without relying on manually preset thresholds. This overcomes the contradiction between noise reduction intensity and edge preservation accuracy when traditional spatial filtering processes TOF depth data.
[0021] The temporal fusion mechanism based on spatial credibility weights calculates the inter-frame depth change per pixel using a spatially optimized depth map and adaptively calculates the temporal fusion weights accordingly. When the depth change is low, it is determined to be a stable state or slow motion, and the weight approaches 1. At this time, the fusion result is dominated by the current frame, thereby reducing the lag in motion response at the algorithm level. When the depth change suddenly increases beyond the prediction range based on credibility scores, it is determined to be an abnormal jump, and the weight approaches 0. At this time, the fusion result is forced to revert to the historical stable value. The continuity of historical data is used to block the transmission of single-frame noise, and the high-weight response ensures the tracking accuracy of the continuous motion trajectory of the object. The low-weight truncation eliminates the temporal flicker caused by sudden sensor noise, effectively solving the problem of the difficulty in balancing stability and real-time performance in dynamic scenes for TOF depth flow.
[0022] By integrating active depth sensing and auxiliary visual sensing, and leveraging the high-fidelity texture capture capability of the auxiliary sensing module in low-light environments, reliable guidance is provided for depth map optimization. The combined effect of spatiotemporal dual stabilization mechanisms effectively overcomes the inherent defects of a single TOF sensor in low-light environments, such as signal-to-noise ratio degradation, failure to match weak texture regions, and depth jumps on highly reflective surfaces, thereby improving the geometric measurement accuracy and data integrity in complex optical environments.
[0023] By providing well-defined and spatiotemporally stable depth data at the underlying level, the development cycle is shortened. The non-generative data processing mechanism eliminates the risk of false geometric structures introduced by interpolation completion, ensuring the physical authenticity and geometric fidelity of the depth data. This reduces the difficulty for upper-layer applications to adapt to the sensor's underlying noise model, eliminating the need for complex preprocessing for specific depth sensor noise models and improving the overall reliability of the vision system in complex scenarios.
[0024] All processing steps are optimized and fused based on the actual data collected by the sensors, adhering to the principle of not generating new depth values. That is, no depth value interpolation or speculation is performed out of thin air, which minimizes the risk of introducing false objects or geometric structures due to incorrect predictions and helps to ensure the reliability of the output depth data. Attached Figure Description
[0025] Figure 1 This is an overall flowchart of the present invention; Figure 2 This is a flowchart of the implementation method of the present invention; Figure 3 This is a flowchart of the data acquisition and alignment process of the present invention; Figure 4 This is a flowchart of the credibility weight calculation process for this invention; Figure 5 This is a flowchart of the spatial domain guided filtering process of the present invention; Figure 6 This is a flowchart illustrating the timing consistency constraint of the present invention. Detailed Implementation
[0026] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0027] This invention provides, for example Figures 1-6 The method for stabilizing depth image data, as shown, includes the following steps: S1: Acquire the original depth map collected by the depth sensor and the auxiliary image collected by the auxiliary sensing module, and perform time synchronization and spatial alignment processing on the original depth map and the auxiliary image to obtain the pixel-level corresponding depth map and auxiliary image. S2: Based on the aligned auxiliary image and depth map, construct a cross-modal structural consistency constraint model, including: Calculate the image gradient of the auxiliary image and the depth gradient of the depth map respectively; Based on the difference between image gradient and depth gradient, pixel-level consistency weights are generated to characterize structural consistency. S3: Based on consistency weights and local spatial neighborhood information of the depth map, perform adaptive spatial optimization processing on the original depth map with structural constraints to obtain a spatially optimized depth map. In the case of structurally consistent regions, the original depth value is maintained, and neighborhood smoothing is performed in structurally inconsistent regions. S4: Construct a motion state determination model based on continuous multi-frame depth map data, classify the depth changes of the current pixel, and obtain the corresponding motion state categories. The motion state categories include at least stable state, continuous motion state, and abnormal jump state. S5: Based on the motion state category, determine the corresponding temporal fusion strategy and perform adaptive temporal fusion processing on the spatial optimized depth map and the historical depth map; S6: Jointly fuse the spatial optimization results, temporal fusion results, and structural consistency constraints to output the final stable depth map and update the historical frame data.
[0028] Step S1 specifically involves acquiring an original depth map using a depth sensor and an auxiliary image using an auxiliary sensing module. The depth sensor is used to output an original depth map matrix containing depth information; the auxiliary sensing module is used to output an auxiliary image that can maintain texture and structure contrast even under low light conditions. The original depth map and the auxiliary image are then synchronized in time and aligned in space to obtain a pixel-level corresponding depth map and auxiliary image. The auxiliary images are used exclusively for depth image stabilization guidance and are not used for subsequent target recognition or classification processing.
[0029] Step S2 specifically involves the system simultaneously acquiring the original depth map from the depth sensor and the auxiliary image from the auxiliary sensing module, and then performing a spatial coordinate system transformation to achieve pixel-level alignment. This is because the optical centers and viewing angles of the two sensors are different, and the pixel coordinates of the same physical point are not the same in the two images.
[0030] It is necessary to ensure that the original depth map and the auxiliary image data capture the scene at the same moment.
[0031] Spatiotemporal synchronization specifically refers to the relative position between cameras based on calibrated external parameters. These parameters are a set of precise parameters obtained through joint calibration between the depth sensor and the auxiliary sensing module after the equipment is installed. Specifically, they include a rotation matrix and a translation vector, which describe the relative position and orientation relationship between the depth camera and the auxiliary camera. Input the relative positions between the calibrated external parameters of the cameras, perform coordinate transformation and reprojection, and use the rotation matrix and translation vector mentioned above to project each three-dimensional point in the depth map into the two-dimensional pixel coordinate system of the auxiliary image through perspective geometric transformation, thereby establishing a one-to-one correspondence between pixels in the two images. The time alignment process is as follows: input a hardware synchronization signal, perform frame buffer pairing, ensure that the two sensors are exposed at the same physical moment through a hardware trigger line, ensure that the original depth map and the auxiliary image for alignment processing are captured at the same physical moment, retrieve the frame depth map and auxiliary image captured at the same moment from their respective buffer queues for pairing, and eliminate motion artifacts caused by sampling time difference. After completing the spatiotemporal synchronization and alignment of the original depth map from the depth sensor and the auxiliary image from the auxiliary sensing module, the system obtains the transformation relationship between the two sensor coordinate systems. During online operation, this transformation relationship is used to project each pixel of the depth map onto the coordinate system of the auxiliary image, thereby establishing a pixel-level mapping relationship and obtaining the aligned depth map and auxiliary image. For any pixel coordinate, depth value, and image brightness value in the output image pair, they all come from the same point in the real world, providing a corresponding relationship for subsequent use of the clear edges of the auxiliary image to guide the depth map filtering. The depth sensor and black light camera are rigidly fixed on the same device to ensure that their relative poses remain unchanged during operation. A specific calibration plate is simultaneously placed in the field of view of both sensors to acquire multiple sets of synchronized image pairs under different poses. The rotation matrix and translation vector, as well as the intrinsic parameters and distortion coefficients of both are calculated and optimized using a calibration algorithm. Connect the trigger interface of the depth camera and the auxiliary camera, configure one of them as the master device and the other as the slave device, and send the exposure trigger signal by the master device. If hardware synchronization is not possible, it is necessary to ensure that the two camera drivers can provide timestamps and implement the time-stamp-based frame matching logic in the software layer. When the system is running, it synchronously triggers two sensors to collect data through hardware signals. It reads depth map and auxiliary image frame data from the drivers of the two sensors respectively, adds timestamps, and puts them into their respective frame buffers. From the buffers, it finds the pair of depth map and auxiliary image with the smallest timestamp difference as the data at the same time. Using the rotation matrix, translation vector, and camera intrinsic parameters obtained from offline calibration, a dense mapping table from the depth map pixel coordinate system to the auxiliary image pixel coordinate system is pre-calculated. For each valid pixel in the depth map, the dense mapping table calculated in the previous step is used to find its corresponding sub-pixel coordinates in the auxiliary image. Methods such as bilinear interpolation are used to obtain the pixel value at the corresponding position from the auxiliary image.
[0032] Finally, an aligned auxiliary image with the same size as the depth map and corresponding pixel positions is generated.
[0033] The process of generating consistency weights in step S2 includes: Calculate the gradient information, including gradient direction and gradient magnitude, at the same pixel position in the aligned depth map and auxiliary image respectively.
[0034] Specifically, this includes determination of directional consistency and determination of amplitude consistency: Directional consistency is determined based on the angle between gradient directions: Calculate the angle θ between the gradient direction of the auxiliary image and the gradient direction of the depth map. A preset angle threshold θ is set, which is the upper limit of the difference between the gradient direction of the auxiliary image and the gradient direction of the depth map. Its value ranges from 0 degrees to 30 degrees. Preferably, the directional angle threshold θ is between 5 and 20 degrees. When the directional angle threshold θ is set to 15 degrees, the angle between the gradient direction of the auxiliary image and the gradient direction of the depth map is less than the preset threshold. It is determined that the two modalities have captured the same physical edge structure at the pixel point, thus satisfying the directional consistency condition.
[0035] Amplitude consistency is determined based on gradient amplitude ratio or amplitude difference: The following two determination methods or a combination thereof shall be used for determination: 1. Gradient magnitude ratio determination: Calculate the magnitude ratio Rmag of the aligned auxiliary image gradient to the gradient magnitude of the depth map. The preset reasonable range for the magnitude ratio Rmag is 0.3 to 3.0. Preferably, the amplitude ratio Rmag is in the range of 0.5 to 2.0; when the amplitude ratio Rmag is set to the range of 0.8 to 1.25, and the calculated amplitude ratio falls within the above-mentioned preset range, it indicates that the sensing intensity of the two modes on the edge structure is matched, and is not affected by sudden noise or strong reflection interference, thus satisfying the amplitude consistency condition.
[0036] 2. Gradient magnitude difference determination: The gradient magnitude difference is evaluated by calculating the absolute difference ΔM between the gradient magnitude of the normalized auxiliary image and the gradient magnitude of the depth map. The gradient magnitude difference threshold ΔM ranges from 5 to 40. Preferably, the gradient magnitude difference threshold ΔM is 10 to 25; when the gradient magnitude difference threshold ΔM is set to 15, the amplitude consistency condition is determined to be met when the absolute difference between the gradient magnitudes of the two is less than the gradient magnitude difference threshold ΔM.
[0037] The weights are adjusted based on the above determination results: When the gradient direction of the auxiliary image and the gradient direction of the depth map meet the direction consistency condition, and the gradient magnitude meets the magnitude consistency condition, that is, when the magnitude ratio is within a preset range or the magnitude difference is less than a preset threshold, the consistency weight of the pixel position is increased. If any of the above conditions are not met, i.e., the directional difference is too large or the amplitude response is mismatched, the consistency weight of the pixel position is reduced.
[0038] The original depth map may also undergo the same coordinate transformation and be resampled during this process, resulting in an aligned depth map as the output. Step S3 specifically involves outputting the aligned auxiliary image, calculating the structural gradient, generating confidence weights, and outputting a spatial confidence weight map. The formula for calculating the structural gradient is:
[0039] The output value represents the gradient magnitude of the auxiliary image at time t and pixel position, with a value range of [0, 500]. The specific value depends on the image contrast and resolution and can be modified. These are pixel coordinates; This is the original depth image; The input value represents the auxiliary image acquired at time t, which is aligned with the depth map and normalized to [0, 1]. To assist in the gradient of the image in the horizontal x-direction, representing the intensity of the edge in the vertical direction; To assist in determining the gradient of the image in the vertical direction y, representing the intensity of the edge in the horizontal direction, the square root of the sum of the squares of the gradients in both directions is taken to obtain the comprehensive gradient magnitude G. The larger the gradient magnitude, the sharper and higher the contrast of the edge in the image.
[0040] Before calculating the structural gradient, the input image can be slightly smoothed to suppress thermal noise or fine textures of the image sensor itself, preventing these noises from being misjudged as important edges in subsequent gradient calculations. A Gaussian convolution kernel is then applied to filter the image, preserving the main edges while filtering out high-frequency noise. That is, step S3 also includes: The auxiliary images are preprocessed to suppress high-frequency noise, and the image gradient is calculated based on the preprocessed images to avoid the influence of noise on the structural consistency determination.
[0041] The formula for generating spatial credibility weights is as follows:
[0042] The output value represents the spatial confidence weight at time t and pixel location, with a value range of [0, 1]; G is the gradient magnitude calculated in the previous step. It is a stability constant used to avoid the denominator being zero.
[0043] When G is large Approximately equal to 1; when G=0, Equals 0; when G is a medium value, Smooth transition between 0 and 1; In the downstream spatial guided filtering step, the algorithm will traverse every pixel of the depth map and read the weight corresponding to that location. ;if A value closer to 1 indicates a more reliable depth at that location, meaning the final depth output will almost entirely use the original depth value, thus preserving the edges; if... A value close to 0 indicates that the depth at that location requires more smoothing, and the final depth output will use more of the average depth of its surrounding pixels to smooth out noise; if... If the value is the median, then mix them proportionally; Credibility weight It is completely data-driven and does not require manually setting fixed thresholds to distinguish between edges and flat areas, thus solving the problem that traditional filters inevitably blur edges during noise reduction. That is, the space optimization process in step S3 includes: Based on consistency weight, the original depth value and the depth statistics in its local neighborhood are weighted and fused. The weighted fusion retains the original depth when the consistency weight is high and enhances neighborhood smoothing when the consistency weight is low.
[0044] Step S4 specifically includes the following steps: by inputting the aligned original depth map and the spatial confidence weight map, the original depth map and the spatial confidence weight map structure are processed pixel by pixel; The aligned raw depth map is the raw data from a TOF or structured light sensor, which has been precisely aligned with the auxiliary image, but contains noise and may have blurred edges; the spatial confidence weight map is the direct output of confidence weight modeling, where each pixel value indicates the confidence level of the depth value at the corresponding location.
[0045] Spatial guided filtering is applied to the aligned depth map using a spatial confidence weight map. The process involves pixel-by-pixel weighted guided filtering, with each pixel in the depth map processed independently. The resulting spatially stabilized depth map is defined by the following formula:
[0046] The output value represents the spatially stabilized depth value at time t and pixel position. Input the value to calculate the current spatial credibility weight; Input value, representing the aligned raw depth value acquired at time t; Local spatial neighborhood centered on the current location Within, the average of all effective depth values; when If the value is approximately 1, it is in a high-weight region. According to the definition of the weight map, this corresponds to a region in the auxiliary image with clear structure and sharp edges. It almost completely preserves the original depth value, strongly protects the clear outline of the object, and prevents the edges from becoming blurred during the filtering process. when If the value is approximately 0, it falls within a low-weight region, corresponding to a flat, uniformly textured area in the auxiliary image. The average depth of the surrounding pixels is used as the output, effectively smoothing out noise in this region. Random errors can be suppressed by averaging.
[0047] when When the value is an intermediate value, the output is a linear mixture of the original value and the local average, achieving a natural transition.
[0048] After this step, the spatially stabilized depth map is processed, which suppresses noise in the spatial domain and preserves the physical edges of objects. This provides the input basis for the next step of temporal processing.
[0049] By inputting the spatially stable result of the current frame, which is the direct output of the spatial domain guided filtering step, it is the depth map after spatial domain denoising and edge preservation processing. At the same time, the historical frame buffer is input, which stores the final stable output of the previous frame and is used to compare with the current frame. Calculate the absolute change T per pixel between the current frame of the spatially stable depth map and the final stable depth map of the previous frame in the history frame buffer; The calculation formula is as follows:
[0050] The output value represents the absolute change in depth between frames at a pixel position at time t; Input value: the spatially stabilized depth value of the current frame; Input value: the spatially stabilized depth value of the previous frame.
[0051] A matrix of the same size as the depth map, with each pixel value T represents the absolute change in depth value at that location between adjacent frames. A large T value indicates a drastic change, which may be due to rapid movement or sudden noise; a small T value indicates a slight change, which may be due to stillness or slow movement.
[0052] That is, the motion state determination model in step S4 includes: Continuity analysis is performed based on the changing trend between the current frame and at least one historical depth map. Local consistency analysis is performed based on the consistency of pixel changes within the spatial neighborhood; The pixel motion states are classified based on the results of continuity analysis and local consistency analysis. The specific judgment rules are as follows: First, calculate the spatially stable depth map of the current frame. Final stable depth map compared to the previous frame absolute change per pixel , where p represents the pixel coordinate; Centered on the current pixel p, define a local neighborhood of size M x M, and calculate the absolute change of all pixels within this neighborhood; The criteria for determining the motion state category are: Steady state: if And the variance of pixel changes within the neighborhood is less than If so, it is determined to be a stable state; Continuous motion state: if Furthermore, the direction of change of the current pixel p is consistent with the direction of change of pixels in its neighborhood, and the change amount is consistent across N consecutive frames. If it shows a monotonically increasing or decreasing trend, it is determined to be a continuous motion state; Abnormal transition state: If Furthermore, it does not meet the continuous motion condition, or the consistency of changes within the neighborhood exceeds a preset threshold. This is determined to be an abnormal transition state; The threshold for distinguishing between a steady state and a state of motion; The threshold for determining abnormal transition states; This represents the upper limit of the variance of the neighborhood change under steady-state conditions. The lower bound of the variance of the neighborhood change under abnormal transition states; N is the number of frames required for continuous motion state determination; where , , , Both N and are preset parameters, and < , < .
[0053] The process involves comparing the current frame with the previous frame, calculating the pixel-by-pixel absolute difference, determining the temporal fusion weights based on the differences, and then making a decision based on these weights. The formula for calculating the adaptive time fusion weights is:
[0054] The output value represents the time-domain fusion weights, with a value range of [value range missing]. The input value T represents the absolute change in depth between frames at time t and pixel position. The attenuation coefficient is an adjustable parameter in this step; it is a constant greater than 0; e is the natural constant. The input value is the spatially stable depth of the current frame; The input value is the spatially stable depth of the previous frame; When T is small Approximately equal to 1; when T is large Approximately equal to 0; The role of controlling time fusion weights The rate of decay as the change in amount T increases The larger the value, the faster the decay, the more conservative the algorithm, the more it tends to believe in history, the stronger the smoothing force, and the higher the latency. The smaller the value, the slower the decay, the more aggressive the algorithm, the more it tends to trust the current frame, the faster the response, but the weaker the noise suppression ability.
[0055] The time fusion strategy in step S5 includes: Different time fusion methods are used for different motion state categories determined in step S4, and the specific correspondence is as follows: Under steady-state conditions, time fusion weights Approaching 1, the fusion result is based on the current frame. The primary goal is to maximize the suppression of minute noise in static scenes; In continuous motion, time fusion weights It decreases as the absolute change between frames increases, so that the fusion result is adaptively balanced between the current frame and the historical frames, ensuring a fast response to continuous motion while taking into account a certain degree of temporal smoothness. In abnormal transition states, time fusion weights The value is forced to 0, and the fusion result depends entirely on historical frames. This effectively blocks the interference of single-frame abnormal noise on the output results.
[0056] The weighted intelligent decision-making and fusion formula is as follows:
[0057] The output value represents the final spacetime stable depth. The input value is the spatially stable depth of the current frame; The input value is the spatially stable depth of the previous frame; The input value is the time fusion weight; If T is small Approximately equal to 1, the change is minimal, it could be noise, and it strongly depends on the current frame. Approximately equal to For regions where the depth value remains almost constant or changes slowly, the small jumps in the current frame are considered to be residual noise. The current frame is favored, and the system responds quickly to real motion to avoid motion blur.
[0058] If T is large Approximately equal to 0, fluctuating drastically, representing real motion or abrupt noise, strongly dependent on the previous frame, output Approximately equal to For pixels whose depth values change drastically, there are two possibilities: first, the object is actually moving rapidly; second, there is sudden noise such as reflections. As a safe, foundational stabilization layer, the preferred strategy in this step is to strongly suppress jumps, that is, to highly trust historical stable values. This is to avoid transmitting abrupt noise from a single frame.
[0059] If the depth change exceeds the normal fluctuation range predicted based on historical data, it is judged as an abnormal jump and suppression processing is performed. At this time, the time fusion weights... When forced to 0, the output depends entirely on historical frames. This prevents the outlier from contaminating the final result; The data is fused according to weights, and the final stable depth map is output, updated, and then stored in the historical frame buffer.
[0060] That is, the joint fusion process in step S6 includes: A fusion model that simultaneously considers spatial constraints, temporal constraints, and structural consistency constraints is constructed, and the spatial optimization results and temporal fusion results are weighted and combined.
[0061] Temporal consistency constraints are implemented to combat flickering and jittering of depth values, distinguish between normal object movement and abnormal noise jumps. For slow or continuous movement, the introduced delay is minimal; for sudden and unreasonable changes in depth values, decisive smoothing suppression is achieved. The output depth map sequence is smoother in time, reduces flickering, and has the ability to respond to the real movement of objects.
[0062] By improving data quality, we provide upper-layer applications with well-defined and spatiotemporally stable depth data, thus improving their input conditions. We also reduce development complexity, as upper-layer applications do not need to handle complex depth sensor noise individually and can focus on the core logic of recognition and navigation. Furthermore, we enhance system robustness, helping to improve the robustness of the vision system under conditions such as low light and strong reflection. The entire process strictly adheres to the principle of not generating new depth data, only optimizing existing data, thus avoiding the risk of introducing false information due to incorrect interpolation.
[0063] A depth image data processing stabilization device includes: The data acquisition and alignment module is used to synchronously acquire the original depth map acquired by the depth sensor and the auxiliary image acquired by the auxiliary sensing module, and to perform spatiotemporal synchronization and pixel-level alignment processing on the original depth map and the auxiliary image to obtain the aligned depth map and the aligned auxiliary image. The weight calculation module is used to calculate the structural gradient of the aligned auxiliary image and generate a spatial confidence weight map based on the structural gradient. The spatial filtering module is used to perform spatial guided filtering on the aligned depth map using the spatial credibility weight map to generate a spatially stable depth map. The temporal fusion module is used to calculate the temporal fusion weight based on the temporal changes between the current frame and historical frames of the spatially stable depth map, and to fuse the spatially stable depth map of the current frame with the historical stable depth map according to the temporal fusion weight, outputting the final spatiotemporally stable depth map and updating the historical stable depth map.
[0064] A readable storage medium, which may be a standalone SDK or embedded firmware, wherein the algorithm may be packaged as a standalone SDK or embedded firmware and deployed in an edge computing box, an on-board SOC or a robot controller, storing multiple applications and configured to be executed by one or more processors, wherein the multiple applications are configured to implement a depth image data processing stabilization method when executed by the processor.
[0065] Finally, it should be noted that the above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art can still modify the technical solutions described in the foregoing embodiments or make equivalent substitutions for some of the technical features. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the protection scope of the present invention.
Claims
1. A method of depth image data processing stabilization, characterized in that, Includes the following steps: S1: Acquire the original depth map collected by the depth sensor and the auxiliary image collected by the auxiliary sensing module, and perform time synchronization and spatial alignment processing on the original depth map and the auxiliary image to obtain the pixel-level corresponding depth map and auxiliary image; S2: Based on the aligned auxiliary image and depth map, construct a cross-modal structural consistency constraint model, including: Calculate the image gradient of the auxiliary image and the depth gradient of the depth map respectively; Based on the difference between the image gradient and the depth gradient, pixel-level consistency weights are generated to characterize structural consistency. S3: Based on the consistency weight and the local spatial neighborhood information of the depth map, perform adaptive spatial optimization processing on the original depth map with structural constraints to obtain a spatially optimized depth map, wherein the original depth value is maintained in the structurally consistent region and neighborhood smoothing is performed in the structurally inconsistent region. S4: Construct a motion state determination model based on continuous multi-frame depth map data, classify the depth changes of the current pixel, and obtain the corresponding motion state category. The motion state category includes at least stable state, continuous motion state, and abnormal jump state. S5: Based on the motion state category, determine the corresponding time fusion strategy and perform adaptive time fusion processing on the spatial optimized depth map and the historical depth map; S6: Jointly fuse the spatial optimization results, temporal fusion results, and the structural consistency constraints to output the final stable depth map and update the historical frame data.
2. The method of claim 1, wherein, Step S2 specifically involves: Based on the calibrated external parameters and relative positions between cameras, coordinate transformation and reprojection are performed through the joint calibration parameters between the depth sensor and the auxiliary sensing module. The original depth map and the auxiliary image are spatially aligned, and the pixel values of the corresponding positions are obtained from the auxiliary image. Finally, an aligned auxiliary image with the same size as the depth map and with corresponding pixel positions is generated. Based on the hardware synchronization signal, the hardware trigger line ensures that the two sensors are exposed at the same physical moment, ensuring that the original depth map and the auxiliary image are captured from the same physical moment for alignment processing. The process of generating consistency weights in step S2 includes: When the gradient direction or magnitude of the auxiliary image meets the preset consistency condition with the gradient direction or magnitude of the depth map, the consistency weight of that pixel position is increased. When the consistency conditions are not met, the consistency weight of the pixel position is reduced. Step S2 also includes: The auxiliary image is preprocessed to suppress high-frequency noise, and the image gradient is calculated based on the preprocessed image to avoid the influence of noise on the structural consistency determination.
3. The method of claim 1, wherein, Step S3 specifically includes calculating the gradient magnitude map G of the auxiliary image, calculating the gradient of the aligned auxiliary image in the horizontal and vertical directions, and calculating the comprehensive gradient magnitude based on the gradient.
4. The depth image data processing stabilization method according to claim 3, characterized in that, Step S3 specifically includes calculating the spatial confidence weight map based on the gradient magnitude map G. The comprehensive gradient magnitude is input into a preset weight function. The output value of the weight function is close to the first value when the gradient magnitude is large and close to the second value when the gradient magnitude is small, and smoothly transitions in between. The output value constitutes a spatial credibility weight map. For any pixel location in the spatial credibility weight map, read the weight corresponding to that pixel location. ; Read the weight corresponding to this position ; The closer the value is to 1, the more reliable the original depth value at that pixel location becomes. The final depth output will almost entirely use the original depth value, thus preserving the edges. The closer the value is to 0, the more the original depth value at that pixel location needs to be smoothed, and the final depth output uses more of the average depth of its surrounding pixels to smooth out noise; Use the intermediate value as a reference and mix them proportionally; The space optimization process in step S3 includes: Based on the consistency weight, the original depth value and the depth statistics in its local neighborhood are weighted and fused. The weighted fusion retains the original depth when the consistency weight is high and enhances neighborhood smoothing when the consistency weight is low.
5. The method for stabilizing depth image data according to claim 1, characterized in that, Step S4 specifically includes: using the spatial confidence weight map to perform spatial guided filtering on the aligned depth map, performing weighted guided filtering on a pixel-by-pixel basis, processing each pixel in the depth map independently, and outputting a spatially stable depth map. For any pixel location, its output depth value is a weighted mixture of the original depth value and the average depth of the local neighborhood, with the weights determined by the value of that location in the spatial confidence weight map; The motion state determination model in step S4 includes: Continuity analysis is performed based on the changing trend between the current frame and at least one historical depth map. Local consistency analysis is performed based on the consistency of pixel changes within the spatial neighborhood; The pixel motion state is classified based on the continuity analysis results and the local consistency analysis results.
6. The method for stabilizing depth image data according to claim 1, characterized in that, Step S5 specifically includes calculating the absolute change T per pixel between the spatially stabilized depth map of the current frame and the final stabilized depth map of the previous frame in the historical frame buffer.
7. The method for stabilizing depth image data according to claim 6, characterized in that, Step S5 specifically includes calculating the time fusion weight based on the absolute change amount T, and then making a decision based on the weight. The absolute change amount T is input into a function based on exponential decay to calculate the time fusion weight. The time fusion weight is close to the first weight value when the change amount is small and close to the second weight value when the change amount is large. Step S5 also includes, based on the time fusion weights The current frame and historical frames of the spatially stabilized depth map are fused together, and the final stable depth map of the current frame is output by fusion using a formula. The time fusion strategy in step S5 includes: Different time fusion methods are used for different motion state categories, including: Increase the weight of historical frames to improve stability in a stable state; Enhance the weight of the current frame during continuous motion to improve response speed; Suppress the impact of outliers in the current frame on the output results during abnormal transition states.
8. The method for stabilizing depth image data according to claim 7, characterized in that, The joint fusion process in step S6 includes: A fusion model that simultaneously considers spatial constraints, temporal constraints, and structural consistency constraints is constructed, and the spatial optimization results and temporal fusion results are weighted and combined.
9. A depth image data processing stabilization apparatus, applied to the depth image data processing stabilization method according to any one of claims 1-8, characterized in that, include: The data acquisition and alignment module is used to simultaneously acquire the raw depth map from the depth sensor and the auxiliary image from the auxiliary sensing module, and perform spatiotemporal alignment processing on the two to obtain the aligned depth map and the aligned auxiliary image. The weight calculation module is used to calculate the spatial confidence weight of each pixel location based on the aligned auxiliary image. The spatial filtering module is used to perform spatial guided filtering on the aligned depth map using the spatial credibility weight map to generate a spatially stable depth map. The temporal fusion module is used to calculate the temporal fusion weights based on the temporal changes between the current frame and historical frames of the spatially stabilized depth map, and then fuse the current frame and historical frames according to the temporal fusion weights to output the final stable depth map.
10. A readable storage medium, characterized in that, The system stores multiple applications and is configured to be executed by one or more processors, wherein the multiple applications are configured to implement a depth image data processing stabilization method as described in any one of claims 1-8 when executed by the processor.