Dynamic adaptive depth camera occlusion detection method, system and device, and storage medium

Through the dynamic adaptive depth camera occlusion detection method, the three-frame timing analysis window and dynamic threshold value are used, combined with speckle feature points and gradient change rate, space-time consistency comparison is performed, which solves the problem of inadaptable threshold setting of occlusion detection and insufficient utilization of space-time information in the prior art, and achieves high-precision and stable occlusion detection.

CN120182813APending Publication Date: 2025-06-20SHENZHEN GUANGJIAN TECH CO LTD +1
View PDF 0 Cites 9 Cited by

Patent Information

Application Number
CN202510192352.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-21
Publication Date
2025-06-20

AI Technical Summary

Technical Problem

The existing depth camera occlusion detection methods have problems such as threshold setting not adapted to complex scenarios, insufficient utilization of space-time information and low feature extraction efficiency, resulting in insufficient detection accuracy and reliability.

Method used

A dynamically adaptive depth camera occlusion detection method is proposed. By constructing a three-frame timing analysis window, a pixel-level spatiotemporal alignment mapping relationship is established, dynamic thresholds are calculated, speckle feature points are located, gradient change rate is calculated, and temporal consistency comparison is performed to mark and retain occlusion areas.

Benefits of technology

It improves the accuracy and stability of occlusion detection, can flexibly respond to changes in depth distribution in different scenarios, reduces misjudgment and missed detection, and meets the high-precision and real-time requirements in complex scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120182813A_ABST
    Figure CN120182813A_ABST
Patent Text Reader

Abstract

A dynamic adaptive depth camera occlusion detection method, system and device, and a storage medium, the method comprising: step T1: constructing a three-frame time sequence analysis window, continuously obtaining infrared speckle image sequences at different moments and corresponding dense depth map sequences, and establishing a pixel-level space-time alignment mapping relationship; t2, calculating a motion vector between adjacent frames, and executing the step T3 when the average motion intensity exceeds a dynamic threshold value; wherein the dynamic threshold is determined by the dense depth map; t3, positioning a speckle feature point in the infrared speckle image, generating an analysis window by taking the speckle feature point as a center, calculating a gradient change rate, and when the gradient change rate is smaller than a gradient threshold value, marking the region as a primary candidate region; and T4, carrying out space-time consistency comparison on the primary candidate region and two adjacent frames, and reserving a shielding region of which the variation value is smaller than a variation threshold value in three continuous frames. The method is high in precision, good in real-time performance and high in stability.
Need to check novelty before this filing date? Find Prior Art

Description

Background Art

[0002] In the fields of computer vision and robotics, depth cameras are widely used in many tasks such as object recognition, scene understanding, and motion tracking. Depth cameras can obtain the depth information of objects in the scene, providing key data support for subsequent analysis. However, in the actual application process, depth cameras face a severe challenge - the occlusion problem.

[0003] Existing depth camera occlusion detection methods have many deficiencies. Many traditional methods use fixed thresholds to judge occlusion situations, and this method cannot adapt to complex and changeable scenes. For example, in an indoor environment, the objects are relatively close to the camera, and the depth change is relatively small; while in an outdoor scene, the objects are far from the camera, and the depth range is wider. Fixed thresholds will either cause a large number of real occlusion areas to be missed in different scenes, or misjudge normal scene changes as occlusions, greatly reducing the accuracy and reliability of detection.

[0004] In addition, some detection methods lack the effective use of time series information. They often only focus on the features of single-frame images and ignore the motion and change rules of objects between consecutive frames. This makes the detection results in dynamic scenes vulnerable to interference and prone to frequent misjudgments and instability when objects move quickly or there are multiple objects occluding each other.

[0005] At the same time, traditional methods also have problems with low efficiency in feature extraction. Some algorithms require complex calculations and a large number of parameter adjustments, resulting in slow detection speeds and being unable to meet the requirements of application scenarios with high real-time requirements, such as autonomous driving and robot navigation.

[0006] The disclosure of the above background art content is only for assisting in understanding the inventive concept and technical solution of the present invention, and it does not necessarily belong to the prior art of this patent application. Without clear evidence indicating that the above content was publicly available on the filing date of this patent application, the above background art should not be used to evaluate the novelty and inventiveness of this application. Summary of the Invention

[0007] Therefore, the present invention proposes a dynamic adaptive depth camera occlusion detection method to solve the deficiencies of traditional methods in threshold setting, spatio-temporal information utilization, and feature extraction, overcome the limitations of existing depth camera occlusion detection methods, and meet the requirements of high precision, real-time performance, and stability in complex scenes.

[0008] In a first aspect, the present invention provides a dynamic adaptive depth camera occlusion detection method, characterized by including:

[0009] Step T1: Construct a three-frame time-series analysis window, continuously acquire infrared speckle image sequences and corresponding dense depth map sequences at different times, and establish a pixel-level spatio-temporal alignment mapping relationship;

[0010] Step T2: Calculate the motion vectors between adjacent frames, and execute Step T3 when the average motion intensity exceeds the dynamic threshold; wherein, the dynamic threshold is determined by the dense depth map;

[0011] Step T3: Locate the speckle feature points in the infrared speckle image, generate an analysis window centered on the speckle feature points, calculate the gradient change rate, and mark it as a primary candidate region when the gradient change rate is less than the gradient threshold;

[0012] Step T4: Perform spatio-temporal consistency comparison on the primary candidate region and the adjacent two frames, and retain the occlusion regions with change values less than the change threshold in three consecutive frames.

[0013] Optionally, in the above-mentioned dynamic adaptive depth camera occlusion detection method, it further includes:

[0014] Step T5: Obtain the sparse depth map of the infrared speckle image; align the infrared speckle image with the sparse depth map;

[0015] Step T6: Remove the depth values of the occlusion regions in the dense depth map, and fill the depth values of the corresponding regions in the sparse depth map.

[0016] Optionally, in the above-mentioned dynamic adaptive depth camera occlusion detection method, when calculating the motion vectors between adjacent frames, an automatic window confirmation matching range is set for the infrared speckle images of adjacent frames; the size of the automatic window is determined by the depth change rate on the dense depth map.

[0017] Optionally, in the above-mentioned dynamic adaptive depth camera occlusion detection method, the dynamic threshold is proportional to the average depth of the dense depth map.

[0018] Optionally, in the above-mentioned dynamic adaptive depth camera occlusion detection method, Step T3 includes:

[0019] Step T31: Binarize the infrared speckle image to obtain the light spot region, and further obtain the light spot feature points;

[0020] Step T32: Generate an analysis window centered on the light spot feature points and bounded by the light spot region;

[0021] Step T33: Within the analysis window, calculate the gradient change value of each pixel point using a gradient operator, and calculate the standard deviation of the gradient change values of all pixel points within the analysis window as the gradient change rate;

[0022] Step T34: When the gradient change rate is less than the gradient threshold, mark it as a primary candidate region.

[0023] Optionally, in the dynamic adaptive depth camera occlusion detection method, Step T4 includes:

[0024] Step T41: According to the spatio-temporal alignment mapping relationship, extract the regions corresponding to the primary candidate regions in two adjacent frames;

[0025] Step T42: Calculate the brightness difference between the primary candidate region and the corresponding regions in two adjacent frames to obtain a change value;

[0026] Step T43: Retain the primary candidate regions with a change value less than the change threshold as occlusion regions.

[0027] Optionally, in the dynamic adaptive depth camera occlusion detection method, in Step T43, morphological analysis is also performed on the primary candidate regions to remove noise interference.

[0028] In a second aspect, the present invention provides a dynamic adaptive depth camera occlusion detection system for implementing the dynamic adaptive depth camera occlusion detection method described in any one of the foregoing, which includes:

[0029] An acquisition module, configured to construct a three-frame time-sequence analysis window, continuously acquire infrared speckle image sequences and corresponding dense depth map sequences at different times, and establish a pixel-level spatio-temporal alignment mapping relationship;

[0030] A motion module, configured to calculate the motion vectors between adjacent frames, and execute the candidate module when the average motion intensity exceeds the dynamic threshold; wherein, the dynamic threshold is determined by the dense depth map;

[0031] A candidate module, configured to locate speckle feature points in the infrared speckle image, generate an analysis window centered on the speckle feature points, calculate the gradient change rate, and mark it as a primary candidate region when the gradient change rate is less than the gradient threshold;

[0032] A determination module, configured to perform spatio-temporal consistency comparison on the primary candidate region and two adjacent frames, and retain the occlusion regions with a change value less than the change threshold in three consecutive frames.

[0033] In a third aspect, the present invention provides a dynamic adaptive depth camera occlusion detection device, which includes:

[0034] A processor;

[0035] A memory, in which executable instructions of the processor are stored;

[0036] Wherein, the processor is configured to execute the steps of the dynamic adaptive depth camera occlusion detection method described in any one of the foregoing by executing the executable instructions.

[0037] In a fourth aspect, the present invention provides a computer-readable storage medium for storing a program, characterized in that when the program is executed, the steps of the dynamic adaptive depth camera occlusion detection method described in any one of the foregoing are implemented.

[0038] Compared with the prior art, the present invention has the following beneficial effects:

[0039] By constructing a three-frame time-series analysis window and establishing a pixel-level spatio-temporal alignment mapping relationship, the present invention can accurately associate the infrared speckle image with the dense depth map data, providing an accurate data basis for subsequent motion analysis and occlusion detection, and greatly improving the detection accuracy. For example, in a complex scene, this accurate association can clearly distinguish the motion and occlusion situations of different objects, avoiding misjudgment.

[0040] The dynamic threshold in the present invention is determined by the dense depth map and adjusted in real time according to the scene depth information. In different scenes, such as indoor close-range scenes and outdoor long-range scenes, the depth distributions are different, and the dynamic threshold can be automatically adapted, enabling the detection method to flexibly cope with various environments, effectively improving the detection accuracy and adaptability, and avoiding missed detection or false detection due to a fixed threshold in different scenes.

[0041] The present invention uses a feature point detection algorithm to locate the speckle feature points and generates an analysis window centered on them to calculate the gradient change rate, which can quickly and accurately capture the areas in the image that may be occluded. Algorithms like the FAST algorithm are fast and can locate a large number of feature points in a short time, providing a rich candidate for subsequent screening of occluded areas and improving the detection efficiency.

[0042] The present invention performs spatio-temporal consistency comparison on the primary candidate area and the adjacent two frames, and retains the occluded areas with a change value less than the change threshold. This method effectively eliminates misjudgments caused by normal object motion or interference factors, ensures the stability and reliability of the detected occluded areas, and improves the credibility of the final detection result. Description of the Drawings

[0043] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the accompanying drawings required for the description of the embodiments or the prior art. Obviously, the accompanying drawings in the following description are only the embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other accompanying drawings can also be obtained according to the provided drawings. By reading the detailed description of the non-restrictive embodiments with reference to the following drawings, other features, purposes, and advantages of the present invention will become more obvious:

[0044] Figure 1 It is a flowchart of the steps of a dynamic adaptive depth camera occlusion detection method in an embodiment of the present invention;

[0045] Figure 2 It is a flowchart of the steps of another dynamic adaptive depth camera occlusion detection method in an embodiment of the present invention;

[0046] Figure 3 It is a flowchart of the steps of marking a primary candidate region in an embodiment of the present invention;

[0047] Figure 4 It is a flowchart of the steps of marking an occlusion region in an embodiment of the present invention;

[0048] Figure 5 It is a schematic structural diagram of a dynamic adaptive depth camera occlusion detection system in an embodiment of the present invention;

[0049] Figure 6 It is a schematic structural diagram of a dynamic adaptive depth camera occlusion detection device in an embodiment of the present invention; and

[0050] Figure 7 It is a schematic structural diagram of a computer-readable storage medium in an embodiment of the present invention. Detailed Embodiments

[0051] The present invention will be described in detail below with reference to specific embodiments. The following embodiments will help those skilled in the art to further understand the present invention, but do not limit the present invention in any form. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present invention, several modifications and improvements can still be made. These all belong to the protection scope of the present invention.

[0052] In the description, claims and above-mentioned drawings of the present invention, terms such as "first", "second", "third", "fourth", etc. (if any) are used to distinguish similar objects and do not necessarily describe a specific order or sequence. It should be understood that the data used in this way can be interchanged under appropriate circumstances, so that the embodiments of the present invention described herein, for example, can be implemented in an order other than those illustrated or described herein. In addition, the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device comprising a series of steps or units does not necessarily limit to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products or devices.

[0053] A dynamic adaptive depth camera occlusion detection method provided by an embodiment of the present invention aims to solve the problems existing in the prior art.

[0054] The technical solutions of the present invention and how the technical solutions of this application solve the above technical problems will be described in detail below with specific embodiments. These several specific embodiments can be combined with each other, and the same or similar concepts or processes may not be repeated in some embodiments. The embodiments of the present invention will be described below with reference to the drawings.

[0055] By constructing a three-frame temporal analysis window and establishing a pixel-level spatio-temporal alignment mapping relationship, the present invention can make full use of the spatio-temporal information of the image sequence and the depth map sequence, laying a foundation for accurate detection; using the dense depth map to determine the dynamic threshold to calculate the motion vector between adjacent frames, which can be dynamically adjusted according to the actual scene, effectively adapting to different environmental changes; locating the speckle feature points in the infrared speckle image and calculating the gradient change rate to mark the primary candidate areas, which can accurately screen out the possible occlusion areas; performing spatio-temporal consistency comparison on the primary candidate areas and retaining the occlusion areas with the change value less than the change threshold, further improving the accuracy and stability of the detection results, and ensuring that the occlusion areas can be reliably detected even in complex dynamic scenes.

[0056] Figure 1 It is a step flow chart of a dynamic adaptive depth camera occlusion detection method in an embodiment of the present invention.

[0057] As Figure 1 shown, the steps of a dynamic adaptive depth camera occlusion detection method in an embodiment of the present invention include:

[0058] Step T1: Construct a three-frame temporal analysis window, continuously obtain infrared speckle image sequences and corresponding dense depth map sequences at different times, and establish a pixel-level spatio-temporal alignment mapping relationship.

[0059] In this step, to dynamically analyze the occlusion situation of the depth camera, three consecutive frames of data are selected as an analysis window. This can utilize the information in the time series and detect changes in occlusion by comparing images at different times. Compared with single-frame analysis, the three-frame window can better capture the movement of objects and the dynamic process of occlusion.

[0060] The depth camera can simultaneously acquire infrared speckle images and corresponding dense depth maps. The infrared speckle images contain the texture information of the object surface, and the movement of the object can be analyzed through the distribution and changes of the speckles. The dense depth map provides the depth information of each pixel point in the scene, that is, the distance between the object and the camera, which is crucial for detecting occlusion because occlusion usually leads to discontinuity in depth information. Continuously acquiring these image and depth map sequences at different times means that the data within the analysis window is updated over time, enabling real-time tracking of occlusion in the scene.

[0061] Since the infrared speckle images and the dense depth maps are different representations of the same physical scene, a pixel-level correspondence relationship needs to be established between them. This means that each pixel point in the infrared speckle image can be accurately corresponded to the pixel point at the same position in the dense depth map, and in the time series, the pixel points between different frames also have an accurate correspondence relationship.

[0062] This spatio-temporal alignment mapping relationship enables subsequent steps to perform accurate joint analysis between different types of data (images and depth maps). For example, using depth map information to assist in occlusion detection on infrared speckle images, and at the same time being able to correspond the information detected on the infrared speckle images to the depth map to further analyze the depth characteristics of the occlusion area.

[0063] Step T2: Calculate the motion vectors between adjacent frames and execute Step T3 when the average motion intensity exceeds the dynamic threshold.

[0064] In this step, the dynamic threshold is determined by the dense depth map. Motion vectors are used to describe the movement direction and displacement magnitude of objects in the image between two adjacent frames. By using specific algorithms (such as optical flow method, etc.), analyze the adjacent two frames of infrared speckle images to calculate the motion vectors of each pixel point or local area. These motion vectors reflect the movement of objects in the scene, such as in which direction the object moves and how much distance it moves.

[0065] The dynamic threshold is determined by the dense depth map. This is because the depth information is closely related to the layout and movement of objects in the scene. For example, in areas with large depth changes (such as the edges of objects or the junctions of objects at different depths), the movement of objects may be more likely to cause visual changes, so a relatively low threshold is required to detect movement in these areas; while in areas with relatively uniform depth, a higher threshold may be required to avoid false detections.

[0066] The threshold is dynamically adjusted based on the statistical characteristics of the depth map (such as the variance of the depth, the distribution of different depth intervals, etc.), so that the detection method can adapt to changes in the scene. For example, if there are a large number of fast-moving objects in the scene, the depth map will change more dramatically. At this time, the dynamic threshold will be adjusted accordingly to ensure that the movement of these objects can be accurately detected.

[0067] Calculate the average motion intensity of the motion vectors of all pixels or regions of interest and compare it with the dynamic threshold. If the average motion intensity exceeds the dynamic threshold, it means that the objects in the scene are moving more obviously and there may be changes such as occlusion. At this time, execute step T3 to further analyze the image to detect the occluded area. If the average motion intensity does not exceed the threshold, it means that the scene is relatively stable and more detailed occlusion detection may not be required. You can continue to acquire new image frames for analysis.

[0068] In some embodiments, the dynamic threshold is proportional to the average depth of the dense depth map. Setting the dynamic threshold to be proportional to the average depth of the dense depth map allows the detection method to adapt to the depth characteristics of the scene. When the average depth is large, it means that the objects in the scene are farther away from the camera as a whole. In order to avoid false detection, a relatively high dynamic threshold is required to determine whether the object has significant movement; conversely, when the average depth is small, a lower dynamic threshold can detect relatively obvious movement of the object.

[0069] When the average depth of the scene is large, the calculated dynamic threshold is correspondingly higher. This means that in the calculation of motion vectors between adjacent frames, only when the average motion intensity of the object exceeds this higher threshold will it be considered that the object in the scene is moving significantly, and the subsequent occlusion detection steps will be performed. For example, in a scene monitoring a distant mountain range, the slight natural movement of the mountain range (such as slight shaking of vegetation caused by wind) will not be misjudged as significant movement at a higher dynamic threshold because its pixel changes in the image are relatively small, thereby avoiding unnecessary processing of these small, non-critical changes.

[0070] For scenes with a small average depth, the dynamic threshold is relatively low. At this time, even if the object moves slightly, its average motion intensity is more likely to exceed the threshold, thereby triggering the subsequent occlusion detection process. For example, in a hand operation scene shot at close range, a slight movement of the hand will produce a more obvious pixel change in the image. The lower dynamic threshold can capture these movements in time, ensuring accurate detection of close-range object movement and occlusion.

[0071] By making the dynamic threshold proportional to the average depth of the dense depth map, the deep camera occlusion detection method can automatically adjust the detection sensitivity according to the depth characteristics of the scene, improving the accuracy and reliability of detection in different scenes.

[0072] Step T3: locating a speckle feature point in the infrared speckle image, generating an analysis window with the speckle feature point as the center, calculating the gradient change rate, and marking it as a primary candidate region when the gradient change rate is less than a gradient threshold.

[0073] In this step, the speckles in the infrared speckle image have unique texture and distribution characteristics. Through feature extraction algorithms (such as SIFT, SURF and other algorithms optimized for speckle images), points with obvious features, namely speckle feature points, are found in the infrared speckle image. These feature points usually have high recognition in the image, such as at the edges and corners of the speckles, and their changes between different frames can reflect the movement of objects and changes in scenes.

[0074] With each located speckle feature point as the center, a window of a specific size is defined as the analysis window. The size of this window can be adjusted according to the actual situation. Generally speaking, the window cannot be too large, otherwise it may contain too much irrelevant information, affecting the calculation efficiency and accuracy; the window cannot be too small to ensure that it can contain enough speckle texture information for subsequent analysis. The role of the analysis window is to perform a detailed analysis of the speckle image in a local range to detect possible occlusion areas.

[0075] In each analysis window, the gradient change rate of the image is calculated. The gradient reflects the change of the gray value in the image. In the speckle image, the area with large gradient change usually corresponds to the edge of the speckle or the place where the texture changes significantly, while the occluded area may cause the speckle texture to change suddenly or disappear, thus changing the gradient change rate. By calculating the gradient change rate in the analysis window, this change can be quantified.

[0076] Compare the calculated gradient change rate with a preset gradient threshold. If the gradient change rate is less than the gradient threshold, it indicates that within this analysis window, the change in the speckle texture is relatively small, and there may be an occlusion situation, resulting in unclear or occluded speckle features. At this time, mark this analysis window as a primary candidate region, that is, initially consider that this region may be an occlusion region, but further verification is still required.

[0077] Step T4: Perform spatio-temporal consistency comparison on the primary candidate region and the adjacent two frames, and retain the occlusion regions with change values less than the change threshold in three consecutive frames.

[0078] In this step, for the primary candidate region marked in step T3, perform spatio-temporal consistency comparison on it in three adjacent frames of images (i.e., the constructed three-frame time-series analysis window). This means not only considering the features of this region in space (i.e., within the same frame of image), but also analyzing its changes in the time series (between different frames).

[0079] For example, check the changes in the position, shape, speckle features, etc. of this region in the three frames of images. If the changes of a region between different frames do not conform to the normal object motion law, or there are abnormal jumps or disappearances in the speckle features in different frames, then it may be an occlusion region.

[0080] Calculate the change value of the primary candidate region in three consecutive frames through a specific algorithm. This change value can comprehensively consider multiple factors such as the position change of the region, the difference in speckle features, and the change in depth information. For example, the displacement of the center position of the region between different frames and the similarity of the speckle texture within the region can be calculated, and these factors are quantified into a change value.

[0081] Compare the calculated change value with a preset change threshold. If in three consecutive frames, the change value of this primary candidate region is less than the change threshold, it indicates that the change of this region in the time series is relatively stable and conforms to the characteristics of an occlusion region (occlusion regions usually do not have large and illogical changes in a short period of time). Therefore, retain it as the finally determined occlusion region. For regions with change values greater than the change threshold, it indicates that their changes may be caused by the normal motion of the object or other non-occlusion factors, and exclude them and do not consider them as occlusion regions. In this way, through spatio-temporal consistency comparison and the judgment of change values, the true occlusion regions can be detected more accurately, improving the accuracy and reliability of occlusion detection.

[0082] Figure 2 This is the step flow chart of another dynamic adaptive depth camera occlusion detection method in the embodiments of the present invention. As Figure 2As shown, compared with the foregoing embodiments, the steps of another dynamic adaptive depth camera occlusion detection method in the embodiments of the present invention further include:

[0083] Step T5: Obtain the sparse depth map of the infrared speckle image; the infrared speckle image is aligned with the sparse depth map.

[0084] In this step, the sparse depth map is a depth map that only contains the depth information of some pixel points. Compared with the dense depth map, it only records the depth values of some key positions in the scene, rather than having depth information for each pixel point. Obtaining the sparse depth map usually can be achieved through specific algorithms, such as the method based on feature point matching. These algorithms will extract feature points in the infrared speckle image and use the geometric relationships of these feature points and the parameters of the camera to calculate their corresponding depth values, thereby generating the sparse depth map. The advantage of the sparse depth map is that the data volume is relatively small, the processing speed is fast, and in some cases, it can provide sufficient depth information to assist subsequent processing.

[0085] Ensuring the alignment of the infrared speckle image and the sparse depth map is crucial. This means establishing an accurate correspondence between the pixel points in the infrared speckle image and the points with depth information in the sparse depth map. The alignment method can be based on the calibration parameters of the camera and image feature matching. Through camera calibration, the internal parameters (such as focal length, principal point position, etc.) and external parameters (such as rotation and translation matrices) of the camera can be obtained. These parameters can help accurately map the depth information in the sparse depth map to the coordinate system of the infrared speckle image. At the same time, using the image feature matching algorithm, common feature points are found between the infrared speckle image and the sparse depth map to further accurately align the two. This alignment enables accurate data interaction and processing between specific regions of the infrared speckle image and the corresponding regions of the sparse depth map in the subsequent process.

[0086] Step T6: Remove the depth values of the occlusion regions in the dense depth map and fill the depth values of the corresponding regions in the sparse depth map.

[0087] In this step, for the occlusion regions determined in the previous steps, in the dense depth map, the depth values of these regions may be inaccurate or unreliable because occlusion will cause incorrect measurement of depth information. Therefore, it is necessary to perform an operation to remove the depth values of these occlusion regions. The depth values of the occlusion regions can be set to specific invalid values (such as -1 or a very large value) to indicate that the depth information of these regions is unavailable. This can avoid using these incorrect depth data in the subsequent analysis and processing based on the dense depth map, thereby improving the accuracy of the entire depth information processing.

[0088] After removing the depth values of the occluded regions in the dense depth map, the sparse depth map is used to fill these regions. Since the sparse depth map is already aligned with the infrared speckle image, the positions corresponding to the occluded regions in the dense depth map can be found. The depth values of the corresponding regions in the sparse depth map are filled into the occluded regions of the dense depth map. Although the depth information of the sparse depth map is incomplete, in the occluded regions, it may contain some valid depth estimates. Through this filling operation, the depth information of the occluded regions can be restored to a certain extent, making the dense depth map more complete and accurate as a whole, and providing a more reliable data basis for subsequent depth map-based applications (such as 3D reconstruction, object recognition, etc.). This method of repairing the occluded regions of the dense depth map using the sparse depth map combines the characteristics of the two depth maps, while ensuring the integrity of the depth information, minimizing the problem of depth data loss caused by occlusion.

[0089] In some embodiments, when calculating the motion vectors between adjacent frames, an automatic window is set for the infrared speckle images of adjacent frames to confirm the matching range; the size of the automatic window is determined by the depth change rate on the dense depth map. When calculating the motion vectors of the infrared speckle images of adjacent frames, the role of the automatic window is to determine the local region for finding matching points in each frame of the image. The traditional fixed window size may not be able to adapt to the changing characteristics of different regions in the scene, while the automatic window can dynamically adjust its size according to the depth change rate, so as to more accurately capture the motion information of the object.

[0090] On the dense depth map, the depth change rate is usually calculated based on the depth values of adjacent pixels. When the depth change rate of a certain region on the dense depth map is large, it means that this region may be at the edge of an object or at the junction of objects with different depths, and the motion of the object in these regions may cause significant visual changes. In order to more accurately capture these complex motion information, the size of the automatic window should be appropriately increased. This can include more possible matching points and improve the accuracy of motion vector calculation. On the contrary, in the regions with a small depth change rate, the scene is relatively stable, and the motion of the object is relatively simple and has little change. At this time, the automatic window can maintain a small size to improve the calculation efficiency and avoid introducing too much irrelevant information. Since the object motion and depth information in the scene are dynamically changing, the size of the automatic window needs to be updated in real time. When calculating the motion vector for each frame of the image, the size of the automatic window should be re-evaluated and adjusted according to the depth change rate of the current frame of the dense depth map. This can ensure that during the processing of the entire video sequence, it can always adapt to the changes in the scene and accurately calculate the motion vector.

[0091] In actual implementation, by traversing each pixel point of the dense depth map and according to its depth change rate, the automatic window size can be set for the corresponding region of the infrared speckle image according to a preset rule (such as adjusting the window size according to the threshold as described above). For example, a mask matrix with the same size as the infrared speckle image can be used. Each element in this matrix corresponds to a pixel region of the infrared speckle image. The automatic window size corresponding to this region is calculated according to the depth change rate and recorded in the mask matrix. When calculating the motion vector, the automatic window size of the corresponding region is determined according to the value of each element in the mask matrix, so as to realize the dynamic setting of the automatic window to confirm the matching range.

[0092] By this way of dynamically setting the automatic window to confirm the matching range according to the depth change rate on the dense depth map, the motion vector between adjacent frames can be calculated more effectively in complex scenes, providing more accurate basic data for subsequent tasks such as occlusion detection.

[0093] Figure 3 This is a flowchart of the steps for marking the primary candidate region in the embodiment of the present invention. As Figure 3 shown, the steps for marking the primary candidate region in the embodiment of the present invention include:

[0094] Step T31: Binarize the infrared speckle image to obtain the light spot region, and then obtain the light spot feature points.

[0095] In this step, binarization is the process of converting a grayscale image like the infrared speckle image (usually with pixel values between 0 - 255) into an image with only two values (usually 0 and 1, or 0 and 255). Its purpose is to highlight the light spot information in the image and simplify subsequent processing. Common binarization methods include the fixed threshold method and the adaptive threshold method.

[0096] The fixed threshold method is to select a fixed grayscale value as the threshold. If the grayscale value of a certain pixel point in the image is greater than the threshold, it is set to one value (such as 255, representing white), otherwise it is set to another value (such as 0, representing black). For example, assuming the threshold is 128, for pixel points with grayscale values greater than 128, they are changed to 255, and pixel points less than or equal to 128 are changed to 0.

[0097] The adaptive threshold method dynamically determines the threshold according to the gray characteristics of the local region of the image. This method is more suitable for infrared speckle images because the lighting conditions in different regions of the image may be different. For example, the local mean method calculates the average gray value within the neighborhood of each pixel point and uses it as the threshold for binarizing this pixel point. After binarization, the light spot part in the image usually shows as a white region, and the background shows as a black region, thus obtaining the light spot region.

[0098] In the obtained binary image of the light spot area, the light spot feature points are those points that can represent the unique attributes of the light spot. Common light spot feature points include the center of the light spot, corner points on the edge, etc. Some feature extraction algorithms can be used to obtain these points.

[0099] For example, for a circular light spot, the geometric center of the light spot area can be calculated as the feature point. For an irregularly shaped light spot, an edge detection algorithm (such as the Canny edge detection) can be used to first detect the edge of the light spot, and then corner points on the edge are found as feature points. These light spot feature points serve as the center of the analysis window in subsequent steps for further analyzing the local characteristics of the light spot.

[0100] Step T32: Generate an analysis window with the light spot feature point as the center and the light spot area as the boundary.

[0101] In this step, the light spot feature points obtained in step T31 are used as the center of the analysis window. The reason for this is that the light spot feature points usually contain the key information of the light spot. Generating a window with it as the center can ensure that the window contains the light spot texture and variation information closely related to this feature point. For example, if the light spot feature point is the center of the light spot, the window generated with it as the center can evenly cover all parts of the light spot; if it is a corner point on the edge, the window can focus on analyzing the variation of the light spot texture near the corner point.

[0102] The size and shape of the analysis window are determined by the light spot area. This means that the analysis window will not exceed the range of the light spot area, ensuring that the information within the window is all related to the light spot and avoiding introducing too much background noise. For example, if the light spot area is an irregular shape, the shape and size of the analysis window will be adapted to the light spot area as much as possible to make full use of the information within the light spot area for subsequent gradient calculation and analysis. The analysis window generated in this way can more accurately perform local analysis on the light spot and improve the accuracy of detecting the occluded area.

[0103] Step T33: Inside the analysis window, use a gradient operator to calculate the gradient change value of each pixel point, and calculate the standard deviation of the gradient change values of all pixel points within the analysis window as the gradient change rate.

[0104] In this step, inside the analysis window, a gradient operator (such as the Sobel operator, Prewitt operator, etc.) is used to calculate the gradient change value of each pixel point. The gradient operator performs a difference operation on the gray value of the pixel point in the image to obtain the gray change rate of the pixel point in the horizontal and vertical directions, and then calculates the gradient amplitude and direction of the pixel point.

[0105] Taking the Sobel operator as an example, it contains two templates, one for detecting the gradient in the horizontal direction (Gx) and the other for detecting the gradient in the vertical direction (Gy). For each pixel point within the analysis window, its neighborhood is convolved with these two templates to obtain the approximate gradient values of the pixel point in the horizontal and vertical directions. Then, the gradient magnitude of the pixel point is calculated through a formula, and this gradient magnitude is the gradient change value of the pixel point. It reflects the degree of intensity change of the light spot image at this pixel point. The more intense the gray level change, the greater the gradient change value.

[0106] After obtaining the gradient change values of all pixel points within the analysis window, the standard deviation of these values is calculated. The standard deviation is a statistic that measures the degree of data dispersion. Here, it can reflect the overall fluctuation of the gradient change values of pixel points within the analysis window. If the light spot texture within the analysis window is relatively uniform and the gradient change values of pixel points are relatively close, the standard deviation will be smaller; conversely, if the light spot texture is complex and there are many regions with intense gray level changes, the gradient change values will vary greatly, and the standard deviation will be larger. By using this standard deviation as the gradient change rate, the change degree of the light spot texture within the analysis window can be quantified, providing a numerical basis for subsequent judgment of whether it is an occluded area.

[0107] Step T34: When the gradient change rate is less than the gradient threshold, it is marked as a primary candidate area.

[0108] In this step, the gradient threshold is a preset empirical value, which is a key indicator for judging whether the analysis window may be an occluded area. The setting of this threshold requires multiple experiments and adjustments according to the specific application scenario and the characteristics of the infrared speckle image. If the threshold is set too high, some real occluded areas may be missed; if the threshold is set too low, some normal light spot areas may be misjudged as occluded areas.

[0109] Compare the gradient change rate calculated in step T33 with the gradient threshold. When the gradient change rate is less than the gradient threshold, it indicates that within this analysis window, the change of the light spot texture is relatively small, and there may be an occlusion situation, resulting in unclear or occluded light spot features. At this time, mark this analysis window as a primary candidate area, that is, initially consider that this area may be an occluded area, but it still needs to be further verified in subsequent steps. For example, if within a certain analysis window, due to the occlusion of an object, part of the light spot area is covered, making the light spot texture within this area relatively smooth, the gradient change rate will decrease. When it is lower than the gradient threshold, it will be marked as a primary candidate area.

[0110] Figure 4 This is a flowchart of the steps for marking an occluded area in an embodiment of the present invention. As Figure 4 shown, the steps for marking an occluded area in an embodiment of the present invention include:

[0111] Step T41: Extract the region corresponding to the primary candidate region in two adjacent frames according to the spatio-temporal alignment mapping relationship.

[0112] In this step, the pixel-level spatio-temporal alignment mapping relationship established in step T1 is a bridge connecting images of different frames and different types of data (such as infrared speckle images and dense depth maps). It clarifies the corresponding relationship of each pixel point in each frame image in the time and space dimensions.

[0113] Based on this spatio-temporal alignment mapping relationship, for the primary candidate region marked in step T3, the corresponding regions can be accurately found in the adjacent previous and next frame infrared speckle images. This is like finding regions with the same geographical location on the "map" of images at different times through an accurate map. For example, if the coordinate range of the primary candidate region in the current frame image is from to, according to the coordinate conversion information provided by the spatio-temporal alignment mapping relationship, the coordinate range of the corresponding region with the same physical position can be determined in two adjacent frame images. The accuracy of this corresponding relationship is crucial for subsequent region-based comparison analysis, because only by ensuring the strict spatio-temporal correspondence of regions can their changes at different times be effectively compared, so as to accurately judge the possibility of occlusion.

[0114] Step T42: Calculate the brightness difference between the primary candidate region and the corresponding regions in two adjacent frames to obtain the change value.

[0115] In this step, brightness is an important feature of infrared speckle images. Between different frames, due to factors such as object movement and occlusion, the brightness of the same region may change. Calculating the brightness difference between the primary candidate region and its corresponding regions in two adjacent frames can help determine whether there is occlusion in this region.

[0116] Specifically in the calculation, multiple methods can be used. A common way is to calculate the average brightness of all pixel points in each region, and then find the difference between the average brightness of the primary candidate region and the average brightness of the corresponding regions in two adjacent frames. For example, let the average brightness of the primary candidate region be, the average brightness of the corresponding region in the previous frame be, and the average brightness of the corresponding region in the next frame be, then two brightness differences and can be calculated.

[0117] In addition to calculating the difference in averages, more complex methods can also be considered, such as calculating the difference in the brightness values of each pixel point and performing a certain statistical operation (such as summing, averaging, etc.) on the difference values of all pixel points. For example, for each pixel point, calculate its brightness values, and in the primary candidate region and the corresponding regions in two adjacent frames, and then calculate as a measure of the brightness difference.

[0118] Through the above brightness difference calculation method, we finally get a value that can comprehensively reflect the brightness change of the primary candidate area in three adjacent frames. This value is the change value. It quantifies the brightness change of the area in the time series. The larger the change value, the more obvious the brightness change of the area between different frames; the smaller the change value, the relatively small brightness change. The change value will serve as one of the important bases for judging whether the area is an occlusion area.

[0119] Step T43: retaining the primary candidate regions whose change values ​​are less than the change threshold as occlusion regions.

[0120] In this step, the change threshold is a pre-set reference value, which is determined based on a large amount of experimental data and the needs of actual application scenarios. This threshold is used to distinguish whether the change value of the primary candidate area is a normal scene change or a change that may be caused by occlusion. If the change threshold is set too high, some real occlusion areas may be misjudged as normal areas and ignored; if it is set too low, some normal scene change areas may be misjudged as occlusion areas.

[0121] Compare the change value calculated in step T42 with the change threshold. When the change value is less than the change threshold, it indicates that the brightness change of the primary candidate area in the three adjacent frames is relatively small, and this smaller change pattern meets the characteristics of the occluded area. Because in the case of occlusion, the physical characteristics of the occluded area are relatively stable, and there will be no large brightness fluctuations (unless the occluder itself has obvious brightness changes). Therefore, at this time, the primary candidate area is retained and confirmed as an occluded area. For the primary candidate area whose change value is greater than the change threshold, it means that its brightness change is more drastic, which may be caused by non-occluded factors such as normal movement of objects and changes in illumination, so it is not retained as an occluded area. Through this judgment method based on the change value and the change threshold, the real occluded area can be further screened out, and the accuracy of the occlusion detection result can be improved.

[0122] In some embodiments, in step T43, morphological analysis is also performed on the primary candidate area to remove noise interference. Morphological analysis is a shape-based technology in image processing, which mainly changes the shape characteristics of the image by interacting with the image through structural elements. The core idea is to use structural elements of specific shapes and sizes (such as rectangles, circles, crosses, etc.) to operate on the image to highlight or suppress certain features in the image. In this step, morphological analysis is mainly used to process the primary candidate area after brightness difference calculation and screening, remove possible noise interference therein, and make the final occlusion area more accurate.

[0123] According to the characteristics of the primary candidate regions and the possible forms of noise, select appropriate structuring elements. For example, if the noise appears as isolated small dots, circular or square structuring elements can be selected, and their sizes should be determined according to the approximate size of the noise dots. If the noise dots are small, the size of the structuring element should also be correspondingly small to avoid over-erosion of the normal candidate regions; if the noise distribution is relatively complex, different shapes and sizes of structuring elements may need to be tried, and the best choice can be determined through experimental comparison.

[0124] Use the selected structuring element to perform an erosion operation on the primary candidate regions. The erosion process is to move the structuring element on the image of the primary candidate regions. For each position covered by the structuring element, if all pixels at that position belong to the candidate region (in a binary image, usually represented as white pixels), the pixels at that position are retained, otherwise they are set to the background (black pixels). Through the erosion operation, isolated noise dots or small noise blocks on the boundaries of the candidate regions can be removed because these noise dots cannot meet the condition that all pixels belong to the candidate region during the movement of the structuring element and are thus removed. For example, for a primary candidate region containing a small number of isolated noise dots, after the erosion operation, these noise dots will be eliminated while the main part of the candidate region remains basically unchanged.

[0125] After the erosion operation, a dilation operation is usually performed. Dilation is the opposite of erosion. It moves the structuring element on the image. For each position covered by the structuring element, as long as one pixel belongs to the candidate region, that position is set as a candidate region pixel. The purpose of the dilation operation is to restore the normal candidate regions that have been eroded and connect some regions that may have been disconnected during the erosion process. For example, after the erosion operation, the boundary of the candidate region may become jagged. The dilation operation can make the boundary smoother and connect some small holes generated due to noise removal, making the shape of the final candidate region more complete.

[0126] In step T43, first calculate the luminance difference between the primary candidate region and the corresponding regions in the adjacent two frames to obtain the change value, and compare it with the change threshold to screen out the primary candidate regions with change values less than the change threshold. Then, perform morphological analysis on these preliminarily screened regions. Such an order arrangement is reasonable because the possible occlusion regions have been preliminarily determined through the change value judgment. Based on this, morphological analysis can be carried out to specifically remove the noise in these regions without performing unnecessary morphological processing on a large number of regions that do not conform to the occlusion characteristics, improving the processing efficiency.

[0127] After morphological analysis, these areas are inspected again to ensure that they still conform to the characteristics of occluded areas. For example, check whether the areas after morphological analysis are still clearly distinguishable from the surrounding environment and whether they are meaningful in subsequent processing. If an area becomes too small or the difference from the surrounding areas is not obvious after morphological analysis, it may be necessary to re-evaluate whether the area should still be retained as an occluded area. By combining morphological analysis with the determination of change values, noise interference can be more effectively removed, the final occluded areas can be accurately determined, and the reliability and accuracy of the entire occlusion detection method can be improved.

[0128] By introducing morphological analysis to process the primary candidate areas in step T43, noise interference can be effectively removed, the finally determined occluded areas can be made more accurate and reliable, and a better data basis can be provided for subsequent analysis and applications based on the occluded areas.

[0129] Figure 5 This is a schematic structural diagram of a dynamic adaptive depth camera occlusion detection system in an embodiment of the present invention.

[0130] As Figure 5 shown, a dynamic adaptive depth camera occlusion detection system in an embodiment of the present invention includes:

[0131] An acquisition module, configured to construct a three-frame time-series analysis window, continuously acquire infrared speckle image sequences and corresponding dense depth map sequences at different times, and establish a pixel-level spatio-temporal alignment mapping relationship;

[0132] A motion module, configured to calculate the motion vectors between adjacent frames, and execute the candidate module when the average motion intensity exceeds a dynamic threshold; wherein, the dynamic threshold is determined by the dense depth map;

[0133] A candidate module, configured to locate speckle feature points in the infrared speckle image, generate an analysis window centered on the speckle feature points, calculate the gradient change rate, and mark it as a primary candidate area when the gradient change rate is less than the gradient threshold;

[0134] A determination module, configured to perform spatio-temporal consistency comparison on the primary candidate area and the adjacent two frames, and retain the occluded areas with change values less than the change threshold in three consecutive frames.

[0135] Specifically, the acquisition module constructs a three-frame timing analysis window, which provides a framework in the time dimension for the subsequent analysis of consecutive image frames. Continuously acquire sequences of infrared speckle images and corresponding dense depth maps at different times, and these two types of data are the basic data for the analysis of the entire system. Establish a pixel-level spatio-temporal alignment mapping relationship to ensure that subsequent analyses based on different image frames and different types of data (infrared speckle images and dense depth maps) can be carried out under accurate pixel correspondence relationships, providing a prerequisite for subsequent accurate detection of occlusions.

[0136] The motion module calculates the motion vectors between adjacent frames and measures the motion of objects by analyzing the changes between adjacent frames. Compare the calculated average motion intensity with the dynamic threshold, and when the average motion intensity exceeds the dynamic threshold, trigger the candidate module. Among them, the dynamic threshold is determined by the dense depth map, which enables the system to adaptively adjust the conditions for triggering subsequent analyses according to the scene depth information, enhancing the adaptability of the system to different scenarios.

[0137] The candidate module locates the speckle feature points in the infrared speckle image, and these feature points are the key identifiers for analyzing scene changes. Generate an analysis window centered on the speckle feature points, which limits the local range of the analysis and improves the analysis efficiency. Calculate the gradient change rate within the analysis window, and the gradient change rate reflects the degree of change of the image information within the window. When the gradient change rate is less than the gradient threshold, mark this area as a primary candidate area, initially screening out the areas where occlusions may exist.

[0138] The determination module performs spatio-temporal consistency comparison on the primary candidate area and the adjacent two frames, and further analyzes the changes in the primary candidate area from both the time and space dimensions. Retain the occlusion areas where the change values are less than the change threshold in three consecutive frames. In this way, the true occlusion areas can be more accurately determined, excluding some misjudged areas caused by noise or other factors.

[0139] The acquisition module is the foundation of the entire system, providing the required sequences of infrared speckle images and dense depth maps as well as the pixel-level spatio-temporal alignment mapping relationship for the subsequent modules. The motion module calculates the motion vectors between adjacent frames based on the image sequences provided by the acquisition module and decides whether to trigger the candidate module according to the dynamic threshold, playing a connecting role and controlling the start of subsequent analyses according to the motion situation. The candidate module conducts preliminary screening in the infrared speckle image and marks the primary candidate areas, providing the basic areas for further analysis for the determination module. The determination module performs spatio-temporal consistency comparison on the primary candidate areas marked by the candidate module and finally determines the occlusion areas, which is the last step of the entire occlusion detection process, relying on the processing results of the previous modules to complete the entire occlusion detection task.

[0140] In this embodiment, by constructing a three-frame time-series analysis window and establishing a pixel-level spatio-temporal alignment mapping relationship, the spatio-temporal information of the image sequence and the depth map sequence can be fully utilized, laying a foundation for accurate detection. Using the dense depth map to determine the dynamic threshold to calculate the motion vector between adjacent frames, it can be dynamically adjusted according to the actual scene, effectively adapting to different environmental changes. Locating the speckle feature points in the infrared speckle image and calculating the gradient change rate to mark the primary candidate regions can accurately screen out the possible occlusion regions. Conducting spatio-temporal consistency comparison on the primary candidate regions and retaining the occlusion regions with change values less than the change threshold further improves the accuracy and stability of the detection results, ensuring that the occlusion regions can be reliably detected even in complex dynamic scenes.

[0141] In an embodiment of the present invention, there is also provided a dynamically adaptive depth camera occlusion detection device, including a processor and a memory in which executable instructions of the processor are stored. Among them, the processor is configured to execute the steps of a dynamically adaptive depth camera occlusion detection method via executing the executable instructions.

[0142] As above, in this embodiment, by constructing a three-frame time-series analysis window and establishing a pixel-level spatio-temporal alignment mapping relationship, the spatio-temporal information of the image sequence and the depth map sequence can be fully utilized, laying a foundation for accurate detection. Using the dense depth map to determine the dynamic threshold to calculate the motion vector between adjacent frames, it can be dynamically adjusted according to the actual scene, effectively adapting to different environmental changes. Locating the speckle feature points in the infrared speckle image and calculating the gradient change rate to mark the primary candidate regions can accurately screen out the possible occlusion regions. Conducting spatio-temporal consistency comparison on the primary candidate regions and retaining the occlusion regions with change values less than the change threshold further improves the accuracy and stability of the detection results, ensuring that the occlusion regions can be reliably detected even in complex dynamic scenes.

[0143] Those skilled in the art of the relevant technical field can understand that various aspects of the present invention can be implemented as a system, a method, or a program product. Therefore, various aspects of the present invention can be specifically implemented in the following forms, namely: a complete hardware implementation manner, a complete software implementation manner (including firmware, microcode, etc.), or an implementation manner combining hardware and software aspects, which can be collectively referred to as "circuit", "module", or "platform" here.

[0144] Figure 6 is a structural schematic diagram of a dynamically adaptive depth camera occlusion detection device in an embodiment of the present invention. The following refers to Figure 6 to describe the electronic device 600 according to this embodiment of the present invention. Figure 6 The shown electronic device 600 is only an example and should not bring any limitation to the functions and usage scope of the embodiments of the present invention.

[0145] As Figure 6As shown, the electronic device 600 is presented in the form of a general-purpose computing device. The components of the electronic device 600 may include, but are not limited to: at least one processing unit 610, at least one storage unit 620, a bus 630 connecting different platform components (including the storage unit 620 and the processing unit 610), a display unit 640, etc.

[0146] Among them, the storage unit stores program code, and the program code can be executed by the processing unit 610, so that the processing unit 610 executes the steps according to various exemplary embodiments of the present invention described in the above-mentioned part of a dynamic adaptive depth camera occlusion detection method in this specification. For example, the processing unit 610 can execute steps as shown in Figure 1 shown.

[0147] The storage unit 620 may include a readable medium in the form of a volatile storage unit, such as a random access storage unit (RAM) 6201 and / or a cache storage unit 6202, and may further include a read-only storage unit (ROM) 6203.

[0148] The storage unit 620 may also include a program / utilities 6204 having a set (at least one) of program modules 6205. Such program modules 6205 include, but are not limited to: an operating system, one or more application programs, other program modules, and program data. Each or some combination of these examples may include the implementation of a grid environment.

[0149] The bus 630 may represent one or more of several types of bus structures, including a storage unit bus or a storage unit controller, a peripheral bus, a graphics acceleration port, a processing unit, or a local bus using any bus structure in a variety of bus structures.

[0150] The electronic device 600 can also communicate with one or more external devices 700 (such as a keyboard, a pointing device, a Bluetooth device, etc.), can also communicate with one or more devices that enable a user to interact with the electronic device 600, and / or communicate with any device that enables the electronic device 600 to communicate with one or more other computing devices (such as a router, a modem, etc.). Such communication can be carried out through an input / output (I / O) interface 650. And, the electronic device 600 can also communicate with one or more grids (such as a local area network (LAN), a wide area network (WAN), and / or a public grid, such as the Internet) through a grid adapter 660. The grid adapter 660 can communicate with other modules of the electronic device 600 through the bus 630. It should be understood that although Figure 6Not shown in the figure, other hardware and / or software modules can be used in combination with the electronic device 600, including but not limited to: microcode, device drivers, redundant processing units, external disk drive arrays, RAID systems, tape drives, and data backup storage platforms, etc.

[0151] An embodiment of the present invention also provides a computer-readable storage medium for storing a program, and the steps of a dynamic adaptive depth camera occlusion detection method are implemented when the program is executed. In some possible implementation manners, various aspects of the present invention can also be implemented in the form of a program product, which includes program code. When the program product runs on a terminal device, the program code is used to cause the terminal device to execute the steps according to various exemplary embodiments of the present invention described in the above-mentioned part of the dynamic adaptive depth camera occlusion detection method of this specification.

[0152] As shown above, by constructing a three-frame timing analysis window and establishing a pixel-level spatio-temporal alignment mapping relationship, this embodiment can make full use of the spatio-temporal information of the image sequence and the depth map sequence, laying a foundation for accurate detection; using the dense depth map to determine the dynamic threshold to calculate the motion vector between adjacent frames, which can be dynamically adjusted according to the actual scene and effectively adapt to different environmental changes; locating the speckle feature points in the infrared speckle image and calculating the gradient change rate to mark the primary candidate regions, which can accurately screen out the possible occlusion regions; performing spatio-temporal consistency comparison on the primary candidate regions and retaining the occlusion regions with the change value less than the change threshold, further improving the accuracy and stability of the detection result and ensuring that the occlusion regions can be reliably detected in complex dynamic scenes.

[0153] Figure 7 is a schematic structural diagram of the computer-readable storage medium in the embodiment of the present invention. Refer to Figure 7 As shown, a program product 800 for implementing the above method according to an embodiment of the present invention is described. It can be a portable compact disc read-only memory (CD-ROM) and includes program code, and can run on a terminal device, such as a personal computer. However, the program product of the present invention is not limited to this. In this document, the readable storage medium can be any tangible medium that contains or stores a program, and the program can be used by or in combination with an instruction execution system, device, or device.

[0154] The program product may employ any combination of one or more readable media. The readable media may be a readable signal medium or a readable storage medium. The readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the foregoing. More specific examples (a non-exhaustive list) of the readable storage medium include: an electrical connection having one or more wires, a portable disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0155] A computer-readable storage medium may include a data signal propagated in a baseband or as part of a carrier wave, in which the readable program code is carried. Such a propagated data signal may take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the foregoing. The readable storage medium may also be any readable medium other than a readable storage medium, which can send, propagate, or transmit a program for use by or in connection with an instruction execution system, apparatus, or device. The program code contained on the readable storage medium may be transmitted using any appropriate medium, including but not limited to wireless, wireline, optical fiber cable, RF, etc., or any suitable combination of the foregoing.

[0156] The program code for performing the operations of the present invention may be written in any combination of one or more programming languages, including object-oriented programming languages such as Java, C++, etc., and also including conventional procedural programming languages such as the "C" language or similar programming languages. The program code may be executed entirely on the user's computing device, partially on the user's device, executed as a stand-alone software package, partially on the user's computing device and partially on a remote computing device, or entirely on a remote computing device or server. In the case of a remote computing device, the remote computing device may be connected to the user's computing device through any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computing device (e.g., using an Internet service provider to connect through the Internet).

[0157] In this embodiment, by constructing a three-frame timing analysis window and establishing a pixel-level spatio-temporal alignment mapping relationship, the spatio-temporal information of the image sequence and the depth map sequence can be fully utilized, laying a foundation for accurate detection. Using the dense depth map to determine the dynamic threshold to calculate the motion vector between adjacent frames, it can be dynamically adjusted according to the actual scene, effectively adapting to different environmental changes. Locating the speckle feature points in the infrared speckle image and calculating the gradient change rate, marking the primary candidate area, can accurately screen out the possible occlusion areas. Conducting spatio-temporal consistency comparison on the primary candidate areas and retaining the occlusion areas with variation values less than the variation threshold further improves the accuracy and stability of the detection results, ensuring that the occlusion areas can be reliably detected even in complex dynamic scenes.

[0158] The various embodiments in this specification are described in a progressive manner. Each embodiment focuses on the differences from other embodiments. The same or similar parts among the various embodiments can be referred to each other. The above description of the disclosed embodiments enables those skilled in the art to implement or use the present invention. Various modifications to these embodiments will be obvious to those skilled in the art. The general principles defined herein can be implemented in other embodiments without departing from the spirit or scope of the present invention. Therefore, the present invention will not be limited to these embodiments shown herein, but rather to the widest scope consistent with the principles and novel features disclosed herein.

[0159] The specific embodiments of the present invention have been described above. It should be understood that the present invention is not limited to the above specific embodiments, and those skilled in the art can make various deformations or modifications within the scope of the claims, which do not affect the essence of the present invention.

Claims

1. A dynamic adaptive depth camera occlusion detection method, characterized in that: include: Step T1: construct a three-frame timing analysis window, continuously acquire infrared speckle image sequences and corresponding dense depth map sequences at different times, and establish a pixel-level spatiotemporal alignment mapping relationship; Step T2: Calculate the motion vector between adjacent frames, and execute step T3 when the average motion intensity exceeds a dynamic threshold; wherein the dynamic threshold is determined by the dense depth map; Step T3: locating a speckle feature point in the infrared speckle image, generating an analysis window with the speckle feature point as the center, calculating the gradient change rate, and marking it as a primary candidate region when the gradient change rate is less than a gradient threshold; Step T4: performing a temporal and spatial consistency comparison between the primary candidate region and two adjacent frames, and retaining the occluded regions whose change values ​​in three consecutive frames are less than a change threshold.

2. The method for dynamically adaptive depth camera occlusion detection according to claim 1, characterized in that: Also includes: Step T5: obtaining a sparse depth map of the infrared speckle image; The infrared speckle image is aligned with the sparse depth map; Step T6: removing the depth value of the blocked area in the dense depth map, and filling the depth value of the corresponding area of ​​the sparse depth map.

3. The method for dynamically adaptive depth camera occlusion detection according to claim 1, characterized in that: When calculating the motion vector between adjacent frames, an automatic window is set for the infrared speckle images of adjacent frames to confirm the matching range; the size of the automatic window is determined by the depth change rate on the dense depth map.

4. The method for dynamically adaptive depth camera occlusion detection according to claim 1, characterized in that: The dynamic threshold is proportional to the average depth of the dense depth map.

5. The method for dynamically adaptive depth camera occlusion detection according to claim 1, characterized in that: Step T3 includes: Step T31: binarizing the infrared speckle image to obtain a spot area, and then obtaining a spot feature point; Step T32: Generate an analysis window with the light spot feature point as the center and the light spot area as the boundary; Step T33: in the analysis window, using the gradient operator to calculate the gradient change value of each pixel point, and obtaining the standard deviation of the gradient change values ​​of all the pixels points in the analysis window as the gradient change rate; Step T34: When the gradient change rate is less than the gradient threshold, it is marked as a primary candidate region.

6. The method for dynamically adaptive depth camera occlusion detection according to claim 1, characterized in that: Step T4 includes: Step T41: extracting a region corresponding to the primary candidate region in two adjacent frames according to the spatiotemporal alignment mapping relationship; Step T42: Calculate the brightness difference between the primary candidate region and the corresponding regions in two adjacent frames to obtain a change value; Step T43: retaining the primary candidate regions whose change values ​​are less than the change threshold as occlusion regions.

7. The method for dynamically adaptive depth camera occlusion detection according to claim 6, characterized in that: In step T43, morphological analysis is also performed on the primary candidate regions to remove noise interference.

8. A dynamically adaptive depth camera occlusion detection system, used to implement the dynamically adaptive depth camera occlusion detection method according to any one of claims 1 to 7, characterized in that: include: The acquisition module is used to construct a three-frame timing analysis window, continuously acquire infrared speckle image sequences and corresponding dense depth map sequences at different times, and establish a pixel-level spatiotemporal alignment mapping relationship; A motion module, configured to calculate motion vectors between adjacent frames and execute a candidate module when an average motion intensity exceeds a dynamic threshold; wherein the dynamic threshold is determined by the dense depth map; A candidate module is used to locate a speckle feature point in the infrared speckle image, generate an analysis window with the speckle feature point as the center, calculate a gradient change rate, and mark a region as a primary candidate region when the gradient change rate is less than a gradient threshold; The determination module is used to compare the temporal and spatial consistency of the primary candidate area with two adjacent frames, and retain the occlusion area whose change value in three consecutive frames is less than the change threshold.

9. A dynamic adaptive depth camera occlusion detection device, characterized in that: include: processor; a memory storing executable instructions of the processor; The processor is configured to execute the steps of the dynamically adaptive depth camera occlusion detection method of any one of claims 1 to 7 by executing the executable instructions.

10. A computer-readable storage medium for storing a program, characterized in that: When the program is executed, the steps of the dynamically adaptive depth camera occlusion detection method described in any one of claims 1 to 7 are implemented.

Citation Information

Cited By

  • Lightweight parking space detection method and system integrating space-time consistency and self-supervised learning

    CN120431553A

  • Lightweight parking space detection method and system integrating spatiotemporal consistency and self-supervised learning

    CN120431553B

  • Dead pixel detection method and system combining spatial domain analysis and time domain analysis

    CN120495308A

  • A bad pixel detection method and system combining spatial analysis and temporal analysis

    CN120495308B

  • Security box video monitoring system and method with intelligent discrimination function

    CN120568026A