Self-adaptive depth camera and intelligent device
By adopting adaptive infrared speckle technology and depth map processing methods in depth cameras, the problem of insufficient imaging quality and dynamic capture capabilities in complex scenes is solved, and deep information acquisition with higher accuracy and completeness is achieved.
Patent Information
- Application Number
- CN202510192460.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-21
- Publication Date
- 2025-05-30
AI Technical Summary
In complex and changing scenes, traditional depth cameras have problems such as imaging quality being affected by light, difficulty in capturing the motion trajectory and morphological changes of dynamically changing objects, and insufficient occlusion area recognition and processing capabilities.
Adaptive depth camera is used to generate infrared speckle map sequences and dense depth map sequences using infrared speckle projectors and receivers. By calculating the motion vectors and gradient change rates between adjacent frames, the initial candidate areas are determined, and the occlusion area is identified based on the spatial and temporal consistency. Finally, the sparse depth map is used to fill the depth value of the occlusion area.
It improves the imaging quality under different lighting conditions, accurately captures the movement trajectory and morphological changes of dynamically changing objects, effectively identify and process occluded areas, reduces information loss, and improves the accuracy and completeness of depth information.
Smart Images

Figure CN120075556A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of depth cameras, and in particular, to an adaptive depth camera and an intelligent device. Background Art
[0002] In today's digital age, depth cameras, as key devices capable of acquiring three-dimensional information of scenes, are widely used in many fields, such as autonomous driving, robot navigation, virtual reality (VR) / augmented reality (AR), and security monitoring.
[0003] Traditional depth cameras often have many limitations when facing complex and changing scenes. On the one hand, under different lighting conditions, their imaging quality will be severely affected. For example, in an environment with direct strong light or dim light, the acquired images are prone to overexposure or excessive noise, resulting in inaccurate depth information, which in turn affects subsequent data analysis and processing. On the other hand, when the shooting object is in a dynamic change process, such as a fast-moving object, traditional depth cameras are difficult to accurately capture its motion trajectory and morphological changes, and cannot generate a continuous and accurate sequence of depth maps in a timely manner, making applications based on depth information difficult to run stably.
[0004] In addition, existing depth cameras have weak capabilities in identifying and processing occluded areas, which will cause a large amount of information loss in complex scenes, greatly limiting their application effects in actual scenes.
[0005] The disclosure of the above background art content is only for assisting in understanding the inventive concept and technical solution of the present invention, and it does not necessarily belong to the prior art of this patent application. Without clear evidence indicating that the above content was publicly available on the filing date of this patent application, the above background art should not be used to evaluate the novelty and inventiveness of this application. Summary of the Invention
[0006] Therefore, the present invention proposes an adaptive depth camera to solve the deficiencies of traditional methods in aspects such as threshold setting, spatio-temporal information utilization, and feature extraction, overcome the limitations of existing depth camera occlusion detection methods, and meet the requirements of high precision, real-time performance, and stability in complex scenes.
[0007] In a first aspect, the present invention provides an adaptive depth camera, characterized by comprising:
[0008] An infrared speckle projector for projecting infrared speckles;
[0009] An infrared receiver for receiving the reflected signal of the infrared speckles;
[0010] A processor for generating a sequence of infrared speckle patterns and a sequence of dense depth maps based on the reflected signal, calculating the motion vectors between adjacent frames, determining an initial candidate region according to the gradient change rate of the speckles, and then determining the occluded region according to the spatio-temporal consistency between adjacent frames.
[0011] Optionally, in the described adaptive depth camera, the processor further generates a sequence of sparse depth maps based on the reflected signal, and fills the occluded region with the depth values of the sequence of sparse depth maps.
[0012] Optionally, in the described adaptive depth camera, the processing by the processor includes:
[0013] Step T1: Construct a three-frame time-series analysis window, continuously acquire a sequence of infrared speckle images and the corresponding sequence of dense depth maps at different times, and establish a pixel-level spatio-temporal alignment mapping relationship;
[0014] Step T2: Calculate the motion vectors between adjacent frames, and execute Step T3 when the average motion intensity exceeds the dynamic threshold; wherein, the dynamic threshold is determined by the dense depth map;
[0015] Step T3: Locate the speckle feature points in the infrared speckle image, generate an analysis window centered on the speckle feature points, calculate the gradient change rate, and mark it as a primary candidate region when the gradient change rate is less than the gradient threshold;
[0016] Step T4: Perform a spatio-temporal consistency comparison on the primary candidate region and the adjacent two frames, and retain the occluded region with a change value less than the change threshold in three consecutive frames.
[0017] Optionally, in the described adaptive depth camera, when calculating the motion vectors between adjacent frames, an automatic window confirmation matching range is set for the infrared speckle images of adjacent frames; the size of the automatic window is determined by the depth change rate on the dense depth map.
[0018] Optionally, in the described adaptive depth camera, Step T3 includes:
[0019] Step T31: Binarize the infrared speckle image to obtain the light spot region, and further obtain the light spot feature points;
[0020] Step T32: Generate an analysis window centered on the light spot feature points and bounded by the light spot region;
[0021] Step T33: Inside the analysis window, use a gradient operator to calculate the gradient change value of each pixel point, and calculate the standard deviation of the gradient change values of all pixel points in the analysis window as the gradient change rate;
[0022] Step T34: When the gradient change rate is less than the gradient threshold, it is marked as a primary candidate region.
[0023] Optionally, for the described adaptive depth camera, step T4 includes:
[0024] Step T41: According to the spatio-temporal alignment mapping relationship, in two adjacent frames, extract the regions corresponding to the primary candidate regions;
[0025] Step T42: Calculate the brightness difference between the primary candidate region and the corresponding regions in two adjacent frames to obtain a change value;
[0026] Step T43: Retain the primary candidate regions with the change value less than the change threshold as occluded regions.
[0027] Optionally, for the described adaptive depth camera, in step T43, morphological analysis is also performed on the primary candidate regions to remove noise interference.
[0028] In a second aspect, the present invention provides an adaptive depth camera, which includes:
[0029] A sparse infrared speckle projector for projecting sparse infrared speckles;
[0030] A dense infrared speckle projector for projecting dense infrared speckles;
[0031] An infrared receiver for receiving the reflected signals of the sparse infrared speckles and the dense infrared speckles;
[0032] A processor for generating an infrared speckle map sequence, a dense depth map sequence, and a sparse depth map sequence according to the reflected signals, calculating the motion vectors between adjacent frames based on the dense depth map, determining the initial candidate regions according to the gradient change rate of the speckles, then determining the occluded regions according to the spatio-temporal consistency of the initial candidate regions between adjacent frames, and finally filling the depth values of the occluded regions with the sparse depth map.
[0033] Optionally, for the described adaptive depth camera, when the processor processes, it includes:
[0034] Step T1: Construct a three-frame time series analysis window, continuously obtain an infrared speckle image sequence and the corresponding dense depth map sequence at different times, and establish a pixel-level spatio-temporal alignment mapping relationship;
[0035] Step T2: Calculate the motion vectors between adjacent frames, and execute step T3 when the average motion intensity exceeds the dynamic threshold; wherein, the dynamic threshold is determined by the dense depth map;
[0036] Step T3: Locate the speckle feature points in the infrared speckle image, generate an analysis window centered on the speckle feature points, calculate the gradient change rate, and mark it as a primary candidate region when the gradient change rate is less than the gradient threshold;
[0037] Step T4: Perform spatio-temporal consistency comparison on the primary candidate region and the adjacent two frames, and retain the occlusion region with a change value less than the change threshold in three consecutive frames;
[0038] Step T5: Obtain the sparse depth map of the infrared speckle image; align the infrared speckle image with the sparse depth map;
[0039] Step T6: Remove the depth values of the occlusion region in the dense depth map and fill the depth values of the corresponding region in the sparse depth map.
[0040] In a third aspect, the present invention provides an intelligent device for implementing the adaptive depth camera described in any one of the foregoing.
[0041] Compared with the prior art, the present invention has the following beneficial effects:
[0042] The present invention projects infrared speckles using an infrared speckle projector, and the infrared band is less affected by changes in light. Even in an environment with direct strong light or dim light, the infrared receiver can still stably receive the reflected signals of the infrared speckles, so as to ensure that the processor can generate accurate infrared speckle map sequences and dense depth map sequences based on these signals, effectively avoiding the decline in imaging quality caused by light problems and greatly improving the accuracy of depth information acquisition.
[0043] In the present invention, the processor can keenly capture the dynamic changes of the photographed object by calculating the motion vectors between adjacent frames. When facing a fast-moving object, it can analyze the motion trajectory and speed of the object based on the motion vectors, combine with the generated depth map sequences, accurately present the morphological changes of the dynamic target, and timely generate continuous and accurate depth map sequences, effectively ensuring the stable operation of applications based on depth information in dynamic scenarios.
[0044] The present invention determines the initial candidate region through the gradient change rate of the speckles and then determines the occlusion region according to the spatio-temporal consistency between adjacent frames. This unique processing method can more effectively identify and process the occlusion region compared with traditional depth cameras. In complex scenarios, it can greatly reduce information loss, provide more complete and accurate depth information for subsequent data analysis and processing, and significantly improve the application effect of the camera in actual scenarios. Description of the Drawings
[0045] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the accompanying drawings required in the description of the embodiments or the prior art. Obviously, the accompanying drawings in the following description are only the embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other accompanying drawings can also be obtained based on the provided drawings. By reading the detailed description of the non-limiting embodiments with reference to the following drawings, other features, objectives, and advantages of the present invention will become more obvious:
[0046] Figure 1 Schematic structural diagram of an adaptive depth camera in an embodiment of the present invention;
[0047] Figure 2 Flowchart of steps when a processor processes in an embodiment of the present invention;
[0048] Figure 3 Flowchart of steps for marking a primary candidate region in an embodiment of the present invention;
[0049] Figure 4 Flowchart of steps for marking an occluded region in an embodiment of the present invention;
[0050] Figure 5 Schematic structural diagram of another adaptive depth camera in an embodiment of the present invention;
[0051] Figure 6 Flowchart of steps when another processor processes in an embodiment of the present invention. Detailed implementation manners
[0052] The present invention will be described in detail below in conjunction with specific embodiments. The following embodiments will help those skilled in the art to further understand the present invention, but do not limit the present invention in any form. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present invention, several modifications and improvements can still be made. These all belong to the protection scope of the present invention.
[0053] In the description and claims of the present invention and the above-mentioned drawings, the terms "first", "second", "third", "fourth", etc. (if any) are used to distinguish similar objects and do not necessarily describe a specific order or sequence. It should be understood that the data used in this way can be interchanged under appropriate circumstances so that the embodiments of the present invention described herein, for example, can be implemented in an order other than those illustrated or described herein. In addition, the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or device comprising a series of steps or units need not be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products, or devices.
[0054] An adaptive depth camera provided by an embodiment of the present invention aims to solve the problems existing in the prior art.
[0055] The technical solution of the present invention and how the technical solution of the present application solves the above technical problems will be described in detail below with specific embodiments. These specific embodiments can be combined with each other, and the same or similar concepts or processes may not be repeated in some embodiments. The embodiments of the present invention will be described below with reference to the drawings.
[0056] The present invention uses an infrared speckle projector to project infrared speckles that are less affected by light. The infrared receiver stably receives the reflected signal to ensure that the processor can generate accurate speckle patterns and depth map sequences, effectively overcoming light interference; it accurately captures dynamic targets by calculating the motion vectors between adjacent frames and presents their morphological changes in combination with the depth map sequence; it determines the initial candidate region by means of the speckle gradient change rate and determines the occlusion region based on the spatio-temporal consistency between adjacent frames, efficiently handling the occlusion problem and reducing information loss; all links from signal transmission and reception to data processing cooperate closely, and can automatically adapt to different working conditions such as light, dynamic, and complex scenes, demonstrating strong environmental adaptability and stability.
[0057] Figure 1 It is a schematic structural diagram of an adaptive depth camera in an embodiment of the present invention. As Figure 1 shown, an adaptive depth camera in an embodiment of the present invention includes:
[0058] An infrared speckle projector for projecting infrared speckles.
[0059] Specifically, the infrared speckle projector is one of the key components of the camera, and its main function is to project infrared speckles onto the target scene. Infrared speckles are infrared light rays with a specific pattern distribution. The reason for choosing the infrared band is that it is minimally affected by changes in ambient light. Under various complex lighting conditions, such as direct strong light, dim light, etc., the infrared speckles can be stably projected onto the surface of the target object, laying the foundation for subsequent depth information acquisition. By projecting a unique speckle pattern, when these speckles illuminate the object surface and are reflected back, different object distances and surface features will cause corresponding changes in the pattern of the reflected speckles, and these changes are important bases for obtaining depth information.
[0060] An infrared receiver for receiving the reflection signal of the infrared speckles.
[0061] Specifically, the infrared receiver is responsible for receiving the signal reflected after the infrared speckles are projected onto the target object. It has high sensitivity and can accurately capture weak reflection signals. After receiving the reflection signal, it converts it into an electrical signal or a digital signal and then transmits it to the processor for subsequent processing. Its stable receiving performance ensures the effective acquisition of the reflection signal by the entire camera system. Regardless of environmental interference, it can ensure the integrity of the signal, providing a reliable data source for the processor to generate accurate depth information.
[0062] A processor for generating a sequence of infrared speckle maps and a sequence of dense depth maps based on the reflection signal, calculating the motion vector between adjacent frames, determining an initial candidate region according to the gradient change rate of the speckles, and then determining the occluded region according to the spatio-temporal consistency between adjacent frames.
[0063] Specifically, as the core control and data processing unit of the camera, the processor undertakes multiple important tasks. First, based on the reflection signal transmitted by the infrared receiver, it generates a sequence of infrared speckle maps and a sequence of dense depth maps through complex algorithm processing. Then, by calculating the motion vector between adjacent frames, it can accurately analyze the position changes and motion trends of the shooting object at different times, thereby achieving precise capture of dynamic targets. In terms of determining the occluded region, the processor first determines the initial candidate region according to the gradient change rate of the speckles. The speckle gradient change reflects the surface characteristics and depth changes of the object, and based on this, it initially screens out the regions where occlusion may occur. Then, it further determines the occluded region according to the spatio-temporal consistency between adjacent frames. By comparing the image information between different frames and comprehensively considering the changes in the time and space dimensions, it accurately identifies the real occluded region, greatly improving the accuracy and integrity of the depth information and reducing the information loss caused by occlusion.
[0064] In some embodiments, the processor further generates a sparse depth map sequence based on the reflected signal and fills the occluded region with the depth values of the sparse depth map sequence. The sparse depth map sequence contains depth data at key positions in the scene. Although these data are sparsely distributed, they have high accuracy and anti-interference ability. Filling the occluded region with the depth values of the sparse depth map sequence can effectively compensate for the missing depth information caused by occlusion. In complex scenes, occlusion between objects is a common phenomenon, which will cause incomplete regions in the depth maps obtained by conventional methods. By filling the occluded region, the depth map can more completely present the overall scene, providing more comprehensive data support for applications based on depth information. By filling the occluded region, the depth judgment error caused by information loss is avoided, making the depth data more conform to the actual scene. At the same time, the integrity of the data is improved, making the depth map maintain high usability in various complex scenes. Whether in the fields of robot navigation, virtual reality or security monitoring, complete depth data can improve the performance and reliability of the system, providing a more reliable basis for subsequent data analysis and decision-making.
[0065] Figure 2 This is a flowchart of the steps of a processor during processing in an embodiment of the present invention. As Figure 2 shown, the steps of a processor during processing in an embodiment of the present invention include:
[0066] Step T1: Construct a three-frame temporal analysis window, continuously obtain an infrared speckle image sequence and a corresponding dense depth map sequence at different times, and establish a pixel-level spatio-temporal alignment mapping relationship.
[0067] In this step, the processor constructs a temporal analysis window containing three frames of images, continuously obtains an infrared speckle image sequence and a corresponding dense depth map sequence at different times. Then, a pixel-level spatio-temporal alignment mapping relationship is established between these images and depth maps through a specific algorithm. By constructing a three-frame temporal analysis window, it is possible to observe and analyze data changes from a continuous time dimension, providing richer and more coherent information for accurately calculating the motion vector and determining the occluded region in the subsequent process. The pixel-level spatio-temporal alignment mapping relationship ensures the precise matching of different types of data (images and depth maps) at the same spatio-temporal scale, enabling subsequent analysis to be carried out based on an accurate correspondence relationship, laying a foundation for the precise processing of depth information.
[0068] Step T2: Calculate the motion vector between adjacent frames, and execute Step T3 when the average motion intensity exceeds the dynamic threshold; wherein, the dynamic threshold is determined by the dense depth map.
[0069] In this step, after completing the spatio-temporal alignment mapping, the processor calculates the motion vectors between adjacent frames to measure the motion of objects at different times. Meanwhile, the dynamic threshold is determined by the dense depth map. When the average motion intensity exceeds this dynamic threshold, it indicates that there are obvious dynamic changes in the scene, thus triggering the execution of step T3. The calculation of motion vectors can capture the motion trajectories and speeds of objects, helping to judge the dynamic situation in the scene. The setting of the dynamic threshold enables the camera to dynamically adjust the judgment criteria according to the actual depth information, improving the system's adaptability to dynamic changes in different scenes. When the average motion intensity exceeds the threshold, it indicates that there may be dynamic objects in the scene that require further analysis, and thus enter the next step for more detailed processing.
[0070] Step T3: Locate the speckle feature points in the infrared speckle image, generate an analysis window centered on the speckle feature points, calculate the gradient change rate, and mark it as a primary candidate region when the gradient change rate is less than the gradient threshold.
[0071] In this step, in the infrared speckle image, the processor uses a specific feature extraction algorithm to locate the speckle feature points. An analysis window is generated centered on these speckle feature points, and the gradient change rate is calculated within this window. When the gradient change rate is less than the pre-set gradient threshold, it indicates that the change in this region is relatively stable, and it is marked as a primary candidate region. These regions may be occluded regions or relatively stable background regions. The speckle feature points contain rich depth and object surface information. By analyzing them, the change situation of different regions in the scene can be judged more accurately. Generating an analysis window centered on the feature points can focus on key regions for detailed analysis, improving the processing efficiency. The gradient change rate as a judgment criterion helps to screen out regions that may be occluded or stable, providing a preliminary screening result for accurately determining the occluded region subsequently.
[0072] Step T4: Perform spatio-temporal consistency comparison on the primary candidate region and the adjacent two frames, and retain the occluded regions with change values less than the change threshold in three consecutive frames.
[0073] In this step, a spatio-temporal consistency comparison is performed on the marked primary candidate region and the adjacent two frames. By comparing the change values of these regions in three consecutive frames and comparing them with the change threshold. Only those regions with change values less than the change threshold in three consecutive frames are retained as the finally determined occluded regions. Spatio-temporal consistency comparison is a key step in determining the occluded region. Through multi-frame comparison, misjudgments caused by noise or short-term interference can be effectively excluded. Retaining the regions with change values less than the threshold as occluded regions ensures the stability and consistency of the determined occluded regions in both time and space, improving the accuracy of occluded region recognition, and providing a reliable basis for subsequent improvement and processing of depth information.
[0074] In some embodiments, when calculating the motion vectors between adjacent frames, an automatic window confirmation matching range is set for the infrared speckle images of the adjacent frames; the size of the automatic window is determined by the depth change rate on the dense depth map. The size of the automatic window is not fixed, but is determined by the depth change rate on the dense depth map, forming a dynamic adaptive matching mechanism. This solution associates depth information with the image matching range. When the depth change rate is fast, it indicates that the objects in the scene are moving violently and their positions change greatly. At this time, the automatic window will increase accordingly, so as to cover a larger range where the objects may appear, ensuring that when calculating the motion vectors, the position change details caused by the rapid movement of the objects will not be missed, effectively improving the ability to capture the motion information of fast-moving objects. For example, when photographing a vehicle traveling at high speed, a larger automatic window can completely track the position changes of the vehicle between different frames.
[0075] Conversely, when the depth change rate is slow, it means that the object movement is relatively stable and the changes are relatively subtle. A smaller automatic window can accurately focus on the key parts of the object, carefully capture its subtle changes, and at the same time reduce the unnecessary calculation range, greatly improving the calculation efficiency. For example, when photographing a slowly rotating mechanical part, a smaller window can accurately track the subtle rotational changes on the surface of the part and avoid invalid calculations of the surrounding irrelevant areas.
[0076] This feature breaks the limitations of traditional fixed-window matching, enabling the camera to intelligently adjust the matching range according to the dynamic changes of the actual scene, significantly improving the accuracy and efficiency of motion vector calculation, providing a more reliable data basis for subsequent processes such as depth information processing and occlusion area judgment based on motion analysis, and enhancing the adaptability and stability of the entire adaptive depth camera system in complex dynamic scenes.
[0077] Figure 3 This is a flowchart of the steps for marking the primary candidate area in the embodiments of the present invention. As Figure 3 shown, the steps for marking the primary candidate area in the embodiments of the present invention include:
[0078] Step T31: Binarize the infrared speckle image to obtain the light spot area, and then obtain the light spot feature points.
[0079] In this step, first, the infrared speckle image is binarized. By setting a suitable threshold, the pixel points in the image are divided into two categories, namely the speckle area (usually represented by white) and the non-speckle area (usually represented by black). In this way, the original infrared speckle image with continuously changing gray values is transformed into an image with only two colors, black and white, making the speckle area more prominent. After obtaining the clear speckle area, specific algorithms are used to extract the speckle feature points from the speckle area. These feature points are usually representative positions such as the center of the speckle and the turning points of the edge.
[0080] The principle of image binarization is based on the distribution of image gray values. By setting a threshold, the pixel points with gray values higher than the threshold are set to one value (such as 255, representing white), and the pixel points with gray values lower than the threshold are set to another value (such as 0, representing black). The algorithm for extracting speckle feature points depends on the geometric features and gray distribution features of the speckle. For example, the feature points can be determined by detecting the centroid of the speckle and the curvature change of the contour.
[0081] Image binarization can simplify the image information, highlight the speckle area, and facilitate the subsequent extraction of speckle feature points. Accurately obtaining the speckle feature points is the key to subsequent analysis. These feature points contain the key information of the infrared speckle and provide a basis for determining the analysis window and calculating the gradient change rate.
[0082] Step T32: Generate an analysis window with the speckle feature points as the center and the speckle area as the boundary.
[0083] In this step, an analysis window is generated with the speckle feature points obtained in step T31 as the center and the boundary of the speckle area as the limit. This analysis window is like a magnifying glass, focusing on the speckle feature points and the surrounding speckle area for subsequent detailed analysis of this area.
[0084] A rectangular or circular window area is defined by determining the center and the boundary, ensuring that the window can completely contain the speckle information related to the speckle feature points, and at the same time, it will not contain too much irrelevant background information, thereby improving the pertinence and efficiency of the analysis.
[0085] Generating the analysis window can focus the analysis scope on the key area, avoid unnecessary calculations for the entire image, and greatly improve the processing efficiency. At the same time, the setting method with the speckle feature points as the center ensures that the analysis window can accurately cover the speckle information related to this feature point, providing a guarantee for accurately calculating the gradient change rate.
[0086] Step T33: Inside the analysis window, use the gradient operator to calculate the gradient change value of each pixel point, and calculate the standard deviation of the gradient change values of all pixel points inside the analysis window as the gradient change rate.
[0087] In this step, within the generated analysis window, each pixel point is calculated using a gradient operator (such as Sobel operator, Prewitt operator, etc.) to obtain the gradient change value of each pixel point. These gradient change values reflect the gray-scale change of the pixel point in the horizontal and vertical directions. Then, the standard deviation of the gradient change values of all pixel points within the analysis window is calculated, and the obtained standard deviation is used as the gradient change rate of this analysis window.
[0088] The gradient operator calculates the gradient value of a pixel point by weighted calculation of the gray-scale values of the pixel point and its neighboring pixel points, thereby reflecting the local change situation of the image. The standard deviation is used to measure the degree of dispersion of a set of data. Here, by calculating the standard deviation of the gradient change values, the overall degree of dispersion of the gradient change within the entire analysis window, that is, the gradient change rate, can be obtained.
[0089] The gradient change rate can quantitatively analyze the change situation of the light spot area within the analysis window. A larger gradient change rate indicates that the gray-scale change in this area is relatively intense, and there may be features such as the edge and texture of an object; a smaller gradient change rate indicates that this area is relatively stable, and it may be an occluded area or a relatively uniform background area, providing a quantitative basis for subsequent judgment of the primary candidate area.
[0090] Step T34: Mark it as a primary candidate area when the gradient change rate is less than the gradient threshold.
[0091] In this step, the gradient change rate calculated in step T33 is compared with a pre-set gradient threshold. When the gradient change rate is less than the gradient threshold, it indicates that the change of the light spot area within this analysis window is relatively stable, conforming to the characteristics of an occluded area or a relatively stable background area, and this area is marked as a primary candidate area.
[0092] By setting a fixed threshold as the judgment criterion, the continuous gradient change rate data is divided into two categories, namely the area that meets the condition (the gradient change rate is less than the threshold) and the area that does not meet the condition, thereby screening out possible primary candidate areas.
[0093] Marking the primary candidate area is an important preliminary step in determining the occluded area. By comparing the gradient change rate with the threshold, the area where occlusion may exist can be initially screened out, providing a candidate range for subsequent further determination of the true occluded area through spatio-temporal consistency comparison, reducing the workload of subsequent processing, and improving the accuracy and efficiency of determining the occluded area.
[0094] Figure 4 This is a flowchart of the steps for marking the occluded area in an embodiment of the present invention. As Figure 4 shown, the steps for marking the occluded area in an embodiment of the present invention include:
[0095] Step T41: Extract the region corresponding to the primary candidate region in two adjacent frames according to the spatio-temporal alignment mapping relationship.
[0096] In this step, based on the pixel-level spatio-temporal alignment mapping relationship established in Step T1, in two adjacent frame images, accurately locate and extract the region corresponding to the primary candidate region marked in Step T34. Since the spatio-temporal alignment mapping relationship details the corresponding relationship of pixel points between different frame images, it can ensure that the extracted regions are consistent in both the time and space dimensions.
[0097] The spatio-temporal alignment mapping relationship is established based on feature matching and geometric transformation of images. By matching the feature points in different frame images, determine their position changes in different frames, and then use this position change information to construct a geometric transformation model, thereby realizing the mapping of pixel-level corresponding relationships. When extracting the corresponding region, according to this mapping model, convert the pixel coordinates of the primary candidate region into two adjacent frame images to determine the corresponding region range.
[0098] Extracting the corresponding region provides a basis for subsequent calculation of brightness differences. By comparing the features of the primary candidate region and the corresponding regions in two adjacent frames, it is possible to analyze the change situation of this region in the time dimension, which helps to determine whether this region is an occluded region, because the change of the occluded region between different frames is usually relatively stable.
[0099] Step T42: Calculate the brightness difference between the primary candidate region and the corresponding regions in two adjacent frames to obtain a change value.
[0100] In this step, for the extracted primary candidate region and the corresponding regions in two adjacent frames, calculate the brightness difference between them through a specific algorithm, and take the obtained brightness difference value as the change value. When calculating the brightness difference, various methods can be used, such as calculating the sum of the squares of the differences in the brightness values of the corresponding pixel points in the two regions and then taking the average value to measure the overall brightness difference degree between the two regions.
[0101] The brightness of an image is determined by the gray values of pixel points, and different object or scene regions have different brightness performances in the image. When a certain region is an occluded region, its brightness change between different frames is relatively small; while if it is a moving object or a region affected by light changes, the brightness change will be relatively large. By calculating the brightness difference, this change degree can be quantified, providing data support for judging the nature of the region.
[0102] The change value can intuitively reflect the changes in the primary candidate regions between two adjacent frames. A smaller change value indicates that the region is relatively stable in the time dimension and is more likely to be an occluded region; while a larger change value indicates that there are obvious changes in the region, which may be a moving object or a region affected by other factors, thus providing a quantitative basis for screening occluded regions.
[0103] Step T43: Retain the primary candidate regions with change values less than the change threshold as occluded regions.
[0104] In this step, the change values calculated in step T42 are compared with a pre-set change threshold. Only those primary candidate regions with change values less than the change threshold will be retained as the finally determined occluded regions. This process is like a filter that filters out the primary candidate regions that do not conform to the characteristics of occluded regions through the change threshold.
[0105] Setting the change threshold is based on a summary of a large amount of experimental data and practical application scenarios. By statistically analyzing the distribution of change values of different types of regions (occluded regions, moving regions, background regions, etc.) under different conditions, a reasonable threshold is determined so that the regions below this threshold are likely to be occluded regions.
[0106] After this step, the true occluded regions can be accurately screened out from the primary candidate regions, providing accurate data for subsequent depth information processing. Accurately determining the occluded regions can avoid information errors or omissions caused by occlusion during the generation and analysis of depth maps, improve the depth camera's understanding and processing capabilities of complex scenes, and enhance the performance of various applications based on depth information.
[0107] In some embodiments, in step T43, morphological analysis is also performed on the primary candidate regions to remove noise interference. Before comparing the change value with the change threshold, morphological analysis is first performed on the primary candidate regions. The specific operation is to use morphological operations such as erosion and dilation. The erosion operation gradually removes the boundary pixels of the primary candidate regions to eliminate those isolated small pixel points that may be caused by noise; the dilation operation expands the remaining regions to restore the real region pixels that were mistakenly deleted due to the erosion operation and fill some holes caused by noise. By alternately performing erosion and dilation operations multiple times, the boundary of the primary candidate region can be effectively smoothed, noise interference can be removed, and the region morphology can be made closer to the real occluded region.
[0108] Morphological operations are based on the morphological structure features of images. The erosion operation scans each pixel point of the primary candidate region with a predefined structuring element (such as a rectangle, a circle, etc.). If all the pixel points covered by the structuring element belong to this region, then the central pixel point is retained; otherwise, it is deleted, which realizes the removal of boundary noise pixels. The dilation operation is the opposite. When the center of the structuring element coincides with a pixel point in the region, all the pixels covered by the structuring element are included in the region, thus expanding the region and filling the holes. Through such operations based on morphological structures, the primary candidate region can be optimized from the perspective of the geometric shape of the image, reducing the influence of noise on subsequent judgments.
[0109] In complex actual scenarios, images are easily interfered by various noises, which may cause the primary candidate region to contain some false regions or abnormal pixel points. If not processed, these noises may misjudge some parts that are not occlusion regions as occlusion regions, or affect the accurate judgment of real occlusion regions. Removing noise interference through morphological analysis can purify the primary candidate region, improve the accuracy of subsequent comparison with the change threshold, ensure that the finally determined occlusion region is more reliable, provide a cleaner and more accurate data basis for subsequent depth information processing, and further improve the performance of the depth camera in complex scenarios.
[0110] Figure 5 This is a schematic structural diagram of another adaptive depth camera in an embodiment of the present invention. As Figure 5 shown, another adaptive depth camera in an embodiment of the present invention includes:
[0111] A sparse infrared speckle projector for projecting sparse infrared speckles.
[0112] Specifically, the sparse infrared speckle projector projects sparse infrared speckles. These speckles form a specific pattern distribution in the scene, which is used to assist in the acquisition of depth information. Based on specific optical designs and infrared light source technologies, infrared light is emitted at a certain spacing and distribution pattern to form a sparse speckle pattern. These speckles will be projected onto the surface of the target object. When encountering the object surface, they will produce different reflections according to factors such as the distance and shape of the object surface. Due to its sparsity, it can be used for some specific depth calculations and analyses in subsequent processing. For example, when dealing with occlusion regions, the sparse depth map can provide key information.
[0113] A dense infrared speckle projector for projecting dense infrared speckles.
[0114] Specifically, the dense infrared speckle projector projects dense infrared speckles, that is, a more dense infrared speckle pattern is formed in the scene. By precisely controlling the emission of the infrared light source and the configuration of the optical elements, the emitted infrared light forms closely arranged speckles in the target area. Compared with sparse speckles, dense speckles can cover the object surface more meticulously and capture richer detailed information on the object surface. During the depth calculation process, the dense infrared speckles can provide a richer data basis for generating a dense depth map with higher accuracy.
[0115] An infrared receiver for receiving the reflection signals of the sparse infrared speckles and the dense infrared speckles.
[0116] Specifically, the infrared receiver receives the reflection signals of the sparse infrared speckles and the dense infrared speckles. It has a detector array sensitive to infrared light, which can sense the infrared speckle light signals reflected from the surface of the target object. These detectors convert the received light signals into electrical signals and perform preliminary signal processing, such as amplification, filtering, etc., to improve the quality and processability of the signals. Subsequently, the processed signals are transmitted to the processor as the original data source for generating sequences of infrared speckle maps, depth maps, etc.
[0117] A processor for generating a sequence of infrared speckle maps, a sequence of dense depth maps, and a sequence of sparse depth maps according to the reflection signals, calculating the motion vectors between adjacent frames based on the dense depth map, determining the initial candidate regions according to the gradient change rate of the speckles, then determining the occlusion regions according to the spatio-temporal consistency of the initial candidate regions between adjacent frames, and finally filling the depth values of the occlusion regions with the sparse depth map.
[0118] Specifically, the processor generates a sequence of infrared speckle maps, a sequence of dense depth maps, and a sequence of sparse depth maps according to the reflection signals transmitted by the infrared receiver; calculates the motion vectors between adjacent frames based on the generated dense depth map to analyze the motion state of the objects in the scene; determines the initial candidate regions by analyzing the gradient change rate of the speckles. The gradient change rate of the speckles reflects the change of the speckles in space and can be used to locate the regions where depth changes or occlusions may occur; further determines the occlusion regions based on the spatio-temporal consistency of the initial candidate regions between adjacent frames. The spatio-temporal consistency analysis considers the change rules of the candidate regions in the time and space dimensions, excludes some misjudged regions caused by noise or other interferences, and more accurately identifies the true occlusion regions; fills the depth values of the determined occlusion regions with the sparse depth map to obtain more complete and accurate depth information.
[0119] When generating the sequence, the reflection signals are decoded and analyzed, and a specific algorithm is used to convert the signals into the corresponding image sequence. For example, through the principle of triangulation and the matching analysis of the speckle pattern, the distance information carried in the reflection signals is converted into a sequence of depth maps.
[0120] When calculating the motion vector, algorithms such as the optical flow method are used to perform feature matching and analysis between adjacent dense depth maps. By comparing the depth changes of corresponding points in different frames, the displacement and direction of the object between two frames are calculated, thereby obtaining the motion vector.
[0121] When determining the candidate region, the gradient of the infrared speckle pattern is calculated. Regions with large changes in speckle gradient may correspond to discontinuities on the object surface or areas with large depth changes, and these regions are initially identified as initial candidate regions.
[0122] When determining the occluded region, in the time dimension, the changes of the initial candidate region in multiple frames of images are tracked, and combined with the spatial position relationship and speckle characteristics, it is judged which regions are occluded during the movement process, thereby determining the final occluded region.
[0123] When filling the depth value, since sparse speckles can better penetrate occluded regions or provide unique depth information in some cases, the depth data corresponding to the occluded region in the sparse depth map is used to fill the depth value of the occluded region according to a certain interpolation or fusion algorithm, optimizing the integrity and accuracy of the depth map.
[0124] Figure 6 This is the flowchart of the steps when another processor in the embodiment of the present invention is processing. As Figure 6 shown, the steps when another processor in the embodiment of the present invention is processing include:
[0125] Step T1: Construct a three-frame time-series analysis window, continuously acquire a sequence of infrared speckle images and corresponding dense depth map sequences at different times, and establish a pixel-level spatio-temporal alignment mapping relationship;
[0126] Step T2: Calculate the motion vector between adjacent frames, and execute Step T3 when the average motion intensity exceeds the dynamic threshold; wherein, the dynamic threshold is determined by the dense depth map;
[0127] Step T3: Locate the speckle feature points in the infrared speckle image, generate an analysis window centered on the speckle feature points, calculate the gradient change rate, and mark it as a primary candidate region when the gradient change rate is less than the gradient threshold;
[0128] Step T4: Perform spatio-temporal consistency comparison on the primary candidate region and the adjacent two frames, and retain the occluded regions with change values less than the change threshold in three consecutive frames;
[0129] The above steps are the same as those in the previous embodiment and will not be elaborated here.
[0130] Step T5: Obtain the sparse depth map of the infrared speckle image; the infrared speckle image is aligned with the sparse depth map.
[0131] In this step, the sparse depth map is a depth map generated from the speckle signals projected by a sparse infrared speckle projector, which contains the depth information of some points in the scene. The processor uses a depth calculation algorithm similar to that for generating a dense depth map (such as triangulation) based on the reflected signals of the sparse infrared speckles to generate the sparse depth map. To ensure that the depth information in the sparse depth map can be accurately filled into the occluded regions of the dense depth map in subsequent steps, the infrared speckle image and the sparse depth map need to be aligned. The processor will use an image registration algorithm similar to that in Step T1 to align the sparse depth map and the infrared speckle image at the pixel level, so that they are consistent in space and time.
[0132] Step T6: Remove the depth values of the occluded regions in the dense depth map and fill in the depth values of the corresponding regions of the sparse depth map.
[0133] In this step, the processor marks or deletes the depth values of the occluded regions determined in Step T4 in the dense depth map for subsequent depth value filling. Since the sparse depth map can provide the depth information of the occluded regions in some cases, the processor extracts the depth values corresponding to the occluded regions in the sparse depth map and fills them into the corresponding positions in the dense depth map. During the filling process, some interpolation or smoothing processing may be required to ensure that the filled depth values transition naturally with the depth values of the surrounding regions and improve the quality of the depth map.
[0134] This specification also provides an intelligent device, including the adaptive depth camera in any of the foregoing embodiments. This specification exemplarily describes the intelligent device. Those skilled in the art can understand that the description of the intelligent device in this specification is only for the purpose of explaining the intelligent device to facilitate those skilled in the art to better understand the role of the adaptive depth camera in the intelligent device, and should not constitute a limitation on the protection scope.
[0135] The intelligent device includes a moving component and an adaptive depth camera.
[0136] The moving component is a key component for intelligent devices to achieve spatial displacement. It can adopt diverse forms according to different application scenarios and design requirements. For example, in a wheeled mobile robot, the moving component is the wheels, as well as the motors and transmission devices that drive the wheels to rotate. By reversing and adjusting the rotation speed of the motors, the robot can perform actions such as moving forward, backward, and turning on a plane. In intelligent devices such as drones, the moving component is the propellers and the supporting power system. Relying on the lift and thrust generated by the high-speed rotation of the propellers, the drone can achieve ascending, descending, hovering, and flying in various directions in the air. The moving component endows the intelligent device with flexible spatial movement capabilities, enabling it to reach different positions and perform various tasks.
[0137] The adaptive depth camera consists of an infrared speckle projector, an infrared receiver, and a processor. The infrared speckle projector projects infrared speckles onto the target scene. These speckles are less affected by light and can stably adhere to the object surface. The infrared receiver receives the reflected signals and transmits them to the processor. The processor generates sequences of infrared speckle maps, dense depth maps, and sparse depth maps based on the signals, calculates the motion vectors between adjacent frames, determines the occluded areas according to the speckle gradient change rate and the spatio-temporal consistency of adjacent frames, and also fills the occluded areas with the depth values of the sparse depth map sequence.
[0138] The functions of the adaptive depth camera mainly include:
[0139] Environmental perception and navigation assistance: Provide key environmental information for the movement of the moving component. By obtaining the depth information of the surrounding environment, the intelligent device can understand whether there are obstacles ahead, the undulation of the terrain, etc. For example, in an autonomous driving vehicle, the adaptive depth camera continuously senses the distances and positions of objects such as roads, vehicles, and pedestrians around the vehicle. The moving component (such as the power and steering systems of the vehicle) adjusts the driving direction and speed based on this information to achieve safe and intelligent driving navigation.
[0140] Target recognition and tracking: Using the generated depth map sequences and motion vectors, the adaptive depth camera can accurately recognize and track specific targets. In a security surveillance drone, the depth camera can identify targets such as people and vehicles and continuously track their movement trajectories. The moving component then adjusts the position and angle of the drone in real time according to the movement of the target to ensure that the target is always within the surveillance range.
[0141] Scene Reconstruction and Data Analysis: The depth information generated by a depth camera can be used to construct a three-dimensional scene model, providing a more comprehensive understanding of the scene for intelligent devices. In an intelligent floor cleaning robot, the depth camera scans the room to generate a three-dimensional map of the room, and the moving components plan the cleaning path based on the map, improving the cleaning efficiency and coverage rate. At the same time, the data obtained by the depth camera can also be used for data analysis to help intelligent devices make more reasonable decisions and optimize their own behaviors and task executions.
[0142] The various embodiments in this specification are described in a progressive manner. Each embodiment focuses on the differences from other embodiments, and the same or similar parts among the embodiments can be referred to each other. The above description of the disclosed embodiments enables those skilled in the art to implement or use the present invention. Various modifications to these embodiments will be obvious to those skilled in the art, and the general principles defined herein can be implemented in other embodiments without departing from the spirit or scope of the present invention. Therefore, the present invention will not be limited to these embodiments shown herein, but will conform to the widest scope consistent with the principles and novel features disclosed herein.
[0143] The specific embodiments of the present invention have been described above. It should be understood that the present invention is not limited to the above specific embodiments, and those skilled in the art can make various deformations or modifications within the scope of the claims, which do not affect the essence of the present invention.
Claims
1. An adaptive depth camera, characterized in that: include: An infrared speckle projector, used for projecting infrared speckles; An infrared receiver, used for receiving a reflection signal of the infrared speckle; The processor is used to generate an infrared speckle image sequence and a dense depth image sequence according to the reflection signal, calculate the motion vector between adjacent frames, determine the initial candidate area according to the gradient change rate of the speckle, and then determine the occlusion area according to the spatiotemporal consistency between adjacent frames.
2. The adaptive depth camera according to claim 1, characterized in that: The processor also generates a sparse depth map sequence according to the reflection signal, and fills the occluded area with depth values of the sparse depth map sequence.
3. The adaptive depth camera according to claim 1, characterized in that: The processor includes: Step T1: construct a three-frame timing analysis window, continuously acquire infrared speckle image sequences and corresponding dense depth map sequences at different times, and establish a pixel-level spatiotemporal alignment mapping relationship; Step T2: Calculate the motion vector between adjacent frames, and execute step T3 when the average motion intensity exceeds a dynamic threshold; wherein the dynamic threshold is determined by the dense depth map; Step T3: locating a speckle feature point in the infrared speckle image, generating an analysis window with the speckle feature point as the center, calculating the gradient change rate, and marking it as a primary candidate region when the gradient change rate is less than a gradient threshold; Step T4: performing a temporal and spatial consistency comparison between the primary candidate region and two adjacent frames, and retaining the occluded regions whose change values in three consecutive frames are less than a change threshold.
4. The adaptive depth camera according to claim 3, characterized in that: When calculating the motion vector between adjacent frames, an automatic window is set for the infrared speckle images of adjacent frames to confirm the matching range; the size of the automatic window is determined by the depth change rate on the dense depth map.
5. The adaptive depth camera according to claim 3, characterized in that: Step T3 includes: Step T31: binarizing the infrared speckle image to obtain a spot area, and then obtaining a spot feature point; Step T32: Generate an analysis window with the light spot feature point as the center and the light spot area as the boundary; Step T33: in the analysis window, using the gradient operator to calculate the gradient change value of each pixel point, and obtaining the standard deviation of the gradient change values of all the pixels points in the analysis window as the gradient change rate; Step T34: When the gradient change rate is less than the gradient threshold, it is marked as a primary candidate region.
6. The adaptive depth camera according to claim 3, characterized in that: Step T4 includes: Step T41: extracting a region corresponding to the primary candidate region in two adjacent frames according to the spatiotemporal alignment mapping relationship; Step T42: Calculate the brightness difference between the primary candidate region and the corresponding regions in two adjacent frames to obtain a change value; Step T43: retaining the primary candidate regions whose change values are less than the change threshold as occlusion regions.
7. The adaptive depth camera according to claim 6, characterized in that: In step T43, morphological analysis is also performed on the primary candidate regions to remove noise interference.
8. An adaptive depth camera, characterized in that: include: A sparse infrared speckle projector, used for projecting sparse infrared speckles; A dense infrared speckle projector, used for projecting dense infrared speckles; An infrared receiver, used for receiving reflection signals of the sparse infrared speckle and the dense infrared speckle; A processor is used to generate an infrared speckle image sequence, a dense depth image sequence and a sparse depth image sequence according to the reflection signal, calculate the motion vector between adjacent frames according to the dense depth image, determine the initial candidate area according to the gradient change rate of the speckle, determine the occlusion area according to the spatiotemporal consistency of the initial candidate area between adjacent frames, and finally fill the depth value of the occlusion area with the sparse depth image.
9. The adaptive depth camera according to claim 8, characterized in that: The processor includes: Step T1: construct a three-frame timing analysis window, continuously acquire infrared speckle image sequences and corresponding dense depth map sequences at different times, and establish a pixel-level spatiotemporal alignment mapping relationship; Step T2: Calculate the motion vector between adjacent frames, and execute step T3 when the average motion intensity exceeds a dynamic threshold; wherein the dynamic threshold is determined by the dense depth map; Step T3: locating a speckle feature point in the infrared speckle image, generating an analysis window with the speckle feature point as the center, calculating the gradient change rate, and marking it as a primary candidate region when the gradient change rate is less than a gradient threshold; Step T4: performing a temporal and spatial consistency comparison between the primary candidate region and two adjacent frames, and retaining the occlusion region whose change value in three consecutive frames is less than the change threshold; Step T5: acquiring a sparse depth map of the infrared speckle image; aligning the infrared speckle image with the sparse depth map; Step T6: removing the depth value of the blocked area in the dense depth map, and filling the depth value of the corresponding area of the sparse depth map.
10. A smart device, characterized in that: The adaptive depth camera comprises any one of claims 1 to 9.