Depth camera and intelligent device

By combining the depth camera technology of speckle projector and flood projector, infrared and depth maps are generated, which solves the problem of low recognition accuracy of transparent objects in the prior art, and realizes high-precision recognition of transparent objects and depth perception capabilities suitable for multi-scene.

CN119936839APending Publication Date: 2025-05-06SHENZHEN GUANGJIAN TECH CO LTD +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510081605.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-20
Publication Date
2025-05-06

AI Technical Summary

Technical Problem

Existing depth camera technology is difficult to obtain the depth information of transparent objects stably, resulting in low recognition accuracy and difficult to meet the needs of industry, consumer electronics and other fields for accurate perception of transparent objects.

Method used

The depth camera technology combining speckle projector and flood projector is used to identify the range and depth of transparent objects through the generation of infrared images and the fusion of depth images. This technology uses a speckle projector to capture the details of the object and the flood projector to supplement the light of the scene, generate dense and sparse depth maps, and identify them in combination with infrared maps.

Benefits of technology

It improves the recognition accuracy of transparent objects and reduces the recognition blind spots. It is suitable for a variety of application scenarios, including industrial inspection, smart home and consumer electronics.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119936839A_ABST
    Figure CN119936839A_ABST
Patent Text Reader

Abstract

A depth camera is characterized in that the depth camera comprises a speckle projector used for projecting infrared speckles; the floodlight projector is used for projecting floodlight; the receiver is used for receiving reflection signals of the infrared speckles and / or the floodlight and generating an infrared image; the processor is used for respectively generating a dense depth map and a sparse depth map according to the infrared map, and inputting the dense depth map, the sparse depth map and the infrared map into a recognition model to recognize a transparent object range; wherein the transparent object range is a point corresponding to the dense depth map. According to the invention, the transparent object range can be quickly and accurately identified.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of depth cameras, and in particular to a depth camera and a smart device. Background Art

[0002] In the wave of modern technology-driven industries, depth cameras have become the core components for accurate perception and intelligent decision-making in many fields. From quality control in industrial manufacturing, automated sorting in logistics and warehousing, to environmental interaction in smart homes and road condition perception in autonomous driving, the three-dimensional spatial information provided by depth cameras lays a key foundation for system operation.

[0003] Traditional depth camera technology has developed a relatively mature system for processing conventional opaque objects. Depth cameras based on binocular vision, lidar, structured light and other principles can efficiently capture the reflected light on the surface of an object, and then accurately calculate the depth information of the object, helping the machine to complete a series of tasks such as recognition, positioning, and grasping. However, with the continuous expansion of application scenarios, the problem of identifying transparent objects has become increasingly prominent.

[0004] Due to their special optical properties, transparent objects have high permeability to light, making the reflected signal of the projected light extremely weak compared to opaque objects. In the food packaging industry, a large number of transparent plastic boxes and glass bottles carry products on the production line; in the field of electronic manufacturing, transparent optical lenses, display cover plates and other components are the focus of quality inspection; in the new retail scenario, the display of goods made of transparent materials and self-service checkout also require accurate identification. Existing depth camera technology is often unable to cope with these scenarios. It is difficult to stably obtain reflection signals of sufficient strength to generate accurate depth maps, resulting in difficulties in the subsequent identification process.

[0005] At present, some temporary solutions to the problem of transparent object recognition have significant drawbacks. Although simply increasing the power of the light source can enhance the reflected signal to a certain extent, it will lead to a surge in energy consumption and serious heating of the equipment, which will not only increase the operating cost, but may also shorten the service life of the equipment and reduce the stability of the system. If we only optimize at the algorithm level, if the original data quality is poor and the reflected light information is missing, no matter how sophisticated the algorithm is, it will be difficult to accurately restore the complete outline and spatial range of the transparent object, causing the recognition accuracy to stagnate.

[0006] Therefore, it is extremely necessary to develop a depth camera technology specifically for transparent object recognition. By integrating speckle projectors and floodlight projectors, ensuring sufficient and diverse light projection from the light source end, using the receiver to collect weak reflection signals in all directions to generate infrared images, and then cooperating with the processor's depth map generation and intelligent recognition model, it is expected to systematically overcome the difficulties in transparent object recognition, meet the urgent needs of various industries for accurate perception and efficient processing of transparent objects, fill the gaps in existing technologies, and inject new vitality into the intelligent upgrading of the industry.

[0007] The disclosure of the above background technology content is only used to assist in understanding the inventive concept and technical solution of the present invention. It does not necessarily belong to the prior art of this patent application. In the absence of clear evidence that the above content has been disclosed on the filing date of this patent application, the above background technology should not be used to evaluate the novelty and creativity of the present application. Summary of the invention

[0008] To this end, the present invention provides a transparent object detection solution based on infrared images to solve the requirements for the range and depth of transparent objects in various application scenarios.

[0009] In a first aspect, the present invention provides a depth camera, characterized in that it includes:

[0010] A speckle projector, used for projecting infrared speckles;

[0011] A floodlight projector, for projecting floodlight;

[0012] A receiver, used for receiving the reflected signal of the infrared speckle and / or the floodlight, and generating an infrared image;

[0013] A processor is used to generate a dense depth map and a sparse depth map according to the infrared map, and input the dense depth map, the sparse depth map and the infrared map into a recognition model to identify the range of transparent objects; wherein the transparent object range is the points corresponding to the dense depth map.

[0014] Optionally, the depth camera is characterized in that the speckle projector further comprises:

[0015] The speckle modulation unit is used to change the density, size and projection angle of the infrared speckle in real time according to the feedback from the processor.

[0016] Optionally, the depth camera is characterized in that the floodlight projector further comprises:

[0017] A floodlight adjustment unit is used to change the intensity of the floodlight in real time according to the feedback from the processor.

[0018] Optionally, the depth camera is characterized in that the processor further generates a high-precision depth map, including:

[0019] Step S1: extracting features from the dense depth map, the sparse depth map and the infrared map respectively to obtain dense features, sparse features and infrared features, and assigning weights according to the dense depth map, the sparse depth map and the infrared map respectively;

[0020] Step S2: Generate a high-precision depth map based on Bayesian probability model fusion.

[0021] Optionally, the depth camera is characterized in that the dense depth map is obtained by using a time-of-flight algorithm or a triangulation method, and the sparse depth map is obtained by using a feature matching method.

[0022] Optionally, the depth camera is characterized in that the processor, when identifying the range of a transparent object, comprises:

[0023] Step S1: extracting features from the dense depth map, the sparse depth map and the infrared map respectively to obtain dense features, sparse features and infrared features, and assigning weights according to the dense depth map, the sparse depth map and the infrared map respectively;

[0024] Step S3: fusing the dense features, the sparse features and the infrared features according to the weights to obtain a fusion graph;

[0025] Step S4: using a recognition model to recognize the fusion image to obtain a transparent object range.

[0026] Optionally, the depth camera is characterized in that the recognition model is a hierarchical multi-scale convolutional neural network, and when the first-layer large-scale convolution kernel is processed, a multi-channel coarse-grained feature map is output; the middle-layer convolution layer receives the coarse-grained feature map, convolves and refines it while extracting key features from the first layer through lateral connections to supplement the output of middle-layer features; the bottom-layer small-scale convolution focuses on the bottom-layer details, merges them with the middle-layer features, and outputs the prediction results through the fully connected layer to determine the range of corresponding points of transparent objects in the dense depth map.

[0027] Optionally, the depth camera is characterized in that the recognition model also includes a self-supervised learning module during training; the self-supervised learning module runs synchronously with the recognition model, and the common features output by the self-supervised module are injected into the recognition model at intervals of a certain number of training steps; the recognition model is based on a back-propagation algorithm, combining the loss of labeled data with the loss of feature consistency transmitted by the self-supervised module, and jointly updates the weights to continuously optimize the transparent object range recognition capability.

[0028] Optionally, the depth camera is characterized in that the speckle projector and the floodlight projector can project infrared light of different bands, and the projection ratio of infrared light of different bands can be adjusted according to the infrared image.

[0029] In a second aspect, the present invention provides a smart device, characterized by comprising a depth camera as described in any one of the above items.

[0030] Compared with the prior art, the present invention has the following beneficial effects:

[0031] The present invention is equipped with both a speckle projector and a floodlight projector. The infrared speckle projected by the speckle projector has a unique speckle pattern. When encountering an object, the reflected speckle pattern will be distorted due to the surface shape and distance of the object, providing fine features for depth information extraction; the floodlight emitted by the floodlight projector can evenly illuminate the scene over a large area, supplementing those areas that are difficult to cover with speckles, especially in the face of transparent objects with complex shapes, as the floodlight can capture reflected signals from more angles. The two work together, and the receiver can collect more comprehensive reflected signals, making the generated infrared image more complete and minimizing the blind spots in transparent object recognition caused by signal loss.

[0032] The processor of the present invention generates a dense depth map and a sparse depth map. The dense depth map is rich in detail information and can accurately reflect the tiny undulations and contour changes on the surface of the object; the sparse depth map focuses on key feature points, while retaining important spatial structure information and reducing data redundancy. Figure 1 With the same input recognition model, the model is able to extract clues from different levels of depth information, just like building a high-rise building requires both fine brick and stone splicing (dense depth map) and stable framework support (sparse depth map). Multi-dimensional data fusion greatly improves the accuracy of transparent object range recognition and effectively avoids misjudgment and missed judgment.

[0033] The present invention uses a complete set of processes including light source, reception and processing to accurately locate points on a dense depth map corresponding to transparent objects, and can clearly outline the range of transparent objects, thus having a wide range of application scenarios. BRIEF DESCRIPTION OF THE DRAWINGS

[0034] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following will briefly introduce the drawings required for use in the embodiments or the prior art descriptions. Obviously, the drawings in the following descriptions are only embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on the provided drawings without creative work. By reading the detailed description of the non-limiting embodiments with reference to the following drawings, other features, purposes and advantages of the present invention will become more obvious:

[0035] Figure 1 is a schematic structural diagram of a depth camera according to an embodiment of the present invention;

[0036] Figure 2 A schematic diagram of a network architecture in an embodiment of the present invention;

[0037] Figure 3 Schematic diagram of various images in an embodiment of the present invention;

[0038] Figure 4 is a schematic diagram of the structure of another depth camera in an embodiment of the present invention;

[0039] Figure 5 is a schematic diagram of the structure of another depth camera in an embodiment of the present invention;

[0040] Figure 6 A flowchart of the steps of generating a high-precision depth map in an embodiment of the present invention; and

[0041] Figure 7 The figure is a flow chart of steps for identifying the range of a transparent object in an embodiment of the present invention. DETAILED DESCRIPTION

[0042] The present invention is described in detail below in conjunction with specific embodiments. The following embodiments will help those skilled in the art to further understand the present invention, but are not intended to limit the present invention in any form. It should be noted that, for those of ordinary skill in the art, several variations and improvements may be made without departing from the concept of the present invention. These all belong to the protection scope of the present invention.

[0043] The terms "first", "second", "third", "fourth", etc. (if any) in the specification and claims of the present invention and the above-mentioned drawings are used to distinguish similar objects, and are not necessarily used to describe a specific order or sequence. It should be understood that the data used in this way can be interchanged where appropriate, so that the embodiments of the present invention described herein, for example, can be implemented in an order other than those illustrated or described herein. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusions, for example, a process, method, system, product or device that includes a series of steps or units is not necessarily limited to those steps or units that are clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.

[0044] An embodiment of the present invention provides a depth camera, which aims to solve the problems existing in the prior art.

[0045] The following specific embodiments are used to describe in detail the technical solution of the present invention and how the technical solution of the present application solves the above technical problems. The following specific embodiments can be combined with each other, and the same or similar concepts or processes may not be repeated in some embodiments. The embodiments of the present invention will be described below in conjunction with the accompanying drawings.

[0046] The present invention simultaneously arranges a speckle projector and a flood projector, thereby generating two different depth maps, namely a dense depth map and a sparse depth map, which are combined with the infrared image and input into a recognition model to obtain the range of transparent objects. The invention has a better recognition effect on transparent objects and can output an accurate range of transparent objects in real time, greatly improving the recognition effect on transparent objects.

[0047] Figure 1 FIG. 1 is a schematic diagram of the structure of a depth camera according to an embodiment of the present invention. Figure 1 As shown, a depth camera in an embodiment of the present invention includes:

[0048] Speckle projector, used to project infrared speckles.

[0049] Specifically, the speckle projector is a component that uses a laser light source to generate infrared speckle patterns. Infrared speckle can be random speckle or coded speckle. The speckle is formed by the projection of multiple independent small light beams, which appear as multiple independent small light spots on the image. Infrared speckle can be used for structured light measurement, and has a strong light projection density and can be projected over a long distance. These speckle patterns are emitted in the infrared band. Infrared light is chosen because it is relatively stable under the interference of ambient light, and is invisible to the human eye and does not cause visual interference to the user. The projected infrared speckle is irradiated onto the surface of objects in the scene. The different shapes, materials, distances and other characteristics of the surface of the objects will cause the speckle to produce unique reflective deformations. The subsequent receiver captures these reflected speckles, which can provide key information for depth calculation and assist in constructing the three-dimensional structure of the scene.

[0050] Flood projector, used to project flood light.

[0051] Specifically, floodlight projectors are usually based on infrared light-emitting diode (LED) arrays. These LEDs can emit a large area of ​​relatively uniform infrared light in a short period of time, covering a certain field of view in front of the camera. Through circuit control, the LED's luminous intensity, pulse frequency and other parameters are adjusted to achieve stable floodlight projection.

[0052] A receiver is used to receive the reflected signal of the infrared speckle and / or the floodlight to generate an infrared image.

[0053] Specifically, the receiver is generally an infrared image sensor, and is specially designed to be sensitive to the infrared band. The surface of the sensor is covered with pixel units. When the reflected signals of infrared speckle and floodlight reach the sensor, the photons hit the photodiodes in the pixels, resulting in the accumulation of electronic charge. After the integration, amplification, analog-to-digital conversion and other processing processes of the circuit, the optical signal is converted into a digital signal, and finally an infrared image representing the intensity distribution of the reflected light is output. It is the key link for the entire depth camera to obtain raw data, accurately capture the reflected signal, and provide a data basis for the subsequent processor to generate a depth map. The performance indicators such as the resolution and sensitivity of the captured image directly affect the overall accuracy of the depth camera.

[0054] A processor is used to generate a dense depth map and a sparse depth map according to the infrared map, and input the dense depth map, the sparse depth map and the infrared map into a recognition model to identify the range of transparent objects; wherein the transparent object range is the points corresponding to the dense depth map.

[0055] Specifically, the processor uses the received infrared image to obtain a depth map based on a time-of-flight algorithm or a structured light algorithm. For example, in the structured light method, the deformation of the speckle pattern reflected back is analyzed, and the depth information of the scene midpoint is calculated by matching the pre-set speckle template and the actual received speckle image, combined with the principle of triangulation, thereby generating a dense or sparse depth map. Whether the structured light technology obtains a dense depth map or a sparse depth map is determined by the density of the speckle midpoint.

[0056] For infrared images, a time-of-flight algorithm can also be used to obtain a depth map, which is a dense depth map.

[0057] A feature point extraction algorithm is used to identify significant feature points from the infrared image, such as pixels at the edges and corners of objects. The depth is calculated only for these key feature points. Compared with a dense depth map, the amount of data is much smaller, forming a sparse depth map.

[0058] Different dense depth maps and sparse depth maps can be selected according to different application scenarios. For example, a sparse depth map is generated by using speckle, and a dense depth map is obtained by using the time-of-flight algorithm; a dense depth map is generated by using speckle, and a sparse depth map is obtained by using the feature point extraction algorithm; a dense depth map is obtained by using the time-of-flight algorithm, and a sparse depth map is obtained by using the feature point extraction algorithm.

[0059] The generated dense depth map, sparse depth map and original infrared Figure 1The pre-trained recognition model is input. The recognition model uses a deep learning algorithm to learn a large amount of data features of scenes containing transparent objects. It can analyze the depth consistency of corresponding points on different depth maps, changes in infrared reflection intensity, and other features to determine which points belong to transparent objects, and then accurately identify the range of transparent objects, making up for the shortcomings of traditional depth perception methods in identifying transparent objects.

[0060] In some embodiments, the speckle projector and the floodlight projector can project infrared light of different bands, and the projection ratio of infrared light of different bands can be adjusted according to the infrared image. The speckle projector and the floodlight projector have the ability to emit infrared light of different bands, thanks to the multi-infrared light source design adopted internally. For example, the speckle projector may be equipped with a plurality of infrared laser sources with different central wavelengths, and through an optical switching device or a combination of beam splitters, it can selectively output infrared lasers of a specific band to generate speckles; the floodlight projector may integrate infrared light emitting diode (LED) arrays of different bands, and rely on circuit control to activate LEDs of corresponding bands to achieve multi-band floodlight projection.

[0061] After receiving the infrared image, the processor will conduct a series of image analysis. On the one hand, the grayscale distribution of different areas is observed. The grayscale value is related to the intensity of reflected light. Objects of different materials and distances have different reflection characteristics for infrared light of different bands, so the grayscale distribution can indirectly reflect the reflection of infrared light of each band. On the other hand, the image texture clarity and noise level will also be analyzed, because inappropriate band combinations may lead to blurred images and increased noise.

[0062] Based on the above analysis, the processor sends instructions to the speckle projector and the floodlight projector to dynamically adjust the projection ratio of infrared light in different bands. For example, when it is found that the infrared light of a certain band makes the outline of the target object in the image blurred, while another band can highlight the outline, the projection ratio of the blurred band is reduced and the projection ratio of the clear outline band is increased. If the scene is strongly disturbed by ambient light, the projection ratio of infrared light in those bands that are less affected by ambient light is increased to ensure the quality of the final infrared image, thereby improving the accuracy of subsequent depth map generation and transparent object recognition.

[0063] This embodiment can adapt to complex scenes. Infrared light of different bands has great differences in penetrability, anti-interference, and interaction characteristics with objects. By flexibly adjusting the projection ratio, the camera can adapt to complex scenes such as smoke, strong light interference, and long-distance shooting, and always obtain usable reflection signals.

[0064] This embodiment can also improve recognition accuracy. Accurately adapted band combinations help to more clearly outline the contours of objects and distinguish materials. For identifying transparent objects, suitable bands can enhance the unique signals of refracted and reflected light of transparent objects, allowing subsequent recognition models to more accurately lock the range of transparent objects.

[0065] Figure 2 FIG. 1 is a schematic diagram of a network architecture in an embodiment of the present invention. Figure 2 As shown in the figure, the recognition model includes an encoder and a decoder. The dense depth map, sparse depth map and infrared image are input into the encoder for feature extraction. Then they are decoded step by step by the decoder to finally obtain the range of transparent objects. Figure 2 The image in the middle is a scene arrangement diagram, which is only used to illustrate the scene and does not participate in the processing in this embodiment.

[0066] Figure 3 Schematic diagram of various images in the embodiments of the present invention. Figure 3 As shown, a is a scene layout diagram, which is only used to illustrate the scene situation and does not participate in the calculation and processing of this embodiment. b is an infrared image, c is a dense depth image, and d is a sparse depth image. b, c, and d are input into the recognition model for processing. Figure 3 It can be seen that the transparent object range is not a strict range of transparent objects, but a depth range corresponding to the depth map, which is different from the range for RGB images in the prior art.

[0067] Figure 4 FIG. 1 is a schematic diagram of the structure of another depth camera in an embodiment of the present invention. Figure 4 As shown, compared with the above-mentioned embodiment, another speckle projector in a depth camera in the embodiment of the present invention further includes:

[0068] The speckle modulation unit is used to change the density, size and projection angle of the infrared speckle in real time according to the feedback from the processor.

[0069] Specifically, a close communication link is established between the speckle modulation unit and the processor, which receives information from the processor at all times. The processor issues instructions to the speckle modulation unit based on the analysis results of the current shooting scene, such as the distance of the target object, the complexity of the scene, and the requirements for the accuracy of the depth information. This feedback control is a dynamic cycle process. While the camera continues to collect images and generate depth maps, the processor continuously evaluates the data quality and adjusts the speckle projection parameters as needed.

[0070] When the speckle density needs to be changed, the speckle modulation unit adjusts the internal optical elements, such as by controlling the micro-movement of elements such as the diffraction grating or switching gratings of different specifications. To increase the density, the grating lines will be made finer, so that more interference spots are generated after the laser passes through, and vice versa. When shooting small objects at close range, a higher speckle density can provide more detailed texture information, assisting in accurately locating the tiny undulations on the surface of the object, thereby improving the accuracy of depth calculation; when shooting large scenes at a long distance, a lower density of speckles is sufficient to cover the target area and reduce the amount of data processing.

[0071] The speckle size can be changed by means of optical zoom, lens combination, etc. For example, through a set of lenses with electrically adjustable spacing, the laser beam is converged or diverged before projection, thereby changing the size of the final speckle projected onto the surface of the object. For large objects, larger speckles help to quickly cover the surface of the object and improve the efficiency of reflecting signal collection; for small and delicate objects, small-sized speckles can fit the contour of the object and obtain reflected light that is more in line with the object's true shape, making subsequent depth restoration more accurate.

[0072] With the help of precise angle adjustment devices such as mechanical rotation, universal joint structure or micro-electromechanical system (MEMS) mirrors. MEMS mirrors can change angles quickly and accurately under the drive of electrical signals, thereby changing the direction of speckle projection. When there are obstructions in the scene or a complex three-dimensional scene needs to be scanned, flexible adjustment of the projection angle can avoid obstacles, fill light and speckle projection are performed on different areas of the scene, ensuring that enough reflection signals can be collected in every corner to generate a complete and accurate depth map.

[0073] Figure 5 FIG. 1 is a schematic diagram of the structure of another depth camera in an embodiment of the present invention. Figure 5 As shown, another flood light projector in a depth camera in an embodiment of the present invention further includes:

[0074] A floodlight adjustment unit is used to change the intensity of the floodlight in real time according to the feedback from the processor.

[0075] Specifically, the floodlight adjustment unit maintains a stable communication connection with the processor and can receive feedback from the processor in real time. The processor will give adjustment instructions based on multiple factors, including the lighting environment of the current scene, such as strong light indoors, cloudy outdoors, and other different lighting conditions, as well as the material characteristics of the photographed object, such as highly reflective metal and light-absorbing dark fabrics.

[0076] The floodlight adjustment unit usually relies on the control of the infrared light-emitting diode (LED) inside the floodlight projector to change the floodlight intensity. If a constant current drive circuit is used, the brightness of the LED can be accurately adjusted by changing the size of the drive current. When the processor determines that the current scene light is dark and the floodlight needs to be enhanced, it will send a command to the adjustment unit to increase the current, causing the LED to emit stronger infrared light; conversely, if the scene light is sufficient, in order to avoid signal interference caused by excessive brightness, the current is reduced to reduce the floodlight intensity.

[0077] In dim environments, increasing the floodlight intensity allows more light to reach the surface of objects, allowing the receiver to capture more reflected signals, making up for the disadvantage of insufficient ambient light, making the generated infrared image clearer and more complete, and reducing the lack or error of depth information caused by too weak light.

[0078] For objects with strong reflective ability, appropriately reducing the floodlight intensity can prevent the receiver from being saturated due to strong light reflection and causing signal distortion. For objects with good light absorption, increasing the floodlight intensity will help obtain sufficient reflected light to assist in subsequent depth calculations, ensuring that objects of different materials can achieve accurate depth detection in the field of view of the depth camera.

[0079] Figure 6 FIG. 1 is a flow chart of steps for generating a high-precision depth map according to an embodiment of the present invention. Figure 6 As shown, a step of generating a high-precision depth map in an embodiment of the present invention includes:

[0080] Step S1: extracting features from the dense depth map, the sparse depth map and the infrared map respectively to obtain dense features, sparse features and infrared features, and assigning weights according to the dense depth map, the sparse depth map and the infrared map respectively.

[0081] In this step, the dense depth map contains the depth information of a large number of points in the scene, and the processor uses algorithms such as convolutional neural networks (CNN) to extract its features. The convolution layer of CNN scans the depth map through filters of different sizes to capture subtle depth change patterns such as continuous undulations on the surface of objects and edge gradients. The output dense features can reflect the rich geometric structure details of the scene.

[0082] For sparse depth maps, the processor focuses on those discrete feature points. Using algorithms such as Scale Invariant Feature Transform (SIFT) or Speeded Up Robust Features (SURF) related derivative algorithms, it extracts key attributes such as local depth gradients and directions around the feature points to form sparse features, which are good at characterizing the key contours and structural turning points of objects.

[0083] For infrared images, the processor extracts the light intensity edge information that reflects the contour of the object on the one hand, and on the other hand, since different materials have different reflection and absorption characteristics of infrared light, it also mines out material-related features, and uses texture analysis methods such as grayscale co-occurrence matrix to capture texture features such as the periodicity and uniformity of the infrared reflected light intensity distribution, and summarizes them into infrared features.

[0084] The processor will assign weights based on the reliability of different data sources. Generally speaking, dense depth maps are given higher weights because they contain rich details and are more reliable in close-up and texture-rich scenes; sparse depth maps have advantages in capturing large-scale object structures and locating long-distance targets, and are assigned reasonable weights accordingly; infrared images play an irreplaceable role in supplementing information related to object materials and lighting compensation, and are also given appropriate weights.

[0085] The complexity of the scene also affects the weight distribution. In simple and regular scenes, the weight of dense depth maps is increased to quickly and accurately restore the scene; in complex occlusion and multi-material mixed scenes, the weight of sparse depth maps and infrared maps is appropriately increased to assist in dealing with complex situations.

[0086] Step S2: Generate a high-precision depth map based on Bayesian probability model fusion.

[0087] In this step, the Bayesian probability model is based on Bayes' theorem: In the depth map fusion scenario, different features are regarded as events. The depth map related events to be fused are recorded as A1 (corresponding to dense depth map), A2 (corresponding to sparse depth map), and A3 (corresponding to infrared image), and the observed current scene data is regarded as B.

[0088] The model first estimates the probability P(A) of different depth maps to generate accurate depth information based on prior knowledge. i ), this prior probability is related to the weight distribution mentioned above. Then calculate the likelihood probability P(B|A) of the current scene data given each depth map. i ), which reflects the fit between each depth map and the actual scene data.

[0089] By using the Bayesian formula, the posterior probability P(A i |B), weighted sum of depth values ​​in different depth maps based on these probabilities. For example, the depth value of a point in a dense depth map is d1, in a sparse depth map is d2, and in an infrared map is d3. The corresponding posterior probabilities are P1, P2, and P3, respectively. The fused high-precision depth value d = P1d1 + P2d2 + P3d3 is traversed through all points to finally generate a high-precision depth map.

[0090] Figure 7 FIG. 1 is a flow chart of steps for identifying the range of a transparent object in an embodiment of the present invention. Figure 7 As shown, in an embodiment of the present invention, a step of identifying a range of a transparent object includes:

[0091] Step S1: extracting features from the dense depth map, the sparse depth map and the infrared map respectively to obtain dense features, sparse features and infrared features, and assigning weights according to the dense depth map, the sparse depth map and the infrared map respectively.

[0092] This step is the same as the previous step and will not be repeated here.

[0093] Step S3: fusing the dense features, the sparse features and the infrared features according to the weights to obtain a fusion graph.

[0094] In this step, based on the weights assigned previously, the processor uses a weighted summation method to fuse dense features, sparse features, and infrared features. Assume that the dense feature is F d , with weight w d ; The sparse feature is F s , with weight w s ; infrared signature is F i , with weight w i , then the calculation formula of the fused feature F is: F = w d F d +w s F s +w i F i Through this fusion method, the advantages of each data source are integrated, and the generated fusion map contains both fine geometric details and key contour and material lighting information.

[0095] Step S4: using a recognition model to recognize the fusion image to obtain a transparent object range.

[0096] In this step, the recognition model is usually a pre-trained neural network based on a deep learning architecture, such as a convolutional neural network (CNN) or a fully connected neural network (FCN). After the fusion image is input into the model, the model classifies and judges the fusion image based on the learned features of a large number of sample data containing transparent objects and non-transparent objects. It identifies which areas in the image correspond to transparent objects, and through pixel-level annotation and classification, it finally accurately outputs the range of transparent objects, overcoming the difficulty of traditional methods that are difficult to identify transparent objects by processing depth images or infrared images alone.

[0097] In some embodiments, the recognition model is a hierarchical multi-scale convolutional neural network. When the first-layer large-scale convolution kernel is processed, a multi-channel coarse-grained feature map is output; the middle-layer convolution layer receives the coarse-grained feature map, refines it through convolution, and extracts key features from the first layer through lateral connections to output middle-layer features; the bottom-layer small-scale convolution focuses on the bottom-layer details, merges them with the middle-layer features, and outputs the prediction results through the fully connected layer to determine the range of corresponding points of transparent objects in the dense depth map.

[0098] First layer: large-scale convolution kernel processing

[0099] Convolution operation: The first layer of the hierarchical multi-scale convolutional neural network uses a large-scale convolution kernel to perform convolution operations. Large-scale convolution kernels, such as 7×7 or 9×9, are convolved with the input fusion map (obtained from the previous feature fusion step). Due to its large coverage area, each convolution operation can capture a large range of image information at one time, quickly aggregate some macro features of the image, ignore small noise interference, and output a multi-channel coarse-grained feature map. Each channel represents features extracted from different dimensions. For example, one channel focuses on the general outline of an object, and another channel reflects the overall brightness distribution. This coarse-grained feature map builds a basic framework for subsequent processing. It quickly outlines the general structure of the scene, so that the subsequent operations of the network can be optimized based on the general direction, reducing unnecessary detail calculations and improving computing efficiency.

[0100] Middle layer: convolution refinement and lateral connections

[0101] Convolution refinement: The middle convolution layer receives the coarse-grained feature map output by the first layer, and uses a relatively small convolution kernel, such as a 3×3 or 5×5 convolution kernel, to gradually perform multiple convolution operations on the coarse-grained feature map. Each convolution deeply explores more details hidden in the feature map, such as more accurate positioning of object edges, refinement of depth change trends, etc., so that the features are gradually refined.

[0102] Lateral connection: At the same time, the middle convolutional layer also extracts key features from the first layer for supplementation. The lateral connection mechanism can integrate the important global features captured by the first layer into the local features being refined by the middle layer. For example, the layout of large objects in the scene identified by the first layer can be used through lateral connections to assist the middle layer in more accurately locating the object boundaries that may be blurred during the refinement process, so that the middle layer does not lose key general direction information during refinement, and then outputs the middle layer features, which not only retains global perception but also details.

[0103] Bottom layer: small-scale convolution and feature fusion

[0104] Small-scale convolution focuses on details: The underlying small-scale convolution uses extremely small convolution kernels, such as 1×1 or 2×2, to "attack" the details of the underlying image. They focus on tiny differences and fine textures at the pixel level, capturing ultra-fine information that is easily overlooked by the upper layers, such as the extremely subtle light and shadow changes caused by refraction on the surface of transparent objects.

[0105] Feature fusion and fully connected output: The detailed features captured by the bottom layer are then fused with the middle-level features output by the middle layer. This fusion integrates the optimal information at different scales. The fused features are then passed through the fully connected layer, which comprehensively weighs and classifies the features and finally outputs the prediction results, accurately locating the range of the corresponding points of the transparent objects in the dense depth map, allowing the entire network to achieve accurate recognition of transparent objects from macro to micro, from coarse to fine.

[0106] In some embodiments, the recognition model also includes a self-supervised learning module during training; the self-supervised learning module runs synchronously with the recognition model, and the common features output by the self-supervised module are injected into the recognition model at intervals of a certain number of training steps; the recognition model is based on a back propagation algorithm, combining the loss of labeled data with the feature consistency loss transmitted by the self-supervised module, and jointly updates the weights to continuously optimize the transparent object range recognition capability.

[0107] Self-supervised learning module

[0108] Operation mechanism: The self-supervised learning module starts the training process synchronously with the recognition model. During the training process, it outputs common features to the recognition model according to the preset training step interval. The advantage of self-supervised learning is that it does not require large-scale manual annotation of data, and it can mine valuable feature information from the structure and distribution characteristics of the data itself. For example, it can use the natural properties of the image such as rotation invariance and color consistency to allow the model to learn the common features of the image under different transformations. These features are sorted and refined to become common features.

[0109] Feature injection: Every certain number of training steps, the self-supervised learning module injects these common features into the recognition model. This injection operation is like bringing new "knowledge nutrients" to the recognition model, expanding the model's feature reserve, and prompting the model to understand image data from more dimensions, especially those potential features that are difficult to learn through labeled data, thereby assisting the recognition model to cope with complex and changeable transparent object recognition scenarios.

[0110] Joint weight update based on back-propagation

[0111] Loss calculation: When the recognition model is trained, on the one hand, the loss is calculated based on the labeled data. This is the loss measurement method in conventional supervised learning. That is, the range of transparent objects predicted by the model is compared with the actual labeled range to obtain the loss of labeled data, which reflects the degree of deviation between the model prediction result and the facts. On the other hand, after the self-supervisory module injects common features, the model also considers the feature consistency loss. This loss measures the difference and fit between the original features of the model and the newly injected self-supervised common features. If the two are not well integrated and the difference is too large, the loss will become larger.

[0112] Joint update: The recognition model uses the back propagation algorithm to adjust the model weights based on the sum of the labeled data loss and the feature consistency loss. The back propagation algorithm starts from the loss end and propagates the error gradient backward along the hierarchical structure of the neural network, updating the weight parameters of each layer step by step. By combining these two losses to update the weights, the model can not only fit the labeled data and accurately locate transparent objects, but also continuously absorb the common features mined by self-supervised learning, continuously optimize the transparent object range recognition ability, and improve the generalization and robustness of the model.

[0113] This specification also provides an embodiment of a smart device. It should be noted that this embodiment is only used to illustrate the smart device and is exemplary and should not constitute any limitation on the scope of protection.

[0114] Smart devices integrate the above-mentioned depth cameras, and their typical structures are as follows:

[0115] shell:

[0116] Materials and design: Lightweight and strong engineering plastics, aluminum alloys and other materials are often used to ensure the overall durability of the equipment, withstand a certain degree of bumps and falls, and take into account portability. The shape of the shell varies according to the function of the equipment. Consumer products often pursue fashion, simplicity, and fit the human body's holding habits; industrial smart devices tend to be tough and regular, easy to install and fix.

[0117] Protection function: The casing is well sealed to prevent dust and water vapor from entering, thus avoiding damage to the internal precision electronic components. For some smart devices used outdoors, the casing also has sun protection, rain protection, cold protection and other protective features to ensure that the equipment can operate stably in harsh environments.

[0118] Sensing layer:

[0119] Depth camera module: As the core sensing component, it includes a speckle projector, a flood projector, a receiver and a processor. The speckle projector accurately projects infrared speckles, and the flood projector supplements uniform infrared light as needed. The two work together, and the receiver collects the reflected signal to generate an infrared map. The processor further processes it to obtain various depth maps and identify the range of transparent objects.

[0120] Other auxiliary sensors: often used with accelerometers and gyroscopes to sense the posture and motion state changes of the device, assist in the calibration of the depth camera data, and make the acquired visual information more accurate; some devices are also equipped with ambient light sensors to measure the current ambient light intensity and cooperate with the depth camera to adjust the projected light intensity in time.

[0121] Processing layer:

[0122] Main processor: Most of them are high-performance, low-power chips, such as customized multi-core processors in smartphones or industrial-grade microcontrollers used in industrial intelligent devices. It coordinates various tasks of the device, receives data from the depth camera, runs complex algorithms, coordinates the interaction between different components, performs in-depth analysis, fusion, and calculation on the data, and finally outputs usable visual information.

[0123] Storage unit: includes cache, random access memory (RAM) and flash memory, etc. Cache accelerates data reading, RAM ensures temporary storage of data when the system is running, and flash memory retains the operating system, application, device configuration information and data collected by the depth camera for a long time to ensure that data is not lost after the device is powered off.

[0124] Interaction layer:

[0125] Display screen: Consumer-grade smart devices are usually equipped with high-definition touch screens, through which users can intuitively view the images captured by the depth camera, set parameters, and preview shooting effects; the screens of industrial equipment focus on displaying key monitoring data and status prompts, and the size and resolution depend on actual needs.

[0126] Input components: physical buttons, touchpad, voice input module, etc., to facilitate users to control the device, adjust the working mode of the depth camera, and issue commands, such as switching shooting scenes, starting the transparent object recognition function, etc.

[0127] Power module:

[0128] Batteries: Lithium batteries are commonly used in consumer devices. They have high energy density and can be charged and discharged cyclically, which can meet the needs of daily use. Industrial smart devices may use large-capacity lead-acid batteries or lithium battery packs to ensure long-term continuous operation.

[0129] Power management chip: responsible for monitoring battery power, adjusting charging current and voltage, ensuring safe and efficient charging of the battery, and reasonably allocating power to various components to achieve device power consumption optimization.

[0130] With the help of the above-mentioned depth camera, this kind of smart device has powerful visual perception capabilities and can be used in many fields:

[0131] Consumer electronics: After embedding the depth camera in consumer electronics devices such as smartphones and tablets, cooler and more practical photo functions can be realized. For example, when taking portraits, the depth camera's precise depth perception and the algorithm can achieve a more natural background blur effect, making the subject more prominent; when taking 3D panoramic photos or videos, dense and sparse depth maps work together to accurately capture the three-dimensional structure of the scene, making the user feel as if they are there, immersing themselves in the moment of shooting. Moreover, the function of identifying the range of transparent objects can also help users take creative photos, such as taking clear photos of exhibits through a glass window.

[0132] Smart home: The smart sweeping robot equipped with this depth camera can carry out detailed 3D modeling of the home environment. It can accurately identify the position and outline of furniture and obstacles, even transparent glass coffee tables, so as to plan a more scientific and efficient cleaning route, reduce collisions, and improve cleaning efficiency. The smart door lock is equipped with a depth camera, which can identify the depth information of the person in front of the door, assist in judging the person's body shape and posture, and combine with transparent object recognition to prevent criminals from using transparent obstructions to deceive the door lock, greatly enhancing security.

[0133] Industrial inspection: On industrial production lines, smart devices carry depth cameras that can quickly scan the three-dimensional appearance of products. For some electronic products with transparent shells or parts, such as the transparent dial of smart watches and the transparent protective shell of mobile phones, they can accurately locate their range and detect whether there are scratches, bubbles and other appearance defects to ensure product quality. In addition, in the assembly of parts, the high-precision depth map obtained by the depth camera helps robots to accurately grasp and place parts, improving assembly accuracy and automation level.

[0134] Logistics and warehousing: Logistics robots use depth cameras to quickly locate goods in warehouses with numerous shelves, identify items in transparent packaging, and accurately calculate the volume and stacking status of goods, facilitating inventory management and cargo handling, optimizing logistics and warehousing processes, saving labor costs, and improving cargo turnover efficiency.

[0135] In this specification, each embodiment is described in a progressive manner, and each embodiment focuses on the differences from other embodiments, and the same and similar parts between the embodiments can be referred to each other. The above description of the disclosed embodiments enables professionals and technicians in this field to implement or use the present invention. Various modifications to these embodiments will be obvious to professionals and technicians in this field, and the general principles defined herein can be implemented in other embodiments without departing from the spirit or scope of the present invention. Therefore, the present invention will not be limited to these embodiments shown in this article, but will comply with the widest range consistent with the principles and novel features disclosed herein.

[0136] The above describes the specific embodiments of the present invention. It should be understood that the present invention is not limited to the above specific embodiments, and those skilled in the art may make various modifications or variations within the scope of the claims, which do not affect the essence of the present invention.

Claims

1. A depth camera, characterized in that: include: A speckle projector, used for projecting infrared speckles; A floodlight projector for projecting floodlight; a receiver, configured to receive reflection signals of the infrared speckle and / or the flood light and generate an infrared image; A processor is configured to generate a dense depth map and a sparse depth map based on the infrared image, and input the dense depth map, the sparse depth map, and the infrared image into a recognition model to identify a range of transparent objects; wherein the range of transparent objects is the points corresponding to the dense depth map.

2. A depth camera according to claim 1, characterized in that: The speckle projector further comprises: The speckle modulation unit is configured to change the density, size, and projection angle of the infrared speckle in real time according to feedback from the processor.

3. The depth camera according to claim 1, wherein: The floodlight projector further comprises: A floodlight adjustment unit is used to change the intensity of the floodlight in real time according to the feedback from the processor.

4. The depth camera according to claim 1, wherein: The processor also generates a high-precision depth map, including: Step S1: extracting features from the dense depth map, the sparse depth map, and the infrared image respectively to obtain dense features, sparse features, and infrared features, and assigning weights according to the dense depth map, the sparse depth map, and the infrared image respectively; Step S2: Generate a high-precision depth map based on Bayesian probability model fusion.

5. The depth camera according to claim 1, wherein: The dense depth map is obtained by using a time-of-flight algorithm or a triangulation method, and the sparse depth map is obtained by using a feature matching method.

6. The depth camera according to claim 1, wherein: When the processor identifies the range of a transparent object, it includes: Step S1: extracting features from the dense depth map, the sparse depth map, and the infrared image respectively to obtain dense features, sparse features, and infrared features, and assigning weights according to the dense depth map, the sparse depth map, and the infrared image respectively; Step S3: fusing the dense features, the sparse features, and the infrared features according to the weights to obtain a fusion graph; Step S4: using a recognition model to identify the fusion image and obtain the range of transparent objects.

7. The depth camera according to claim 1, wherein: The recognition model is a hierarchical multi-scale convolutional neural network. When the first layer is processed by a large-scale convolution kernel, it outputs a multi-channel coarse-grained feature map; the middle convolution layer receives the coarse-grained feature map, refines it while performing convolution, and extracts key features from the first layer through lateral connections to output middle-level features; the bottom layer's small-scale convolution focuses on the bottom-level details, fuses them with the middle-level features, and outputs the prediction results through the fully connected layer to determine the range of the corresponding points of the transparent object in the dense depth map.

8. The depth camera according to claim 1, wherein: During training, the recognition model also includes a self-supervised learning module; the self-supervised learning module runs synchronously with the recognition model, and the common features output by the self-supervised module are injected into the recognition model every certain number of training steps; the recognition model is based on a back-propagation algorithm, combining the loss of labeled data with the loss of feature consistency transmitted by the self-supervised module, jointly updating the weights, and continuously optimizing the transparent object range recognition capability.

9. The depth camera according to claim 1, wherein: The speckle projector and the floodlight projector can project infrared light of different wavelength bands, and the projection ratio of infrared light of different wavelength bands can be adjusted according to the infrared image.

10. A smart device, characterized in that: A depth camera comprising any one of claims 1 to 9.