Transparent object detection segmentation and depth reconstruction method, system and device based on 3D camera and storage medium

Through a 3D camera-based method, the speckle infrared map and depth map are obtained, and the transparent object range is identified and the depth value is corrected, which solves the accuracy problems of transparent object recognition and depth reconstruction in the prior art, and realizes high-precision transparent object detection and depth reconstruction.

CN119963620APending Publication Date: 2025-05-09SHENZHEN GUANGJIAN TECH CO LTD +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510143035.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-10
Publication Date
2025-05-09

AI Technical Summary

Technical Problem

The prior art is difficult to accurately identify and divide transparent objects, and there are errors in the depth reconstruction of transparent objects.

Method used

Using a 3D camera-based method, by obtaining speckle infrared map and depth map, using the first deep learning network to identify the range of transparent objects, and correct the depth value through the correction module to finally obtain an accurate transparent object depth map.

Benefits of technology

Accurate recognition of the range of transparent objects and accurate correction of depth values ​​are achieved, improving the recognition accuracy of transparent objects in various application scenarios and the accuracy of depth reconstruction.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119963620A_ABST
    Figure CN119963620A_ABST
Patent Text Reader

Abstract

The invention discloses a transparent object detection segmentation and depth reconstruction method, system and device based on a 3D camera, and a storage medium. The method comprises the following steps: S1, obtaining a speckle infrared image and a depth image; s2, recognizing a transparent object range through a first deep learning network; wherein the first deep learning network obtains the range of the transparent object according to the astigmatism characteristic of the speckle irradiated on the transparent object; and S3, correcting the depth value in the transparent object range according to the depth map and the speckle infrared map to obtain a final depth map. According to the invention, the depth values of various transparent objects can be accurately identified.
Need to check novelty before this filing date? Find Prior Art

Description

Background Art

[0002] In recent years, with the development of automation and intelligent technology, the demand for transparent object recognition and corresponding depth reconstruction has become increasingly high, such as in the fields of sweeping robot obstacle avoidance and mapping, autonomous driving, and industrial inspection.

[0003] However, most existing object detection technologies are mainly used for opaque objects, and it is difficult to detect and segment transparent objects. There are also some transparent object detection technologies based on traditional methods or deep learning methods, but most of them are based on a single RGB image for transparent object segmentation, which has great limitations in effect stability and usage scenarios.

[0004] Unlike traditional cameras, 3D cameras can collect more environmental information during actual use, including RGB images, depth images and different types of infrared images. Transparent objects have richer and more obvious features in these images, especially depth and infrared images, compared with single RGB images. In this way, a more stable and accurate range can be obtained through the algorithm. After obtaining the range, the surrounding environment information and the depth characteristics of the transparent object itself can be used to reconstruct its depth.

[0005] The disclosure of the above background technology content is only used to assist in understanding the inventive concept and technical solution of the present invention. It does not necessarily belong to the prior art of this patent application. In the absence of clear evidence that the above content has been disclosed on the filing date of this patent application, the above background technology should not be used to evaluate the novelty and creativity of the present application. Summary of the invention

[0006] To this end, the present invention provides a transparent object detection, segmentation and depth reconstruction solution based on a 3D camera to solve the requirements for the range and depth of transparent objects in various application scenarios.

[0007] In a first aspect, the present invention provides a transparent object detection, segmentation and depth reconstruction method based on a 3D camera, characterized in that it includes:

[0008] Step S1: Obtaining a speckle infrared image and a depth image;

[0009] Step S2: identifying the range of transparent objects through a first deep learning network; wherein the first deep learning network obtains the range of transparent objects according to the astigmatism characteristics of speckle irradiation on the transparent objects;

[0010] Step S3: Correcting the depth value within the range of the transparent object according to the depth map and the speckle infrared map to obtain a final depth map.

[0011] Optionally, the transparent object detection, segmentation and depth reconstruction method based on a 3D camera is characterized in that step S1 comprises:

[0012] Step S11: acquiring two speckle infrared images;

[0013] Step S12: Calculate the disparity according to the two speckle infrared images to obtain a depth map.

[0014] Optionally, the 3D camera-based transparent object detection, segmentation and depth reconstruction method is characterized in that, in step S1, an RGB image is also acquired, and in step S2, at least the RGB image, the speckle infrared image and the depth map are input into the first deep learning network to identify the range of transparent objects.

[0015] Optionally, the transparent object detection, segmentation and depth reconstruction method based on a 3D camera is characterized in that the first deep learning network is a Unet network, comprising an encoder, a first decoder and a second decoder; the first decoder outputs the range of the transparent object; the second decoder outputs the boundary of the transparent object; and there is feature interaction between features of each level of the first decoder and the second decoder.

[0016] Optionally, the transparent object detection, segmentation and depth reconstruction method based on a 3D camera is characterized in that, in step S3, the depth map, the transparent object range and the speckle infrared image are input into a second deep learning network to obtain a final depth map; wherein the second deep learning network corrects the depth value within the transparent object range.

[0017] Optionally, the transparent object detection, segmentation and depth reconstruction method based on a 3D camera is characterized in that step S3 comprises:

[0018] Step S31: performing three-dimensional reconstruction on the portion outside the range of the transparent object according to the depth map and the speckle infrared map;

[0019] Step S32: obtaining a transparent object boundary and identifying the shape of the transparent object boundary;

[0020] Step S33: performing plane fitting according to the boundary of the transparent object to obtain a depth value corresponding to each pixel point within the range of the transparent object.

[0021] Optionally, the transparent object detection, segmentation and depth reconstruction method based on a 3D camera is characterized in that step S33 comprises:

[0022] Step S331: performing straight line fitting on the boundary of the transparent object to obtain a plurality of straight lines in a three-dimensional space;

[0023] Step S332: fitting the plane of three consecutive straight lines to obtain a transparent object;

[0024] Step S333: Calculate the depth value corresponding to each pixel point within the range of the transparent object according to the transparent object.

[0025] In a second aspect, the present invention provides a transparent object detection, segmentation and depth reconstruction system based on a 3D camera, which is used to implement any of the above-mentioned transparent object detection, segmentation and depth reconstruction methods based on a 3D camera, and is characterized by comprising:

[0026] Acquisition module for speckle infrared image and depth image;

[0027] An identification module, configured to identify a range of transparent objects through a first deep learning network; wherein the first deep learning network obtains the range of transparent objects according to the astigmatism characteristics of speckle irradiation on the transparent objects;

[0028] A correction module is used to correct the depth value within the range of the transparent object according to the depth map and the speckle infrared map to obtain a final depth map.

[0029] In a third aspect, the present invention provides a transparent object detection, segmentation and depth reconstruction device based on a 3D camera, characterized in that it includes:

[0030] processor;

[0031] a memory storing executable instructions of the processor;

[0032] The processor is configured to execute the steps of any of the aforementioned methods for transparent object detection, segmentation and depth reconstruction based on a 3D camera by executing the executable instructions.

[0033] In a fourth aspect, the present invention provides a computer-readable storage medium for storing a program, characterized in that when the program is executed, the steps of any of the aforementioned 3D camera-based transparent object detection, segmentation and depth reconstruction methods are implemented.

[0034] Compared with the prior art, the present invention has the following beneficial effects:

[0035] The present invention adopts speckle infrared images and utilizes the fact that the reflection of light spots by transparent objects is different from that of non-transparent objects, so that the reflection signals of the light spots present different characteristics on the image; at the same time, the calculated depth values ​​also present messy characteristics; the present invention can accurately identify the range of transparent objects through deep learning of light spots and depth values, and is suitable for identification scenarios of various transparent objects. BRIEF DESCRIPTION OF THE DRAWINGS

[0036] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following will briefly introduce the drawings required for use in the embodiments or the prior art descriptions. Obviously, the drawings in the following descriptions are only embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on the provided drawings without creative work. By reading the detailed description of the non-limiting embodiments with reference to the following drawings, other features, purposes and advantages of the present invention will become more obvious:

[0037] Figure 1 This is a flowchart of a method for detecting, segmenting and reconstructing a transparent object based on a 3D camera in an embodiment of the present invention;

[0038] Figure 2 This is a flowchart of the steps of obtaining a speckle infrared image and a depth image in an embodiment of the present invention;

[0039] Figure 3 A schematic diagram of a network architecture in an embodiment of the present invention;

[0040] Figure 4 Schematic diagram of various images in an embodiment of the present invention;

[0041] Figure 5 is a flow chart of steps for obtaining a final depth map in an embodiment of the present invention;

[0042] Figure 6 This is a flowchart of the steps of calculating the depth value corresponding to each pixel in an embodiment of the present invention;

[0043] Figure 7 Schematic diagram of a transparent object detection, segmentation and depth reconstruction system based on a 3D camera in an embodiment of the present invention;

[0044] Figure 8 is a schematic structural diagram of a transparent object detection, segmentation and depth reconstruction device based on a 3D camera in an embodiment of the present invention; and

[0045] Fig. 9 Schematic diagram of the structure of a computer-readable storage medium in an embodiment of the present invention. DETAILED DESCRIPTION

[0046] The present invention is described in detail below in conjunction with specific embodiments. The following embodiments will help those skilled in the art to further understand the present invention, but are not intended to limit the present invention in any form. It should be noted that, for those of ordinary skill in the art, several variations and improvements may be made without departing from the concept of the present invention. These all belong to the protection scope of the present invention.

[0047] The terms "first", "second", "third", "fourth", etc. (if any) in the specification and claims of the present invention and the above-mentioned drawings are used to distinguish similar objects, and are not necessarily used to describe a specific order or sequence. It should be understood that the data used in this way can be interchanged where appropriate, so that the embodiments of the present invention described herein, for example, can be implemented in an order other than those illustrated or described herein. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusions, for example, a process, method, system, product or device that includes a series of steps or units is not necessarily limited to those steps or units that are clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.

[0048] The embodiment of the present invention provides a transparent object detection, segmentation and depth reconstruction method based on a 3D camera, aiming to solve the problems existing in the prior art.

[0049] The following specific embodiments are used to describe in detail the technical solutions of the present invention and how the technical solutions of the present application solve the above-mentioned technical problems. The following specific embodiments can be combined with each other, and the same or similar concepts or processes may not be repeated in some embodiments. The embodiments of the present invention will be described below in conjunction with the accompanying drawings.

[0050] The present invention is based on an RGBD camera, obtains a speckle infrared image and a depth map, identifies the range of transparent objects through a first deep learning network, and then corrects the depth value to obtain a final depth map. It can accurately identify transparent objects in various application scenarios and solve the problem of inaccurate identification in the prior art.

[0051] Figure 1 FIG. 1 is a flow chart of the steps of a transparent object detection, segmentation and depth reconstruction method based on a 3D camera in an embodiment of the present invention. Figure 1 As shown, the steps of a transparent object detection, segmentation and depth reconstruction method based on a 3D camera in an embodiment of the present invention include:

[0052] Step S1: Obtain a speckle infrared image and a depth image.

[0053] In this step, the 3D camera has a built-in speckle infrared emitter, which emits a specific pattern of infrared speckle light to the scene. The speckle pattern can be coded structured light or random speckle, etc. The infrared sensor receives the speckle infrared light reflected from the scene, converts it into an electrical signal based on the difference in light intensity, and then generates a speckle infrared image through analog-to-digital conversion, image algorithm processing, etc. This image is critical for the subsequent analysis of transparent objects using speckle characteristics.

[0054] Based on the depth perception technology of 3D cameras, the depth map can be calculated based on the principle of structured light according to the individual speckle infrared image. This can accurately obtain the depth value for non-transparent objects. However, when the target object is a transparent object, when the light spot is irradiated on the transparent object, the light spot is affected by the transparent object, and its reflected signal will be eliminated to a certain extent and will be relatively weak. At the same time, the transparent object will also change the light path, causing the position and shape of the light spot to shift, and ultimately causing a large deviation in the calculated depth value.

[0055] You can also use binocular technology to calculate the depth map using parallax. In this case, the 3D camera needs to be equipped with two structured light receivers to calculate the depth map using two speckle infrared images.

[0056] Step S2: Identify the range of transparent objects through the first deep learning network.

[0057] In this step, the first deep learning network obtains the range of the transparent object according to the scattered light characteristics of the speckle irradiated on the transparent object. The speckle infrared image and the depth map are input into the first deep learning network. The speckle infrared image is selected because the scattering characteristics of infrared light of transparent objects are significantly different from those of ordinary non-transparent objects. This difference can present a unique texture and light intensity distribution pattern on the speckle infrared image, providing key clues for the network to identify transparent objects.

[0058] The first deep learning network can be a convolutional neural network (CNN) that is good at processing image data. The speckle infrared image and the depth map serve as the input data of the first deep learning network. During the training phase, a large number of labeled samples are used to train the network, and the labeled information accurately indicates the boundaries and ranges of transparent objects in the image. The network continuously extracts local features of the speckle infrared image in the convolution layer, performs feature dimensionality reduction in the pooling layer, and integrates global features in the fully connected layer, and gradually learns the mapping relationship between the speckle pattern and the existence of transparent objects, that is, it can distinguish the part belonging to the transparent object from the complex texture and light intensity distribution of the speckle infrared image.

[0059] When a speckle infrared image of an unknown scene is input, the trained network will output a binary or probability map result. In the binary result, for example, white pixels represent the area belonging to transparent objects, and black pixels represent the background, thus clearly outlining the range of transparent objects; the probability map gives the probability of each pixel belonging to a transparent object. By setting a suitable threshold, the range of transparent objects can also be divided.

[0060] Step S3: Correcting the depth value within the range of the transparent object according to the depth map and the speckle infrared map to obtain a final depth map.

[0061] In this step, due to the optical properties of transparent objects, the depth map obtained conventionally often has errors in the transparent object area. This is because the light penetrates the transparent object and the reflection is complicated, making the depth value generated based on the common principle inaccurate, so correction is required.

[0062] Speckle infrared images have unique astigmatism characteristics on transparent objects, which can reflect some laws of light propagation on the surface and inside of transparent objects. Combined with the range of transparent objects that have been identified, focus on these areas and use the additional information in the speckle infrared image to assist in judging the true depth.

[0063] For each pixel within the range of the transparent object, the depth value in the depth map is adjusted by a preset correction algorithm based on the relevant features of the speckle infrared image. The correction algorithm can be a deep learning model, such as obtaining an accurate depth value through training of a second deep learning network. The correction algorithm can also be a spatial calculation model to perform calculations in three-dimensional space.

[0064] In some embodiments, in step S3, the depth map, the transparent object range and the speckle infrared map are input into a second deep learning network to obtain a final depth map; wherein the second deep learning network corrects the depth value within the transparent object range. When entering step S3, a preliminary depth map, a transparent object range identified by the first deep learning network, and an original speckle infrared map have been obtained. These data are fed into the second deep learning network as a whole to provide a comprehensive information basis for subsequent depth correction work. The depth map carries the original distance information of the scene, although there are errors in the transparent object area; the transparent object range clearly defines the target area that needs to be corrected, allowing the network to focus on it; the speckle infrared map contains the unique astigmatism characteristics of transparent objects, providing key clues for correcting the depth value. The second deep learning network can use a convolutional neural network architecture suitable for depth estimation and correction tasks, such as an architecture improved based on the residual network (ResNet). It has multiple convolutional layers, pooling layers and fully connected layers. The convolutional layers are used to extract the features of the input data, the pooling layers are responsible for reducing the data dimension and the amount of calculation, and the fully connected layers integrate the features extracted at different levels to complete the final correction decision. After the network front end receives the input data, each layer processes it step by step and learns the mapping relationship from the combination of input multi-source data to the accurate correction depth value.

[0065] After the input data enters the network, each part is first extracted by different convolutional layers. The features of the depth map reflect the overview of the scene distance, the features of the transparent object range focus on the geometric shape of the target area, and the features of the speckle infrared map capture the scattering characteristics of light on transparent objects. Subsequently, the network merges these features from different sources through a specific fusion layer, so that the information required for correction complements each other and forms a comprehensive judgment basis.

[0066] The fused features are sent to the subsequent network layer, where the depth value within the transparent object range is adjusted based on the pre-learned correction model. The model refers to the astigmatism characteristics in the speckle infrared image, such as speckle deformation and light intensity distribution, and combines the boundary and shape information of the transparent object range to determine the deviation direction and amplitude of the current depth value. For each pixel within the transparent object range, a more accurate depth value is recalculated, and a corrected depth map is gradually constructed.

[0067] After multi-layer network processing, the depth values ​​of all pixels within the transparent object range are corrected, and a complete and accurate final depth map is generated. This depth map not only maintains the original accuracy in the non-transparent object area, but more importantly, solves the problem of inaccurate depth measurement in the transparent object area due to optical characteristics, providing reliable data support for subsequent tasks that rely on accurate depth information, such as 3D scene reconstruction, robot visual navigation, etc.

[0068] In some embodiments, in step S1, an RGB image is also obtained, and in step S2, at least the RGB image, the speckle infrared image and the depth map are input into the first deep learning network to identify the range of transparent objects. The RGB image is obtained using the color imaging module of the 3D camera. The module is usually composed of a color filter such as a Bayer array and a photosensitive element. After the light enters the camera through the lens, it is filtered by the color filter, and the light of different colors is captured by the corresponding photosensitive units respectively. After a series of photoelectric conversion, signal amplification and processing processes, a two-dimensional image containing three color channel information of red (Red), green (Green), and blue (Blue) is finally generated, that is, an RGB image. It intuitively reflects the color and texture information of the scene, is a color visual presentation familiar to the human eye, and can be used to capture the appearance characteristics of the scene, such as the color of the object, the surface pattern, etc. The RGB image, the speckle infrared image and the depth map are used as inputs to the first deep learning network.

[0069] Figure 2 FIG. 1 is a flowchart of steps for obtaining a speckle infrared image and a depth image in an embodiment of the present invention. Figure 2 As shown, in an embodiment of the present invention, a step of obtaining a speckle infrared image and a depth image includes:

[0070] Step S11: Acquire two speckle infrared images.

[0071] In this step, for 3D camera systems, binocular or multi-camera configurations are often used to obtain scene information more accurately. In the case of binocular cameras, two cameras capture the scene synchronously from slightly different perspectives. Due to the difference in perspective, the imaging position of the same object in the two speckle infrared images will be offset.

[0072] At the hardware level, the binocular camera device will be precisely calibrated to ensure that the optical axis parallelism and focal length consistency of the two lenses meet the requirements, so as to accurately analyze the correspondence between the images later. The image sensor and signal processing circuit inside the camera will perform photoelectric conversion and preliminary processing on the light from different lenses, and output two RGB images or speckle infrared images in a standardized format.

[0073] Step S12: Calculate the disparity according to the two speckle infrared images to obtain a depth map.

[0074] In this step, first, the two acquired speckle infrared images are used to find significant feature points using a feature extraction algorithm. Common methods include corner detection algorithms, such as Harris corner detection, which calculates the grayscale changes of pixels in a local window of the image to find those points with drastic grayscale changes, that is, corners. These corners are easier to match and track in two images from different perspectives. In addition, the scale-invariant feature transform (SIFT) is also a common method, which can extract feature descriptors with scale and rotation invariance, making it easier to accurately locate the same features in images from different perspectives.

[0075] Matching feature points is the key to calculating disparity. Between two images, feature points in one image are paired with corresponding feature points in another image based on the similarity of feature descriptors. For example, the normalized cross correlation (NCC) algorithm is used to calculate the correlation of pixel grayscale values ​​in the neighborhood window of two feature points. The higher the correlation, the more likely the two points are images of the same point in space at different viewing angles. After finding the matching point pair, the disparity is determined. Disparity refers to the pixel offset of the feature point in the horizontal direction of the two images.

[0076] According to the principle of triangulation, the depth of scene points can be derived if the camera baseline (the distance between the optical axes of the two cameras), the camera focal length, and the calculated parallax are known. All matching feature point pairs in the two images are traversed, the corresponding depth values ​​are calculated, and filled into the corresponding pixel positions, and finally a depth map representing the distance from each point in the scene to the camera is generated.

[0077] This embodiment combines speckle infrared images with the principle of stereo vision, and the method can provide high-precision and robust depth information, especially for the recognition and depth reconstruction of transparent objects. This embodiment is not only suitable for the detection and segmentation of transparent objects, but also can be applied to various scenarios that require high-precision depth information, such as augmented reality, industrial automation, and medical imaging.

[0078] Compared with structured light technology, binocular technology uses two images of light spots to calculate depth values. Both images used to calculate parallax contain reflection signals of transparent objects, making them more sensitive to signals of transparent objects. The speckle features in the two speckle infrared images are paired through a feature matching algorithm to calculate the parallax. For transparent objects, parallax calculation is particularly critical because the refraction and reflection inside transparent objects are complex, which causes large deviations in the calculation of parallax. Based on the principle of triangulation, the depth values ​​of each point within the range of the transparent object are calculated using the known camera baseline, focal length, and calculated parallax. At this time, the depth values ​​of the transparent object range obtained often have large deviations and have obvious characteristics.

[0079] In some embodiments, Figure 3 As shown, the first deep learning network is a Unet network, which includes an encoder, a first decoder and a second decoder; the first decoder outputs the range of transparent objects; the second decoder outputs the boundary of transparent objects; and the first decoder and the second decoder interact with each other at each level. The Unet network is a classic encoder-decoder architecture that performs well in image segmentation tasks. Its encoder part is responsible for downsampling the input image and gradually extracting high-level semantic features of the image; the decoder performs upsampling operations to restore the feature map to the size of the original image, thereby achieving pixel-level prediction.

[0080] When an image is input into the Unet network, the encoder first processes the image using a convolutional layer. Usually, multiple continuous convolution kernels, such as 3×3 convolution kernels, are used in conjunction with activation functions (such as ReLU) to extract local features of speckle infrared images. As the network layers deepen, downsampling is performed through pooling layers (such as maximum pooling), and the image size gradually decreases, but the feature dimension continues to increase, capturing more and more abstract and semantically rich feature information, which reflects the common characteristics of speckle patterns related to transparent objects at different scales.

[0081] The function of the first decoder is to output the range of transparent objects. It receives high-level features from the encoder, and then reversely performs the upsampling process, gradually restoring the feature map size using transposed convolution or interpolation methods. At each stage of upsampling, it fuses features from the corresponding encoder level. This jump connection can supplement the detail information lost during the downsampling process. For example, through the splicing operation, the high-resolution features of a certain layer of the encoder are merged with the low-resolution features of the current layer of the decoder, so that the output result has both high-level semantics and details. Finally, after a convolution layer, the predicted transparent object range is output, which is presented in the form of a probability map or a binary map.

[0082] The second decoder is responsible for outputting the boundaries of transparent objects. Similar to the first decoder, it also obtains key features from the encoder and performs upsampling and feature fusion at each level. The difference is that it focuses more on the precise positioning of the boundaries. When training the network, the weights are optimized by focusing on the labeled boundary information so that the output results can accurately outline the contours of transparent objects. For example, a more sophisticated loss function is used to impose higher penalties on boundary pixel prediction errors, thereby prompting the network to learn more acute boundary feature discrimination capabilities.

[0083] The feature interaction between each level of the first decoder and the second decoder is the key design of the Unet network in this task. At the same level, the two achieve information complementarity by sharing some features. For example, at a certain middle layer, the range-related features extracted by the first decoder can help the second decoder better understand the boundary position, because the boundary must be at the edge of the transparent object range; conversely, the sharp boundary features captured by the second decoder also provide clues for the first decoder to optimize the range prediction and avoid excessive expansion or contraction of the range prediction. This multi-level feature interaction synergy enhances the accuracy and robustness of the entire network for transparent object segmentation and boundary determination.

[0084] Figure 4 In the figure, a is the speckle infrared image, b is the depth map, c is the transparent object range, and d is the final depth map. As can be seen from a, the speckle infrared image includes multiple scattered spots, which can effectively detect transparent objects. b is the calculated depth map. Due to the existence of transparent objects, the boundaries in the depth map are blurred and the depth values ​​are disordered. The transparent object range c obtained by the first deep learning model can accurately obtain the transparent object range. The final depth map obtained in d can accurately express the depth values ​​including transparent objects in the target scene.

[0085] Figure 5 FIG. 1 is a flowchart of steps for obtaining a final depth map in an embodiment of the present invention. Figure 5 As shown, in an embodiment of the present invention, a step of obtaining a final depth map includes:

[0086] Step S31: performing three-dimensional reconstruction on the portion outside the range of the transparent object according to the depth map and the speckle infrared map.

[0087] In this step, pixel data outside the range of the corresponding transparent objects is extracted from the existing depth map and speckle infrared map. Since the optical properties of transparent objects are complex, excluding them first can reduce the difficulty of subsequent processing and focus on the conventional object area that is easier to process. The filtered data is subjected to noise reduction processing to remove outliers caused by sensor noise and ambient light interference, improve data quality, and ensure the accuracy of subsequent 3D reconstruction.

[0088] Use a point cloud-based 3D reconstruction algorithm, such as the Poisson reconstruction method. First, convert the 2D pixel information into point cloud data in 3D space based on the depth map. The depth value corresponds to the Z coordinate of the point, and the XY coordinate is calculated from the pixel position. In this process, the speckle infrared image assists in determining the surface characteristics of the point cloud and supplements the texture information. Based on these point cloud data, the Poisson reconstruction algorithm solves the Poisson equation to fit a smooth and continuous 3D surface, thereby achieving 3D reconstruction of the scene outside the range of transparent objects.

[0089] Step S32: Obtain the boundary of the transparent object and identify the shape of the boundary of the transparent object.

[0090] In this step, with the help of the transparent object range information output by the first deep learning network, the image morphology algorithm is used to accurately extract the boundary, or the boundary directly output by the first deep learning network in the above embodiment is used. For example, the binary image representing the transparent object range is first corroded, and then dilated. The small noise blocks are removed and the object range is refined by corrosion. The subsequent dilation operation restores the original size of the object, and the boundary is highlighted. You can also use edge detection algorithms, such as Canny edge detection, set a suitable threshold, capture pixels with drastic intensity changes on the transparent object range map, and outline precise boundaries.

[0091] After the boundary is extracted, shape descriptors are used to identify its shape. Commonly used ones include Hu moments, which construct 7 invariant moments by calculating the normalized central moment of the image area. These invariant moments can uniquely describe a shape regardless of its translation, rotation, and scaling. By comparing the Hu moment values ​​corresponding to the boundary of a transparent object with the Hu moment values ​​of various pre-stored shape templates, the shape of the boundary can be quickly identified, such as a circle, square, irregular polygon, etc.

[0092] Step S33: performing plane fitting according to the boundary of the transparent object to obtain a depth value corresponding to each pixel point within the range of the transparent object.

[0093] In this step, considering that the surfaces of many transparent objects can be approximated as planes locally, pixels within a certain neighborhood range within the identified transparent object boundary are collected. The most suitable plane equation is determined based on the boundary. For each pixel within the transparent object range, its coordinates are substituted into the fitted plane equation to solve the corresponding depth value. These calculated depth values ​​are updated to the corresponding pixel positions, and finally a complete and corrected depth map is obtained.

[0094] This embodiment performs plane fitting on transparent objects according to their boundaries in three-dimensional space, so that the transparent objects can be restored more accurately; it can achieve accurate segmentation, depth value correction and three-dimensional reconstruction of transparent objects, which not only improves the overall accuracy but also ensures the integrity of the reconstruction results; through plane fitting, this embodiment can calculate the depth value corresponding to each pixel point within the range of the transparent object, taking into account the special properties of the transparent object (such as light refraction and reflection), thereby improving the accuracy of the depth value.

[0095] Figure 6 FIG. 1 is a flowchart of a step of calculating the depth value corresponding to each pixel in an embodiment of the present invention. Figure 6 As shown, in an embodiment of the present invention, a step of calculating a depth value corresponding to each pixel point includes:

[0096] Step S331: performing straight line fitting on the boundary of the transparent object to obtain a plurality of straight lines in three-dimensional space.

[0097] In this step, based on the transparent object boundary obtained in step S32, this boundary is composed of a series of discrete two-dimensional pixel points. A linear fitting algorithm such as the least squares method is used to process these three-dimensional space points. The least squares method aims to find a straight line that minimizes the sum of the squares of the vertical distances of all points to this straight line. By constructing an error function and solving the parameters that minimize the error, multiple fitting straight lines are determined, which will serve as the basic elements for subsequent plane construction.

[0098] Step S332: Fit the plane to three consecutively connected straight lines to obtain a transparent object.

[0099] In this step, three connected straight lines are selected from the many straight lines that have been fitted. A plane is determined by three points that are not on the same straight line, and the three connected straight lines can provide enough geometric information to determine a plane. The above plane fitting process is repeated continuously, and three connected straight lines in different combinations are selected to obtain multiple local planes. These local planes are spliced ​​together to gradually outline the surface morphology of the entire transparent object in three-dimensional space, simulate the shape of the transparent object from a geometric level, and provide a geometric model basis for the subsequent calculation of the depth value.

[0100] Step S333: Calculate the depth value corresponding to each pixel point within the range of the transparent object according to the transparent object.

[0101] In this step, for each pixel point within the transparent object range, its corresponding position on the fitted three-dimensional transparent object surface needs to be determined. Through coordinate transformation and projection relationship, the two-dimensional coordinates of the pixel point are combined with the depth map and mapped to the three-dimensional space to find the local plane where it is located. For example, the inverse operation of perspective projection is used to restore the pixel coordinates to the three-dimensional space, and then determine which fitting plane area the point falls in.

[0102] Once the plane to which the pixel belongs is determined, its 3D coordinates are substituted into the plane equation to solve the corresponding depth value. After all the pixels within the transparent object range have been processed in this way, the depth value calculation is completed and the corrected depth map is obtained.

[0103] This embodiment can obtain the three-dimensional shape of a transparent object by connecting three straight-line fitting planes. This method takes into account the spatial continuity of the transparent object and can reconstruct the three-dimensional structure of the object more accurately. Even if the transparent object has a complex shape or surface features, a relatively accurate three-dimensional reconstruction result can be obtained by connecting three straight-line fitting planes in sequence. After obtaining the three-dimensional shape of the transparent object, the depth value corresponding to each pixel point within the range of the transparent object can be calculated based on the shape. This depth value calculation method takes into account the three-dimensional structure of the object and is therefore more accurate.

[0104] Figure 7 FIG. 1 is a schematic diagram of a transparent object detection, segmentation and depth reconstruction system based on a 3D camera in an embodiment of the present invention. Figure 7 As shown, in an embodiment of the present invention, a transparent object detection, segmentation and depth reconstruction system based on a 3D camera includes:

[0105] An acquisition module, used to acquire a speckle infrared image and a depth image;

[0106] An identification module, configured to identify a range of transparent objects through a first deep learning network; wherein the first deep learning network obtains the range of transparent objects according to the astigmatism characteristics of speckle irradiation on the transparent objects;

[0107] A correction module is used to correct the depth value within the range of the transparent object according to the depth map and the speckle infrared map to obtain a final depth map.

[0108] This embodiment is based on an RGBD camera to obtain a speckle infrared image and a depth map, identifies the range of transparent objects through a first deep learning network, and then corrects the depth value to obtain a final depth map. It can accurately identify transparent objects in various application scenarios and solve the problem of inaccurate recognition in the prior art.

[0109] The embodiment of the present invention also provides a transparent object detection, segmentation and depth reconstruction device based on a 3D camera, including a processor and a memory, in which executable instructions of the processor are stored. The processor is configured to execute the steps of a transparent object detection, segmentation and depth reconstruction method based on a 3D camera by executing the executable instructions.

[0110] As mentioned above, this embodiment is based on an RGBD camera to obtain a speckle infrared image and a depth map, identifies the range of transparent objects through a first deep learning network, and then corrects the depth value to obtain a final depth map, which can accurately identify transparent objects in various application scenarios and solve the problem of inaccurate recognition in the prior art.

[0111] It will be appreciated by those skilled in the art that various aspects of the present invention may be implemented as systems, methods or program products. Therefore, various aspects of the present invention may be specifically implemented in the following forms, namely: complete hardware implementation, complete software implementation (including firmware, microcode, etc.), or a combination of hardware and software implementations, which may be collectively referred to herein as "circuits", "modules" or "platforms".

[0112] Figure 8 Schematic diagram of a transparent object detection, segmentation and depth reconstruction device based on a 3D camera in an embodiment of the present invention. Figure 8 The electronic device 600 according to this embodiment of the present invention is described. Figure 8 The electronic device 600 shown is only an example and should not bring any limitation to the functions and scope of use of the embodiments of the present invention.

[0113] like Figure 8 As shown, the electronic device 600 is in the form of a general computing device. The components of the electronic device 600 may include, but are not limited to: at least one processing unit 610, at least one storage unit 620, a bus 630 connecting different platform components (including the storage unit 620 and the processing unit 610), a display unit 640, etc.

[0114] The storage unit stores program codes, which can be executed by the processing unit 610, so that the processing unit 610 executes the steps of various exemplary embodiments of the present invention described in the above-mentioned transparent object detection, segmentation and depth reconstruction method based on a 3D camera. For example, the processing unit 610 can execute the following steps: Figure 1 Follow the steps shown in .

[0115] The storage unit 620 may include a readable medium in the form of a volatile storage unit, such as a random access memory unit (RAM) 6201 and / or a cache memory unit 6202 , and may further include a read-only memory unit (ROM) 6203 .

[0116] The storage unit 620 may also include a program / utility 6204 having a set (at least one) of program modules 6205, such program modules 6205 including but not limited to: an operating system, one or more application programs, other program modules, and program data, each of which or some combination may include an implementation of a grid environment.

[0117] Bus 630 may represent one or more of several types of bus structures, including a memory unit bus or memory unit controller, a peripheral bus, an accelerated graphics port, a processing unit, or a local bus using any of a variety of bus architectures.

[0118] The electronic device 600 may also communicate with one or more external devices 700 (e.g., keyboards, pointing devices, Bluetooth devices, etc.), one or more devices that enable a user to interact with the electronic device 600, and / or any device that enables the electronic device 600 to communicate with one or more other computing devices (e.g., routers, modems, etc.). Such communication may be performed through an input / output (I / O) interface 650. Furthermore, the electronic device 600 may also communicate with one or more grids (e.g., a local area network (LAN), a wide area network (WAN), and / or a public grid, such as the Internet) through a grid adapter 660. The grid adapter 660 may communicate with other modules of the electronic device 600 through the bus 630. It should be understood that although Figure 8 Not shown, other hardware and / or software modules may be used in conjunction with electronic device 600, including but not limited to: microcode, device drivers, redundant processing units, external disk drive arrays, RAID systems, tape drives, and data backup storage platforms.

[0119] In an embodiment of the present invention, a computer-readable storage medium is also provided for storing a program, and when the program is executed, the steps of a method for detecting, segmenting, and reconstructing a transparent object based on a 3D camera and performing depth reconstruction are implemented. In some possible implementations, various aspects of the present invention may also be implemented in the form of a program product, which includes a program code. When the program product is run on a terminal device, the program code is used to enable the terminal device to perform the steps of various exemplary implementations of the present invention described in the above-mentioned method for detecting, segmenting, and reconstructing a transparent object based on a 3D camera section of this specification.

[0120] As shown above, this embodiment is based on an RGBD camera to obtain a speckle infrared image and a depth map, identifies the range of transparent objects through a first deep learning network, and then corrects the depth value to obtain a final depth map, which can accurately identify transparent objects in various application scenarios and solve the problem of inaccurate recognition in the prior art.

[0121] Fig. 9 Schematic diagram of the structure of a computer-readable storage medium in an embodiment of the present invention. Fig. 9 As shown, a program product 800 for implementing the above method according to an embodiment of the present invention is described, which can adopt a portable compact disk read-only memory (CD-ROM) and include program code, and can be run on a terminal device, such as a personal computer. However, the program product of the present invention is not limited thereto. In this document, a readable storage medium can be any tangible medium containing or storing a program, which can be used by or in combination with an instruction execution system, an apparatus or a device.

[0122] The program product may use any combination of one or more readable media. The readable medium may be a readable signal medium or a readable storage medium. The readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, device or device, or any combination of the above. More specific examples of readable storage media (a non-exhaustive list) include: an electrical connection with one or more wires, a portable disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above.

[0123] Computer readable storage media may include data signals propagated in baseband or as part of a carrier wave, wherein readable program codes are carried. Such propagated data signals may take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. The readable storage medium may also be any readable medium other than a readable storage medium, which may send, propagate, or transmit a program for use by or in conjunction with an instruction execution system, device, or device. The program codes contained on the readable storage medium may be transmitted using any suitable medium, including but not limited to wireless, wired, optical cable, RF, etc., or any suitable combination thereof.

[0124] Program code for performing the operations of the present invention may be written in any combination of one or more programming languages, including object-oriented programming languages ​​such as Java, C++, etc., and conventional procedural programming languages ​​such as "C" or similar programming languages. The program code may be executed entirely on the user computing device, partially on the user device, as a separate software package, partially on the user computing device and partially on a remote computing device, or entirely on a remote computing device or server. In cases involving a remote computing device, the remote computing device may be connected to the user computing device through any type of grid, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computing device (e.g., via the Internet using an Internet service provider).

[0125] This embodiment is based on an RGBD camera to obtain a speckle infrared image and a depth map, identifies the range of transparent objects through a first deep learning network, and then corrects the depth value to obtain a final depth map. It can accurately identify transparent objects in various application scenarios and solve the problem of inaccurate recognition in the prior art.

[0126] In this specification, each embodiment is described in a progressive manner, and each embodiment focuses on the differences from other embodiments, and the same and similar parts between the embodiments can be referred to each other. The above description of the disclosed embodiments enables professionals and technicians in this field to implement or use the present invention. Various modifications to these embodiments will be obvious to professionals and technicians in this field, and the general principles defined herein can be implemented in other embodiments without departing from the spirit or scope of the present invention. Therefore, the present invention will not be limited to these embodiments shown in this article, but will comply with the widest range consistent with the principles and novel features disclosed herein.

[0127] The above describes the specific embodiments of the present invention. It should be understood that the present invention is not limited to the above specific embodiments, and those skilled in the art may make various modifications or variations within the scope of the claims, which do not affect the essence of the present invention.

Claims

1. A transparent object detection, segmentation and depth reconstruction method based on a 3D camera, characterized in that: include: Step S1: Obtaining a speckle infrared image and a depth image; Step S2: identifying the range of transparent objects through a first deep learning network; wherein the first deep learning network obtains the range of transparent objects according to the astigmatism characteristics of speckle irradiation on the transparent objects; Step S3: Correcting the depth value within the range of the transparent object according to the depth map and the speckle infrared map to obtain a final depth map.

2. The method for detecting, segmenting and reconstructing a transparent object based on a 3D camera according to claim 1, characterized in that: Step S1 includes: Step S11: acquiring two speckle infrared images; Step S12: Calculate the disparity according to the two speckle infrared images to obtain a depth map.

3. The method for detecting, segmenting and reconstructing a transparent object based on a 3D camera according to claim 1, characterized in that: In step S1, an RGB image is also acquired, and in step S2, at least the RGB image, the speckle infrared image and the depth image are input into the first deep learning network to identify the range of transparent objects.

4. The method for detecting, segmenting and reconstructing a transparent object based on a 3D camera according to claim 1, characterized in that: The first deep learning network is a Unet network, which includes an encoder, a first decoder and a second decoder; the first decoder outputs the range of the transparent object; the second decoder outputs the boundary of the transparent object; and there is feature interaction between features of each level of the first decoder and the second decoder.

5. The method for detecting, segmenting and reconstructing a transparent object based on a 3D camera according to claim 1, characterized in that: In step S3, the depth map, the transparent object range and the speckle infrared map are input into a second deep learning network to obtain a final depth map; wherein the second deep learning network corrects the depth value within the transparent object range.

6. The method for detecting, segmenting and reconstructing a transparent object based on a 3D camera according to claim 1, characterized in that: Step S3 includes: Step S31: performing three-dimensional reconstruction on the portion outside the range of the transparent object according to the depth map and the speckle infrared map; Step S32: obtaining a transparent object boundary and identifying the shape of the transparent object boundary; Step S33: performing plane fitting according to the boundary of the transparent object to obtain a depth value corresponding to each pixel point within the range of the transparent object.

7. The method for detecting, segmenting and reconstructing a transparent object based on a 3D camera according to claim 6, characterized in that: Step S33 includes: Step S331: performing straight line fitting on the boundary of the transparent object to obtain a plurality of straight lines in a three-dimensional space; Step S332: fitting the plane of three consecutive straight lines to obtain a transparent object; Step S333: Calculate the depth value corresponding to each pixel point within the range of the transparent object according to the transparent object.

8. A transparent object detection, segmentation and depth reconstruction system based on a 3D camera, used to implement the transparent object detection, segmentation and depth reconstruction method based on a 3D camera according to any one of claims 1 to 7, characterized in that: include: An acquisition module, used to acquire a speckle infrared image and a depth image; An identification module, configured to identify a range of transparent objects through a first deep learning network; wherein the first deep learning network obtains the range of transparent objects according to the astigmatism characteristics of speckle irradiation on the transparent objects; A correction module is used to correct the depth value within the range of the transparent object according to the depth map and the speckle infrared map to obtain a final depth map.

9. A transparent object detection, segmentation and depth reconstruction device based on a 3D camera, characterized in that: include: processor; a memory storing executable instructions of the processor; The processor is configured to execute the steps of the 3D camera-based transparent object detection, segmentation and depth reconstruction method as described in any one of claims 1 to 7 by executing the executable instructions.

10. A computer-readable storage medium for storing a program, characterized in that: When the program is executed, the steps of the transparent object detection, segmentation and depth reconstruction method based on a 3D camera described in any one of claims 1 to 7 are implemented.