Intelligent device for identifying transparent object

By combining RGB cameras and speckle projection receiving systems on smart devices, using timing image differences to identify transparent objects and perform depth corrections, the problem of inaccurate recognition of transparent objects in the prior art is solved, and high-precision recognition and deep reconstruction on mobile devices are realized.

CN119984091APending Publication Date: 2025-05-13SHENZHEN GUANGJIAN TECH CO LTD +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510143091.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-10
Publication Date
2025-05-13

AI Technical Summary

Technical Problem

The prior art is difficult to accurately identify and divide transparent objects, especially in mobile device applications, where there is a problem of inaccurate identification.

Method used

Using a time-based transparent object detection segmentation and depth correction scheme, the RGB camera, speckle projector and speckle receiver are used to identify the range and boundaries of transparent objects by acquiring image sets of different distances, and performing depth correction.

Benefits of technology

It realizes accurate identification and deep reconstruction of transparent objects, which is suitable for mobile devices in the detection of fast moving objects, improving the stability and accuracy of recognition.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119984091A_ABST
    Figure CN119984091A_ABST
Patent Text Reader

Abstract

An intelligent device for identifying a transparent object is characterized by comprising an RGB camera used for obtaining an RGB image; the speckle projector is used for projecting infrared speckles; the speckle receiver is used for receiving reflection signals of the infrared speckles; the moving part is used for moving the intelligent equipment; the controller is used for controlling the RGB camera, the speckle projector and the speckle receiver to work synchronously and controlling the moving part to move so as to obtain a first image set and a second image set at different positions; identifying a transparent object range and a transparent object boundary according to the difference between the first image set and the second image set; each of the first image set and the second image set comprises an RGB image and a speckle infrared image. The method can be used for detecting the transparent object, especially for detecting the transparent object by a moving intelligent device.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of intelligent devices, and in particular to an intelligent device for identifying transparent objects. Background Art

[0002] In recent years, with the development of automation and intelligent technology, the demand for transparent object recognition and corresponding depth reconstruction has become increasingly high, such as in the fields of sweeping robot obstacle avoidance and mapping, autonomous driving, and industrial inspection.

[0003] However, most existing object detection technologies are mainly used for opaque objects, and it is difficult to detect and segment transparent objects. There are also some transparent object detection technologies based on traditional methods or deep learning methods, but most of them are based on a single RGB image for transparent object segmentation, which has great limitations in effect stability and usage scenarios.

[0004] Unlike traditional cameras, 3D cameras can collect more environmental information during actual use, including RGB images, depth images and different types of infrared images. Transparent objects have richer and more obvious features in these images, especially depth and infrared images, compared with single RGB images. In this way, a more stable and accurate range can be obtained through the algorithm. After obtaining the range, the surrounding environment information and the depth characteristics of the transparent object itself can be used to reconstruct its depth.

[0005] In the measurement of transparent objects, there is an outstanding demand for mobile smart devices. Mobile smart devices need to perform obstacle avoidance and navigation, which all require the recognition of transparent objects.

[0006] The disclosure of the above background technology content is only used to assist in understanding the inventive concept and technical solution of the present invention. It does not necessarily belong to the prior art of this patent application. In the absence of clear evidence that the above content has been disclosed on the filing date of this patent application, the above background technology should not be used to evaluate the novelty and creativity of the present application. Summary of the invention

[0007] To this end, the present invention provides a time-series-based transparent object detection, segmentation and depth correction solution to address the requirements for the range and depth of transparent objects in various application scenarios.

[0008] The present invention provides an intelligent device for identifying transparent objects, characterized in that it includes:

[0009] RGB camera, used to obtain RGB images;

[0010] A speckle projector, used for projecting infrared speckles;

[0011] A speckle receiver, used for receiving a reflection signal of the infrared speckle;

[0012] Mobile components, used for mobile smart devices;

[0013] A controller is used to control the RGB camera, the speckle projector, and the speckle receiver to work synchronously, and control the movement of the moving component to obtain a first image set and a second image set at different positions; according to the difference between the first image set and the second image set, the range and boundary of the transparent object are identified; the first image set and the second image set both include an RGB image and a speckle infrared image.

[0014] Optionally, the intelligent device for identifying transparent objects is characterized in that there are two speckle receivers for calculating a depth map.

[0015] Optionally, the intelligent device for identifying transparent objects is characterized in that, when the controller obtains the first image set and the second image set, it includes:

[0016] Step M1: acquiring a first RGB image, a first speckle infrared image, and a second speckle infrared image at a first distance, and calculating a first depth map according to the first speckle infrared image and the second speckle infrared image;

[0017] Step M2: acquiring a second RGB image, a third speckle infrared image, and a fourth speckle infrared image at a second distance, and calculating a second depth map according to the third speckle infrared image and the fourth speckle infrared image.

[0018] Optionally, the intelligent device for identifying transparent objects is characterized in that the RGB image and the speckle infrared image are reprojected and aligned at the pixel level.

[0019] Optionally, the intelligent device for identifying transparent objects is characterized in that the RGB camera is located on the optical axis of the two speckle receivers so that the RGB image and the depth map are aligned.

[0020] Optionally, the intelligent device for identifying transparent objects is characterized in that the controller, when processing the first image set and the second image set, comprises:

[0021] Step S1: performing feature matching on the first image set and the second image set to obtain an area representing the same target object;

[0022] Step S2: inputting the first image set and the second image set into a comparison model, and identifying the transparent object range and the transparent object boundary according to the difference between the same target object area in the first image set and the second image set;

[0023] Step S3: Correcting the depth value within the range of the transparent object according to the depth value of the boundary of the transparent object to obtain a final depth map of the transparent object.

[0024] Optionally, the intelligent device for identifying transparent objects is characterized in that, in step S2, a preliminary transparent object area is identified based on the difference between the speckle infrared images in the first image set and the second image set, and the final transparent object range and transparent object boundary are identified based on feature matching based on the RGB images in the first image set and the second image set.

[0025] Optionally, the intelligent device for identifying transparent objects is characterized in that the comparison model is a Unet network, comprising an encoder, a first decoder and a second decoder; the first decoder outputs the range of the transparent object; the second decoder outputs the boundary of the transparent object; and there is feature interaction between features of each level of the first decoder and the second decoder.

[0026] Optionally, the intelligent device for identifying transparent objects is characterized in that step S3 comprises:

[0027] Step S31: performing three-dimensional reconstruction on the portion outside the range of the transparent object according to the depth map and the RGB map;

[0028] Step S32: obtaining three-dimensional information of the boundary of the transparent object, and identifying the shape of the boundary of the transparent object in three-dimensional space;

[0029] Step S33: performing plane fitting according to the boundary of the transparent object to obtain a depth value corresponding to each pixel point within the range of the transparent object.

[0030] Optionally, the intelligent device for identifying transparent objects is characterized in that step S33 comprises:

[0031] Step S331: performing straight line fitting on the boundary of the transparent object to obtain a plurality of straight lines in a three-dimensional space;

[0032] Step S332: fitting the plane of three consecutive straight lines to obtain a transparent object;

[0033] Step S333: Calculate the depth value corresponding to each pixel point within the range of the transparent object according to the transparent object.

[0034] Compared with the prior art, the present invention has the following beneficial effects:

[0035] The present invention utilizes the mobility of smart devices and the light-eliminating effect of transparent objects, and utilizes the differences of the same object in images obtained at different distances to identify transparent objects. The present invention can be applied to the detection of transparent objects by moving objects, and is particularly suitable for the detection of transparent objects by fast-moving objects. The present invention can accurately identify the range of transparent objects through deep learning of light spots and depth values, and can effectively correct the depth values ​​of transparent objects. The present invention is suitable for various transparent object recognition scenarios. BRIEF DESCRIPTION OF THE DRAWINGS

[0036] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following will briefly introduce the drawings required for use in the embodiments or the prior art descriptions. Obviously, the drawings in the following descriptions are only embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on the provided drawings without creative work. By reading the detailed description of the non-limiting embodiments with reference to the following drawings, other features, purposes and advantages of the present invention will become more obvious:

[0037] Figure 1 This is a schematic diagram of the structure of an intelligent device for identifying transparent objects in an embodiment of the present invention;

[0038] Figure 2 A flowchart of steps for obtaining a first image set and a second image set in an embodiment of the present invention;

[0039] Figure 3 is a flowchart of steps for processing a first image set and a second image set in an embodiment of the present invention;

[0040] Figure 4 A schematic diagram of a network architecture in an embodiment of the present invention;

[0041] Figure 5 Schematic diagram of various images in an embodiment of the present invention;

[0042] Figure 6 is a flow chart of steps for obtaining a final depth map in an embodiment of the present invention;

[0043] Figure 7 The following is a flowchart of the steps for calculating the depth value corresponding to each pixel in an embodiment of the present invention. DETAILED DESCRIPTION

[0044] The present invention is described in detail below in conjunction with specific embodiments. The following embodiments will help those skilled in the art to further understand the present invention, but are not intended to limit the present invention in any form. It should be noted that, for those of ordinary skill in the art, several variations and improvements may be made without departing from the concept of the present invention. These all belong to the protection scope of the present invention.

[0045] The terms "first", "second", "third", "fourth", etc. (if any) in the specification and claims of the present invention and the above-mentioned drawings are used to distinguish similar objects, and are not necessarily used to describe a specific order or sequence. It should be understood that the data used in this way can be interchanged where appropriate, so that the embodiments of the present invention described herein, for example, can be implemented in an order other than those illustrated or described herein. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusions, for example, a process, method, system, product or device that includes a series of steps or units is not necessarily limited to those steps or units that are clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.

[0046] An embodiment of the present invention provides an intelligent device for identifying transparent objects, aiming to solve the problems existing in the prior art.

[0047] The following specific embodiments are used to describe in detail the technical solutions of the present invention and how the technical solutions of the present application solve the above-mentioned technical problems. The following specific embodiments can be combined with each other, and the same or similar concepts or processes may not be repeated in some embodiments. The embodiments of the present invention will be described below in conjunction with the accompanying drawings.

[0048] The present invention utilizes the light-eliminating effects of smart devices, mobility, and transparent objects, and utilizes the differences of the same object in images obtained at different distances to identify transparent objects. The present invention can be applied to transparent object detection of moving objects, and is particularly suitable for transparent object detection of fast-moving smart devices, thereby solving the problem of inaccurate recognition in the prior art.

[0049] Figure 1 FIG. 1 is a schematic diagram of the structure of an intelligent device for identifying transparent objects in an embodiment of the present invention. Figure 1 As shown, in an embodiment of the present invention, an intelligent device for identifying transparent objects includes:

[0050] RGB camera, used to obtain RGB images.

[0051] Specifically, the RGB camera captures images based on the color model of the three primary colors of red, green, and blue. It contains a photosensitive element inside. When light enters the lens, the pixels on the photosensitive element will record the corresponding RGB values ​​according to the intensity of different colors of light, and finally combine them into a colorful RGB image. For example, when shooting a red apple, the surface of the apple reflects more red light, and the R value of the corresponding pixel is higher, thereby accurately presenting the color information of the object's appearance. In the process of identifying transparent objects, the RGB image provides basic appearance color data for subsequent processing, and assists in judging the color characteristics of the surrounding environment of transparent objects, which is of reference significance for locating transparent objects and subsequent boundary outlining.

[0052] Speckle projector, used to project infrared speckles.

[0053] Specifically, the speckle projector can have EEL, Vcsel or other infrared lasers, which can project random speckle or coded structured light speckle. When the speckle is irradiated on the surface of an object, it will be reflected and scattered due to the surface characteristics of the object. For example, the final shape of the speckle will be significantly different when projected on a transparent object such as flat and smooth glass and when projected on a rough and opaque object, especially for the brightness and shape of the speckle returned at a longer distance. The speckle projector actively emits infrared speckle, creating conditions for the subsequent use of speckle reflection information to identify transparent objects, and providing additional depth and structure information to the entire recognition system.

[0054] The speckle receiver is used to receive the reflected signal of the infrared speckle.

[0055] Specifically, the speckle receiver is a light sensor that is sensitive to the infrared band and can receive infrared speckle signals reflected from the surface of an object. After receiving the reflected light, the optical signal is converted into an electrical signal, which is processed by the internal circuit and finally outputs a digital signal reflecting the speckle information, such as the brightness distribution and shape change of the speckle. The speckle receiver collects the reflected signal of the infrared speckle and cooperates with the speckle projector to provide key data for determining the three-dimensional structure, surface undulations and other key features of transparent objects. These data are extremely important for accurately locating the boundaries of transparent objects.

[0056] Mobile components for mobile smart devices.

[0057] Specifically, the moving part can be a mechanical structure composed of a motor, a transmission device (such as a belt, a gear, etc.), and a wheel. After the controller issues a movement command, the motor drives the transmission device to drive the wheels of the smart device to move in a specified direction, thereby changing the shooting position of the device, so that the smart device can acquire images at different spatial positions.

[0058] A controller is used to control the RGB camera, the speckle projector, and the speckle receiver to work synchronously, and control the movement of the moving component to obtain a first image set and a second image set at different positions; according to the difference between the first image set and the second image set, the range and boundary of the transparent object are identified; the first image set and the second image set both include an RGB image and a speckle infrared image.

[0059] Specifically, the controller, as the "brain" of the device, has built-in control programs and algorithm logic. On the one hand, it can send synchronous trigger signals to the RGB camera, speckle projector, and speckle receiver to ensure that the three start working at the same time and obtain the image and speckle information at the same time; on the other hand, it directs the movement of the moving parts according to the preset movement trajectory and parameters. Moreover, after acquiring multiple sets of image data, a special image analysis algorithm is run to compare the first image set and the second image set, and the range and boundaries of transparent objects are accurately located based on the differences in the RGB image color and speckle infrared image structure information in the two sets of data.

[0060] In some embodiments, there are two speckle receivers for calculating the depth map. The two speckle receivers are arranged at a certain position interval in space. They each independently receive the infrared speckle signal reflected from the surface of the object and convert the optical signal into an electrical signal. Due to the different positions of the two, the observed speckle pattern will be slightly different due to parallax. From the principle of optical geometry, based on the triangulation method, the position of the corresponding point of the speckle in space can be inferred by using the difference in speckle patterns received by the two receivers, combined with known parameters such as the receiver spacing. By capturing the speckle reflection signal with parallax, basic data is provided for the subsequent calculation of the depth map. The depth map can present the distance information of different objects in the scene from the device, which is particularly critical for the recognition of transparent objects, because the depth change of the transparent object and the front and back layer relationship with the background are the essential basis for accurately outlining its boundaries and determining its range.

[0061] In some embodiments, the RGB image and the speckle infrared image are reprojected and aligned at the pixel level. The pixel-level reprojection alignment of the RGB image and the speckle infrared image is to accurately match the two different types of images in spatial position to ensure the accuracy of subsequent processing and analysis. Since the RGB camera and the speckle projection-receiving system have different imaging principles and coordinate system settings. The RGB image is based on visible light imaging and usually follows the conventional two-dimensional image coordinate system; while the speckle infrared image is captured and processed by the speckle receiver using the reflection of infrared light, and its coordinate system is related to the imaging optical path and the layout of the speckle receiver. When the two are initially collected, even if the same object is photographed, the actual spatial positions represented by the corresponding pixels in their respective images are not completely consistent. The core of the pixel-level reprojection alignment is to "project" the pixel coordinates of the speckle infrared image to the corresponding correct position of the RGB image through mathematical transformation according to the optical imaging model and geometric relationship. This process needs to consider the internal parameters of the device, such as the focal length of the camera, the position of the optical center, as well as many factors such as the speckle projection angle and the position of the receiver, to construct a complex projection matrix. Specifically, it includes camera calibration, speckle system calibration, feature extraction, feature matching, calculation of transformation matrix, reprojection transformation, etc.

[0062] In some embodiments, the RGB camera is located on the optical axis of the two speckle receivers so that the RGB image and the depth map are aligned. The optical axis is the central axis of light propagation in an optical system. For a speckle receiver, the optical axis determines the main path direction of the reflected speckle signal received by it. When the RGB camera is located on the optical axis of the two speckle receivers, it means that the light reflected from the target object, whether it is captured by the RGB camera in the visible light band or by the speckle receiver in the infrared band, has a high degree of overlap in the spatial propagation path. From a geometric point of view, the three observation "lines of sight" are almost parallel, reducing the image deviation caused by the difference in observation angles. The imaging planes of the camera and the receiver are easier to establish an association under this layout. Under the ideal optical model, along the optical axis direction, the imaging planes of different sensors can be matched in a more regular way. That is to say, for the same spatial point, its projection position on the imaging plane of the RGB camera and its corresponding position in the depth map generated by the speckle receiver theoretically only have a simple linear translation or scaling relationship, rather than complex irregular transformations such as distortion and rotation. Therefore, coordinate transformation can be simplified, alignment errors can be reduced, and hardware collaborative optimization can be achieved.

[0063] Figure 2 FIG. 1 is a flow chart of steps for obtaining a first image set and a second image set in an embodiment of the present invention. Figure 2 As shown, in an embodiment of the present invention, a step of acquiring a first image set and a second image set includes:

[0064] Step M1: acquiring a first RGB image, a first speckle infrared image, and a second speckle infrared image of a first distance, and calculating a first depth map according to the first speckle infrared image and the second speckle infrared image.

[0065] In this step, there are two speckle receivers. When the device starts working, the distance between the smart device and the target scene is the first distance. The RGB camera captures the first RGB image of the scene at this moment. At the same time, the speckle projector projects infrared speckles, and the two speckle receivers receive the reflected speckle signals respectively to form the first speckle infrared image and the second speckle infrared image. The first RGB image here records the color distribution of the scene at this distance, while the two speckle infrared images carry the preliminary information of the infrared speckle reflection of the object surface.

[0066] Since the two speckle receivers are in different positions, based on the principle of triangulation, the controller uses the first speckle infrared image and the second speckle infrared image they receive to infer the depth information of each point in the scene to obtain a first depth map. For example, the larger the parallax of the speckle pattern on the two receivers, the closer the corresponding object point is to the device; the smaller the parallax, the farther the distance. Through a series of complex geometric calculations and data processing algorithms, the first depth map is finally generated, which intuitively shows the distance between different objects in the scene and the smart device at the first distance.

[0067] Step M2: acquiring a second RGB image, a third speckle infrared image, and a fourth speckle infrared image at a second distance, and calculating a second depth map according to the third speckle infrared image and the fourth speckle infrared image.

[0068] In this step, after completing the image acquisition at the first distance, the controller drives the moving component to move the smart device to a new position, that is, a position at a second distance from the target scene. The RGB camera shoots again to obtain the second RGB image at this position, the speckle projector continues to project speckles, and the two speckle receivers capture the reflected signals to obtain the third speckle infrared image and the fourth speckle infrared image, in preparation for the subsequent depth calculation of the new position.

[0069] Consistent with the depth map calculation principle in step M1, the controller uses the third speckle infrared image and the fourth speckle infrared image collected by the two speckle receivers to calculate the second depth map corresponding to the scene at the second distance based on the triangulation method, taking into account key factors such as the distance between the two receivers and the speckle pattern parallax, thereby further enriching the depth data at different positions of the scene and providing a multi-dimensional reference for the subsequent accurate identification of the range and boundaries of transparent objects.

[0070] Figure 3 FIG. 1 is a flowchart of the steps of processing the first image set and the second image set in an embodiment of the present invention. Figure 3As shown, in an embodiment of the present invention, a step of processing a first image set and a second image set includes:

[0071] Step S1: performing feature matching on the first image set and the second image set to obtain regions representing the same target object.

[0072] In this step, the controller extracts features from the first image set and the second image set respectively. For the RGB image, algorithms such as SIFT (scale-invariant feature transform), SURF (accelerated robust features), and ORB (orientation-fast rotation brief) can be used to identify significant feature points such as corner points and edge contours; for speckle infrared images, given their unique speckle texture, a texture-based feature extraction method is used to capture key features such as texture changes and density of the speckle pattern, and convert them into feature descriptors that can be calculated.

[0073] Matching algorithms, such as nearest neighbor matching and bidirectional matching, are used to pair the features extracted from the two sets of images. For example, the distance between feature descriptors in different images is calculated (commonly used Euclidean distance), and the feature points with the closest distance and meeting certain threshold conditions are considered as corresponding points of the same actual object. Multiple pairs of matching points are thus found, and the area enclosed by these matching points is likely to represent the same target object.

[0074] Step S2: Input the first image set and the second image set into a comparison model, and identify the transparent object range and the transparent object boundary according to the difference between the same target object area in the first image set and the second image set.

[0075] In this step, the comparison model can be a convolutional neural network architecture based on deep learning, or a traditional image difference algorithm. If a deep learning model is used, a large number of image sets with transparent objects and non-transparent objects labeled will be used for training in the early stage, so that the model can learn the characteristic change rules of transparent object areas between different images; the traditional difference algorithm directly calculates the difference of pixel values ​​in corresponding areas.

[0076] The first and second image sets that match the same target object area are input into the model. In the deep learning model, after multiple layers of convolution, pooling, full connection and other operations, the prediction results about the range and boundary of the transparent object are output; under the traditional algorithm, the pixel difference is analyzed. If the RGB value and speckle pattern changes of a certain area between the two sets of images meet the optical characteristics of transparent objects (such as small color changes and special speckle deformation), then the area is determined to be a transparent object, and its boundary range is gradually outlined.

[0077] Step S3: Correcting the depth value within the range of the transparent object according to the depth value of the boundary of the transparent object to obtain a final depth map of the transparent object.

[0078] In this step, the depth value of each point on the boundary of the transparent object is accurately extracted from the depth map calculated previously. These boundary depth values ​​are relatively accurate and are the key reference for correcting the depth of the internal area. Based on the known boundary depth values, interpolation algorithms such as bilinear interpolation and spline interpolation are used to adjust the depth values ​​within the range of the transparent object. Due to the special optical properties of the transparent object itself, the internal depth changes may be relatively smooth. The use of boundary value interpolation can correct local depth anomalies caused by light refraction and reflection interference, and finally generate a depth map that is more in line with the real spatial form of the transparent object, providing accurate depth information for the subsequent accurate identification and positioning of transparent objects.

[0079] In some embodiments, in step S2, a preliminary transparent object region is identified based on the difference between the speckle infrared images in the first image set and the second image set, and the final transparent object range and transparent object boundary are identified based on feature matching based on the RGB images in the first image set and the second image set. The speckle infrared image reflects the situation where the infrared speckle is reflected back after being projected onto the surface of the object. When light encounters a transparent object, its reflection and refraction characteristics are significantly different from those of an opaque object. The surface of a transparent object will cause the speckle to produce a unique deformation, distortion, and intensity change pattern. For example, when light penetrates a transparent object, part of the speckle will deviate from the original propagation path due to refraction, so that the reflected speckle pattern will change regularly in shape, density, and brightness distribution compared to when it was projected. Compared with floodlight, speckle has a high light intensity density and a long range, so the image of speckle is more sensitive to transparent objects and can be quickly identified in the speckle infrared image. The controller will perform pixel-level comparison on the speckle infrared images in the first image set and the second image set. Calculate the gray value difference, texture feature change and other parameters of the corresponding position pixels, set a certain threshold, and once the difference exceeds the threshold and shows a law that conforms to the optical characteristics of transparent objects, mark the area where these pixels are located as a preliminary transparent object area. This method can quickly screen out the approximate range where transparent objects may exist and reduce the amount of subsequent calculations.

[0080] After determining the preliminary area, we start focusing on the RGB image. Using feature extraction algorithms, such as SIFT and SURF, we extract feature points in the corresponding preliminary area of ​​the RGB images of the first and second image sets. These feature points cover the key positions of corners and edges, and they carry key visual information such as the appearance shape and color transition of the object.

[0081] Using feature matching algorithms, such as nearest neighbor matching, the feature points in the preliminary area of ​​the two sets of RGB images are paired to find the feature point pairs corresponding to the same actual object. Based on the successfully matched feature point pairs, the positional relationship between them is calculated through geometric algorithms, and then the real range and boundaries of the transparent object are accurately outlined. The RGB image provides rich color and detail information, so the range and boundaries determined in this way can best fit the actual appearance of the transparent object, make up for the recognition errors that may occur only by relying on the speckle infrared image, and improve the accuracy of the final recognition result.

[0082] In some embodiments, Figure 4 As shown, the comparison model is a Unet network, which includes an encoder, a first decoder and a second decoder; the first decoder outputs the range of the transparent object; the second decoder outputs the boundary of the transparent object; and there is feature interaction between the features of each level of the first decoder and the second decoder.

[0083] Unet is a convolutional neural network with a classic encoder-decoder architecture, which is particularly suitable for image segmentation tasks. Its encoder part is used to gradually extract high-level semantic features of the image, which can capture the complex patterns in the speckle infrared image, and the decoder is responsible for converting these abstract features back to feature maps with a size close to the original image to accurately locate the target area.

[0084] The encoder is usually composed of multiple convolutional modules stacked in sequence, each of which contains a convolutional layer, an activation function (such as ReLU), and a pooling layer. As the number of network layers increases, the convolution kernel continuously extracts more and more abstract features in the speckle infrared image, and the pooling layer gradually reduces the size of the feature map while increasing the receptive field, allowing the network to focus on a wider range of image context information, making it easier to identify the overall astigmatism characteristics associated with transparent objects.

[0085] Corresponding to the encoder, the first decoder gradually upsamples the high-level semantic features transmitted by the encoder to restore the size of the feature map. For example, it uses transposed convolution or up-pooling to gradually move the feature map from abstract to concrete. It finally outputs the range of transparent objects, which means that the features learned by the network can accurately identify which pixels belong to transparent objects after a series of decoding transformations, providing key area information for subsequent processing.

[0086] The second decoder also receives features from the encoder and performs an upsampling process similar to the first decoder. However, it focuses on outputting the boundaries of transparent objects. During the learning process, it focuses on capturing the key features that distinguish transparent objects from the surrounding environment. These features are decoded to outline clear boundary lines, separating transparent objects from the background and other objects.

[0087] The first decoder and the second decoder interact with each other at each level, which is a key design to improve performance. High-level features are rich in semantic information, and low-level features retain more details. The interaction between the two allows the decoder to take into account both semantic understanding and detail restoration. For example, at a certain level, the feature map of the first decoder is spliced ​​with the feature map of the corresponding level of the second decoder, and then subjected to convolution operation. The newly generated feature map combines the advantages of both sides, making the output transparent object range more accurate and the boundaries sharper, reducing misjudgments and unclear segmentation.

[0088] Figure 5 In the figure, a is the speckle infrared image, b is the depth map, c is the transparent object range, and d is the final depth map. As can be seen from a, the speckle infrared image includes multiple scattered spots, which can effectively detect transparent objects. b is the calculated depth map. Due to the existence of transparent objects, the boundaries in the depth map are blurred and the depth values ​​are disordered. The transparent object range c obtained by the first deep learning model can accurately obtain the transparent object range. The final depth map obtained in d can accurately express the depth values ​​including transparent objects in the target scene.

[0089] Figure 6 FIG. 1 is a flowchart of steps for obtaining a final depth map in an embodiment of the present invention. Figure 6 As shown, in an embodiment of the present invention, a step of obtaining a final depth map includes:

[0090] Step S31: performing three-dimensional reconstruction on the portion outside the range of the transparent object according to the depth map and the RGB map.

[0091] In this step, the depth map records the distance between each object in the scene and the smart device, and the grayscale value of each pixel corresponds to a certain depth value. The RGB image contains rich color and texture details, and can accurately present the appearance characteristics of the object. The combination of the two can supplement key visual and spatial information for three-dimensional reconstruction. For example, the depth map clarifies the front and back position relationship of the object, while the RGB image can help distinguish objects of different materials and colors, allowing the reconstruction algorithm to better identify the outline of the object. During the calibration stage of the smart device or its camera, the RGB image and the depth map are aligned, and three-dimensional reconstruction can be performed directly.

[0092] Step S32: Obtain three-dimensional information of the boundary of the transparent object, and identify the shape of the boundary of the transparent object in three-dimensional space.

[0093] In this step, the depth values ​​corresponding to the boundary pixels are accurately extracted from the partially reconstructed depth map and the previously identified transparent object boundaries. Due to the special optical properties of transparent objects, the depth information at the boundary is particularly critical, and it is the key anchor point for subsequently determining the complete three-dimensional shape of the transparent object. With the help of the extracted boundary depth data, some geometric shape recognition algorithms are used. For example, for transparent objects with regular shapes, by fitting geometric models such as quadratic surfaces and cylindrical surfaces, it is determined whether the boundary meets specific shape characteristics; for irregular shapes, point cloud-based geometric feature analysis is used, such as calculating curvature changes and normal vector distribution, to outline the real curvature, undulations and other shape characteristics of the boundary in three-dimensional space.

[0094] Step S33: performing plane fitting according to the boundary of the transparent object to obtain a depth value corresponding to each pixel point within the range of the transparent object.

[0095] In this step, after obtaining the three-dimensional shape of the boundary of the transparent object, it is assumed that the depth change inside the transparent object is relatively gentle, and a plane fitting technique is used. The depth value of the boundary point is used as a constraint condition, and a mathematical optimization algorithm such as the least square method is used to determine the plane equation that best fits the internal area of ​​the transparent object. For example, if the transparent object is similar to flat glass, its internal depth value should roughly satisfy a plane relationship.

[0096] Substitute the coordinates of each pixel point within the range of the transparent object into the fitted plane equation and calculate the corresponding depth value, so as to correct the inaccurate depth data in the previous depth map that may be affected by the refraction and reflection characteristics of the transparent object, and generate a depth map that more accurately reflects the actual depth of the transparent object.

[0097] This embodiment fits the transparent object according to the boundary of the transparent object in three-dimensional space, so that the transparent object can be restored more accurately; it can realize accurate segmentation, depth value correction and three-dimensional reconstruction of the transparent object, which not only improves the overall accuracy but also ensures the integrity of the reconstruction result; through fitting, this embodiment can calculate the depth value corresponding to each pixel point within the range of the transparent object, taking into account the special properties of the transparent object (such as light refraction and reflection), thereby improving the accuracy of the depth value.

[0098] Figure 7 FIG. 1 is a flowchart of a step of calculating the depth value corresponding to each pixel in an embodiment of the present invention. Figure 7 As shown, in an embodiment of the present invention, a step of calculating a depth value corresponding to each pixel point includes:

[0099] Step S331: performing straight line fitting on the boundary of the transparent object to obtain a plurality of straight lines in three-dimensional space.

[0100] In this step, after obtaining the three-dimensional information of the boundary of the transparent object, since the boundary is often composed of a series of discrete points, in order to simplify subsequent processing, these points need to be fitted into straight lines. The boundary point cloud data carries rich spatial position information. Based on the principle of least squares, the least squares method aims to find a straight line so that the sum of the squares of the distances from all boundary points to this straight line is minimized, so as to determine the optimal fitting straight line. In three-dimensional space, for point clusters with relatively linear distribution on the boundary, repeating this operation can obtain multiple fitting straight lines, which can abstractly summarize the local linear features of the boundary. For example, for a transparent object in the shape of a rectangular parallelepiped, the boundary points corresponding to its edge parts can be accurately extracted after straight line fitting to represent each edge. Even if the object is slightly deformed or interfered by noise, a suitable fitting algorithm can grasp the main linear trend and prepare for subsequent plane fitting.

[0101] Step S332: Fit the plane to three consecutively connected straight lines to obtain a transparent object.

[0102] In this step, after having multiple fitting lines, it is considered that a plane can be determined by three points that are not on the same straight line, and three connected straight lines can provide sufficient points to determine a plane. Select three straight lines that are connected in sequence, and their intersecting or adjacent endpoints together constitute the point set required to determine the plane. Using these points, the least squares method is used again to calculate a plane equation so that the sum of the squares of the distances from these points to the plane is minimized, thereby fitting a plane that fits the local surface of the transparent object. Repeat this "three-in-a-group" plane fitting operation for the straight lines on the boundary of the entire transparent object. When multiple fitting planes are spliced ​​together, the surface of the entire transparent object can be approximately represented. This method is particularly suitable for transparent objects with polyhedral shapes. Fitting the plane with connected straight lines can fit the original geometric shape of the object well and reduce the fitting error caused by complex shapes.

[0103] Step S333: Calculate the depth value corresponding to each pixel point within the range of the transparent object according to the transparent object.

[0104] In this step, after completing a series of plane fittings and obtaining an approximate surface model of the transparent object, each pixel within the range of the transparent object is theoretically located in the space enclosed by these fitting planes. Knowing the equation of the fitting plane, substituting the coordinates of the pixel into the corresponding plane equation, the depth value of the pixel can be quickly calculated. Compared with the initial depth map, the calculation result of this depth value takes into account the geometric shape characteristics of the transparent object itself, corrects the depth deviation that may be caused by optical phenomena such as refraction and reflection, and makes the final transparent object depth information more accurate and reliable.

[0105] This embodiment can obtain the three-dimensional shape of a transparent object by connecting three straight-line fitting planes. This method takes into account the spatial continuity of the transparent object and can reconstruct the three-dimensional structure of the object more accurately. Even if the transparent object has a complex shape or surface features, a relatively accurate three-dimensional reconstruction result can be obtained by connecting three straight-line fitting planes in sequence. After obtaining the three-dimensional shape of the transparent object, the depth value corresponding to each pixel point within the range of the transparent object can be calculated based on the shape. This depth value calculation method takes into account the three-dimensional structure of the object and is therefore more accurate.

[0106] In this specification, each embodiment is described in a progressive manner, and each embodiment focuses on the differences from other embodiments, and the same and similar parts between the embodiments can be referred to each other. The above description of the disclosed embodiments enables professionals and technicians in this field to implement or use the present invention. Various modifications to these embodiments will be obvious to professionals and technicians in this field, and the general principles defined herein can be implemented in other embodiments without departing from the spirit or scope of the present invention. Therefore, the present invention will not be limited to these embodiments shown in this article, but will comply with the widest range consistent with the principles and novel features disclosed herein.

[0107] The above describes the specific embodiments of the present invention. It should be understood that the present invention is not limited to the above specific embodiments, and those skilled in the art may make various modifications or variations within the scope of the claims, which do not affect the essence of the present invention.

Claims

1. An intelligent device for identifying transparent objects, characterized in that: include: RGB camera, used to obtain RGB images; A speckle projector, used for projecting infrared speckles; A speckle receiver, used for receiving a reflection signal of the infrared speckle; Mobile components, used for mobile smart devices; A controller is used to control the RGB camera, the speckle projector, and the speckle receiver to work synchronously, and control the movement of the moving component to obtain a first image set and a second image set at different positions; according to the difference between the first image set and the second image set, the range of the transparent object and the boundary of the transparent object are identified; the first image set and the second image set both include an RGB image and a speckle infrared image.

2. The intelligent device for identifying transparent objects according to claim 1, characterized in that: There are two speckle receivers, which are used to calculate the depth map.

3. The intelligent device for identifying transparent objects according to claim 2, characterized in that: The controller, when obtaining the first image set and the second image set, comprises: Step M1: acquiring a first RGB image, a first speckle infrared image, and a second speckle infrared image at a first distance, and calculating a first depth map according to the first speckle infrared image and the second speckle infrared image; Step M2: acquiring a second RGB image, a third speckle infrared image, and a fourth speckle infrared image at a second distance, and calculating a second depth map according to the third speckle infrared image and the fourth speckle infrared image.

4. The intelligent device for identifying transparent objects according to claim 2, characterized in that: The RGB image and the speckle infrared image are aligned by pixel-level reprojection.

5. The intelligent device for identifying transparent objects according to claim 2, characterized in that: The RGB camera is located on the optical axes of the two speckle receivers so that the RGB image and the depth image are aligned.

6. The intelligent device for identifying transparent objects according to claim 1, characterized in that: The controller, when processing the first image set and the second image set, includes: Step S1: performing feature matching on the first image set and the second image set to obtain an area representing the same target object; Step S2: inputting the first image set and the second image set into a comparison model, and identifying the transparent object range and the transparent object boundary according to the difference between the first image set and the second image set in the same target object area; Step S3: Correcting the depth value within the range of the transparent object according to the depth value of the boundary of the transparent object to obtain a final depth map of the transparent object.

7. The intelligent device for identifying transparent objects according to claim 6, characterized in that: In step S2, a preliminary transparent object region is identified based on the difference between the speckle infrared images in the first image set and the second image set, and a final transparent object range and transparent object boundary are identified based on feature matching based on the RGB images in the first image set and the second image set.

8. The intelligent device for identifying transparent objects according to claim 6, characterized in that: The comparison model is a Unet network, which includes an encoder, a first decoder and a second decoder; the first decoder outputs the range of the transparent object; the second decoder outputs the boundary of the transparent object; and there is feature interaction between features of each level of the first decoder and the second decoder.

9. The intelligent device for identifying transparent objects according to claim 6, characterized in that: Step S3 includes: Step S31: performing three-dimensional reconstruction on the portion outside the range of the transparent object according to the depth map and the RGB map; Step S32: obtaining three-dimensional information of the boundary of the transparent object, and identifying the shape of the boundary of the transparent object in three-dimensional space; Step S33: performing plane fitting according to the boundary of the transparent object to obtain a depth value corresponding to each pixel point within the range of the transparent object.

10. The intelligent device for identifying transparent objects according to claim 9, characterized in that: Step S33 includes: Step S331: performing straight line fitting on the boundary of the transparent object to obtain a plurality of straight lines in a three-dimensional space; Step S332: fitting the plane of three consecutive straight lines to obtain a transparent object; Step S333: Calculate the depth value corresponding to each pixel point within the range of the transparent object according to the transparent object.