Transparent object detection method, system and device based on infrared image and storage medium

By using infrared maps in transparent object detection, a multimodal data fusion method combined with dense and sparse depth maps, the problems of environmental impact, insufficient computing resources and high error detection rates in the prior art are solved, and accurate identification and real-time detection of the range and depth of transparent objects are achieved.

CN119942142APending Publication Date: 2025-05-06SHENZHEN GUANGJIAN TECH CO LTD +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510081566.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-20
Publication Date
2025-05-06

Smart Images

  • Figure CN119942142A_ABST
    Figure CN119942142A_ABST
Patent Text Reader

Abstract

The invention discloses a transparent object detection method, system and device based on an infrared image, and a storage medium. The method comprises the following steps: S1, obtaining a dense depth image, a sparse depth image and an infrared image; s2, inputting the dense depth image, the sparse depth image and the infrared image into a recognition model, and recognizing a transparent object range; wherein the transparent object range is a point corresponding to the dense depth map. According to the invention, the depth values of various transparent objects can be accurately identified.
Need to check novelty before this filing date? Find Prior Art

Description

Background Art

[0002] The infrared image-based transparent object detection method is an emerging technology that combines infrared light and computer vision technology to identify transparent objects in a scene. This technology is currently in the research and development stage. Although some progress has been made, it still faces many challenges and technical difficulties. The following is the current status of the technology in this field:

[0003] Technical principle

[0004] Speckle phenomenon: When laser or infrared light is irradiated onto a rough surface, the interference of light waves will form randomly distributed bright and dark spots on the observation plane. This phenomenon is called speckle. For transparent objects, light will pass through the object and undergo refraction, reflection, etc. inside it, causing the speckle pattern to change. These changes can be used to infer the existence and characteristics of transparent objects.

[0005] Depth perception: By using time-of-flight (ToF), structured light, or other types of depth sensors, the distance information from each pixel in the scene to the camera can be obtained, i.e., a dense depth map. However, for transparent objects, traditional depth sensing technology often cannot provide accurate measurement results because light passes through the object without being reflected back to the camera.

[0006] Data fusion: To overcome the limitations of a single sensor, researchers have begun to explore fusing different types of sensor data, such as combining infrared speckle images with depth maps to improve the detection of transparent objects.

[0007] Existing research results

[0008] Algorithm development: Some machine learning and deep learning-based algorithms have been proposed to extract features from speckle patterns and identify transparent objects. These algorithms usually require a large amount of training data and may include convolutional neural networks (CNNs) and other advanced image processing techniques.

[0009] Hardware innovation: Some new hardware designs have also been introduced, such as special speckle projectors and high-resolution infrared cameras, which can generate higher-quality speckle patterns and capture more details. In addition, depth sensors optimized for transparent objects are under development.

[0010] Challenges

[0011] Environmental influence: In practical applications, factors such as lighting conditions, background complexity, and the color and texture of surrounding objects will affect the quality of the speckle pattern and thus the detection accuracy of transparent objects. Therefore, how to improve the robustness and adaptability of the system is an important research direction.

[0012] Computing resources: Processing large amounts of image data in real time and performing complex mathematical operations requires powerful computing power, especially when implemented on mobile devices or embedded systems. To this end, researchers are looking for more efficient data processing methods and lightweight model architectures.

[0013] False detection rate: Due to the special properties of transparent objects, they are easily confused with other non-transparent objects, resulting in a high false detection rate. Reducing the false detection rate is the key to improving user experience.

[0014] Application prospects

[0015] Industrial Automation: In the manufacturing industry, transparent object detection technology can be applied to product quality control to ensure that products containing transparent materials such as glass and plastic meet the specifications. For example, the packaging line can automatically check whether the bottle is intact.

[0016] Intelligent Transportation: For self-driving cars, it is crucial to accurately perceive transparent obstacles around them (such as flooded roads, icy surfaces, etc.). Transparent object detection technology can help vehicles better understand road conditions, thereby improving driving safety.

[0017] Security monitoring: In terms of security protection in public places, transparent object detection technology can be used to monitor suspicious objects or behaviors, providing safer protection for people's lives.

[0018] Future development trends

[0019] Multimodal perception: Future research may further explore how to fuse more types of sensor data together to achieve more comprehensive and accurate environmental perception. This is not limited to optical sensors, but also includes acoustic, thermal imaging and other perception methods.

[0020] Adaptive learning: With the advancement of artificial intelligence technology, future transparent object detection systems are expected to have the ability of adaptive learning, which can automatically adjust parameters according to different application scenarios to achieve optimal performance.

[0021] Edge computing: To meet the requirements of real-time performance and low power consumption, researchers may focus on developing miniaturized neural networks suitable for running on edge devices, enabling transparent object detection technology to work efficiently in resource-constrained environments.

[0022] The disclosure of the above background technology content is only used to assist in understanding the inventive concept and technical solution of the present invention. It does not necessarily belong to the prior art of this patent application. In the absence of clear evidence that the above content has been disclosed on the filing date of this patent application, the above background technology should not be used to evaluate the novelty and creativity of the present application. Summary of the invention

[0023] To this end, the present invention provides a transparent object detection solution based on infrared images to solve the requirements for the range and depth of transparent objects in various application scenarios.

[0024] In a first aspect, the present invention provides a transparent object detection method based on infrared images, characterized by comprising:

[0025] Step S1: Obtain a dense depth map, a sparse depth map and an infrared map;

[0026] Step S2: Input the dense depth map, the sparse depth map and the infrared image into a recognition model to identify a transparent object range; wherein the transparent object range is the points corresponding to the dense depth map.

[0027] Optionally, the infrared image-based transparent object detection method is characterized in that step S1 comprises:

[0028] Step S11: Acquire LED infrared image and speckle infrared image;

[0029] Step S12: obtaining a dense depth map according to a time-of-flight algorithm; and obtaining a sparse depth map using the speckle infrared image according to a triangulation method.

[0030] Optionally, the infrared image-based transparent object detection method is characterized in that the dense depth map is obtained by using a time-of-flight algorithm or a triangulation method, and the sparse depth map is obtained by using a feature matching method.

[0031] Optionally, the infrared image-based transparent object detection method is characterized in that step S2 comprises:

[0032] Step S21: extracting features from the dense depth map, the sparse depth map, and the infrared map to obtain dense features, sparse features, and infrared features, and assigning weights to the dense depth map, the sparse depth map, and the infrared map, respectively;

[0033] Step S22: fusing the dense features, the sparse features and the infrared features according to the weights to obtain a fusion graph;

[0034] Step S23: using a recognition model to recognize the fusion image to obtain a transparent object range.

[0035] Optionally, the infrared image-based transparent object detection method is characterized in that the recognition model is a layered multi-scale convolutional neural network, and when the first-layer large-scale convolution kernel is processed, a multi-channel coarse-grained feature map is output; the middle-layer convolution layer receives the coarse-grained feature map, convolves and refines it while laterally connecting to extract key features from the first layer to supplement it, and outputs the middle-layer features; the bottom-layer small-scale convolution focuses on the bottom-layer details, merges them with the middle-layer features, and outputs the prediction results through the fully connected layer to determine the range of the corresponding points of the transparent object in the dense depth map.

[0036] Optionally, the infrared image-based transparent object detection method is characterized in that the recognition model also includes a self-supervised learning module during training; the self-supervised learning module runs synchronously with the recognition model, and the common features output by the self-supervised module are injected into the recognition model at intervals of a certain number of training steps; the recognition model is based on a back-propagation algorithm, combining the loss of labeled data with the feature consistency loss transmitted by the self-supervised module, and jointly updates the weights to continuously optimize the transparent object range recognition capability.

[0037] Optionally, the infrared image-based transparent object detection method further comprises:

[0038] Step S3: taking the transparent object range as a boundary, using the depth value of the non-transparent object to correct the depth value of the transparent object.

[0039] In a second aspect, the present invention provides a transparent object detection system based on infrared images, which is used to implement any of the transparent object detection methods based on infrared images described above, and is characterized by comprising:

[0040] An acquisition module, used to acquire dense depth maps, sparse depth maps and infrared images;

[0041] A recognition module is used to input the dense depth map, the sparse depth map and the infrared map into a recognition model to identify the range of transparent objects; wherein the range of transparent objects is the points corresponding to the dense depth map.

[0042] In a third aspect, the present invention provides a transparent object detection device based on infrared images, characterized in that it includes:

[0043] processor;

[0044] a memory storing executable instructions of the processor;

[0045] Wherein, the processor is configured to perform the steps of any of the aforementioned infrared image-based transparent object detection methods by executing the executable instructions.

[0046] In a fourth aspect, the present invention provides a computer-readable storage medium for storing a program, characterized in that when the program is executed, the steps of any of the aforementioned infrared image-based transparent object detection methods are implemented.

[0047] Compared with the prior art, the present invention has the following beneficial effects:

[0048] The present invention uses infrared bands for detection. Even when facing transparent objects, slight differences between their surface and interior will cause changes in the speckle pattern. By analyzing these speckle patterns, rich object information can be obtained. It has certain penetrating and anti-interference capabilities and high detection accuracy.

[0049] During the detection process, the present invention does not need to make direct contact with the transparent object, and at the same time, the detection speed is fast, thereby improving the detection efficiency.

[0050] The present invention can combine multi-dimensional information such as dense depth map and sparse depth map to describe and locate transparent objects from different angles. The depth map can provide the position and distance information of the object in space. Combined with the infrared speckle map, it can more accurately determine the range and posture of the transparent object, and realize the three-dimensional reconstruction and positioning of the transparent object. BRIEF DESCRIPTION OF THE DRAWINGS

[0051] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following will briefly introduce the drawings required for use in the embodiments or the prior art descriptions. Obviously, the drawings in the following descriptions are only embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on the provided drawings without creative work. By reading the detailed description of the non-limiting embodiments with reference to the following drawings, other features, purposes and advantages of the present invention will become more obvious:

[0052] Figure 1 This is a flow chart of the steps of a transparent object detection method based on infrared images in an embodiment of the present invention;

[0053] Figure 2 A schematic diagram of a network architecture in an embodiment of the present invention;

[0054] Figure 3 Schematic diagram of various images in an embodiment of the present invention;

[0055] Figure 4 This is a flowchart of the steps of obtaining a dense depth map, a sparse depth map and an infrared map in an embodiment of the present invention;

[0056] Figure 5 This is a flow chart of steps for identifying the range of a transparent object in an embodiment of the present invention;

[0057] Figure 6 is a flowchart of another method for detecting transparent objects based on infrared images in an embodiment of the present invention;

[0058] Figure 7 Schematic diagram of the structure of a transparent object detection system based on infrared images in an embodiment of the present invention;

[0059] Figure 8 is a schematic structural diagram of a transparent object detection device based on infrared images in an embodiment of the present invention; and

[0060] Fig. 9 Schematic diagram of the structure of a computer-readable storage medium in an embodiment of the present invention. DETAILED DESCRIPTION

[0061] The present invention is described in detail below in conjunction with specific embodiments. The following embodiments will help those skilled in the art to further understand the present invention, but are not intended to limit the present invention in any form. It should be noted that, for those of ordinary skill in the art, several variations and improvements may be made without departing from the concept of the present invention. These all belong to the protection scope of the present invention.

[0062] The terms "first", "second", "third", "fourth", etc. (if any) in the specification and claims of the present invention and the above-mentioned drawings are used to distinguish similar objects, and are not necessarily used to describe a specific order or sequence. It should be understood that the data used in this way can be interchanged where appropriate, so that the embodiments of the present invention described herein, for example, can be implemented in an order other than those illustrated or described herein. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusions, for example, a process, method, system, product or device that includes a series of steps or units is not necessarily limited to those steps or units that are clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.

[0063] An infrared image-based transparent object detection method provided in an embodiment of the present invention aims to solve the problems existing in the prior art.

[0064] The following specific embodiments are used to describe in detail the technical solutions of the present invention and how the technical solutions of the present application solve the above-mentioned technical problems. The following specific embodiments can be combined with each other, and the same or similar concepts or processes may not be repeated in some embodiments. The embodiments of the present invention will be described below in conjunction with the accompanying drawings.

[0065] The present invention utilizes two different depth maps, namely a dense depth map and a sparse depth map, in combination with an infrared image, and inputs a recognition model to obtain the range of transparent objects. It has a better recognition effect on transparent objects and can output the accurate range of transparent objects in real time, greatly improving the recognition effect on transparent objects.

[0066] Figure 1 FIG. 1 is a flow chart of the steps of a transparent object detection method based on infrared images in an embodiment of the present invention. Figure 1 As shown, the steps of a transparent object detection method based on infrared images in an embodiment of the present invention include:

[0067] Step S1: Obtain a dense depth map, a sparse depth map, and an infrared image.

[0068] In this step, two depth maps are obtained: a dense depth map and a sparse depth map, and an infrared map is also obtained.

[0069] Dense depth maps are designed to provide depth information for almost every pixel in the scene. Data can usually be collected using depth cameras, such as structured light cameras or time-of-flight (ToF) cameras. Structured light cameras project specific patterns (such as stripes, gray codes, etc.) into the scene, and the camera captures the pattern modulated by the scene objects, and calculates the depth value corresponding to each pixel based on the principle of triangulation. Time-of-flight cameras measure the flight time of light pulses from emission to reflection, and then convert the distance to generate a depth map. This method can quickly obtain high-resolution, wide-coverage dense depth information.

[0070] The collected raw data often contains noise and needs to be filtered, such as Gaussian filtering, to reduce the interference of random noise on the depth value. Distortion correction may also be required to compensate for the image distortion caused by the camera lens and ensure that the depth value accurately corresponds to the actual scene, ultimately obtaining a high-quality dense depth map.

[0071] There are many ways to obtain sparse depth maps. A common method is a stereo vision method based on feature point matching. Using a binocular camera system, the left and right cameras shoot the same scene from different perspectives. By extracting feature points in the image (such as SIFT, ORB and other feature points), the two images are matched, and then the depth of the matching feature points is calculated based on the triangulation principle. Since the number of feature points is small compared to the pixels of the entire image, a sparse depth map is formed. In addition, it can also be combined with LiDAR. The point cloud data obtained by the LiDAR scan is projected onto the image plane through coordinate transformation. The depth information corresponding to these discrete points constitutes a sparse depth map. Sparse depth maps also have to deal with noise problems, and in the feature point matching link, mismatches may occur. It is necessary to use algorithms such as RANSAC to eliminate mismatched point pairs and improve the accuracy of sparse depth maps.

[0072] Infrared image acquisition relies on infrared thermal imaging cameras, whose detectors can sense infrared radiation in the scene. Objects have different temperatures and emit different intensities of infrared light. Thermal imaging cameras convert these differences in infrared radiation intensity into differences in image grayscale values, thereby generating infrared images. In some application scenarios, although transparent objects are transparent to visible light, they still show thermal characteristics different from the surrounding environment in the infrared band, which provides a basis for subsequent detection. The original infrared image may have a low contrast, so enhancement methods such as histogram equalization are often used to stretch the grayscale range of the image and highlight the difference between the target object and the background, which is convenient for subsequent analysis.

[0073] Step S2: Input the dense depth map, the sparse depth map and the infrared image into a recognition model to identify a transparent object range; wherein the transparent object range is the points corresponding to the dense depth map.

[0074] In this step, the recognition model is designed to simultaneously receive three different modal data as inputs: dense depth map, sparse depth map, and infrared image. For example, a multimodal fusion convolutional neural network (CNN) architecture can be used to set independent input branches for each modal data. These branches extract the features of the modality in the shallow layer of the network, and then merge the features extracted by different branches through the feature fusion layer.

[0075] Before training, we collected a large amount of scene data containing transparent objects, and manually annotated the accurate range of transparent objects on the corresponding dense depth map, which was used as the true value label for training. These annotated data were divided into training sets, validation sets, and test sets for model training, parameter tuning, and performance evaluation.

[0076] During training, choose a suitable loss function, such as cross entropy loss (if it is a classification task to determine whether a pixel belongs to the range of a transparent object) or mean square error loss (if it is a regression task to predict the boundary of a transparent object, etc.), and use optimization algorithms such as stochastic gradient descent (SGD) and its variants Adagrad, Adam, etc. to repeatedly iterate the training model on the training set, so that the model's predicted output is continuously close to the annotated true value label, thereby improving the model's accuracy in identifying the range of transparent objects.

[0077] When the trained model receives the newly acquired dense depth map, sparse depth map and infrared image, it first extracts the features of the corresponding modality in each input branch, fuses these features, and then processes them through a series of convolution, pooling and fully connected layers to output a prediction result corresponding to the size of the dense depth map. This result is presented in the form of a probability map or directly a binary mask, marking which points are most likely to belong to the transparent object range. These points are exactly the points corresponding to the dense depth map, thereby accurately locating the position of the transparent object in the scene.

[0078] The preliminary prediction results may contain some tiny noise points or small holes, which require morphological processing, such as opening operations to remove isolated noise points and closing operations to fill small holes, so that the final output of the transparent object range is smoother and more accurate.

[0079] Figure 2 FIG. 1 is a schematic diagram of a network architecture in an embodiment of the present invention. Figure 2 As shown in the figure, the recognition model includes an encoder and a decoder. The dense depth map, sparse depth map and infrared image are input into the encoder for feature extraction. Then they are decoded step by step by the decoder to finally obtain the range of transparent objects. Figure 2 The image in the middle is a scene arrangement diagram, which is only used to illustrate the scene and does not participate in the processing in this embodiment. Figure 2 It can be seen that the transparent object range is not a strict range of transparent objects, but a depth range corresponding to the depth map, which is different from the range for RGB images in the prior art.

[0080] Figure 3 Schematic diagram of various images in the embodiments of the present invention. Figure 3 As shown, a is a scene layout diagram, which is only used to illustrate the scene situation and does not participate in the calculation and processing of this embodiment. b is an infrared image, c is a dense depth image, and d is a sparse depth image. b, c, and d are input into the recognition model for processing.

[0081] In some embodiments, the recognition model is a hierarchical multi-scale convolutional neural network. When the first-layer large-scale convolution kernel is processed, a multi-channel coarse-grained feature map is output; the middle-layer convolution layer receives the coarse-grained feature map, refines it through convolution, and extracts key features from the first layer through lateral connections to output middle-layer features; the bottom-layer small-scale convolution focuses on the bottom-layer details, merges them with the middle-layer features, and outputs the prediction results through the fully connected layer to determine the range of corresponding points of transparent objects in the dense depth map.

[0082] Design principle of the first-layer large-scale convolution kernel: In the first layer of the hierarchical multi-scale convolutional neural network, the use of large-scale convolution kernels is to quickly capture the macroscopic and coarse-grained information in the input image (i.e., the fusion image). The receptive field of large-scale convolution kernels is large. For example, using a 7×7 or 9×9 convolution kernel, a single convolution operation can cover a larger range of pixel areas in the image compared to a small-size convolution kernel. When it slides on the fusion image, it can cross multiple object boundaries and texture areas, and quickly aggregate information related to the overall outline of the transparent object and the large-scale spatial layout, without being limited by subtle textures or small noise interference.

[0083] After the convolution operation of the large-scale convolution kernel, a multi-channel coarse-grained feature map is output. Each channel corresponds to features of different dimensions. For example, one channel may focus on the overall difference in depth between a transparent object and the background, while another channel captures the approximate shape characteristics of the object's thermal radiation distribution. The multi-channel setting allows the network to preliminarily characterize the key information in the fusion map from multiple angles, laying the foundation for subsequent mid-level refinement processing.

[0084] Middle-layer convolution refinement: After receiving the coarse-grained feature map output by the first layer, the middle-layer convolution layer begins to gradually refine the features. Here, a series of smaller-sized convolution kernels, such as 3×3 convolution kernels, are used for multiple convolution operations. Each convolution can extract finer local features from the coarse-grained feature map, such as the precise curvature of the edge of a transparent object, the subtle slope of the depth gradient, etc., continuously improving the accuracy and recognition of the features.

[0085] At the same time, the middle layer also extracts key features from the first layer through lateral connections to supplement itself. Lateral connections can reintroduce some important large-scale features captured by the first layer but which may be weakened in subsequent convolutions into the middle layer. For example, the layout and positioning features of transparent objects in large scenes can prevent the middle layer from focusing too much on local details and losing the overall layout cognition, so that the middle-layer features output by the middle layer contain both accurate local details and macro positioning information.

[0086] The bottom-level small-scale convolution focuses on the bottom-level details: The bottom-level uses a small-scale convolution kernel, such as a 1×1 convolution kernel, to focus on mining the subtle details of the bottom layer of the image. After the middle-level processing, although there are relatively fine features, some tiny information that is critical to the precise positioning of transparent objects may still be missed. The small-scale convolution kernel can capture these ultra-fine depth changes and small differences in thermal radiation, supplementing the feature dimension.

[0087] Fusion with middle-layer features: The bottom-level detail features obtained by the bottom-level small-scale convolution are fused with the middle-level features output by the middle-level. This fusion can organically combine the microscopic details captured by the bottom-level with the comprehensive features of the middle-level through splicing, weighted summation, etc., making the feature representation more complete, including information related to transparent objects from macro to micro, from the whole to the details.

[0088] Function of the fully connected layer: The fused feature map enters the fully connected layer, which reduces the dimension and integrates the high-dimensional feature vector extracted by the previous convolutional layer. It aggregates and associates the information of all pixels in the feature map, and maps it to the final prediction result based on the weight parameters learned by the network. In this process, the fully connected layer plays a role in the global "understanding" and decision-making of the features.

[0089] Determine the range of transparent objects: The final prediction result will determine the range of transparent objects in the corresponding points of the dense depth map. The prediction result can be in the form of a probability map, that is, each pixel corresponds to a probability value belonging to a transparent object, and a suitable threshold is set to define the pixels above the threshold as the range of transparent objects; it can also directly output a binary mask map to clearly mark the range of transparent objects.

[0090] In some embodiments, the recognition model also includes a self-supervised learning module during training; the self-supervised learning module runs synchronously with the recognition model, and the common features output by the self-supervised module are injected into the recognition model at intervals of a certain number of training steps; the recognition model is based on a back propagation algorithm, combining the loss of labeled data with the feature consistency loss transmitted by the self-supervised module, and jointly updates the weights to continuously optimize the transparent object range recognition capability.

[0091] Throughout the training process, the self-supervised learning module works synchronously with the recognition model. Self-supervised learning aims to utilize the intrinsic structural information of the data itself, so that the model can learn some common features without a large amount of manual annotation of data. Every certain number of training steps, the self-supervised module will inject the extracted and processed common features into the recognition model. For example, it may mine common visual patterns in the data set based on the underlying attributes of the image, such as color, texture, and shape. These patterns are encoded into common features, which provide additional knowledge supplements for the recognition model.

[0092] Self-supervised learning modules often use contrastive learning and other methods to obtain common features. For example, it can perform data enhancement operations such as random cropping, rotation, and flipping on the input image to generate a series of image variants from different perspectives. Then, the model is asked to learn whether these variants come from the same original image. In this way, the model can learn the essential features of the image that are not affected by the transformation, that is, the common features. These features contain some common structures and texture laws of scene objects, which are helpful for the subsequent recognition of transparent objects.

[0093] Labeled data loss: The recognition model was originally trained based on manually labeled transparent object range data, using common loss functions such as cross entropy loss or mean square error loss to measure the deviation between the model prediction result and the labeled true value. For example, if the range of transparent objects predicted by the model is far from the actual labeled range, the label data loss will be very high, driving the model to adjust the weights and optimize in the direction of reducing the loss.

[0094] Feature consistency loss: After the universal features from the supervision module are injected, feature consistency loss will be introduced. This loss is used to measure the consistency of the recognition model's own feature representation before and after accepting the newly injected features. If the feature difference of the model output is too large before and after the injection of features, it means that the model has not been able to integrate the new features well. The feature consistency loss will prompt the model to adjust the weights so that it can absorb and use these universal features more smoothly and maintain the stability and coherence of the feature space.

[0095] Back propagation algorithm: With the help of the back propagation algorithm, the above two losses - label data loss and feature consistency loss are aggregated. Back propagation updates the weights layer by layer in the reverse direction of the network based on the gradient of the composite loss to the model weight. In other words, the model will not only change because the prediction results do not match the annotations, but also be optimized due to the failure to properly integrate the common features of self-supervised learning. This two-pronged approach continuously improves the ability to identify the range of transparent objects, allowing the model to more accurately lock on transparent objects when facing complex scenes and scenes with scarce data.

[0096] Figure 4 FIG. 1 is a flowchart of steps for obtaining a dense depth map, a sparse depth map, and an infrared map in an embodiment of the present invention. Figure 4 As shown, in an embodiment of the present invention, a step of obtaining a dense depth map, a sparse depth map and an infrared map includes:

[0097] Step S11: Obtaining LED infrared images and speckle infrared images.

[0098] In this step, the acquisition of LED infrared images relies on an imaging system equipped with an infrared LED light source. Infrared LEDs emit infrared light of a specific wavelength to illuminate the target scene. The camera lens captures the infrared light reflected from the scene and focuses it on the image sensor. The sensor converts the light signal into an electrical signal based on the difference in light intensity, and then further digitizes it into image data, ultimately forming an LED infrared image. Since different objects have different reflection characteristics for infrared rays, even in low light or complex lighting conditions, it is possible to capture information such as the outline and texture of the object, providing basic data for subsequent depth calculations.

[0099] The generation of speckle infrared images is based on speckle projection technology. First, a speckle projector projects randomly distributed speckle patterns composed of a large number of tiny bright and dark spots onto the target scene. An infrared camera is used to shoot the scene covered by speckles from a specific angle. When the speckles are irradiated on the surface of an object, the speckles will be deformed due to the different depth, shape and surface material of the object. The camera captures the deformed speckle image and obtains the speckle infrared image. This deformation of the speckle pattern contains the depth information of the object and is the key to the subsequent use of triangulation to infer the depth.

[0100] Step S12: obtaining a dense depth map according to a time-of-flight algorithm; and obtaining a sparse depth map using the speckle infrared image according to a triangulation method.

[0101] In this step, the time-of-flight (ToF) algorithm is based on the temporal characteristics of light propagation. Infrared light pulses emitted from an LED infrared light source are directed toward the scene, and the light pulses are reflected back to the camera after contacting the surface of the object. The camera records the emission and reception times of the light pulses, and by measuring the time difference between the two times (flight time), combined with the speed of light propagation in the air (approximately 299792458m / s), the distance from each point in the scene to the camera is calculated, thereby constructing a dense depth map. Each pixel can obtain the corresponding depth value, so the generated depth map has a high resolution and rich information.

[0102] In actual measurement, due to factors such as response delay of electronic components and noise interference, the measured flight time will have errors. It is necessary to use high-precision timing chips to reduce timing errors; filtering algorithms, such as median filtering, will also be used to filter out abnormal time values ​​caused by noise, ensure the accuracy of depth calculation, and ultimately output high-quality dense depth maps.

[0103] The triangulation method uses the geometric relationship of similar triangles to determine the depth of an object. For speckle infrared images, the relative position and posture of the speckle projector and the camera are known, forming a fixed triangulation system. When the speckle is projected to a point on the surface of the object, the camera captures the deformed speckle image, and by identifying the feature points of the speckle, the original projected speckle and the deformed speckle are compared to find the corresponding relationship. Based on the principle of similar triangles, the depth of the feature point on the surface of the object can be calculated if the baseline length (the distance between the projector and the camera) and the angle information are known. Since the depth is calculated only for the feature points of the speckle, not all pixels, a sparse depth map is obtained.

[0104] Improving the accuracy of speckle feature point extraction and matching is crucial. Using high-precision feature point extraction algorithms, such as feature extraction networks based on deep learning, can more accurately capture subtle differences in deformed speckles; at the same time, accurately calibrating the relative position and posture parameters between the speckle projector and the camera can reduce the depth calculation deviation caused by inaccurate system parameters and optimize the quality of sparse depth maps.

[0105] In some embodiments, the dense depth map is obtained by using a time-of-flight algorithm or a triangulation method, and the sparse depth map is obtained by using a feature matching method. The time-of-flight algorithm can be used to obtain the depth value of each pixel, and the triangulation method can be used to obtain the depth value of each scattered speckle. When the density of scattered speckles is high, a dense depth map can be obtained. The feature matching method can use key feature points to calculate the depth value, has high reliability, and is a sparse depth map.

[0106] Figure 5 FIG. 1 is a flow chart of steps for identifying the range of a transparent object in an embodiment of the present invention. Figure 5 As shown, in an embodiment of the present invention, a step of identifying a range of a transparent object includes:

[0107] Step S21: extracting features from the dense depth map, the sparse depth map and the infrared map respectively to obtain dense features, sparse features and infrared features, and assigning weights according to the dense depth map, the sparse depth map and the infrared map respectively.

[0108] In this step, the features of dense depth map, sparse depth map and infrared image are extracted respectively.

[0109] Convolutional neural networks (CNNs) are often used to mine features for dense depth maps. The convolution layer of a CNN uses convolution kernels of different sizes to slide on the depth map to capture the depth change pattern of the local neighborhood, such as edges, gradients, and other information. This information will be combined into a more abstract and advanced feature representation as the number of network layers increases. The pooling layer is responsible for reducing the dimensionality of the features, reducing the amount of data while retaining key features. After stacking multiple convolution and pooling layers, the dense features corresponding to the dense depth map are finally output.

[0110] Since the data distribution of sparse depth maps is relatively discrete, the sparse depth maps are first interpolated to make their data distribution more regular and convenient for subsequent calculations. The CNN architecture is also used, but in the design of convolution kernels, more emphasis may be placed on larger convolution kernels so that multiple sparse points are covered at one time and features that reflect the spatial relationship between sparse points are extracted. After convolution and pooling operations, the sparse features of the sparse depth map are obtained.

[0111] For infrared images, the CNN-based method is also very effective. Considering that infrared images reflect the difference in thermal radiation of objects, the extracted features include information such as texture and grayscale gradient corresponding to the temperature boundaries of different objects. In the shallow layer of the network, the convolution kernel focuses on extracting small-scale texture features, while the deep layer integrates features related to the overall thermal contour of the object and outputs infrared features.

[0112] Weights are assigned based on information reliability. Dense depth maps often contain more comprehensive and detailed depth information, so they are usually given higher weights. Although sparse depth maps have sparse data, they can provide accurate depth at key feature points and are also given a certain weight. Infrared images provide additional information on the thermal characteristics of objects, and the weights are determined based on the specific scenario. If the thermal characteristics of transparent objects are highly distinguishable, the weight will increase accordingly, and vice versa.

[0113] Step S22: fusing the dense features, the sparse features and the infrared features according to the weights to obtain a fusion graph.

[0114] In this step, the weighted fusion method can adopt a pixel-by-pixel weighting method or a feature dimension weighting method.

[0115] Pixel-by-pixel weighting: The corresponding pixels of the dense feature, sparse feature, and infrared feature maps are weighted and summed according to the pre-set weights. Assume that the value of a pixel in the dense feature map is a, the sparse feature map is b, and the infrared feature map is c. The corresponding weights are w1, w2, and w3 respectively. The fused pixel value d = w1a + w2b + w3c. By performing this operation on all pixels, a fused image is generated.

[0116] Feature dimension weighting: The three extracted features are concatenated in the feature dimension to form a high-dimensional feature vector, and then the weight matrix is ​​multiplied by the high-dimensional feature vector to achieve fusion. This method can better preserve the high-order relationship between features, and is particularly suitable for deep learning frameworks. It can be directly connected to the fully connected layer for further processing.

[0117] Step S23: using a recognition model to recognize the fusion image to obtain a transparent object range.

[0118] In this step, the recognition model can be a convolutional neural network (CNN) or a fully convolutional network (FCN).

[0119] Convolutional Neural Network (CNN): You can use classic CNN architectures, such as FasterR-CNN, MaskR-CNN, etc. The fusion image is input into the network, and the convolution layer of the network continuously extracts features from the fusion image. The region proposal network (RPN, for FasterR-CNN) screens out candidate regions that may contain transparent objects, and then through subsequent classification and regression layers, accurately determine the category (if there is a classification requirement) and boundary range of the transparent object, and output the transparent object range.

[0120] Fully Convolutional Network (FCN): FCN directly takes the fusion image as input, uses convolution operations throughout the process, and outputs a prediction result with the same size as the fusion image. Each pixel corresponds to a prediction probability, which represents the possibility that the pixel belongs to a transparent object. By setting a threshold, pixels above the threshold are marked as transparent objects.

[0121] The training and optimization process is described as follows:

[0122] Data preparation: A large amount of fusion graph data with transparent object annotations is collected and divided into training set, validation set and test set. The annotation information accurately outlines the range of transparent objects in the graph and serves as a supervisory signal for training.

[0123] Training process: Select a suitable loss function, such as a combination of cross entropy loss and dice loss, and use an optimization algorithm, such as the Adam optimizer, to repeatedly train the model on the training set, and adjust the model parameters according to the performance of the validation set until the model achieves satisfactory recognition accuracy and recall on the test set.

[0124] Figure 6 FIG. 1 is a flowchart of another method for detecting transparent objects based on infrared images according to an embodiment of the present invention. Figure 6 As shown, compared with the above-mentioned embodiment, another transparent object detection method based on infrared image in the embodiment of the present invention further includes:

[0125] Step S3: taking the transparent object range as a boundary, using the depth value of the non-transparent object to correct the depth value of the transparent object.

[0126] In this step, after obtaining the range of transparent objects, the depth values ​​corresponding to non-transparent objects and transparent objects are first extracted from the dense depth map. The depth value of non-transparent objects is relatively accurate because the reflection and refraction characteristics of non-transparent objects to light are relatively stable, and the depth measurement technology can more accurately measure its distance information. However, due to the optical properties of transparent objects, such as light transmittance and refraction effects, the depth values ​​obtained by conventional depth measurement methods often have large deviations.

[0127] Correction can be performed based on domain relationships and statistical methods.

[0128] Based on neighborhood relationship: Take the boundary of the transparent object as a clue and observe the depth value of the surrounding non-transparent objects. Assuming that a certain boundary area of ​​the transparent object is adjacent to a non-transparent object, since the depth is usually continuous and gradual in the same local scene, the depth value of the adjacent non-transparent object can be referred to, and a more reasonable depth value at the boundary of the transparent object can be inferred through interpolation algorithms such as linear interpolation and bilinear interpolation. Then gradually extend from the boundary to the inside of the transparent object to correct the depth value of the entire transparent object.

[0129] Statistical method: Collect the depth values ​​of multiple non-transparent objects around the transparent object and calculate their statistics, such as mean, median, etc. Take these statistics as a reference, combine the geometric factors such as the shape and size of the transparent object itself, and adjust the depth value inside the transparent object according to certain rules. For example, if the transparent object has a regular shape and is small, the mean of the depth values ​​of the surrounding non-transparent objects can be directly used as its corrected depth value; if the object is large and complex in shape, it needs to be corrected by using statistics in different sections and situations.

[0130] Before performing the correction, the extracted depth value data is filtered to remove noise interference. Because noise may cause abnormal depth values, if not processed, the correction results will deviate from the actual ones. Gaussian filtering, median filtering and other methods are used to ensure data quality.

[0131] Calibration is not a one-time process, and the depth value of transparent objects after the first calibration may not be accurate enough. The calibration process can be repeated multiple times, and each time the depth relationship with the surrounding non-transparent objects is re-evaluated based on the newly calibrated depth value, and the calibration strategy is fine-tuned until the depth value of the transparent object reaches a relatively stable and reasonable state, making it as consistent as possible with the real scene depth.

[0132] Figure 7 FIG. 1 is a schematic diagram of a transparent object detection system based on infrared images according to an embodiment of the present invention. Figure 7 As shown, in an embodiment of the present invention, a transparent object detection system based on infrared images includes:

[0133] An acquisition module, used to acquire dense depth maps, sparse depth maps and infrared images;

[0134] A recognition module is used to input the dense depth map, the sparse depth map and the infrared map into a recognition model to identify the range of transparent objects; wherein the range of transparent objects is the points corresponding to the dense depth map.

[0135] This embodiment uses two different depth maps, namely a dense depth map and a sparse depth map, combined with an infrared image, and inputs a recognition model to obtain the range of transparent objects. It has a better recognition effect on transparent objects and can output the accurate range of transparent objects in real time, greatly improving the recognition effect on transparent objects.

[0136] The present invention also provides a transparent object detection device based on infrared images, including a processor and a memory in which executable instructions of the processor are stored. The processor is configured to execute the steps of a transparent object detection method based on infrared images by executing the executable instructions.

[0137] As mentioned above, this embodiment uses two different depth maps, namely a dense depth map and a sparse depth map, combined with an infrared map, and inputs a recognition model to obtain the range of transparent objects. It has a better recognition effect on transparent objects and can output the accurate range of transparent objects in real time, greatly improving the recognition effect on transparent objects.

[0138] It will be appreciated by those skilled in the art that various aspects of the present invention may be implemented as systems, methods or program products. Therefore, various aspects of the present invention may be specifically implemented in the following forms, namely: complete hardware implementation, complete software implementation (including firmware, microcode, etc.), or a combination of hardware and software implementations, which may be collectively referred to herein as "circuits", "modules" or "platforms".

[0139] Figure 8 Schematic diagram of the structure of a transparent object detection device based on infrared image in an embodiment of the present invention. Figure 8 The electronic device 600 according to this embodiment of the present invention is described. Figure 8 The electronic device 600 shown is only an example and should not bring any limitation to the functions and scope of use of the embodiments of the present invention.

[0140] like Figure 8 As shown, the electronic device 600 is in the form of a general computing device. The components of the electronic device 600 may include, but are not limited to: at least one processing unit 610, at least one storage unit 620, a bus 630 connecting different platform components (including the storage unit 620 and the processing unit 610), a display unit 640, etc.

[0141] The storage unit stores program codes, which can be executed by the processing unit 610, so that the processing unit 610 executes the steps of various exemplary embodiments of the present invention described in the above-mentioned transparent object detection method based on infrared images of this specification. For example, the processing unit 610 can execute the following steps: Figure 1 Follow the steps shown in .

[0142] The storage unit 620 may include a readable medium in the form of a volatile storage unit, such as a random access memory unit (RAM) 6201 and / or a cache memory unit 6202 , and may further include a read-only memory unit (ROM) 6203 .

[0143] The storage unit 620 may also include a program / utility 6204 having a set (at least one) of program modules 6205, such program modules 6205 including but not limited to: an operating system, one or more application programs, other program modules, and program data, each of which or some combination may include an implementation of a grid environment.

[0144] Bus 630 may represent one or more of several types of bus structures, including a memory unit bus or memory unit controller, a peripheral bus, an accelerated graphics port, a processing unit, or a local bus using any of a variety of bus architectures.

[0145] The electronic device 600 may also communicate with one or more external devices 700 (e.g., keyboards, pointing devices, Bluetooth devices, etc.), one or more devices that enable a user to interact with the electronic device 600, and / or any device that enables the electronic device 600 to communicate with one or more other computing devices (e.g., routers, modems, etc.). Such communication may be performed through an input / output (I / O) interface 650. Furthermore, the electronic device 600 may also communicate with one or more grids (e.g., a local area network (LAN), a wide area network (WAN), and / or a public grid, such as the Internet) through a grid adapter 660. The grid adapter 660 may communicate with other modules of the electronic device 600 through the bus 630. It should be understood that although Figure 8 Not shown, other hardware and / or software modules may be used in conjunction with electronic device 600, including but not limited to: microcode, device drivers, redundant processing units, external disk drive arrays, RAID systems, tape drives, and data backup storage platforms.

[0146] In an embodiment of the present invention, a computer-readable storage medium is also provided for storing a program, and when the program is executed, the steps of a transparent object detection method based on infrared images are implemented. In some possible implementations, various aspects of the present invention can also be implemented in the form of a program product, which includes a program code. When the program product is run on a terminal device, the program code is used to enable the terminal device to execute the steps of various exemplary embodiments of the present invention described in the above-mentioned transparent object detection method based on infrared images section of this specification.

[0147] As shown above, this embodiment uses two different depth maps, namely a dense depth map and a sparse depth map, combined with an infrared map, and inputs a recognition model to obtain the range of transparent objects. It has a better recognition effect on transparent objects and can output the accurate range of transparent objects in real time, greatly improving the recognition effect on transparent objects.

[0148] Fig. 9 Schematic diagram of the structure of a computer-readable storage medium in an embodiment of the present invention. Fig. 9 As shown, a program product 800 for implementing the above method according to an embodiment of the present invention is described, which can adopt a portable compact disk read-only memory (CD-ROM) and include program code, and can be run on a terminal device, such as a personal computer. However, the program product of the present invention is not limited thereto. In this document, a readable storage medium can be any tangible medium containing or storing a program, which can be used by or in combination with an instruction execution system, an apparatus or a device.

[0149] The program product may use any combination of one or more readable media. The readable medium may be a readable signal medium or a readable storage medium. The readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, device or device, or any combination of the above. More specific examples of readable storage media (a non-exhaustive list) include: an electrical connection with one or more wires, a portable disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above.

[0150] Computer readable storage media may include data signals propagated in baseband or as part of a carrier wave, wherein readable program codes are carried. Such propagated data signals may take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. The readable storage medium may also be any readable medium other than a readable storage medium, which may send, propagate, or transmit a program for use by or in conjunction with an instruction execution system, device, or device. The program codes contained on the readable storage medium may be transmitted using any suitable medium, including but not limited to wireless, wired, optical cable, RF, etc., or any suitable combination thereof.

[0151] Program code for performing the operations of the present invention may be written in any combination of one or more programming languages, including object-oriented programming languages ​​such as Java, C++, etc., and conventional procedural programming languages ​​such as "C" or similar programming languages. The program code may be executed entirely on the user computing device, partially on the user device, as a separate software package, partially on the user computing device and partially on a remote computing device, or entirely on a remote computing device or server. In cases involving a remote computing device, the remote computing device may be connected to the user computing device through any type of grid, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computing device (e.g., via the Internet using an Internet service provider).

[0152] This embodiment uses two different depth maps, namely a dense depth map and a sparse depth map, combined with an infrared image, and inputs a recognition model to obtain the range of transparent objects. It has a better recognition effect on transparent objects and can output the accurate range of transparent objects in real time, greatly improving the recognition effect on transparent objects.

[0153] In this specification, each embodiment is described in a progressive manner, and each embodiment focuses on the differences from other embodiments, and the same and similar parts between the embodiments can be referred to each other. The above description of the disclosed embodiments enables professionals and technicians in this field to implement or use the present invention. Various modifications to these embodiments will be obvious to professionals and technicians in this field, and the general principles defined herein can be implemented in other embodiments without departing from the spirit or scope of the present invention. Therefore, the present invention will not be limited to these embodiments shown in this article, but will comply with the widest range consistent with the principles and novel features disclosed herein.

[0154] The above describes the specific embodiments of the present invention. It should be understood that the present invention is not limited to the above specific embodiments, and those skilled in the art may make various modifications or variations within the scope of the claims, which do not affect the essence of the present invention.

Claims

1. A transparent object detection method based on infrared images, characterized in that: include: Step S1: Obtain a dense depth map, a sparse depth map and an infrared image; Step S2: Input the dense depth map, the sparse depth map and the infrared image into a recognition model to identify a transparent object range; wherein the transparent object range is the points corresponding to the dense depth map.

2. The method for detecting transparent objects based on infrared images according to claim 1, characterized in that: Step S1 includes: Step S11: Acquire LED infrared image and speckle infrared image; Step S12: obtaining a dense depth map according to a time-of-flight algorithm; and obtaining a sparse depth map using the speckle infrared image according to a triangulation method.

3. The method for detecting transparent objects based on infrared images according to claim 1, characterized in that: The dense depth map is obtained by using a time-of-flight algorithm or a triangulation method, and the sparse depth map is obtained by using a feature matching method.

4. The method for detecting transparent objects based on infrared images according to claim 1, characterized in that: Step S2 includes: Step S21: extracting features from the dense depth map, the sparse depth map, and the infrared map to obtain dense features, sparse features, and infrared features, and assigning weights to the dense depth map, the sparse depth map, and the infrared map, respectively; Step S22: fusing the dense features, the sparse features and the infrared features according to the weights to obtain a fusion graph; Step S23: using a recognition model to recognize the fusion image to obtain a transparent object range.

5. The method for detecting transparent objects based on infrared images according to claim 1, characterized in that: The recognition model is a hierarchical multi-scale convolutional neural network. When the first-layer large-scale convolution kernel is processed, a multi-channel coarse-grained feature map is output; the middle-layer convolution layer receives the coarse-grained feature map, refines it through convolution, and extracts key features from the first layer through lateral connections to output middle-layer features; the bottom-layer small-scale convolution focuses on the bottom-layer details, merges them with the middle-layer features, and outputs the prediction results through the fully connected layer to determine the range of the corresponding points of the transparent object in the dense depth map.

6. The method for detecting transparent objects based on infrared images according to claim 1, characterized in that: During training, the recognition model also includes a self-supervised learning module; the self-supervised learning module runs synchronously with the recognition model, and the common features output by the self-supervised module are injected into the recognition model at intervals of a certain number of training steps; the recognition model is based on a back-propagation algorithm, combining the loss of labeled data with the loss of feature consistency transmitted by the self-supervised module, jointly updating weights, and continuously optimizing the transparent object range recognition capability.

7. The method for detecting transparent objects based on infrared images according to claim 1, characterized in that: Also includes: Step S3: taking the transparent object range as a boundary, using the depth value of the non-transparent object to correct the depth value of the transparent object.

8. A transparent object detection system based on infrared images, used to implement the transparent object detection method based on infrared images as claimed in any one of claims 1 to 7, characterized in that: include: An acquisition module, used to acquire dense depth maps, sparse depth maps and infrared images; A recognition module is used to input the dense depth map, the sparse depth map and the infrared map into a recognition model to identify the range of transparent objects; wherein the range of transparent objects is the points corresponding to the dense depth map.

9. A transparent object detection device based on infrared images, characterized in that: include: processor; a memory storing executable instructions of the processor; Wherein, the processor is configured to perform the steps of the infrared image-based transparent object detection method described in any one of claims 1 to 7 by executing the executable instructions.

10. A computer-readable storage medium for storing a program, characterized in that: When the program is executed, the steps of the infrared image-based transparent object detection method described in any one of claims 1 to 7 are implemented.