A Transparent Object Detection Method Based on Polarization Imaging and Deep Learning
Through methods based on polarization imaging and deep learning, characterization images of transparent targets are constructed and feature fusion and supervision training are combined with FPN networks, which solves the problem of inaccurate transparent target detection in the prior art and achieves higher precision transparent target detection and positioning.
Patent Information
- Application Number
- CN202211118323.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-09-11
- Publication Date
- 2025-06-24
- Estimated Expiration
- 2042-09-11
AI Technical Summary
The existing technology is difficult to accurately detect transparent targets, resulting in inaccurate identification and positioning problems in applications such as intelligent manufacturing and intelligent driving.
Using transparent object detection methods based on polarization imaging and deep learning, we train the deep learning network detection model, use polarized image data to construct S0, DoLP, IE and L1 characterization images, and combine the FPN network for feature fusion and supervision training to achieve accurate detection of transparent objects.
It improves the accuracy of transparent boundaries and the positioning accuracy of transparent objects, enhances the details and accuracy of detection results, and is suitable for applications in fields such as intelligent manufacturing.
Smart Images

Figure CN115424115B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of transparent target detection, and particularly relates to a transparent target detection method based on polarization imaging and deep learning. Background Art
[0002] Transparent targets widely exist in the real world, such as many items made of plastic and glass like transparent plastic bottles and transparent cups. Therefore, the detection of transparent items has extremely important practical significance and application value in fields such as intelligent manufacturing and intelligent driving. For example, on the production line of transparent items, how to detect the accurate position of transparent items to help the manipulator capture the items; how intelligent driving identifies transparent objects and avoids them. Therefore, it is crucial to correctly locate each transparent target from the input image. However, compared with traditional objects, the images of transparent targets lack relevant information such as color, texture, shape, and edges that traditional target recognition and detection rely on extremely. There are few existing detection methods for transparent targets. Due to the inaccurate recognition and positioning of transparent targets, these transparent targets will have a key impact on existing machine vision systems, further affecting decision-making and costs in many applications. In 2003, Osadchy et al. used specular highlights as a positive information source for identifying luminous objects, but this process requires a bright light source. TransCut proposed an energy function based on LF-Linear and occlusion detection to optimize the segmentation result from 4D light field images. TOM-Net defined the extinction of transparent objects as a refraction estimation problem, and proposed a multi-scale encoder-decoder network to generate a rough input, and then used a residual network to refine it into a detailed rough input. It should be noted that TOM-Net requires a refraction flow map as a label during the training process, which is difficult to obtain from the real world, so only synthetic training data can be relied on. Therefore, it is very necessary to introduce transparent target detection into the field of polarization imaging.
[0003] At the same time, deep learning is to learn the internal laws and representation levels of sample data. The information obtained during these learning processes is very helpful for the interpretation of data such as text, images, and sounds. Its ultimate goal is to enable the machine to have the ability of analysis and learning like a human, and be able to identify data such as text, images, and sounds. Deep learning is a complex machine learning algorithm, and the effects achieved in speech and image recognition far exceed those of previous related technologies. Therefore, inventing a transparent target detection method based on polarization imaging and deep learning has extremely important practical significance and application value in fields such as intelligent manufacturing. Summary of the Invention
[0004] In order to solve the above-mentioned existing technical problems, the present invention designs a transparent target detection method based on polarization imaging and deep learning.
[0005] To solve the above-mentioned existing technical problems, the present invention adopts the following solutions:
[0006] A transparent target detection method based on polarization imaging and deep learning, comprising the following steps:
[0007] Step 1, training a deep learning network detection model;
[0008] Step 11, building a transparent target polarization image data acquisition system;
[0009] Step 12, using the acquisition system to collect polarization image data and constructing a training data set;
[0010] Step 13, demosaicking each original polarization image data in the training data set to obtain corresponding linearly polarized images I0(x,y), I 45 (x,y), I 90 (x,y) and I 135 (x,y);
[0011] Step 14, according to the four-angle linearly polarized images I0(x,y), I 45 (x,y), I 90 (x,y) and I 135 (x,y) obtained from each original polarization image data in Step 13, constructing a set of S0, DoLP, I E and L1 images corresponding to each original polarization image data according to the polarization principle using the Stokes vector, linear polarization angle, and degree of linear polarization formulas. S0 is the total detection intensity, usually the sum of two mutually orthogonal linearly polarized images, and the formula is as follows:
[0012] S0 = I0(x,y) + I 90 (x,y) = I 45 (x,y) + I 135 (x,y) #(1);
[0013] The concept of the degree of linear polarization DoLP is proposed based on the differences in the internal structure, appearance, material, and roughness of different objects, reflecting the differences in the radiation characteristics of the polarization states of objects, and the formula is as follows:
[0014]
[0015] S1 is a parameter calculated based on the Stokes vector parameters for two polarized images I0(x,y), I 90 (x,y) of the same original polarization image data, representing the difference between the light intensity in the horizontal direction and the light intensity in the vertical direction; S2 is based on the Stokes vector parameters for two polarized images I 45(x, y), I 135 The parameter obtained by calculating (x, y), representing the difference between the light intensity in the 45° direction and the light intensity in the 135° direction;
[0016] I E Represents the enhanced linearly polarized image, and the formula is as follows:
[0017]
[0018] AoLP represents the linearly polarized angle image, that is, the polarization azimuth angle, which can well reflect the attribute information such as illumination, shadow, surface flatness, stress birefringence, etc. in the light field, and more contains the edge information of transparent objects. The formula is as follows:
[0019]
[0020] To solve the aliasing problem of 0 and π in AoLP, we proposed L1. Because the division operation in the linearly polarized angle formula makes it very sensitive to noise, when S0 and S1 are close to 0, calculation errors will occur. The aliasing problem of 0 and π in AoLP is caused by the inability to distinguish 0 and π in the arctan operation. In the new representation, mainly to solve the aliasing problem of 0 and π, that is, to remove the new representation of the arctan operation, we only use ρcos2θ, that is, L1, where ρ is DoLP and θ is AoLP:
[0021] L1 = ρcos2θ = S1 / S0#(5);
[0022] Step 15, perform image annotation on each original polarized image data in the training dataset to obtain the Mask mask dataset; make the corresponding Edge edge image according to each Mask mask image in the Mask mask dataset;
[0023] Step 16, input each group of S0, DoLP, I E and the L1 representation image into the deep learning network detection model, train the deep learning network detection model, use the Mask mask image labeled with the corresponding original polarized image data in the training dataset as the training target, and perform edge supervision with the corresponding Edge edge image during the training process; until all the original polarized image data in the training dataset are used once, that is, one round of training is completed;
[0024] Step 17, set the number of training rounds; repeat Step 16 until all the training rounds are completed to obtain the final deep learning network detection model;
[0025] Step 18: Set up a test set containing the original transparent target image data, and use two evaluation metrics, MAE and S-measure, to evaluate the deep learning network detection model obtained in Step 17. If the test is qualified, a trained deep learning network detection model is obtained; otherwise, repeat Steps 16-17 to retrain the deep learning network detection model until a trained deep learning network detection model is obtained.
[0026] Step 2: Perform target detection on the transparent target through the trained deep learning network detection model.
[0027] Detection result = Trained deep learning network detection model(S0, DoLP, I E , L1).
[0028] Furthermore, in Step 13, Newton interpolation method or nearest neighbor interpolation method or bilinear interpolation method is used to demosaic the original polarization image data.
[0029] Furthermore, in Step 13, Newton interpolation method is used to demosaic the original polarization image data.
[0030] Furthermore, Step 14 specifically includes:
[0031] Step 141: The Stokes parameters S = [S0, S1, S2, S3] T is a set of values defined by George Gabriel Stokes to describe the polarization state of electromagnetic radiation; the four parameters of the Stokes vector represent different meanings respectively. S0 represents the total intensity of light, S1 represents the difference between the horizontally polarized light intensity and the vertically polarized light intensity, S2 represents the difference between the 45° polarized light intensity and the 135° polarized light intensity, and S3 represents the difference between the right-handed polarized light intensity and the left-handed polarized light intensity.
[0032]
[0033] Step 142: Degree of linear polarization DoLP and angle of linear polarization AoLP are two key metrics widely used in target detection tasks. These two metrics can be calculated from the first three parameters of the Stokes vector. The formulas are as follows:
[0034]
[0035] Step 143: The probing light is usually partially polarized light, composed of natural light and polarized light; the circular polarization component in natural targets is very small and can be ignored. Therefore, the polarized light is usually mainly linearly polarized light. The linearly polarized light I p is determined by S 1 and S2, and its calculation expression is:
[0036]
[0037] In nature, under the sun environment, the linearly polarized light in the reflected light is mainly horizontally polarized; therefore, S1 is far more significant than S2. For this reason, a linearly polarized enhancement parameter I for linearly polarized enhancement is proposed. E , In order to obtain more target polarization characteristics, the S1 parameter is added to I. p The second squared term of the parameter, and its calculation expression is:
[0038]
[0039] Step 144, based on the Umov effect, in a light field with weak light intensity, the scattering of light is more significant, but the noise in this part is also relatively high. Therefore, the linearly polarized angle image is very sensitive to noise. The noise sensitivity characteristic of AoLP is mainly caused by the division operation introduced therein. When S0 and S1 are close to 0, calculation errors will occur. The aliasing problem of 0 and π in AoLP is caused by the inability to distinguish 0 and π in the arctan operation; in the new representation, mainly to solve the aliasing problem of 0 and π in AoLP, an equivalent expression of ρ and θ is proposed, that is, a new representation that removes the arctan operation. We only use ρcos2θ, that is, L1, and the formula is as follows, where ρ is DoLP and θ is AoLP:
[0040] L1 = ρcos2θ = S1 / S0#(5).
[0041] Furthermore, the deep learning network detection model in step 14 is mainly based on the ResNet-101 network or the VGG network as the backbone network.
[0042] Furthermore, the deep learning network detection model in step 14 is mainly based on the ResNet-101 network as the backbone network;
[0043] S0, DoLP, I E , The four characterization images of L1 are input into the deep learning network detection model; each input image passes through a unique backbone network, and each input will obtain five side path features Conv1-2, Conv2-2, Conv3-3, Conv4-3, Conv5-3. Discard the side path Conv1-2, and then fuse the four pictures convolutional out by each layer of the backbone network starting from the second layer of Conv2-2, Conv3-3, Conv4-3, Conv5-3 into one picture through a one-to-one fusion module; finally, the four inputs S0, DoLP, I EFour images respectively fused from [[ID=]] and L1 through the corresponding backbone networks are input into the FPN network to fuse an output image, and the output image is supervised by the Mask mask of the corresponding original image data; this not only improves the positioning accuracy of high-order prediction, but more importantly, improves the segmentation details and obtains better prediction results;
[0044] The Feature Pyramid Networks (FPN) is a network proposed in 2017. The FPN algorithm simultaneously utilizes the high resolution of low-level features and the high semantic information of high-level features, and achieves the prediction effect by fusing the features of these different layers. Without substantially increasing the computational complexity of the original model, it greatly improves the performance of small object detection. And the prediction is carried out separately on each fused feature layer, and the effect is very good.
[0045] Furthermore, Edge edge supervision is added to the features on the side of Conv2-2 to obtain better edge features; the boundary features after Edge edge supervision of Conv2-2 are integrated into the transparent object features in the subsequent branches through FPN; the fused feature maps of the Conv3-3, Conv4-3, and Conv5-3 branches are supervised by the Mask masks of the corresponding original polarization image data;
[0046] Final model training result = Deep learning network (Mask, Edge, S0, DoLP, I E , L1);
[0047] Mask: The Mask mask image obtained by making corresponding mask markings on the original polarization image data;
[0048] Edge: That is, the edge image of the original polarization image data, which is calculated according to the Mask mask image of the corresponding original polarization image data.
[0049] Furthermore, the loss function of the deep learning network detection model is determined by the loss functions of three parts:
[0050] The first part: To explicitly model the transparent edge features, we add additional transparent edge supervision to enhance the transparent edge features; use cross-entropy loss, which can be defined as:
[0051]
[0052] where Z + and Z - respectively represent the set of edge pixels and the set of background pixels, Pr(y j = 1|C (2)) is a prediction map, and each value in the map represents the transparency edge confidence of a pixel;
[0053] The second part: Mask supervision for the fused feature maps of the Conv3-3, Conv4-3, and Conv5-3 branches:
[0054]
[0055] Among them, Y + and Y - represent the set of transparent region pixels and the set of opaque pixels respectively; is a prediction map, and each value in the map represents the transparency region confidence of a pixel;
[0056] The third part: Mask supervision for the final result obtained finally
[0057]
[0058] Among them, K + and K - represent the set of transparent region pixels and the set of opaque pixels in the final transparent target segmentation result output by FPN; Pr(y j =1|G) is the final transparent target segmentation prediction map, and each value in the map represents the transparency region confidence of a pixel;
[0059] Therefore, we define the total training loss function as follows:
[0060]
[0061] Furthermore, the mechanism of the fusion module of the deep learning network detection model is pixel weighting.
[0062] Furthermore, in step 18, a test set containing 93 original transparent target image data is set up to evaluate the trained deep learning network detection model;
[0063] MAE is a metric for evaluating the average difference between the prediction map and the Mask, and is expressed by the following formula:
[0064]
[0065] P(x, y) and Y(x, y) represent the prediction map and the Mask mapped and normalized to [0, 1];
[0066] W and H are the width and height of the image respectively;
[0067] The S-measure focuses on evaluating structural information, which is closer to the human visual system than the F-measure; therefore, the S-measure is added for a more comprehensive evaluation, and the S-measure can be calculated using the following formula:
[0068] S-measure = γO i +(1 - γ)O r #(13)
[0069] O i and O r are the regional perception and object perception structural similarities respectively; γ is default set to 0.5.
[0070] Furthermore, the specific steps of step 2 include:
[0071] Step 21, obtaining the polarization image of the transparent target to be measured using polarization imaging technology;
[0072] Step 22, demosaicing the polarization image of the transparent target to be measured to obtain the linearly polarized images I0(x,y), I 45 (x,y), I 90 (x,y) and I 135 (x,y);
[0073] Step 23, according to I0(x,y), I 45 (x,y), I 90 (x,y) and I 135 (x,y) in step 22, using the polarization physical formula to construct the S0, DoLP, I E and L1 characterization images of the transparent target to be measured;
[0074] Step 24, inputting the four characterization images of S0, DoLP, I E and L1 obtained in step 23 into the trained deep learning network detection model for detection;
[0075] Detection result = Trained deep learning network detection model(S0, DoLP, I E , L1).
[0076] The transparent target detection method based on polarization imaging and deep learning has the following beneficial effects:
[0077] (1) The transparent object detection method based on polarization imaging and deep learning of the present invention is different from other methods of synthesizing multi-scale features or using post-processing. It focuses on the complementarity between edge information and transparent object information in polarization images, proposes a deep learning network detection model, applies these polarization and complementary feature information in the network, and fuses these polarization and complementary features together, greatly improving the accuracy of transparent boundaries and the localization of transparent objects.
[0078] (2) In the present invention, the original polarization image is demosaicked into I0(x,y), I 45 (x,y), I 90 (x,y) and I 135 (x,y). The polarization physical equations are applied to construct S0, DoLP, I E and L1, and these four characterization images are normalized to the range of [0-255] and then input into the network. The features of the same resolution are weighted for each input and each layer. The feature maps of each same layer in the backbone network passed by each input are weighted and fused, and then the four fused feature maps are input into the FPN network to fuse the final detection result map, which not only improves the localization accuracy of high-order prediction, but more importantly, improves the segmentation details and obtains better detection results.
[0079] (3) In the present invention, the detection work focuses on the edge detection of transparent objects. To make the detection more accurate, we model the edges of transparent objects. In this module, we aim to model the transparent edge information and extract transparent edge features. As mentioned above, Conv2-2 retains better edge information. Therefore, we extract local edge information from Conv2-2 and strengthen the edge features of the transparent object through the enhanced linear polarization angle image in the one-to-one fusion module, that is, the edge features in L1.
[0080] (4) In the present invention, in order to obtain more accurate transparent object detection results, for the progressive extraction of transparent object features, the fusion feature maps of the Conv3-3, Conv4-3, and Conv5-3 branches and the final segmentation results are supervised by the mask of the transparent object. The boundary features supervised by Conv2-2 are integrated into the transparent object features in the subsequent branches through FPN, and better detection results are obtained.
[0081] (5) The present invention realizes the detection of transparent items, which has extremely important practical significance and application value in the fields of intelligent manufacturing and so on. BRIEF DESCRIPTION OF THE DRAWINGS
[0082] Figure 1 : Flowchart of the transparent object detection method based on polarization imaging and deep learning in the embodiment of the present invention;
[0083] Figure 2: Schematic diagram of the optical system for collecting raw data in the embodiment of the present invention;
[0084] Figure 3 : Schematic diagram of the structure of the deep learning network detection model in the embodiment of the present invention;
[0085] Figure 4 : Schematic diagram of the method for detecting transparent objects based on polarization imaging and deep learning in the embodiment of the present invention. Detailed implementation manners
[0086] The following further describes the present invention with reference to the accompanying drawings:
[0087] Figures 1 to 4 Shows a specific implementation manner of a method for detecting transparent objects based on polarization imaging and deep learning according to the present invention. Figure 1 Is the flowchart of the method for detecting transparent objects based on polarization imaging and deep learning in this embodiment; Figure 2 Is the schematic diagram of the optical system structure for collecting raw data in this embodiment; Figure 3 Is the schematic diagram of the structure of the deep learning network detection model in this embodiment; Figure 4 Is the schematic diagram of the method for detecting transparent objects based on polarization imaging and deep learning in the embodiment of the present invention.
[0088] As Figures 1 to 4 shown, the method for detecting transparent objects based on polarization imaging and deep learning in this embodiment includes the following steps:
[0089] Step 1, training the deep learning network detection model;
[0090] Step 11, building a data acquisition system for polarized images of transparent objects, and adjusting the polarizing film to meet S1 significantly;
[0091] As Figure 2 shown, a polarizer is covered on the light source, the polarized light is projected onto the transparent object, and then a polarization camera is used to collect images;
[0092] Step 12, collecting polarized image data, saving it in bmp format, and constructing a training data set;
[0093] In this embodiment, the training data set contains 400 images from 3 possible categories: glassware, plastic balls, and other plastic products;
[0094] Step 13, demosaicking each original polarized image data in the training data set to obtain the linearly polarized images I0(x, y), I 45 (x, y), I 90 (x, y) and I 135 (x, y);
[0095] Step 14. Based on the four-angle linearly polarized images I0(x,y), I 45 (x,y), I 90 (x,y) and I 135 (x,y) obtained from each piece of original polarized image data in Step 13, construct a set of S0, DoLP, I E and L1 images corresponding to each piece of original polarized image data according to the polarization principle using the Stokes vector, linear polarization angle, and degree of linear polarization formulas. S0 is the detected total intensity, usually the sum of two mutually perpendicular linearly polarized images, and the formula is as follows:
[0096] S0 = I0(x,y) + I 90 (x,y) = I 45 (x,y) + I 135 (x,y) #(1);
[0097] The concept of the degree of linear polarization DoLP is proposed based on the differences in the internal structure, appearance, material, and roughness of different objects, reflecting the differences in the radiation characteristics of the polarization states of objects. The formula is as follows:
[0098]
[0099] S1 is a parameter calculated based on the Stokes vector parameters for two polarized images I0(x,y), I 90 (x,y) of the same piece of original polarized image data, representing the difference between the light intensity in the horizontal direction and the light intensity in the vertical direction; S2 is a parameter calculated based on the Stokes vector parameters for two polarized images I 45 (x,y), I 135 (x,y) of the same piece of original polarized image data, representing the difference between the light intensity in the 45° direction and the light intensity in the 135° direction;
[0100] I E represents the enhanced linearly polarized image, and the formula is as follows:
[0101]
[0102] AoLP represents the linearly polarized angle image, that is, the polarization azimuth angle, which can well reflect the illumination, shadow, surface flatness, stress birefringence and other attribute information in the light field, and more contains the edge information of transparent targets. The formula is as follows:
[0103]
[0104] To solve the aliasing problem of 0 and π in AoLP, we propose L1. Since the division operation in the linear polarization angle formula makes it very sensitive to noise, calculation errors occur when S0 and S1 are close to 0. The aliasing problem of 0 and π in AoLP is caused by the inability to distinguish 0 and π in the arctan operation. In the new representation, mainly to solve the aliasing problem of 0 and π, that is, a new representation that removes the arctan operation, we only use ρcos2θ, namely L1, where ρ is DoLP and θ is AoLP:
[0105] L1 = ρcos2θ = S1 / S0#(5);
[0106] Step 15, perform image annotation on each original polarization image data in the training dataset to obtain a Mask mask dataset; make corresponding Edge edge images according to each Mask mask image in the Mask mask dataset;
[0107] Step 16, input each group of S0, DoLP, I E and L1 characterization images in the training dataset into the deep learning network detection model, train the deep learning network detection model, use the labeled Mask mask image corresponding to the original polarization image data in the training dataset as the training target, and perform edge supervision with the corresponding Edge edge image during the training process; until all the original polarization image data in the training dataset are used once, that is, one round of training is completed;
[0108] Step 17, set the number of training rounds; repeat Step 16 until all the training rounds are completed to obtain the final deep learning network detection model;
[0109] Step 18, set a test set containing the original image data of the transparent target, and use two evaluation metrics, MAE and S-measure, to evaluate the deep learning network detection model obtained in Step 17; if the test is qualified, the trained deep learning network detection model is obtained; otherwise, repeat Steps 16 - 17 to retrain the deep learning network detection model until the trained deep learning network detection model is obtained;
[0110] In this embodiment, the structure of the deep learning network detection model is as Figure 3 shown. After training the model for 34 rounds, we constructed a test set containing 93 images to evaluate the performance of the model. The average values of the two evaluation metrics of the test set were calculated as MAE: 0.0076 and S-measure: 0.9628 respectively, and the trained deep learning network detection model was obtained;
[0111] Step 2, perform target detection on the transparent target through the trained deep learning network detection model;
[0112] Detection result = the trained deep learning network detection model (S0, DoLP, I E , L1), as Figure 4 shown.
[0113] Preferably, in step 13, Newton interpolation method or nearest neighbor interpolation method or bilinear interpolation method is used to demosaic the original polarization image data. In this embodiment, Newton interpolation method is used to demosaic the original polarization image data.
[0114] Preferably, step 14 specifically includes:
[0115] Step 141, Stokes parameters S = [S0, S1, S2, S3] T is a set of values that describe the polarization state of electromagnetic radiation defined by George Gabriel Stokes; the four parameters of the Stokes vector represent different meanings respectively. S0 represents the total intensity of light, S1 represents the difference between the horizontally polarized light intensity and the vertically polarized light intensity, S2 represents the difference between the 45° polarized light intensity and the 135° polarized light intensity, and S3 represents the difference between the right-handed polarized light intensity and the left-handed polarized light intensity;
[0116]
[0117] Step 142, Degree of linear polarization DoLP and Angle of linear polarization AoLP are two key metrics widely used in target detection tasks. These two metrics can be calculated from the first three parameters of the Stokes vector, and the formulas are as follows:
[0118]
[0119]
[0120] Step 143, The detection light is usually partially polarized light, which is composed of natural light and polarized light; the circular polarization component in natural targets is very small and can be ignored. Therefore, the polarized light is usually mainly linearly polarized light. The linearly polarized light I p is determined by S1 and S2, and its calculation expression is:
[0121]
[0122] In nature, under the sun environment, the reflected light is mainly horizontally linearly polarized light; therefore, S1 is much more significant than S2. For this reason, a linearly polarized enhancement parameter I E is proposed. In order to obtain more target polarization features, the S1 parameter is added to the second square term of the I p parameter, and its calculation expression is:
[0123]
[0124] Step 144, based on the Umov effect, in a light field with weak light intensity, the scattering of light is relatively significant, but the noise in this part is also relatively high. Therefore, the linear polarization angle image is very sensitive to noise. The noise-sensitive characteristic of AoLP is mainly introduced by the division operation therein. When S0 and S1 are close to 0, calculation errors will occur. The aliasing problem of 0 and π in AoLP is caused by the inability to distinguish 0 and π in the arctan operation; in the new representation, mainly to solve the aliasing problem of 0 and π in AoLP, an equivalent expression of ρ and θ is proposed, that is, a new representation that removes the arctan operation. We only use ρcos2θ, that is, L1. The formula is as follows, where ρ is DoLP and θ is AoLP:
[0125] L1 = ρcos2θ = S1 / S0#(5).
[0126] Preferably, the deep learning network detection model in step 14 uses the ResNet-101 network or the VGG network as the backbone network; in this embodiment, the ResNet-101 network is used as the backbone network;
[0127] Input the four representation images of S0, DoLP, I E and L1 into the deep learning network detection model; each input image passes through a unique backbone network, and each input will obtain five side path features Conv1-2, Conv2-2, Conv3-3, Conv4-3, Conv5-3. Discard the side path Conv1-2, and then fuse the four pictures output from each layer of the backbone network starting from the second layer of Conv2-2, Conv3-3, Conv4-3, Conv5-3 into one picture through a one-to-one fusion module; finally, input the four pictures output from the corresponding backbone network fusions of the four inputs S0, DoLP, I E and L1 into the FPN network, and fuse them into an output picture, which is supervised by the Mask mask image of the corresponding original polarization image data; this not only improves the positioning accuracy of high-order prediction, but more importantly, improves the segmentation details and obtains better prediction results.
[0128] Preferably, add Edge edge detection to the Conv2-2 side feature to obtain better edge features; integrate the boundary features supervised by Conv2-2 edges into the transparent target features in the subsequent branches through FPN; the fusion feature maps of the Conv3-3, Conv4-3, and Conv5-3 branches are supervised by the Mask mask image of the corresponding original polarization image data;
[0129] Final model training result = Deep learning network detection model (Mask, Edge, S0DoLP, I E , L1);
[0130] Mask: The Mask mask image obtained by making corresponding mask markings on the original polarization image data;
[0131] Edge: That is, the edge image of the original polarization image data, calculated based on the Mask mask image of the corresponding original polarization image data.
[0132] Preferably, the loss function of the deep learning network detection model is determined by the loss functions of three parts:
[0133] The first part: To explicitly model the transparent edge features, we added additional transparent edge supervision to enhance the transparent edge features; using cross - entropy loss, which can be defined as:
[0134]
[0135] where Z + and Z - represent the edge pixel set and the background pixel set respectively, and Pr(y j = 1∣C (2) ) is the prediction map, and each value in the map represents the transparent edge confidence of the pixel;
[0136] The second part: Mask supervision for the fused feature maps of the Conv3 - 3, Conv4 - 3, and Conv5 - 3 branches:
[0137]
[0138] where Y + and Y - represent the transparent region pixel set and the opaque pixel set respectively; is the prediction map, and each value in the map represents the transparent region confidence of the pixel;
[0139] The third part: Mask supervision for the finally obtained final result
[0140]
[0141] where K + and K - represent the transparent region pixel set and the opaque pixel set in the final transparent target segmentation result output by FPN; Pr(y j = 1∣G) is the final transparent target segmentation prediction map, and each value in the map represents the transparent region confidence of the pixel;
[0142] Therefore, we define the total training loss function as follows:
[0143]
[0144] Preferably, the mechanism of the fusion module of the deep learning network detection model is pixel weighting.
[0145] Preferably, step 2 specifically includes:
[0146] Step 21, using polarization imaging technology to obtain the polarization image of the transparent target to be measured;
[0147] Step 22, demosaicking the polarization image of the transparent target to be measured to obtain linearly polarized images I0(x,y), I 45 (x,y), I 90 (x,y) and I 135 (x,y);
[0148] Step 23, according to I0(x,y), I 45 (x,y), I 90 (x,y) and I 135 (x,y) in step 22, using polarization physical formulas to construct the S0, DoLP, I E and L1 images of the transparent target to be measured;
[0149] Step 24, inputting the four characterization images of S0, DoLP, I E and L1 obtained in step 23 into the trained deep learning network detection model for detection;
[0150] Detection result = trained deep learning network detection model(S0, DoLP, I E ,L1).
[0151] In this embodiment, our model is implemented under the PyTorch architecture, and the model is trained for 34 rounds. During training and testing, the images are resized to a resolution of 612×512.
[0152] The transparent target detection method of the present invention based on polarization imaging and deep learning is different from other methods of synthesizing multi-scale features or using post-processing. It focuses on the complementarity between edge information and transparent object information in the polarization image, proposes a deep learning network detection model, applies these polarization and complementary feature information in the network, and fuses these polarization and complementary features together, greatly improving the accuracy of the transparent boundary and the localization of the transparent object.
[0153] In the present invention, the original polarization image is demosaicked into I0(x,y), I 45 (x,y), I 90 (x,y) and I135 (x, y), construct S0, DoLP, I by applying the polarization physical equation E And L1, normalize these four characterization images to the range of [0 - 255] and then input them into the network. For each input and each layer, weight the features with the same resolution. Weight and fuse the feature maps of each same layer in the backbone network passed by each input, and then input the four fused feature maps into the FPN network to fuse out the final detection result map, which not only improves the positioning accuracy of high-order prediction, but more importantly, improves the segmentation details and obtains better detection results.
[0154] In the present invention, the detection work focuses on the edge detection of transparent objects. To make the detection more accurate, we model the edges of transparent objects. In this module, we aim to model the transparent edge information and extract transparent edge features. As mentioned above, Conv2-2 retains better edge information. Therefore, we extract local edge information from Conv2-2 and strengthen the edge features of the transparent object through the enhanced linear polarization angle image in the one-to-one fusion module, that is, the edge features in L1.
[0155] In the present invention, in order to obtain more accurate detection results of transparent objects, for the progressive extraction of transparent object features, the fused feature maps of the Conv3-3, Conv4-3, and Conv5-3 branches and the final segmentation results are supervised by the mask of the transparent object. The boundary features supervised by Conv2-2 are integrated into the transparent object features in the subsequent branches through FPN, and better detection results are obtained.
[0156] The present invention realizes the detection of transparent objects, which has extremely important practical significance and application value in the fields such as intelligent manufacturing.
[0157] The present invention has been described exemplarily above in conjunction with the accompanying drawings. Obviously, the implementation of the present invention is not limited by the above-mentioned manner. As long as various improvements are made by adopting the method concept and technical solution of the present invention, or the concept and technical solution of the present invention are directly applied to other occasions without improvement, they are all within the protection scope of the present invention.
Claims
1. A transparent target detection method based on polarization imaging and deep learning, characterized in that, Including the following steps: Step 1, training a deep learning network detection model; Step 11, building a transparent target polarization image data acquisition system; Step 12, using the acquisition system to collect polarization image data and constructing a training data set; Step 13: Demosaic each original polarization image data in the training dataset to obtain corresponding linear polarization images I0(x, y), I45(x, y), I 90 (x, y) and I 135 (x, y); Step 14, for the four angular linearly polarized images I0(x, y), I 45 (x, y), I 90 (x, y) and I 135 (x, y) obtained from each of the original polarization image data in Step 13, construct a set of S0, DoLP, I E and L1 images corresponding to each of the original polarization image data according to the polarization principle using the Stokes vector, linear polarization angle, and degree of linear polarization formulas. S0 is the detected total intensity, which is the sum of two mutually orthogonal linearly polarized images, and the formula is as follows: S0 = I0(x, y) + I 90 (x, y) = I 45 (x, y) + I 135 (x, y) # (1); The degree of linear polarization DoLP reflects the difference in the radiation characteristics of the polarization state of an object. The formula is as follows: S1 is a parameter calculated based on the Stokes vector parameters for two polarization images I0(x, y) and I 90 (x, y) of the same original polarization image data, representing the difference between the light intensity in the horizontal direction and the light intensity in the vertical direction; S2 is a parameter calculated based on the Stokes vector parameters for two polarization images I 45 (x, y) and I 135 (x, y) of the same original polarization image data, representing the difference between the light intensity in the 45° direction and the light intensity in the 135° direction; I E represents the linearly polarized image after enhancement, and the formula is as follows: AoLP represents the angle-of-linear-polarization image, which contains more edge information of the transparent target. The formula is as follows: To solve the aliasing problem of 0 and π in AoLP, we propose L1. In the formula, ρ is DoLP and θ is AoLP: L1 = ρcos2θ = S1 / S0#(5); Step 15, performing image annotation on each original polarization image data in the training data set to obtain a Mask mask data set; Making a corresponding Edge edge image according to each Mask mask image in the Mask mask data set; Step 16: Input each group of S0, DoLP, I E and L1 characterization images in the training dataset into the deep learning network detection model, and train the deep learning network detection model. Use the annotated Mask mask image corresponding to the original polarization image data in the training dataset as the training target, and perform edge supervision with the corresponding Edge edge image during the training process; until all the original polarization image data in the training dataset have been used once, that is, one round of training is completed; Wherein the deep learning network detection model is based on a ResNet-101 network or a VGG network as the backbone network, and its features are: Input the four representation images of S0, DoLP, I E , and L1 into the deep learning network detection model; each input image passes through a unique backbone network, and each input obtains five side-path features Conv1-2, Conv2-2, Conv3-3, Conv4-3, and Conv5-3. Discard the side path Conv1-2, and then fuse the four pictures convolutionalized from each layer starting from the second layer of the backbone network by Conv2-2, Conv3-3, Conv4-3, and Conv5-3 into one picture through a one-to-one fusion module; add Edge edge supervision to the Conv2-2 side-path feature to obtain better edge features; integrate the boundary features after Conv2-2 edge supervision into the transparent target features in the subsequent branches through FPN; the fused feature maps of the Conv3-3, Conv4-3, and Conv5-3 branches are supervised by the Mask mask images of the corresponding original polarization image data; Finally, the four inputs S0, DoLP, I E and L1 are respectively input into the FPN network through four pictures fused by the corresponding backbone networks, and an output picture is fused. The output picture is supervised by the Mask mask picture of the corresponding original polarization image data; Step 17, setting the number of training epochs; repeating Step 16 until all training epochs are completed to obtain the final deep learning network detection model; Step 18, setting a test set containing the original image data of the transparent target, and using two evaluation metrics, MAE and S-measure, to evaluate the deep learning network detection model obtained in Step 17; if the test is qualified, the trained deep learning network detection model is obtained; otherwise, repeat Steps 16-17 to retrain the deep learning network detection model until the trained deep learning network detection model is obtained; Final model training result = Deep learning network detection model (Mask, Edge, S0, DoLP, I E , L1); Mask: A Mask mask image obtained by making corresponding mask markings on the original polarization image data; Edge: That is, the edge image of the original polarization image data, calculated according to the Mask mask image of the corresponding original polarization image data; Step 2, performing target detection on the transparent target through the trained deep learning network detection model: Detection result = trained deep learning network detection model (S0, DoLP, I E , L1).
2. The transparent object detection method based on polarization imaging and deep learning according to claim 1, wherein, In Step 13, the Newton interpolation method, the nearest neighbor interpolation method, or the bilinear interpolation method is used to demosaic the original polarization image data.
3. The transparent target detection method based on polarization imaging and deep learning according to claim 1, characterized in that The specific content of Step 14 includes: Step 141, Stokes parameter S = [S0, S1, S2, S3] T is a set of values defined by George Gabriel Stokes to describe the polarization state of electromagnetic radiation; the four parameters of the Stokes vector represent different meanings respectively. S0 represents the total intensity of light, S1 represents the difference between the horizontally polarized light intensity and the vertically polarized light intensity, S2 represents the difference between the 45° polarized light intensity and the 135° polarized light intensity, and S3 represents the difference between the right-handed polarized light intensity and the left-handed polarized light intensity; Step 142, the degree of linear polarization DoLP and the angle-of-linear-polarization AoLP are two key metrics widely used in target detection tasks. These two metrics can be calculated from the first three parameters of the Stokes vector. The formula is as follows: Step 143, the probing light is usually partially polarized light, which consists of natural light and polarized light; the circularly polarized component in natural targets is very small and can be ignored. Therefore, the polarized light usually mainly consists of linearly polarized light, and the linearly polarized light I p is determined by S1 and S2, and its calculation expression is: In nature, under the sun's environment, the reflected light is mainly linearly polarized light in the horizontal direction; therefore, S1 is far more significant than S2. For this reason, a linearly polarized enhancement parameter I for linearly polarized enhancement is proposed. E , in order to obtain more target polarization characteristics, the S1 parameter is added to the second squared term of the I p parameter, and its calculation expression is: Step 144, based on the Umov effect, in a light field with weak light intensity, the scattering of light is more significant, but the noise in this part is also relatively high. Therefore, the angle-of-linear-polarization image is very sensitive to noise. The noise sensitivity characteristic of AoLP is mainly caused by the division operation introduced therein. When S0 and S1 are close to 0, calculation errors will occur. The aliasing problem of 0 and π in AoLP is caused by the inability to distinguish 0 and π in the arctan operation; in the new representation, mainly to solve the aliasing problem of 0 and π, an equivalent expression of ρ and θ is proposed, that is, a new representation that removes the arctan operation, and only uses ρcos2θ, that is, L1. The formula is as follows, where ρ is DoLP and θ is AoLP: L1 = ρcos2θ = S1 / S0 #(5).
4. The transparent target detection method based on polarization imaging and deep learning according to claim 1, characterized in that The loss function of the deep learning network detection model is determined by the loss functions of three parts: The first part: To explicitly model the transparent edge features, we add additional transparent edge supervision to enhance the transparent edge features; the cross - entropy loss is used, which can be defined as: where Z + and Z - represent the edge pixel set and the background pixel set respectively, and Pr(y j = 1|C (2) ) is the prediction map, where each value represents the transparent edge confidence of the pixel; The second part: Mask supervision for the fused feature maps of the Conv3 - 3, Conv4 - 3, and Conv5 - 3 branches: where Y + and Y - represent the transparent region pixel set and the opaque pixel set respectively; is the prediction map, where each value represents the transparency region confidence of the pixel; The third part: Mask supervision for the finally obtained final result Among which K + and K - represent the transparent region pixel set and the opaque pixel set in the final transparent object segmentation result output by the FPN; Pr(y j = 1|G) is the final transparent object segmentation prediction map, where each value represents the transparency region confidence of a pixel; Therefore, we define the total training loss function as follows:
5. The transparent object detection method based on polarization imaging and deep learning according to claim 1, characterized in that The mechanism of the fusion module of the deep learning network detection model is pixel weighting.
6. The transparent object detection method based on polarization imaging and deep learning according to claim 1, characterized in that, In step 18, a test set containing 93 original transparent target image data is set up.
7. The transparent target detection method based on polarization imaging and deep learning according to claim 1, characterized in that Step 2 specifically includes: Step 21, using polarization imaging technology to obtain the polarization image of the transparent target to be measured; Step 22, demosaic the polarization image of the transparent target to be measured to obtain linearly polarized images I0(x, y), I 45 (x, y), I 90 (x, y) and I1 35 (x, y); Step 23, according to I0(x, y) in Step 22, I 45 (x, y), I 90 (x, y) and I 135 (x, y), use the polarization physical formula to construct the S0, DoLP, I E and L1 characterization images of the transparent target to be measured; Step 24, input the four characterization images of S0, DoLP, I E and L1 obtained in Step 23 into the trained deep learning network detection model for detection; Detection result = Trained deep learning network detection model (S0, DoLP, I E , L1).