An unmanned vehicle vision detection method considering the influence of transparent objects
By introducing edge detection module and preliminary segmentation module in the process of unmanned vehicle drawing construction, combining confidence factor and multi-scale features, input into the Transformer-based encoder and decoder channels, the problem of insufficient shape information of transparent objects in unmanned vehicle drawing construction is solved, and the accuracy and robustness of detection and segmentation are significantly improved.
Patent Information
- Application Number
- CN202411154232.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-08-22
- Publication Date
- 2025-07-01
- Estimated Expiration
- 2044-08-22
AI Technical Summary
The prior art fails to fully consider the shape information of transparent objects during the drawing construction process of unmanned vehicles, resulting in unsatisfactory drawing construction results, unsmooth edges and low accuracy.
An edge detection module and a preliminary segmentation module are introduced to increase confidence factor, improve the accuracy of object detection segmentation through the fusion of edge features, and input it into the Transformer-based encoder and decoder channels in combination with multi-scale features.
The robustness and accuracy of unmanned vehicles in the process of transparent object detection and segmentation are improved. The average cross-convergence ratio (mIoU) of the model is 74.43%, and the pixel accuracy (ACC) is 96.35%. The visual detection ability of unmanned vehicles is significantly improved.
Smart Images

Figure CN119068299B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of intelligent perception of unmanned vehicles and artificial intelligence applications, especially the sensor detection technology, machine learning, and computer vision technology in the process of map building for intelligent unmanned vehicles. In particular, it relates to a vision detection method for unmanned vehicles considering the influence of transparent objects. Background Art
[0002] With the rapid development of intelligent manufacturing industries such as automated guided vehicles (AGVs), laser-guided vehicles (LGVs), unmanned ships, and intelligent unmanned aerial vehicles in the field of unmanned vehicles, it also stimulates the common development of artificial intelligence fields such as intelligent perception and control devices, and computer vision and auditory technologies. There is a rich variety of laser measurement instruments, and it is particularly important to select a suitable model and install it on a laser-guided vehicle to quickly and accurately construct an external environment map. At the same time, using a vision sensor to complete object recognition and assist in constructing the external environment map can avoid serious consequences caused by the failure of a single sensor, such as transparent objects like glass doors, glass windows, and glass curtain walls. Therefore, using deep learning methods in machine learning and computer vision can quickly and accurately achieve good results.
[0003] Chinese Patent CN 117451029A, "A Robot Mapping Method for Multi-Sensor Fusion in a Glass Environment", proposes a multi-sensor fusion method that uses a vision camera to detect the presence of glass and combines ultrasonic data to compensate the lidar to achieve the construction of a two-dimensional grid map. However, this method only simply identifies the glass and does not know the shape of the glass, which will lead to errors in the construction of a three-dimensional point cloud map due to the lack of glass shape information, resulting in a poor mapping effect.
[0004] The literature "Segmenting transparent object in the wild with transformer" proposes a vision segmentation method Trans2Seg, which introduces a structure based on Transformer encoders and decoders on the basis of the CNN architecture. The input image is feature-extracted, and the extracted features are fed into the encoder for self-attention. A learnable class prototype is introduced in the decoder stage, and finally, a smaller convolution is used to combine the decoder output and the shallow features of a higher resolution to output the segmentation result. However, this method lacks the edge information of the image, which will result in good internal annotation of the image but uneven edges of the segmented image and a low feature matching rate.
[0005] The literature "Research and Application of Semantic Segmentation of Transparent Objects Incorporating Edge Detection Networks" proposed a semantic segmentation network integrating an edge detection module, and jointly trained the semantic segmentation module and the edge detection module through multi-task learning to complete the segmentation task of semantic glass objects. Although this method combines edge features at different stages to assist semantic segmentation in generating a segmentation prediction map, it does not consider the issue of confidence at each stage. At the same time, it only has a simple decoder structure and lacks learnable class prototypes for reference, which may lead to inaccurate segmented objects due to pixel annotation errors.
[0006] The literature "An Open-pit Mine Road Network Extraction Method Based on Improved DeepLabv3+ Network" and "Apple Planting Area Extraction Based on Improved DeepLab V3+" both proposed using the ASPP module to enhance the effect of large-scale feature extraction. However, these studies did not fully consider the importance of edge features in small object segmentation and ignored the possible role of the global receptive field in transparent object segmentation, which may cause deviations in edge determination when segmenting transparent objects.
[0007] In real indoor or outdoor working scenarios, due to the special material of transparent objects, the light beam of laser measuring instruments will be refracted or penetrated. This characteristic will lead to misjudgment of the external environment during the mapping process of autonomous vehicles. The present invention identifies and segments transparent objects through deep learning methods to solve the problem of unsatisfactory mapping results of transparent objects encountered during the mapping process of autonomous vehicles equipped with laser measuring instruments. An edge detection module and a preliminary segmentation module are introduced, and a confidence factor is added to obtain more edge features. The accuracy of object detection and segmentation is improved through the fusion of edge features, the detection ability of the model is enhanced, more information is obtained by combining multi-scale features, and the most accurate prediction results are obtained through a self-attention encoder and a decoder with learnable class prototypes. Summary of the Invention
[0008] In view of the deficiencies considered in the existing method technologies, the present invention proposes a vision detection method for autonomous vehicles considering the influence of transparent objects, focusing on solving the problem that existing traditional deep learning methods do not fully consider the information required during the mapping process of autonomous vehicles in the process of transparent object detection and segmentation. An edge detection module is introduced and a confidence factor is added to obtain more edge features. The edge features of the preliminary segmentation map obtained by the preliminary segmentation module are combined to obtain more accurate new edge features, which are fused with multi-scale features and input into the Transformer-based encoder and decoder channels, improving the robustness and accuracy of the model in detecting and segmenting transparent objects.
[0009] To achieve the above objectives, the present invention adopts the following specific technical solutions to solve the problem:
[0010] S1: Select a suitable dataset based on the actual application scenario of the unmanned vehicle, and perform deletion, modification, and addition on the selected dataset. Specifically, it includes the following sub-steps:
[0011] S1.1: When the unmanned vehicle uses a 2D lidar to work in a scenario with transparent objects, due to the special nature of the material, the laser beam emitted and received by the lidar will be reflected and refracted, resulting in errors in the laser data returned to the lidar. More seriously, it will pass through the transparent object and mislabel the location that should be detected as an obstacle as a path that can be normally passed in the map, which will cause serious consequences such as collisions and inability to locate during the task execution of the unmanned vehicle;
[0012] S1.2: Based on the similarity and difference between the maps established considering 2D lidar and 3D lidar, select the transparent glass dataset Trans10k-v2;
[0013] S1.3: Considering the actual application scenario, delete objects that are rarely or almost never encountered during the map building process of the unmanned vehicle in the selected dataset, such as glasses, glass cups, glass bowls, etc., and retain 5 class prototypes of common transparent plastic bottles, glass windows, glass doors, glass walls, and transparent cabinets in daily life. Based on the above 5 class prototypes, add corresponding data pictures;
[0014] S2: Introduce an edge detection module, introduce a confidence factor on the basis of considering the need to pay more attention to the edges of transparent objects, and increase the confidence at different stages to obtain stage features more conducive to the segmentation of transparent objects. Specifically, it includes the following sub-steps:
[0015] S2.1: Use ResNet50 based on convolutional neural network as the backbone network in the feature extraction stage;
[0016] S2.2: Input the image into the convolutional neural network to extract the features of each stage of the image
[0017] S2.3: The image features extracted at each stage of the backbone network Are input into the edge detection module;
[0018] S2.4: The edge detection module extracts the image features of the corresponding stage. Consider designing the edge detection module in the form of a similar image pyramid. The edge loss function is described using DiceLoss, as shown in Equation (1):
[0019]
[0020] In the formula, y i And Represent the label value and predicted value of pixel i respectively, and N represents the total number of pixels;
[0021] S2.5: Considering that different stages have different edge features, a confidence factor α is added in each stage i , to ensure that the edge loss calculation in each stage does not change the scale and to more conveniently control the input and output of each stage, the confidence factor is normalized to ensure that the confidence factor satisfies where the confidence factor of the second stage is given a higher value than other stages;
[0022] S2.6: The mean of the sum of the losses of each stage is used to represent the total loss function L of the edge detection module s_edm , and its calculation formula is shown in Equation (2):
[0023]
[0024] In the formula, n represents the total number of each stage, and α i represents the confidence factor added to each stage, and the confidence factor satisfies
[0025] S2.7: The edge feature V of the feature map output by the edge detection module 1_features ;
[0026] S3: A preliminary segmentation module is proposed. The feature map generated by the preliminary segmentation module is subjected to edge feature extraction and analyzed and corrected with the edge feature map generated by the edge detection module to generate a new and more accurate edge feature, which specifically includes the following sub-steps:
[0027] S3.1: The features extracted by the backbone network are input into the preliminary segmentation module, which consists of an atrous spatial pyramid pooling (ASPP) layer and a preliminary segmentation module;
[0028] S3.2: The feature map is first input into the ASPP structure to obtain a new feature map z;
[0029] S3.3: The newly obtained feature map z is input into the segmentation module. The segmentation module consists of a 1×1 convolutional layer (Conv1×1), upsampling, and a Softmax activation function. First, the 1×1 convolutional layer (Conv1×1) is used to adjust the number of channels, and the number of channels of the feature map output by the ASPP module is adjusted to the number of classes, such as: background + target classes; secondly, through the bilinear interpolation upsampling method, the feature map with adjusted channels is restored to the same spatial resolution as the input image; finally, the Softmax activation function is used to generate a preliminary segmentation prediction map. The formula for the Softmax activation function to calculate the probability that the i-th pixel belongs to the j-th class is shown in Equation (3):
[0030]
[0031] where z i,j represents the feature value of the i-th pixel in the j-th class in the feature map, C represents the total number of classes, k represents the index value starting from the first class, and exp(·) represents the exponential function.
[0032] S3.4: Extract the edge features from the preliminary segmentation prediction map obtained in step S3.3. Use the Canny edge detection algorithm to extract the edge features of the feature map output by the segmentation module;
[0033] S3.5: Combine the edge features obtained in steps S2.7 and S3.4, and use weighted feature fusion for the edge features of the two steps. Its calculation formula is shown in Equation (4):
[0034] correct_edge(i,j) = β × first_edge(i,j) + (1 - β) × canny_edge(i,j) (4)
[0035] In the formula, correct_edge(i,j) represents the pixel value of the corrected edge feature map at position (i,j), β is the weighting coefficient, first_edge(·) represents the pixel value of the edge feature obtained by the method in step S2.4, and canny_edge(·) represents the pixel value of the edge feature obtained by the method in step S3.4;
[0036] S3.6: For the obtained corrected edge features, splice and fuse them with the multi-scale features obtained by the ASPP module using the Concat operation. Its calculation formula is shown in Equation (5):
[0037] fused_features(i,j) = concat(correct_edge(i,j), aspp_features(i,j)) (5)
[0038] In the formula, fused_features(i,j) represents the pixel value of the fused feature map at position (i,j), concat(·) represents the splicing operation in the channel dimension, correct_edge(i,j) represents the pixel value of the corrected edge feature map at position (i,j), and aspp_features(i,j) is the pixel value of the multi-scale feature at position (i,j);
[0039] S3.7: Output the feature map V that contains more edge information and multi-scale information after fusion features ;
[0040] S4: Use the feature map V featuresInput into the Transformer-based encoder and decoder channels for generating the final detection and segmentation results, which specifically include the following sub-steps:
[0041] S4.1: Flatten the feature V features into a shape suitable for the Transformer channel;
[0042] S4.2: Calculate the positional embedding and add it to the flattened feature V features_flat and input it into the Transformer encoder structure for self-attention, and output the feature V o ;
[0043] S4.3: Introduce a learnable class prototype E cls as a query for the Transformer decoder structure, combine with the output feature V in step S4.2 o for calculation and output the attention map V attention ;
[0044] S4.4: In step S2.5, a higher confidence value is given to the second stage of feature extraction, and in the operation of step S3, edge feature analysis and correction are performed and fused with the preliminary segmentation module. In this step, using a small convolution to fuse image features is omitted, and directly the attention map V output in step S4.3 attention is used to output the final prediction map through per-pixel classification.
[0045] Use the present invention for comparative testing adapted to the unmanned vehicle dataset.
[0046] The present invention has the following beneficial effects:
[0047] 1. Based on the background of lidar mapping, in view of the unsatisfactory mapping of transparent objects encountered by unmanned vehicles, the selected dataset suitable for unmanned vehicles is deleted and added to improve the engineering practice ability of this method. And a vision-based edge feature fusion and correction is proposed to complete the update and correction of edge information, combined with multi-scale features to complete the task of transparent object detection and segmentation, solving the problems of uneven edges, low accuracy and inability to combine with the actual unmanned vehicle mapping work in the existing vision detection and segmentation methods;
[0048] 2. In the edge detection module, in order to better obtain the edge features of transparent objects, edge detection at each stage is completed by assigning confidence values to different stages, and the features obtained at different stages are combined to output the edge features extracted using the backbone network for subsequent edge correction and update processes, enabling this method to focus on the edge features in the initial stage. The mean intersection over union mIoU of the model is 74.43%, improving the robustness of vision detection adapted to unmanned vehicles for transparent objects;
[0049] 3. In the initial segmentation module, the ASPP module is used to capture context information at different scales, enhance the receptive field and semantic information of the feature map, and achieve multi-scale feature extraction. The segmentation module therein is used to receive the data output by the ASPP module to complete the initial segmentation, and a lightweight and mature edge detection algorithm is adopted to extract the initial edge features. The two edge features are weighted and fused to obtain more reliable edge features. Then, the newly obtained edge features are combined with the multi-scale information obtained by the ASPP through the Concat splicing fusion operation, enabling this method to obtain more and more accurate image information, improving the accuracy of pixel filling in the visual detection of the unmanned vehicle, and the pixel accuracy ACC is 96.35%;
[0050] 4. In the Transformer-based encoder and decoder channels, considering that the network has already focused on the shallow features of the high-resolution image, the use of small convolutional layers to continue fusing the aforementioned features is omitted, enabling this method to reduce the computational amount, reduce the operation burden of the unmanned vehicle, and improve the working efficiency on the premise of ensuring stability and accuracy. Brief Description of the Drawings
[0051] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0052] Figure 1 It is the overall structural flowchart of the present invention;
[0053] Figure 2 It is the actual mapping effect diagram when the unmanned vehicle encounters a transparent glass door;
[0054] Figure 3 It is the schematic diagram of the newly added picture data based on the actual working scenario of the unmanned vehicle;
[0055] Figure 4 It is the schematic diagram of the edge detection module structure;
[0056] Figure 5 It is the schematic diagram of the initial segmentation module structure;
[0057] Figure 6 It is the schematic diagram of the Transformer-based encoder-decoder structure;
[0058] Figure 7 It is the comparison result diagram between the present invention and the reference literature. Detailed Embodiments
[0059] To make the above objects, features, and advantages of the present invention more obvious and understandable, a vision detection method for an autonomous vehicle considering the influence of transparent objects, the structural flowchart of which is as follows Figure 1 shown. First, through feature extraction, the features in each stage are first input into the edge detection module to obtain more edge features by increasing the confidence factor. At the same time, the information in the feature extraction stage is input into the preliminary segmentation module, which consists of a dilated convolution with a dilation rate and a segmentation module, and the preliminary segmentation edge information is obtained through feature extraction. Then, new and accurate edge features are obtained through weighted fusion, and feature fusion is performed by combining multi-scale features and input into the Transformer encoder-decoder structure. Finally, the prediction results are output, including the following steps:
[0060] S1: Select a suitable dataset based on the actual application scenario of the autonomous vehicle, and perform deletion, modification, and addition on the selected dataset. Specifically, it includes the following sub-steps:
[0061] S1.1: Considering that the lidar carried by the autonomous vehicle is different, the mapping results will also be different. The most obvious ones are the 2D grid map and the 3D point cloud map. Use a 2D lidar to map in a scene with transparent objects. The result is as follows Figure 2 shown. Figure 2 The left picture in is the working scene of the autonomous vehicle. At this time, the front of the vehicle is facing the glass door. Four arrows in the left picture indicate that the laser beam passes through the glass door. The actual result is as follows Figure 2 shown in the right picture, and the environment map behind the door will be ignored by ignoring the glass door.
[0062] When carrying a 2D lidar, only need to consider whether there are glass material objects in the front. During the process of building a grid map, it is known that its mainstream mapping result is a 2D grid map, also known as a 2D occupancy grid map. There are usually three states to represent the current grid, namely: occupied, unknown, and free. When there is an obstacle within the detection range of the lidar, the grid where it is located will be set to the occupied state, that is, usually black is used to indicate that there is an obstacle in this grid. There are only two states in the known map, occupied and free, and its calculation formula is shown in Equation (1):
[0063]
[0064] In the formula, Odd(s|z) represents the state s of the grid under the condition of the measurement value z, p(s = 0) represents the occupancy rate of the free state, and p(s = 1) represents the occupancy rate of the occupied state. It means that the measurement value model is divided into two states, occupied and free.
[0065] 3D lidar can establish the height information that 2D lidar lacks and can describe the external environment more accurately. At this time, it is necessary to consider the specific shape of glass material objects to complete the calibration of mapping. However, no matter which type of lidar, when encountering transparent objects, especially glass materials, due to the special nature of the materials, the emitted and received light beams of the lidar will be reflected and refracted, resulting in errors in the laser data returned to the lidar. More seriously, the lidar beam will pass through the transparent object and mislabel the location that should be detected as an obstacle as a path that can be normally passed in the map, which will cause serious consequences such as collisions and inability to locate during the subsequent mission execution of the unmanned vehicle;
[0066] S1.2: There are many currently known and publicly available datasets. Considering the similarities and differences between the maps established based on 2D lidar and 3D lidar, in order to achieve better training results, the transparent glass dataset Trans10k-v2 is selected, which is currently the dataset with the largest amount of data and the richest scenes;
[0067] S1.3: Considering the actual application scenario, objects that are rarely or almost never encountered during the mapping work of the unmanned vehicle, such as glasses, glass cups, glass bowls, etc., are deleted from the selected dataset, and 5 class prototypes of common transparent plastic bottles, glass windows, glass doors, glass walls, and transparent cabinets in daily life are retained. Based on the above 5 class prototypes, corresponding data pictures are added. The added data pictures are as Figure 3 shown;
[0068] S2: Introduce an edge detection module, and introduce a confidence factor on the basis of considering that more attention needs to be paid to the edges of transparent objects. Its structure diagram is as Figure 4 shown. Confidence levels are assigned at different stages to obtain stage features that are more conducive to the segmentation of transparent objects. After the edge detection module extracts the edge features of transparent objects, it specifically includes the following sub-steps:
[0069] S2.1: Use ResNet50 based on a convolutional neural network as the backbone network in the feature extraction stage, which optimizes the data propagation in the neural network through the idea of residual learning;
[0070] S2.2: Input the image into the convolutional neural network to extract the features of the image at each stage
[0071] S2.3: The image features extracted at each stage by the backbone network are passed into the edge detection module;
[0072] S2.4: The edge detection module processes image features in four corresponding stages to extract the image features at the corresponding stages. Each of the four stages of feature extraction has its own advantages and disadvantages. Starting from the second stage, a larger range of context information is combined. While the resolution is improved, more detailed information of the image can be retained, and it has the ability to filter image noise, improving the reliability and stability of edge detection. Consider designing the edge detection module in a form similar to an image pyramid. The edge loss is described using Dice Loss, and its calculation formula is shown in Equation (2):
[0073]
[0074] In the formula, y i and represent the label value and the predicted value of pixel i respectively, and N represents the total number of pixels.
[0075] S2.5: Considering that different stages have different edge features, a confidence factor α i is added at each stage. To ensure that the edge loss calculation at each stage does not change the scale and to more conveniently control the input and output of each stage, the confidence factor is normalized to ensure that the confidence factor satisfies where the confidence factor of the second stage is given a higher value than other stages;
[0076] S2.6: The mean of the losses of each stage is used to represent the total loss function L s_edm of the edge detection module, and its calculation formula is shown in Equation (3):
[0077]
[0078] In the formula, n represents the total number of each stage, and α i represents the confidence factor added to each stage, satisfying
[0079] S2.7: The edge feature V 1_features of the feature map is output after passing through the edge detection module;
[0080] S3: A preliminary segmentation module is proposed. The feature map generated by the preliminary segmentation module is subjected to edge feature extraction and analyzed and corrected with the edge feature map generated by the edge detection module to generate a new and more accurate edge feature. Its structural diagram is as Figure 5 shown, including a 1*1 convolutional layer to adjust the number of channels, bilinear interpolation for upsampling and generating a preliminary segmentation map through an activation function, and extracting the edge features of another module through Canny edge detection, specifically including the following sub-steps:
[0081] S3.1: Input the features extracted by the backbone network selected in S2.1 into the preliminary segmentation module, which consists of an Atrous Spatial Pyramid Pooling (ASPP) layer and a segmentation module;
[0082] S3.2: The feature map is first input into the ASPP structure, and atrous convolutions with different dilation rates are used for multi-scale feature extraction to obtain a new feature map z;
[0083] S3.3: To achieve better edge features as described in this paper, the newly obtained feature map z is input into the segmentation module. The preliminary segmentation module consists of a 1×1 convolutional layer (Conv1×1), upsampling, and a Softmax activation function. First, use the 1×1 convolutional layer (Conv1×1) to adjust the number of channels, adjusting the number of channels of the feature map output by the ASPP module to the number of classes, e.g., background + target classes; second, through the bilinear interpolation upsampling method, restore the feature map with adjusted number of channels to the same spatial resolution as the input image; finally, use the Softmax activation function to generate the preliminary segmentation prediction map. For the input feature map z, where z i is the feature vector of the i-th pixel, and the Softmax activation function calculates the probability that the i-th pixel belongs to the j-th class, and its calculation formula is shown in Equation (4):
[0084]
[0085] In the formula, z i,j represents the feature value of the i-th pixel in the j-th class in the feature map, C represents the total number of classes, k represents the index value starting from the first class, and exp(·) represents the exponential function.
[0086] The Softmax activation function converts the feature vector of each pixel into a probability distribution of classes, ensuring that the feature of each pixel is mapped to the range of [0, 1], and the sum of the probabilities of all classes is 1;
[0087] S3.4: To achieve the edge feature fusion mentioned in S3.5, in this step, it is necessary to extract the edge features from the preliminary segmentation prediction map obtained in step S3.3. Considering the requirements of the network structure and computational complexity, in this step, different from the edge detection method introduced in step S2.4, the Canny edge detection algorithm is used to directly extract the edge features from the feature map output by the segmentation module;
[0088] S3.5: Next, combine the edge features obtained in steps S2.7 and S3.4, and perform analysis and correction to generate a more accurate segmentation result. Use weighted feature fusion to combine the edge features of the two steps, and its calculation formula is shown in Equation (5):
[0089] correct_edge(i,j) = β × first_edge(i,j) + (1 - β) × canny_edge(i,j) (5)
[0090] Wherein, correct_edge(i,j) represents the pixel value of the corrected edge feature map at position (i,j), β is the weighting coefficient, first_edge(·) represents the edge feature pixel value obtained by the method of step S2.4, and canny_edge(·) represents the edge feature pixel value obtained by the method of step S3.4;
[0091] S3.6: After obtaining the corrected edge features obtained in step S3.5, fuse them with the multi-scale features obtained by the ASPP module in step S3.2, and splice and fuse them together using the Concat operation. Its calculation formula is shown in formula (6):
[0092] fused_features(i,j) = concat(correct_edge(i,j), aspp_features(i,j)) (6)
[0093] Wherein, fused_features(i,j) represents the pixel value of the fused feature map at position (i,j), concat(·) represents the splicing operation in the channel dimension, correct_edge(i,j) represents the pixel value of the corrected edge feature map at position (i,j), and aspp_features(i,j) is the pixel value of the multi-scale feature at position (i,j);
[0094] S3.7: Output the feature map V that contains more edge information and multi-scale information after fusion features ;
[0095] S4: Input the feature map V output in step S3.7 features into the Transformer-based encoder and decoder channels for generating the final detection and segmentation results. Its structure diagram is as Figure 6 shown, omitting the use of 1*1 convolutional layers to fuse high-resolution shallow features and attention maps, and specifically including the following sub-steps:
[0096] S4.1: Flatten the feature V features into a shape suitable for the Transformer channel;
[0097] S4.2: Calculate the positional embedding and add it to the flattened feature V features_flat and input it into the Transformer encoder structure for self-attention, and output the feature V o ;
[0098] S4.3: Introduce the learnable class prototype E cls As a query based on the Transformer decoder structure, combine the output feature V in step S4.2 o Perform calculations and output the attention map V attention ;
[0099] S4.4: In step S2.5, a higher confidence value is given to the second stage of feature extraction, and in the operation of step S3, edge feature analysis, correction, and fusion are performed with the preliminary segmentation module. In this step, instead of using a small convolution to fuse image features, directly use the attention map V output in step S4.3 attention Output the final prediction map through per-pixel classification, and its calculation formula is as shown in Equation (7):
[0100] Class(i,j) = arg max c P n,c,i,j (7)
[0101] In the formula, P n,c,i,j represents the probability that the nth sample belongs to the cth category at pixel (i,j), Class(i,j) represents the predicted category of the nth sample at pixel (i,j), and arg max c represents the index that can find the maximum probability on the category dimension c;
[0102] Use the present invention to conduct comparative tests on adapting the unmanned vehicle dataset with different methods, and measure by using the mean intersection over union (mIoU) and pixel accuracy (ACC). Among them, mIoU is used to measure the overall prediction effect of the model, and ACC represents the proportion of correctly classified pixels in the total number of image pixels. Its calculation formulas are as shown in Equation (8) and Equation (9):
[0103]
[0104] In the formula, k represents the number of categories in the dataset, i represents the true value of a certain category of pixels, j represents the predicted value of the deep learning network, p ij represents the number of sample pixels that should actually be category i but are predicted as category j, p ii represents the number of correct predictions for a certain category, represents all the prediction results of category i, that is, represents all the prediction results of category j, represents the sum of all correctly predicted categories, represents the sum of all pixels.
[0105] In actual application scenarios, driverless vehicles often encounter the presence of transparent objects. The method proposed in the present invention is tested and compared with Document 1, "Segmenting transparent object in the wild with transformer" and Document 2, "Research and application of semantic segmentation of transparent objects integrating edge detection networks". The test results are as follows Figure 7 shown. The test objects are glass windows and glass doors. These two types of transparent objects are very common in the working environment of driverless vehicles. As can be clearly seen from Figure 7 subfigure a) in, compared with Document 1 and Document 2, the present invention can more effectively filter out the noise data in the surrounding environment, and the edge processing is more regular and smooth. Figure 7 Subfigure b) of shows that the present invention is more accurate in detecting transparent objects, can filter out non-transparent objects such as door handles, and the edge processing is also more regular and smooth. The results show that the present invention improves the stability and accuracy of transparent object detection and can be more practically applied to the working scenario of driverless vehicles.
[0106] At the same time, based on the selected dataset adapted to driverless vehicles, it can be obtained from the data in Table 1 that the mean intersection over union (mIoU) of the present invention is increased by 2.28% and 2.47% respectively compared with Document 1 and Document 2, and the pixel accuracy (ACC) is increased by 2.21% and 2.60% respectively, improving the accuracy and robustness of detecting and segmenting transparent objects.
[0107] Table 1 Test comparison
[0108]
[0109] The above-mentioned specific implementation solutions further illustrate the invention purpose and technical solutions of the present invention. The above embodiments are only used to illustrate the technical solutions of the present invention, rather than limiting the protection scope of the present invention. Those of ordinary skill in the art should understand that any modification and equivalent replacement of the technical solutions of the present invention are included in the protection scope of the present invention.
Claims
1. A visual inspection method for unmanned vehicles considering the influence of transparent objects, characterized in that: The following steps are involved: S1: Delete, modify and add data based on the actual application scenarios of unmanned vehicles; S2: Introduce edge detection module. Considering that more attention should be paid to the introduction of confidence factor on the edge of transparent objects, it includes the following sub-steps: S2.1: ResNet50 based on convolutional neural network as the backbone network; S2.2: Input the image into the convolutional neural network to extract features S2.3: Image features Pass to edge detection module; S2.4: edge detection module extracts image features at the corresponding stage; S2.5: Considering that different stages have different edge features, the confidence factor α is increased at each stage i In order to ensure that the edge loss calculation at each stage does not change the scale, it is more convenient to control the input and output of each stage. The confidence factor is normalized to ensure that the confidence factor satisfies Among them, the confidence factor of the second stage is assigned a higher value than that of other stages; S2.6: The total loss function L of the edge detection module is represented by averaging the sum of the losses of each stage s_edm , and its calculation formula is shown in formula (1): In the formula, n represents the total number of each stage, α i It is represented by the confidence factor added at each stage, and the confidence factor satisfies S3: Propose a preliminary segmentation module, analyze and correct the edge feature map generated by the preliminary segmentation module and the edge feature map generated by the edge detection module, which specifically includes the following sub-steps: S3.1: The features extracted by the backbone network Input to the preliminary segmentation module; S3.2: The feature map is input into the ASPP structure to obtain a new feature map z; S3.3: The newly obtained feature map z is input into the segmentation module. The segmentation module consists of a 1*1 convolution layer (Conv1*1), upsampling and Softmax activation function. First, the 1*1 convolution layer (Conv1*1) is used to adjust the number of channels. The number of channels of the feature map output by the ASPP module is adjusted to the number of categories, including background + target categories. Secondly, the feature map after adjusting the number of channels is restored to the same spatial resolution as the input image through the bilinear interpolation upsampling method. Finally, the Softmax activation function is used to generate a preliminary segmentation prediction map. S3.4: Extract edge features from the preliminary segmentation prediction image obtained in step S3.3 using the Canny edge detection algorithm; S3.5: Combine the edge features obtained in step S2.7 and step S3.4, and use weighted features to fuse the edge features of the two steps. The calculation formula is shown in formula (2): correct_edge(i,j)=β×first_edge(i,j)+(1-β)×canny_edge(i,j) (2) Wherein, correct_edge(i,j) represents the pixel value of the corrected edge feature map at position (i,j), β is the weighting coefficient, first_edge(·) represents the edge feature pixel value obtained by the method in step S2.4, and canny_edge(·) represents the edge feature pixel value obtained by the method in step S3.4; S3.6: The corrected edge features and multi-scale features are concatenated and fused using the Concat operation; S4: Feature map V features Input to the Transformer-based encoder and decoder channels to generate the final detection segmentation results, which specifically includes the following sub-steps: S4.1: Set feature V features Flatten; S4.2: Compute the position embedding and add it to the flattened features V features_flat Input to the encoder; S4.3: Decoder outputs attention map; S4.4: Directly convert the attention map V of S4.3 attention The final prediction map is output through pixel-by-pixel classification.
Citation Information
Patent Citations
Multi-sensor fusion robot mapping method in glass environment
CN117451029A
Lightweight image segmentation method based on multi-level feature parallel interactive fusion
CN117197457A
Battery defect detection method based on richer convolutional feature edge detection network
CN117808748A