Transparent object detection method, device and equipment
By introducing semantic-boundary co-evolution module into neural networks, fusion and weighting semantic features and boundary features, the accuracy problem of transparent object detection in complex scenarios in the prior art is solved, and higher detection accuracy and robustness are achieved.
Patent Information
- Application Number
- CN202510277661.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-10
- Publication Date
- 2025-06-06
- Estimated Expiration
- 2045-03-10
AI Technical Summary
Existing transparent object detection methods are difficult to accurately identify transparent objects in complex scenarios, especially because the network model's receptive field limitation and the semantic information and boundary information are not fully integrated, resulting in unsmooth edges of the segmentation results and low feature matching rate.
Building a neural network, including parallel residual network modules and boundary detection modules, uses the semantic-boundary coevolution module to fusion and cross-attention mechanisms to enhance the coevolution of semantic information and boundary information.
Through the design of the semantic-boundary co-evolution module, the detailed semantic information of transparent object boundaries can be captured more accurately, the model's understanding of transparent object boundaries can be improved, and detection accuracy and robustness can be improved.
Smart Images

Figure CN120107597A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of transparent object detection, and in particular to a transparent object detection method, device and equipment. Background Art
[0002] In the field of computer vision, transparent object detection is a challenging problem. Service robots need to accurately identify transparent objects, such as glass cups, transparent packaging bags or plastic boxes, in order to safely grasp and operate them. Due to their physical properties, transparent objects reflect and refract light, resulting in unclear boundaries in the image and lack of sufficient texture information, making it difficult for traditional image processing methods to effectively identify and locate them. With the development of deep learning technology, deep learning-based methods have been widely used in the detection of transparent objects due to their high efficiency in image recognition and processing.
[0003] The current transparent object detection methods mainly use convolutional neural networks to automatically learn the features of transparent objects in images. This will result in limited receptive fields of network models, making it difficult to cope with the recognition and detection of transparent objects in more complex scenes. The paper "Segmenting transparent object in the wild with transformer" proposes a visual segmentation method Trans2Seg, which first applies Transformer to transparent object segmentation scenes. This method first feeds the image into an encoder based on a convolutional neural network, extracts features, and then feeds them into the Transformer module for self-attention operation, and then feeds the enhanced features into the decoder to obtain the final segmentation result. However, this method does not make full use of the boundary information of transparent objects, resulting in uneven edges of the segmentation results and low feature matching rates. The paper "A visual detection method for unmanned vehicles considering the influence of transparent objects" also pays attention to this point. After first extracting features from the input image, it passes through the boundary detection module to extract the edge information of the transparent object. Then, the edge features are used to perform feature fusion correction with the results of the preliminary segmentation module, and finally the prediction results are obtained through the encoder-decoder mechanism based on Transformer.
[0004] However, the above method only fuses and corrects the semantic information of transparent objects with their boundary information in the last step, ignoring the potential mutual reinforcement relationship between the two, which can easily lead to the model's misunderstanding of the boundaries of transparent objects, resulting in the need to improve the accuracy of transparent object detection. Summary of the invention
[0005] Based on this, in order to solve the technical problems in the prior art, the present invention provides a transparent object detection method, device and equipment.
[0006] The present invention provides a transparent object detection method, comprising:
[0007] Constructing a neural network, the neural network comprising a parallel residual network module and a boundary detection module, a semantic-boundary co-evolution module connected to an output end of the residual network module and an output end of the boundary detection module, and a decoding module connected to an output end of the semantic-boundary co-evolution module; wherein the semantic-boundary co-evolution module comprises a fusion layer and a cross-attention mechanism layer connected in sequence;
[0008] Collect images containing transparent objects to construct a data set, use the data set to train a neural network, and obtain a transparent object detection model;
[0009] An image to be detected containing a transparent object is input into the detection model, and the semantic features of the transparent object in the image to be detected are extracted through the residual network module; the boundary features of the transparent object in the image to be detected are extracted through the edge detection algorithm in the boundary detection module; the semantic features and the boundary features are fused through the fusion layer in the semantic-boundary co-evolution module to obtain the semantic-boundary fusion features of the transparent object; the semantic features and the boundary features are mapped to semantic queries and boundary queries respectively through the cross-attention mechanism layer in the semantic-boundary co-evolution module, and the semantic-boundary fusion features are mapped to semantic keys and semantic values corresponding to the semantic queries, and boundary keys and boundary values corresponding to the boundary queries, and based on the obtained queries, keys and values, the semantic features and boundary features are weighted by the cross-attention mechanism respectively to obtain weighted semantic features and weighted boundary features; the weighted semantic features and weighted boundary features are decoded through the decoding module to obtain the transparent object in the image to be detected.
[0010] Furthermore, the mapping of semantic features and boundary features into semantic queries and boundary queries respectively, and mapping semantic-boundary fusion features into semantic keys and semantic values corresponding to the semantic queries, and boundary keys and boundary values corresponding to the boundary queries, specifically includes:
[0011] The shapes are uniformly transformed into C×H×W semantic features, boundary features, and semantic-boundary fusion features, which are flattened into 2D semantic features F s ', 2D boundary feature F e ' and 2D semantic-boundary fusion feature F c Among them, F s '、F e ' and F c '∈R N×C , N = H × W, C is the number of channels;
[0012] The 2D semantic features F are respectively mapped by the pre-trained semantic mapping matrix s ' and 2D semantic-boundary fusion feature F c Mapping is performed to obtain the semantic query Q s, semantic key K s and semantic value V s :
[0013]
[0014] in, and is the pre-trained semantic mapping matrix;
[0015] The 2D boundary features F are respectively mapped by the pre-trained boundary mapping matrix e ' and 2D semantic-boundary fusion feature F c 'Map and obtain boundary query Q e , boundary key K e and the boundary value V e :
[0016]
[0017] in, and is the pre-trained boundary mapping matrix.
[0018] Furthermore, the boundary features are weighted by a cross-attention mechanism, specifically including:
[0019] The boundary query Q e , boundary key K e and the boundary value V e Reshape to H×W×C and obtain the reshaped boundary query f Q , reshape the boundary key f K and reshape the boundary value f V ;
[0020] In the H×W dimension, the boundary query f is reshaped Q and reshape the boundary key f K Divide into n×n regions, and perform average pooling operation on each region to obtain the first weight representation matrix Q and the second weight representation matrix K; based on the first weight representation matrix Q and the second weight representation matrix K, obtain the regional matching score matrix G:
[0021] G=Q(K) T
[0022] Among them, the superscript T represents the transpose operation;
[0023] Select the top k scores in each row of the region matching score matrix G, integrate the indexes of the k scores to obtain the matrix S; and take out the reshape boundary key f according to the index K and reshape the boundary value f V The corresponding area in:
[0024] f′ K =Gather(f K , S)
[0025] f′ V =Gather(f V ,S)
[0026] Gather(·) is the torch.gather operation, which is used to extract elements from the input tensor according to the index in the specified dimension and construct a new tensor; f' K and f' V are respectively Q The most correlated key and value features;
[0027] F' K and f' V After changing to 2D shape, calculate the spatial attention A e :
[0028]
[0029] Among them, K' e and V′ e f' K and f' V 2D flat key features and 2D flat value features obtained after transformation into 2D shape; d = C / h, h is the number of attention heads;
[0030] The 2D boundary feature F is transformed into e 'Add spatial attention A e The calculation result of weighted boundary feature f is obtained ei :
[0031]
[0032] in, It is the element-by-element addition operation in the residual connection.
[0033] Furthermore, the semantic features are weighted by a cross-attention mechanism, specifically including:
[0034] Use the cross attention mechanism to s Weighted, get the attention score A s :
[0035]
[0036] Where d = C / h, h is the number of attention heads; the superscript T represents the transposition operation;
[0037] The 2D semantic features F are transformed intos 'Add attention score A s The calculation result of , obtains the weighted semantic feature f si :
[0038]
[0039] in, It is the element-by-element addition operation in the residual connection.
[0040] Furthermore, the output end of the cross attention mechanism layer in the semantic-boundary co-evolution module is connected to a Mix-FFN network layer; the weighted semantic features and weighted boundary features output by the cross attention mechanism layer are enhanced by the Mix-FFN network layer, specifically including:
[0041]
[0042] Among them, MPL() is the feedforward operation in the Mix-FFN network, Conv 3×3 () is the convolution operation in the Mix-FFN network; GeLU() is the activation function operation in the Mix-FFN network; f ei is the weighted boundary feature, f eo is the enhanced boundary feature; f si is the weighted semantic feature, f so It is to enhance semantic features.
[0043] Furthermore, the residual network module is formed by a ResNeXt-101 network; a multi-scale feature F of a transparent object in the image to be detected is extracted through multiple residual convolutional layers of the ResNeXt-101 network. i , i∈{0, 1, 2, 3, 4, 5}, with the feature F that has the most semantic information 5 As the semantic feature of transparent objects in the image to be detected.
[0044] Furthermore, the output end of the boundary detection module and the output end of the residual network module are connected to a boundary refinement module, and the boundary refinement module includes an aggregation layer, a residual layer, a spatial attention mechanism layer and a fusion layer; the boundary features output by the boundary detection module are refined by the boundary refinement module, specifically including
[0045] The feature F output by the ResNeXt-101 network is aggregated through the aggregation layer 0 , Feature F 1 , Feature F 2 , Feature F 3 After upsampling to the same dimension, aggregate operations are performed with boundary features:
[0046] f i =Upsample(Fi ),i=0,1,2,3
[0047] sum =[f 0 ,f 1 , f 2 , f 3 , f m ]
[0048] Among them, Upsample() is the upsampling operation, f m is the boundary feature extracted by the Canny edge detection algorithm in the boundary detection module; [] is the feature aggregation operation;
[0049] The residual layer is used to enhance the upsampled feature map f with the most boundary features and the least semantic information. 0 :
[0050]
[0051] Among them, ReLU() is the activation function, BN() is the batch normalization operation; Conv 3×3 () is the convolution operation, f' 0 It is the enhanced boundary feature; is the element-by-element addition operation in the residual connection; f' 0 is the initial enhanced boundary feature f' 0 ;
[0052] The initial enhanced boundary feature f' is obtained through the spatial attention mechanism layer 0 Enhanced again:
[0053]
[0054] α j =Sigmoid([f' 0 , f j ]) j∈{1,2,sum}
[0055] Among them, Sigmoid() is the sigmoid function used for normalization, α j is the attention map obtained from the j-th upsampled feature map, It is an element-by-element multiplication operation; is the jth boundary feature after further enhancement;
[0056] The fourth boundary feature output by the spatial attention mechanism layer is merged into The feature F with the most semantic information output by the ResNeXt-101 network 5 Perform feature aggregation to obtain refined boundary features
[0057]
[0058] Furthermore, the semantic-boundary co-evolution module includes a first semantic-boundary co-evolution module, a second semantic-boundary co-evolution module, a third semantic-boundary co-evolution module and a fourth semantic-boundary co-evolution module connected in sequence; the first, second, third and fourth semantic-boundary co-evolution modules enhance the semantic features and boundary features in the feature transfer process by connecting in series.
[0059] The present invention provides a transparent object detection device, comprising:
[0060] A model building module, used to build a neural network, wherein the neural network includes a parallel residual network module and a boundary detection module, a semantic-boundary co-evolution module connected to the output end of the residual network module and the output end of the boundary detection module, and a decoding module connected to the output end of the semantic-boundary co-evolution module; wherein the semantic-boundary co-evolution module includes a fusion layer and a cross-attention mechanism layer connected in sequence;
[0061] A model training module is used to collect images containing transparent objects to construct a data set, and use the data set to train a neural network to obtain a transparent object detection model;
[0062] The detection module is used to input an image to be detected containing a transparent object into a detection model, extract the semantic features of the transparent object in the image to be detected through a residual network module; extract the boundary features of the transparent object in the image to be detected through an edge detection algorithm in a boundary detection module; fuse the semantic features and the boundary features through a fusion layer in a semantic-boundary co-evolution module to obtain the semantic-boundary fusion features of the transparent object; map the semantic features and the boundary features to semantic queries and boundary queries respectively through a cross-attention mechanism layer in the semantic-boundary co-evolution module, map the semantic-boundary fusion features to semantic keys and semantic values corresponding to the semantic queries, and boundary keys and boundary values corresponding to the boundary queries, and based on the obtained queries, keys and values, respectively weight the semantic features and the boundary features through a cross-attention mechanism to obtain weighted semantic features and weighted boundary features; decode the weighted semantic features and the weighted boundary features through a decoding module to obtain the transparent object in the image to be detected.
[0063] The present invention provides a computer device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the transparent object detection method when executing the program.
[0064] At least one of the above technical solutions adopted by the present invention can achieve the following beneficial effects:
[0065] After the present invention fuses the semantic features and the boundary features, the query obtained based on the semantic features is matched with the key value obtained based on the fused features, so that the semantic features can obtain the contextual information related to the boundary features in the fused features, thereby being able to capture the detailed semantic information of the boundary of the transparent object; similarly, the query obtained based on the boundary features is matched with the key value obtained based on the fused features, so that the boundary information can obtain more contextual support from the semantic features, making the pixels around the boundary features more continuous. In the attention mechanism, through the attention weighting of query-key-value, the semantic information and the boundary information can enhance each other, promoting the co-evolution of the semantic features and the boundary features, that is: the semantic features can enhance their semantic expression through the boundary information, and the boundary features can also locate and describe the boundaries more accurately with the help of the semantic information, thereby enabling the semantic information and the boundary information to enhance the model's understanding of the boundary of the transparent object under the action of each other, making the final transparent object features more comprehensive and accurate, thereby obtaining more accurate detection results. BRIEF DESCRIPTION OF THE DRAWINGS
[0066] The drawings described herein are used to provide a further understanding of the present invention and constitute a part of the present invention. The exemplary embodiments of the present invention and their descriptions are used to explain the present invention and do not constitute an improper limitation of the present invention. In the drawings:
[0067] Figure 1 A schematic diagram of the transparent object detection method provided by the present invention;
[0068] Figure 2 A schematic diagram of a neural network framework is provided for the present invention;
[0069] Figure 3 A schematic diagram of a boundary refinement module provided by the present invention;
[0070] Figure 4 A schematic diagram of a semantic-boundary co-evolution module based on a cross-attention mechanism provided by the present invention;
[0071] Figure 5 A schematic diagram of transparent object detection results provided by the present invention;
[0072] Figure 6 This is a schematic diagram of the transparent object detection results provided by the present invention. DETAILED DESCRIPTION
[0073] In order to make the purpose, technical solution and advantages of the present invention clearer, the technical solution of the present invention will be clearly and completely described below in combination with the specific embodiments of the present invention and the corresponding drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments in the specification, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0074] Existing methods do not take into account that the semantic information and boundary information of image features can enhance each other when performing transparent object detection. Using boundary features to strengthen semantic features can help semantic features have more detailed information, and using semantic features to strengthen boundary features can make the pixels around boundary features more continuous, while improving the model's ability to distinguish foreground from background. Therefore, the present invention proposes to co-evolve boundary features and semantic features to make the segmentation results smoother, more complete, and more robust.
[0075] In view of the shortcomings of the existing methods, the present invention proposes a transparent object detection method based on the co-evolution of semantic and boundary features, focusing on solving the problem that transparent objects are mixed with the background in complex scenes and are difficult to segment completely. The present invention uses Transformer to expand the receptive field of the model and uses the cross-attention mechanism to promote the co-evolution of semantic features and boundary features. First, a convolutional neural network is used to extract the high-dimensional features of the image, and the shallow features are fused to form a boundary feature stream, and the deep features are fused to form a semantic feature stream. Then, a semantic boundary co-evolution module is used to enhance the semantic features and boundary features. In order to reduce the noise of the boundary feature stream, a boundary refinement module based on multi-scale feature fusion is introduced to refine the boundary features. Finally, the result is input into the decoder to generate the segmentation result. The present invention proposes to co-evolve the boundary features with the semantic features, so that the segmentation result is smoother, more complete, and more robust.
[0076] Example 1
[0077] Figure 1 The process of the transparent object detection method of this embodiment is shown. Figure 1 The method is described in detail and specifically comprises the following steps:
[0078] S1: Construct a neural network, the neural network includes a parallel residual network module and a boundary detection module, a semantic-boundary co-evolution module connected to the output of the residual network module and the output of the boundary detection module, and a decoding module connected to the output of the semantic-boundary co-evolution module; wherein the semantic-boundary co-evolution module includes a fusion layer and a cross-attention mechanism layer connected in sequence. The input of the first semantic-boundary co-evolution module is simultaneously connected to the output of the residual network module and the output of the boundary refinement module, and the features output by the ResNeXt-101 network and the refined boundary features output by the boundary refinement module are used as input; the output of the fourth semantic-boundary co-evolution module is used as the input of the decoding module.
[0079] The constructed neural network structure is as follows Figure 2 As shown, it includes a residual network module formed by a ResNeXt-101 network, a boundary detection module formed by a Canny edge detection algorithm, a boundary refinement module formed by an aggregation layer, a residual layer, a spatial attention mechanism layer, and a fusion layer, 4 semantic-boundary co-evolution modules connected in series, and a decoding module for transparent object boundary detection. Among them, the input end of the boundary refinement module is connected to the output end of the residual network module and the output end of the boundary detection module, and the input end of the first semantic-boundary co-evolution module is connected to the output end of the residual network module and the output end of the boundary refinement module. The neural network adopts a customized network framework based on feature fusion and evolution. By pre-training on the ImageNet dataset, a large number of sample set features are learned to complete multi-scale image feature extraction, and high-level image features can be more accurately extracted in more generalized scenarios. The working process of the neural network is:
[0080] (1) An input image of size 3×H×W is passed through a replaceable backbone feature extraction network to obtain a feature map of size C×H×W containing high-level information, where C is the number of feature map channels. The replaceable backbone feature extraction network can be any backbone feature extraction network based on a deep neural network. Deep neural networks have made great progress in image information extraction. The present invention uses the ResNeXt-101 network. Multi-scale features F are extracted from the shallow features to the deep features of the network. i , i∈{0, 1, 2, 3, 4, 5}, used for multi-scale feature fusion. ResNeXt-101 is an improved version of the neural network based on the residual network (ResNet) architecture, which performs well in image classification and other computer vision tasks. ResNeXt introduces the concept of "grouped convolution" and improves performance by increasing the width of the network instead of the depth based on the original ResNet network. The advantage of this structure is that it reduces the increase in the number of parameters while improving the network's expressiveness and computational efficiency.
[0081] (2) The multi-layer features extracted from the backbone feature extraction network and the image boundary f m Send it to the boundary refinement module to obtain the initialization feature F of the boundary feature flow 1 e .
[0082] (3) F 1 e With f 5 They are used as the input features of the boundary branch and semantic branch of the first semantic-boundary co-evolution module, respectively, and are refined to obtain the refined boundary feature stream and semantic feature stream, which are then used as the feature input of the next semantic-boundary co-evolution module. After the feature refinement of four semantic-boundary co-evolution modules, the final semantic feature stream is obtained, and after passing through the CNN-based decoder, the final result is obtained.
[0083] The structure of each module in the neural network is described in detail below.
[0084] Residual network module.
[0085] It is formed by the ResNeXt-101 network; the multi-scale features F of transparent objects in the image to be detected are extracted through multiple residual convolutional layers of the ResNeXt-101 network. i , i∈{0, 1, 2, 3, 4, 5}, with the feature F that has the most semantic information 5 As the semantic feature of transparent objects in the image to be detected.
[0086] Boundary refinement module.
[0087] like Figure 3 As shown, the present invention adopts a customized boundary refinement module based on multi-scale feature fusion, uses multi-scale features extracted from the backbone feature extraction network to perform aggregation operations and progressive refinement operations, generates boundary refinement features of transparent objects, and reduces background noise. Considering that deep features have rich high-level semantic features and shallow features have rich fine-grained features, aggregating features at different levels can make full use of context information. The specific approach is as follows:
[0088] The multi-scale features output by the ResNeXt-101 network are upsampled to the same dimension through the aggregation layer to facilitate the aggregation operation:
[0089] f i =Upsample(F i ),i=0,1,2,3.
[0090] Upsample is the upsampling operation. Then aggregate the multi-scale features:
[0091] f sum=[f 0 , f 1 , f 2 ,f 3 , f m ]
[0092] Among them, Upsample() is the upsampling operation, f m It is the boundary feature extracted by the Canny edge detection algorithm in the boundary detection module; [] is the feature aggregation operation.
[0093] By using spatial attention, the deep features are used to continuously refine the shallow features, remove the noise of the shallow features, and continuously refine the pixels around the boundaries of transparent objects. The specific steps are as follows:
[0094] Use residual convolution module to strengthen shallow features f 0 :
[0095]
[0096] Among them, ReLU() is the activation function, BN() is the batch normalization operation; Conv 3×3 () is the convolution operation, f' 0 It is the enhanced boundary feature; is the element-by-element addition operation in the residual connection; f' 0 is the initial enhanced boundary feature f' 0 The feature representation can be further enhanced through the residual module without discarding the original feature details.
[0097] Use f in sequence 1 、f 2 、f sum Strengthen, use spatial attention operation to strengthen boundary features. First generate the attention map α:
[0098] α j =Simgoid([f' 0 , f j ]) j∈{1,2,sum};
[0099] Among them, Sigmoid() is the sigmoid function used for normalization, α j is the attention map obtained from the jth upsampled feature map. Then the attention map α is used to strengthen f' 0 :
[0100]
[0101] in, is element-wise multiplication, is the jth boundary feature after further enhancement. 1、f 2 、f sum After repeating the above operation, we get
[0102] The fourth boundary feature output by the spatial attention mechanism layer is merged into The feature F with the most semantic information output by the ResNeXt-101 network 5 Perform feature aggregation to obtain refined boundary features
[0103]
[0104] This feature will be used as the initialization feature of the boundary feature stream and sent to the first semantic-boundary co-evolution module for enhancement.
[0105] Semantic-boundary co-evolution module.
[0106] The structure of the semantic-boundary co-evolution module is as follows Figure 4 As shown. The boundary feature flow initial feature F 1 e (i.e., refined boundary features) and the initial features F of the semantic feature stream 1 s (i.e. F 5 ) as the input of the first semantic-boundary co-evolution module. For each semantic-boundary co-evolution module, firstly, the boundary feature flow F e With the semantic feature flow F s Perform feature aggregation to obtain F c :
[0107] F c =[F s ,F e ].
[0108] F e , F s , F c After the shape is unified into C×H×W, it is serialized into a flattened 2D feature F s ',F e ',F c '∈Ρ N×C , N = H × W, C is the number of channels.
[0109] The semantic-boundary co-evolution module is divided into a boundary branch and a semantic branch. The input of the boundary branch is the boundary feature flow F e , the input of the semantic branch is the semantic feature stream F s . Use the mapping matrix to map the 2D features. For the boundary branch, we have:
[0110] Q e =Fe 'W e Q ,K e =F c 'W e K ,V e =F c 'W e V .
[0111] and is the pre-trained semantic mapping matrix. Similarly, for the semantic branch:
[0112] Q s =F s 'W s Q ,K s =F c 'W s K ,V s =F c 'W s V .
[0113] in, and is the pre-trained boundary mapping matrix.
[0114] For the semantic branch, a cross-attention mechanism is used:
[0115]
[0116] d = C / h, where h is the number of attention heads. Use residual connections and add F s ':
[0117]
[0118] Use Mix-FFN in "SegFormer: Simple and Efficient Design for Semantic Segmentation with Transformers" to provide feature position information, as follows:
[0119]
[0120] Get f soAfter that, it is restored to its original shape to obtain the enhanced semantic features, which are used as the input of the semantic feature stream of the next semantic-boundary co-evolution module. Mix-FFN no longer uses the position encoding in ViT to provide position information. Position encoding is not necessary for semantic segmentation. Instead, 3×3 convolution is used in the forward network FFN to provide position information.
[0121] For boundary branches, customize the region dynamic selection mechanism.
[0122] The boundary query Q e , boundary key K e and the boundary value V e The dimensions are reshaped to H×W×C, denoted by f Q , f K , f V Then it is divided into n×n regions in the H×W dimension, and Q e and K e After the division operation, each region is subjected to an average pooling operation to obtain the first weight representation matrix Q and the second weight representation matrix K. Finally, the region matching score matrix is calculated:
[0123] G=Q(K) T .
[0124] Select the k highest scores in each row of G, record the index to get the matrix S, and then take out V according to the index e With K e The corresponding area in:
[0125] f′ K =Gather(f K ,S),f′ V =Gather(f V ,S).
[0126] Where Gather(·) is the torch.gather operation. K and f' V The 2D shape is recorded as K' e , V' e , which is convenient for subsequent attention operations.
[0127] Calculate spatial attention as follows:
[0128]
[0129] Then the original boundary feature flow F e 'join in:
[0130]
[0131] Same as the semantic branch, the final boundary branch output is obtained:
[0132]
[0133] f eo It is the enhanced boundary feature.
[0134] In summary, the workflow of the entire neural network is as follows: the input image of size 3×H×W is sent to the ResNeXt-101 network to extract multi-scale features F 0 ,F 1 ,F 2 ,F 3 ,F 4 ,F 5 . Using semantically rich F 5 As the initial feature F of the semantic feature stream 1 s Four cascaded semantic boundary co-evolution modules (SBMA) are used to strengthen the two feature streams. The input of each SBMA module is the semantic feature stream F i s and boundary characteristic flow F i e , where i = 1, 2, 3, 4. The output of the previous SBMA is the input of the next SBMA.
[0135] For the boundary feature flow F i e , the expression is as follows:
[0136]
[0137] Where MSBR refers to the boundary refinement module, f m are image boundaries extracted using a boundary detection algorithm. is the boundary branch of the i-th SBMA, [] is the feature aggregation operation, as follows:
[0138] [a,b]=Conv 3×3 (Concat(a,b));
[0139] Conv 3×3 () refers to the convolution operation using a 3×3 convolution kernel, and Concat() refers to the concatenation operation based on the channel dimension.
[0140] For the semantic feature flow F i s , the expression is as follows:
[0141]
[0142] F isum The is the sum of boundary flow features and semantic flow features:
[0143]
[0144] in It is element-by-element addition. After passing through the fourth SBMA, the semantic feature stream is passed through the decoder to generate the final result. In view of the problems of limited receptive field and incomplete segmentation of transparent objects in complex scenes in existing methods, in order to effectively extract the blurred boundaries of transparent objects, a boundary refinement module based on multi-scale feature fusion is designed in the method proposed in the present invention. Utilizing the characteristics of features at different levels, the boundary information of transparent objects is refined, and background noise is removed at the same time. In order to enhance the integrity of transparent object segmentation in complex scenes, a semantic-boundary co-evolution module based on the cross-attention mechanism is designed. It is used to realize the interaction between semantic information and images, and works simultaneously with the designed regional dynamic selection mechanism to realize the co-evolution of semantic information and boundary information.
[0145] S2: Collect images containing transparent objects to build a data set, use the data set to train a neural network, and obtain a transparent object detection model.
[0146] The customized multi-scale feature loss function and the final loss function specifically include:
[0147] (1) Calculate the cross entropy loss L between the output of the semantic branch of each semantic-boundary co-evolution module and the true label s .
[0148] (2) Calculate the cross entropy loss L between the output of the boundary branch of each semantic-boundary co-evolution module and the true boundary label b , the true boundary label is 8 pixel values selected from the boundary of the true label as the boundary true value.
[0149] (3) Calculate the cross entropy loss between the final result and the true label The final loss function is as follows:
[0150]
[0151] where λ s =0.01,λ b =0.25.
[0152] The experimental results show that this method can improve the segmentation effect of transparent objects to a certain extent, and its performance is better than the most advanced methods. The following is an analysis of the experimental results:
[0153] The experimental results are analyzed using the indicators mIoU and F β, mAE and mBER. A total of three comparative experiments were conducted on three datasets: GSD, HSO and GDD datasets.
[0154] The comparison methods include MirrorNet, GDNet, TransLab, Trans2Seg, EBLNet, GSDNet, PGSNet, RFENet, etc. The test results on GSD, GDD and HSO datasets are shown in Table 1, Table 2 and Table 3 respectively.
[0155] Table 1 Test results on the GSD dataset
[0156]
[0157]
[0158] Table 2 Test results on the GDD dataset
[0159] Method QUR <![CDATA[F β ]]> oeI mBER MirrorNet 85.07 0.866 0.083 7.67 GDNet 87.63 0.898 0.063 5.62 TransLab 81.64 0.849 0.097 9.70 Trans2Seg 84.41 0.872 0.078 7.36 GSDNet 88.07 0.932 0.059 5.71 PGSNet 87.81 0.901 0.062 5.56 GDNet-B 87.83 0.939 0.061 5.52 RFENet 87.84 0.940 0.063 6.37 GlassFormer(Ours) 88.60 0.943 0.056 5.51
[0160] Table 3 Test results on the HSO dataset
[0161]
[0162]
[0163] Experimental results show that the present invention is ahead of existing methods on a large number of benchmark datasets for transparent object detection. Figure 5 and Figure 6 , Figure 5 The first column is the original image including transparent objects, and the second to seventh columns are the detection results obtained using the Trang2Seg model, the GDNet model, the EBLNet model, the GSDNet model, the RFENet model, the model of the present invention, and the GT model respectively; Figure 6 The first column is the original image including the transparent object, and the second to fourth columns are the detection results obtained using the TransLab model, the EBLNet model, the model of the present invention, and the GT model, respectively. It can be found that in many complex scenes, such as when the background and the boundary of the transparent object are mixed together, and there are similar transparent objects in the background, the segmentation effect and robustness of the present invention have been significantly improved. The reason is that the present invention has a global receptive field, and adaptively selects local features and global features to enhance the boundary features and semantic features. The present invention makes full use of the co-evolution of local fine-grained features and high-level semantic features, and enables the network to dig deeper differences between transparent objects and background objects without losing local details, so that the foreground and background can be better distinguished, and the robustness is stronger.
[0164] S3: Input the image to be detected containing transparent objects into the detection model, and extract the semantic features of the transparent objects in the image to be detected through the residual network module; extract the boundary features of the transparent objects in the image to be detected through the edge detection algorithm in the boundary detection module; fuse the semantic features and boundary features through the fusion layer in the semantic-boundary co-evolution module to obtain the semantic-boundary fusion features of the transparent objects; map the semantic features and boundary features to semantic queries and boundary queries respectively through the cross-attention mechanism layer in the semantic-boundary co-evolution module, map the semantic-boundary fusion features to the semantic keys and semantic values corresponding to the semantic queries, and the boundary keys and boundary values corresponding to the boundary queries, and based on the obtained queries, keys and values, cross-attention mechanism weights are performed on the semantic features and boundary features respectively to obtain weighted semantic features and weighted boundary features; decode the weighted semantic features and weighted boundary features through the decoding module to obtain the transparent objects in the image to be detected.
[0165] The present invention discloses a transparent object detection method based on the co-evolution of semantic and boundary features, which further promotes effective context and boundary learning in transparent object images. The method comprises: a novel learning framework based on the co-evolution of semantic and boundary features based on Transformer is proposed, which adopts a boundary-semantic dual-stream structure and utilizes the complementary and mutually enhanced features of the two. To achieve this purpose, a semantic-boundary co-evolution module based on a cross-attention mechanism is proposed, which utilizes the cross-attention mechanism to realize the interaction and fusion of boundary features and semantic features. In order to effectively utilize local features and global features, a regional dynamic selection mechanism is proposed, so that the boundary feature stream pays more attention to the boundary part of the transparent object, mines more detailed features, and effectively improves the integrity and smoothness of the boundary of the transparent object. In order to obtain purer boundary features, a boundary refinement module based on multi-scale feature fusion is proposed, which utilizes the multi-scale features extracted from the backbone network and the image boundary information, and utilizes the strategy of feature aggregation and feature refinement to continuously refine the boundary information. Experiments have proved that the method can effectively and perfectly detect transparent objects in complex scenes, and has strong robustness.
[0166] The above is a transparent object detection method provided by one or more embodiments of the present invention. Based on the same idea, the present invention also provides a corresponding transparent object detection device, including:
[0167] A model building module is used to build a neural network, which includes a parallel residual network module and a boundary detection module, a semantic-boundary co-evolution module connected to the output end of the residual network module and the output end of the boundary detection module, and a decoding module connected to the output end of the semantic-boundary co-evolution module; wherein the semantic-boundary co-evolution module includes a fusion layer and a cross-attention mechanism layer connected in sequence.
[0168] The model training module is used to collect images containing transparent objects to build a data set, use the data set to train the neural network, and obtain a transparent object detection model.
[0169] The detection module is used to input an image to be detected containing a transparent object into a detection model, extract the semantic features of the transparent object in the image to be detected through a residual network module; extract the boundary features of the transparent object in the image to be detected through an edge detection algorithm in a boundary detection module; fuse the semantic features and the boundary features through a fusion layer in a semantic-boundary co-evolution module to obtain the semantic-boundary fusion features of the transparent object; map the semantic features and the boundary features to semantic queries and boundary queries respectively through a cross-attention mechanism layer in the semantic-boundary co-evolution module, map the semantic-boundary fusion features to semantic keys and semantic values corresponding to the semantic queries, and boundary keys and boundary values corresponding to the boundary queries, and based on the obtained queries, keys and values, respectively weight the semantic features and the boundary features through a cross-attention mechanism to obtain weighted semantic features and weighted boundary features; decode the weighted semantic features and the weighted boundary features through a decoding module to obtain the transparent object in the image to be detected.
[0170] For the specific definition of the transparent object detection device, please refer to the definition of the transparent object detection method above, which will not be repeated here. Each module in the above transparent object detection device can be implemented in whole or in part by software, hardware and a combination thereof. The above modules can be embedded in or independent of the processor in the computer device in the form of hardware, or can be stored in the memory of the computer device in the form of software, so that the processor can call and execute the operations corresponding to the above modules.
[0171] The present invention also provides a computer device structure. At the hardware level, the computer device includes a processor, an internal bus, a network interface, a memory, and a non-volatile memory. Of course, it may also include other hardware required for the business. The processor reads the corresponding computer program from the non-volatile memory into the memory and then runs it to achieve the above Figure 1 Provided transparent object detection method.
[0172] Those skilled in the art can understand that all or part of the processes in the above-mentioned embodiment methods can be completed by instructing the relevant hardware through a computer program, and the computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above-mentioned methods. Among them, any reference to memory, storage, database or other media used in the embodiments provided by the present invention can include at least one of non-volatile and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory or optical memory, etc. Volatile memory can include random access memory (RAM) or external cache memory. As an illustration and not limitation, RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM).
[0173] The technical features of the above embodiments may be arbitrarily combined. To make the description concise, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of the present invention.
Claims
1. A transparent object detection method, characterized in that: include: Constructing a neural network, the neural network comprising a parallel residual network module and a boundary detection module, a semantic-boundary co-evolution module connected to an output end of the residual network module and an output end of the boundary detection module, and a decoding module connected to an output end of the semantic-boundary co-evolution module; wherein the semantic-boundary co-evolution module comprises a fusion layer and a cross-attention mechanism layer connected in sequence; Collect images containing transparent objects to construct a data set, use the data set to train a neural network, and obtain a transparent object detection model; An image to be detected containing a transparent object is input into the detection model, and the semantic features of the transparent object in the image to be detected are extracted through the residual network module; the boundary features of the transparent object in the image to be detected are extracted through the edge detection algorithm in the boundary detection module; the semantic features and the boundary features are fused through the fusion layer in the semantic-boundary co-evolution module to obtain the semantic-boundary fusion features of the transparent object; the semantic features and the boundary features are mapped to semantic queries and boundary queries respectively through the cross-attention mechanism layer in the semantic-boundary co-evolution module, and the semantic-boundary fusion features are mapped to semantic keys and semantic values corresponding to the semantic queries, and boundary keys and boundary values corresponding to the boundary queries, and based on the obtained queries, keys and values, the semantic features and boundary features are weighted by the cross-attention mechanism respectively to obtain weighted semantic features and weighted boundary features; the weighted semantic features and weighted boundary features are decoded through the decoding module to obtain the transparent object in the image to be detected.
2. The transparent object detection method according to claim 1, characterized in that: The mapping of semantic features and boundary features to semantic queries and boundary queries respectively, and mapping semantic-boundary fusion features to semantic keys and semantic values corresponding to the semantic queries, and boundary keys and boundary values corresponding to the boundary queries specifically include: The shapes are uniformly transformed into C×H×W semantic features, boundary features, and semantic-boundary fusion features, which are flattened into 2D semantic features F s ', 2D boundary feature F e ' and 2D semantic-boundary fusion feature F c Among them, F s ', F e ' and F c '∈R N×C , N = H × W, C is the number of channels; The 2D semantic features F are respectively mapped by the pre-trained semantic mapping matrix s ' and 2D semantic-boundary fusion feature F c Mapping is performed to obtain the semantic query Q s , semantic key K s and semantic value V s : in, and is the pre-trained semantic mapping matrix; The 2D boundary features F are respectively mapped by the pre-trained boundary mapping matrix e ' and 2D semantic-boundary fusion feature F c 'Map and obtain boundary query Q e , boundary key K e and the boundary value V e : in, and is the pre-trained boundary mapping matrix.
3. The transparent object detection method according to claim 2, characterized in that: The boundary features are weighted by a cross-attention mechanism, specifically including: The boundary query Q e , boundary key K e and the boundary value V e Reshape to H×W×C and obtain the reshaped boundary query f Q , reshape the boundary key f K and reshape the boundary value f V ; In the H×W dimension, the boundary query f is reshaped Q and reshape the boundary key f K Divide into n×n regions, and perform average pooling operation on each region to obtain the first weight representation matrix Q and the second weight representation matrix K; based on the first weight representation matrix Q and the second weight representation matrix K, obtain the regional matching score matrix G: G=Q(K) T Among them, the superscript T represents the transpose operation; Select the top k scores in each row of the region matching score matrix G, integrate the indexes of the k scores to obtain the matrix S; and take out the reshape boundary key f according to the index K and reshape the boundary value f V The corresponding area in: f′ K =Gather(f K ,S) f′ V =Gather(f V ,S) Gather(·) is the torch.gather operation, which is used to extract elements from the input tensor according to the index in the specified dimension and construct a new tensor; f' K and f' V are respectively Q The most correlated key and value features; F' K and f' V After changing to 2D shape, calculate the spatial attention A e : Among them, K' e and V e 'respectively f' K and f' V 2D flat key features and 2D flat value features obtained after transformation into 2D shape; d = C / h, h is the number of attention heads; The 2D boundary feature F is transformed into e 'Add spatial attention A e The calculation result of weighted boundary feature f is obtained ei : in, It is the element-wise addition operation in the residual connection.
4. The transparent object detection method according to claim 3, characterized in that: The semantic features are weighted by a cross-attention mechanism, specifically including: Use the cross attention mechanism to s Weighted, get the attention score A s : Where d = C / h, h is the number of attention heads; the superscript T represents the transposition operation; The 2D semantic features F are transformed into s 'Add attention score A s The calculation result of , obtains the weighted semantic feature f si : in, It is the element-wise addition operation in the residual connection.
5. The transparent object detection method according to claim 4, characterized in that: The output end of the cross attention mechanism layer in the semantic-boundary co-evolution module is connected to the Mix-FFN network layer; the weighted semantic features and weighted boundary features output by the cross attention mechanism layer are enhanced by the Mix-FFN network layer, specifically including: Among them, MPL() is the feedforward operation in the Mix-FFN network, Conv 3×3 () is the convolution operation in the Mix-FFN network; GeLU() is the activation function operation in the Mix-FFN network; f ei is the weighted boundary feature, f eo is the enhanced boundary feature; f si is the weighted semantic feature, f so It is to enhance semantic features.
6. The transparent object detection method according to claim 1, characterized in that: The residual network module is formed by a ResNeXt-101 network; the multi-scale features F of the transparent object in the image to be detected are extracted through multiple residual convolutional layers of the ResNeXt-101 network. i , i∈{0, 1, 2, 3, 4, 5}, the feature F5 with the most semantic information is taken as the semantic feature of the transparent object in the image to be detected.
7. The transparent object detection method according to claim 6, characterized in that: The output end of the boundary detection module and the output end of the residual network module are connected to a boundary refinement module, and the boundary refinement module includes an aggregation layer, a residual layer, a spatial attention mechanism layer and a fusion layer; the boundary features output by the boundary detection module are refined by the boundary refinement module, specifically including The features F0, F1, F2, and F3 output by the ResNeXt-101 network are upsampled to the same dimension through the aggregation layer and then aggregated with the boundary features: f i =Upsample(F i ),i=0,1,2,3 <h2 style=";text-align:left;direction:ltr">f<h2 style=";text-align:left;direction:ltr"> sum <h2 style=";text-align:left;direction:ltr"> (f0, f1, f2, f3, f<h2 style=";text-align:left;direction:ltr"> m <h2 style=";text-align:left;direction:ltr"> ] Among them, Upsample() is the upsampling operation, f m is the boundary feature extracted by the Canny edge detection algorithm in the boundary detection module; [] is the feature aggregation operation; The residual layer is used to enhance the upsampled feature map f0 with the most boundary features and the least semantic information: Among them, ReLU() is the activation function, BN() is the batch normalization operation; Conv 3×3 () is the convolution operation, f'0 is the enhanced boundary feature; is the element-by-element addition operation in the residual connection; f'0 is the initial enhanced boundary feature f'0; The initial enhanced boundary feature f'0 is further enhanced through the spatial attention mechanism layer: α j =Sigmoid([f'0,f j ])j∈{1,2,sum} Among them, Sigmoid() is the sigmoid function used for normalization, α j is the attention map obtained from the j-th upsampled feature map, It is an element-by-element multiplication operation; is the jth boundary feature after further enhancement; The fourth boundary feature output by the spatial attention mechanism layer is fused through the layer The feature F5 with the most semantic information output by the ResNeXt-101 network is aggregated to obtain the refined boundary feature 8. The transparent object detection method according to claim 7, characterized in that: The semantic-boundary co-evolution module includes a first semantic-boundary co-evolution module, a second semantic-boundary co-evolution module, a third semantic-boundary co-evolution module and a fourth semantic-boundary co-evolution module connected in sequence; the first, second, third and fourth semantic-boundary co-evolution modules enhance the semantic features and boundary features in the feature transfer process by connecting in series.
9. A transparent object detection device, characterized in that: include: A model building module, used to build a neural network, wherein the neural network includes a parallel residual network module and a boundary detection module, a semantic-boundary co-evolution module connected to the output end of the residual network module and the output end of the boundary detection module, and a decoding module connected to the output end of the semantic-boundary co-evolution module; wherein the semantic-boundary co-evolution module includes a fusion layer and a cross-attention mechanism layer connected in sequence; A model training module is used to collect images containing transparent objects to construct a data set, and use the data set to train a neural network to obtain a transparent object detection model; The detection module is used to input an image to be detected containing a transparent object into a detection model, extract the semantic features of the transparent object in the image to be detected through a residual network module; extract the boundary features of the transparent object in the image to be detected through an edge detection algorithm in a boundary detection module; fuse the semantic features and the boundary features through a fusion layer in a semantic-boundary co-evolution module to obtain the semantic-boundary fusion features of the transparent object; map the semantic features and the boundary features to semantic queries and boundary queries respectively through a cross-attention mechanism layer in the semantic-boundary co-evolution module, map the semantic-boundary fusion features to semantic keys and semantic values corresponding to the semantic queries, and boundary keys and boundary values corresponding to the boundary queries, and based on the obtained queries, keys and values, respectively weight the semantic features and the boundary features through a cross-attention mechanism to obtain weighted semantic features and weighted boundary features; decode the weighted semantic features and the weighted boundary features through a decoding module to obtain the transparent object in the image to be detected.
10. A computer device, characterized in that: The method comprises a memory, a processor and a computer program stored in the memory and executable on the processor, wherein when the processor executes the program, the method described in any one of claims 1 to 8 is implemented.
Citation Information
Patent Citations
Transparent object image segmentation method and system
CN115082675A
Urban streetscape advertisement image segmentation method
CN116189180A
Semantic segmentation method and device for transparent object in image and electronic equipment
CN116993973A
System and method for occluding contour detection
US20190050667A1