OANet-based occlusion perception instance segmentation method

By introducing the OANet model into the instance segmentation technology, combining GCN, Transformer and multiple attention feature fusion modules, the problem of segmentation accuracy reduction in complex overlap and occlusion is solved, and a more efficient instance segmentation effect is achieved.

CN120070887APending Publication Date: 2025-05-30ZHONGBEI UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510117693.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-24
Publication Date
2025-05-30

AI Technical Summary

Technical Problem

The existing instance segmentation technology is difficult to accurately locate and segment target boundaries when dealing with complex overlapping instances, resulting in reduced segmentation accuracy, especially when feature confusion and key features are occluded in the case of occlusion, the complexity of segmentation task increases.

Method used

Using OANet-based occlusion-aware instance segmentation method, through the joint use of graph convolution networks GCN and Transformer, the interactive relationship between the occluded object and the occluded object is learned, and combined with the multiple attention feature fusion module EFF and the occluded pixel regression module PSA, the global features are integrated to achieve more accurate instance segmentation.

Benefits of technology

The accuracy of instance segmentation in complex overlap and occlusion is improved, the ability to identify and segment objects of different scales is enhanced, the ability to detect small and large targets is improved, and the problem of segmentation of occluded areas and obstructed objects is better handled.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120070887A_ABST
    Figure CN120070887A_ABST
Patent Text Reader

Abstract

The invention discloses an OANet (Occlusion-Aware Network)-based occlusion perception instance segmentation method, and aims to solve the problem of instance segmentation caused by object overlapping in a complex scene. According to the method, modeling is performed on the occlusion relation explicitly by constructing a three-layer network architecture, so that the accuracy and robustness of instance segmentation in a complex scene are improved. Wherein the top-layer GCN + Transform network detects an occluded object, the bottom-layer GCN + Transform network deduces an occluded instance, and an occlusion regression branch PSA + RefConv of the newly added occluded object is used for integrating global information and optimizing a segmentation boundary. Experimental results show that compared with a BCNet network, the performance of the OANet model is remarkably improved on COCO and COCOA instance segmentation reference data sets, the AP (average precision) is improved by 1.2 compared with that of a reference model, and the effectiveness of the OANet model in processing serious shielding conditions is proved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of computer vision, and particularly relates to an occlusion-aware instance segmentation method based on OANet. Background Art

[0002] With the rapid development of deep learning technology, the performance of image segmentation technology has been significantly improved. How to effectively solve complex image segmentation tasks has become an important issue in the field of computer vision. The existing latest research on instance segmentation generally follows the Mask R-CNN paradigm, that is, first detect candidate bounding boxes, and then finely segment the masks of each instance. In the COCO instance segmentation task, Mask R-CNN and its variants have won wide recognition in the industry with their excellent performance. However, this technology still has deficiencies. First, most improvements mainly focus on optimizing the backbone architecture design, and pay less attention to the instance mask regression after extracting region of interest (ROI) features from object detection, which will lead to a decrease in accuracy during the mask regression process, especially in complex scenarios where there is a high similarity and overlap between objects; second, since each instance mask is regressed separately, and it is implicitly assumed that the objects in the ROI have almost complete contours during the regression process, resulting in segmentation errors.

[0003] In actual application scenarios, there are a large number of overlapping and occluding objects. When dealing with complex overlapping instances, it is difficult for the model to accurately locate and segment the boundaries of the targets, and it is easy to cause feature confusion between occluding objects and occluded objects, thus affecting the segmentation accuracy. In addition, in the case of occlusion, the key features of some targets may be completely covered, resulting in the segmentation model lacking effective visual information and unable to infer the complete object contour, which further increases the complexity of the segmentation task. Summary of the Invention

[0004] Aiming at the problem of affecting segmentation accuracy and increasing segmentation difficulty in the case of overlapping occlusion, the present invention provides an occlusion-aware instance segmentation method based on OANet (Occlusion-Aware Network).

[0005] To achieve the above object, the present invention adopts the following technical solutions:

[0006] Step 1, image input and preprocessing: Input the original image, and perform preprocessing on the image, including operations such as image normalization and data augmentation, so as to provide a standardized input for subsequent network processing.

[0007] Step 2, Extract multi-scale features: Input the preprocessed image into the backbone network of OANet model with FPN to extract features at different levels of the image, and use the multi-attention feature fusion module EFF to fuse the multi-level image features to reduce the problem of single-scale information loss.

[0008] Step 3, Construct an occlusion-aware branch to obtain the feature information of the occluder and the occluded object; By jointly using the graph convolutional network GCN and Transformer, learn the interaction relationship between the occluding object and the occluded object, and then obtain the respective feature information of the occluder and the occluded object.

[0009] Step 4, Global feature integration and optimization: Input the feature information of the occluder and the occluded object into the occlusion pixel regression module PSA, and combine the reparameterized convolution module RefConv to globally integrate all the information in the image to achieve more accurate instance segmentation.

[0010] Step 5, Training and evaluation: Train on the COCO dataset and evaluate the model performance on the COCO-OCC dataset to ensure that the model can effectively handle highly overlapping object scenes.

[0011] Furthermore, the specific steps of the preprocessing operation in Step 1 are as follows:

[0012] Step 1.1, Obtain the input image, and the image format can be common image formats such as JPEG, PNG, BMP, etc.

[0013] Step 1.2, Perform image data augmentation and normalization preprocessing on the input image. Data augmentation includes operations such as rotation, translation, and flipping to improve the generalization ability of the model for different samples. The normalization process scales the image pixel values to a unified range to ensure the consistency of the network input.

[0014] Furthermore, the specific steps of the multi-scale feature extraction in Step 2 are as follows:

[0015] Step 2.1, Feature extraction: Input the preprocessed image into the backbone network to extract features at different scales, generating multi-level and multi-scale feature maps;

[0016] Step 2.2, Construct the Feature Pyramid Network FPN: Select feature maps at different levels from the backbone network as the input of the Feature Pyramid Network FPN. Use the Feature Pyramid Network FPN to upsample the high-level feature maps to have the same spatial resolution as the low-level feature maps, then unify the number of channels of the upsampled high-level feature maps and the low-level feature maps through 1×1 convolution, add them for feature fusion, and perform 3×3 convolution operations on the fused feature maps to obtain the final feature maps.

[0017] Step 2.3, use the multi - attention feature fusion module EFF to perform multi - scale feature fusion on the obtained feature maps: The multi - attention feature fusion module EFF includes three modules: the enhanced attention gate module EAG, the efficient channel attention module ECA, and the spatial attention module SA;

[0018] Among them, the enhanced attention gate module EAG fuses the global feature g and the local feature x, enhances the important features in the occluded area and suppresses the irrelevant areas, and realizes feature fusion through the combination of grouped convolution and residual connection. The calculation formula is as follows:

[0019] EAG skip = ReLU(BN(GroupConv 32 (g)))+ReLU(BN(GroupConv 32 (x)))(1)

[0020] ψ = σ(Conv1×1(EAG skip )) (2)

[0021] Output EAG = x·ψ + x (3)

[0022] Among them, g represents the global feature; x represents the local feature; EAG skip represents the enhanced feature after fusion; σ represents the sigmoid function, ψ represents the attention weight calculated based on the global feature g and the local feature x, which is used to adjust the importance of the local feature; output EAG represents the output result of the EAG module.

[0023] The efficient channel attention module ECA introduces local interaction between channels to capture the local context information of each channel. Its core is to use global average pooling and adaptive convolutional kernels to learn the weight distribution of different channels. The calculation formula is:

[0024] X ECA = Concat(x, Output EAG ) (4)

[0025] v = σ(Conv1D(GlobalAvgPool(X ECA ))) (5)

[0026] Output ECA = X ECA ·v (6)

[0027] Among them, v represents the weight vector of channel attention, and the weight distribution of different channels is learned through global average pooling and adaptive convolution kernels. GlobalAvgPool represents global average pooling for each channel of the input local feature x, and output EAG represents the output of the ECA module;

[0028] The spatial attention module SA dynamically adjusts the saliency of each pixel in the feature map by calculating the spatial attention weight, strengthening the attention to key regions. The calculation formula is:

[0029] X SA = Concat(MaxPool(X ECA ), AvgPool(X ECA )) (7)

[0030] Output SA = X SA ·σ(Conv2D(X SA )) (8)

[0031] Among them, Concat represents the feature concatenation operation.

[0032] Furthermore, the specific steps of constructing the occlusion perception model in step 3 are as follows:

[0033] Step 3.1, use the graph convolutional network GCN to perform local feature modeling on the objects in the image to better process the detailed features of overlapping objects.

[0034] The graph convolutional network aggregates local features through the adjacency matrix. The formula is:

[0035] e = RELU(σ(A ij XW)) + X (9)

[0036] Among them, X is the input feature matrix, e is the local output feature matrix, W is the learnable weight matrix, σ is the activation function, and A ij is the adjacency matrix representing the similarity between nodes, defined as:

[0037] A ij = sigmoid(θ(x i )) T , φ(x j )) (10)

[0038] Among them, x i , x j are graph nodes, and θ, φ are linear transformations implemented by 1×1 convolutions;

[0039] Step 3.2: Use the Transformer model to capture long-range dependencies to assist in the accurate prediction of occluders and occluded objects.

[0040] Furthermore, input the local output feature matrix e extracted by the GCN module into the Transformer. The propagation process of each layer is as follows:

[0041]

[0042] Among them, H l is the node embedding of the l-th layer, Q = H l-1 W Q , K = H l-1 W K , V = H l-1 W V represent the query, key, and value of the features respectively. W Q , W K , W V are learnable weight matrices; when l = 0, H l = e. The formula for the global feature aggregation of the transformer is:

[0043] E = H l (12)

[0044] E is the feature matrix with global information;

[0045] Finally, after multi-layer propagation, the final output of the Transformer is the feature matrix with global information.

[0046] Furthermore, the specific steps of the global feature optimization in step 4 are as follows:

[0047] Step 4.1: In the occlusion pixel regression module PSA, optimize the feature information in the channel dimension and the spatial dimension through two methods of channel attention and spatial attention. First, calculate the channel attention weight and weight the channels through 1×1 convolution to enhance important features; then use spatial attention to further strengthen the attention to key regions and optimize the spatial features.

[0048] Input the feature matrix E with global information into the occlusion pixel regression module PSA. Apply the attention mechanism on the channels. The calculation formula for the channel attention weight is:

[0049] W ch (E) = sigmoid[W z (σ 1 (W v (E))·softMax(σ 2 (W q (E))))] (13)

[0050] Among them, W q , W v , W z ∈R C×C represents a 1×1 convolutional matrix for linear transformation of channel information. W q (E), W v (E), W z (E)∈R C×1×1 , representing the compressed channel features. σ 1 , σ 2 are tensor shape transformation operations for matching the dimensions of matrix dot product.

[0051] The calculation formula for the output channel-weighted feature is:

[0052] Z ch =W ch (E)⊙E (14)

[0053] Similarly, the calculation formula for the spatial attention weight is:

[0054] W sp (E)=sigmoid[σ 3 (softMax(σ 1 (FGP(W q (E))))·σ 2 (W v (E)))] (15)

[0055] Among them, represents global pooling.

[0056] The spatially weighted output feature is expressed as:

[0057] Z sp =W sp (E)⊙E (16)

[0058] The enhancement results of the channel attention and spatial attention branches are fused:

[0059] Z=Z ch +Z sp (17)

[0060] Step 4.2, use the reparameterized convolution module RefConv to refine the output features of the occlusion pixel regression module PSA and the original features, further optimize the segmentation effect of the occlusion area, and finally obtain accurate instance segmentation results.

[0061] Finally, the output result of the occlusion area is:

[0062] Z last =σ(WRef1 *Z input +W Ref2 *Z SA ) (18)

[0063] Among them, W Ref1 and W Ref2 are trainable weights, and X represents the original input features.

[0064] Compared with the prior art, the present invention has the following advantages:

[0065] (1) Improving the detection ability of small and large targets: Through the multi-attention feature fusion module EFF, the present invention effectively fuses low-level detailed features and high-level semantic features, enhancing the recognition and segmentation ability of objects at different scales. Especially in complex backgrounds, it can handle small and large targets simultaneously, improving the overall detection performance.

[0066] (2) Better occlusion perception ability: The present invention effectively addresses the segmentation problem of occluded regions and occluded objects by combining the graph convolutional network GCN and Transformer. GCN can model the local relationships between objects, and Transformer captures global dependencies through the self-attention mechanism. The combination of the two enables the model to accurately distinguish objects in the case of occlusion or overlap, improving the segmentation accuracy.

[0067] (3) More accurate segmentation of occluded regions: In the detection and segmentation of occluding and occluded objects, the PSA module proposed by the present invention enhances the model's attention to occluded regions through channel attention and spatial attention mechanisms, improving the segmentation accuracy of objects in the case of occlusion. Description of the Drawings

[0068] Figure 1 is the overall model block diagram of OANet.

[0069] Figure 2 is the network framework diagram of the EFF model;

[0070] Figure 3 is the network framework diagram of the PSA model;

[0071] Figure 4 is the network framework diagram of the RefConv model. Detailed Embodiments

[0072] To understand the present invention in depth, we will describe it comprehensively and meticulously. However, the present invention has multiple implementation manners and is not limited to the specific examples listed herein. The presentation of these examples aims to deepen the comprehensive understanding of the disclosed content of the present invention.

[0073] Embodiment 1

[0074] The OANet model in an occlusion-aware instance segmentation method based on OANet consists of three modules:

[0075] (1) The backbone network with Feature Pyramid Network (FPN) and the Multiple Attention Feature Fusion Module (EFF): The backbone network with FPN is used to extract multi-level features from the input image, and the EFF is used to deeply fuse the multi-level features;

[0076] (2) The detection branch is used to generate instance proposals (ROIs) and identify the target regions;

[0077] (3) The three-layer convolutional network structure is used to achieve the final mask prediction.

[0078] Among them, the top-level GCN + Transformer network detects the position and contour of the occluded object and provides the initial features of the occluded area for the downstream network;

[0079] The bottom-level GCN + Transformer network obtains the effective features of the occluded instance and infers its implicit features;

[0080] The newly added occlusion-aware branch then fuses the occlusion area features and the interaction relationship of the occlusion through the Pixel Segmentation of Occlusion (PSA) module and the Reparameterized Convolution (RefConv) module to enhance the perception ability of the occluded area and optimize the segmentation effect; as Figure 1 shown.

[0081] The method includes the following steps:

[0082] Step 1, Image input and preprocessing: Input the original image and preprocess the image, aiming to convert the original image into a format suitable for network training to ensure the consistency and standardization of the input. The specific steps of Step 1, image input and preprocessing are as follows:

[0083] Step 1.1, Input the original image: The input image can come from various sources, such as public datasets (e.g., COCO, VOC, ADE20K, etc.), or devices such as cameras, sensors, drones, and smartphones in practical applications; the images can have different resolutions and formats, and common image formats include JPEG, PNG, BMP, TIFF, etc. To process these images, it is first necessary to convert them into a unified color space (RGB) and a standard input format (the value of each pixel is 0 - 255).

[0084] Step 1.2, Image preprocessing: Normalize the image, requiring that the pixel values be mapped to the range [0, 1] so that the pixel distributions of different images are consistent.

[0085] Step 2, Extract multi-scale features: Input the preprocessed image into the backbone network of the OANet model with a Feature Pyramid Network (FPN) to extract features at different levels of the image, and use a Multiple Attention Feature Fusion Module (EFF) to fuse the multi-level image features; The Multiple Attention Feature Fusion Module (EFF) is as Figure 2 shown; The specific steps for extracting multi-scale features in Step 2 are as follows:

[0086] Step 2.1, Feature extraction. The backbone network adopts a residual learning mechanism. After the preprocessed input image passes through the first convolutional layer, a series of feature maps are generated, and these feature maps are then propagated through a series of residual blocks. Among them, each residual block contains multiple convolutional layers and a skip connection for adding the input directly to the output. As the feature maps propagate through the residual blocks, through more convolutional operations and non-linear activation functions, higher-level features are gradually extracted.

[0087] Step 2.2, To solve the problem of different object scales, the image features are used to extract features at different scales through the Feature Pyramid Network (FPN). First, the Feature Pyramid Network (FPN) upsamples the high-level feature maps to have the same spatial resolution as the low-level feature maps. Secondly, the upsampled high-level feature maps and the low-level feature maps are unified in the number of channels through 1×1 convolution and added for feature fusion; Finally, a 3×3 convolution operation is performed on the fused feature maps to obtain the final feature maps.

[0088] Step 2.3, Use the Multiple Attention Feature Fusion Module (EFF) to perform multi-scale feature fusion on the obtained feature maps to improve the comprehensive processing ability of different-scale features in the image.

[0089] The Multiple Attention Feature Fusion Module (EFF) includes three modules: an Enhanced Attention Gate Module (EAG), an Efficient Channel Attention Module (ECA), and a Spatial Attention Module (SA);:

[0090] Among them, the Enhanced Attention Gate Module (EAG) fuses the global feature g and the local feature x, enhances the important features in the occluded area and suppresses the irrelevant areas, and realizes feature fusion through the combination of grouped convolution and residual connection. The calculation formula is as follows:

[0091] EAG skip = ReLU(BN(GroupConv 32 (g)))+ReLU(BN(GroupConv 32 (x))) (1)

[0092] ψ = σ(Conv1×1(EAG skip )) (2)

[0093] OutputEAG = x·ψ + x (3)

[0094] Among them, g represents the global feature; x represents the local feature; EAG skip represents the enhanced feature after fusion; σ represents the sigmoid function, ψ represents the attention weight calculated based on the global feature g and the local feature x, and is used to adjust the importance of the local feature; Output EAG represents the output result of the EAG module.

[0095] The efficient channel attention module ECA introduces local interaction between channels to capture the local context information of each channel. Its core is that global average pooling and adaptive convolutional kernels learn the weight distribution of different channels. The calculation formula is:

[0096] X ECA = Concat(x, Output EAG ) (4)

[0097] v = σ(Conv1D(GlobalAvgPool(X ECA ))) (5)

[0098] Output ECA = X ECA ·v (6)

[0099] Among them, v represents the weight vector of channel attention, and learns the weight distribution of different channels through global average pooling and adaptive convolutional kernels. GlobalAvgPool represents global average pooling for each channel of the input local feature x, and Output ECA represents the output of the ECA module;

[0100] The spatial attention module SA calculates the spatial attention weight to dynamically adjust the saliency of each pixel in the feature map and strengthen the attention to the key area. The calculation formula is:

[0101] X SA = Concat(MaxPool(X ECA ), AvgPool(X ECA )) (7)

[0102] Output SA = X SA ·σ(Conv2D(X SA )) (8)

[0103] Among them, Concat represents the feature concatenation operation.

[0104] Step 3: Construct an occlusion perception branch to obtain the feature information of the occluder and the occluded object; by jointly using a graph convolutional network (GCN) and a Transformer, learn the spatial relationships and context information between different objects in the image, distinguish and predict the occluder and the occluded object, and further solve the instance segmentation problem in the occluded area.

[0105] Step 3.1: Use a graph convolutional network (GCN) to perform local feature modeling on the objects in the image;

[0106] The graph convolutional network aggregates local features through an adjacency matrix, and the formula is:

[0107] e = RELU(σ(A ij XW)) + X (9)

[0108] where X is the input feature matrix, e is the local output feature matrix, W is the learnable weight matrix, σ is the activation function, and A ij is the adjacency matrix representing the similarity between nodes, defined as:

[0109] A ij = sigmoid(θ(x i ), φ(x T )) (10) j

[0110] where x i , x j are the graph nodes, and θ, φ are linear transformations implemented by 1×1 convolutions;

[0111] Step 3.2: Use the Transformer model to capture long-range dependencies and accurately predict the occluder and the occluded object;

[0112] Input the local output feature matrix e extracted by the GCN module into the Transformer model, and the propagation process of each layer is:

[0113]

[0114] where H l is the node embedding of the l-th layer, Q = H l-1 W Q , K = H l-1 W K , V = H l-1 W V represent the query, key, and value of the feature respectively, and W Q , W K , W V are learnable weight matrices; when l = 0, H l = e; ​

[0115] The formula for the global feature aggregation of the Transformer is as follows:

[0116] E = H l (12)

[0117] E is the feature matrix with global information;

[0118] Finally, after multiple layers of propagation, the final output of the Transformer is the feature matrix with global information.

[0119] Step 4, Global feature integration and optimization: Input the obtained feature information of the occluder and the occluded object into the occlusion pixel regression module PSA, and combine it with the reparameterized convolution module RefConv to globally integrate all the information in the image, thereby improving the accuracy of instance segmentation.

[0120] The occlusion pixel regression module PSA is a module that combines pixel-level attention mechanism and regression strategy. By simultaneously modeling channel attention and spatial attention, it infers the features of the occluded area, thereby restoring the integrity of the image features, as shown in the appendix Figure 3 shown; the reparameterized convolution module RefConv refines these global features through the reparameterization mechanism, focusing on enhancing the representation of the boundary region and fine-grained features, as shown in the appendix Figure 4 shown.

[0121] Step 4.1, Obtain the prediction information of the occluder and the occluded object from the previous modules GCN and Transformer. This information includes the position, category, and local and global feature embeddings of the object. Input this information into the occlusion pixel regression module PSA to extract global semantic information.

[0122] Input the feature matrix E with global information into the PSA module, and apply the attention mechanism on the channels. The calculation formula for the channel attention weight is:

[0123] W ch (E) = sigmoid[W z (σ 1 (W v (E))·softMax(σ 2 (W q (E))))] (13)

[0124] Among them, W q 、W v 、W z ∈R C×C represents a 1×1 convolution matrix for the linear transformation of channel information; W q (E), W v (E), Wz (E) ∈ R C×1×1 represents the compressed channel features; σ 1 、σ 2 is a tensor shape transformation operation used to match the dimensions for matrix dot product;

[0125] The calculation formula for the output channel-weighted features is:

[0126] Z ch = W ch (E) ⊙ E (14)

[0127] The calculation formula for the spatial attention weight is:

[0128] W sp (E) = sigmoid[σ 3 (softMax(σ 1 (FGP(W q (E)))) · σ 2 (W v (E)))] (15)

[0129] Where, represents global pooling;

[0130] The spatially weighted output features are represented as:

[0131] Z sp = W sp (E) ⊙ E (16)

[0132] The enhancement results of the channel attention and spatial attention branches are fused:

[0133] Z = Z ch + Z sp (17)

[0134] Step 4.2, use the reparameterized convolution module RefConv to refine the output features of the occlusion pixel regression module PSA and the original features;

[0135] Finally, the output result of the occlusion area is:

[0136] Z last = σ(W Ref1 * X + W Ref2 * Z psa ) (18)

[0137] Where, W Ref1 and W Ref2 are trainable weights, and X represents the original input features.

[0138] Step 5, Model Training and Evaluation. Train on the COCO dataset and evaluate the model performance on the COCO-OCC dataset to ensure that the model can effectively handle highly overlapping object scenarios. The specific steps of the model training and evaluation in Step 5 are as follows:

[0139] Step 5.1, Use the COCO dataset for model training and evaluation. This dataset contains 2017train (118k images), 2017val (5k images), and 2017test-dev (40k images). In addition, the additional COCO-OCC dataset is used to propose an instance segmentation model for occlusion situations between multiple objects.

[0140] Step 5.2, Evaluate the results on 2017test-dev using standard object detection and segmentation task metrics. Compare the performance of TCNet with SOTA methods, including segmentation performance on large and small objects. Especially on the COCO-OCC dataset, which contains highly overlapping objects and is more challenging for the segmentation task.

[0141] Table 1 shows the comparison between this method and the baseline model on the COCO dataset, and Table 2 shows the effects of the proposed three-layer composed of the occluder GCN+TransFormer, the occluded GCN+TransFormer, and the occlusion-aware PSA+RefConv layer.

[0142] Example 2

[0143] Dataset: When constructing and validating the OANet model, a large-scale dataset, Microsoft COCO, was selected. This dataset is widely used in the field of object detection and segmentation. Specifically, the experimental dataset covers three key parts of COCO 2017: the training set (including 118,000 images), the validation set 2017val (5,000 images), and the test development set 2017test-dev (40,000 images). To further explore the model's segmentation ability in the face of occlusion, the COCO-OCC dataset was additionally introduced. This dataset was carefully selected and divided from the Microsoft COCO dataset by the original authors of the baseline model and is specifically used to evaluate the model's performance in dealing with occlusion problems. The COCO-OCC dataset focuses on highly occluded scenarios. It contains 1,005 images carefully selected from 2017val, and the overlap rate of the object bounding boxes in these images is not less than 0.2, thus constituting a very challenging test set.

[0144] Evaluation Metrics: When evaluating the performance of object detection and segmentation tasks on 2017 test-dev, the widely adopted standard metric in this field - Average Precision (AP) - is followed. This metric is used to measure the comprehensive performance of the model at different recall rates. Generally, as the recall rate increases, the number of objects that the model needs to detect or segment increases, which may lead to false detections or a decrease in segmentation accuracy, thus forming a trade-off between precision and recall. To quantify this trade-off relationship, this experiment uses Average Precision (AP) as the core evaluation metric. AP provides a single value that can intuitively reflect the overall performance of the model by calculating the precision values at different recall rates and averaging these values. This metric not only helps to understand the overall performance of the model but also guides the optimization of the model to find the best balance between precision and recall. In addition, for multi-class tasks, the AP value for each class can be calculated and the average of these values can be taken to obtain mAP (mean AP). However, in the official evaluation standard of the Microsoft COCO dataset, to simplify the expression, this multi-class average metric is abbreviated as AP to distinguish it from the AP value of a single class. Therefore, when referring to the Microsoft COCO dataset, this special naming convention needs to be noted.

[0145] Experiment Details: In the model training stage, the Detectron2 platform (an object detection and semantic segmentation framework based on PyTorch, developed by FAIR) is adopted, and the resource configuration is optimized. The configuration is a GeForce RTX 3090 GPU with a total batch size of 4. The optimizer is SGD with momentum, the momentum is 0.1, the initial learning rate is 0.01, and the training process includes 180,000 iterations and 1,000 warm-up iterations. The model contains an occluder GCN + Transformer layer for explicitly modeling occluded regions. However, due to the fact that occlusions are not common in the training data, the learning effect is limited, so a filtering process is adopted to balance the sampling.

[0146] In the ablation study, by comparing the OANet model with other baseline models in Table 1, it can be seen that the AP has improved by 1.2 in precision, demonstrating the advantage of the OANet model in terms of segmentation accuracy. By comparing the performance of the model using three-layer occluder-occluder modeling with the model without adopting this strategy in Table 2, it shows the advantages of three-layer modeling in enhancing the understanding ability of complex occlusion relationships in the case of occlusions and optimizing the processing of occlusion edges, etc.

[0147] Table 1 Comparison of the proposed method with SOTA methods on the COCO test-dev set

[0148]

[0149]

[0150] Table 2 Effects of the three-layer occluder-occluded object modeling

[0151]

[0152] The content not described in detail in the specification of the present invention belongs to the prior art well known to those skilled in the art. Although the illustrative specific embodiments of the present invention have been described above for the convenience of those skilled in the art to understand the present invention, it should be clear that the present invention is not limited to the scope of the specific embodiments. For those of ordinary skill in the art, as long as various changes are within the spirit and scope of the present invention defined and determined by the appended claims, these changes are obvious, and all inventions created using the concept of the present invention are within the scope of protection.

Claims

1. An occlusion-aware instance segmentation method based on OANet, characterized in that: The method comprises the following steps: Step 1, image input and preprocessing: input the original image and preprocess the image; Step 2, extract multi-scale features: input the preprocessed image into the backbone network of the OANet model with the feature pyramid network FPN, extract features at different levels of the image, and use the multi-attention feature fusion module EFF to fuse multi-level image features; Step 3: Build an occlusion perception branch to obtain the feature information of the occluder and the occluded object: By using the graph convolutional network (GCN) and the Transformer model together, learn the interaction between the occluder and the occluded object, and then obtain the feature information of each of the occluder and the occluded object; Step 4: Global feature integration and optimization: Input the feature information of the occluder and the occluded object into the occluded pixel regression module PSA, and combine it with the re-parameterized convolution module RefConv to globally integrate all the information in the image to achieve instance segmentation; Step 5, training and evaluation: Train and test on the COCO dataset, and evaluate the model performance on the COCO-OCC dataset to ensure that the model can effectively handle highly overlapping object scenes.

2. The occlusion-aware instance segmentation method based on OANet according to claim 1, characterized in that: The specific steps of step 1 image input and preprocessing: inputting the original image and preprocessing the image are as follows: Step 1.1, obtain the input image; Step 1.2, perform image data enhancement and normalization preprocessing on the input image.

3. The occlusion-aware instance segmentation method based on OANet according to claim 2, characterized in that: The step 2, extracting multi-scale features: inputting the preprocessed image into the backbone network of the OANet model with the feature pyramid network FPN, extracting features at different levels of the image, and using the multi-attention feature fusion module EFF to fuse multi-level image features. The specific steps are as follows: Step 2.1, feature extraction: input the preprocessed image into the backbone network, extract features of different scales, and generate multi-level and multi-scale feature maps; Step 2.2, construct the feature pyramid network FPN: select feature maps of different levels from the backbone network as the input of the feature pyramid network FPN, use the feature pyramid network FPN to upsample the high-level feature map to make it have the same spatial resolution as the low-level feature map, and then unify the number of channels of the upsampled high-level feature map and the low-level feature map through 1×1 convolution, add them for feature fusion, and perform 3×3 convolution operation on the fused feature map to obtain the final feature map; Step 2.3, use the multi-attention feature fusion module EFF to perform multi-scale feature fusion on the obtained feature map: the multi-attention feature fusion module EFF includes three modules: enhanced attention gate module EAG, efficient channel attention module ECA and spatial attention module SA; Among them, the enhanced attention gate module EAG fuses the global feature g and the local feature x, enhances the important features of the occluded area and suppresses the irrelevant area, and realizes feature fusion through the combination of grouped convolution and residual connection. The calculation formula is as follows: eAG of skip =ReLU(BN(GroupConv 32 (g)))+ReLU(BN(GroupConv 32 (x)))(1) ψ=σ(Conv1×1(EAG) skip )) (2) Output EAG =x·ψ+x (3) Among them, g represents the global feature; x represents the local feature; EAG skip represents the enhanced features after fusion; σ represents the sigmoid function, ψ represents the attention weight calculated based on the global feature g and the local feature x, which is used to adjust the importance of the local feature; Output EAG Represents the output result of the EAG module; The efficient channel attention module ECA introduces local interactions between channels to capture the local context information of each channel. Its core is global average pooling and adaptive convolution kernel learning weight distribution of different channels. The calculation formula is: X ECA =Concat(x,Output EAG ) (4) v=σ(Conv1D(GlobalAvgPool(X ECA ))) (5) Output ECA =X ECA ·v (6) Among them, v represents the weight vector of channel attention. The weight distribution of different channels is learned through global average pooling and adaptive convolution kernel. GlobalAvgPool represents global average pooling for each channel of the input local feature x. Output ECA Represents the output of the ECA module; The spatial attention module SA dynamically adjusts the significance of each pixel in the feature map by calculating the spatial attention weight, and strengthens the focus on the key area. The calculation formula is: X SA =Concat(MaxPool(X ECA ),AvgPool(X ECA )) (7) Output SA =X SA ·σ(Conv2D(X SA ))· (8) Among them, Concat represents the feature concatenation operation.

4. The occlusion-aware instance segmentation method based on OANet according to claim 3, characterized in that: The step 3, constructing an occlusion perception branch to obtain feature information of the occluder and the occluded object: by using the graph convolutional network GCN and the Transformer model together, learning the interaction between the occluder and the occluded object, and then obtaining the feature information of each of the occluder and the occluded object, the specific steps are: Step 3.1, use the graph convolutional network GCN to model the local features of objects in the image; The graph convolutional network aggregates local features through the adjacency matrix, and the formula is: e=RELU(σ(A ij XW))+X (9) Among them, X is the input feature matrix, e is the local output feature matrix, W is the learnable weight matrix, σ is the activation function, and A ij The adjacency matrix representing the similarity between nodes is defined as: A ij =sigmoid(θ(x i ) T ,φ(x j )) (10) Among them, x i 、x j is a graph node, θ and φ are linear transformations implemented by 1×1 convolution; Step 3.2: Use the Transformer model to capture long-range dependencies and accurately predict occluders and occluded objects. The local output feature matrix e extracted by the GCN module is input into the Transformer model, and the propagation process of each layer is: Among them, H l is the node embedding of the lth layer, Q = H l-1 W Q ,K=H l-1 W K ,V=H l-1 W V Represented as the query, key and value of the feature respectively, W Q , W K , W V is a learnable weight matrix; when l = 0, H l =e; The transformer performs global feature aggregation as follows: E=H l (12) E is the feature matrix with global information; Finally, after multiple layers of propagation, the final output of Transformer is a feature matrix with global information.

5. The occlusion-aware instance segmentation method based on OANet according to claim 4, characterized in that: The step 4 global feature integration and optimization: input the feature information of the occluder and the occluded object into the occluded pixel regression module PSA, and combine with the re-parameterized convolution module RefConv to globally integrate all the information in the image. The specific steps to achieve instance segmentation are: Step 4.1, obtain the prediction information of occluders and occluded objects from the previous modules GCN and Transformer, input this information into the occluded pixel regression module PSA, extract global semantic information; input the feature matrix E with global information into the occluded pixel regression module PSA, apply the attention mechanism on the channel, and the calculation formula of the channel attention weight is: W ch (E)=sigmoid[W z (σ1(W v (E))·softMax(σ2(W q (E))))] (13) Among them, W q , W v , W z ∈R C×C Represents a 1×1 convolution matrix, which is used for linear transformation of channel information; W q (E), W v (E), W z (E)∈R C×1×1 , represents the compressed channel features; σ1, σ2 are tensor shape transformation operations, used to match the matrix dot multiplication dimensions; The output channel weighted feature calculation formula is: WITH ch =In ch (E)⊙E (14) The calculation formula of spatial attention weight is: W sp (E)=sigmoid[σ3(softMax(σ1(FGP(W q (E))))·σ2(W v (E)))] (15) in, Represents global pooling; The spatially weighted output feature is expressed as: WITH sp =In sp (E)⊙E (16) Finally, the enhanced results of the channel attention and spatial attention branches are fused: WITH psa =Z ch +Z sp (17) Step 4.2, using the re-parameterized convolution module RefConv to refine the output features and original features of the occluded pixel regression module PSA; Finally, the output result of the occluded area is: WITH last =σ(W Ref1 *X+W Ref2 *WITH psa ) (18) Among them, W Ref1 and W Ref2 are trainable weights and X represents the original input features.