Semi-Supervised Infrared Cloud Segmentation Method Based on Pixel Context Information Mining

By designing a semi-supervised infrared cloud segmentation method based on pixel context information mining, using Resnet38 network and conditional random field to refine the pseudo-label, the problem of low performance of segmentation model caused by low pseudo-label quality is solved, and efficient segmentation of infrared image clouds is achieved.

CN117253040BActive Publication Date: 2025-07-18NORTHWESTERN POLYTECHNICAL UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202311254414.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-09-27
Publication Date
2025-07-18
Estimated Expiration
2043-09-27

AI Technical Summary

Technical Problem

The existing infrared image cloud segmentation method based on semi-supervised learning has low performance in the segmentation model due to the low quality of pseudo-labels, making it difficult to accurately determine the target boundaries and there is a problem of segmentation errors.

Method used

A semi-supervised infrared cloud segmentation method based on pixel context information mining is designed, and a semi-supervised semantic segmentation network is adopted, including the backbone network Resnet38, classification branches and segmentation branches. The context information is mined through the hybrid pooling module and the expanded differential convolution module, and pseudo-label fine-tune processing is performed using conditional random fields to improve segmentation accuracy.

Benefits of technology

It significantly improves the accuracy and reliability of infrared image cloud segmentation, can obtain high-quality segmentation results, and is suitable for other computer vision tasks.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117253040B_ABST
    Figure CN117253040B_ABST
Patent Text Reader

Abstract

The present invention provides a semi-supervised infrared cloud segmentation method based on pixel context information mining. First, a training data set is constructed; then, the semi-supervised semantic segmentation network is trained using the training data set, where the semi-supervised semantic segmentation network mainly includes a backbone network, a classification branch, and a segmentation branch, and the designs of the classification branch and the segmentation branch fully mine context information and texture information to improve the accuracy of the segmentation result; finally, the trained network is used to process the infrared image to be processed to obtain the segmentation result. The present invention can obtain better infrared image segmentation results.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of infrared image segmentation, and particularly relates to a semi-supervised infrared cloud segmentation method based on pixel context information mining. Background Art

[0002] Infrared imaging has the advantages of long detection range, all-weather operation, and high concealment. Therefore, infrared images are widely used in military monitoring, reconnaissance, early warning and other tasks. With the development of deep learning, significant achievements have been made in the method of infrared image cloud segmentation based on neural networks. During the training process of fully supervised infrared image cloud segmentation, there are pixel-by-pixel semantic annotation images as labels to assist learning, and a high accuracy can be obtained. Compared with the bounding box level and the fully supervised infrared image semantic segmentation task, it is a huge challenge to supervise the training of the segmentation model only relying on image-level semantic information. In addition, due to the lack of pixel-level accurate semantic annotation, it is difficult for the model that loses the boundary information between the target and the background to accurately determine the target boundary; due to the lack of strong supervision information, the segmentation error problem caused by image blur and target occlusion will be further aggravated, making the problems of image blur leading to segmentation errors, information loss, and occlusion leading to incorrect prediction categories in semantic segmentation itself more obvious. Summary of the Invention

[0003] In order to overcome the deficiency that the performance of the segmentation model is low due to the low quality of pseudo-labels in the existing semi-supervised learning-based infrared image cloud segmentation method, the present invention provides a semi-supervised infrared cloud segmentation method based on pixel context information mining. First, a training data set is constructed; then, the semi-supervised semantic segmentation network is trained using the training data set, where the semi-supervised semantic segmentation network mainly includes a backbone network, a classification branch, and a segmentation branch, and the design of the classification branch and the segmentation branch fully mines context information and texture information to improve the accuracy of the segmentation result; finally, the trained network is used to process the infrared image to be processed to obtain the segmentation result. The present invention can obtain better infrared image segmentation results.

[0004] A semi-supervised infrared cloud segmentation method based on pixel context information mining, characterized by the following steps:

[0005] Step 1, construct a training data set: divide the infrared image data into a training data set and a test data set according to a ratio of 9:1, where the training data set includes a class-labeled data set and an unlabeled data set, the class-labeled data set consists of labeled images and their class labels, and the ratio of the number of images contained in the class-labeled data set to the unlabeled data set is 1:4;

[0006] Step 2, Network Training: Use the training dataset obtained in Step 1 to train the semi-supervised semantic segmentation network to obtain a trained network;

[0007] The semi-supervised semantic segmentation network mainly includes a feature extraction backbone network Resnet38, a classification branch, and a segmentation branch. Among them, the network Resnet38 is pre-trained on ImageNet;

[0008] The specific processing process of the semi-supervised semantic segmentation network is as follows: Input the images in the training dataset into the network Resnet38 to obtain image features F. The image features F are input into the classification branch to output the classification activation map of the global semantic features and the class scores p c , and the image features F are input into the segmentation branch to output the multi-resolution activation map Add and average the activation map output by the classification branch and the activation map output by the segmentation branch to obtain the final multi-resolution semantic activation map Each pixel value in the map is the activation value corresponding to the pixel at that position. Use the L1 norm to normalize the activation map so that the activation value of each pixel is between 0 and 1, and then use the Argmax function to find the maximum activation value corresponding to each pixel in the normalized activation map to form a rough semantic segmentation pseudo-label I CAM , and use conditional random fields to perform segmentation refinement on the rough semantic segmentation pseudo-label to obtain a fine pseudo-label I CRF ; where, F3 represents the global semantic feature, and c represents the class serial number;

[0009] The classification branch includes a hybrid pooling module, a Softmax function, and a max pooling layer. The image features F are input into the hybrid pooling module to output the global semantic feature F3. The global semantic feature F3 passes through the Softmax function to obtain the classification activation map The global semantic feature F3 passes through the max pooling layer to obtain the class scores p c ; The hybrid pooling module includes a local hybrid pooling module, a 1×1 convolutional layer, and a global hybrid pooling module. The local hybrid pooling module contains a parallel mean pooling layer AP p1 and a max pooling layer MP p1 , and the global hybrid pooling module contains a parallel global average pooling layer GAP p2 and a global max pooling layer GMP p2 , and the processing process expression of the hybrid pooling module is as follows:

[0010]

[0011] F2 = fconv (F1) (2)

[0012]

[0013] Among them, F1 represents the features output by the local mixing pooling module, F2 represents the features output by the 1×1 convolutional layer, F3 represents the global semantic features output by the global mixing pooling module, and AP p1 (·) represents the operation of the average pooling layer, MP p1 (·) represents the operation of the max pooling layer, f conv (·) represents the 1×1 convolution operation, GAP p2 (·) represents the operation of the global average pooling layer, GMP p2 (·) represents the operation of the global max pooling layer;

[0014] The segmentation branch mainly includes a multi-scale sampling module, a 4-neighborhood dilated differential convolution module, a dilated diagonal differential convolution module, and a two-dimensional convolutional layer. The specific processing process is as follows: The multi-scale sampling module performs sampling processing on the input image features F at four scales s∈[0.5, 1, 1.5, 2]. After sampling, the features at each scale pass through the dilated 4-neighborhood differential convolution module. Centered on each pixel, the differential convolution is performed on the center point feature and the features of the points in its second-order 4-neighborhood to obtain the context information of the corresponding scale features After sampling, the features at each scale pass through the dilated diagonal differential convolution module. Centered on each pixel, the differential convolution is performed on the center point feature and the features of the points in its second-order diagonal neighborhood to obtain the context information of the corresponding scale features After sampling, the features at each scale pass through the two-dimensional convolutional layer to obtain the low-dimensional information of the corresponding scale features For the information of the features at each scale The sum is taken and averaged to obtain the final context features of each scale The context features of each scale Pass through the Softmax function to obtain the class activation maps corresponding to each scale The sum is taken and averaged for all class activation maps to obtain the final multi-resolution activation map

[0015] The loss function of the semi-supervised semantic segmentation network is expressed as follows:

[0016] L total = L cls + L seg (4)

[0017]

[0018]

[0019] Among them, L total represents the total network loss, and L cls represents the classification branch loss, and L seg represents the segmentation branch loss, and y c represents the classification label of the c-th class;

[0020] Step 3, data segmentation: Input the infrared image to be processed into the trained semi-supervised semantic segmentation network obtained in Step 2, and output its cloud layer segmentation result.

[0021] The beneficial effects of the present invention are as follows: Since a semi-supervised semantic segmentation framework model based on pixel context information mining is designed, and a hybrid pooling module and a dilated differential neural network module are used to mine context information, the segmentation performance of the model is significantly improved; the segmentation performance of the method of the present invention is reliable, and the adopted image context information mining method is of great significance for semi-supervised semantic segmentation tasks and can also be applied to other computer vision tasks. Description of the Drawings

[0022] Figure 1 is the structural diagram of the semi-supervised semantic segmentation network of the present invention;

[0023] Figure 2 is the structural diagram of the hybrid pooling module network;

[0024] Figure 3 is the schematic diagram of the processing of the dilated differential convolution module;

[0025] Figure 4 is an example of the original image of the 38-cloud dataset;

[0026] Figure 5 is an example of the manual annotation result of the 38-cloud dataset image;

[0027] Figure 6 is the segmentation result of the 38-cloud dataset image using the method of the present invention. Specific Embodiments

[0028] The present invention will be further described below in conjunction with the drawings and embodiments, and the present invention includes but is not limited to the following embodiments.

[0029] To improve the pseudo-label quality and model segmentation performance of semi-supervised image infrared image cloud layer segmentation, the present invention provides a semi-supervised infrared cloud layer segmentation method based on pixel context information mining. Starting from the perspective of pixel context information mining, a semi-supervised semantic segmentation network is designed to obtain a more accurate semantic segmentation result of the infrared image. The specific implementation process is as follows:

[0030] 1. Construct the training dataset

[0031] The infrared image data is divided into a training dataset and a test dataset according to a ratio of 9:1, ensuring that images of different ground background types, different cloud thicknesses, and shapes are covered. In this embodiment, the 38-cloud dataset is used. The 38-cloud dataset contains 38 Landsat 8 satellite remote sensing images and includes pixel-level manual annotations. To ensure the capacity of the dataset, the original images are cropped into non-overlapping small pieces of 384×384. Finally, 4500 cloud images are divided into the training set and 500 images are divided into the test set. Among them, the training dataset includes a class-labeled dataset and an unlabeled dataset. The class-labeled dataset consists of labeled images and their class labels, and the ratio of the number of images in the class-labeled dataset to the number of images in the unlabeled dataset is 1:4.

[0032] 2. Network training

[0033] This step mainly uses the training dataset to train the semi-supervised semantic segmentation network to obtain a trained network.

[0034] To obtain good segmentation performance and fully exploit the context information, the present invention designs a semi-supervised semantic segmentation network, as Figure 1 shown, which mainly includes a feature extraction backbone network, a classification branch, and a segmentation branch. Among them, the Resnet38 network pre-trained on ImageNet is selected as the feature extraction backbone network. The classification branch includes a hybrid pooling module, a Softmax function, and a max pooling layer. The hybrid pooling module, as Figure 2 shown, includes a local hybrid pooling module, a 1×1 convolutional layer, and a global hybrid pooling module. The local hybrid pooling module consists of a parallel mean pooling layer and a max pooling layer, and the global hybrid pooling module consists of a parallel global average pooling layer and a global max pooling layer. The segmentation branch mainly includes a multi-scale sampling module, a dilated 4-neighborhood difference convolution module, a dilated diagonal difference convolution module, and a two-dimensional convolutional layer, Figure 3 shows the schematic diagram of the processing of the dilated difference convolution module.

[0035] The specific processing process of the semi-supervised semantic segmentation network is as follows: The images in the training dataset are input into the network Resnet38 to obtain image features F. The image features F are input into the classification branch to output the classification activation map of the global semantic features and the class scores p c of each class. The image features F are input into the segmentation branch to output a multi-resolution activation map The activation map output by the classification branch and the activation map output by the segmentation branch are added and averaged to obtain the final multi-resolution semantic activation map Each pixel value in the figure is the activation value corresponding to the pixel at that position. Using the L1 norm to normalize the activation map so that the activation value of each pixel is between 0 and 1, and then using the Argmax function to find the maximum activation value corresponding to each pixel in the normalized activation map, forming a rough semantic segmentation pseudo-label I CAM , and using conditional random field (CRF) to perform segmentation refinement on the rough semantic segmentation pseudo-label to obtain a fine pseudo-label I CRF ; where, F3 represents the global semantic feature, and c represents the category number.

[0036] The specific processing process of the classification branch is as follows: The image feature F is input into the hybrid pooling module, and the global semantic feature F3 is output. The global semantic feature F3 passes through the Softmax function to obtain the classification activation map The global semantic feature F3 passes through the maximum pooling layer to obtain the category score p c ; where, the expression of the processing process of the hybrid pooling module is as follows:

[0037]

[0038] F2 = f conv (F1)(8)

[0039]

[0040] where, F1 represents the feature output by the local hybrid pooling module, F2 represents the feature output by the 1×1 convolutional layer, F3 represents the global semantic feature output by the global hybrid pooling module, AP p1 (·) represents the mean pooling layer operation, MP p1 (·) represents the maximum pooling layer operation, f conv (·) represents the 1×1 convolutional operation, GAP p2 (·) represents the global average pooling layer operation, GMP p2 (·) represents the global maximum pooling layer operation;

[0041] The specific processing process of the segmentation branch is as follows: The multi-scale sampling module performs sampling processing on the input image feature F at four scales s ∈ [0.5, 1, 1.5, 2]. After sampling, the features at each scale pass through the dilated 4-neighborhood difference convolutional module. Taking each pixel as the center, the difference convolution is performed on the center point feature and the features of the points in its second-order 4-neighborhood to obtain the context information of the corresponding scale feature After sampling, the features at each scale pass through the dilated diagonal difference convolutional module. Taking each pixel as the center, the difference convolution is performed on the center point feature and the features of the points in its second-order diagonal neighborhood to obtain the context information of the corresponding scale feature After sampling, the features at each scale pass through a two-dimensional convolutional layer to obtain the low-dimensional information of the corresponding scale features For the information of the features at each scale Sum and average them to obtain the final context features at each scale For the context features at each scale Use the Softmax function to obtain the class activation maps corresponding to each scale Sum and average all the class activation maps to obtain the final multi-resolution activation map

[0042] The loss function of the semi-supervised semantic segmentation network is expressed as follows:

[0043] L total = L cls + L seg (10)

[0044]

[0045]

[0046] Among them, L total represents the total loss of the network, L cls represents the classification branch loss, L seg represents the segmentation branch loss, c represents the class serial number, y c represents the classification label of the c-th class, I CRF represents the refined pseudo-label processed by the conditional random field (CRF), represents the multi-resolution activation map

[0047] 3. Data segmentation

[0048] Input the infrared image to be processed into the trained semi-supervised semantic segmentation network obtained in step 2, and output its cloud segmentation result

[0049] In this embodiment, the test set in step 1 is used as the dataset to be segmented and input into the trained network, and the segmentation result output by the network is compared and verified with its manually annotated result Figure 4 The original image example of the 38-cloud dataset is given Figure 5 is its corresponding manually annotated result Figure 6 is the segmentation result obtained by using the method of the present invention. It can be seen that the present invention has a high quality of cloud segmentation for infrared images

Claims

1. A semi-supervised infrared cloud segmentation method based on pixel context information mining, characterized in that The steps are as follows: Step 1, construct a training dataset: Divide the infrared image data into a training dataset and a test dataset according to a ratio of 9:

1. Among them, the training dataset includes a class-labeled dataset and an unlabeled dataset. The class-labeled dataset consists of labeled images and their class labels, and the ratio of the number of images contained in the class-labeled dataset to the unlabeled dataset is 1:4; Step 2, network training: Use the training dataset obtained in Step 1 to train the semi-supervised semantic segmentation network to obtain a trained network; The semi-supervised semantic segmentation network mainly includes a feature extraction backbone network Resnet38, a classification branch, and a segmentation branch. Among them, the network Resnet38 is pre-trained on ImageNet; The specific processing process of the semi-supervised semantic segmentation network is as follows: The images in the training dataset are input into the network Resnet38 to obtain image features F. The image features F are input into the classification branch, and the classification activation map of the global semantic features and the class scores p are output. and class scores p c The image features F are input into the segmentation branch, and the multi-resolution activation map is output. The activation map output by the classification branch and the activation map output by the segmentation branch are added and averaged to obtain the final multi-resolution semantic activation map. Each pixel value in the map is the activation value corresponding to the pixel at that position. The activation map is normalized using the L1 norm so that the activation value of each pixel is between 0 and 1. Then, the Argmax function is used to find the maximum activation value corresponding to each pixel in the normalized activation map, forming a rough semantic segmentation pseudo-label I. The conditional random field is used to refine the rough semantic segmentation pseudo-label to obtain the refined pseudo-label I. CAM ; where F3 represents the global semantic feature and c represents the class serial number. CRF ​ The classification branch includes a hybrid pooling module, a Softmax function, and a max pooling layer. The image feature F is input into the hybrid pooling module, and the global semantic feature F3 is output. The global semantic feature F3 passes through the Softmax function to obtain the classification activation map. The global semantic feature F3 passes through the max pooling layer to obtain the class score p. c The hybrid pooling module includes a local hybrid pooling module, a 1×1 convolutional layer, and a global hybrid pooling module. The local hybrid pooling module contains a parallel mean pooling layer AP p1 and a max pooling layer MP p1 . The global hybrid pooling module contains a parallel global average pooling layer GAP p2 and a global max pooling layer GMP p2 . The processing process expression of the hybrid pooling module is as follows: F2 = f conv (F1) (2) Among them, F1 represents the feature output by the local mixed pooling module, F2 represents the feature output by the 1×1 convolutional layer, F3 represents the global semantic feature output by the global mixed pooling module, and AP p1 (·) represents the mean pooling layer operation, MP p1 (·) represents the max pooling layer operation, f conv (·) represents the 1×1 convolution operation, GAP p2 (·) represents the global average pooling layer operation, GMP p2 (·) represents the global max pooling layer operation; The segmentation branch mainly includes a multi-scale sampling module, a 4-neighborhood dilated differential convolution module, a dilated diagonal differential convolution module, and a two-dimensional convolution layer. The specific processing process is as follows: The multi-scale sampling module performs sampling processing on the input image feature F at four scales s ∈ [0.5, 1, 1.5, 2]. After sampling, the features at each scale pass through the dilated 4-neighborhood differential convolution module. Taking each pixel as the center, the differential convolution is performed on the feature of the center point and the features of the points in its second-order 4-neighborhood to obtain the context information of the corresponding scale feature. After sampling, the features at each scale pass through the dilated diagonal differential convolution module. Taking each pixel as the center, the differential convolution is performed on the feature of the center point and the features of the points in its second-order diagonal neighborhood to obtain the context information of the corresponding scale feature. After sampling, the features at each scale pass through the two-dimensional convolution layer to obtain the low-dimensional information of the corresponding scale feature. For the information of each scale feature The sum is taken and averaged to obtain the final context feature of each scale. The context features of each scale Pass through the Softmax function to obtain the class activation map corresponding to each scale. The sum is taken and averaged for all class activation maps to obtain the final multi-resolution activation map. The loss function of the semi-supervised semantic segmentation network is expressed as follows: L total = L cls + L seg (4) Among them, L total represents the total network loss, and L cls represents the classification branch loss, and L seg represents the segmentation branch loss, and y c represents the classification label of the c-th class; Step 3, data segmentation: Input the infrared image to be processed into the trained semi-supervised semantic segmentation network obtained in Step 2, and output its cloud segmentation result.

Citation Information

Patent Citations

  • Semi-supervised remote sensing image semantic segmentation method and device, and computer equipment

    CN113298815A

  • CSM image segmentation method and apparatus, terminal device, and storage medium

    WO2022205657A1