An Infrared Image Semantic Segmentation Algorithm Based on Improved Deeplabv3+

Through the improved Deeplabv3+ network structure, combined with the hollow convolution layer, residual module and cross-resolution attention module, the problems of edge blur and texture information loss in infrared image segmentation are solved, and the precise segmentation effect of infrared images is achieved.

CN116486084BActive Publication Date: 2025-08-05ANHUI UNIV OF SCI & TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310466461.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-04-26
Publication Date
2025-08-05
Estimated Expiration
2043-04-26

AI Technical Summary

Technical Problem

When processing infrared images, existing semantic segmentation algorithms are prone to inaccurate segmentation due to edge blur and texture information loss, especially in low-resolution infrared image tasks.

Method used

The improved Deeplabv3+ network structure is adopted, combining the hollow convolution layer, residual module and cross-resolution attention module, and the dual resolution module aggregates low-level features and advanced features, and uses the GPU-friendly attention module and multi-axis gate module to capture local and global information to perform precise segmentation of infrared images.

Benefits of technology

The precise segmentation of infrared images is achieved, which avoids edge blur and texture information loss, and generates higher quality segmentation results.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116486084B_ABST
    Figure CN116486084B_ABST
Patent Text Reader

Abstract

An infrared image semantic segmentation algorithm based on an improved DeepLabv3+ specifically includes the following steps: using an infrared camera to capture images and construct a dataset; using DeepLabv3+ as a baseline network, using high-level features from the backbone network as input to a dilated convolutional layer, and performing a first upsampling on the output of the dilated convolutional layer; using the high-level features from the first upsampling and low-level features processed by a residual module as input to a dual resolution module, which aggregates and outputs the low-level features and high-level features, performs a second upsampling, and then performs a splicing operation with the low-level features, and finally performs a bilinear interpolation operation to output the result; training the infrared image semantic segmentation network; and using a validation set for verification, parameter adjustment, and selection of the optimal model. The present invention can achieve the purpose of accurately segmenting the input image and avoid inaccurate segmentation caused by reasons such as blurred edges and loss of texture information when processing infrared images.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to vision technology and belongs to the field of image processing, and particularly to an infrared image semantic segmentation algorithm based on improved DeepLabv3+. Background Art

[0002] Image semantic segmentation is the process of classifying each pixel in a given image to achieve both object segmentation and classification. It has important applications in areas such as autonomous driving, pose estimation, image search, and medical diagnosis. For example, in autonomous driving, image semantic segmentation can equip cars with the necessary perception capabilities, allowing them to "observe" road conditions and the surrounding environment, enabling autonomous vehicles to safely navigate the road.

[0003] Currently, semantic segmentation algorithms are mostly targeted at visible light images. Compared to visible light images, infrared images can store more information, have stronger penetration, and are not affected by harsh environments (rain, snow, fog, etc.) and lighting. However, there is little research on semantic segmentation for infrared images.

[0004] Infrared image semantic segmentation and visible light image semantic segmentation are similar in that both aim to classify each pixel in a given image, thereby achieving the effect of target segmentation and classification. However, there are still some problems in semantic segmentation of infrared images, such as the blurred edges and serious loss of texture information in infrared images. Current semantic segmentation algorithms are prone to losing edge information when processing low-resolution infrared image-related tasks, and therefore cannot accurately segment them. Summary of the Invention

[0005] The purpose of the present invention is to provide an infrared image semantic segmentation algorithm based on the improved Deeplabv3+, so as to achieve the purpose of accurately segmenting the input image and avoid inaccurate segmentation caused by reasons such as edge blur and texture information loss when segmenting infrared graphics.

[0006] To achieve the above objectives, this paper proposes an infrared image semantic segmentation algorithm based on the improved Deeplabv3+, which is characterized by comprising the following steps:

[0007] S1: Collect images with an infrared camera and build a dataset;

[0008] S2: Constructing the semantic segmentation network structure: Using DeepLabv3+ as the baseline network, the high-level features of the backbone network are used as the input of the dilated convolution layer, and the output of the dilated convolution layer is subjected to the first upsampling;

[0009] The high-level features after the first upsampling and the low-level features processed by a residual module are used as the input of the dual resolution module. The dual resolution module aggregates and outputs the low-level features and high-level features. After the second upsampling, they are concatenated with the low-level features and finally output after bilinear interpolation.

[0010] S4: training infrared image semantic segmentation network;

[0011] S5: Use the validation set to verify, adjust parameters, and select the optimal model;

[0012] S6: Use the test set to test the selected optimal model and evaluate the model performance.

[0013] In some embodiments, in step S2, the low-level features processed by a residual module are used as input to the low-resolution branch, and the high-level features processed by the first upsampling are used as input to the high-resolution branch;

[0014] The dual-resolution module has a GPU-friendly attention module that acts on the low-resolution branch to capture high-level global context information, and a cross-resolution attention module that acts on the high-resolution branch to propagate high-level global context information;

[0015] GFA uses a multi-axis gating module to capture local and global information of input features in parallel.

[0016] In some embodiments, the GPU-friendly attention module is a set of matrix operations, and the calculation formula is:

[0017]

[0018] Where X∈R N×d represents the input feature, N is the number of pixels in the image, and d is the feature dimension;

[0019] K g ,V g is a learnable matrix, M g =M×H, where M is the parameter dimension, H is the number of heads, and GDN() stands for double group renormalization, which splits the second normalization of the original double normalization into H groups.

[0020] In some embodiments, the process formula of the cross-resolution attention module is:

[0021]

[0022] K c ,V c =φ(X c )

[0023] X c =θ(X l )

[0024] Among them, X h 、X l denote the feature maps on the high-resolution branch and the low-resolution branch respectively, φ() is a set of matrix operations, including splitting, permutation and reshaping, d h is the feature dimension of the high-resolution branch; θ() is a function composed of the pooling layer and the convolution layer, X c It's X l The cross-features calculated by the θ() function, K c and V c are two sets of learnable matrices.

[0025] In some embodiments, the multi-axis gating module has two sets of convolution operations, where in the local branch, the feature half-head of size (H, W, C / 2) is blocked into a tensor of shape (H / b×W / b, b×b, C / 2);

[0026] In the global branch, the other half of the head is meshed into shape (d×d, H / d×W / d, C / 2) using a fixed (d×d) grid, and each window has size (H / d×W / d);

[0027] A gated MLP (gMLP) block is applied on a single axis of each branch to make it fully convolutional and share parameters on other spatial axes, so that operations in local and global branches are performed in parallel.

[0028] In some embodiments, in step S1, the infrared camera may be a Hikvision thermal imaging dual-spectrum hemispheric camera, and when collecting infrared images, the open source annotation tool labelme is used for annotation;

[0029] The dataset is divided into training set, validation set and test set in the ratio of 1:0.9:0.1.

[0030] Compared with the existing technology, this infrared image semantic segmentation algorithm based on the improved DeepLabv3+ uses DeepLabv3+ as the baseline network. The high-level features after the first upsampling and the low-level features processed by a residual module are used as the input of the dual resolution module. The dual resolution module aggregates and outputs the low-level features and high-level features to achieve the purpose of accurate segmentation of the input image, avoiding the inaccurate segmentation caused by edge blur and texture information loss when processing infrared images.

[0031] In some embodiments, GFA is used in the low-resolution branch, and multi-axis gated MAGBlock is used to capture local and global information in parallel. CRA is used in the high-resolution branch. The dual-resolution module aggregates the features of the two resolutions through the corresponding attention module to extract the edge information and global information of the infrared image, so as to achieve the purpose of accurate segmentation of the input image and generate higher quality renderings. BRIEF DESCRIPTION OF THE DRAWINGS

[0032] Figure 1 It is the overall network structure diagram of the present invention;

[0033] Figure 2 It is the MRBlock process structure diagram of the present invention;

[0034] Figure 3 It is a structural diagram of the MAGBlock module in the present invention;

[0035] Figure 4 It is a structural diagram of the GFA module in the present invention;

[0036] Figure 5 This is a comparison chart of the detection effects of the three exemplary models in the present invention. DETAILED DESCRIPTION

[0037] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the present invention.

[0038] like Figure 1 As shown in the figure, this infrared image semantic segmentation algorithm based on the improved Deeplabv3+ includes the following steps:

[0039] S1: Collect images with an infrared camera and build a dataset;

[0040] Specifically, the infrared camera can use Hikvision's thermal imaging dual-spectrum hemispheric camera, model DS-2TD1217-3 / PA. When collecting infrared images, the open source annotation tool labelme is used for annotation.

[0041] The dataset is divided into training set, validation set and test set in the ratio of 1:0.9:0.1 and used as network input;

[0042] S2: Constructing the semantic segmentation network structure: Using DeepLabv3+ as the baseline network, the high-level features of the backbone network are used as the input of the dilated convolution layer, and the output of the dilated convolution layer is subjected to the first upsampling;

[0043] The high-level features after the first upsampling and the low-level features processed by a residual module are used as the input of the dual resolution module (MRBlock for short). The dual resolution module aggregates the low-level features and high-level features and outputs them. The output of the dual resolution module is upsampled for the second time, then spliced with the low-level features, and finally output after bilinear interpolation.

[0044] Specifically, the backbone network can use a MobileNetv2 network composed of ResNet modules to perform feature extraction and obtain high-level features. To achieve context aggregation at different levels, retain more spatial information, and solve the problem of multi-scale object segmentation, the DeepLabv3+ baseline network can use cascaded or parallel dilated convolutions to capture multi-scale context by adopting multiple dilated rates.

[0045] S4: training infrared image semantic segmentation network;

[0046] S5: Use the validation set to verify, adjust parameters, and select the optimal model;

[0047] S6: Use the test set to test the selected optimal model and evaluate the model performance.

[0048] In some embodiments, in step S2, the low-level features processed by a residual module are used as input to the low-resolution branch, and the high-level features processed by the first upsampling are used as input to the high-resolution branch;

[0049] MRBlock has a GPU-friendly attention module (GFA) that acts on low-resolution branches to capture high-level global context information, and a cross-resolution attention module (CRA) that acts on high-resolution branches to propagate high-level global context information;

[0050] Among them, GFA uses a multi-axis gating module MAGBlock to capture local and global information of input features in parallel;

[0051] The MRBlock module aggregates the low-level features of the MAGBlock and the high-level features of the CRA module through the attention module.

[0052] Specifically, refer to Figure 2 , low-level features serve as input to the low-resolution branch, and high-level features serve as input to the high-resolution branch;

[0053] The low-level features are first processed by the GFA module ①, added to the initial downsampling operation of the high-level features ②, then go through convolution processing ③, multi-axis gating module MAGBlock, convolution processing ④, and then undergo initial upsampling operation ⑥, and then go through the GFA module again ①, add to the second downsampling operation ② of the high-level features, convolution processing ③, multi-axis gating module MAGBlock, convolution processing ④, and finally upsampling operation ⑥;

[0054] The high-level features are first downsampled (②), convolved (⑤), and added to the low-level features (⑥). After CRA processing (⑦), they are again downsampled (②) and convolved (⑤). Finally, they are added to the low-level features (⑥) for fusion processing (⑧) and output.

[0055] In this embodiment, GFA is used for the low-resolution branch, and multi-axis gated MAGBlock is used to capture local and global information in parallel. CRA is used for the high-resolution branch. The MRBlock module aggregates the features of the two resolutions through the corresponding attention module to achieve the purpose of accurate segmentation of the input image, avoiding inaccurate segmentation due to blurred edges, loss of texture information, etc. in infrared graphics.

[0056] In some embodiments, the GFA module is a set of matrix operations that replace the grouped matrix operations in the multi-head mechanism with ordinary matrix operations and introduce group normalization;

[0057] Let X∈R N×d represents the input features, where N is the number of elements (or pixels in the image) and d is the feature dimension. The original external attention module (EA for short) is calculated as follows:

[0058] EA(X,K,V)=DN(X·K T )·V

[0059] Where K,V∈R M×d is a learnable parameter, M is the parameter dimension, and DN() is a double normalization operation.

[0060] refer to Figure 4 (a) Multi-head EA has two matrix multiplications. The multiple heads obtained by channel slicing at the input undergo the first matrix multiplication I, then the attention calculation, and then the second matrix multiplication II, and finally spliced together for output. The calculation formula of multi-head EA is as follows:

[0061] MHEA(X)=Concat(h1,h2,...,h H )

[0062] hi =EA(X i ,K',V'),i∈[1,H]

[0063] Where K', V'∈R M×d ', d' = d / H, H is the number of heads, X i is the i-th head of X, Concat() is the concatenation function;

[0064] The present invention uses GFA without slicing and directly calculates the attention with K matrix, such as Figure 4 As shown in (b), the input does not undergo slicing operations, but directly performs the initial matrix multiplication III with the K matrix, then performs the attention calculation, and finally directly outputs it after another matrix multiplication IV. GFA simplifies the operation and makes good use of the GPU's advantage of large matrix parallel speed. The GFA calculation formula is as follows:

[0065]

[0066] Among them, K g ,V g is a learnable matrix, M g =M×H, GDN() represents double group renormalization, which splits the second normalization of the original double normalization into H groups;

[0067] In some embodiments, the CRA module is a cross-attention mechanism that propagates the global features learned on the low-resolution branch to the high-resolution branch. That is, the high-level features on the high-resolution branch are subjected to convolution processing ⑤, added with the initial upsampling operation ⑥ of the low-level features, and then subjected to CRA processing ⑦ to complete the propagation of the global features to the high-resolution branch. The process formula is:

[0068]

[0069] K c ,V c =φ(X c )

[0070] X c =θ(X l )

[0071] Among them, X h 、X l denote the feature maps on the high-resolution branch and the low-resolution branch respectively, φ() is a set of matrix operations, including splitting, permutation and reshaping, d h is the feature dimension of the high-resolution branch; θ() is a function composed of the pooling layer and the convolution layer, X c It's X lThe cross-features calculated by the θ() function, K c and V c are two sets of learnable matrices;

[0072] refer to Figure 3 In some embodiments, MAGBlock captures local and global information of input features in parallel. Figure 2 The low-level features after convolution operation ③ are used as the input of MAGBlock, such as Figure 3 As shown, MAGBlock has two sets of convolution operations. The LN operation doubles the number of input feature channels and then divides them into two halves for local feature extraction and global feature extraction respectively. In the local branch, the features of size (H, W, C / 2) are divided into blocks of shape (H / b×W / b, b×b, C / 2) tensors; in the global branch, a fixed (d×d) grid is used to grid the other half into a tensor of shape (d×d, H / d×W / d, C / 2), and each window has a size of (H / d×W / d);

[0073] A gated MLP (gMLP) block is applied on a single axis of each branch to make it fully convolutional, that is, a set of residual operations, while sharing parameters on other spatial axes. The operations in the local branch and the global branch are performed in parallel, that is, the local and global information of the input features are obtained synchronously, and finally the splicing operation is performed.

[0074] like Figure 5 As shown in the figure, three embodiments were captured by Hikvision thermal imaging dual-spectrum hemispherical camera, where a is the original image, b is annotated using the open source annotation tool labelme; c is the test set image of the segmentation after training before improvement, and d is the semantic segmentation test set image after network training of Deeplabv3+ after improvement of the present invention; by comparing the test set images in c and d, the three embodiments in c have the problems of blurred edges and serious loss of texture information, such as Figure 5 In each example, the segmentation results in c have the following problems compared to b: 1. Inaccurate segmentation targets, 2. Blurred edges, and 3. Incomplete segmentation targets. However, the corresponding locations in d have better segmentation results compared to b. It is clearly observed that the segmentation targets in the test set image in d are more accurate and the edge segmentation is more complete. Therefore, the improved model of this invention can output more accurate segmentation results compared to DeepLabv3+.

[0075] The above description is merely an exemplary embodiment of the present invention and is not intended to limit the scope of protection of the present invention. The scope of protection of the present invention is determined by the appended claims.

[0076] The embodiments described above merely represent several implementations of the present invention, and their descriptions are relatively specific and detailed, but they should not be construed as limiting the scope of the patented invention. It should be noted that while the present invention has been shown and described with reference to various embodiments, it will be apparent to those skilled in the art that various variations and improvements in form and detail may be made without departing from the spirit of the present invention, and without departing from the scope of the present invention as defined by the appended claims, all of which fall within the scope of protection of the present invention. Therefore, the scope of protection of the patented invention shall be based on the appended claims.

Claims

1. An infrared image semantic segmentation algorithm based on improved Deeplabv3+, characterized by: The specific steps include: S1: Collect images with an infrared camera and build a dataset; S2: Constructing the semantic segmentation network structure: Using DeepLabv3+ as the baseline network, the high-level features of the backbone network are used as the input of the dilated convolution layer, and the output of the dilated convolution layer is subjected to the first upsampling; The high-level features after the first upsampling and the low-level features processed by a residual module are used as the input of the dual resolution module. The dual resolution module aggregates and outputs the low-level features and high-level features. After the second upsampling, they are concatenated with the low-level features and finally output after bilinear interpolation. S4: training infrared image semantic segmentation network; S5: Use the validation set to verify, adjust parameters, and select the optimal model; S6: Use the test set to test the selected optimal model and evaluate the model performance; In step S2, the low-level features processed by a residual module are used as the input of the low-resolution branch, and the high-level features processed by the first upsampling are used as the input of the high-resolution branch; The dual-resolution module has a GPU-friendly attention module that acts on the low-resolution branch to capture high-level global context information, and a cross-resolution attention module that acts on the high-resolution branch to propagate high-level global context information; The GPU-friendly attention module uses a multi-axis gating module to capture local and global information of input features in parallel.

2. The infrared image semantic segmentation algorithm based on the improved Deeplabv3+ according to claim 1, characterized in that: The GPU-friendly attention module is a set of matrix operations, and the calculation formula is: Where X∈R N×d represents the input feature, N is the number of pixels in the image, and d is the feature dimension; K g ,V g is a learnable matrix, M g =M×H, where M is the parameter dimension, H is the number of heads, and GDN() stands for double group renormalization, which splits the second normalization of the original double normalization into H groups.

3. The infrared image semantic segmentation algorithm based on the improved Deeplabv3+ according to claim 2, characterized in that: The process formula of the cross-resolution attention module is: K c ,V c =φ(X c ) X c =θ(X l ) Among them, X h 、X l denote the feature maps on the high-resolution branch and the low-resolution branch respectively, φ() is a set of matrix operations, including splitting, permutation and reshaping, d h is the feature dimension of the high-resolution branch; θ() is a function composed of the pooling layer and the convolution layer, X c It's X l The cross-features calculated by the θ() function, K c and V c are two sets of learnable matrices.

4. The infrared image semantic segmentation algorithm based on the improved Deeplabv3+ according to claim 3, characterized in that: The multi-axis gating module has two sets of convolution operations. In the local branch, the feature half-head of size (H, W, C / 2) is blocked to a shape tensor of shape (H / b×W / b, b×b, C / 2); In the global branch, the other half of the head is meshed into shape (d×d, H / d×W / d, C / 2) using a fixed (d×d) grid, and each window has size (H / d×W / d); A gated MLP (gMLP) block is applied on a single axis of each branch to make it fully convolutional and share parameters on other spatial axes. Operations in local and global branches are performed in parallel.

5. The infrared image semantic segmentation algorithm based on the improved Deeplabv3+ according to any one of claims 1 to 4, characterized in that: In step S1, the infrared camera can use a Hikvision thermal imaging dual-spectrum hemispheric camera. When collecting infrared images, the open source annotation tool labelme is used for annotation. The dataset is divided into training set, validation set and test set in the ratio of 1:0.9:0.1.

Citation Information

Patent Citations

  • Glandular cell image segmentation method based on selective multi-branch cavity convolution

    CN115457061A

  • Apparatus and method for segmenting medical image using MLP based architecture

    KR102419270B1