Infrared image real-time semantic segmentation method based on wavelet pooling three-branch network

By introducing wavelet pooling module and three-branch network structure into the deep learning real-time semantic segmentation algorithm, the problem of information loss during downsampling is solved, and the accuracy and efficiency of infrared image segmentation are improved.

CN120070890AInactive Publication Date: 2025-05-30HENAN UNIV OF SCI & TECH
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510134725.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-07
Publication Date
2025-05-30
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

Existing deep learning real-time semantic segmentation algorithms are prone to lose key information in infrared images during downsampling, affecting segmentation accuracy.

Method used

Using a three-branch network based on wavelet pooling, the Haar wavelet pooling module reduces information loss during the downsampling process, retains significant parts of the features, and introduces pixel attention-guided fusion module and parallel aggregation module into the three-branch network to improve segmentation accuracy.

Benefits of technology

In the infrared image segmentation task, the segmentation accuracy is improved, the real-time processing needs are met, and a good balance is achieved in terms of parameter quantity and calculation quantity.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120070890A_ABST
    Figure CN120070890A_ABST
Patent Text Reader

Abstract

The invention relates to the field of deep learning infrared image segmentation, in particular to an infrared image real-time semantic segmentation method based on a wavelet pooling three-branch network, and the method comprises the following steps: S1, carrying out the modeling of interference radiation and motion rules on the basis of a public infrared target data set, and making an infrared airplane data set under interference; s2, designing a Haar wavelet pooling module, and keeping key frequency domain information while reducing the resolution of the feature map by means of the advantage of frequency decomposition of the Haar wavelet pooling module; and S3, constructing a three-branch real-time semantic segmentation network model for respectively extracting space, semantic and boundary information. And S4, introducing a Haar wavelet pooling module into the three-branch network model as a down-sampling layer of the model to reduce information loss during down-sampling. On an infrared aircraft data set under self-made interference, the method not only can meet the real-time processing requirement of the infrared image, but also can improve the segmentation precision.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of deep learning infrared image segmentation, and specifically to a real-time semantic segmentation method for infrared images based on a wavelet pooling three-branch network. Background Art

[0002] Infrared imaging guidance is a technology that uses an infrared sensor to obtain the thermal radiation information of a target to achieve precise guidance. Its basic principle is to detect the infrared radiation of the target through an infrared imaging device, calculate and control the guidance device to approach the target, and finally achieve effective destruction. Infrared imaging technology has the advantages of strong anti-interference ability, long operating distance, being unaffected by weather conditions, working all-weather, and a powerful ability to penetrate smoke. These advantages enable it to provide accurate target information for the guidance system in complex environments. In this process, infrared image segmentation is a key technology in infrared automatic target recognition and is the basis for target feature extraction, recognition, and tracking. Therefore, under infrared imaging data, how to quickly and accurately segment infrared targets such as airplanes and ships has important practical significance for improving the guidance accuracy.

[0003] However, currently, convolutional neural networks generally use downsampling operations such as pooling or strided convolution to aggregate local features, expand the receptive field, and minimize the computational cost. However, these downsampling operations may cause the loss of important spatial details. Especially in pixel-level prediction tasks, the lack of spatial details will significantly affect the accuracy of semantic segmentation. And due to the characteristics of scarce texture features, weak signals, and susceptibility to noise interference in infrared images themselves, key information is particularly likely to be lost during the downsampling process.

[0004] To solve this problem, the present invention proposes a real-time semantic segmentation algorithm for infrared images based on a wavelet pooling three-branch network, which further improves the segmentation accuracy on the basis of meeting the real-time segmentation requirements of infrared images. Summary of the Invention

[0005] The purpose of the present invention is to provide a real-time semantic segmentation method for infrared images based on a wavelet pooling three-branch network to solve the problem that key information of the target is easily lost during the downsampling process of the real-time semantic segmentation algorithm based on deep learning, thereby affecting the segmentation accuracy. To achieve the above purpose, the present invention provides the following technical solutions: A real-time semantic segmentation method for infrared images based on a wavelet pooling three-branch network, including the following steps:

[0006] S1: Simulate and model the laws of interference such as radiation, motion, and shape, and make an infrared aircraft dataset under interference on the basis of a public infrared target dataset;

[0007] S2: Design the Haar wavelet pooling module. By leveraging its advantage in frequency decomposition, reduce the information loss during the downsampling process when reducing the resolution of the feature map, and maximize the retention of the significant parts of the features.

[0008] S3: Construct a three-branch real-time semantic segmentation network model that extracts spatial, semantic, and boundary information, achieving the best balance between the model inference speed and segmentation accuracy.

[0009] S4: Introduce the Haar wavelet pooling module into the three-branch network as the downsampling layer of the network, forming a real-time semantic segmentation network for infrared images based on the wavelet pooling three-branch network, which is more suitable for scenarios with scarce detailed information such as infrared images.

[0010] S5: Validate the real-time semantic segmentation network model for infrared images based on the wavelet pooling three-branch network on the self-made infrared aircraft dataset under interference. This model can not only meet the real-time processing requirements of infrared images, but also improve the segmentation accuracy compared with other classic real-time semantic segmentation networks.

[0011] The self-made infrared aircraft dataset under interference is derived from the publicly available infrared dataset LSOTB-TIR. Since the images in this dataset have a relatively high resolution and the aircraft targets have contours and certain physical characteristics, 22,043 aircraft images are selected from this dataset. Based on this, combined with the radiation characteristics and operation forms of the interference, artificial interference is added to create the infrared aircraft dataset under interference.

[0012] When simulating the interference, the infrared radiation characteristics after the interference is released follow a Gaussian distribution, that is, the gray value in the central region is large and gradually weakens outward. And under the action of inertia and its own gravity, the interference may drop straight down or move around in a parabolic form, gradually moving away from the aircraft.

[0013] The said Haar wavelet pooling module decomposes the image into four different components through two-dimensional Haar wavelet transform. Subsequently, the components are merged, and the merged feature map is further processed by the feature enhancement module. That is, first use two-dimensional Haar wavelet transform to reduce the resolution of the feature map and retain the detailed information. In this process, the image undergoes two-stage downsampling. In the first stage, apply the low-pass filter h low , high-pass filter h high and the downsampling operation to decompose the image of size C×H into low-frequency and high-frequency components along the horizontal axis.

[0014] In the second stage, a low-pass filter and a high-pass filter in the vertical direction are respectively applied to the obtained low-frequency component and high-frequency component, and further decomposed into four components, which respectively correspond to the low-frequency information of the image and the high-frequency information in different directions. The low-frequency component (LL) represents the overall structure and main information of the image, mainly highlighting the main body contour, such as large-range brightness change areas like the fuselage and interference; the horizontal component (LH) mainly captures the edge and detail changes in the horizontal direction, mainly highlighting the horizontal features of the aircraft fuselage and propeller; the vertical component (HL) focuses on extracting the changes in the vertical direction, such as the wing edge and the vertical texture in the background; the diagonal component (HH) reflects the high-frequency changes along the diagonal direction, capturing the intersection points of the aircraft propeller and its complex edge features. The size of each component is H / 2×W / 2, and these four components are stacked along the channel direction to form a new feature map. The Haar wavelet transform operation effectively encodes part of the information in the spatial dimension into the channel dimension, ensuring lossless transmission of information.

[0015] After completing the wavelet decomposition, these features are further enhanced through a feature enhancement module. First, the number of channels of the new feature map formed by stacking along the channels is adjusted through 1×1 convolution. Then, batch normalization and ReLU activation function are used to further enhance the expression ability and fitting performance of the model. The feature enhancement process is shown as follows:

[0016] F 0 =σ(BN(conv(F)))

[0017] where conv(.) represents a 1×1 convolution, BN represents batch normalization, σ(.) represents the ReLU activation function, and F represents the feature map stacked along the channel direction.

[0018] Through this design that combines wavelet decomposition and feature enhancement, the Haar wavelet pooling module can not only retain the key information in the image during the downsampling process, but also make full use of the details in different frequency directions, enhancing the diversity and integrity of the features, thereby improving the performance of the model in various scenarios.

[0019] As Figure 4 shown, the output feature map of the HWP module is compared with the other three common downsampling methods (max pooling, average pooling, and strided convolution). The original image is fed into a simple convolutional neural network. After two rounds of convolution, batch normalization, ReLU activation, and downsampling operations, four different output feature maps are obtained based on different downsampling methods. Compared with the other three downsampling methods, HWP can more clearly retain the edge information of the object, such as the contour of the propeller blade and the edge of the main body structure.

[0020] The three-branch real-time semantic segmentation network model first enters the shared branch to perform preliminary feature extraction on the image. Subsequently, the feature map downsampled by 1 / 8 is sent to three branches to extract information from feature maps of different resolutions. The spatial branch is responsible for retaining detailed information in the high-resolution feature map. The semantic branch I aggregates local and global context information to resolve long-range dependencies. The boundary branch extracts high-frequency features to predict the boundary region and guides the fusion of semantic and spatial information by integrating boundary information. Among them, the loss of the entire model includes the semantic loss l 0 、l 2 , the boundary loss l 1 and the boundary-aware cross-entropy loss l 3 .

[0021] The model introduces a Pixel-attention-guided fusion module (Pag) for the spatial branch to selectively extract useful semantic features from the semantic branch and avoid information inundation. In the semantic branch, a Parallel Aggregation PPM (PAPPM) module is proposed to capture rich global semantic information from multiple scales. Finally, a Boundary-attention-guided fusion module (Bag) is used to fuse the information of the three branches, further improving the segmentation accuracy.

[0022] Introducing the Haar wavelet pooling module into the three-branch network model, there are a total of six downsampling layers in the original three-branch network framework from Stage-1 to Stage-5. The sizes of the feature maps obtained by each downsampling layer are 1 / 2, 1 / 4, 1 / 8, 1 / 16, 1 / 32, and 1 / 64 of the original image respectively. Using the HWP module for downsampling of the feature maps in Stage-1 to Stage-5 can retain more high-frequency information while effectively reducing the dimension of the input image, thereby retaining more detailed features.

[0023] Compared with the prior art, the beneficial effects of the present invention are as follows:

[0024] In the present invention, by introducing the Haar wavelet pooling module as the downsampling layer in the three-branch network framework, the significant part of the feature can be retained to the greatest extent while reducing the resolution of the feature map. On the infrared aircraft dataset under self-made interference, the infrared image real-time semantic segmentation network model based on the wavelet pooling three-branch network is verified. This model can not only meet the real-time processing requirements of infrared images, but also improve the segmentation accuracy compared with other classical real-time semantic segmentation networks. Brief Description of the Drawings

[0025] Figure 1 This is the flowchart of the present invention;

[0026] Figure 2 This is the schematic diagram of the three-branch network framework structure based on wavelet pooling of the present invention;

[0027] Figure 3 This is the schematic diagram of the HWP module structure of the present invention;

[0028] Figure 4 This is the example diagram of the downsampling operation of the present invention;

[0029] Figure 5 This is part of the images of the anti-jamming infrared aircraft dataset of the present invention;

[0030] Figure 6 This is the comparison diagram of the detailed shallow-layer features of the network of the present invention;

[0031] Figure 7 This is the heat map of the output of the PAPPM module after different downsampling operations of the present invention;

[0032] Figure 8 This is the comparison diagram of the segmentation results of different algorithms of the present invention. Detailed implementation manners

[0033] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0034] Embodiment

[0035] 1. Experimental platform and parameter settings:

[0036] In this embodiment, the experimental operating system is Windows10, the CPU model is Intel(R) Core(TM) i7-10700 CPU @ 2.90GHz, the running memory is 32GB, the GPU model is NVIDIA3090, and the video memory size is 24GB. All algorithms are built and run based on the Pytorch framework. The programming language is Python3.8. CUDA11.3.1 and CUDNN8.2.1 are used to accelerate the GPU. During the algorithm training stage, the image resolution is uniformly set to 640×640, the batch-size is set to 8, the network uses the Adam optimizer, the learning rate is 0.004, the weight decay is 0.0001, and a total of 100 rounds of training are performed. The learning rate gradually decreases linearly;

[0037] 2. Dataset production:

[0038] The experimental data of this embodiment is sourced from the publicly available infrared dataset LSOTB-TIR. The images in this dataset have a relatively high resolution, and the aircraft targets have contours and certain physical characteristics. 22,043 aircraft images are selected from this dataset. Based on this, combined with the radiation characteristics and operating patterns of the interference, interference is artificially added to create an infrared aircraft dataset under interference. As Figure 5 shown, the first row represents the original infrared aircraft data in LSOTB-TIR, and the second row represents the infrared aircraft data after artificially adding interference. The added interference follows the following rules:

[0039] (1) The radiation of the interference is modeled according to a Gaussian distribution, that is, the gray value is the largest at the center and gradually weakens outward;

[0040] (2) After the interference is released, it first increases from small to large, then decreases from large to small, and gradually moves away from the aircraft;

[0041] (3) The shape is modeled by the equation of an irregular circle. Using the labelme software, annotation is achieved by generating a JSON file containing target pixel category information and position information. For the infrared aircraft targets under interference, a total of four categories are divided during annotation and represented by different colors: red represents the nose (handpiece), green represents the tail (empennage), yellow represents the propeller (airscrew), and blue represents the interference (interference). The dataset is divided into a training set, a validation set, and a test set according to a ratio of 8:1:1;

[0042] 3. Evaluation metrics:

[0043] (1) Mean Intersection over Union (MIoU)

[0044] MIoU first calculates the intersection over union for each class, and then takes the mean of the intersection over unions of all classes, as shown in the following formula:

[0045]

[0046] (2) Mean Pixel Accuracy (MPA)

[0047] MPA first calculates the pixel accuracy for each class, and then takes the mean of the pixel accuracies of all classes, as shown in the following formula:

[0048]

[0049] (3) Frames per Second (FPS)

[0050] The FPS represents the number of frames per second and is used to evaluate the inference speed of the model, as shown in the following formula:

[0051]

[0052] 4. Ablation experiments:

[0053] 4.1 Analysis of the effectiveness of HWP:

[0054] To verify the effectiveness of the HWP module, the present invention uses strided convolution, average pooling, max pooling, and the proposed HWP module as the downsampling layers in the PIDNet network respectively, and the obtained segmentation results are shown in Table 1.

[0055] Table 1 Segmentation results of different pooling algorithms of the present invention

[0056] Model MIoU / % MPA / % Parameters / MB GFLOP / s FPS frames / s PIDNet 90.14 93.85 27.21 3482 3457 PIDNet_Maxpool 89.83 93.59 23.62 3234 38.3 PIDNet_Averagepool 90.03 93.78 23.62 32.34 37.93 PIDNet HWP 90.49 94.47 25.13 33.29 31.44

[0057] It can be seen that in terms of the number of parameters and computational complexity, the PIDNet network using max / average pooling is superior to the PIDNet network using HWP. However, compared with the original PIDNet network using strided convolution, the number of parameters of the PIDNet network using the HWP module has decreased by 7.6 percentage points, and the computational complexity has decreased by 4.4 percentage points. Although the FPS is slightly lower than that of other downsampling algorithms, it still meets the real-time requirement.

[0058] In terms of accuracy, compared with the other three downsampling methods, the PIDNet algorithm using the HWP module reaches the highest in both MIoU and MPA metrics. Compared with the original model using strided convolution, MIou has increased by 0.35% to reach 90.49%; MPA has increased by 0.62% to reach 94.47%.

[0059] Thus, different from methods such as strided convolution and pooling that cause information compression, the HWP module decomposes and recombines high-frequency and low-frequency information by means of Haar wavelet transform, effectively retaining details, reducing information loss, and improving segmentation accuracy. In summary, the HWP module achieves a good balance between segmentation accuracy and computational efficiency while meeting the real-time requirement.

[0060] Figure 6 Shows the detailed comparison effect of the shallow feature maps obtained after downsampling twice by the PIDNet algorithm with and without using the HWP module. It can be seen that the algorithm using the HWP module retains more edge detail information in the propeller and tail parts, further verifying its advantage in detail retention.

[0061] Figure 7It shows the heatmaps extracted after downsampling using strided convolution and the HWP module in the PIDNet algorithm. Among them, the first row represents the original infrared image, the second row represents the heatmap output by the parallel aggregation PPM module after downsampling using strided convolution in the original model, and the third row represents the heatmap output by the PAPPM module after downsampling using the HWP module. The heatmaps are generated by the Grad-CAM++ algorithm, which shows the contributions of different regions to the model prediction. The red and yellow regions represent the regions with the greatest and relatively large impacts on the model segmentation decision, while the blue and green regions represent the regions with relatively small impacts. By analyzing these color regions, the key feature regions that the real-time semantic segmentation algorithm focuses on during the decision-making process can be intuitively revealed.

[0062] From Figure 7 It can be seen that in the first column of data, compared with strided convolution, the heatmap after downsampling by the HWP module shows more accurate feature boundaries, reduces background interference, and makes the outline of the aircraft clearer. In the second column of data, the model using the HWP module significantly enhances the feature responses of the fuselage and the tail. Compared with the model using strided convolution, the heat is more concentrated on the key parts of the aircraft, reducing the interference of the background area on the decision-making. In the third column of data, the features in the fuselage and propeller areas are more prominent and the feature responses are stronger after downsampling by the HWP module, which makes the model focus more on the target. In the fourth column of data, the model using the HWP module further clarifies the target boundary and the model pays more attention to the aircraft target.

[0063] 4.2 Analysis of the number of HWP:

[0064] The original PIDNet algorithm uses ResNet as the feature extraction network, constructs the main trunk before the branch and the subsequent three branches of space, semantics, and boundary. The downsampling operation is mainly concentrated in the semantic branch, and the feature dimension is gradually reduced through multiple downsamplings to extract high-level semantic information. In the present invention, the downsampling layer in each stage of the ResNet architecture is gradually replaced by the HWP module, and the segmentation results are shown in Table 2.

[0065] Table 2 Analysis of the number of HWP in the present invention

[0066] No. of HWP MIoU / % MPA / % Parameters / MB GFLOP / s 0 90.14 93.85 27.21 34.82 2 90.29 94.11 27.19 34.2 3 90.27 93.95 27.16 33.94 4 90.45 94.01 27.00 33.68 5 90.39 93.88 26.37 33.42 6 90.49 94.47 25.13 33.29

[0067] Specifically, the number of HWPs in the table being 0 means all downsampling layers of the original architecture are retained without any replacement; 2 means replacing two convolutional layers with a stride of 2 in the backbone before the branch with HWP modules; 3 means replacing the downsampling layers in the backbone before the branch and the second layer of the ResNet architecture with HWP modules; 4 means replacing the downsampling layers including those in the backbone before the branch, the second layer, and the third layer with HWP modules; 5 means replacing the downsampling layers in the backbone before the branch, the second layer, the third layer, and the fourth layer with HWP modules; 6 means global replacement.

[0068] As can be seen from Table 2, compared with the original PIDNet algorithm without using HWP modules, using HWP as the downsampling layer can improve the segmentation performance. In addition, it is observed that the operation of replacing all downsampling layers in the ResNet architecture is more effective than the operation of only partial replacement. As the number of replaced layers increases, the two accuracy metrics of MIoU and MPA are continuously improved, reaching up to 90.49% and 94.47% respectively. In terms of the two metrics of the number of parameters and the amount of computation, they also decrease as the number of HWP replacement layers increases. The number of parameters decreases from 27.21MB to 25.13MB, and the amount of computation also decreases from 34.82 to 33.29, achieving a good balance between efficiency and accuracy.

[0069] 4.3 High - and low - frequency component analysis:

[0070] Based on the decomposition ability of the HWP module, this module divides the feature map into a low - frequency component and three high - frequency components. To verify the importance of the high - frequency and low - frequency components in the HWP module, the present invention combines and applies the high - frequency and low - frequency components, and the segmentation results are shown in Table 3.

[0071] Table 3 High - and low - frequency information analysis of the present invention

[0072] Model MIoU / % MPA / % PIDNet 90.14 93.85 PIDNet LL 90.08 94.09 PIDNet_LH - HL - HH 89.94 93.64 PIDNet_HWP 90.49 94.47

[0073] The results show that both low-frequency components and high-frequency components are crucial for improving the segmentation performance of the target. In the case of only retaining the low-frequency components, the MIoU is 90.08% and the MPA is 94.09%. Although the MIoU slightly decreases by 0.06% compared with the original PIDNet model, the MPA increases by 0.24%, indicating that low-frequency information plays an important role in improving the pixel classification accuracy. When only the high-frequency components are retained, the MIoU drops to 89.94% and the MPA drops to 93.64%. This shows that although the high-frequency components make a certain contribution to the segmentation edges and detailed features, the lack of support from low-frequency information will lead to a decline in the overall segmentation performance. When both high-frequency and low-frequency components are retained, the MIoU increases to 90.49% and the MPA increases to 94.47%. Compared with the original PIDNet model, the PIDNet_HWP model combined with the HWP module has an MIoU increase of 0.35% and an MPA increase of 0.62%. This indicates that the combination of high-frequency and low-frequency information can not only provide the detailed information and edge features of the image, but also provide the overall structure and general contour information of the image, which helps to improve the segmentation accuracy;

[0074] 5. Comparison of different algorithms:

[0075] The algorithm of this application is compared with SegNet, ERFNet, ESPNet, LEDNet, CGNet, BiSeNetv2, DDRNet, SeaFormer, and PIDNet in comparative experiments. The results are shown in Table 4. It can be found from it that in the infrared aircraft segmentation task, the real-time semantic segmentation algorithm WPTNet for infrared images based on the wavelet pooling three-branch network has an MIoU of 90.49% and an MPA of 94.47%, and the accuracy reaches the optimal; although it is not the best in terms of the number of parameters, computational complexity, and FPS, it still meets the requirements of real-time performance.

[0076] Table 4 Segmentation results of different algorithms of the present invention

[0077] Model MIoU / % MPA / % Parameters / MB GFLOP / s FPS frames / s SeggNet 75.00 79.39 28.08 91.79 24.99 ERFNet 79.14 84.99 27.54 81.16 30.51 ESPNet 81.25 87.24 24.24 75.96 32.42 LEDNet 82.02 87.46 20.46 71.92 33.81 CGNet 84.53 89.29 16.29 63.31 34.53 BiSeNetv2 86.54 91.36 13.36 59.66 36.84 DDRNet 87.09 91.81 10.81 46.83 37.92 SeaFormer 89.22 93.08 8.58 13.73 28.87 PIDNet 90.14 93.85 27.21 34.82 34.57 WPTNet 90.49 94.47 25.13 33.29 31.44

[0078] Visualization results of different real-time semantic segmentation algorithms are as Figure 8As shown. The first and second rows are the infrared aircraft nose detection results. It can be seen from the first row of data that only some algorithms can detect the nose, such as DDRNet, PIDNet, etc., but there are still some phenomena of missing information in these algorithms. The result of the algorithm of the present application for segmenting the nose is closer to the true label and can identify the nose structure more completely. It can be seen from the data in the second row that the algorithm of the present application can also clearly distinguish the nose from the propeller, effectively avoiding misclassification of details. For other algorithms such as Seaformer and BiseNetv2, although they are relatively close to the true label in the overall contour, when dealing with the boundary area between the propeller and the nose segmentation, there are often some misclassifications or confusions, resulting in less refined results.

[0079] The third and fourth rows are the infrared aircraft tail detection results. It can be seen from the third row of data that most algorithms are difficult to detect the infrared aircraft tail. The algorithm of the present application can accurately segment the complete contour of the tail, with clear edges and rich details, showing higher accuracy. It can be seen from the data in the fourth row that other algorithms such as DDRNet and BiSeNetv2 can detect the tail, but the structure is incomplete. The tail segmented by the algorithm of the present application is detailed and complete, with clear boundaries and closer to the true label.

[0080] The fifth and sixth rows are the infrared aircraft propeller detection results. It can be seen from these two rows of data that all algorithms can detect the propeller, but there are obvious breaks and detail losses in the segmentation results of ESPNet and CGNet, especially in the middle area of the propeller. Although the overall segmentation result of PIDNet is good, there are still a small number of blurred or broken edges. In contrast, the algorithm of the present application performs the most prominently. It can not only accurately segment the complete contour of the propeller, but also maintain the continuity and integrity of the propeller structure, avoiding breakage phenomena.

[0081] The above shows and describes the basic principles, main features and advantages of the present invention. Those skilled in the art of this industry should understand that the present invention is not limited by the above embodiments. The above embodiments and the descriptions in the specification are only preferred examples of the present invention and are not used to limit the present invention. Without departing from the spirit and scope of the present invention, the present invention will have various changes and improvements, and these changes and improvements all fall within the scope of the present invention claimed. The scope of protection claimed by the present invention is defined by the appended claims and their equivalents.

Claims

1. A real-time semantic segmentation method for infrared images based on a wavelet pooling three-branch network, characterized in that: The steps include: S1: Simulate and model the radiation, movement, and shape of interference, and create an infrared aircraft dataset under interference based on the public infrared target dataset; S2: Design the Haar Wavelet Pooling (HWP) module, which takes advantage of its frequency decomposition to reduce the information loss in the downsampling process when reducing the resolution of the feature map, and retain the significant part of the feature to the greatest extent; S3: Build a three-branch real-time semantic segmentation network model that extracts spatial, semantic, and boundary information, achieving the best balance between model inference speed and segmentation accuracy; S4: The Haar wavelet pooling module is introduced into the three-branch network as the downsampling layer of the network to form a real-time semantic segmentation network for infrared images based on the wavelet pooling three-branch network, which is more suitable for scenes with scarce detail information such as infrared images; S5: The infrared image real-time semantic segmentation network model based on wavelet pooling three-branch network is verified on the infrared aircraft dataset under self-made interference. The model can not only meet the real-time segmentation requirements of infrared images, but also has improved segmentation accuracy compared with other classic real-time semantic segmentation networks.

2. The infrared image real-time semantic segmentation method based on wavelet pooling three-branch network according to claim 1 is characterized in that: The infrared aircraft dataset under self-made interference specifically includes: 22043 aircraft images are screened out from the infrared public dataset LSOTB-TIR with high image resolution and aircraft targets having contours and certain shape features, and based on this, the radiation characteristics and operation form of the interference are combined, and interference is artificially added to produce the infrared aircraft dataset under interference. When imitating interference, the infrared radiation characteristics after the interference is released follow the Gaussian distribution, that is, the gray value in the center area is large and gradually weakens outward. And the interference may drop straight down or move around in the form of a parabola under the action of inertia and its own gravity, gradually moving away from the aircraft.

3. The infrared image real-time semantic segmentation method based on wavelet pooling three-branch network according to claim 1 is characterized in that: The Haar wavelet pooling module specifically includes: firstly, using a two-dimensional Haar wavelet transform to reduce the resolution of the feature map and retain the detail information. In this process, the image undergoes two stages of downsampling. In the first stage, a low-pass filter h is applied. low , high pass filter h high and a downsampling operation that decomposes the image into low-frequency and high-frequency components along the horizontal axis. In the second stage, the obtained low-frequency component and high-frequency component are further decomposed into four components by applying a low-pass filter and a high-pass filter in the vertical direction, respectively, which correspond to the low-frequency information of the image and the three high-frequency information in the horizontal, vertical and diagonal directions. Then the four components are stacked along the channel direction to form a new feature map, and the number of channels of the new feature map formed by stacking along the channel is adjusted by 1×1 convolution. Next, batch normalization and ReLU activation function are used to enhance the wavelet domain features.

4. The infrared image real-time semantic segmentation method based on wavelet pooling three-branch network according to claim 1 is characterized in that: The three-branch real-time semantic segmentation network model specifically includes: the network first enters the shared branch to perform preliminary feature extraction on the image, and then sends the downsampled 1 / 8 feature map to three branches to extract information from feature maps of different resolutions. The spatial branch is responsible for retaining detailed information in the high-resolution feature map, the semantic branch aggregates local and global context information to resolve long-distance dependencies, and the boundary branch extracts high-frequency features to predict boundary areas and guides the fusion of semantic and spatial information by integrating boundary information. Among them, the loss of the entire model includes semantic loss l0, l2, boundary loss l1, and boundary-aware cross entropy loss l3.

5. The infrared image real-time semantic segmentation method based on wavelet pooling three-branch network according to claim 1 is characterized in that: The method of introducing the Haar wavelet pooling module into the three-branch network model specifically includes: there are six downsampling layers in the three-branch network framework Stage-1 to Stage-5, and the size of the feature map obtained by each downsampling layer is 1 / 2, 1 / 4, 1 / 8, 1 / 16, 1 / 32, and 1 / 64 of the original image, respectively. The HWP module is used to downsample the feature map in Stage-1 to Stage-5, which can effectively reduce the dimension of the input image while retaining more high-frequency information, thereby retaining more detailed features.

Citation Information

Cited By

  • Cut tobacco production process key index prediction and rolling and packaging quality guarantee method based on AI

    CN120318219A