An infrared blind hole segmentation method for constructing a neural network by using atmospheric transmission and thermal inertia effect

By constructing a neural network that combines atmospheric transmission and thermal inertia effects, the problem of tactile paving segmentation under low light conditions was solved, achieving high-precision infrared tactile paving segmentation, which is suitable for tactile paving detection in complex environments.

CN116168044BActive Publication Date: 2025-12-05BEIHANG UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202310196115.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-03-03
Publication Date
2025-12-05
Estimated Expiration
2043-03-03

AI Technical Summary

Technical Problem

Existing visible light-based tactile paving segmentation technologies have limited performance under low light conditions and cannot effectively address the needs of tactile paving segmentation in complex environments such as nighttime or heavy fog.

Method used

An infrared blind path segmentation method is designed, which utilizes atmospheric transmission and thermal inertia effects to construct a neural network. By building a neural network and combining atmospheric transmission and thermal inertia effects in thermal infrared imaging, the network's ability to analyze the influence of atmospheric transmission is enhanced, and high-precision blind path segmentation is achieved through a feature fusion module.

Benefits of technology

High-precision tactile paving segmentation was achieved under low-light conditions, which can effectively handle tactile paving segmentation tasks in complex environments and improve segmentation accuracy and interpretability.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116168044B_ABST
    Figure CN116168044B_ABST
Patent Text Reader

Abstract

The application provides an infrared blind road segmentation method using atmospheric transmission and thermal inertia effect to construct a neural network, and the method comprises the following steps: step one: constructing a neural network based on a thermal infrared imaging process. First, the atmospheric transmission module is used to add atmospheric transmission information of the image, and then a backbone network is used to perform multi-scale extraction on target features; meanwhile, according to the thermal inertia effect of a thermal radiometer at an imaging end, a thermal inertia module is used to calculate the original signal value of the target; finally, the obtained multi-scale target features and the original signal value are fused, and the fused features are input into a decoder to obtain a probability prediction map at a pixel level of the whole image. Step two: constructing a loss function to train the network. The loss calculation is performed on the prediction result and a pixel-level label to realize the training of the network parameters. The trained neural network is used to process an infrared image. After the constructed neural network is fully iteratively trained using the training data, the trained network is used to segment a target image.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to an infrared blind path segmentation method that utilizes atmospheric transmission and thermal inertia effects to construct a neural network. It belongs to the fields of digital image processing and computer vision, and mainly involves deep learning and image segmentation technology. It has broad application prospects in various image-based application systems. Background Technology

[0002] According to WHO statistics, in 2020, an estimated 43.3 million people worldwide were blind, 295 million had moderate to severe visual impairment, 258 million had mild visual impairment, and 510 million had visual impairment due to uncorrected presbyopia. Globally, the prevalence of moderate to severe visual impairment among adults aged 50 and older increased slightly between 1990 and 2020. During this period, the number of blind people increased by 50.6%, and the number of people with moderate to severe visual impairment increased by 91.7%.

[0003] Tactile paving detection technology is a crucial component of mobility assistance systems for visually impaired individuals. Current tactile paving detection and segmentation techniques primarily rely on visible light images, which are ill-suited to poor lighting conditions such as darkness or heavy fog. Furthermore, uneven lighting and complex imaging conditions significantly increase the risk of failure for visible light tactile paving segmentation technologies. In contrast, thermal infrared imaging, with its unique active imaging mode, can overcome imaging difficulties under challenging lighting conditions and obtain stable images. Therefore, accurate segmentation of tactile paving from thermal infrared images is a challenging and significant research task.

[0004] Conventional image segmentation techniques primarily use models constructed from various features to build classifiers, with most features being based on color and texture. Alvarez et al. used a light-invariant feature space and road similarity to construct a classifier for road detection (see: Alvarez et al., Road Detection Based on Illuminant Invariance, IEEE Transactions on Intelligent Transportation Systems, 2011, 12(1): 184-193.). Horn et al. designed a semantic labelling method (see: Semantic Labelling to Aid Navigation in Prosthetic Vision, Horn et al., Engineering in Medicine & Biology Society, IEEE, 2015.) for path navigation. It divides each pixel in the image into a semantic category and then uses a classifier to highlight these categories to assist navigation. Tang Zhichao et al. proposed a method for identifying tactile paving and pedestrian crossings using Gaussian vectors and multi-color spaces (see reference: Research on Traffic Sign Visual Recognition Technology of Guiding Robot, Tang Zhichao et al., Computer Technology and Development, 2014.). However, it performs poorly in complex environments. Peng Yuqing et al. designed a tactile paving segmentation method based on color histograms and gray-level co-occurrence matrices, considering the color and texture information of tactile paving (see reference: Blind road recognition algorithm based on color and texture information, Peng Yuqing et al., Journal of Computer Applications, 2014.). It has better light shielding effect, but requires a longer computation time.This indicates that color-based segmentation methods are susceptible to changes in illumination. The viewing angle also affects texture information extraction and its performance. Cao et al. proposed a lightweight semantic segmentation network (see reference: Rapid Detection of Blind Roads and Crosswalks by Using a Lightweight Semantic Segmentation Network[J].IEEE Transactions on Intelligent Transportation Systems, PP(99):1-10.), which quickly and accurately segments blind roads and crosswalks. This method uses depthwise separable convolutions to improve the network's segmentation speed and dense dilated spatial pyramid pooling modules and contextual feature fusion modules to ensure segmentation accuracy.

[0005] In recent years, due to the development of deep learning, semantic segmentation has been widely used in the field of image segmentation. Unlike traditional algorithms, these algorithms do not require manual feature design. Therefore, many research focuses have shifted to semantic understanding of traffic scenes based on deep learning. Yang et al. constructed a deep learning network in the form of an encoder-decoder (see reference: Unifying Terrain Awareness for the Visually Impaired through Real-Time Semantic Segmentation, Yang Kailun et al., Sensors, Yang K, Wang K, Bergasa LM, et al. Unifying Terrain Awareness for the Visually Impaired through Real-Time Semantic Segmentation[J]. Sensors, 2018, 18(5): 1506.). It extracts features from RGB-D images captured by IR and RGB cameras to segment stairs, sidewalks, etc. In their subsequent work, Yang et al. optimized the network by proposing a compact panoramic environment lens system composed of a fisheye camera and other devices (see reference: Intersection Perception Through Real-Time Semantic Segmentation to Assist Navigation of Visually Impaired Pedestrians, Yang K, Cheng R, Bergasa LM, et al., 2018 IEEE International Conference on Robotics and Biomimetics (ROBIO). IEEE, 2018.). They then used deep learning methods for semantic segmentation to achieve panoramic scene analysis. Lomer et al. combined residual modules with decomposition convolution to construct a semantic segmentation network (see reference: Efficient Residual Decomposition Convolutional Neural Network for Real-Time Semantic Segmentation, Lomer et al., IEEE Transactions on Intelligent Transportation Systems, Romera E, Alvarez JM, Bergasa LM, et al. ERFNet: Efficient ResidualFactorized ConvNet for Real-Time Semantic Segmentation[J].IEEE Transactions on Intelligent Transportation Systems, 2017, PP(1):1-10.).Wei et al. designed a tactile paving segmentation method based on machine learning and marked watershed segmentation (see reference: Blind sidewalk image location based on machine learning recognition and marked watershed segmentation, Wei T, Zhou Y H. Blind sidewalk image location based on machine learning recognition and marked watershed segmentation[J]. Guangxue Jingmi Gongcheng / Optics and Precision Engineering, 2019, 27(1):201-210.), which can identify tactile paving of different colors. Li et al. segmented tactile paving by combining various salient features in a lightweight deep convolutional model (see reference: Blind road classification based on transfer learning and salient object detection, Li Lin et al. Blind road classification based on transfer learning and salient object detection[J]. Computer Engineering and Applications, 2018.). It can complete the segmentation quickly, but its recognition accuracy is relatively low.

[0006] Visible light-based tactile paving segmentation and detection algorithms cannot meet the requirements of nighttime thermal infrared segmentation, exhibiting limited performance in thermal infrared tactile paving segmentation tasks. To achieve effective thermal infrared tactile paving segmentation, this invention designs a deep learning network model based on atmospheric transmission and thermal inertia effects in infrared imaging, proposing a neural network-based infrared tactile paving segmentation method that utilizes atmospheric transmission and thermal inertia effects to construct a neural network. Summary of the Invention

[0007] 1. Objective: To address the difficulty of tactile paving segmentation under low light conditions, this invention proposes an infrared tactile paving segmentation method that utilizes atmospheric transmission and thermal inertia effects to construct a neural network. This method effectively combines atmospheric transmission and thermal inertia effects in thermal infrared imaging for network design, resulting in a significant improvement in segmentation accuracy.

[0008] 2. Technical Solution: To achieve the above objectives, the overall approach of this invention is to utilize the characteristics of thermal infrared radiation and imaging, and specifically target key influencing factors such as atmospheric transmission and thermal inertia, to construct a neural network for thermal infrared blind path segmentation. This effectively improves the ability to segment targets. The algorithmic approach of this invention is mainly reflected in the following three aspects:

[0009] 1) For the imaging process of this application, a network module based on atmospheric transmission is proposed to describe the atmospheric transmission process and enhance the network's ability to analyze the impact of atmospheric transmission.

[0010] 2) Based on the imaging characteristics of thermal infrared microradiometers, the thermal inertial effect in the imaging process is analyzed, and a thermal inertial module is proposed. The original signal value of the target is calculated through the thermal radiation effect, and a backbone network is added to provide information supplementation for the segmentation process, thereby enhancing the interpretability and information acquisition capability of the module.

[0011] 3) Design a feature fusion module to perform feature fusion based on the characteristics of atmospheric transmission and thermal inertia in the thermal imaging process, so as to achieve a higher accuracy segmentation effect.

[0012] This invention relates to an infrared blind path segmentation method that utilizes atmospheric transport and thermal inertia effects to construct a neural network. The specific steps of this method are as follows:

[0013] Step 1: Construct a neural network based on the thermal infrared imaging process. First, atmospheric transmission information is added to the image through an atmospheric transmission module. Then, multi-scale extraction of target features is performed through a backbone network. Simultaneously, based on the thermal inertia effect of the thermal radiometer at the imaging end, the original target signal value is calculated through a thermal inertia module. Finally, the obtained multi-scale target features are fused with the original signal value, and the fused features are sent to the decoder to obtain a probability prediction map at the pixel level of the entire image.

[0014] Step 2: Construct a loss function to train the neural network. Calculate the loss using the prediction results and pixel-level labels to train the neural network parameters.

[0015] Output: Process infrared images using a trained neural network. After sufficient iterative training of the constructed neural network using training data, a trained neural network is obtained for segmenting target images.

[0016] Specifically, step one is as follows:

[0017] 1.1: Atmospheric transport information is added to the input image using an atmospheric transport module. The input image I undergoes data augmentation and has a size of 3×512×512. The atmospheric transport module contains three trainable tensors: α... abs ,α sca d and represent the atmospheric absorption factor, atmospheric scattering factor, and target depth, respectively. The magnitudes of these three tensors are the same as those of the input image, such as... Figure 2 As shown, for α abs ,α sca The following operations are performed on the input image and d:

[0018]

[0019] I air For the image after adding information, i,j represent the coordinates of the corresponding pixel. This invention processes the input image by combining two main factors in atmospheric transmission: atmospheric attenuation and scattering. I air Input into the backbone network, such as Figure 1 As shown in the dashed box, after being embedded by overlapping blocks and operated by four multi-scale deformer modules, four layers of atmospheric sensing features F1, F2, F3, and F4 are obtained, with sizes of 128×128×64, 64×64×128, 32×32×320, and 16×16×512, respectively.

[0020] 1.2: Calculate the original information from the input image using the thermal inertia module. Pass the input image I to the thermal inertia module and initialize six thermal inertia factor tensors: C th G th , t0, t e α and Φ0, each with a size of 3×512×512. First, use C... th Point division G th The sensor's thermal effect factor τ is obtained, which is also 3×512×512. Then, as... Figure 3 As shown, the following calculations are performed:

[0021]

[0022] Wherein, Φ1 is the obtained original radiation intensity information, and its size is 3×512×512.

[0023] 1.3: The feature fusion module fuses the output features of the atmospheric transmission module and the thermal inertial module. The operation process is as follows: the output features F1, F2, F3, and F4 of the atmospheric transmission module and the output feature Φ1 of the thermal inertial module are input into the multi-layer fusion decoder, corresponding to the five decoders DE1, DE2, DE3, DE4, DE5, and DE6 respectively. phi .like Figure 4 As shown, these five decoders are all multilayer perceptron layers. First, the dimensions of F1, F2, F3, and F4 are transformed, unifying the third dimension to 768. Then, F1, F2, F3, and F4 are upsampled, resulting in a unified size of 128. Therefore, the sizes of F1, F2, F3, and F4 are now all 128×128×768. Φ1 is input into the corresponding decoder, and the decoder outputs a dimension of 768. Then, Φ1 is downsampled, reducing its size from 512 to 128, resulting in a size of 128×128×768. Finally, F1, F2, F3, and F4 are concatenated with Φ1 in the last dimension; the concatenated result is the fused total feature F. final Its size is 128×128×3840. The total feature F after fusion... finalThe input is fed into the final multilayer perceptron to obtain the two-dimensional prediction probability, which is the prediction probability value P output by the neural network, and its size is 512×512×2.

[0024] Step two is as follows:

[0025] 2.1: The neural network undergoes multiple iterations of training to ensure that the predicted output class is as close as possible to the correct class in the true label. The network parameters obtained when the iteration stops are the final result. The loss function measures the difference between the predicted value and the true label. The network reduces the loss function using gradient descent to make the predicted value closer to the true label. This invention uses cross-entropy as the loss function, and its calculation method is as follows:

[0026]

[0027] Where Y is the true label value, P is the predicted value, m and c represent the pixel index and label category, respectively. n represents the total number of pixels in the image. The optimal parameters of the neural network proposed in this invention are obtained by iteratively optimizing this loss function.

[0028] 2.2: During the iteration process, the neural network is optimized using gradient descent. The parameter settings for gradient descent in this invention are as follows: the AdamW optimizer is used for gradient optimization, the initial learning rate for neural network gradient descent is 0.0002, and the weight decay coefficient is 10. -4 During training, the learning rate is adaptively updated using a polynomial decay method, and the neural network parameters are adjusted through gradient backpropagation to reduce the corresponding loss function.

[0029] 3. Advantages and effects:

[0030] This invention proposes an infrared blind path segmentation method utilizing atmospheric transmission and thermal inertia effects to construct a neural network. For the imaging process in this application, a network module based on atmospheric transmission is proposed to describe the atmospheric transmission process, enhancing the network's analytical capability regarding the impact of atmospheric transmission. Furthermore, based on the imaging characteristics of thermal infrared microradiometers, the thermal inertia effect during the imaging process is analyzed, and a thermal inertia module is proposed to calculate the original target signal value through thermal radiation effects, enhancing the module's interpretability and information acquisition capability. A feature fusion module is designed to efficiently fuse the features extracted by the atmospheric transmission and thermal inertia modules. The model design, starting from the entire thermal imaging process, can achieve high-precision segmentation results, demonstrating good interpretability and performance, and has broad application prospects. Attached Figure Description

[0031] Figure 1 This is a schematic diagram of the principle of the infrared blind path segmentation method proposed in this invention, which utilizes atmospheric transmission and thermal inertia effects to construct a neural network.

[0032] Figure 2 This is a schematic diagram of the atmospheric transmission module in this invention.

[0033] Figure 3 This is a schematic diagram of the thermal inertia module in this invention.

[0034] Figure 4 This is the basic structure of the feature fusion module in this invention.

[0035] Figure 5 a-5h demonstrates the segmentation results of this invention in a real-world scenario; among them, Figure 5 a, 5b, 5e, and 5f are the original infrared images. Figure 5 c, 5d, 5g, and 5h are the segmentation results of the method of this invention.

[0036] Figure 6 This invention is compared with other cutting-edge methods. Detailed Implementation

[0037] To better understand the technical solution of the present invention, the embodiments of the present invention will be further described below with reference to the accompanying drawings.

[0038] This invention proposes an infrared blind path segmentation method that utilizes atmospheric transport and thermal inertia effects to construct a neural network. The principle block diagram is as follows: Figure 1 As shown, the specific implementation steps are as follows:

[0039] Step 1: Build a neural network based on the principle of thermal infrared imaging. The basic structure of the neural network is as follows: Figure 1 As shown;

[0040] Step 2: Construct a loss function to train the neural network.

[0041] Output: Process infrared images using a trained neural network. After iteratively training the constructed thermal infrared imaging characteristic neural network using training data, a trained network is obtained for segmenting target pixels.

[0042] Specifically, step one is as follows:

[0043] 1.1: Atmospheric transport information is added to the input image using the atmospheric transport module. The input image I has undergone data augmentation and is 3×512×512 in size. The atmospheric transport module contains three trainable tensors: α abs ,α sca d and represent the atmospheric absorption factor, atmospheric scattering factor, and target depth, respectively. All are initialized to 0 so that the initial value after the exponential operation is 1. The size of these three tensors is the same as the input image, all being 3×512×512, as shown below. Figure 2 As shown. Input I into the dot product layer, for α abs ,αsca The following operations are performed on the input image and d:

[0044]

[0045] I air For the image after adding atmospheric transmission information, i and j represent the coordinates of the corresponding pixels, and d shares the same set of parameters. This invention processes the input image by combining two main factors in atmospheric transmission: atmospheric attenuation and scattering phenomena, to obtain atmospheric transmission information. I air Input into the backbone network, such as Figure 1 As shown in the dashed box, after overlapping block embedding, with the number of overlapping blocks per row set to 4, I... air The data is expanded into 4×4=16 overlapping small blocks, and each small block is operated on using a deformer module based on a self-attention mechanism. Deformer modules of different scales downsample and augment the dimensions of the original data. The downsampled dimensions are 128, 64, 32, and 16, respectively, and the augmented dimensions are 64, 128, 320, and 512, respectively. By operating with the four multi-scale deformer modules, the input data I can be transformed... air The encoding consists of four multi-scale atmospheric sensing features F1, F2, F3, and F4, with sizes of 128×128×64, 64×64×128, 32×32×320, and 16×16×512, respectively.

[0046] 1.2: Calculate the original information of the input image using the thermal inertia module. Pass the input image I to the thermal inertia module and initialize six thermal inertia factor tensors: C th G th , t0, t e α, Φ0, and Φ0 represent heat capacity, thermal conductivity, initial exposure time, exposure duration, microradiometer absorptivity, and incoming radiation intensity, respectively. The tensors of these six thermal inertia factors are all 3 × 512 × 512. First, using C... th Point division G th The sensor's thermal effect factor τ is obtained, which is also 3×512×512. Then, as... Figure 3 As shown, C th G th , t0, t e Substitute α, Φ0, τ and I into the following formula for calculation:

[0047]

[0048] Φ1 represents the original radiation intensity information. The multiplication and exponentiation calculations are performed point-to-point for each corresponding pixel. The final output size of Φ1 is 3×512×512.

[0049] 1.3: The feature fusion module fuses the output features of the atmospheric transmission module and the thermal inertia module. The operation process is as follows: First, the output features F1, F2, F3, and F4 of the atmospheric transmission module and the output feature Φ1 of the thermal inertia module are input into the multilayer fusion decoder, corresponding to the five decoders DE1, DE2, DE3, DE4, DE5, and DE6 respectively. phi .like Figure 4 As shown, these five decoders are all multilayer perceptron layers. First, F1, F2, F3, and F4 are processed through multilayer perceptrons MLP1, MLP2, MLP3, and MLP4 at the corresponding scales. Dimension transformation is performed on the last dimension of the data to unify the last dimension of F1, F2, F3, and F4 after transformation. Here, the output dimension is uniformly set to 768, so the sizes of F1, F2, F3, and F4 after transformation become 128×128×768, 64×64×768, 32×32×768, and 16×16×768, respectively. Then, the transformed F1, F2, F3, and F4 are upsampled to maintain a uniform data size. The output size is set to 128×128, so the sizes of F1, F2, F3, and F4 after upsampling all become 128×128×768. Adjust the dimensional order of Φ1 to make its size 512×512×3, and then input it into the corresponding multilayer perceptron (MLP). phi The output dimension is 768, so the size of Φ1 is 512×512×3. Then, Φ1 is downsampled to make its size the same as F1, F2, F3, and F4. After downsampling, the size of Φ1 is 128×128×768. Finally, F1, F2, F3, F4, and Φ1 are merged in the last dimension. The feature dimension after merging is 128×128×3840, which is the total feature F after fusion. final The total feature F after fusion final The data is input into the final multilayer perceptron and upsampled to increase the data size from 128 to 512. The final multilayer perceptron input is 3840 dimensions, and the output dimension is the number of label categories. In this invention, there are two categories: blind path and background. Therefore, the output is a 2-dimensional prediction probability, which is the prediction probability value P output by the neural network, and its size is 512×512×2.

[0050] Step two is as follows:

[0051] 2.1: The neural network undergoes multiple iterations of training to make the predicted output class as close as possible to the correct class in the true label. The network parameters obtained when the iteration stops are the final result. The loss function is used to measure the difference between the predicted value and the true label. The network reduces the loss function using gradient descent to make the predicted value closer to the true label. This invention uses cross-entropy as the loss function to calculate the difference between the output probability and the true label. The calculation method is as follows:

[0052]

[0053] Where Y is the true label value, P is the predicted value, m and c represent the pixel index and label category, respectively. n represents the total number of pixels in the image. The optimal parameters of the neural network proposed in this invention are obtained by iteratively optimizing this loss function.

[0054] 2.2: During the iteration process, the neural network is optimized using gradient descent. The parameter settings for gradient descent in this invention are as follows: For the gradient optimizer, this invention uses the AdamW optimizer. The initial learning rate of the neural network is 0.0002, and the weight decay coefficient is 10. -4 During training, the learning rate is adaptively updated using a polynomial decay method, and the neural network parameters are adjusted through gradient backpropagation to reduce the corresponding loss function. In this process, gradient descent is used for backpropagation, and the chain rule is used to update the parameters by taking the partial derivative of the loss function with respect to the neural network parameters. Where θ i The parameters of the neural network before backpropagation, θ′ i Here, η represents the network parameters updated after backpropagation, η is the learning rate, and L is the loss function.

[0055] Figure 5 a-5h represents the application of this invention in a real-world infrared scenario. The white area represents the segmented tactile paving area. Figure 5 a, 5b, 5e, and 5f are the original infrared images. Figure 5 c, 5d, 5g, and 5h are the corresponding segmentation results. Figure 6 This invention is compared with other cutting-edge methods. The images used in the experiment were acquired between 8 PM and 10 PM from different cities and seasons. The tactile paving was relatively dark with significant depth variations, and there was interference from other road infrastructure in the background. However, the experimental results effectively eliminated noise interference and accurately segmented the target area. Accurate segmentation results were achieved across various difficulty levels, fully demonstrating the effectiveness of this invention. It can be widely applied to various infrared tactile paving segmentation systems and has broad market prospects and application value.

Claims

1. An infrared blind-hole segmentation method for constructing a neural network using atmospheric transmission and thermal inertia effects, characterized in that, The method comprises the following specific steps: Step one: build a neural network based on the thermal infrared imaging process; first, through the atmospheric transmission module, increase the image atmospheric transmission information, then through the backbone network to extract the target features in multiple scales; at the same time, according to the thermal inertia effect of the imaging end thermal radiometer, through the thermal inertia module, the original signal value of the target is calculated; finally, the multi-scale target features and the original signal value are fused, and the fused features are sent into the decoder to obtain the probability prediction map of the whole image pixel level; Step two: construct a loss function to train the neural network; use the prediction result and the pixel level label to calculate the loss, so as to realize the training of the neural network parameters; Output: process the infrared image with the trained neural network; after the constructed neural network is fully iteratively trained using the training data, the trained neural network is obtained for segmenting the target image.

2. The method of claim 1, wherein the method is an infrared blind-salve segmentation method using atmospheric transmission and thermal inertia effects. In step one, the following is specific: 1.1: Add atmospheric transmission information to the input image using an atmospheric transmission module; the input image I is data enhanced, with a size of 3x512x512, and the atmospheric transmission module contains three trainable tensors, which are: α abs , α sca and d, representing the atmospheric absorption factor, the atmospheric scattering factor and the target depth respectively; the sizes of the three tensors are consistent with the input image, and the following operations are performed on α abs , α sca and d and the input image: I air For the image after adding information, i, j represent the corresponding pixel point coordinates, and the input image is processed by combining the atmospheric attenuation and scattering phenomena in atmospheric transmission. I air In the input backbone network, it is sequentially operated through the operation of the overlap block embedding and the four multi-scale deformer modules to obtain four atmospheric perception features F1, F2, F3 and F4, and the sizes thereof are respectively: 128x128x64, 64x64x128, 32x32x320 and 16x16x512. 1.2: Compute raw information from input image using thermal inertia module; pass input image I to thermal inertia module, initialize 6 thermal inertia factor tensors: C th , G th , t0, t e , a, F0, all of size 3x512x512; first, divide G th by C th to get sensor thermal effect factor t, also of size 3x512x512; perform following calculations: Wherein, Φ1 is the obtained original radiation intensity information, and the size is 3*512*512; 1.3: The construction feature fusion module fuses the output features of the atmospheric transmission module and the thermal inertia module; the operation process is: the output features F1, F2, F3, F4 of the atmospheric transmission module and the output feature Φ1 of the thermal inertia module are input into a multi-layer fusion decoder, and correspond to five decoders DE1, DE2, DE3, DE4, DE phi respectively. The five decoders are all multi-layer perceptron layers, F1, F2, F3, F4 are first changed in dimension, and the third dimension is unified to 768, then F1, F2, F3, F4 are respectively up-sampled, and the size after up-sampling is unified to 128, so the size of F1, F2, F3, F4 is all 128x128x768; Φ1 is input into the corresponding decoder, the output dimension of the decoder is 768, then Φ1 is down-sampled, and the size is reduced from 512 to 128, so the size of Φ1 is also 128x128x768; F1, F2, F3, F4 and Φ1 are finally spliced in the last dimension, and the splicing result is the total fused feature F final , whose size is 128x128x3840; the total fused feature F final is input into the final multi-layer perceptron, and a two-dimensional prediction probability is obtained, that is, the prediction probability value P output by the neural network, whose size is 512x512x2.

3. The method of claim 1, wherein the method is an infrared blind-salve segmentation method using atmospheric transmission and thermal inertia effects. In step two, the following is specific: 2.1: the neural network is trained through multiple iterations to make the network prediction output category as close as possible to the correct category in the real label, and the network parameters obtained when the iteration stops are the final results, the loss function is used to measure the difference between the prediction value and the real label, and the network reduces the loss function through gradient descent method to make the prediction value close to the real label; cross entropy is used as the loss function, and the calculation method is as follows: Wherein Y is the real label value, P is the prediction value, m and c represent the pixel serial number and label category respectively; n represents the number of all pixels in the image; the optimal parameters of the neural network are obtained by iteratively optimizing the loss function; 2.2: In the iteration process, the neural network is optimized by gradient descent method, and the parameter settings in the gradient descent method are as follows: the AdamW optimizer is used for gradient optimization, the initial learning rate of neural network gradient descent is 0.0002, and the weight decay coefficient is 10 -4 , the polynomial attenuation method is used to adaptively update the learning rate during training, and the neural network parameters are adjusted by gradient back propagation to reduce the corresponding loss function.

4. The infrared blind-sound segmentation method of claim 3, wherein: In step 2.2, backpropagation is performed using the gradient descent method, and the parameter update is performed by taking the partial derivative of the loss function with respect to the neural network parameters by deriving the chain rule: where θ i is the neural network parameter before backpropagation, θ′ i is the updated network parameter after backpropagation, η is the learning rate, and L is the loss function.