Infrared Video Salient Object Detection Method Based on Deep Learning and Differentiable Clustering

By constructing an infrared video significant object detection model based on deep learning and differentiable clustering, the problem of poor results in infrared video significance detection is solved, and high-precision significant object detection is achieved.

CN116385752BActive Publication Date: 2025-06-03FUZHOU UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202310386196.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-04-12
Publication Date
2025-06-03
Estimated Expiration
2043-04-12

AI Technical Summary

Technical Problem

The prior art is not effective in infrared video significance detection, especially when dealing with background clutter and noise, it is difficult to accurately detect significant targets.

Method used

Using deep learning and differentiable clustering methods, an infrared video significant object detection model is constructed, including feature extraction based on Vgg16 network, significance detection model based on attention and ConvLSTM, and significant object segmentation model based on differentiable clustering, through these models, the detection accuracy is gradually improved.

Benefits of technology

It significantly improves the accuracy of infrared video significance detection, can clearly and accurately detect the salient areas in infrared video objects, and is suitable for industrial production.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116385752B_ABST
    Figure CN116385752B_ABST
Patent Text Reader

Abstract

The present invention relates to an infrared video salient object detection method based on deep learning and differentiable clustering, comprising: 1. acquiring infrared video image frames and constructing an infrared video saliency dataset; 2. constructing an infrared video salient object detection model, which mainly consists of a feature extraction network based on the Vgg16 network, a saliency detection model based on attention and ConvLSTM, and a salient object fine segmentation model based on differentiable clustering. The input image is subjected to feature extraction by the feature extraction network, and the extracted features are then input into the saliency detection model to obtain a dynamic saliency map. Finally, the salient object fine segmentation model is used for refined segmentation to obtain the image segmentation result. The infrared video salient object detection model is trained with the dataset; 3. inputting the image to be detected into the trained model to obtain the detection result. This method can improve the accuracy of infrared video saliency detection and clearly and accurately detect the salient regions in infrared video objects.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of image saliency detection, and particularly to an infrared video salient object detection method based on deep learning and differentiable clustering. Background Art

[0002] Currently, saliency detection is mainly carried out on visible light images. Such as frame difference method, background difference method and optical flow method, which effectively fuse the spatial saliency features of single-frame images and the motion features of videos, and can well highlight the salient objects in the videos. However, if these methods are directly applied to infrared videos, the effect will be much reduced. Compared with the saliency detection of visible light, infrared imaging has its special mechanism, and the attributes used for visible light videos will be invalid when used for infrared videos, and the infrared videos contain strong background clutter, which will cause the salient objects to be submerged by the background.

[0003] In recent years, there are many methods for infrared video object segmentation. Based on classical single-threshold segmentation, the image is divided into two parts: background and object, and then the object gray level and average value are selected to perform two-dimensional double-threshold segmentation on the image, which can effectively extract the object, but the anti-noise effect is poor. The interactive graph cut method based on geodesic distance, appearance overlap information and edge information to adaptively adjust each constraint term reduces noise interference and optimizes the segmentation effect, but it requires manual selection of foreground and background. To automatically distinguish the background from the foreground and segment the pixel level of the salient object, researchers introduced an image saliency detection method, which fuses spatial information into the traditional Gaussian mixture model using saliency mapping, and uses the saliency map as the spatial information of the Gaussian mixture model to effectively improve the object segmentation accuracy, but the computational cost of the convergence process is large and the iteration speed for large-scale data is slow, which is not suitable for the industrial production field. Summary of the Invention

[0004] The purpose of the present invention is to provide an infrared video salient object detection method based on deep learning and differentiable clustering, which can improve the accuracy of infrared video saliency detection and clearly and accurately detect the salient regions in infrared video objects.

[0005] To achieve the above purpose, the technical solution adopted by the present invention is: an infrared video salient object detection method based on deep learning and differentiable clustering, comprising the following steps:

[0006] Step 1, obtain infrared video image frames and construct an infrared video saliency data set;

[0007] Step 2: Construct an infrared video salient object detection model, which mainly consists of a feature extraction network based on the Vgg16 network, a saliency detection model based on attention and ConvLSTM, and a fine segmentation model of salient objects based on differentiable clustering. The infrared video salient object detection model extracts features from the input image through the feature extraction network to obtain a series of convolutional features, and then inputs the obtained series of convolutional features into the saliency detection model to obtain the hidden state H t , and then the hidden state H t is convolved with a 1×1 convolutional kernel to obtain a dynamic saliency map. Finally, the obtained dynamic saliency map is refined and segmented by the fine segmentation model of salient objects to obtain the image segmentation result; the infrared video salient object detection model is trained through the infrared video saliency dataset to obtain a trained infrared video salient object detection model;

[0008] Step 3: Input the image to be detected into the trained infrared video salient object detection model to obtain the final detection result.

[0009] Furthermore, the feature extraction network based on the Vgg16 network consists of 13 convolutional layers and 3 fully connected layers. The convolutional layers are divided into 5 convolutional blocks, and behind each convolutional block is a max pooling layer with a stride of 2;

[0010] The feature extraction network introduces dilated convolution, and the calculation method of the dilated convolution is as follows:

[0011] y(i) = ∑ k x(i + rk)w(k) (1)

[0012] where x represents the input feature map, y represents the output feature map, w represents the filter, and r represents the stride of sampling the input signal in the dilated convolution;

[0013] Assume that the convolutional kernel size of the dilated convolution is k and the dilation rate is d, then its equivalent convolutional kernel size k', and its formula is as follows:

[0014] k' = k + (k - 1)×(d - 1) (2)

[0015] The calculation formula of the receptive field of the current layer is as follows:

[0016] RF i+1 = RF i + (k' - 1)×(d - 1) (3)

[0017] where RF i+1 represents the receptive field of the current layer, RF i represents the receptive field of the previous layer, and k' represents the size of the convolutional kernel;

[0018] S i represents the product of the strides of all previous layers, and its formula is as follows:

[0019]

[0020] Furthermore, the size of the convolution kernel of the Vgg16 convolutional layer is set to 5, the dilation rate is set to 3, and the stride of all layers is set to 1, thus obtaining a receptive field of 17×17.

[0021] Furthermore, the saliency detection model based on attention and ConvLSTM includes an attention module and a ConvLSTM module;

[0022] The attention module is built on the conv5-3 layer of VGG-16 and is interleaved with pooling layers and upsampling operations; given the input feature X, after passing through the pooling layer, the attention module generates a downsampled attention map with an enlarged receptive field; then the downsampled attention map is multiplied by 4, and let M∈[0,1] 28×28 be the upsampled attention map, and the feature X∈R 28×28×512 is further enhanced in the following way:

[0023]

[0024] where c∈{1,......,512} is the index of the channels, and the attention module is used as a feature selector to enhance features; the above formula is improved in the form of residuals:

[0025]

[0026] The original convolutional neural network features and the enhanced features are combined using residual connections and input into the ConvLSTM module. ConvLSTM is used to model the temporal dynamics of sequential problems, which is achieved by combining memory cells with gating operations. In addition, the ConvLSTM retains spatial information by replacing the dot product with a convolutional operation; the ConvLSTM module uses three convolutional gates including input, output, and forget gates to control the signal flow within the unit; the ConvLSTM outputs a hidden state H t , and uses a memory cell C t to control the update and output of the state:

[0027]

[0028] where i t , f t , o t are the input gate, forget gate, and output gate in the LSTM model respectively, and σ=1 / [1+exp(-x)] is i t, f t , o t The activation function; the hyperbolic tangent function tanh(x) = [exp(x) - exp(-x)] / [exp(x) + exp(-x)] is c t The activation function; * represents the convolution operator, and ο represents the Hadamard product;

[0029] Convolve the hidden state H t with a 1×1 convolution kernel to obtain a dynamic saliency map.

[0030] Further, the salient object fine segmentation model is implemented based on a differentiable clustering algorithm, and the differentiable clustering algorithm performs differential clustering on a single image in a convolutional neural network to achieve fine segmentation of the saliency image;

[0031] The differentiable clustering algorithm first uses a CNN for feature extraction to obtain infrared image features {x n }, then uses a q-dimensional convolutional layer to calculate a q-dimensional feature map {r n }, and then performs normalization using BatchNorm to obtain a normalized feature map {r' n }; then, the ID of the maximum value in the feature vector corresponding to each pixel is obtained through the argmax function as the pixel category, and thus the final clustering label {C n } is obtained; finally, the direction propagation method of gradient descent is used to update the parameters to find a label assignment solution that minimizes the loss function L; the loss function L is expressed as follows:

[0032]

[0033]

[0034] where N represents the set of pixel points of the input image; r' n,i represents the eigenvalue of the i-th element in the mapping {r' n }; r' ξ,η represents the eigenvalue at (ξ, η) in the mapping {r' n }; μ represents the weight for balancing these two constraints;

[0035] For the case where the rough segmentation map contains the background, the region corresponding to the pixel value with the smallest pixel ratio in the object is defined as the background to obtain an optimized image saliency map S, and the mapping relationship is as follows:

[0036]

[0037] where D(x, y) represents the pixel value at (x, y) in the rough segmentation map; I minRepresents the pixel value with the least proportion in the statistical grayscale histogram of the region corresponding to the target region in the rough segmentation map in the differentiable clustering map; I(x, y) represents the pixel value at (x, y) in the differentiable clustering map.

[0038] For the case where the target region in the rough segmentation map is too small, the region of interest of the differentiable clustering map is selected using the largest connected region of the target binary map as a mask to obtain a new possible target region. The pixel value with the highest proportion in the grayscale histogram of this region, excluding the 0 pixel value, is used as the seed point I. max To obtain the optimized saliency map S, the mapping relationship is as follows:

[0039]

[0040] Compared with the prior art, the present invention has the following beneficial effects: It improves an infrared video salient target detection method based on deep learning and differentiable clustering. This method adopts the idea from "coarse" to "fine", divides the salient target detection problem into two steps: target localization and segmentation refinement. It detects target saliency through a saliency detection model based on attention and ConvLSTM, and then refines the segmentation of the saliency map through a salient target fine segmentation model based on differentiable clustering, thereby effectively improving the accuracy of infrared video saliency detection and clearly and accurately detecting the salient regions in infrared video objects. BRIEF DESCRIPTION OF THE DRAWINGS

[0041] Figure 1 is the flowchart of the method implementation in the embodiment of the present invention;

[0042] Figure 2 is the comparison diagram of ordinary convolution and dilated convolution in the embodiment of the present invention;

[0043] Figure 3 is the structural diagram of the saliency detection model based on attention and ConvLSTM in the embodiment of the present invention;

[0044] Figure 4 is the structural diagram of the differentiable clustering algorithm in the embodiment of the present invention;

[0045] Figure 5 is the optimized effect diagram of the infrared foam image in the embodiment of the present invention;

[0046] Figure 6 is the optimized effect diagram of the SegTrack-v2 dataset in the embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0047] The present invention will be further described below with reference to the accompanying drawings and embodiments.

[0048] It should be noted that the following detailed description is exemplary and is intended to provide further explanation of the present application. Unless otherwise specified, all technical and scientific terms used herein have the same meaning as commonly understood by those of ordinary skill in the technical field to which the present application belongs.

[0049] It should be noted that the terms used herein are only for describing specific embodiments and are not intended to limit the exemplary embodiments according to the present application. As used herein, unless the context clearly indicates otherwise, the singular forms are also intended to include the plural forms. In addition, it should be understood that when the terms "comprising" and / or "including" are used in this specification, they indicate the presence of features, steps, operations, devices, components, and / or combinations thereof.

[0050] As Figure 1 shown, this embodiment provides an infrared video salient object detection method based on deep learning and differentiable clustering, including the following steps:

[0051] Step 1: Obtain infrared video image frames and construct an infrared video saliency dataset.

[0052] Step 2: Construct an infrared video salient object detection model. The infrared video salient object detection model mainly consists of a feature extraction network based on the Vgg16 network, a saliency detection model based on attention and ConvLSTM, and a salient object fine segmentation model based on differentiable clustering. The infrared video salient object detection model extracts features from the input image through the feature extraction network to obtain a series of convolutional features, and then inputs the obtained series of convolutional features into the saliency detection model to obtain the hidden state H t , and then the hidden state H t is convolved with a 1×1 convolutional kernel to obtain a dynamic saliency map. Finally, the obtained dynamic saliency map is refined and segmented by the salient object fine segmentation model to obtain an image segmentation result; the infrared video salient object detection model is trained through the infrared video saliency dataset to obtain a trained infrared video salient object detection model.

[0053] Step 3: Input the image to be detected into the trained infrared video salient object detection model to obtain the final detection result.

[0054] The relevant content involved in the above method steps will be further described below.

[0055] 1. Improvement of the feature extraction network based on the Vgg16 network

[0056] The feature extraction network based on the Vgg16 network consists of 13 convolutional layers and 3 fully connected layers. The convolutional layers are divided into 5 convolutional blocks, and each convolutional block is followed by a max pooling layer with a stride of 2. Since the purpose of Vgg16 is to extract feature maps and deep convolutional layers have been proven to be an effective method for extracting image feature representations, we only need to consider the convolutional layers. However, pooling layers may cause loss of image semantic information. To overcome this limitation, the present invention introduces dilated convolution into the feature extraction network.

[0057] Different from ordinary convolution, dilated convolution introduces a hyperparameter called 'dilation rate', which defines the spacing of each value when the convolutional kernel processes data. To improve the computational efficiency of wavelet transform, the calculation method of dilated convolution is as follows:

[0058] y(i)=∑ k x(i+rk)w(k) (1)

[0059] where x represents the input feature map, y represents the output feature map, w represents the filter, and r represents the stride of sampling the input signal in dilated convolution. Generally, standard convolution is a special case with rate r = 1. Taking a 3×3 convolution as an example, Figure 2 shows the difference between ordinary convolution and dilated convolution.

[0060] Figure 2 From left to right are subgraphs a, b, and c. The three graphs are convolved independently. The large box represents the input image (the receptive field is defaulted to 1), the black dots represent the 3×3 convolutional kernel, and the gray area represents the receptive field after convolution. It can be seen that a is an ordinary convolution process (dilation rate = 1), and the receptive field after convolution is 3. b is a dilated convolution with dilation rate = 2, and the receptive field after convolution is 5. c is a dilated convolution with dilation rate = 3, and the receptive field after convolution is 8. From Figure 2 it can be seen that the same 3×3 convolution can achieve the effects of convolutions such as 5×5 and 7×7. Dilated convolution can increase the receptive field without increasing the number of parameters (the number of parameters = convolutional kernel size + bias). Assuming the convolutional kernel size of dilated convolution is k and the dilation number is d, then its equivalent convolutional kernel size k'. For example, for a 3×3 convolutional kernel, then k = 3, and its formula is as follows:

[0061] k'=k+(k - 1)×(d - 1) (2)

[0062] The calculation formula for the receptive field of the current layer is as follows:

[0063] RF i+1 =RF i +(k' - 1)×(d - 1) (3)

[0064] Among them, RF i+1 represents the receptive field of the current layer, and RF i represents the receptive field of the previous layer, and k' represents the size of the convolutional kernel. S i represents the product of the strides of all previous layers, and its formula is as follows:

[0065]

[0066] In this embodiment, the size of the convolutional kernel of the Vgg16 convolutional layer is set to 5, the dilation rate is set to 3, and the stride of all layers is set to 1. According to the above formula, a receptive field of 17×17 can be obtained. For the Vgg16 model, to retain more spatial details, pool5 is removed.

[0067] 2. Salience Detection Model Based on Attention and ConvLSTM

[0068] The salience detection model based on attention and ConvLSTM includes an attention module and a ConvLSTM module. This model is a hybrid prediction model that contains a time series analysis module and an attention mechanism. Attention is learned in an automatic, top-down, task-specific manner, which enables the network to focus on the most relevant parts of the image. In the present invention, attention is used to enhance the intra-frame salience features, making it easier for the LSTM to perform dynamic modeling. The difference from the previous attention module is that the attention module in this model can encode static salience information and learn from the static salience dataset in a supervised manner, and a supervised attention mechanism is added to the network structure to achieve advanced results in dynamic fixation prediction, improving the generality and prediction performance of the network model.

[0069] The attention module is built on the conv5-3 layer of VGG-16 and is interleaved with pooling layers and upsampling operations. As Figure 3 shown, given the input feature X, after passing through the pooling layer, the attention module generates a downsampled attention map (7×7) with an enlarged receptive field (260×260). Then the downsampled attention map is multiplied by 4, and let M∈[0,1] 28×28 be the upsampled attention map, and the feature X∈R 28×28×512 is further enhanced in the following way:

[0070]

[0071] Among them, c ∈ {1,......, 512} is the index of the channel. Among them, the attention module is used as a feature selector to enhance features. However, the above attention module may lose some useful information for learning dynamic saliency because the attention module only considers the static saliency information in static video frames. Therefore, the above formula is improved in the form of a residual:

[0072]

[0073] The original convolutional neural network features and the enhanced features are combined using residual connections and input into the ConvLSTM module. In the present invention, ConvLSTM is used to model the temporal dynamics of sequential problems, which is achieved by combining memory cells with gating operations. In addition, the ConvLSTM retains spatial information by replacing the dot product with a convolutional operation; the ConvLSTM module uses three convolutional gates (input, output, and forget) to control the signal flow within the unit. The ConvLSTM outputs a hidden state H t , and uses a memory cell C t to control the update and output of the state:

[0074]

[0075] where, i t , f t , o t are the input gate, forget gate, and output gate in the LSTM model respectively. σ = 1 / [1 + exp(-x)] is the activation function of i t , f t , o t ; the hyperbolic tangent function tanh(x) = [exp(x) - exp(-x)] / [exp(x) + exp(-x)] is the activation function of c t ; * represents the convolution operator, and ο represents the Hadamard product.

[0076] The hidden state H t is convolved with a 1×1 convolutional kernel to obtain a dynamic saliency map.

[0077] 3. Fine-grained segmentation model of salient objects based on differentiable clustering

[0078] To obtain better image segmentation effects, the present invention introduces a differentiable clustering algorithm, and the fine-grained segmentation model of salient objects is implemented based on the differentiable clustering algorithm. This algorithm performs differential clustering on a single image in a convolutional neural network to achieve fine-grained segmentation of the salient image. The structure of the differentiable clustering algorithm is as Figure 4 shown.

[0079] The differentiable clustering algorithm first uses a CNN for feature extraction to obtain infrared image features {x n}, then uses a q-dimensional convolutional layer to calculate a q-dimensional feature map {r n}, and then uses BatchNorm for normalization to obtain a normalized feature map {r' n}; then, the ID of the maximum value in the feature vector corresponding to each pixel is obtained through the argmax function as the pixel category, and the final clustering label {C n} is obtained in this way; finally, the direction propagation method of gradient descent is used to update the parameters to find a label assignment solution that minimizes the loss function L; the loss function L is expressed as follows:

[0080]

[0081]

[0082] Among them, N represents the set of pixel points of the input image; r' n,i represents the eigenvalue of the i-th element in the mapping {r' n}; r' ξ,η represents the eigenvalue at (ξ, η) in the {r' n} mapping; μ represents the weight balancing these two constraints.

[0083] For the case where the rough segmentation map contains the background, the region corresponding to the pixel value with the smallest pixel ratio in the target is defined as the background to obtain an optimized image saliency map S, and the mapping relationship is as follows:

[0084]

[0085] Among them, D(x, y) represents the pixel value at (x, y) in the rough segmentation map; I min represents the pixel value with the smallest ratio in the statistical gray histogram of the region of the differentiable clustering map corresponding to the target region in the rough segmentation map; I(x, y) represents the pixel value at (x, y) in the differentiable clustering map.

[0086] For the case where the target region in the rough segmentation map is too small, the maximum connected region of the target binary map is used as a mask to select the region of interest in the differentiable clustering map to obtain a new possible target region, and the pixel value with the highest ratio except the 0 pixel value in the gray histogram of this region is used as the seed point I max to obtain an optimized saliency map S, and the mapping relationship is as follows:

[0087]

[0088] In this embodiment, experiments are carried out using the infrared foam video saliency dataset collected by a thermal infrared imager and the publicly available RGB-T infrared video saliency dataset. As Figure 5 , 6 shown, the images 5(a) and 5(a) of the dataset are input into the trained network model, thereby obtaining the preliminary saliency Figure 5 (b) and 6(b). The differentiable clustering Figure 5 (c) and 6(c) is obtained through the differentiable clustering algorithm. For the missing detail parts in the preliminary saliency map, the differentiable clustering map is used for differentiable clustering optimization, and the optimization effect is as Figure 5 (d) and 6(d) shown.

[0089] To better illustrate the performance of the method of the present invention, a quantitative comparison of the images segmented by various methods is shown in Table 1. In Table 1, the "coarse saliency map" represents the coarse saliency map obtained by the saliency detection model, and the "refined segmentation map" represents the saliency map after the differentiable clustering optimization of the coarse saliency map. It can be seen from Table 1 that compared with the coarse saliency map, the IOU index and pixel accuracy (PA) of the refined segmentation map are significantly improved, and the segmentation error rate (Error) is significantly reduced.

[0090] Table 1 Comparison of Coarse Saliency Map and Refined Segmentation Map

[0091]

[0092] Although the traditional saliency detection model has a simple model and low time cost, the segmentation accuracy is generally not high. The deep learning method theoretically has extremely high accuracy under the training of a sufficient amount of samples, but it is relatively difficult to collect and label data, and it is difficult to form a large dataset. The present invention proposes an infrared video object detection method combining deep learning and differentiable clustering. This method adopts the idea from "coarse" to "fine", divides the saliency object detection problem into two steps: object localization and segmentation refinement. Multiple comparison experiments on the dataset show that this method can effectively improve the accuracy of saliency detection and clearly and accurately detect the saliency region in the infrared video object.

[0093] Those skilled in the art should understand that the embodiments of the present application can be provided as a method, a system, or a computer program product. Therefore, the present application can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present application can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk memory, CD-ROM, optical memory, etc.) containing computer-usable program code.

[0094] This application is described with reference to the flowcharts and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the present application. It should be understood that each flow and / or block in the flowchart and / or block diagram, and the combination of flows and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to generate a machine, such that the instructions executed by the processor of the computer or other programmable data processing device generate means for implementing the functions specified in one or more flows and / or blocks Figure 1 in one or more flows and / or blocks Figure 1 or in one or more blocks.

[0095] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, such that the instructions stored in the computer-readable memory generate a manufactured article including instruction means that implement the functions specified in one or more flows and / or blocks Figure 1 in one or more flows and / or blocks Figure 1 or in one or more blocks.

[0096] These computer program instructions can also be loaded onto a computer or other programmable data processing device, such that a series of operation steps are executed on the computer or other programmable device to generate a computer-implemented process, so that the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in one or more flows and / or blocks Figure 1 in one or more flows and / or blocks Figure 1 or in one or more blocks.

[0097] As described above, it is only the preferred embodiments of the present invention, and it is not a limitation to the present invention in other forms. Any person skilled in the art may use the disclosed technical content to make changes or modifications into equivalent embodiments with equivalent changes. However, any simple modification, equivalent change, and modification made to the above embodiments based on the technical essence of the present invention without departing from the technical solution content of the present invention still fall within the protection scope of the technical solution of the present invention.

Claims

1. An infrared video salient object detection method based on deep learning and differentiable clustering, characterized in that, it includes the following steps: Step 1: Obtain infrared video image frames and construct an infrared video saliency dataset; Step 2: Construct an infrared video salient object detection model, which mainly consists of a feature extraction network based on the Vgg16 network, a saliency detection model based on attention and ConvLSTM, and a salient object fine segmentation model based on differentiable clustering. The infrared video salient object detection model extracts features from the input image through the feature extraction network to obtain a series of convolutional features, and then inputs the obtained series of convolutional features into the saliency detection model to obtain the hidden state H t , and then the hidden state H t is convolved with a 1×1 convolutional kernel to obtain a dynamic saliency map. Finally, the obtained dynamic saliency map is refined and segmented by the salient object fine segmentation model to obtain the image segmentation result; the infrared video salient object detection model is trained by the infrared video saliency dataset to obtain a trained infrared video salient object detection model; Step 3: Input the image to be detected into the trained infrared video salient object detection model to obtain the final detection result; The salient object fine segmentation model is implemented based on the differentiable clustering algorithm, and the differentiable clustering algorithm performs differential clustering on a single image in the convolutional neural network to achieve fine segmentation of the saliency image; The differentiable clustering algorithm first uses a CNN for feature extraction to obtain infrared image features {x n}, then uses a q-dimensional convolutional layer to calculate a q-dimensional feature map {r n}, and then uses BatchNorm for normalization to obtain a normalized feature map {r' n}; then, the ID of the maximum value in the feature vector corresponding to each pixel is obtained through the argmax function as the pixel category, and the final clustering label {C n} is obtained in this way; finally, the direction propagation method of gradient descent is used to update the parameters to find a label assignment solution that minimizes the loss function L; the loss function L is expressed as follows: Among them, N represents the set of pixel points of the input image; r' n,i represents the eigenvalue of the i-th element in the mapping {r' n}; r' ξ,η represents the eigenvalue at (ξ, η) in the mapping {r' n}; μ represents the weight for balancing these two constraints; For the case where the rough segmentation map contains the background, the region corresponding to the pixel value with the smallest pixel ratio in the target is defined as the background to obtain the optimized image saliency map S, and the mapping relationship formula is as follows: Among them, D(x, y) represents the pixel value at (x, y) in the rough segmentation map; I min represents the pixel value with the least proportion in the statistical gray histogram of the region corresponding to the target region in the rough segmentation map in the differentiable clustering map; I(x, y) represents the pixel value at (x, y) in the differentiable clustering map; For the case where the target region of the rough segmentation map is too small, the maximum connected region of the target binary map is used as a mask to select the region of interest from the differentiable clustering map, and a new possible target region is obtained. The pixel value with the highest proportion in the gray histogram of this region, excluding the pixel value of 0, is used as the seed point I max to obtain the optimized saliency map S, and the mapping relationship is as follows:

2. The infrared video salient object detection method based on deep learning and differentiable clustering according to claim 1, characterized in that, The feature extraction network based on the Vgg16 network consists of 13 convolutional layers and 3 fully connected layers. The convolutional layers are divided into 5 convolutional blocks, and each convolutional block is followed by a max pooling layer with a stride of 2; The feature extraction network introduces dilated convolution, and the calculation method of the dilated convolution is as follows: y(i) = ∑ k x(i + rk)w(k) (1) where x represents the input feature map, y represents the output feature map, w represents the filter, and r represents the stride of sampling the input signal in the dilated convolution; Assume that the convolution kernel size of the dilated convolution is k and the dilation rate is d, then its equivalent convolution kernel size k' is given by the formula: k' = k+(k - 1)×(d - 1) (2) The formula for calculating the receptive field of the current layer is as follows: RF i+1 = RF i + (k' - 1) × (d - 1) (3) Among them, RF i+1 represents the receptive field of the current layer, and RF i represents the receptive field of the previous layer, where k' represents the size of the convolutional kernel; S i represents the product of the strides of all previous layers, and its formula is as follows:

3. The infrared video salient object detection method based on deep learning and differentiable clustering according to claim 2, characterized in that, The convolution kernel size of the Vgg16 convolutional layer is set to 5, the dilation size is set to 3, and the stride of all layers is set to 1, so as to obtain a receptive field of 17×17.

4. The infrared video salient object detection method based on deep learning and differentiable clustering according to claim 1, characterized in that, The saliency detection model based on attention and ConvLSTM includes an attention module and a ConvLSTM module; The attention module is built on the conv5-3 layer of VGG-16 and interleaved with pooling layers and upsampling operations; given the input feature X, after passing through the pooling layer, the attention module generates a downsampled attention map with an enlarged receptive field; then the downsampled attention map is multiplied by 4, and let M ∈ [0, 1] 28×28 be the upsampled attention map, and the feature X ∈ R 28×28×512 is further enhanced in the following way: where c ∈ {1,......,512} is the index of the channel. Among them, the attention module is used as a feature selector to enhance features; the formula (5) is improved in the form of a residual: The original convolutional neural network features and the enhanced features are combined using residual connections and input into the ConvLSTM module. The ConvLSTM is used to model the temporal dynamics of sequential problems, which is achieved by incorporating memory cells with gating operations. In addition, the ConvLSTM retains spatial information by replacing the dot product with convolutional operations. The ConvLSTM module uses three convolutional gates, including input, output, and forget gates, to control the signal flow within the unit. The ConvLSTM outputs a hidden state H t , and uses a memory cell C t to control the update and output of the state: where i t , f t , o t are the input gate, forget gate, and output gate in the LSTM model respectively, and σ = 1 / [1 + exp(-x)] is the activation function of i t , f t , o t ; the hyperbolic tangent function tanh(x) = [exp(x) - exp(-x)] / [exp(x) + exp(-x)] is the activation function of c t ; * represents the convolution operator, represents the Hadamard product; The hidden state H t is convolved with a 1×1 convolutional kernel to obtain a dynamic saliency map.