Infrared scene classification method based on flexible deformable convolution and attention mechanism

Through the AFDC-IRNet model with flexible deformable convolution and attention mechanism, the problem of insufficient mining of global features and semantic features in infrared image scene classification is solved, and the accuracy and effect of infrared image scene classification is improved.

CN120564014APending Publication Date: 2025-08-29西安应用光学研究所
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510664083.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-22
Publication Date
2025-08-29

AI Technical Summary

Technical Problem

The existing infrared image scene classification methods lack mining of the global features and semantic context features of the image, resulting in poor classification results.

Method used

The AFDC-IRNet model based on flexible deformable convolution and attention mechanism is adopted to improve the receptive field through flexible deformable convolution, expand the global feature perception ability, and pay attention to the wide-area feature information of infrared images through attention mechanism, and infrared scene classification is carried out by combining activation functions, pooling operations and full connection layer.

Benefits of technology

The task performance of infrared scene classification is improved, the ability to mine global features and semantic context features of infrared images is enhanced, and the accuracy and effect of classification is improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120564014A_ABST
    Figure CN120564014A_ABST
Patent Text Reader

Abstract

The invention provides an infrared scene classification method based on flexible deformable convolution and an attention mechanism, and the method comprises the steps: inputting an infrared image into an AFDC-IRNet model, and outputting the classification information of an infrared scene to which the infrared image belongs through the model; wherein the AFDC-IRNet model processing process comprises the following steps: an infrared image passes through an FDCN1, an ReLU1 and an MAXPool1 in sequence, and a first feature map is output; the first feature map outputs a first network feature image through DPANet1; the first network feature image passes through an FDCN2, a ReLU2, a MAXPool2, a DPANet2, an FDCN3, a ReLU3 and a MAXPool3 in sequence, and a third feature map is output; and the third feature map enters Dense after passing through a Dropout layer, and infrared scene classification information to which the infrared image belongs is output. According to the method, the receptive field of the AFDC-IRNet model is improved through flexible variability convolution, the global feature perception capability of the AFDC-IRNet model is expanded, the AFDC-IRNet model pays attention to the wide-area feature information of the infrared image through an attention mechanism, the image global perception and image semantic context feature mining capability of the AFDC-IRNet model is improved, and the image semantic context feature mining capability of the AFDC-IRNet model is improved. And the task performance of infrared scene classification is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of infrared image classification, and in particular relates to an infrared scene classification method based on flexible deformable convolution and attention mechanism. Background Art

[0002] In the field of computer vision, scene classification is the process of determining the scene category of an image based on its features. Existing scene classification methods, including convolutional neural networks and attention mechanisms, achieve scene classification by extracting multi-dimensional image features. However, single deformable convolutional networks or attention mechanisms are unable to extract global information about the image scene, resulting in poor understanding of the image's context, which in turn affects the effectiveness of scene classification.

[0003] Correctly assigning scene classifications to infrared images can provide valuable reference information for tasks such as target recognition and image semantic segmentation. For example, in target recognition, infrared images depicting the sky are more likely to contain flying objects, and the interference and impact of weather and cloud cover on the target must be carefully considered. In image semantic segmentation, infrared images depicting urban scenes prioritize the accurate segmentation of intersections between straight roads and buildings. Therefore, accurate infrared image scene classification is crucial for subsequent segmentation of image information.

[0004] Since the detailed texture information of infrared images is weaker than that of visible light images, the fine-grained feature information of objects in infrared scenes is not rich enough. Therefore, it is particularly important to use the global feature information and image context association of infrared images to classify the overall scene of the image.

[0005] With the rapid development of artificial intelligence and computer vision technologies, and the exponential expansion of GPU computing power, it has become possible to train complex deep learning neural network models using large-scale infrared image sample data. However, existing scene classification methods mostly rely on traditional convolutional neural networks, which focus more on image features within the local receptive field and lack the mining of global image features and semantic context, thus affecting the classification effect of infrared image scenes.

[0006] It can be seen that the existing image scene classification methods are not sufficient for extracting infrared image features and understanding semantic information, especially the mining and extraction of multi-scale information and global features of infrared images are not thorough enough. Summary of the Invention

[0007] The purpose of the present invention is to solve the shortcomings of the existing technology in that the scene category cannot be accurately judged in complex infrared scenes, and to provide an infrared scene classification method based on flexible deformable convolution and attention mechanism. The flexible deformable convolution is used to improve the receptive field of the AFDC-IRNet model and expand the global feature perception ability of the AFDC-IRNet model. The attention mechanism is used to enable the AFDC-IRNet model to focus on the wide-area feature information of the infrared image, thereby improving the AFDC-IRNet model's ability to perceive the image globally and mine the image semantic context features, thereby improving the task performance of infrared scene classification.

[0008] To achieve the above objectives, the technical solutions provided by the present invention are:

[0009] A method for infrared scene classification based on flexible deformable convolution and attention mechanism, comprising: an infrared image is input into an AFDC-IRNet model, and the AFDC-IRNet model outputs infrared scene classification information to which the infrared image belongs; wherein the AFDC-IRNet model performs infrared scene classification on the infrared image based on flexible deformable convolution and attention mechanism, and the process includes:

[0010] Step 1: The infrared image passes through the first flexible deformable convolution layer FDCN1, the first activation function layer ReLU1, and the first pooling layer MAXPool1 in sequence, and outputs the first feature map showing the most significant image features;

[0011] Step 2: The first feature map output in step 1 is passed through the first weighted feature extraction network DPANet1. Based on the attention mechanism, the second flexible deformable convolutional layer FDCN2 focuses on the global semantic feature information of the first network feature map and outputs the first network feature image.

[0012] Step 3: The first network feature image output in step 2 passes through the second flexible deformable convolution layer FDCN2, the second activation function layer ReLU2, and the second pooling layer MAXPool2 in sequence to output a second feature map; the second feature map passes through the second weighted feature extraction network DPANet2 to output a second network feature image; the second network feature image passes through the third flexible deformable convolution layer FDCN3, the third activation function layer ReLU3, and the third pooling layer MAXPool3 in sequence to output a third feature map;

[0013] Step 4: The third feature map outputted in step 3 passes through the Dropout layer and then enters the fully connected layer Dense to output the infrared scene classification information to which the infrared image belongs.

[0014] As a further limitation of the present invention, the process of performing infrared scene classification on infrared images based on flexible deformable convolution by the AFDC-IRNet model specifically includes:

[0015] (1) In the first flexible deformable convolution layer FDCN1 of step 1, the second flexible deformable convolution layer FDCN2 and the third flexible deformable convolution layer FDCN3 of step 3, in the flexible deformable convolution process, with the starting sampling point as the center, the feature point with the largest numerical difference from the starting sampling point is searched in the area of ​​(2M-1)×(2M-1) as the sampling point after the shift; if the sampling point found is occupied by other starting sampling points, the sampling point with the smallest numerical difference is searched among the remaining sampling points in the area of ​​(2M-1)×(2M-1) as the sampling point after the shift;

[0016] Among them, the values ​​with the largest and smallest numerical differences represent the offset vectors between the offset sampling point and the starting sampling point, and the expression is:

[0017] E={(x0,y0),(x1,y1),(x2,y2),…,(x6,y6),(x7,y7),(x8,y8)} Formula (1)

[0018] In formula (1), E represents the offset vector, (x0, y0) represents the offset of the sampling point after the member variable (x0, y0) in the vector E is offset relative to the starting sampling point in the horizontal and vertical directions, (x1, y1) represents the offset of the sampling point after the member variable (x1, y1) in the vector E is offset relative to the starting sampling point in the horizontal and vertical directions, (x2, y2) represents the offset of the sampling point after the member variable (x2, y2) in the vector E is offset relative to the starting sampling point in the horizontal and vertical directions, (x6, y6) represents the offset of the sampling point after the member variable (x6, y6) in the vector E is offset relative to the starting sampling point in the horizontal and vertical directions, (x7, y7) represents the offset of the sampling point after the member variable (x7, x7) in the vector E is offset relative to the starting sampling point in the horizontal and vertical directions, (x8, y8) represents the offset of the sampling point after the member variable (x8, y8) in the vector E is offset relative to the starting sampling point in the horizontal and vertical directions;

[0019] (2) The first flexible deformable convolution layer FDCN1 of step 1, the second flexible deformable convolution layer FDCN2 of step 3, and the third flexible deformable convolution layer FDCN3 of step 3, after the flexible deformable convolution completes the sampling point offset, perform convolution operation on the offset sampling point to obtain a convolution result, add the convolution result to the convolution result obtained by the initial sampling point convolution operation, and output the final convolution result y f (c0); Specifically, the final result of convolution y f(c0) is expressed as:

[0020]

[0021] In formula (2), y(c0) represents the convolution result obtained by the regional convolution mapping centered on the center point c0; w(a n ) represents the convolution weight coefficient of the corresponding point position in the region, c n Indicates the convolution position of the change in the region, Δc fn Indicates the offset position of the sampling point in the flexible deformable convolution.

[0022] As a further limitation of the present invention, in the AFDC-IRNet model:

[0023] The ReLU function is introduced into the first activation function layer ReLU1, the second activation function layer ReLU2, and the third activation function layer ReLU3 to enable the AFDC-IRNet model to learn complex nonlinear relationships. The expression of the ReLU function is:

[0024] f(x)=max(0,x) Formula (3)

[0025] In formula (3), max(0,x) means that if the input x is greater than 0, the output is x, and if the input x is less than or equal to 0, the output is 0.

[0026] As a further limitation of the present invention, the process of the AFDC-IRNet model performing infrared scene classification on infrared images based on the attention mechanism specifically includes:

[0027] The first weighted feature extraction network DPANet1 allows the second flexible deformable convolution layer FDCN2, and the second weighted feature extraction network DPANet2 allows the third flexible deformable convolution layer FDCN3 to focus on the global semantic feature information of the network feature image, specifically:

[0028] (1) Calculate the attention score of a single feature node in the input feature graph and each feature node in the weighted feature extraction network;

[0029] (2) Adding the attention scores and activating them through the ReLU function to obtain the global attention scores of all feature nodes in the feature graph;

[0030] (3) obtaining a global attention score matrix according to the global attention score, wherein the size of the global attention score matrix is ​​consistent with the size of the feature map;

[0031] (4) Calculate the dot product of the global attention score matrix and the feature map to obtain the network feature image.

[0032] As a further limitation of the present invention, the step 4 includes:

[0033] The third feature map outputted in step 3 is processed by a Dropout layer to obtain feature data;

[0034] The feature data is input into the fully connected layer Dense, where it is first multiplied by the weights and the bias term is added. It is then nonlinearly transformed using the ReLU activation function to obtain the output of the fully connected layer Dense. The Softmax function is then used to calculate the infrared scene classification information to which the infrared image belongs. Each neuron in the fully connected layer Dense has a set of weights w and a bias term b.

[0035] As a further limitation of the present invention, the function expression of the fully connected layer Dense is:

[0036] Dense(x)=ReLU(w·x+b) Formula (4)

[0037] In formula (4), w represents weight, x represents input, and b represents bias term;

[0038] The expression of the Softmax function is:

[0039]

[0040] In formula (5), z i represents the i-th element in the input vector, K represents the total number of elements in the input vector, Represents the exponential sum of all input elements, Softmax(z i ) represents the probability that the infrared image belongs to the i-th scene category.

[0041] As a further limitation of the present invention, the AFDC-IRNet model is obtained through training, and the training process includes:

[0042] (1) preprocessing the infrared image to form an image data set, and digitally processing the category label of the infrared image based on the scene category to which the infrared image belongs to form label data;

[0043] (2) Inputting the image dataset and label data into the AFDC-IRNet model in pairs for training;

[0044] (3) Using a binary cross entropy function to evaluate the loss between the output of the AFDC-IRNet model and the category label of the input infrared image; wherein: the expression of the binary cross entropy function is:

[0045] Binary_crossentropy(T,O)=-(T(log(O)+(1-T)log(1-O)) Formula (6)

[0046] In formula (6), T represents the label data of the infrared image, and O represents the output of the AFDC-IRNet model;

[0047] (4) The AFDC-IRNet model evaluates the loss result output by the binary cross entropy function. If the loss result is lower than the convergence threshold, it is considered that the parameters of the AFDC-IRNet model have achieved the ideal training effect. Otherwise, the AFDC-IRNet model updates the weight parameters of each layer in the AFDC-IRNet model using the gradient descent method through the back propagation algorithm, and repeatedly iterates the AFDC-IRNet model until the AFDC-IRNet model training reaches the convergence condition.

[0048] As a further limitation of the present invention, the backpropagation algorithm calculates the gradient of the binary cross entropy function with respect to the weights of the AFDC-IRNet model by the chain rule, updates the weights of the AFDC-IRNet model along the direction of gradient descent, and obtains the optimal solution of the weights of the AFDC-IRNet model through multiple rounds of iterations;

[0049] When updating the weights of the AFDC-IRNet model, the update rule is expressed as:

[0050]

[0051] In formula (7), W n Represents the updated model weight, W o represents the model weight before updating, L represents the loss function, η represents the learning rate, Represents the partial derivative of the loss function L with respect to the weight.

[0052] As a further limitation of the present invention, the training process further includes:

[0053] (5) The infrared thermal imager connects the video interface to the GPU computing platform to transmit the collected infrared image in real time, loads the model weight file obtained after the AFDC-IRNet model reaches the convergence condition into the GPU computing platform, performs inference calculation on the infrared image, and the AFDC-IRNet model outputs the scene category with the highest probability value.

[0054] The advantages of the present invention are:

[0055] The present invention improves the receptive field of the AFDC-IRNet model through flexible variability convolution, expands the global feature perception capability of the AFDC-IRNet model, and enables the AFDC-IRNet model to focus on the wide-area feature information of infrared images through the attention mechanism, thereby improving the AFDC-IRNet model's ability to perceive the image globally and mine the image semantic context features, thereby improving the task performance of infrared scene classification.

[0056] Additional aspects and advantages of the present invention will be set forth in part in the description which follows and, in part, will be obvious from the description which follows, or may be learned by practice of the present invention. BRIEF DESCRIPTION OF THE DRAWINGS

[0057] The above and / or additional aspects and advantages of the present invention will become apparent and readily understood from the following description of the embodiments with reference to the accompanying drawings, in which:

[0058] Figure 1 :The present invention provides a flow chart of an infrared scene classification method based on flexible deformable convolution and attention mechanism;

[0059] Figure 2 : Illustration of sampling point offset of flexible deformable convolution provided by the present invention;

[0060] Figure 3 : Comparison diagram of sampling point selection for ordinary convolution, traditional deformable convolution, and flexible deformable convolution provided by the present invention;

[0061] Figure 4 : A schematic diagram of the pooling operation provided by the present invention;

[0062] Figure 5 : Flowchart of global attention score calculation of weighted feature extraction network provided by the present invention;

[0063] Figure 6 : Schematic diagram of the output feature map calculation flow of the weighted feature extraction network DPANet provided by the present invention;

[0064] Figure 7 : A structural diagram of the AFDC-IRNet model provided by the present invention;

[0065] Figure 8 : Working principle diagram of the Dropout layer in the AFDC-IRNet model provided by the present invention;

[0066] Figure 9 : Training flow chart of the AFDC-IRNet model provided by the present invention;

[0067] Figure 10 : Illustration of the training process of the AFDC-IRNet model provided by the present invention. DETAILED DESCRIPTION

[0068] The following describes in detail embodiments of the present invention. The embodiments are exemplary and intended to explain the present invention, but are not to be construed as limiting the present invention.

[0069] To address the problem that existing infrared image scene classification methods are insufficient in extracting infrared image features and understanding semantic information, especially in the incomplete mining and extraction of multi-scale information and global features of infrared images, an embodiment of the present invention provides an infrared scene classification method based on flexible deformable convolution and attention mechanism, including: an infrared image is input into the AFDC-IRNet model, and the AFDC-IRNet model outputs infrared scene classification information to which the infrared image belongs; wherein the AFDC-IRNet model performs infrared scene classification on the infrared image based on flexible deformable convolution and attention mechanism. Figure 7 The process includes steps one to four:

[0070] Step 1: The infrared image passes through the first flexible deformable convolution layer FDCN1, the first activation function layer ReLU1 and the first pooling layer MAXPool1 in sequence, and outputs the first feature map showing the most significant image features.

[0071] In step 2 and step 1, the first feature map output is passed through the first weighted feature extraction network DPANet1. Based on the attention mechanism, the second flexible deformable convolutional layer FDCN2 focuses on the global semantic feature information of the first network feature map and outputs the first network feature image.

[0072] In step 3, the first network feature image outputted in step 2 passes through the second flexible deformable convolution layer FDCN2, the second activation function layer ReLU2 and the second pooling layer MAXPool2 in sequence, and outputs the second feature map; the second feature map passes through the second weighted feature extraction network DPANet2, and outputs the second network feature image; the second network feature image passes through the third flexible deformable convolution layer FDCN3, the third activation function layer ReLU3 and the third pooling layer MAXPool3 in sequence, and outputs the third feature map.

[0073] Step 4: The third feature map output from step 3 passes through the Dropout layer and then enters the fully connected layer Dense to output the infrared scene classification information to which the infrared image belongs.

[0074] The difference between traditional deformable convolution and ordinary convolution is that the convolution position is deformable, rather than performing convolution on a fixed M×M region of the feature map. The advantage of deformable convolution is that the model can more accurately extract interesting features in the image. In contrast, the offset position of each sampling point in traditional deformable convolutional networks is random and is not designed to map global features. However, the core of infrared scene classification tasks is to extract and classify features from the entire image. Therefore, traditional deformable convolutional networks are not conducive to wide-area image feature mining for infrared scene classification tasks.

[0075] In the flexible deformable convolution of the present invention, the offset position of each sampling point is not random during the deformable convolution process. Instead, the offset position of each sampling point is searched for the feature point with the largest numerical difference from the starting sampling point in the (2M-1)×(2M-1) area with the starting sampling point as the center, and the feature point with the largest numerical difference from the starting sampling point is used as the offset sampling point. If the sampling point found is already occupied by another starting sampling point, the sampling point with the smallest numerical difference among the remaining sampling points is used as the offset sampling point. For the sampling point selection rules, please refer to Figure 2 Assuming that the size of the convolution kernel of the deformable convolution is (M×M), if the sampling point found has been occupied by another starting sampling point, then the sampling point with the smallest value among the remaining sampling points in the (2M-1)×(2M-1) area is found as the sampling point after offset.

[0076] Specifically, in the infrared scene classification method described above in the embodiment of the present invention, the process of performing infrared scene classification on infrared images based on the flexible deformable convolution by the AFDC-IRNet model specifically includes:

[0077] (1) The first flexible deformable convolution layer FDCN1 of step one, the second flexible deformable convolution layer FDCN2 of step three, and the third flexible deformable convolution layer FDCN3, in the flexible deformable convolution process, with the starting sampling point as the center, search for the feature point with the largest numerical difference from the starting sampling point in the area of ​​(2M-1)×(2M-1) as the sampling point after offset. If the sampling point found is occupied by other starting sampling points, search for the sampling point with the smallest numerical difference among the remaining sampling points in the area of ​​(2M-1)×(2M-1) as the sampling point after offset.

[0078] The relationship between the offset sampling point and the initial sampling point is represented by a set of vectors E. Each member variable (x, y) in vector E represents the horizontal and vertical offset of the offset sampling point relative to the initial sampling point. When the size of the convolution kernel is 3×3, the offset sampling point of a sampling point needs to be searched within the 5×5 range around it, that is, (-2, -2) represents the upper left corner, and (2, 2) represents the lower right corner. Among them, the values ​​with the largest and smallest numerical differences represent the offset vectors between the offset sampling point and the initial sampling point. The offset vector E is expressed as:

[0079] E={(x0,y0),(x1,y1),(x2,y2),…,(x6,y6),(x7,y7),(x8,y8)} Formula (1)

[0080] In formula (1), E represents the offset vector, (x0, y0) represents the offset of the sampling point after the member variable (x0, y0) in vector E is offset relative to the starting sampling point in the horizontal and vertical directions, (x1, y1) represents the offset of the sampling point after the member variable (x1, y1) in vector E is offset relative to the starting sampling point in the horizontal and vertical directions, (x2, y2) represents the offset of the sampling point after the member variable (x2, y2) in vector E is offset relative to the starting sampling point in the horizontal and vertical directions, (x6, y6) represents the offset of the sampling point after the member variable (x6, y6) in vector E is offset relative to the starting sampling point in the horizontal and vertical directions, (x7, x7) represents the offset of the sampling point after the member variable (x7, y7) in vector E is offset relative to the starting sampling point in the horizontal and vertical directions, and (x8, y8) represents the offset of the sampling point after the member variable (x8, y8) in vector E is offset relative to the starting sampling point in the horizontal and vertical directions.

[0081] See also Figure 3 , ordinary convolution, traditional deformable convolution and flexible deformable convolution of the embodiment of the present invention, the selection rules of sampling points are compared as follows Figure 3 shown. Figure 3 (a) represents the sampling point selection of ordinary convolution, and its area is fixed. Figure 3 (b) represents the sampling point selection of traditional deformable convolution, where the sampling point selection has a random offset. Figure 3 The dark green boxes represent the sample points after migration. Figure 3 (c) represents the sampling point selection of flexible deformable convolution. The sampling point selection rules are the same as Figure 3 (a) and Figure 3 (b) Different, by calculating the difference of feature points in the (2M-1)×(2M-1) area around the initial sampling point, the feature point with the largest difference and not repeatedly selected is taken as the sampling point after offset.

[0082] (2) The first flexible deformable convolution layer FDCN1 of step 1, the second flexible deformable convolution layer FDCN2 of step 3, and the third flexible deformable convolution layer FDCN3, after the flexible deformable convolution completes the sampling point offset, the convolution operation is performed on the offset sampling point to obtain the convolution result, the convolution result is added to the convolution result obtained by the initial sampling point convolution operation, and the final convolution result y is output. f (c0).

[0083] For the input features Figure X , taking a 3×3 ordinary convolution kernel as an example, each output y(c0) obtained by convolution mapping in the area centered on c0 is obtained from the feature Figure X Upsample 9 positions. These 9 positions are all spread out from the center position x(c0). (-1,-1) represents the upper left corner of x(c0), (1,1) represents the lower right corner of x(c0), and the position change is represented by c n Indicates. Use w(c n ) represents the convolution weight coefficient of the corresponding point position, then the convolution output result y(c0) is expressed as: y(c0)=∑w(c n )·x(c0+c n ). Similarly, for traditional deformable convolution, the convolution output result y t (c0) is expressed as: y t (c0)=∑w(c n )·x(c0+c n +Δc tn ), where Δc tn Represents the random offset position of the sampling point in the traditional deformable convolution.

[0084] After completing the maximum difference sampling point offset, the flexible deformable convolution of the embodiment of the present invention performs a convolution operation on the offset sampling point, and adds the result of the convolution operation to the result of the convolution on the initial sampling point in the 3×3 area as the final result of the flexible deformable convolution after one convolution. Specifically, the final result y output after convolution is f (c0) is expressed as:

[0085]

[0086] In formula (2), y(c0) represents the convolution result obtained by the regional convolution mapping centered on the center point c0; w(a n ) represents the convolution weight coefficient of the corresponding point position in the region, c n Indicates the convolution position of the change in the region, Δc fnIndicates the offset position of the sampling point in the flexible deformable convolution. That is, the feature point with the largest difference and not repeatedly selected is taken as the sampling point position after the offset. In a single convolution operation, the flexible deformable convolution of the embodiment of the present invention not only maps the initial sampling points of a fixed 3×3 area, but also maps the feature points in a 5×5 area near the initial sampling point whose values ​​differ greatly from the initial sampling point. This fully considers the influence of those feature points in the wide-area features of the infrared scene image that are significantly different from the initial sampling points on the convolution results, which is beneficial to the classification of the entire infrared image scene.

[0087] Specifically: In the above-mentioned AFDC-IRNet model of the embodiment of the present invention: for a high-resolution single-channel infrared image, after performing flexible deformable convolution to extract features, in order to introduce nonlinear characteristics into the image feature extraction process, an activation function is introduced to enable the AFDC-IRNet model to have the ability to learn complex nonlinear relationships. Therefore, the embodiment of the present invention introduces the ReLU function as an activation function. Specifically, the first activation function layer ReLU1, the second activation function layer ReLU2, and the third activation function layer ReLU3 introduce the ReLU function to enable the AFDC-IRNet model to have the ability to learn complex nonlinear relationships; wherein, the expression of the ReLU function is:

[0088] f(x)=max(0,x) Formula (3)

[0089] In formula (3), max(0,x) means that if the input x is greater than 0, the output is x, and if the input x is less than or equal to 0, the output is 0.

[0090] After introducing the activation function, the embodiment of the present invention introduces the maximum pooling (MAXPool) operation to reduce the dimension of the feature map and the number of parameters in the model to reduce the complexity of the AFDC-IRNet model and improve the ability of the AFDC-IRNet model to extract prominent features in the scene. The pooling layer of the embodiment of the present invention calculates the maximum value in the 2×2 square area as the output. The schematic diagram of the pooling operation is as follows Figure 4 shown.

[0091] In order to further improve the weighted feature extraction network's ability to mine scene elements and extract global features in infrared images, the weighted fusion feature extraction network DPANet in the embodiment of the present invention is used. Figure 6 , calculate the input features of size M×N Figure X A single feature node such as x 00 The attention score (α 00 , α 01 , α 02 ...α mn), add these attention scores and activate them through the ReLU function to get x 00 The global attention score a 00 , feature node x 00 The global attention score a 00 The calculation process is as follows Figure 5 As shown. The embodiment of the present invention first performs an operation on the node x in the input feature graph. 00 Multiply by the weight factor w q 00 Weighted, each feature node in the feature graph (including node x 00 ) multiplied by the weight factor w k , through the weight factor w q Weighted x 00 With the weight factor w k All weighted feature nodes are multiplied together to obtain node x 00 With the attention score of all feature nodes (α 00 , α 01 , α 02 ...α mn ). And node x 00 The global attention score a 00 The attention score is passed through the weight factor w a The weighted sum is activated by the ReLU function.

[0092] Specifically: In the process of infrared scene classification of infrared images based on the attention mechanism, the above-mentioned AFDC-IRNet model of the present invention allows the first weighted feature extraction network DPANet1 to allow the second flexible deformable convolution layer FDCN2, and the second weighted feature extraction network DPANet2 allows the third flexible deformable convolution layer FDCN3 to focus on the global semantic feature information of the network feature image, specifically: (1) calculating the attention score of a single feature node in the input feature map and each feature node in the weighted feature extraction network; (2) adding the attention scores and activating them through the ReLU function to obtain the global attention score of all feature nodes in the feature map; (3) obtaining a global attention score matrix based on the global attention score, and the size of the global attention score matrix is ​​consistent with the size of the feature map; (4) calculating the dot product of the global attention score matrix and the feature map to obtain the network feature image.

[0093] See also Figure 6 In the embodiment of the present invention, the weighted feature extraction network DPANet calculates the global attention scores of all feature nodes to obtain the global attention score matrix A. The size of the global attention score matrix A is related to the feature Figure X The size of the global attention score matrix A is consistent with the feature Figure XCalculate the dot product and get the output feature map O of the weighted feature extraction network DPANet. The weighted feature extraction network DPANet is constructed by Figure X The weighted global attention score is calculated for the nodes in , and the relationship between a single node and other feature points is obtained to mine the pixel-level contextual semantic relationship features in the infrared scene image. Figure X Dot product fusion is performed to fuse the global semantic relationship features of the infrared image with the initial infrared image features to obtain reliable global image information, which facilitates infrared scene classification.

[0094] Please continue reading Figure 7 , a single-channel infrared image is input into the AFDC-IRNet model, and first passes through FDCN1, ReLU1, and MAXPool1. The convolution kernel size of FDCN1 is 1*7*7, that is, the number of channels of the convolution kernel is 1, the length and width are 7 respectively, and the convolution step is 1. The larger convolution kernel size of the embodiment of the present invention allows the AFDC-IRNet model to mine local features with a larger range in the initial convolution operation, reducing the subsequent feature map size and computational complexity of the AFDC-IRNet model. Compared with ordinary convolutional neural networks, the flexible deformable convolution layer FDCN adds an offset to each sampling point in its convolution kernel, extracts features from the infrared image through arbitrary deformation, and improves the information mining capability. ReLU1 delinearizes the output of FDCN1 and retains the nonlinear characteristics of the infrared image. MAXPool1 removes redundant information in FDCN1 and compresses the features of the infrared image, so that the output feature map presents the most significant image features.

[0095] Secondly, the AFDC-IRNet model goes through a cycle of DPANet1-FDCN2-ReLU2-MAXPool2-DPANet2-FDCN3-ReLU3-MAXPool3. DPANet1 and DPANet2 are weighted feature extraction networks based on the attention mechanism, while FDCN2 and FDCN3 focus on the global semantic feature information of infrared images to better classify infrared image scenes. Then, the AFDC-IRNet model goes through MAXPool3 and then connects to the Dropout layer. Figure 8 , Figure 8 The working principle diagram of the Dropout layer is as follows. The Dropout layer randomly drops some parameters during the model training process to prevent the AFDC-IRNet model from falling into overfitting and failing to obtain the global optimal solution.

[0096] More specifically, step 4 of the embodiment of the present invention includes:

[0097] The third feature map output in step 3 is processed by the Dropout layer to obtain feature data;

[0098] The feature data is input into the fully connected layer Dense, where it is first multiplied by the weights and the bias term is added. It is then nonlinearly transformed using the ReLU activation function to obtain the output of the fully connected layer Dense. The Softmax function is then used to calculate the infrared scene classification information to which the infrared image belongs. Each neuron in the fully connected layer Dense has a set of weights w and a bias term b.

[0099] The function expression of the fully connected layer Dense in the embodiment of the present invention is:

[0100] Dense(x)=ReLU(w·x+b) Formula (4)

[0101] In formula (4), w represents weight, x represents input, and b represents bias term;

[0102] The Softmax function in the embodiment of the present invention is expressed as:

[0103]

[0104] In formula (5), z i Represents the i-th element in the input vector, and K represents the total number of elements in the input vector, that is, the total number of categories; Represents the exponential sum of all input elements. The exponential sum is used as a normalization constant to ensure that the sum of all output values ​​is 1. Softmax(z i ) represents the probability that the infrared image belongs to the i-th scene category. The scene category with the largest probability value is the prediction result of the AFDC-IRNet model for the category to which the infrared image belongs.

[0105] More specifically, the AFDC-IRNet model in the embodiment of the present invention is obtained through training, see Figure 9 , the training process of the embodiment of the present invention includes:

[0106] (1) Preprocess the infrared image to form an image dataset. Based on the scene category to which the infrared image belongs, digitize the category label of the infrared image to form label data. Specifically, the image dataset of the infrared image is preprocessed, and the image is resized to a fixed size of 960*512 for input into the AFDC-IRNet model. When digitizing the category label of the infrared image, it is preferred to use one-hot encoding to distinguish different categories. For example, scene category 1 (sky) is digitized as 100000...0, and scene category 2 (grass) is digitized as 010000...0, etc., so as to facilitate input into the AFDC-IRNet model for machine calculation.

[0107] (2) The image dataset and label data are input into the AFDC-IRNet model in pairs for training; specifically, the resized image dataset and digitized label data are input into the backbone network in pairs for training, and the binary cross entropy function is used to evaluate the loss between the output of the AFDC-IRNet model and the input category label.

[0108] (3) Use the binary cross entropy function to evaluate the loss between the output of the AFDC-IRNet model and the category label of the input infrared image; where: the expression of the binary cross entropy function is:

[0109] Binary_crossentropy(T,O)=-(T(log(O)+(1-T)log(1-O)) Formula (6)

[0110] In formula (6), T represents the label data of the infrared image, that is, the true value of the category to which the infrared image belongs, and O represents the output of the AFDC-IRNet model, that is, the predicted value of the infrared image by the AFDC-IRNet model.

[0111] (4) The AFDC-IRNet model evaluates the loss result output by the binary cross entropy function. If the loss result is lower than the convergence threshold (the convergence threshold is pre-set according to the image data scale), it is considered that the parameters of the AFDC-IRNet model have achieved the ideal training effect. Otherwise, the AFDC-IRNet model uses the back-propagation algorithm and the gradient descent method to update the weight parameters of each layer in the AFDC-IRNet model, and repeatedly iterates the AFDC-IRNet model until the AFDC-IRNet model training reaches the convergence condition.

[0112] Furthermore, the backpropagation algorithm described above in the embodiment of the present invention calculates the gradient of the binary cross entropy function with respect to the AFDC-IRNet model weights using the chain rule, updates the AFDC-IRNet model weights along the gradient descent direction at a certain learning rate, and allows the AFDC-IRNet model weights to reach the optimal solution after multiple rounds of iteration.

[0113] Furthermore, when updating the weights of the AFDC-IRNet model in the embodiment of the present invention, the update rule for each model weight is expressed as follows:

[0114]

[0115] In formula (7), W n Represents the updated model weight, W o represents the model weight before updating, L represents the loss function, η represents the learning rate, Represents the partial derivative of the loss function L with respect to the weight.

[0116] For further information, see Figure 10 The training process of the embodiment of the present invention is based on (1)-(4), and further includes: (5) the infrared thermal imager connects the video interface to the GPU computing platform to transmit the collected infrared image in real time, loads the model weight file obtained after the AFDC-IRNet model reaches the convergence condition into the GPU computing platform, performs inference calculation on the infrared image, and the AFDC-IRNet model outputs the scene category with the highest probability value. Please continue to refer to Figure 10 In this embodiment of the present invention, infrared images and corresponding scene labels are first prepared to form a training dataset. The training dataset is then input into the AFDC-IRNet model for training. After the AFDC-IRNet model converges, an ideal weight file for the AFDC-IRNet model is obtained (the weight file contains the network structure of the AFDC-IRNet model). Next, an infrared thermal imager is prepared, and its video interface is connected to a GPU computing platform to transmit the captured infrared images in real time. The weight file of the trained AFDC-IRNet model is loaded into the GPU computing platform, and the AFDC-IRNet model is inferred and calculated on the input infrared image. The AFDC-IRNet model in this embodiment of the present invention does not require backpropagation during the inference calculation process and directly outputs the scene category with the highest probability value, thereby realizing the classification function of the infrared image scene.

[0117] The embodiments of the present invention improve the receptive field of the AFDC-IRNet model through flexible variability convolution, expand the global feature perception capability of the AFDC-IRNet model, and use the attention mechanism to enable the AFDC-IRNet model to focus on the wide-area feature information of infrared images, thereby improving the AFDC-IRNet model's ability to perceive the image globally and mine the image semantic context features, thereby improving the task performance of infrared scene classification.

[0118] The above description is only a specific embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any technician familiar with this technical field can easily think of various equivalent modifications or replacements within the technical scope disclosed in the present invention, and these modifications or replacements should all be included in the scope of protection of the present invention.

Claims

1. An infrared scene classification method based on flexible deformable convolution and attention mechanism, characterized by: include: The infrared image is input into the AFDC-IRNet model, which outputs the infrared scene classification information to which the infrared image belongs. The AFDC-IRNet model classifies the infrared image into infrared scenes based on flexible deformable convolution and attention mechanism. The process includes: Step 1: The infrared image passes through the first flexible deformable convolution layer FDCN1, the first activation function layer ReLU1, and the first pooling layer MAXPool1 in sequence, and outputs the first feature map showing the most significant image features; Step 2: The first feature map output in step 1 is passed through the first weighted feature extraction network DPANet1. Based on the attention mechanism, the second flexible deformable convolutional layer FDCN2 focuses on the global semantic feature information of the first network feature map and outputs the first network feature image. Step 3: The first network feature image output in step 2 passes through the second flexible deformable convolution layer FDCN2, the second activation function layer ReLU2, and the second pooling layer MAXPool2 in sequence to output a second feature map; the second feature map passes through the second weighted feature extraction network DPANet2 to output a second network feature image; the second network feature image passes through the third flexible deformable convolution layer FDCN3, the third activation function layer ReLU3, and the third pooling layer MAXPool3 in sequence to output a third feature map; Step 4: The third feature map outputted in step 3 passes through the Dropout layer and then enters the fully connected layer Dense to output the infrared scene classification information to which the infrared image belongs.

2. The infrared scene classification method based on flexible deformable convolution and attention mechanism according to claim 1 is characterized in that: The process of infrared scene classification of infrared images based on flexible deformable convolution by the AFDC-IRNet model specifically includes: (1) In the first flexible deformable convolution layer FDCN1 of step 1, the second flexible deformable convolution layer FDCN2 and the third flexible deformable convolution layer FDCN3 of step 3, in the flexible deformable convolution process, with the starting sampling point as the center, the feature point with the largest numerical difference from the starting sampling point is searched in the area of ​​(2M-1)×(2M-1) as the sampling point after the shift; if the sampling point found is occupied by other starting sampling points, the sampling point with the smallest numerical difference is searched among the remaining sampling points in the area of ​​(2M-1)×(2M-1) as the sampling point after the shift; Among them, the values ​​with the largest and smallest numerical differences represent the offset vectors between the offset sampling point and the starting sampling point, and the expression is: E={(x0,y0),(x1,y1),(x2,y2),…,(x6,y6),(x7,y7),(x8,y8)} Formula (1) In formula (1), E represents the offset vector, (x0, y0) represents the offset of the sampling point after the member variable (x0, y0) in the vector E is offset relative to the starting sampling point in the horizontal and vertical directions, (x1, y1) represents the offset of the sampling point after the member variable (x1, y1) in the vector E is offset relative to the starting sampling point in the horizontal and vertical directions, (x2, y2) represents the offset of the sampling point after the member variable (x2, y2) in the vector E is offset relative to the starting sampling point in the horizontal and vertical directions, (x6, y6) represents the offset of the sampling point after the member variable (x6, y6) in the vector E is offset relative to the starting sampling point in the horizontal and vertical directions, (x7, y7) represents the offset of the sampling point after the member variable (x7, y7) in the vector E is offset relative to the starting sampling point in the horizontal and vertical directions, (y8, y8) represents the offset of the sampling point after the member variable (x8, y8) in the vector E is offset relative to the starting sampling point in the horizontal and vertical directions; (2) The first flexible deformable convolution layer FDCN1 of step 1, the second flexible deformable convolution layer FDCN2 of step 3, and the third flexible deformable convolution layer FDCN3 of step 3, after the flexible deformable convolution completes the sampling point offset, perform convolution operation on the offset sampling point to obtain a convolution result, add the convolution result to the convolution result obtained by the initial sampling point convolution operation, and output the final convolution result y f (c0); Specifically, the final result of convolution y f (c0) is expressed as: In formula (2), y(c0) represents the convolution result obtained by the regional convolution mapping centered on the center point c0; w(a n ) represents the convolution weight coefficient of the corresponding point position in the region, c n Indicates the convolution position of the change in the region, Δc fn Indicates the offset position of the sampling point in the flexible deformable convolution.

3. The infrared scene classification method based on flexible deformable convolution and attention mechanism according to claim 1 is characterized in that: In the AFDC-IRNet model: The ReLU function is introduced into the first activation function layer ReLU1, the second activation function layer ReLU2, and the third activation function layer ReLU3 to enable the AFDC-IRNet model to learn complex nonlinear relationships. The expression of the ReLU function is: f(x)=max(0,x) Formula (3) In formula (3), max(0,x) means that if the input x is greater than 0, the output is x, and if the input x is less than or equal to 0, the output is 0.

4. The infrared scene classification method based on flexible deformable convolution and attention mechanism according to claim 1 is characterized in that: The process of infrared scene classification of infrared images based on the attention mechanism of the AFDC-IRNet model specifically includes: The first weighted feature extraction network DPANet1 allows the second flexible deformable convolution layer FDCN2, and the second weighted feature extraction network DPANet2 allows the third flexible deformable convolution layer FDCN3 to focus on the global semantic feature information of the network feature image, specifically: (1) Calculate the attention score of a single feature node in the input feature graph and each feature node in the weighted feature extraction network; (2) Adding the attention scores and activating them through the ReLU function to obtain the global attention scores of all feature nodes in the feature graph; (3) obtaining a global attention score matrix according to the global attention score, wherein the size of the global attention score matrix is ​​consistent with the size of the feature map; (4) Calculate the dot product of the global attention score matrix and the feature map to obtain the network feature image.

5. The infrared scene classification method based on flexible deformable convolution and attention mechanism according to claim 1, characterized in that: The fourth step includes: The third feature map outputted in step 3 is processed by a Dropout layer to obtain feature data; The feature data is input into the fully connected layer Dense, where it is first multiplied by the weights and the bias term is added. It is then nonlinearly transformed using the ReLU activation function to obtain the output of the fully connected layer Dense. The Softmax function is then used to calculate the infrared scene classification information to which the infrared image belongs. Each neuron in the fully connected layer Dense has a set of weights w and a bias term b.

6. The infrared scene classification method based on flexible deformable convolution and attention mechanism according to claim 5, characterized in that: The function expression of the fully connected layer Dense is: Dense(x)=ReLU(w·x+b) Formula (4) In formula (4), w represents weight, x represents input, and b represents bias term; The expression of the Softmax function is: In formula (5), z i represents the i-th element in the input vector, K represents the total number of elements in the input vector, Represents the exponential sum of all input elements, Softmax(z i ) represents the probability that the infrared image belongs to the i-th scene category.

7. The infrared scene classification method based on flexible deformable convolution and attention mechanism according to claim 1, characterized in that: The AFDC-IRNet model is obtained through training, and the training process includes: (1) preprocessing the infrared image to form an image data set, and digitally processing the category label of the infrared image based on the scene category to which the infrared image belongs to form label data; (2) Inputting the image dataset and label data into the AFDC-IRNet model in pairs for training; (3) Using a binary cross entropy function to evaluate the loss between the output of the AFDC-IRNet model and the category label of the input infrared image; wherein: the expression of the binary cross entropy function is: Binary_crossentropy(T,O)=-(T(log(O)+(1-T)log(1-O)) Formula (6) In formula (6), T represents the label data of the infrared image, and O represents the output of the AFDC-IRNet model; (4) The AFDC-IRNet model evaluates the loss result output by the binary cross entropy function. If the loss result is lower than the convergence threshold, it is considered that the parameters of the AFDC-IRNet model have achieved the ideal training effect. Otherwise, the AFDC-IRNet model updates the weight parameters of each layer in the AFDC-IRNet model using the gradient descent method through the back propagation algorithm, and repeatedly iterates the AFDC-IRNet model until the convergence condition is reached.

8. The infrared scene classification method based on flexible deformable convolution and attention mechanism according to claim 7, characterized in that: The backpropagation algorithm calculates the gradient of the binary cross entropy function relative to the AFDC-IRNet model weights using the chain rule, updates the AFDC-IRNet model weights along the gradient descent direction, and allows the AFDC-IRNet model weights to reach the optimal solution after multiple rounds of iterations. When updating the weights of the AFDC-IRNet model, the update rule is expressed as: In formula (7), W n Represents the updated model weight, W o represents the model weight before updating, L represents the loss function, η represents the learning rate, Represents the partial derivative of the loss function L with respect to the weight.

9. The infrared scene classification method based on flexible deformable convolution and attention mechanism according to claim 7, characterized in that: The training process also includes: (5) The infrared thermal imager connects the video interface to the GPU computing platform to transmit the collected infrared image in real time, loads the model weight file obtained after the AFDC-IRNet model reaches the convergence condition into the GPU computing platform, performs inference calculation on the infrared image, and the AFDC-IRNet model outputs the scene category with the highest probability value.