A training method and device for an infrared image detection model, and an infrared image detection method
Through the training method of infrared image detection model, the convolutional neural network and the Transformer encoder decoder are used, combined with multi-layer perceptual loss function and binary graph matching, small object detection of photovoltaic panels is optimized, detection accuracy and feature extraction capabilities are improved, and the problem of insufficient accuracy of existing models is solved.
Patent Information
- Application Number
- CN202510289005.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-12
- Publication Date
- 2025-07-08
- Estimated Expiration
- 2045-03-12
AI Technical Summary
The existing deep learning models have low detection accuracy of small target defects of photovoltaic panels in photovoltaic power plants, making it difficult to accurately identify defects of small targets such as heat spots and diode failures.
The infrared image detection model is adopted, and the training is carried out through convolutional neural network, spectral domain Transformer encoder and multi-scale Transformer decoder. Combining the multi-layer perceptual loss function and binary graph matching results, the model is optimized to improve the detection accuracy of small objects.
The detection accuracy of the infrared image detection model for small targets is improved, feature extraction ability and detail reduction degree are enhanced, and defects on photovoltaic panels can be better identified.
Smart Images

Figure CN119785127B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of artificial intelligence technology, and in particular, to a training method and device for an infrared image detection model and an infrared image detection method. Background Art
[0002] In a photovoltaic power station, defects in photovoltaic panels can lead to energy loss and a decrease in system efficiency. With the continuous expansion of the scale of photovoltaic power stations, how to detect the energy loss areas caused by defects such as cracks, hot spots, blockages, and corrosion in photovoltaic panels has attracted increasing attention.
[0003] Through the unmanned aerial vehicle (UAV) infrared inspection technology, large areas of photovoltaic panels can be quickly inspected at high altitude, which has the advantages of high efficiency and non-destruction, and has been widely used in the defect detection of photovoltaic panels. Infrared thermography technology can detect the temperature distribution by measuring the infrared radiation of an object. After an infrared image is captured by the UAV infrared inspection technology, a deep learning model needs to be used to identify the abnormal areas on the photovoltaic panels. However, due to the small number of label files that can be provided by photovoltaic power stations, the current deep learning models have low detection accuracy for small targets such as hot spots and diode failures, and it is difficult to detect the defects of small targets.
[0004] Therefore, how to achieve a relatively accurate detection of small targets in infrared images has become a problem to be solved. Summary of the Invention
[0005] Based on the above problems, the present application provides a training method and device for an infrared image detection model and an infrared image detection method, which can achieve a relatively accurate detection of small targets in infrared images.
[0006] The embodiments of the present application disclose the following technical solutions:
[0007] In a first aspect, an embodiment of the present application provides a training method for an infrared image detection model, the method including:
[0008] Obtaining training pictures including infrared images;
[0009] Based on the training pictures, outputting a defect detection result through a pre-constructed infrared image detection model; the number of prediction boxes generated by the infrared image detection model is greater than the number of real targets; the class label in the defect detection result is selected from the real class to which the real target belongs and an empty class representing no real target;
[0010] Calculating a first loss function based on the bipartite graph matching result between the defect detection result and the defect label corresponding to the training pictures;
[0011] Calculate the sum of the first loss function and the multi-layer perception loss term to obtain the loss function; the multi-layer perception loss term includes the sum of the weighted perception losses corresponding to the multi-layer activation functions in the infrared image detection model;
[0012] Iteratively train the infrared image detection model based on the loss function.
[0013] Optionally, the infrared image detection model includes a convolutional neural network, a spectral domain Transformer encoder, a multi-scale Transformer decoder, and a fully connected layer;
[0014] The convolutional neural network is used to extract the feature markers of the training pictures to obtain a feature map;
[0015] The spectral domain Transformer encoder is used to encode the feature map to obtain an encoded feature;
[0016] The multi-scale Transformer decoder is used to obtain a pixel-by-pixel defect prediction result based on the input encoded feature;
[0017] The fully connected layer is used to output the defect detection result of the training picture based on the defect prediction result.
[0018] Optionally, the depths of both the spectral domain Transformer encoder and the multi-scale Transformer decoder are 6; the spectral domain Transformer encoder and the multi-scale Transformer decoder are symmetric structures;
[0019] The last layer of the spectral domain Transformer encoder inputs the encoded feature into the decoding modules at each depth in the multi-scale Transformer decoder.
[0020] Optionally, the spectral domain Transformer encoder based on the self-attention mechanism is specifically used for:
[0021] Map the feature map from the spatial domain to the spectral domain to obtain an encoded feature; the encoded feature includes the global dependencies between different feature markers.
[0022] Optionally, the multi-scale Transformer decoder includes a multi-scale feature extraction module and a multi-layer regression prediction module; the multi-scale Transformer decoder is specifically used for:
[0023] Through the multi-scale feature extraction module, based on the multi-head attention mechanism, layer normalization operation, and fully connected operation, multi-scale features are extracted from the encoded features; the multi-scale features are fused through a convolutional layer with a convolution kernel size of 1 to obtain fused features;
[0024] Through the multi-layer regression prediction module, a defect prediction result is obtained based on the fused features.
[0025] Optionally, the number of layers of the multi-scale Transformer decoder is 4, and the corresponding spatial resolution levels of different layers are 1, 1 / 2, 1 / 4, and 1 / 8 respectively.
[0026] Optionally, before obtaining the training pictures containing infrared images, the method further includes:
[0027] Obtain initial pictures containing infrared images; the initial pictures are collected by an infrared camera;
[0028] Based on the distortion coefficient of the infrared camera and the camera internal parameter matrix, coordinate transformation is performed on the initial pictures to obtain training pictures.
[0029] Optionally, the number of prediction boxes generated by the infrared image detection model is 50.
[0030] In a second aspect, an embodiment of the present application provides a training device for an infrared image detection model, the device includes: an acquisition module, an output module, a first calculation module, a second calculation module, and a training module;
[0031] The acquisition module is used to acquire training pictures containing infrared images;
[0032] The output module is used to output a defect detection result based on the training pictures through a pre-constructed infrared image detection model; the number of prediction boxes generated by the infrared image detection model is greater than the number of real targets; the defect detection result includes the real category to which the real target belongs and an empty category indicating no real target;
[0033] The first calculation module is used to calculate a first loss function based on the bipartite graph matching result between the defect detection result and the defect label corresponding to the training pictures;
[0034] The second calculation module is used to calculate the sum of the first loss function and a multi-layer perception loss term to obtain a loss function; the multi-layer perception loss term includes the sum of the weighted perception losses corresponding to the multi-layer activation functions in the infrared image detection model;
[0035] The training module is used to iteratively train the infrared image detection model based on the loss function.
[0036] In a third aspect, an embodiment of the present application provides an infrared image detection method, and the method includes:
[0037] Obtain a to-be-inspected picture including an infrared image;
[0038] Based on the to-be-inspected picture, output a defect detection result through an infrared image detection model trained by the training method of the infrared image detection model according to any implementation manner in the first aspect.
[0039] Compared with the prior art, the present application has the following beneficial effects:
[0040] An embodiment of the present application provides a training method for an infrared image detection model. In this method, first, obtain training pictures including infrared images; then, based on the training pictures, output a defect detection result through a pre-constructed infrared image detection model; the number of prediction boxes generated by the infrared image detection model is greater than the number of real targets; the class labels in the defect detection result are selected from the real classes to which the real targets belong and the empty class representing no real targets; then, calculate a first loss function based on the bipartite graph matching result between the defect detection result and the defect label corresponding to the training picture; next, calculate the sum of the first loss function and the multi-layer perception loss term to obtain a loss function; the multi-layer perception loss term includes the sum of the weighted perception losses corresponding to the multi-layer activation functions in the infrared image detection model; finally, perform iterative training on the infrared image detection model based on the loss function. Thus, the first loss function and the multi-layer perception loss term obtained based on the bipartite graph matching result are used to train the infrared image detection model with a composite loss function. The trained infrared image detection model can achieve more accurate detection of small targets in infrared images. Among them, the first loss function adjusted based on the bipartite graph matching result can minimize the difference between the prediction box and the real target, which helps to improve the detection accuracy of the infrared image detection model for small target objects; and through the multi-layer perception loss term, the feature extraction ability and detail restoration degree of the infrared image detection model can be effectively improved. Through the way of layer-by-layer weighting, it ensures that the influence of the perception losses at different levels on the final result is reasonably balanced. Thus, while retaining the high-level information, the ability to capture low-level details is enhanced, and better constraints can be imposed between small real targets and prediction targets. Description of the Drawings
[0041] To more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the following drawings are only some embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0042] Figure 1Flowchart of a training method for an infrared image detection model provided by an embodiment of this application;
[0043] Figure 2 Schematic diagram of radial distortion provided by an embodiment of this application;
[0044] Figure 3 Schematic diagram of the formation reason of tangential distortion provided by an embodiment of this application;
[0045] Figure 4 Schematic diagram of imaging projection provided by an embodiment of this application;
[0046] Figure 5 Architecture diagram of an infrared image detection model provided by an embodiment of this application;
[0047] Figure 6 Structural diagram of a spectral domain Transformer encoder provided by an embodiment of this application;
[0048] Figure 7 Structural diagram of a multi-scale Transformer decoder provided by an embodiment of this application;
[0049] Figure 8 Schematic diagram of a multi-level encoding and decoding connection mechanism provided by an embodiment of this application;
[0050] Figure 9 Schematic diagram of a training device for an infrared image detection model provided by an embodiment of this application. Detailed implementation manners
[0051] A training method, device, and infrared image detection method for an infrared image detection model provided by this application can be used in the field of artificial intelligence. The above is only an example and does not limit the application fields of a training method, device, and infrared image detection method for an infrared image detection model provided by this application.
[0052] Terms such as "first", "second", "third", and "fourth" in the specification, claims, and drawings of this application are used to distinguish different objects, rather than to limit a specific order.
[0053] In the embodiments of this application, words such as "as an example" or "for example" are used to represent examples, illustrations, or explanations. Any embodiment or design solution described as "as an example" or "for example" in the embodiments of this application should not be interpreted as being more preferred or having more advantages than other embodiments or design solutions. Rather, using words such as "as an example" or "for example" is intended to present related concepts in a specific manner.
[0054] The terms used in the embodiments section of this application are only for explaining the specific embodiments of this application and are not intended to limit this application.
[0055] To enable those skilled in the art to better understand the solution of this application, the following will clearly and completely describe the technical solutions in the embodiments of this application with reference to the accompanying drawings in the embodiments of this application. Obviously, the described embodiments are only a part of the embodiments of this application, rather than all of the embodiments. Based on the embodiments in this application, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the scope of protection of this application.
[0056] See Figure 1 , which is a flowchart of a method for training an infrared image detection model provided by an embodiment of this application. The method includes:
[0057] S101: Obtain training pictures containing infrared images.
[0058] As an example, initial pictures containing infrared images captured by an infrared camera such as an infrared thermal imager carried by a device such as a drone can be used to obtain training pictures; the training pictures carry defect labels, and the training pictures can be divided into a training set, a test set, and a validation set for training the infrared image detection model. As an example, the infrared image can be an infrared image of a photovoltaic panel taken from a high altitude, and the defect labels can include but are not limited to the defect areas and defect categories of the photovoltaic panel in the infrared image.
[0059] Optionally, after obtaining the initial pictures containing infrared images, coordinate transformation can be performed on the initial pictures based on the distortion coefficients and the camera internal parameter matrix of the infrared camera, so as to obtain the training pictures after distortion correction. Among them, the data format of the initial pictures is JPG, and the color palette of the infrared thermal imager (infrared camera) is defaulted to the black-red mode and can be adjusted to the iron-red mode according to actual needs.
[0060] Specifically, on the one hand, due to the lens of the infrared camera carried by the drone, radial and tangential distortions will appear around the initial pictures, and correction needs to be performed through the distortion coefficients. The distortion coefficients [k1, k2, p1, p2] include the radial distortion coefficients k1 and k2, and the tangential distortion coefficients p1 and p2.
[0061] Among them, radial distortion is caused by the irregularity of the lens shape of the lens and the modeling method, resulting in different focal lengths in different regions of the lens. At places farther from the center of the lens, the degree of deflection of the light will change, and it may deflect more, forming a pincushion distortion, as shown in Figure 2 in (a), or deflect less, forming a barrel distortion, as shown in Figure 2As shown in Fig. (b). Among them, the barrel distortion coefficient k1 > 0, and the pincushion distortion coefficient k1 < 0. The tangential distortion comes from the assembly process of the infrared camera. It is caused by defects in lens manufacturing, which makes the lens itself not parallel to the image plane (camera sensor), as Figure 3 shown.
[0062] On the other hand, it is necessary to establish the conversion relationship between the world coordinate, camera coordinate, and image coordinate through the camera internal parameter matrix. Refer to Figure 4 , regarding the infrared camera as a pinhole, a point P in the real world passes through the optical center O of the infrared camera and projects onto the physical imaging plane, becoming point P'. Assume that the coordinates (world coordinates) of point P among the real-world points are [X, Y, Z], and the coordinates (camera coordinates) of the imaged point P' are [X', Y', Z']. The distance between the physical imaging plane and the optical center O is f (i.e., the focal length). Introduce the image coordinate system. There is a pixel plane o-u-v fixed on the physical imaging plane, and the coordinates of P' in the image coordinate system (i.e., image coordinates) are denoted as [u, v]. The image coordinate system is usually defined such that the origin o' is located at the upper left corner of the image, the u-axis is parallel to the x-axis to the right, the v-axis is parallel to the y-axis downward, and K is the camera internal parameter matrix.
[0063] Thus, through the coordinate conversion formula shown below, the three-dimensional world coordinates can be converted into two-dimensional image coordinates:
[0064]
[0065] As an example, first project the point in the world coordinate system onto the physical imaging plane to obtain the coordinates of the point in the camera coordinate system; then, project the point in the camera coordinate system onto the normalized plane. On this plane, the Z-direction coordinate is normalized to 1 to obtain the normalized coordinates; then, based on the distortion coefficient and the distortion correction formula on the normalized coordinates, calculate the undistorted normalized coordinates of the point; finally, use the internal parameter matrix K to map the point to the image coordinate system based on the corrected normalized coordinates, so as to convert the three-dimensional world coordinates into two-dimensional image coordinates and obtain an undistorted training image.
[0066] S102: Based on the training pictures, output the defect detection results through a pre-constructed infrared image detection model.
[0067] Among them, the number of prediction boxes generated by the infrared image detection model is greater than the real target; the class labels in the defect detection results are selected from the real classes to which the real target (Ground Truth) belongs and the empty class Φ representing no real target. The real target refers to the object or area accurately identified in the annotation data.
[0068] Specifically, for an infrared image, assuming that the number of real targets contained therein is m, and the number of targets that can be detected by a pre-set infrared image detection model is N, in the embodiments of the present application, N>m, preferably N>>m, that is, the number of prediction boxes generated by the infrared image detection model is much larger than the number of real targets. As an example, the number of prediction boxes generated by the infrared image detection model can be 50.
[0069] In order to quickly match the positions of the predicted targets and the real targets, the embodiments of the present application construct an empty category Φ representing no real targets and add it to the target categories. Then, the N-m predicted targets that do not correspond to any real targets can be paired with the empty category Φ. Thus, the pairing task of the predicted targets and the real targets can be regarded as a bipartite graph matching problem of two equal-capacity sets, simplifying the matching process and improving the efficiency of model training. Among them, the target categories include real categories and empty categories.
[0070] S103: Calculate a first loss function based on the bipartite graph matching result between the defect detection result and the defect label corresponding to the training picture.
[0071] Specifically, for the defect detection result (prediction box W) and the defect label (real target U) corresponding to the training picture, two node sets can be created to construct a bipartite graph. Among them, for each pair (u j , w j ), if the prediction box w j covers the real target u j , then an edge is added between these two nodes, and a weight is assigned to this edge based on the intersection over union (IoU). In the matching, if a prediction box overlaps with multiple real targets, the real target with the highest IoU value is selected for matching; if a real target is covered by multiple prediction boxes, the prediction box with the highest IoU value is selected; for the redundant prediction boxes, by matching them with the empty category, the influence on the results of other correct matches can be avoided. During the training process, adjusting the first loss function based on the bipartite graph matching result can minimize the difference between the prediction box and the real target, which helps to improve the detection accuracy of the model for small target objects.
[0072] Specifically, the first loss function is:
[0073] ; .
[0074] Among them, L iou is the intersection over union IoU; is a hyperparameter for controlling the IoU loss weight; is the L1 norm loss; It is a hyperparameter for controlling the L1 norm loss weight and can be set to 0.1; is an indicator function that returns 1 when c i is not empty and returns 0 when c i is empty; is the predicted probability distribution, where is a permutation function used to match the class labels in the defect detection results with the corresponding defect labels in the training images; is the bounding box loss function used to measure the predicted box and the true bounding box b i the difference between.
[0075] S104: Calculate the sum of the first loss function and the multi-layer perception loss term to obtain the loss function.
[0076] Among them, the multi-layer perception loss term includes the sum of the weighted perception losses corresponding to the multi-layer activation functions in the infrared image detection model.
[0077] Specifically, the multi-layer perception loss term L percep is:
[0078] .
[0079] Among them, represents the perception loss corresponding to the first layer activation function in the infrared image detection model, and its weight is 1 / 8; represents the perception loss corresponding to the second layer activation function in the infrared image detection model, and its weight is 1 / 4; represents the perception loss corresponding to the third layer activation function in the infrared image detection model, and its weight is 1 / 2; represents the perception loss corresponding to the fourth layer activation function in the infrared image detection model, and its weight is 1.
[0080] The loss function L total is: .
[0081] Thus, by introducing the multi-layer perception loss term into the loss function, the feature extraction ability and detail restoration degree of the infrared image detection model are effectively improved. Through the way of layer-by-layer weighting, it can ensure that the influence of the perception losses at different levels on the final result is reasonably balanced, so as to enhance the ability to capture low-level details while retaining high-level information. This fusion strategy of multi-level perception loss can better constrain between small targets and predicted targets, and achieve more accurate detection of small targets in infrared images.
[0082] S105: Iteratively train the infrared image detection model based on the loss function.
[0083] Specifically, after obtaining the loss function, the internal parameters of the infrared image detection model can be adjusted based on the loss function to complete the training of the current iteration. If the current iteration is greater than or equal to the total number of iterations or the accuracy is less than or equal to the preset accuracy, the training can be ended. If the current iteration is less than the total number of iterations and the accuracy is greater than the preset accuracy, the next iteration can be entered, and steps S101 to S104 are repeated to obtain the training images for the next iteration. Then, after calculating the loss function for the next iteration, the infrared image detection model is trained based on the loss function for the next iteration until the training end condition is reached (reaching the total number of iterations or reaching the preset accuracy).
[0084] As an example, the Adam optimizer can be used to iteratively train the infrared image detection model based on the loss function, and the learning rate can be set to 1×10 -4 .
[0085] Thus, in the embodiment of the present application, the first loss function and the multi-layer perception loss term obtained based on the bipartite graph matching result are used to train the infrared image detection model with a composite loss function. The trained infrared image detection model can more accurately detect small targets in infrared images. Among them, the first loss function adjusted based on the bipartite graph matching result can minimize the difference between the predicted bounding box and the real target, which helps to improve the detection accuracy of the infrared image detection model for small target objects; and through the multi-layer perception loss term, the feature extraction ability and detail restoration degree of the infrared image detection model can be effectively improved. By the way of layer-by-layer weighting, it ensures that the influence of the perception losses at different levels on the final result is reasonably balanced, so that while retaining the high-level information, the ability to capture low-level details is enhanced, and better constraints can be imposed between the small real target and the predicted target.
[0086] See Figure 5 , which is an architecture diagram of an infrared image detection model provided by an embodiment of the present application. The infrared image detection model includes: Convolutional Neural Networks (CNN), Spectral Domain Transformer Encoder, Multiscale Transformer Decoder, and Feedforward Neural Network (FNN).
[0087] Among them, a convolutional neural network is used to extract feature tokens of training images; a spectral domain Transformer encoder is used to encode the feature tokens to obtain encoded features; a multi-scale Transformer decoder is used to obtain pixel-by-pixel defect prediction results based on the input encoded features; and a fully connected layer is used to output the defect detection results of the training images based on the defect prediction results.
[0088] Specifically, the training images input to the infrared image detection model can be first resized so that the training images are adjusted to a preset size, such as 640×512 pixels; then, feature tokens of the adjusted images are extracted through a Restrictive CNN, and a feature map is output; the feature vectors in the feature map are flattened and added with position embeddings to retain spatial information; subsequently, the position-embedded feature vectors (Embedded Tokens) are sent to a TransformerEncoder for encoding; after the output of the TransformerEncoder is rearranged into a two-dimensional matrix form (Unflatten), it enters a Spectral Domain Processing module, and the processed feature map is sent to a Transformer Decoder, and the decoder outputs pixel-by-pixel defect prediction results, including prediction boxes and class labels; finally, final processing is performed through a fully connected layer (FNN) to obtain defect detection results. In the defect detection results, each prediction box corresponds to a class label, and the class labels can include but are not limited to labels such as classbox and no object.
[0089] Assume that X represents the photovoltaic infrared image after distortion correction. represents the output image with k classified defect detection boxes. The detection model aims to generate the output image with k classified defect detection boxes , and be as close as possible to the true defect classification image Y. Its working process is defined as follows:
[0090] .
[0091] Where CNN(.) represents the Restrictive CNN feature extractor, specifically restricting the convolution kernel size of the convolutional layer to ensure the precision of detecting small target objects. Encoder(.) represents the spectral domain Transformer encoder, Decoder(.) represents the multi-scale Transformer decoder, and FFN(.) represents the classifier as a standard fully connected layer. Represents the feature map extracted by the encoder. Represents the features captured by the feature extractor.
[0092] Among them, the spectral domain Transformer encoder includes a multi-head attention mechanism (Multi-Head Attention), a normalization layer (Norm), and a fully connected layer (MLP).
[0093] See Figure 6 , this figure is a structural diagram of a spectral domain Transformer encoder provided by an embodiment of the present application. The feature map input to the spectral domain Transformer encoder is first encoded through an encoder module; then, the encoded feature map is sent to a series of linear transformation modules and spectral domain transformation modules to perform linear Transformer and spectral Transformer respectively. Through Fourier operations, the input is modeled, the input image is mapped from the spatial domain to the spectral domain, the associations between different positions in the image are established, and the global dependencies between different feature tokens are captured; the data after linear Transformer and spectral Transformer are sent to the decoder for final feature reconstruction and output generation, so as to realize the conversion of the input image into a high-dimensional feature representation for target recognition, positioning, and classification through these feature representations.
[0094] In the embodiment of the present application, the core calculation formula for the spectral domain Transformer encoder to perform spectral domain transformation is: .
[0095] Among them, v represents the data input, F represents the Fourier operation, F -1 represents the inverse Fourier operation, K represents the kernel function, and · represents matrix multiplication.
[0096] The spectral domain Transformer encoder provided by the embodiment of the present application adopts the general form of the attention mechanism for exhaustion. Except for the details of position encoding, this mechanism follows the traditional Transformer. The multi-head attention is the concatenation of M single heads and then a linear projection (L), and the specific mathematical representation is:
[0097] ;
[0098] Among them, attn represents the attention mechanism; T represents the attention head; represents the output of the M-th attention head; the [] operation represents concatenation in the channel dimension.
[0099] Optionally, residual connections can be used to help the spectral domain Transformer encoder better learn long-term dependencies; dropout is used to prevent overfitting; layer normalization is used to ensure that the inputs of each layer have similar distributions to stabilize and accelerate the training process. Thus, the final output can be:
[0100] .
[0101] Among them, ; represents the final output, layernorm refers to the layer normalization operation; dropout represents the dropout operation; L is the linear projection matrix.
[0102] See Figure 7 , which is a structural diagram of a multi-scale Transformer decoder provided in an embodiment of this application. The multi-scale Transformer decoder is an architecture that enhances the decoder ability of the Transformer model, including a multi-scale feature extraction module and a multi-layer regression prediction module, which can process information at different scales, better capture multi-level information in the input image, and can effectively identify and locate small targets in the image.
[0103] Specifically, in the multi-scale feature extraction module, the Backbone is a pre-trained convolutional neural network (CNN) used to extract the basic features of an image, which can include convolutional neural networks such as ResNet or VGG; the feature map extracted by the Backbone is divided into multiple small patches, each patch is flattened and converted into a vector form to extract multi-scale features; a global token is added as an additional token to the segmented patch token sequence, and the patch after adding the global token is converted into a higher-dimensional representation Embedded Patches through an embedding layer, and then fused with the local features extracted by the Backbone to obtain fused features, which are output to the multi-layer regression prediction module. Among them, Transformer Layer 1 to Transformer Layer N represent multiple Transformer layers connected in series, and each Transformer layer contains a multi-head attention mechanism, a normalization layer, and a multi-layer perceptron (MLP). The multi-scale feature extraction module extracts multi-scale features from the encoded features provided by the spectral domain Transformer encoder based on the multi-head attention mechanism, layer normalization operation, and fully connected operation.
[0104] The process of multi-scale feature extraction is defined as follows:
[0105] ; ; .
[0106] Among them, MSA represents the multi-head attention mechanism, LN represents the layer normalization operation; MLP represents the fully connected operation; represents the fusion of different features, and the Fusion operation is a convolutional layer with a kernel size of 1.
[0107] In the multi-layer regression prediction module, different levels of features in the fused features are classified through Level Classify, and Level 1, Level 2,..., Level M are feature representations of multiple levels; finally, the different levels of features are integrated by weighted summation (WeightSum) to generate the final prediction result.
[0108] The calculation process of multi-level regression prediction is defined as follows:
[0109] .
[0110] On the basis of obtaining the fused features through multi-scale feature extraction, first, perform softmax weight scaling through Equation and then, perform feature mapping through Equation . Finally, perform weighted summation through Equation to obtain the final feature representation, that is, the defect prediction result. This feature representation is decoded by the fully connected layer into an infrared image at the pixel level.
[0111] See Figure 8 . This figure is a schematic diagram of a multi-level encoding and decoding connection mechanism provided by an embodiment of the present application. Taking the depths of the spectral domain Transformer encoder and the multi-scale Transformer decoder both set to 6 as an example, the spectral domain Transformer encoder and the multi-scale Transformer decoder are set to a symmetric structure. Among them, for the last layer of the spectral domain Transformer encoder, the encoded features are respectively input into the decoding modules at 6 different levels (depths) in the multi-scale Transformer decoder, and a channel splicing operation is performed. Thereby, the problem of important feature loss existing in the spatial features during the encoding and decoding process can be reduced, the spatial features in the image can be more accurately retained, and thus it is more convenient to more accurately identify numerous small targets in the infrared image.
[0112] Optionally, for the infrared image detection model provided by an embodiment of the present application, the number of input feature channels is 256.
[0113] Optionally, in the infrared image detection model provided by an embodiment of the present application, the depth of the Fourier module of the spectral domain Transformer encoder is 4, and the number of layers of the multi-scale Transformer decoder is 4, representing four different levels of spatial resolution of 1, 1 / 2, 1 / 4, and 1 / 8 respectively.
[0114] See Figure 9 . This figure is a schematic diagram of a training device for an infrared image detection model provided by an embodiment of the present application. The device includes: an acquisition module 901, an output module 902, a first calculation module 903, a second calculation module 904, and a training module 905;
[0115] The acquisition module 901 is used to acquire training pictures containing infrared images;
[0116] The output module 902 is used to output defect detection results based on the training pictures through a pre-constructed infrared image detection model; the number of prediction boxes generated by the infrared image detection model is greater than the number of real targets; the defect detection results include the real category to which the real target belongs and an empty category representing the absence of a real target;
[0117] The first calculation module 903 is configured to calculate a first loss function based on the bipartite graph matching result between the defect detection result and the defect label corresponding to the training image.
[0118] The second calculation module 904 is configured to calculate the sum of the first loss function and the multi-layer perception loss term to obtain a loss function; the multi-layer perception loss term includes the sum of the weighted perception losses corresponding to the multi-layer activation functions in the infrared image detection model.
[0119] The training module 905 is configured to perform iterative training on the infrared image detection model based on the loss function.
[0120] Thus, in the embodiments of the present application, based on the first loss function and the multi-layer perception loss term obtained from the bipartite graph matching result, a composite loss function is used to train the infrared image detection model. The trained infrared image detection model can more accurately detect small targets in infrared images. Among them, the first loss function adjusted based on the bipartite graph matching result can minimize the difference between the prediction box and the real target, which helps to improve the detection accuracy of the infrared image detection model for small target objects; and through the multi-layer perception loss term, the feature extraction ability and detail restoration degree of the infrared image detection model can be effectively improved. By the way of layer-by-layer weighting, it ensures that the influence of the perception losses at different levels on the final result is reasonably balanced, so that while retaining the high-level information, the ability to capture low-level details is enhanced, and better constraints can be made between the small real target and the prediction target.
[0121] Optionally, another training device for an infrared image detection model provided in the embodiments of the present application further includes a preprocessing module, configured to obtain an initial image including an infrared image; the initial image is collected by an infrared camera; based on the distortion coefficient of the infrared camera and the camera internal parameter matrix, coordinate transformation is performed on the initial image to obtain a training image.
[0122] In addition, the present application also provides an infrared image detection method, which includes:
[0123] S1: Obtain a to-be-detected image including an infrared image.
[0124] Optionally, the to-be-detected image may be obtained by performing distortion correction on an initial image collected by an infrared camera.
[0125] S2: Based on the to-be-detected image, output a defect detection result through the infrared image detection model trained by the training method of the infrared image detection model described in any of the above embodiments.
[0126] It should be noted that the various embodiments in this specification are described in a progressive manner. For the same or similar parts among the embodiments, reference can be made to each other, and the key points of each embodiment are the differences from other embodiments. In particular, for the device embodiments, since they are basically similar to the method embodiments, the description is relatively simple, and for the relevant parts, reference can be made to the description of the method embodiments. The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, and the components referred to as units may or may not be physical units, that is, they may be located in one place or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment. Those of ordinary skill in the art can understand and implement it without creative efforts.
[0127] As described above, it is only a specific implementation manner of this application, but the protection scope of this application is not limited thereto. Any changes or substitutions that can be easily thought of by those skilled in the art within the technical scope disclosed in this application should be covered by the protection scope of this application. Therefore, the protection scope of this application should be subject to the protection scope of the claims.
Claims
1. A training method for an infrared image detection model, characterized in that The method includes: Obtaining training pictures containing infrared images; Based on the training pictures, outputting a defect detection result through a pre-constructed infrared image detection model; the number of prediction boxes generated by the infrared image detection model is greater than the number of real targets; the class labels in the defect detection result are selected from the real classes to which the real targets belong and the empty class representing no real targets; Calculating a first loss function based on the bipartite graph matching result between the defect detection result and the defect label corresponding to the training picture; Calculating the sum of the first loss function and the multi-layer perception loss term to obtain a loss function; the multi-layer perception loss term includes the sum of the weighted perception losses corresponding to the multi-layer activation functions in the infrared image detection model; Iteratively training the infrared image detection model based on the loss function; The infrared image detection model includes a convolutional neural network, a spectral domain Transformer encoder, a multi-scale Transformer decoder, and a fully connected layer; the depths of both the spectral domain Transformer encoder and the multi-scale Transformer decoder are 6; the spectral domain Transformer encoder and the multi-scale Transformer decoder are symmetric structures; the last layer of the spectral domain Transformer encoder inputs the encoded features into the decoding modules at each depth in the multi-scale Transformer decoder; The spectral domain Transformer encoder based on the self-attention mechanism is used to encode the feature map, model the encoded feature map through Fourier operations to map the encoded feature map from the spatial domain to the spectral domain, and unify the representations of the spatial domain and frequency domain features through a residual cascade architecture to obtain encoded features; the encoded features include the global dependencies between different feature tokens; The multi-scale Transformer decoder includes a multi-scale feature extraction module and a multi-layer regression prediction module; the multi-scale Transformer decoder is specifically used for: Through the multi-scale feature extraction module, extracting a feature map from the encoded features through a pre-trained convolutional neural network; based on the multi-head attention mechanism, layer normalization operation, and fully connected operation, extracting multi-scale features based on the feature map extracted by the convolutional neural network; fusing the multi-scale features through a convolutional layer with a kernel size of 1 to obtain fused features; Through the multi-layer regression prediction module, obtaining a defect prediction result based on the fused features.
2. The method according to claim 1, characterized in that, The infrared image detection model includes a convolutional neural network, a spectral domain Transformer encoder, a multi-scale Transformer decoder, and a fully connected layer; The convolutional neural network is used to extract feature tokens of the training picture to obtain a feature map; The spectral domain Transformer encoder is used to encode the feature map to obtain encoded features; The multi-scale Transformer decoder is used to obtain a pixel-by-pixel defect prediction result based on the input encoded features; The fully connected layer is used to output the defect detection result of the training image based on the defect prediction result.
3. The method according to claim 2, characterized in that The number of layers of the multi-scale Transformer decoder is 4, and the corresponding spatial resolution levels of different layers are 1, 1 / 2, 1 / 4, and 1 / 8 respectively.
4. The method according to claim 1, wherein Before obtaining the training image containing the infrared image, the method further includes: Obtaining an initial image containing an infrared image; the initial image is collected by an infrared camera; Based on the distortion coefficient of the infrared camera and the camera internal parameter matrix, performing coordinate transformation on the initial image to obtain a training image.
5. The method according to claim 1, wherein The number of prediction boxes generated by the infrared image detection model is 50.
6. A training device for an infrared image detection model, characterized in that, The device includes: an acquisition module, an output module, a first calculation module, a second calculation module, and a training module; The acquisition module is used to acquire a training image containing an infrared image; The output module is used to output a defect detection result based on the training image through a pre-constructed infrared image detection model; the number of prediction boxes generated by the infrared image detection model is greater than the number of real targets; the defect detection result includes the real category to which the real target belongs and an empty category representing no real target; The first calculation module is used to calculate a first loss function based on the bipartite graph matching result between the defect detection result and the defect label corresponding to the training image; The second calculation module is used to calculate the sum of the first loss function and a multi-layer perception loss term to obtain a loss function; the multi-layer perception loss term includes the sum of the weighted perception losses corresponding to the multi-layer activation functions in the infrared image detection model; The training module is used to perform iterative training on the infrared image detection model based on the loss function; The infrared image detection model includes a convolutional neural network, a spectral domain Transformer encoder, a multi-scale Transformer decoder, and a fully connected layer; the depths of both the spectral domain Transformer encoder and the multi-scale Transformer decoder are 6; the spectral domain Transformer encoder and the multi-scale Transformer decoder are symmetric structures; the last layer of the spectral domain Transformer encoder inputs the encoded features into the decoding modules at each depth in the multi-scale Transformer decoder; The spectral domain Transformer encoder based on the self-attention mechanism is used to encode the feature map, model the encoded feature map through Fourier operations to map the encoded feature map from the spatial domain to the spectral domain, and unify the representations of the spatial and frequency domain features through a residual cascade architecture to obtain encoded features; the encoded features include the global dependencies between different feature tokens; The multi-scale Transformer decoder includes a multi-scale feature extraction module and a multi-layer regression prediction module; specifically, the multi-scale Transformer decoder is used for: Through the multi-scale feature extraction module, a feature map is extracted from the encoded features by a pre-trained convolutional neural network; based on the multi-head attention mechanism, layer normalization operation, and fully connected operation, multi-scale features are extracted based on the feature map extracted by the convolutional neural network; the multi-scale features are fused through a convolutional layer with a kernel size of 1 to obtain fused features; Through the multi-layer regression prediction module, a defect prediction result is obtained based on the fused features.
7. An infrared image detection method, characterized in that, The method includes: Obtaining a to-be-inspected picture including an infrared image; Based on the to-be-inspected picture, an infrared image detection model trained by the training method of the infrared image detection model according to any one of claims 1 to 5 is used to output a defect detection result.
Citation Information
Patent Citations
Aero-engine blade defect detection method based on multi-scale DETR
CN117173449A
Vehicle target detection method based on novel frequency domain encoder
CN118570611A