Infrared image target detection method based on YOLOv8 improvement

By introducing coordinate attention module and channel knowledge distillation technology into the yolov8 object detection algorithm, the infrared image object detection model is optimized to achieve high-precision detection on edge devices, solving the problem of insufficient performance of infrared image object detection on edge devices.

CN120032105APending Publication Date: 2025-05-23SHANGHAI UNIV
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510082444.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-17
Publication Date
2025-05-23

AI Technical Summary

Technical Problem

When existing infrared image object detection algorithms run on edge devices, due to computing power and memory, it is difficult to achieve high-precision detection, especially when the infrared image resolution is not high.

Method used

By applying knowledge distillation technology, the yolov8 object detection algorithm is optimized, the coordinate attention module is introduced, and the channel knowledge distillation method is used to train the student network model, so that the model improves the performance of the infrared object detection model without increasing the amount of parameters and calculations.

Benefits of technology

It realizes the accuracy and performance of infrared image object detection on edge devices, so that the model can more efficiently identify targets in infrared images, and is suitable for detection tasks of low-resolution infrared images.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120032105A_ABST
    Figure CN120032105A_ABST
Patent Text Reader

Abstract

The invention relates to an infrared image target detection method based on yov8 improvement. The infrared image target detection method comprises the steps of loading an infrared image target detection data set, zooming resolution, standardizing a labeling file format, converting the labeling file format into a yolo training format, and dividing a training set and a test set; constructing a deep neural network model based on the yolov8, and introducing a coordinate attention module into a backbone network to obtain an improved yolov8 target detection neural network; training a teacher model: training the teacher model by adopting a yolov8-s network, storing the model, and training a yolov8-x teacher model as a comparison group; training a student model: adopting a minimum-scale network yolov8-n as the student model, extracting the score features of the teacher model from the teacher model to the student model through knowledge distillation, then training, and storing the model; and finally, identifying the infrared image by using a student model obtained by training the improved yov8 network, and obtaining a detection result. A coordinate attention module is introduced into a backbone network, and a target detection model is obtained through knowledge distillation training.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of target detection in computer vision, and in particular to an infrared image target detection method based on an improvement of YOLOv8. Background Art

[0002] Object detection is a classic task in computer vision, which aims to locate the position of objects in an image and identify the specific category of the objects. It is widely used in intelligent driving, transportation, military, fault detection and other fields, and is an important research direction in the field of computer vision.

[0003] The premise of target detection is to use an optical camera to obtain scene information from an optical sensor. However, the imaging performance of the optical camera will drop sharply at night and in natural environments with snow, fog, smoke or dust, resulting in a sharp drop in detection accuracy. In the above environment, the imaging effect of the infrared sensor is more stable and more suitable for complex and changeable harsh imaging environments. With the continuous development of science and technology, infrared target detection technology is increasingly widely used in military, aviation, medical and other fields. Infrared target detection is a method of detecting targets in infrared images using infrared thermal imaging technology. However, the infrared focal plane array of civilian thermal imaging cameras has a low resolution, resulting in low resolution and lack of details in infrared images, which is a test for the accuracy of target detection algorithms. In addition, in actual applications, in order to reduce costs, edge devices can only be equipped with low-performance processors, resulting in computing power and memory that can only support compressed small-scale target detection algorithms. Therefore, it is very important to achieve high-precision detection of infrared images.

[0004] In recent years, deep neural networks have developed rapidly, and various deep learning methods have emerged. Deep neural networks have made great progress in the performance of target detection tasks, and have basically replaced traditional target detection algorithms. The more famous algorithms include the R-CNN series of two-stage target detection algorithms, and the YOLO (You Only Look Once) series of networks and SSD (Single Shot Multi-box Detector), which are representatives of single-stage networks. They have achieved performance that exceeds traditional methods in target detection tasks.

[0005] With the development of technology and the upgrading of demand, target detection algorithms need to be separated from cloud server scenarios and run in edge devices such as self-driving cars, drones, and robot dogs. Edge devices do not have sufficient computing power and storage space on cloud servers, so the selection and optimization of algorithms are particularly important for edge devices. The YOLO series of algorithms use a single-stage neural network to complete target detection, so the detection speed is fast, the memory usage is small, and it is suitable for edge devices. In addition, in order to help neural network models be deployed on edge devices, researchers have proposed model acceleration and compression technologies, including pruning, quantization, lightweight model design, and knowledge distillation.

[0006] Knowledge distillation, also known as teacher-student learning, transfers the knowledge of a large and complex teacher model to a smaller and more concise student model. During the training process, the student model can learn the soft targets output by the teacher model detection head, that is, the predicted probability distribution of the teacher model, and the activation values ​​of the intermediate layer features of the teacher model. Both methods can help the student model reduce computing resources and storage requirements while maintaining high accuracy. Due to the simplicity and effectiveness of knowledge distillation, it is widely used in model compression and model accuracy improvement.

[0007] In response to the problems faced when deep neural networks are deployed on edge devices, this paper optimizes the yolov8 target detection algorithm by applying knowledge distillation technology. By allowing a smaller and more computationally lightweight student model to learn the knowledge in a pre-trained complex teacher model, the student model can be helped to search for a distribution closer to real data, thereby reducing the number of model parameters while improving its detection accuracy, and ultimately achieving performance improvement of the infrared target detection algorithm on edge devices. Summary of the invention

[0008] In view of the existing problems and deficiencies of the prior art, the present invention proposes an infrared image target detection method based on the improved yolov8. The method adopted by the present invention can improve the accuracy of infrared image target detection without introducing additional parameters and calculations, and ultimately improve the performance of the infrared target detection model on the edge device.

[0009] The method of the present invention first converts the existing infrared image data set into a resolution of 640×640, converts the annotation format into a format supported by yolov8, and divides the training set and the test set. A deep neural network model based on yolov8 is constructed, including a backbone network (feature extraction module), a neck (feature fusion module) and a detection head (feature detection module), and a coordinate attention module (CA) is introduced into the backbone network to obtain an improved yolov8 target detection neural network. According to the model structure, the hyperparameters are set, the yolov8-s (small) network is used to train the teacher model Teacher, and the yolov8-x (super large) teacher model is trained as a comparison group, and the model is saved. The model distillation hyperparameters are set, and the yolov8-n network is used as the student model Student. The logit (score) features of the teacher model are distilled to the student model Student for training, and the student model network weights are updated by the back propagation algorithm until the maximum number of iterations set is reached, and the student network model is saved. Finally, the student model Student obtained by training the improved yolov8 network is used to recognize the infrared image to obtain the detection result of the infrared image. The parameter comparison of the teacher model Teacher and the student model Student is shown in Table 1:

[0010]

[0011]

[0012] Table 1 Comparison of parameters of yolov8 teacher and student models

[0013] In the above method, preparing the infrared image training data set includes: loading the existing Flir ThermalDataset infrared image data set, adjusting it to a 640×640 resolution image through letter box resize, converting the annotation file into yolov8 format, and dividing it into a training set and a test set.

[0014] In the above method, a deep neural network model based on yolov8 is constructed: including the backbone network, neck and detection head, and the coordinate attention module is introduced into the backbone network to obtain an improved yolov8 target detection neural network. The module structure diagram is as follows Figure 2As shown in the figure. The backbone network adopts the CSP (Cross Stage Partial) structure, including 5 CBS (Convolution + Batch Normalization + SiLU) modules, 4 C2f (Cross Stage Partial and 2-Fold Aggregation) modules and an SPPF (Spatial Pyramid Pooling Fast) module; the backbone network mainly plays the role of feature extraction. Through the CBS and C2F modules, the first feature p 1 , the second feature p 2 and the third feature p 3 are obtained. Coordinate attention is added to p 1 , p 2 and p 3 respectively; the attention mechanism has been widely studied and deployed to improve the performance of deep neural networks. The commonly used attention mechanism is the attention based on compression and excitation (Squeeze and Excitation Attention, SE), which calculates channel attention with the help of 2D global pooling and provides a significant performance improvement at a low computational cost. However, the SE attention only considers the information between encoded channels and ignores the location information, which is crucial for identifying target objects in visual tasks. Different from the channel attention that transforms the input into a single feature vector through 2D global pooling, CA uses two one-dimensional global pooling operations to aggregate the input features in the vertical and horizontal directions into two independent direction-aware feature maps respectively. These two feature maps embedded with direction-specific information are then encoded into two attention maps respectively, each of which captures the long-range dependencies of the input feature map along one spatial direction, and thus the location information can be retained in the generated attention map, ultimately enhancing the performance of the model in identifying target features. Given the input (where c, h, and w are the number of channels, height, and width of the feature map respectively), first use a pooling kernel of size (H, 1) to encode each channel along the horizontal coordinate, and use a pooling kernel of size (1, W) to encode each channel along the vertical coordinate. The output of the c-th channel with height h can be expressed as:

[0015]

[0016] The output of the c-th channel with width w can be expressed as:

[0017]

[0018] The feature maps generated above are concatenated and then transformed using the convolution transformation function, expressed as:

[0019] f=δ(F 1 (Concat(z h ,z w )))

[0020] Among them, Concat represents the concatenation operation along the column direction, F 1 represents a 1×1 convolution operation, and δ represents a nonlinear activation function. f is an intermediate feature map with a dimension of r is the scaling factor, which is set to 32 during training. Next, f is split into two separate tensors along the column direction: Using two more 1×1 convolution transformations F h and F w f h and f w Transformed into a tensor g with the same number of channels h and g w :

[0021] g h =σ(F h (f h )),

[0022] g w =σ(F w (f w ))

[0023] Finally, the output g h and g w Expand them separately and use them as attention weights. The output of the coordinate attention block Y can be written as:

[0024]

[0025] From the above steps, we can get the feature p with coordinate attention 1 ca 、p 2 ca and p 3 ca .

[0026] The neck network consists of a path aggregation-feature pyramid network (PAN-FPN) module, which plays a role in further extracting and fusing the features of the backbone network to improve accuracy and robustness. The whole process can be expressed as follows:

[0027] t 1 ,t 2 ,t 3 =F fusion (p 1 ca ,p 1 ca ,p 1 ca )

[0028] F fusion Represents the feature fusion operation, the neck network outputs the first fusion feature t 1 , the second fusion feature t 2 and the third fusion feature t 3 ; Through the detection head network, t 1 ,t 2 and t 3 Classification and positioning are performed to obtain a first target detection result, a second target detection result, and a third target detection result respectively. The first target detection result, the second target detection result, and the third target detection result are formed into a detection result set to obtain a target detection result.

[0029] In the above method, the teacher model is trained by constructing an improved deep neural network model, using the yolov8-s network to train the teacher model Teacher, and training the yolov8-x teacher model as a comparison group. Determine the parameters such as the optimizer, learning rate, and maximum number of iterations, and start training the teacher network; after each forward propagation, calculate the loss between the training result and the true label (Ground Truth), and then update the network parameters through the back propagation algorithm; repeat the above steps until the preset maximum number of iterations is reached, thereby completing the teacher model Teacher training.

[0030] In the above method, the student model is trained: the module structure diagram is as follows Figure 3As shown in the figure, compared with the constructed neural network structure, the minimum-scale network yolov8-n is used as the student model Student, and the yolov8-s and yolov8-x trained in the previous step are used as the teacher model Teacher. Before the training starts, the teacher model is passed in, and during the training process, the features of the teacher model are distilled into the student model through the channel distillation (Channel-wise Knowledge) strategy. The process is as follows: During the forward propagation process, the activation maps corresponding to the detection head modules of the student model and the teacher model are softly aligned, and the activation maps are converted into probability distributions through the Softmax function. In this case, the channel distillation loss function can be written as:

[0031]

[0032] φ(·) is used to convert the activation value into a probability distribution, that is, normalization (softmax normalization), which is expressed as follows:

[0033]

[0034] Indicates the channel corresponding to the activation map of the teacher model and the student model, c = 1, 2, ... C represents the channel index. i represents the spatial position index of a feature map channel, represents the temperature hyperparameter, with a value of 2.5. Normalization can eliminate the influence of different network scales and complexities, which is effective for knowledge distillation. If the number of channels of the student network and the teacher network do not match, a 1×1 convolution is used to upsample the number of student network channels. To evaluate the difference between the channel distributions of the teacher network and the student network, the KL divergence is used:

[0035]

[0036] KL divergence is not symmetric. As can be seen from the above formula, when When it is very big, It should also be as large as the teacher network to minimize the KL divergence. When it is smaller, the KL divergence value will also be smaller, and there will be no Too much emphasis. Therefore, the student network can learn the distribution of the foreground salient area better, but less about the background area. The infrared target image has a low resolution, so it is more expected that the model can focus on the pixel values ​​of the foreground salient area. This asymmetry can help improve the performance of infrared image target detection. After calculating the distillation loss, update the network parameters through the back propagation algorithm; repeat the above steps until the preset maximum number of iterations is reached to complete the student model Student training.

[0037] In the above method, save the student model: solidify the set of network weights with the highest evaluation index during the training process and save it as the final student network model. represents the student model trained by yolov8-x distillation, It represents the student model trained by yolov8-s distillation. Through evaluation and comparison, the best one is selected as the final deployable model.

[0038] Compared with the prior art, the method of the present invention has the following obvious outstanding substantive features and significant technical progress:

[0039] 1) The high-resolution infrared image is adjusted to 640×640 by letterbox scaling, which adapts to the model training while maintaining the original aspect ratio of the image and avoiding image distortion.

[0040] 2) Introduce the coordinate attention mechanism to save the location information into the feature map and enhance the model's performance in identifying target features.

[0041] 3) The channel knowledge distillation method is used to train the student network model, so that the model pays more attention to the foreground salient areas, and improves the performance of the infrared target detection model on edge devices without increasing the number of parameters and computational complexity. BRIEF DESCRIPTION OF THE DRAWINGS

[0042] Figure 1 A flowchart of the method of the present invention;

[0043] Figure 2 The neural network structure diagram of the coordinate attention introduced by the method of the present invention;

[0044] Figure 3 Flow chart of channel distillation of the method of the present invention. DETAILED DESCRIPTION

[0045] The preferred embodiments of the present invention are described in detail as follows in conjunction with the accompanying drawings:

[0046] See also Figure 1 The present invention is based on the improved infrared image target detection method of yolov8, and the specific operation steps are as follows:

[0047] 1) Prepare training data set: Convert the existing Flir Thermal Dataset high-resolution infrared image data set to a resolution of 640×640, convert the annotation format to a format supported by yolov8, and divide it into training set and test set.

[0048] 2) Build a deep neural network model based on yolov8: The network model structure diagram is as follows Figure 2 As shown in Figure 1, it includes a backbone network, a neck and a detection head, and introduces a coordinate attention module into the backbone network to obtain an improved yolov8 target detection neural network. The module structure diagram is shown in Figure 1. Figure 2 The backbone network adopts the CSP structure, including 5 CBS modules, 4 C2f modules and SPPF modules; the backbone network mainly plays the role of feature extraction. Through the CBS and C2F modules, the first feature p is obtained. 1 , the second feature p 2 and the third feature p 3 . 1 、p 2 and p 3 Add coordinate attention separately. Attention mechanisms have been widely studied and deployed to improve the performance of deep neural networks. The commonly used attention mechanism is based on Squeeze and Excitation Attention (SE), which calculates channel attention with the help of 2D global pooling and provides significant performance improvement at a lower computational cost. However, SE attention only considers encoding inter-channel information and ignores position information, which is crucial for identifying target objects in visual tasks. Unlike channel attention, which transforms the input into a single feature vector through two-dimensional global pooling, CA utilizes two one-dimensional global pooling operations to aggregate the input features in the vertical and horizontal directions into two independent direction-aware feature maps, respectively. These two feature maps embedded with direction-specific information are then encoded into two attention maps, each of which captures the long-range dependencies of the input feature maps along one spatial direction. The position information can therefore be retained in the generated attention map, ultimately enhancing the performance of the model in identifying target features. Given an input First, each channel is encoded along the horizontal coordinate using a pooling kernel of size (H, 1), and each channel is encoded along the vertical coordinate using a pooling kernel of size (1, W). The output of the cth channel with height h can be expressed as:

[0049]

[0050] The output of the cth channel with width w can be expressed as:

[0051]

[0052] The feature maps generated above are concatenated and then transformed using the convolution transformation function, expressed as:

[0053] f=δ(F 1 (Concat(z h ,z w )))

[0054] Among them, Concat represents the concatenation operation along the column direction, F 1 represents a 1×1 convolution operation, and δ represents a nonlinear activation function. f is an intermediate feature map with a dimension of r is the scaling factor, which is set to 32 during training. Next, f is split into two separate tensors along the column direction: Using two more 1×1 convolution transformations F h and F w f h and f w Transformed into a tensor g with the same number of channels h and g w :

[0055] g h =σ(F h (f h )),

[0056] g w =σ(F w (f w ))

[0057] Finally, the output g h and g w Expand them separately and use them as attention weights. The output of the coordinate attention block Y can be written as:

[0058]

[0059] From the above steps, we can get the feature p with coordinate attention 1 ca 、p 2 ca and p 3 ca .

[0060] The neck network consists of a path aggregation network-feature pyramid network (PAN-FPN) module, which plays a role in further extracting and fusing the features of the backbone network to improve accuracy and robustness. The whole process can be expressed as follows:

[0061] t 1 ,t 2 ,t 3 =F fusion (p 1 ca ,p 1 ca ,p 1 ca )

[0062] F fusion Represents the feature fusion operation, the neck network outputs the first fusion feature t 1 , the second fusion feature t 2 and the third fusion feature t 3 ; Through the detection head network, t 1 ,t 2 and t 3 Classification and positioning are performed to obtain a first target detection result, a second target detection result, and a third target detection result respectively. The first target detection result, the second target detection result, and the third target detection result are formed into a detection result set to obtain a target detection result.

[0063] 4) Train the teacher model: Build an improved deep neural network model, use the yolov8-s network to train the teacher model Teacher, and train the yolov8-x teacher model as a comparison group. Determine the optimizer, learning rate, maximum number of iterations and other parameters, and start training the teacher network; after each forward propagation, calculate the loss between the training result and the true label (Ground Truth), and then update the network parameters through the back propagation algorithm; repeat the above steps until the preset maximum number of iterations is reached to obtain the teacher model Teacher. x and Teacher s .

[0064] 5) Train the student model: Figure 3 As shown in the figure, compared with the constructed neural network structure, the minimum parameter network yolov8-n is used as the student model Student, and the Teacher trained in the previous step is used respectively. x and Teacher s As the teacher model Teacher. Before the training starts, the teacher model Teacher is passed in, and during the training process, the features of the teacher model are distilled into the student model through the channel distillation strategy. The process is as follows: During the forward propagation process, the activation maps corresponding to the detection head modules of the student model and the teacher model are soft-aligned, and the activation maps are converted into probability distributions through the Softmax function. In this case, the channel distillation loss function can be written as:

[0065]

[0066] φ(·) is used to convert the activation value into a probability distribution, which is expressed as follows:

[0067]

[0068] Indicates the channel corresponding to the activation map of the teacher model and the student model, c = 1, 2, ... C represents the channel index. i represents the spatial position index of a feature map channel, Represents the temperature hyperparameter, which is 2.5 in this example. Normalization can eliminate the influence of different network scales and complexities, which is effective for knowledge distillation. If the number of channels of the student network and the teacher network does not match, a 1×1 convolution is used to upsample the number of student network channels. To evaluate the difference between the channel distributions of the teacher network and the student network, the KL divergence is used:

[0069]

[0070] KL divergence is not symmetric. As can be seen from the above formula, when When it is very big, It should also be as large as the teacher network to minimize the KL divergence. When it is smaller, the KL divergence value will also be smaller, and there will be no Therefore, the student network can learn the distribution of the foreground salient area better, but less about the background area. The infrared target image has a low resolution, and it is more expected that the model can focus on the pixel values ​​of the foreground salient area. This asymmetry can help improve the performance of infrared image target detection. After calculating the distillation loss, update the network parameters through the back propagation algorithm; repeat the above steps until the preset maximum number of iterations is reached to complete the training of the student model Student.

[0071] 6) Save the network: solidify the set of network weights with the highest evaluation index during the training process and save it as the final student network model. represents the student model trained by yolov8-x distillation, Represents the student model trained by yolov8-s distillation. Through evaluation and comparison, the best one is selected as the final deployable model.

[0072] The evaluation results are shown in Table 2:

[0073]

[0074]

[0075] Table 2 The evaluation results are shown in Table 2 compared with Table 1. The model trained by introducing the coordinate attention mechanism and adopting the channel distillation strategy is All of them have higher accuracy than training the yolov8-n model alone. The improvement in accuracy is very significant for the smallest scale yolov8-n network, which quantitatively demonstrates the effectiveness of the present invention. The distillation results show that the teacher model with a large gap between the scale and the student model is not necessarily the most suitable as the teacher model. The yolov8-n student model trained by distillation of the yolov8-s teacher model has higher detection accuracy while maintaining the same parameter scale and computational complexity, and can be used as the final deployment model.

[0076] The preferred specific embodiments of the present invention are described in detail above. It should be understood that ordinary technicians in the field can make many modifications and changes based on the concept of the present invention without creative work. Therefore, all technical solutions that can be obtained by technicians in the technical field based on the concept of the present invention through logical analysis, reasoning or limited experiments on the basis of the prior art should be within the scope of protection determined by the claims.

Claims

1. An infrared image target detection method based on yolov8 improvement, characterized in that: The following steps are involved: S1) Convert the existing infrared image dataset into a resolution of 640×640, convert the annotation format into a format supported by yolov8, and divide it into a training set and a test set; S2) constructing a deep neural network model based on yolov8, wherein the deep neural network model includes a feature extraction module, a feature fusion module and a detection head feature detection module, and introducing a coordinate attention module into the feature extraction module; S3) using the yolov8-s network to train the teacher model, and training the yolov8-x teacher model as a comparison group, and saving the teacher model; S4) using the yolov8-n network as the student model, distilling the score features of the teacher model to the student model for training, updating the student model network weights by the back propagation algorithm, and saving the student model after reaching the set maximum number of iterations; S5) Using the student model obtained by training the improved yolov8 network to recognize the infrared image, and obtaining the detection result of the infrared image.

2. The infrared image target detection method based on yolov8 improvement according to claim 1 is characterized in that, The step S1 also includes loading the existing Flir Thermal Dataset infrared image dataset, adjusting it to a 640×640 resolution image by letterbox scaling, converting the annotation file into a yolov8 format, and dividing it into a training set and a test set.

3. The infrared image target detection method based on yolov8 improvement according to claim 1 is characterized in that, The feature extraction module in step S2 adopts a CSP structure, including 5 CBS modules, 4 C2f modules and an SPPF module; through the CBS and C2F modules, the first feature p1, the second feature p2 and the third feature p3 are obtained, and coordinate attention is added to p1, p2 and p3 respectively, and the input is converted into a single feature vector through two-dimensional global pooling, and the input features in the vertical and horizontal directions are respectively aggregated into two independent direction-aware feature maps using two one-dimensional global pooling operations; given the input First, use a pooling kernel of size (H, 1) to encode each channel along the horizontal coordinate, and use a pooling kernel of size (1, W) to encode each channel along the vertical coordinate; the output of the cth channel with height h can be expressed as: The output of the cth channel with width w can be expressed as: The feature maps generated above are concatenated and then transformed using the convolution transformation function, expressed as: f=δ(F1(Concat(z h ,z w ))) Among them, Concat represents the concatenation operation along the column direction, F1 represents the 1×1 convolution operation, δ represents the nonlinear activation function, and f is an intermediate feature map with a dimension of r is the scaling factor, which is set to 32 during training; Next, we split f into two separate tensors along the column direction: and Using two more 1×1 convolution transformations F h and F w f h and f w Transformed into a tensor g with the same number of channels h and g w : g h =σ(F h (f h )), g w =σ(F w (f w )) Finally, the output g h and g w Expanded separately and used as attention weights, the output of the coordinate attention block Y can be written as: From the above steps, we can get the feature p1 with coordinate attention ca 、p2 ca and p3 ca .

4. The infrared image target detection method based on yolov8 improvement according to claim 1 is characterized in that, The feature fusion module is composed of a path aggregation-feature pyramid network module, and the process can be expressed as follows: t1,t2,t3=F fusion (p1 ca ,p1 ca ,p1 ca ) F fusion It represents a feature fusion operation. The neck network outputs a first fusion feature t1, a second fusion feature t2 and a third fusion feature t3. The detection head network classifies and locates t1, t2 and t3 respectively, and obtains a first target detection result, a second target detection result and a third target detection result accordingly. The first target detection result, the second target detection result and the third target detection result form a detection result set to obtain a target detection result.

5. The infrared image target detection method based on yolov8 improvement according to claim 1 is characterized in that: The step S3 adopts the neural network model constructed in step S2, adopts the yolov8-s network to train the teacher model, and trains the yolov8-x teacher model as a comparison group; determines the optimizer, learning rate and maximum number of iterations, and starts training the teacher network; after each forward propagation is completed, calculates the loss between the training result and the true label, and then updates the network parameters through the back propagation algorithm; Repeat the above steps until the preset maximum number of iterations is reached to complete the teacher model training.

6. The infrared image target detection method improved based on yolov8 according to claim 5, characterized in that: The step S4 adopts the neural network structure constructed in step S2, uses the minimum scale network yolov8-n as the student model, and uses the yolov8-s and yolov8-x trained in the previous step as the teacher models respectively; before the training starts, the teacher model is passed in, and the features of the teacher model are distilled into the student model through the channel distillation strategy during the training process. In the forward propagation process, the activation maps corresponding to the detection head modules of the student model and the teacher model are soft-aligned, and the activation maps are converted into probability distributions through the Softmax function. In this case, the channel distillation loss function can be written as: φ(·) is used to convert the activation value into a probability distribution, that is, normalization, which is expressed as follows: and represents the channel corresponding to the activation map of the teacher model and the student model, c = 1, 2, ... C represents the channel index; i represents the spatial position index of a feature map channel, represents the temperature hyperparameter, with a value of 2.5; if the number of channels of the student network and the teacher network do not match, a 1×1 convolution is used to upsample the number of student network channels. To evaluate the difference between the channel distributions of the teacher network and the student network, the KL divergence is used: After calculating the distillation loss, the network parameters are updated through the back propagation algorithm; the above steps are repeated until the pre-set maximum number of iterations is reached to complete the student model training.

7. The infrared image target detection method based on yolov8 improvement according to claim 1 is characterized in that: The step S5 solidifies the set of network weights with the highest evaluation index during the training process and saves them as the final student network model. and represents the student model trained by yolov8-x distillation, Represents the student model trained by yolov8-s distillation. Through evaluation and comparison, the best one is selected as the final deployable model.

Citation Information

Cited By

  • Light-weight unmanned aerial vehicle infrared remote sensing target detection method and device based on knowledge distillation

    CN121280703A