Infrared target identification method for heterogenous image generalization
By using heterologous image and reparameterization technology, a convolution kernel model adapted to the target size was constructed, which solved the problem of poor results caused by the lack of data in infrared object detection algorithm, and achieved better generalization and detection capabilities of infrared object recognition.
Patent Information
- Application Number
- CN202311535771.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2023-11-17
- Publication Date
- 2025-05-20
AI Technical Summary
The infrared object detection algorithm based on deep learning cannot obtain ideal results due to the lack of infrared image data, and the prior art is difficult to effectively improve the generalization of the model.
Using heterologous images, a convolution kernel model adapted to the target size is constructed through deep learning and reparameterization technology, and a model for infrared target recognition is generated through weights and heavy parameterization of multiple models.
It improves the detection ability of infrared image targets, improves the generalization of heterologous images by the model, and achieves better target recognition effect.
Smart Images

Figure CN120020891A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of infrared target recognition, and more specifically, to an infrared target recognition method for heterogeneous image generalization. Background Art
[0002] Infrared imaging has the advantages of strong smoke penetration ability, high resolution, and can work all day long. Moreover, since infrared imaging is passive in receiving radiation, it has good concealment and stronger security. Therefore, algorithms for infrared target recognition have been widely studied.
[0003] With the development of deep learning in recent years, a large number of infrared target recognition algorithms based on deep learning have been proposed. Although deep learning technology has achieved remarkable results in the visible light field, due to the problems of lack of color, lack of texture in infrared images compared with visible light images, and the difficulty in obtaining infrared images and serious shortage of data volume, the infrared target detection algorithms based on deep learning cannot achieve ideal results. Since the data volume of visible light is relatively rich, the research on the generalization of heterogeneous images in deep learning has gradually attracted attention. Currently, the main means to improve the generalization performance of the model are data augmentation, participation of different data sources in training, and fine-tuning with target source data. However, infrared images are often scarce or even non-existent. In view of this situation, we propose an infrared target recognition method for heterogeneous image generalization. Summary of the Invention
[0004] Aiming at the deficiencies of the prior art, the present invention invents an infrared target recognition method for heterogeneous image generalization. This method adopts heterogeneous images and improves the detection ability of infrared image targets through deep learning and reparameterization technology in the case of lack or even non-existence of infrared image data.
[0005] The technical solution adopted by the present invention to achieve the above object is: an infrared target recognition method for heterogeneous image generalization, comprising the following steps:
[0006] Step 1, statistically analyze the target size of the data set and judge the target size;
[0007] Step 2, construct a convolutional kernel model adapted to the target size;
[0008] Step 3, complete the pre-training of the convolutional kernel model on the data set;
[0009] Step 4, adjust multiple models by different training methods on the data set;
[0010] Step 5, by means of parameter reparameterization, use the weights of multiple models to reparameterize into a single model weight to obtain a model for infrared target recognition;
[0011] Step 6: Input the image to be recognized into the model to achieve target recognition.
[0012] In step 1, the target sizes in the dataset are statistically analyzed as follows: First, the width-to-height ratios of all targets in the dataset to the width and height of the image are statistically analyzed, and then the median of the width-to-height ratios is compared with a reference value to determine the size of the target.
[0013] In step 2, a convolutional kernel model adapted to the target size is constructed. LSKNet is used as the backbone network of the model, and the size of the convolutional kernel is determined according to the target size obtained in step 1.
[0014] In step 3, the pre-training of the convolutional kernel model on the dataset is completed. The default configuration of LSKNet is used, and the network model is initialized with the trained weights.
[0015] In step 4, on the dataset, multiple models are adjusted using different training methods, including the following steps:
[0016] The model to be adjusted uses LSKNet as the backbone network and OrientedRCNN as the detection head; different training methods are set, including data augmentation, number of training steps, and learning rate;
[0017] Among them, the training of data augmentation includes one or more of random flipping augmentation, random rotation augmentation, random shearing augmentation, and mosaic augmentation;
[0018] The training of the number of training steps is as follows: Under the condition that other training conditions are the same, the model is trained with different numbers of training steps;
[0019] The training of the learning rate is as follows: Under the condition that other training conditions are the same, the model is trained with different learning rates;
[0020] For the same training method, multiple models are trained, or at least one model is trained using different training methods.
[0021] In step 5, by means of parameter reparameterization, the weights of multiple models are reparameterized into the weight of a single model to obtain a model for infrared target recognition, including the following steps:
[0022] 1) For the training of multiple models with the same training method and the training of at least one model with different training methods, the multiple trained models are selected, and the weight parameters of different models are obtained;
[0023] 2) The weight parameters are recalculated using different combinations and saved;
[0024] 3) The initialized model with the calculated weight parameters is obtained as a model for infrared target recognition.
[0025] The re - calculation of the weight parameters using different combinations includes the following steps:
[0026] Combine the parameters of two or more models;
[0027] Calculate the average value of the model weight parameters in the combination by averaging;
[0028] An infrared target recognition system for heterogeneous image generalization includes:
[0029] A model construction module, used to count the target sizes in the data set, judge the target sizes; construct a convolutional kernel model adapted to the target sizes;
[0030] A model adjustment module, used to complete the pre - training of the convolutional kernel model on the data set; on the data set, adjust multiple models using different training methods; through the method of parameter re - parameterization, use the weights of multiple models to re - parameterize into a single model weight to obtain a model for infrared target recognition;
[0031] A target recognition module, used to input the image to be recognized into the model to achieve target recognition.
[0032] An infrared target recognition device for heterogeneous image generalization includes a memory and a processor; the memory is used to store a computer program; the processor is used to implement the infrared target recognition method for heterogeneous image generalization when executing the computer program.
[0033] A computer - readable storage medium stores a computer program, and when the computer program is executed by a processor, the infrared target recognition method for heterogeneous image generalization is implemented.
[0034] The present invention has the following beneficial effects and advantages:
[0035] 1. Based on the actual scene target size, construct a large convolutional kernel network model, which can effectively capture the shape features and context information of the target, and thus improve the target detection ability of the model;
[0036] 2. Adopt the pre - training method to pre - train the network model on the large - scale public data set ImageNet, effectively enrich the features learned by the model, and thus improve the expression ability of the model;
[0037] 3. Use the re - parameterization technology to re - parameterize the weights of multiple models trained in different ways, so that the model has better generalization ability in heterogeneous image target recognition. Description of the Drawings
[0038] Figure 1Flowchart of an infrared target recognition method for cross-source image generalization;
[0039] Figure 2 Schematic diagram of the reparameterization process of multiple models;
[0040] Figure 3 Detection result diagram. Specific implementation manner
[0041] In order to make the objectives, technical solutions, and advantages of the present invention clearer, the present invention will be further described below in conjunction with the accompanying drawings and implementation manners.
[0042] An infrared target recognition method for cross-source image generalization includes the following steps:
[0043] Step 1: Statistically analyze the target size of the data set and determine the target size;
[0044] Step 2: Construct a convolutional kernel model adapted to the target size;
[0045] Step 3: Complete the pre-training of the model on the ImageNet data set;
[0046] Step 4: On the visible light data set, fine-tune multiple models using different training methods;
[0047] Step 5: By means of parameter reparameterization, use the weights of multiple models to reparameterize into the weight of a single model.
[0048] Preferably, the step 1 of determining the target size includes step 1.1: Statistically analyze the ratios of the width and height of the target in the data set to the width and height of the image, and obtain the median of the width-to-height ratio; step 1.2: Determine the target size of the data set according to the size reference value, and use the target category of this data set as the label.
[0049] Preferably, the size reference value in the step 1.2 is 0.04. If the value is greater than 0.04, it indicates that the target in the data set is a large-size target; otherwise, it is a small-size target.
[0050] Preferably, in the step 2 of constructing a convolutional kernel model adapted to the target size, LSKNet is used as the backbone network, and an appropriate convolutional kernel size is selected according to the target size of the data set. The size of the convolutional kernel in the LSKblock is 5×5 and 7×7 according to the target size obtained in step 1.
[0051] Preferably, in step 3, the model is pre-trained on the ImageNet dataset, and the newly constructed network model is initialized with weights by pre-training on a large dataset. Since the default configuration of LSKNet is adopted in step 2 and the trained weights are used to initialize the network model; if the model structure is changed and the trained pre-trained weights cannot be adopted, the pre-training on ImageNet needs to be completed again;
[0052] Preferably, in step 4, on the visible light dataset, multiple models are fine-tuned in different ways. The model uses LSKNet as the backbone network and OrientedRCNN as the detection head, and mainly uses data augmentation, learning rate, and training steps to complete the fine-tuning of different models.
[0053] Data augmentation includes random flipping augmentation, random rotation augmentation, random shearing augmentation, and mosaic augmentation. Random flipping augmentation is a widely used augmentation method, including flipping augmentation in the vertical direction, horizontal direction, and diagonal direction. Random rotation augmentation realizes diverse augmentation at different angles by randomly rotating the image by an angle. Random cropping augmentation achieves the effect of data augmentation by cropping the image with different cropping sizes to form diverse samples. Mosaic augmentation refers to stitching four different images into one image, and the new image contains the targets and corresponding labels in all four images. By using different augmentation methods, the training of multiple models is realized.
[0054] The different learning rates affect the fitting effect of the model, and multiple models are fine-tuned by setting different learning rates.
[0055] The different training steps also have a certain impact on the features learned by the model. Therefore, multiple models can be obtained through training with different steps.
[0056] Preferably, in step 5, by means of parameter reparameterization, the weights of multiple models are reparameterized into the weight of a single model, including the following steps:
[0057] 1) Obtain the weight parameters of different models;
[0058] 2) Recalculate the weight parameters by different combinations and save them;
[0059] 3) Initialize the network model with the new weight parameters.
[0060] The different combinations refer to combining the parameters of two or more models and selecting the best combination; recalculating the weights means calculating the average value of the model weight parameters in the combination by averaging.
[0061] Such as Figure 1The infrared target recognition method for heterogeneous image generalization of the present invention is as follows:
[0062] Step 1: Statistically analyze the target sizes in the dataset and determine the target sizes. Statistically analyze the ratios of the widths and heights of all targets in the dataset to the width and height of the image, then obtain the corresponding median, and finally, based on the size reference value of 0.04, determine the target sizes in the dataset. If the median is less than 0.04, it indicates that the target is a small target; otherwise, it is a large target.
[0063] Step 2: Construct a convolutional kernel model adapted to the target size. Use the LSKNet model as the backbone. The model structure adopts the currently popular stacked structure layout. Among them, the repeatedly stacked module is the LSK module. This module includes large convolutional kernel convolution and a spatial convolutional kernel selection mechanism. The large convolutional kernel convolution operation is a depth convolution sequence, and in the sequence, the convolutional kernel size is continuously increased and the dilation rate of the dilated convolution is increased to capture long-range context information as much as possible. For the i-th depth convolution in the sequence, the relationship between the convolutional kernel size k, the dilation rate d, and the receptive field RF is expressed as follows:
[0064] k i-1 ≤k i ; d 1 =1,d i-1 <d i ≤RF i-1 ,
[0065] RF 1 =k 1 ,RF i =d i (k i -1)+RF i-1 .
[0066] Through a cascading method, a large-sized large convolutional kernel can be decomposed into two smaller convolutional kernels, which can reduce the computational amount while ensuring the same receptive field. The spatial convolutional kernel selection mechanism enables the network to more effectively focus on the context space regions related to the target, generates a spatial attention feature map using average pooling and maximum pooling operations, and finally multiplies the generated feature attention map with the original input feature map to achieve the selection of relevant spatial scales.
[0067] Furthermore, according to the target sizes in the dataset, the convolutional kernel size can be reasonably set to ensure better recognition effects while reducing computational redundancy.
[0068] Step 3: Complete the pre-training of the model on the ImageNet dataset. Pre-training on a large dataset can effectively improve the generalization performance of the model.
[0069] Preferably, step 4: On the visible light data set, fine-tune multiple models using different training methods. By using different data augmentation methods, learning rates, and training steps, the training of multiple models is completed.
[0070] Data augmentation uses random flipping augmentation, random rotation augmentation, random cropping augmentation, and mosaic augmentation. Random flipping augmentation is a widely used augmentation method, including flipping augmentation in the vertical, horizontal, and diagonal directions. Random rotation augmentation realizes diverse augmentation at different angles by randomly rotating the image by an angle. Random cropping augmentation achieves the effect of data augmentation by cropping the image with different cropping sizes to form diverse samples. Mosaic augmentation refers to stitching four different images into one image, and the new image contains the objects and corresponding labels in all four images. During training, multiple models are trained using different combinations of data augmentation, and the model weights are saved.
[0071] Different learning rates affect the fitting effect of the model. Multiple models are fine-tuned by setting different learning rates. During training, the learning rates are 0.0001, 0.0002, and 0.0004 respectively, and the training steps are all 6. Three models trained with different learning rates are obtained, and the model weights are saved.
[0072] Different training steps also have a certain impact on the features learned by the model. Therefore, multiple models can be obtained through training with different steps. During training, when the learning rate is 0.0001, the training steps are set to 6 and 12 respectively, and two different models are obtained and the model weights are saved; when the learning rate is 0.0002, the training steps are also set to 6 and 12 respectively, and two different models are obtained and the model weights are saved.
[0073] Preferably, step 5: By means of parameter reparameterization, use the weights of multiple models to reparameterize into a single model weight. The process is as Figure 2 shown. It includes the following steps:
[0074] Step 5.1 Load the weights of different models obtained by training through the deep learning framework;
[0075] Step 5.2 Combine the weights of two or more different models, and recalculate the weight parameters by averaging the weight parameters. The formula is expressed as AVE represents the weight parameters after averaging, param1, param2, paramM represent the weight parameters of different models, and M is the number of models participating in the combination. The calculated weight parameters are re-initialized into the network for testing, and the model with the best recognition effect is selected and saved;
[0076] Step 5.3 Load the obtained optimal weight parameters into the network model for actual detection.
[0077] The recognition result of the infrared target recognition method for heterogeneous image generalization proposed by the present invention is as Figure 3 shown. Compared with the best detection result of the original model, the precision of the model after weight combination has increased by 1.4 percentage points, and the recall has increased by 2.4 percentage points.
[0078] Among them, the definition of precision: TP represents the true positive sample, and FP is the false positive sample;
[0079] The definition of recall: FN represents the false negative sample.
Claims
1. An infrared target recognition method for generalization of heterogeneous images, characterized in that: The following steps are involved: Step 1: Count the target size of the statistical data set and determine the target size; Step 2: Build a convolution kernel model that adapts to the target size; Step 3: Complete the pre-training of the convolution kernel model on the data set; Step 4: Use different training methods to adjust multiple models on the data set; Step 5: Re-parameterize multiple model weights into a single model weight by means of parameter re-parameterization to obtain a model for infrared target recognition; Step 6: Input the image to be identified into the model to achieve target recognition.
2. The infrared target recognition method for generalization of heterogeneous images according to claim 1 is characterized in that: The step 1, counting the size of the target in the statistical data set, is specifically as follows: firstly, the width and height of all the targets in the statistical data set and the aspect ratio of the image are counted, and then the median of the aspect ratio is compared with the reference value to determine the size of the target.
3. The infrared target recognition method for generalization of heterogeneous images according to claim 1 is characterized in that: In the step 2, a convolution kernel model adapted to the target size is constructed, LSKNet is used as the backbone network of the model, and the size of the convolution kernel is determined according to the target size obtained in step 1.
4. The infrared target recognition method for generalization of heterogeneous images according to claim 1 is characterized in that: In step 3, the convolution kernel model is pre-trained on the data set, the default configuration of LSKNet is adopted, and the trained weights are used to initialize the network model.
5. The infrared target recognition method for generalization of heterogeneous images according to claim 1 is characterized in that: The step 4, using different training methods to adjust multiple models on the data set, includes the following steps: The model to be adjusted uses LSKNet as the backbone network and OrientedRCNN as the detection head; different training methods are set including: data enhancement, training steps, and learning rate; The data enhancement training includes one or more of random flip enhancement, random rotation enhancement, random shear enhancement and mosaic enhancement; The training steps are as follows: when other training conditions are the same, different training steps are used to train the model; The learning rate is used for training, as follows: when other training conditions are the same, different learning rates are used to train the model; Multiple models are trained using the same training method, or at least one model is trained using different training methods.
6. The infrared target recognition method for generalization of heterogeneous images according to claim 1 is characterized in that: The step 5, using a parameter re-parameterization method to re-parameterize multiple model weights into a single model weight to obtain a model for infrared target recognition, includes the following steps: 1) Training multiple models using the same training method and training at least one model using different training methods, selecting multiple models obtained through training, and obtaining weight parameters of different models; 2) Recalculate the weight parameters using different combinations and save them; 3) Initialize the model with the calculated weight parameters to obtain a model for infrared target recognition.
7. The infrared target recognition method for generalization of heterogeneous images according to claim 6 is characterized in that: The method of recalculating weight parameters using different combinations includes the following steps: Combine parameters of two or more models; The average value of the model weight parameters in the combination is calculated by averaging.
8. An infrared target recognition system for generalization of heterogeneous images, characterized in that: include: Model building module, used to count the target size of the data set and determine the target size; Construct a convolution kernel model that adapts to the target size; The model adjustment module is used to complete the pre-training of the convolution kernel model on the data set; on the data set, different training methods are used to adjust multiple models; By means of parameter reparameterization, multiple model weights are reparameterized into a single model weight to obtain a model for infrared target recognition; The target recognition module is used to input the image to be recognized into the model to realize target recognition.
9. An infrared target recognition device for generalization of heterogeneous images, characterized in that: It comprises a memory and a processor; the memory is used to store a computer program; the processor is used to implement an infrared target recognition method for generalization of heterogeneous images as described in any one of claims 1 to 7 when executing the computer program.
10. A computer-readable storage medium, characterized in that: The storage medium stores a computer program, and when the computer program is executed by a processor, an infrared target recognition method for generalization of heterogeneous images as described in any one of claims 1 to 7 is implemented.