Domain-adaptive infrared target identification method and device, equipment and medium
Through improved teacher-student framework and field adaptation technology, combined with Nice GAN and Faster R-CNN models, the problems of unsupervised field adaptation in infrared image recognition are solved, and high-precision and robust infrared target recognition are achieved.
Patent Information
- Application Number
- CN202510155245.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-12
- Publication Date
- 2025-06-27
AI Technical Summary
The existing target recognition technology has problems of instability in the unsupervised field adaptation and reliance on high-quality generative models, especially in infrared image recognition. Due to the high cost of data set annotation, it is difficult to achieve effective recognition.
The improved teacher-student framework is adopted, combined with domain adaptive technology, and the tag-consistent class-target domain data is generated through the Nice GAN model, and trained using the Faster R-CNN model to achieve infrared target recognition. This method improves the model's learning ability and cross-domain adaptability of subtle features through probability reconstruction and adaptive weighted loss function of Gaussian distribution.
It improves the accuracy and reliability of infrared target recognition, reduces the cost and workload of manual annotation, and enhances the cross-domain application flexibility and generalization performance of the model.
Smart Images

Figure CN120219701A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of target recognition, and more particularly, to a domain-adaptive infrared target recognition method, apparatus, device, and medium that combines an improved teacher-student framework. Background Art
[0002] Target recognition technology has a wide range of applications in multiple fields, such as airport security, road safety, security monitoring, etc. In these application scenarios, target recognition technology needs to have a high degree of dynamic adaptability to cope with challenges in different environments. For example, in the field of autonomous driving, vehicles need to accurately identify targets in various weather conditions and terrain features, which poses strict requirements on the adaptability of target detection models. Domain adaptation technology has emerged, which can enable target detection models trained only under specific conditions to adjust their own parameters and then adapt to other different driving environments, aiming to enhance the generalization ability of the model in different environments and ensure the stable operation of the system under changing conditions.
[0003] However, existing target recognition technologies have some limitations in unsupervised domain adaptation. Traditional adversarial-based methods introduce a domain classifier and a gradient reversal component into the model to disrupt the judgment of the domain discriminator in order to achieve domain invariance of features. However, this method is prone to causing model instability during training, thereby affecting the sensitivity to fine-grained features. Another method is progressive generation of image reconstruction, which expands the training set by constructing intermediate domain data, but the performance of this method highly depends on the quality of the generation model. In addition, although comprehensive methods attempt to combine multiple technical means, it is still difficult to fully meet the requirements for robustness, generalization, and universality in practical applications.
[0004] In practical applications, it is relatively easy to obtain visible light datasets. However, in the aspect of infrared image recognition, due to the high cost of annotating infrared image datasets and the complex and changeable application scenarios, it is difficult to bear the cost of annotation before each application. Summary of the Invention
[0005] The present invention provides a domain-adaptive infrared target recognition method, apparatus, device, and medium that combines an improved teacher-student framework to improve at least one of the above technical problems.
[0006] In a first aspect, the present invention provides a domain-adaptive infrared target recognition method that combines an improved teacher-student framework, which includes steps S10 to S10.
[0007] S01. Obtain labeled training data of the source domain and unlabeled training data of the target domain. Among them, the source domain is visible light data, and the target domain is infrared light data.
[0008] S02. Obtain a pre-trained Nice GAN model. Among them, the Nice GAN model is used to generate class target domain images with consistent labels according to source domain images.
[0009] S03. Obtain a Faster R-CNN model as the backbone network. Among them, the output results probability of Faster R-CNN is reconstructed as: class label cls, Gaussian distribution reg-a of the first axial coordinate, Gaussian distribution reg-b of the second axial coordinate, Gaussian distribution reg-w of the width, and Gaussian distribution reg-h of the height.
[0010] S04. Input the labeled training data into the Nice GAN model to generate labeled class target domain data consistent with the original labels.
[0011] S05. Use the labeled training data and the labeled class target domain data to train the Faster R-CNN model.
[0012] S06. Use the trained Faster R-CNN model as a single backbone network to establish a teacher model and a student model.
[0013] S07. Input the unlabeled training data into the teacher model to generate pseudo-labels.
[0014] S08. Combine the labeled training data, the labeled class target domain data, and the unlabeled training data with the pseudo-labels into a batch, and input them into the student model to perform consistency training under uncertainty guidance to obtain the trained student model.
[0015] S09. Obtain an image to be recognized. Among them, the image to be recognized is a visible light image or an infrared light image.
[0016] S10. Input the image to be recognized into the trained student model to obtain the recognition information of the image to be recognized.
[0017] In an optional implementation manner, during the training of the student model, there is also an adversarial training module connected to the student model. The adversarial training module includes a gradient reversal layer GRL and a domain discriminator composed of a convolutional layer and a multi-layer perceptron.
[0018] During the forward propagation process, the feature encoding extracted by the student model is transmitted to the domain discriminator through the gradient reversal layer GRL to identify whether the feature comes from the source domain or the target domain. During this process, the gradient reversal layer GRL acts as an identity transformation layer and does not change the input data.
[0019] During the backpropagation process, the gradient generated by the classification loss of the domain discriminator first passes through the domain discriminator to complete the weight update of the domain discriminator. Then it is passed to the student model through the gradient reversal layer GRL. During this process, the gradient reversal layer GRL performs an inversion operation on the gradient to maximize the loss of the backbone network of the student model learning the domain discriminator.
[0020] The inversion operation is: 。
[0021] In the formula, represents the differential, is the classification loss of the domain discriminator, is the output of the network layer before GRL, is a negative constant factor, is the output of GRL. represents the gradient.
[0022] In an optional implementation, the adversarial training module includes a gradient reversal layer GRL, two convolutional layers or one convolutional layer, and a multi-layer perceptron MLP. Among them, two convolutional layers or one convolutional layer and an MLP form a domain classifier.
[0023] In an optional implementation, the classification loss of the domain discriminator during training is: 。
[0024] Where, is the classification loss of the domain discriminator, is the number of pictures, is the domain loss function, is the output of the domain classifier of the th picture, is the th picture's feature encoding.
[0025] In an optional implementation, the total loss function during consistency training is: 。
[0026] In the formula, is the total loss function, is the exponential function with the base of the natural logarithm, is the current training iteration number, is the upper limit of the iteration number, is the number of converted pictures, is the supervised source domain loss of the model trained on source domain data, is the unsupervised target domain loss of the model trained on target domain data using pseudo-labels, is the classification loss of the domain discriminator, is the fourth hyperparameter.
[0027] In an alternative embodiment, during consistency training, the weights of the teacher model are updated based on the weights of the student model by an exponential moving average strategy.
[0028] In an alternative embodiment, the loss function during the training of the Nice GAN model includes: .
[0029] .
[0030] .
[0031] .
[0032] In the formula, is the overall loss function of the Nice GAN, is the first generator, is the second generator, is the first balancing hyperparameter, is the adversarial loss, is the second balancing hyperparameter, is the consistency loss, is the third balancing hyperparameter, is the image data reconstruction loss, is the adversarial loss of the first generator, is the adversarial loss of the second generator, is the consistency loss of the first generator, is the consistency loss of the second generator, is the image data reconstruction loss of the first generator. is the image data reconstruction loss of the second generator.
[0033] In an alternative embodiment, during the training of the Nice GAN model, the training of the first feature encoder is decoupled from the training of the first feature generator .
[0034] The loss functions of the decoupled Nice GAN model are: .
[0035] .
[0036] .
[0037] 。
[0038] In the formula, is the adversarial loss of the first generator, is the first generator, is the image, is the converted image, is the second feature discriminator, is the first feature encoding, is the number of images before conversion, is the second feature encoder, is the number of converted images, is the second generator, is the consistency loss of the first generator, is the image data reconstruction loss of the first generator.
[0039] Modify the in the loss function of the first generator to , Modify , and the loss function of the second generator can be obtained.
[0040] In a second aspect, the present invention provides a domain adaptive infrared target recognition device combining an improved teacher-student framework, which includes a training data acquisition module, a domain conversion module, a backbone network module, an image conversion module, a backbone training module, a model construction module, a pseudo-label module, a training module, an image to be recognized acquisition module, and a recognition module.
[0041] The training data acquisition module is used to acquire labeled training data in the source domain and unlabeled training data in the target domain. Among them, the source domain is visible light data, and the target domain is infrared light data.
[0042] The domain conversion module is used to acquire a pre-trained Nice GAN model. Among them, the Nice GAN model is used to generate label-consistent class target domain images according to source domain images.
[0043] The backbone network module is used to acquire the Faster R-CNN model as the backbone network. Among them, the output result probability of Faster R-CNN is reconstructed as: class label cls, Gaussian distribution reg-a of the first axial coordinate, Gaussian distribution reg-b of the second axial coordinate, Gaussian distribution reg-w of the width, and Gaussian distribution reg-h of the height.
[0044] The image conversion module is used to input the labeled training data into the Nice GAN model to generate labeled class target domain data consistent with the original label.
[0045] A backbone training module for training the Faster R-CNN model using the labeled training data and the labeled class target domain data.
[0046] A model construction module for establishing a teacher model and a student model with the trained Faster R-CNN model as a single backbone network.
[0047] A pseudo-label module for inputting the unlabeled training data into the teacher model to generate pseudo-labels.
[0048] A training module for combining the labeled training data, the labeled class target domain data, and the unlabeled training data with the pseudo-labels into a batch and inputting them into the student model to perform consistency training under uncertainty guidance to obtain the trained student model.
[0049] An image to be recognized acquisition module for acquiring an image to be recognized. Wherein, the image to be recognized is a visible light image or an infrared light image.
[0050] A recognition module for inputting the image to be recognized into the trained student model to obtain recognition information of the image to be recognized.
[0051] In a third aspect, the present invention provides an infrared target recognition device with domain adaptation combined with an improved teacher-student framework, which includes a processor, a memory, and a computer program stored in the memory. The computer program can be executed by the processor to implement an infrared target recognition method with domain adaptation combined with an improved teacher-student framework as described in any paragraph of the first aspect.
[0052] In a fourth aspect, the present invention provides a computer-readable storage medium. The computer-readable storage medium includes a stored computer program, wherein when the computer program runs, it controls the device where the computer-readable storage medium is located to execute an infrared target recognition method with domain adaptation combined with an improved teacher-student framework as described in any paragraph of the first aspect.
[0053] By adopting the above technical solutions, the present invention can achieve the following technical effects: An infrared target recognition method of the present invention improves the accuracy and reliability of infrared target recognition by improving the teacher-student framework and combining domain adaptation technology, and provides a more robust and efficient solution for infrared target recognition.
[0054] In traditional methods, models often rely on hard pseudo-labels for learning, which limits the learning ability of subtle features. In contrast, the present invention transforms traditional hard pseudo-labels into soft pseudo-labels in the form of probability distributions by implementing probability reconstruction of Gaussian distribution on the backbone network, thereby improving the learning of subtle features. Additionally, an adaptive weighted loss function is redesigned to better match the differences between different domains. This innovation ensures effective recognition of infrared targets even when training on unlabeled infrared image datasets, significantly reducing the cost and workload of manual annotation.
[0055] Data of the target domain is generated with the help of a generative adversarial network, enabling the model to adapt to the data distribution of the target domain during both the supervised learning stage and the co-training stage, thus enhancing the flexibility of the model for cross-domain applications.
[0056] The invention has been deeply optimized at both the image level and the feature level. At the image level, an adversarial training mechanism is adopted, combined with a domain discriminator and a gradient reversal layer, continuously strengthening the model's ability to extract cross-domain invariant features and ensuring the stability and generalization performance of the model under different environmental conditions. At the feature level, the problem of the inherent distribution difference between the source domain and the target domain is solved through the generation of style-converted images, greatly improving the adaptability of the model to complex and variable application scenarios. BRIEF DESCRIPTION OF THE DRAWINGS
[0057] To more clearly illustrate the technical solutions of the present invention, the drawings required for use in the specific embodiments of the present invention will be briefly introduced below. It should be understood that the following drawings only show certain specific embodiments of the present invention and should not be regarded as limiting the scope. For those of ordinary skill in the art, other relevant drawings can be obtained based on these drawings without creative efforts.
[0058] Figure 1 is a schematic flowchart of an infrared target recognition method.
[0059] Figure 2 is a schematic diagram of the probability reconstruction of the Faster R-CNN model.
[0060] Figure 3 is a model framework diagram of an infrared target recognition method.
[0061] Figure 4 is a framework diagram of a domain adversarial training module.
[0062] Figure 5 is a framework diagram of the NiceGAN model. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0063] The technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0064] Embodiment 1. Please refer to Figures 1 to 5 , the first embodiment of the present invention provides a domain - adaptive infrared target recognition method combining an improved teacher - student framework, which can be executed by a domain - adaptive infrared target recognition device combining an improved teacher - student framework (hereinafter referred to as: recognition device). Specifically, it is executed by one or more processors in the recognition device to implement steps S10 to S10.
[0065] It can be understood that the recognition device can be an electronic device with computing performance such as a portable notebook computer, a desktop computer, a server, a smart phone or a tablet computer.
[0066] S01. Obtain the labeled training data of the source domain and the unlabeled training data of the target domain. Among them, the source domain is visible - light data, and the target domain is infrared - light data.
[0067] S02. Obtain a pre - trained Nice GAN model. The Nice GAN model is used to generate class - target - domain images with consistent labels according to the source - domain images.
[0068] Specifically, currently, visible - light data sets are very common and easy to obtain, and it is very convenient to train the model using the labeled visible - light data set. However, in the actual application of the model, infrared images also need to be recognized. If the labeled infrared - image data set is used for training, although the model can achieve a high recognition accuracy, manually annotating the infrared images in the actual application scenario will bring additional costs.
[0069] In this embodiment, the labeled visible - light image data set is the source domain, and the unlabeled infrared - image data set is the target domain. The source - domain and target - domain data sets are used as a unified training set for domain - adaptive training of the model. By using the Nice GAN model to generate class - target - domain images with consistent labels according to the source - domain images, the model can also achieve the recognition of infrared targets without manual annotation.
[0070] S03. Obtain the Faster R-CNN model as the backbone network. Among them, the output result probability of Faster R-CNN is reconstructed as: class label cls, Gaussian distribution reg-a of the first axial coordinate, Gaussian distribution reg-b of the second axial coordinate, Gaussian distribution reg-w of the width, and Gaussian distribution reg-h of the height.
[0071] Specifically, Faster R-CNN is a benchmark detector for unsupervised domain adaptation object detection tasks (UDA-OD). This detector decomposes object detection into a classification branch based on cross-entropy and a bounding box regression branch based on the L1 loss. In the classification branch, predicting the probability distribution of the class label space can naturally capture the uncertainty of the class. However, the bounding box regression branch modeled using the Dirac δ function cannot obtain the uncertainty of localization.
[0072] To solve this problem, this embodiment extends the existing Faster R-CNN detector to a probabilistic detector. Specifically, both the instance category and the localization label of the bounding box are represented as probability distributions. As Figure 2 shown, the left side is the output of the traditional Faster R-CNN. The output of the traditional Faster R-CNN includes the class label cls and each coordinate information of the bounding box ; among them, represents the coordinate value of the center point of the bounding box in the first axis direction; represents the coordinate value of the center point of the bounding box in the second axis direction; represents the logarithmic scale change of the bounding box width relative to the anchor box width; represents the logarithmic scale change of the bounding box height relative to the anchor box height; and in this embodiment, each coordinate information of the bounding box is represented as a single Gaussian distribution model, and its specific visual representation is as shown on the right side in Figure 4 .
[0073] S04. Input the labeled training data into the Nice GAN model to generate labeled class target domain data consistent with the original label.
[0074] S05. Use the labeled training data and the labeled class target domain data to train the Faster R-CNN model.
[0075] Specifically, steps S04 and S05 are the initial stages of training.
[0076] The training objective in the initial stage is as follows: Initialize a single backbone network of Faster R-CNN and perform supervised training on the source domain data. At the same time, use the pre-trained Nice GAN model to generate class target domain images that are consistent with the original labels. By combining the supervised training of the source domain and class target domain images, enhance the detection ability of the model and improve the lower limit of the model performance.
[0077] The specific operations in the initial stage are as follows: First, initialize the Faster R-CNN backbone network. Then, use the labeled source domain data to perform supervised training on it so that it can learn the features and annotation information of the source domain data. At the same time, use the pre-trained Nice GAN model to convert the source domain images into class target domain style images. These class target domain images are consistent with the original labels and can be used as additional training data to help the model better adapt to the data distribution of the target domain.
[0078] S06: Use the trained Faster R-CNN model as a single backbone network to establish a teacher model and a student model.
[0079] Specifically, step S06 is the generation stage of training.
[0080] The training objective in the generation stage is as follows: Initialize two backbone networks with exactly the same structure, which are used as the teacher model and the student model respectively. And copy the model weights pre-trained in the first stage to the teacher and student models to complete the model initialization and prepare for the subsequent collaborative training stage.
[0081] The specific operations in the generation stage are as follows: Create two identical backbone networks. One is used as the teacher model, which is specifically used to generate pseudo-labels for the student model to train and does not involve gradient transmission. The other is used as the student model, which is used to receive the pseudo-labels generated by the teacher model and the source domain data for training. Then copy the weights of the Faster R-CNN model trained in the initial stage to these two models so that both the teacher model and the student model have a certain initial detection ability, laying a foundation for the subsequent collaborative training.
[0082] S07: Input the unlabeled training data into the teacher model to generate pseudo-labels.
[0083] S08: Combine the labeled training data, labeled class target domain data, and unlabeled training data with the pseudo-labels into a batch and input them into the student model to perform consistency training under uncertainty guidance to obtain the trained student model.
[0084] Specifically, steps S07 and S08 are the collaborative training stage.
[0085] The training objective in the co-training stage is as follows: generate pseudo-labels through the teacher model, combine the labeled source domain data and class target domain data, as well as the target domain data with pseudo-labels into a batch, perform strong data augmentation and then input it into the student model to implement consistency training under uncertainty guidance, thereby improving the model's recognition ability for target domain data and enhancing the model's cross-domain transfer ability and generalization performance.
[0086] The specific operations in the co-training stage are as follows: First, input the target domain data into the teacher model after weak data augmentation. The teacher model generates pseudo-labels for the target domain data based on the knowledge it has learned. Then, combine the strongly data-augmented target domain data with pseudo-labels, the labeled class target domain images (data generated by Nice GAN with a style similar to the target domain), and the labeled training data after conventional data augmentation into a batch Batch together to increase the diversity and robustness of the data. Finally, input the processed data into the student model, and the student model performs consistency training under the guidance of this data. By learning the annotation information of the source domain data, the style information of the class source domain data, and the pseudo-label information of the target domain data, the model continuously optimizes its own parameters to improve the recognition accuracy for the target domain data.
[0087] Weak data augmentation mainly increases the diversity of data through some minor transformations, which have little impact on the data. Specifically, it includes operations such as random cropping, random flipping, random rotation, and color jitter. These methods can provide more training samples for the model without significantly changing the image content, helping the model learn more robust features.
[0088] Based on weak data augmentation, conventional data augmentation adds more types and greater degrees of transformations. For example, random scaling, random cropping and scaling, Gaussian blur, random noise, and random cropping and padding. These transformations further increase the diversity of the data, enabling the model to better adapt to different image conditions, thereby improving the model's generalization ability.
[0089] Strong data augmentation includes more complex and drastic transformations, which have a greater impact on the data. Common strong data augmentation methods include Mixup, CutMix, RandAugment, AutoAugment, CutOut, GridMask, and elastic deformation. These methods enable the model to learn more complex features through more drastic transformations, thereby significantly improving the model's generalization ability and robustness, and are suitable for tasks and datasets that require higher robustness.
[0090] S09. Obtain the image to be recognized. Among them, the image to be recognized is a visible light image or an infrared light image.
[0091] S10. Input the image to be recognized into the trained student model to obtain the recognition information of the image to be recognized.
[0092] A domain - adaptive infrared target recognition method of the present invention improves the accuracy and reliability of infrared target recognition by improving the teacher - student framework and combining domain - adaptive technology, providing a more robust and efficient solution for infrared target recognition.
[0093] In traditional methods, the model often relies on hard pseudo - labels for learning, which limits the learning ability of fine features. In the present invention, by implementing probability reconstruction of Gaussian distribution on the backbone network, the traditional hard pseudo - labels are transformed into soft pseudo - labels in the form of probability distribution, improving the learning of fine features. In addition, an adaptive weighted loss function is redesigned to better match the differences between different domains. This innovation ensures effective recognition of infrared targets even when training on an unlabeled infrared image dataset, greatly reducing the cost and workload of manual annotation.
[0094] Generate data of the target domain with the help of a generative adversarial network, enabling the model to adapt to the data distribution of the target domain during both the supervised learning stage and the co - training stage, thus enhancing the flexibility of the model for cross - domain applications.
[0095] The invention has been deeply optimized at both the image level and the feature level. At the image level, an adversarial training mechanism is adopted, combined with a domain discriminator and a gradient reversal layer, continuously strengthening the model's ability to extract cross - domain invariant features and ensuring the stability and generalization performance of the model under different environmental conditions. At the feature level, by generating images after style transfer, the inherent distribution difference problem between the source domain and the target domain is solved, greatly improving the adaptability of the model to complex and variable application scenarios.
[0096] During the mutual learning process under the teacher - student framework, one of the main challenges is the domain bias caused by the limitation of training data (only source - domain images are equipped with labels). This bias means that the teacher model mainly relies on its understanding of the source - domain annotation information when generating pseudo - labels for target - domain images, which may lead to a relatively high noise ratio in the pseudo - labels generated for the target domain. If this phenomenon is not solved, it will cause the student model to deviate from the correct learning trajectory during the learning process in the target domain, and then fall into a local optimum and be unable to fully explore the potential knowledge of the target domain, ultimately affecting the generalization ability of the entire model.
[0097] Therefore, on the basis of the above - mentioned embodiments, in an optional embodiment of the present invention, as Figure 3 and Figure 4As shown, during the training of the student model, an adversarial training module connected to the student model is also included. The aim is to reduce the distribution difference between the source domain and the target domain. The domain adversarial training module discriminates whether the features come from the source domain or the target domain by introducing a domain discriminator. At the same time, the backbone network (feature extractor) tries to deceive this domain discriminator so that it cannot distinguish the differences between the two domains. This method simulates a "chasing game". In this process, the features learned by the model not only contain information beneficial to the task but also achieve better alignment between the two domains, thus effectively reducing the impact brought by domain shift. The student model obtains images from two domains during the joint training phase, so adversarial learning can be used to make the student model align the distributions between the two domains.
[0098] Specifically, the adversarial training module includes a gradient reversal layer GRL and a domain discriminator composed of a convolutional layer and a multi-layer perceptron. Preferably, as Figure 4 shown, the adversarial training module consists of a gradient reversal layer GRL, two convolutional layers, and a multi-layer perceptron MLP. Among them, the two convolutional layers and an MLP form a domain classifier. Specifically, the number of channels of the first convolutional layer in the two convolutional layers is 512, and the convolutional kernel is 1. The number of channels of the second convolutional layer is 256, and the convolutional kernel is 1. In other embodiments, only one convolutional layer may also be included.
[0099] In the domain classifier, the input feature data will pass through two convolutional layers (or one convolutional layer). The role of the convolutional layer is to perform spatially local perception and feature extraction on the input features. By sliding the convolutional kernel on the feature map, local patterns and texture information in the feature map can be captured. For example, in image processing tasks, the convolutional layer can detect basic visual elements such as edges and corners. The output of the convolutional layer is a feature map with a higher-level feature representation, and these feature maps contain the abstract information of the input data, providing a basis for subsequent classification tasks.
[0100] The output feature map of the convolutional layer will be flattened or rearranged into a one-dimensional vector and then input into the multi-layer perceptron. The multi-layer perceptron consists of multiple fully connected layers, and each layer contains multiple neurons. In the MLP, the input data will be processed through successive linear transformations and non-linear activation functions (such as ReLU, sigmoid, etc.). The linear transformation can perform weighted summation on the input features, and the non-linear activation function introduces non-linearity to the model, enabling the model to learn complex function mapping relationships. Through successive abstractions and transformations of the multi-layer perceptron, a probability value is finally output, indicating the probability that the input features belong to the target domain.
[0101] During the forward propagation of the training process, the feature encoding extracted by the student model is transmitted to the domain discriminator through the Gradient Reversal Layer (GRL) to identify whether the features come from the source domain or the target domain. During this process, the GRL acts as an identity transformation layer and does not change the input data.
[0102] Specifically, during the forward propagation process, the GRL acts as an identity transformation layer and does not change the input data. Assume that the output of the network layer before the GRL is , then the output of the GRL is simply equal to its input, that is: . This means that for forward propagation, the GRL does not affect the data flow and the prediction output of the network.
[0103] During the backpropagation of the training process, the gradient generated by the classification loss of the domain discriminator first passes through the domain discriminator to complete the weight update of the domain discriminator. Then it is passed to the student model through the GRL. During this process, the GRL performs an inversion operation on the gradient to maximize the loss of the backbone network of the student model learning the domain discriminator.
[0104] Specifically, during the backpropagation process, the main role of the GRL is to "flip" the gradient passed to it, that is, multiply it by a negative constant factor . In this way, the GRL realizes the direction flipping of the gradient.
[0105] The inversion operation is: .
[0106] In the formula, represents the differential, is the classification loss of the domain discriminator, is the output of the network layer before the GRL, is the negative constant factor, is the output of the GRL. represents the gradient.
[0107] The above is the method of the GRL's action in the loss function. For the loss calculation of the domain discriminator, assume that the image after being processed by the backbone network of the student model has the feature encoding . Each batch of data input to the student model consists of source domain and target domain data in a one-to-one ratio (the target domain data source belongs to the source domain data). The domain label of the input image can be defined as a binary label (in the specific implementation of the model, the images in the source domain are marked as , and the images in the target domain are marked as ), so the domain loss function of each input image can be based on The binary cross-entropy loss is used to update the domain classifier. The output of the domain classifier represents the probability that the picture belongs to the domain label.
[0108] Generally speaking, the classification loss of the domain discriminator is defined as: .
[0109] Among them, is the classification loss of the domain discriminator, is the number of pictures, is the domain loss function, is the output of the domain classifier for the th picture, and
[0110] is the feature encoding of the th picture.
[0110] Specifically, for the domain classifier, its training purpose is to better judge which domain the input features belong to. For the feature encoding of the input image, its training purpose is to extract the domain-invariant features of the image as much as possible to deceive the domain classifier. These two contradictory training purposes can be unified by adding a gradient reversal layer GRL between the domain discriminator and the Backbone.
[0111] In the backpropagation stage, the generated gradients are backpropagated. The gradients first pass through the domain discriminator with several convolutional layers and a multi-layer perceptron to complete the weight update of the domain discriminator and optimize the ability of the domain discriminator to distinguish the source domain and the target domain. Before the gradients are passed to the backbone network Backbone, they will pass through the gradient reversal layer, and the gradient reversal layer will reverse the passed gradients. This operation will maximize the loss of the Backbone learning the domain discriminator. In other words, it makes the Backbone learn the domain-invariant features of the image to confuse the discriminator's distinction of the domain.
[0112] Based on the above embodiments, in an optional embodiment of the present invention, during consistency training, the weights of the teacher model are updated based on the weights of the student model through the exponential moving average strategy (EMA). Through the above gradient transfer operation, the domain shift problem of the teacher model in the pseudo-label generation stage can be solved, the quality of the pseudo-label generation can be improved, and the update process of the teacher model under the EMA strategy is more stable.
[0113] Based on the above embodiments, in an optional embodiment of the present invention, the total loss function mainly consists of three parts, namely the supervised source domain loss trained by the detection model on the source domain data and the unsupervised target domain loss trained on the target domain data using pseudo-labels, and the classification loss of the domain discriminator . To coordinate the losses, this embodiment adopts the idea of dynamic weighted loss based on the number of training iterations to dynamically adjust the and . During the co-training phase, the model tends to optimize the supervised loss on the source domain and reduce the perturbation of the model by the inferior pseudo-labels generated in the initial stage of training. As the training iterates, in the later stage of model training, the model tends to optimize the unsupervised loss on the target domain to fully exploit the potential of the teacher model. And the classification loss of the domain discriminator is fixed using the fourth hyperparameter β.
[0114] Generally speaking, the total loss function of the model during consistency training is: .
[0115] In the formula, is the total loss function, is the exponential function with the base of the natural logarithm, is the current training iteration number, is the upper limit of the iteration number, the number of converted pictures, is the supervised source domain loss of the model trained on the source domain data, is the unsupervised target domain loss of the model trained on the target domain data using pseudo-labels, is the classification loss of the domain discriminator, is the fourth hyperparameter.
[0116] Based on the above embodiment, in an optional embodiment of the present invention, the framework diagram of the Generative Adversarial Network without Independent Encoding Component (No-Independent-Component-for-Encoding GAN, Nice-GAN model) is as shown in Figure 5 . In this model, the adversarial loss is responsible for achieving point-to-point migration of the image data style. Its training architecture is carefully divided into three core loss parts: adversarial loss, consistency loss, and image reconstruction loss. Among them, the image data reconstruction loss and the consistency loss cooperate in the model to ensure the correlation of different style semantic information during the conversion and generation process of the image data. Specifically, if we want to complete a conversion from picture to picture , and during the inverse conversion process from picture back to picture again, the model must follow the loss function in the same unified form.
[0117] Based on the above embodiments, in an optional embodiment of the present invention, the Nice-GAN model includes two generators, two encoders, and two discriminators. For example, Cycle GAN uses two generators and two discriminators to achieve a two-way domain conversion task.
[0118] The loss function during the training of the Nice GAN model includes: 。
[0119] 。
[0120] 。
[0121] 。
[0122] In the formula, is the overall loss function of Nice GAN, is the first generator, is the second generator, is the first balancing hyperparameter, is the adversarial loss, is the second balancing hyperparameter, is the consistency loss, is the third balancing hyperparameter, is the image data reconstruction loss, is the adversarial loss of the first generator, is the adversarial loss of the second generator, is the consistency loss of the first generator, is the consistency loss of the second generator, is the image data reconstruction loss of the first generator. is the image data reconstruction loss of the second generator.
[0123] As Figure 5 shown, delving into the framework of the Nice-GAN model, its first feature encoder not only serves as one of the components of the first feature discriminator but also is a component of the first generator 。
[0124] During the training of the Nice GAN model, the training of the first feature encoder is decoupled from the training of the first feature generator 。This effectively avoids the inconsistency problem during the training process, and ultimately reflects in a significant reduction in the failure rate of model training.
[0125] The loss functions of the decoupled Nice GAN model are: 。
[0126] 。
[0127] 。
[0128] 。
[0129] In the formula, is the adversarial loss of the first generator, is the first generator, is the picture, is the converted picture, is the second feature discriminator, is the first feature encoding, is the number of pictures before conversion, is the second feature encoder, is the number of converted pictures, is the second generator, is the consistency loss of the first generator, is the image data reconstruction loss of the first generator.
[0130] Modify the in the loss function of the first generator to , Modify to , and the loss function of the second generator can be obtained.
[0131] The first and second terms in each loss function describe the calculation method of the adversarial loss of the first generator of the model. This adversarial loss is obtained based on the least squares method and is fixed while maximizing and is being trained. When minimizing , and and are both fixed.
[0132] The consistency loss is as shown in the third term of each loss function and is calculated from the L1 loss (i.e., the mean absolute error). This loss constrains the generation process of the model by calculating the consistency between the original image and the image after style transfer (two style transfers).
[0133] The image reconstruction loss is to calculate the difference between two images. For example, use the first generator to complete the style transfer of the image to the image . If it is necessary to prove whether the first generator has truly generated an image styled image, it can be done by taking the image The first generator is used for judgment because of the image should still possess the characteristics of the image style after passing through the first generator, and the difference between the image before passing through the first generator and after passing through the first generator is the image reconstruction loss. The specific calculation is shown in the fourth item of each loss function and is calculated based on the L1 loss. The difference between the image before passing through the first generator and after passing through the first generator is the image reconstruction loss. The specific calculation is shown in the fourth item of each loss function and is calculated based on the L1 loss. The difference between the image before passing through the first generator and after passing through the first generator is the image reconstruction loss. The specific calculation is shown in the fourth item of each loss function and is calculated based on the L1 loss.
[0134] A domain adaptive infrared target recognition method combining an improved teacher-student framework in the present invention aims to explore an efficient unsupervised domain adaptation strategy by studying the teacher-student framework and the two-stage detection model to enhance the performance of target detection. Using unsupervised domain adaptation technology to optimize the two-stage model, promoting the improvement of the model's cross-domain transfer ability, and enhancing the detection accuracy by learning general features. Compared with the existing mainstream domain adaptive infrared target recognition technologies, the present invention has the following differences: 1. The present invention is improved based on the Faster-RCNN detector on the basis of the teacher-student distillation framework, which replaces the past method of using an image reconstruction network for visible -> infrared image conversion. 2. Starting from the teacher-student framework, the present invention realizes the conversion from traditional hard-pseudo labels to soft pseudo labels in the form of probability distribution by probabilistically reconstructing the backbone network with a Gaussian distribution, effectively improving the model's ability to learn subtle features, and redesigning the adaptive weighted loss function. Thereby improving the model's multi-scale cross-domain recognition ability. 3. At the image level, the present invention uses a generative adversarial network to generate data similar to the target domain, so that the model can adapt to the data distribution of the target domain in both the supervised learning stage and the co-training stage; at the feature level, by adopting an adversarial training mechanism, combining a domain discriminator and a gradient reversal layer, continuously enhancing the model's ability to extract cross-domain invariant features.
[0135] Embodiment 2: The present invention provides a domain adaptive infrared target recognition device combining an improved teacher-student framework, which includes a training data acquisition module, a domain conversion module, a backbone network module, an image conversion module, a backbone training module, a model construction module, a pseudo label module, a training module, a to-be-recognized image acquisition module, and a recognition module.
[0136] The training data acquisition module is used to acquire labeled training data in the source domain and unlabeled training data in the target domain. Among them, the source domain is visible light data, and the target domain is infrared light data.
[0137] The domain conversion module is used to acquire a pre-trained Nice GAN model. Among them, the Nice GAN model is used to generate class target domain images with consistent labels according to the source domain images.
[0138] The backbone network module is used to obtain the Faster R-CNN model as the backbone network. Among them, the output result probability of Faster R-CNN is reconstructed as: class label cls, Gaussian distribution reg-a of the first axial coordinate, Gaussian distribution reg-b of the second axial coordinate, Gaussian distribution reg-w of the width, and Gaussian distribution reg-h of the height.
[0139] The image conversion module is used to input the labeled training data into the Nice GAN model to generate labeled class target domain data consistent with the original label.
[0140] The backbone training module is used to train the Faster R-CNN model using the labeled training data and the labeled class target domain data.
[0141] The model construction module is used to establish a teacher model and a student model with the trained Faster R-CNN model as a single backbone network.
[0142] The pseudo-label module is used to input the unlabeled training data into the teacher model to generate pseudo-labels.
[0143] The training module is used to combine the labeled training data, the labeled class target domain data, and the unlabeled training data with the pseudo-labels into a batch, input them into the student model to perform consistency training under uncertainty guidance, and obtain the trained student model.
[0144] The image to be recognized acquisition module is used to acquire the image to be recognized. Among them, the image to be recognized is a visible light image or an infrared light image.
[0145] The recognition module is used to input the image to be recognized into the trained student model to obtain the recognition information of the image to be recognized.
[0146] Embodiment 3: The present invention provides an infrared target recognition device combining an improved teacher-student framework for domain adaptation, which includes a processor, a memory, and a computer program stored in the memory. The computer program can be executed by the processor to implement an infrared target recognition method combining an improved teacher-student framework for domain adaptation as described in any paragraph of Embodiment 1.
[0147] Embodiment 4: The present invention provides a computer-readable storage medium. The computer-readable storage medium includes a stored computer program, wherein when the computer program runs, it controls the device where the computer-readable storage medium is located to execute an infrared target recognition method combining an improved teacher-student framework for domain adaptation as described in any paragraph of Embodiment 1.
[0148] In several embodiments provided by the embodiments of the present invention, it should be understood that the disclosed devices and methods can also be implemented in other ways. The device and method embodiments described above are merely illustrative. For example, the flowcharts and block diagrams in the accompanying drawings show the possible architectures, functions, and operations of devices, methods, and computer program products according to multiple embodiments of the present invention. In this regard, each block in the flowchart or block diagram may represent a module, a program segment, or a part of code, and the module, program segment, or part of code contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order than that marked in the accompanying drawings. For example, two consecutive blocks may actually be executed substantially in parallel, and they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram and / or flowchart, as well as the combination of blocks in the block diagram and / or flowchart, can be implemented by a dedicated hardware-based system for performing the specified functions or actions, or can be implemented by a combination of dedicated hardware and computer instructions.
[0149] In addition, each functional module in various embodiments of the present invention may be integrated together to form an independent part, or each module may exist alone, or two or more modules may be integrated to form an independent part.
[0150] If the described functions are implemented in the form of software functional modules and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or a part of this technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, an electronic device, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present invention. The aforementioned storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memories (ROMs), random access memories (RAMs), magnetic disks, or optical discs that can store program codes. It should be noted that in this article, the term "including", "comprising", or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article, or device including a series of elements not only includes those elements, but also includes other elements not explicitly listed, or further includes elements inherent to such process, method, article, or device. Without further limitations, an element defined by the statement "including one..." does not exclude the existence of additional identical elements in the process, method, article, or device including the said element.
[0151] The terms used in the embodiments of the present invention are for the purpose of describing specific embodiments only and are not intended to limit the present invention. The singular forms "a", "the" and "said" used in the embodiments of the present invention and the appended claims are also intended to include the plural forms unless the context clearly dictates otherwise.
[0152] It should be understood that the term "and / or" used herein is merely a description of the relationship between associated objects, indicating that there can be three relationships. For example, A and / or B can mean: A exists alone, A and B exist simultaneously, and B exists alone. In addition, the character " / " herein generally indicates that the associated objects before and after are in an "or" relationship.
[0153] Depending on the context, the word "if" as used herein can be interpreted as "when" or "while" or "in response to determining" or "in response to detecting". Similarly, depending on the context, the phrase "if determined" or "if detected (stated condition or event)" can be interpreted as "when determined" or "in response to determining" or "when detected (stated condition or event)" or "in response to detecting (stated condition or event)".
[0154] The "first / second" mentioned in the embodiments is only to distinguish similar objects and does not represent a specific order for the objects. It can be understood that the "first / second" can be interchanged in order or sequence under allowable circumstances. It should be understood that the objects distinguished by "first / second" can be interchanged appropriately so that the embodiments described herein can be implemented in an order other than those illustrated or described herein.
[0155] The foregoing are only the preferred embodiments of the present invention and are not used to limit the present invention. For those skilled in the art, the present invention can have various changes and modifications. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of the present invention shall be included within the protection scope of the present invention.
Claims
1. A domain-adaptive infrared target recognition method, characterized in that: Include: Obtain labeled training data of a source domain and unlabeled training data of a target domain; wherein the source domain is visible light data and the target domain is infrared light data; Obtain a pre-trained Nice GAN model; wherein the Nice GAN model is used to generate a target domain image with a consistent label according to a source domain image; Obtain a Faster R-CNN model as a backbone network; wherein the output result probability of Faster R-CNN is reconstructed into: category label cls, Gaussian distribution reg-a of the first axial coordinate, Gaussian distribution reg-b of the second axial coordinate, Gaussian distribution reg-w of the width, and Gaussian distribution reg-h of the height; Inputting the labeled training data into the Nice GAN model to generate labeled target domain data consistent with the original labels; Training the Faster R-CNN model using the labeled training data and the labeled target domain data; Use the trained Faster R-CNN model as a single backbone network to build a teacher model and a student model; Inputting the unlabeled training data into a teacher model to generate pseudo labels; The labeled training data, the labeled target domain data, and the unlabeled training data with the pseudo labels are combined into a batch, and input into the student model to perform consistency training under uncertainty guidance, so as to obtain a trained student model; Acquire an image to be identified; wherein the image to be identified is a visible light image or an infrared light image; The image to be recognized is input into the trained student model to obtain recognition information of the image to be recognized.
2. A domain-adaptive infrared target recognition method according to claim 1, characterized in that: The student model training also includes an adversarial training module connected to the student model; the adversarial training module includes a gradient reversal layer GRL, and a domain discriminator composed of a convolutional layer and a multi-layer perceptron; In the forward propagation process, the feature encoding extracted by the student model is transmitted to the domain discriminator through the gradient reversal layer GRL to identify whether the feature is from the source domain or the target domain; In this process, the gradient reversal layer GRL acts as an identity transformation layer and does not change the input data; In the back propagation process, the gradient generated by the classification loss of the domain discriminator first passes through the domain discriminator to complete the weight update of the domain discriminator; then it is passed to the student model through the gradient reversal layer GRL; in this process, the gradient reversal layer GRL reverses the gradient to maximize the loss of the backbone network learning domain discriminator of the student model; The negation operation is: ; In the formula, Represents differential, is the classification loss of the domain discriminator, is the output of the network layer before GRL, is a negative constant factor, is the output of GRL; Represents the gradient.
3. A domain-adaptive infrared target recognition method according to claim 2, characterized in that: The adversarial training module includes a gradient reversal layer GRL, two convolutional layers or one convolutional layer, and a multi-layer perceptron MLP; wherein the two convolutional layers or one convolutional layer and an MLP form a domain classifier; The classification loss of the domain discriminator during training is: ; in, is the classification loss of the domain discriminator, is the number of pictures, is the domain loss function, For the The output of the domain classifier for images, For the Feature encoding of an image.
4. A domain-adaptive infrared target recognition method according to claim 1, characterized in that: The total loss function during consistency training is: ; In the formula, is the total loss function, is an exponential function based on the base of natural logarithms, is the current training iteration number, is the upper limit of the number of iterations, The number of converted images, is the supervised source domain loss for model training on source domain data, The unsupervised target domain loss for models trained on target domain data using pseudo labels, is the classification loss of the domain discriminator, is the fourth hyperparameter.
5. A domain-adaptive infrared target recognition method according to any one of claims 1 to 4, characterized in that: During consistency training, the weights of the teacher model are updated based on the weights of the student model using an exponential moving average strategy.
6. A domain-adaptive infrared target recognition method according to any one of claims 1 to 4, characterized in that: The loss function of the Nice GAN model includes: ; ; ; ; In the formula, is the overall loss function of Nice GAN, For the first generator, For the second generator, is the first balancing hyperparameter, To combat losses, is the second balancing hyperparameter, For consistency loss, is the third balancing hyperparameter, Reconstruction loss for image data, is the adversarial loss of the first generator, is the adversarial loss of the second generator, is the consistency loss of the first generator, is the consistency loss of the second generator, Reconstruction loss of image data for the first generator; Reconstruction loss for the image data of the second generator.
7. A domain-adaptive infrared target recognition method according to claim 6, characterized in that: When training the Nice GAN model, the first feature encoder The training of the same feature first generator Decoupling of the training phase; The loss functions of the decoupled Nice GAN model are: ; ; ; ; In the formula, is the adversarial loss of the first generator, For the first generator, For pictures, For the converted image, is the second feature discriminator, Encode the first feature, is the number of images before conversion, is the second feature encoder, The number of converted images, For the second generator, is the consistency loss of the first generator, Reconstruction loss of image data for the first generator; The loss function of the first generator is Modified to , Modified to , we can get the loss function of the second generator.
8. A domain-adaptive infrared target recognition device combined with an improved teacher-student framework, characterized in that: Include: A training data acquisition module, used to acquire labeled training data of a source domain and unlabeled training data of a target domain; wherein the source domain is visible light data and the target domain is infrared light data; A domain conversion module, used to obtain a pre-trained Nice GAN model; wherein the Nice GAN model is used to generate a target domain image with consistent labels according to a source domain image; The backbone network module is used to obtain the Faster R-CNN model as the backbone network; wherein the output result probability of Faster R-CNN is reconstructed into: category label cls, Gaussian distribution reg-a of the first axial coordinate, Gaussian distribution reg-b of the second axial coordinate, Gaussian distribution reg-w of the width, and Gaussian distribution reg-h of the height; An image conversion module, used to input the labeled training data into the Nice GAN model to generate labeled target domain data consistent with the original labels; A backbone training module, configured to train the Faster R-CNN model using the labeled training data and the labeled target domain data; The model building module is used to build the teacher model and the student model using the trained Faster R-CNN model as a single backbone network; A pseudo label module, used for inputting the unlabeled training data into a teacher model to generate a pseudo label; A training module, used for combining the labeled training data, the labeled target domain data, and the unlabeled training data with the pseudo labels into a batch, inputting the batch into the student model to perform consistency training under uncertainty guidance, and obtaining a trained student model; An image acquisition module to be identified, used to acquire an image to be identified; wherein the image to be identified is a visible light image or an infrared light image; The recognition module is used to input the image to be recognized into the trained student model to obtain recognition information of the image to be recognized.
9. A domain-adaptive infrared target recognition device combined with an improved teacher-student framework, characterized in that: It comprises a processor, a memory, and a computer program stored in the memory; the computer program can be executed by the processor to implement a field-adaptive infrared target recognition method as described in any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that: The computer-readable storage medium includes a stored computer program, wherein when the computer program is executed, the device where the computer-readable storage medium is located is controlled to execute a field-adaptive infrared target recognition method as described in any one of claims 1 to 7.