Infrared target detection method and device for self-supervised learning, equipment and storage medium
By pre-training the infrared image training set using a self-supervised learning model and loading encoder weights into the infrared target detection model, the problems of scarce infrared image data and complex annotation are solved, improving the accuracy and robustness of infrared target detection and enhancing the model's detection capability in complex infrared scenes.
Patent Information
- Application Number
- CN202510945019.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-09
- Publication Date
- 2025-11-14
AI Technical Summary
In existing technologies, pre-trained models trained on visible light data are difficult to adapt to the characteristics of infrared images. Infrared image data is scarce and its annotation is complex, resulting in poor infrared target detection performance.
A masked autoencoder (MAE) architecture based on a self-supervised learning model is used to pre-train the infrared image training set to obtain encoder weights, which are then loaded into the feature extraction network for target detection. The optimized self-supervised learning model and infrared target detection model are combined for feature extraction and target detection.
It significantly improves the accuracy and robustness of infrared target detection, especially performing well in low-resolution, high-noise infrared images. It narrows the inter-domain differences between visible light and infrared images, and enhances the model's detection and generalization capabilities in complex infrared scenes.
Smart Images

Figure CN120953574A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of computer vision technology, and in particular to a self-supervised learning method, apparatus, device, and storage medium for infrared target detection. Background Technology
[0002] Existing deep learning-based object detection models typically rely on large-scale labeled datasets for training. Widely used standard datasets (such as MS COCO and PASCAL VOC) contain target data derived from visible light images. These datasets possess rich color, texture, and contextual information, providing high-quality feature representations for the models. However, infrared images differ significantly from visible light images in their imaging principles and feature distributions: infrared images, based on the thermal radiation characteristics of the target, typically exhibit lower resolution, more noise, and less distinct contrast between the target and background. Therefore, when directly using pre-trained models trained on visible light data (such as You Only Look Once (YOLO) and Detection Transformer (DETR)) for infrared target detection tasks, the models struggle to fully adapt to the characteristics of infrared images, resulting in poor detection performance.
[0003] Furthermore, the field of infrared target detection has long lacked dedicated, high-quality datasets. Existing datasets are generally small in size, and most lack accurate bounding box annotations required for model training. Their sample size and annotation quality are still insufficient to meet the needs of deep learning models for large-scale labeled data. This data scarcity severely limits the development of deep learning-based infrared target detection research.
[0004] Self-supervised learning methods (such as masked autoencoders (MAE)) have made significant progress in feature learning tasks for visible light images. By randomly masking image regions and reconstructing the masked parts, MAE can learn robust feature representations from unlabeled data, reducing the dependence on large-scale labeled data. However, the application of self-supervised learning methods in the field of infrared target detection has not been fully explored.
[0005] The existing technology has the following problems: 1. Existing pre-trained models for object detection are typically trained on large-scale visible light labeled datasets. However, visible light images and infrared images differ significantly in their imaging principles and feature distributions. Visible light images rely on ambient light reflection and possess rich color and texture information, while infrared images are based on the thermal radiation characteristics of the target and typically exhibit lower resolution, more noise, and less contrast between the target and the background. Therefore, directly using pre-trained models trained on visible light data to train infrared object detection tasks will make it difficult for the model to fully adapt to the characteristics of infrared images, thus limiting the detection performance. This inter-domain difference makes the model's performance in infrared object detection tasks far inferior to its performance in visible light scenes.
[0006] 2. The acquisition cost of infrared images is high and the annotation process is complex, resulting in the generally small size of existing infrared datasets. In addition, most existing infrared datasets lack high-quality bounding box annotations, which are a key element in training target detection models. The problems of insufficient data and incomplete annotation severely limit the performance of models in infrared target detection tasks, making it difficult for models to fully learn the feature representation of targets in infrared images, thus affecting detection accuracy and robustness. Summary of the Invention
[0007] The main objective of this invention is to provide a self-supervised learning infrared target detection method, apparatus, device, and storage medium, aiming to solve the technical problems in the prior art where models are difficult to fully adapt to the characteristics of infrared images, thus limiting the detection effect; the high cost of acquiring infrared images and the complex annotation process; and the limitations of infrared target detection performance caused by insufficient data and incomplete annotation, making it difficult for models to fully learn the feature representation of targets in external images, thus affecting detection accuracy and robustness.
[0008] In a first aspect, the present invention provides a self-supervised learning method for infrared target detection, the self-supervised learning method for infrared target detection comprising the following steps: The infrared image training set is pre-trained based on the masked autoencoder (MAE) architecture of the self-supervised learning model to obtain the pre-trained encoder weights. The encoder weights are loaded into the feature extraction network for target detection to construct an infrared target detection model; The self-supervised learning model and the infrared target detection model are optimized. The optimized self-supervised learning model and the optimized infrared target detection model are used to extract features and detect targets in the input infrared image to obtain detection results.
[0009] Optionally, before pre-training the infrared images in the training set according to the masked autoencoder (MAE) architecture of the self-supervised learning model to obtain the pre-trained encoder weights, the self-supervised learning infrared target detection method further includes: A preset number of infrared images are collected, and each infrared image is preprocessed to obtain an infrared image training set corresponding to the preprocessed infrared training images.
[0010] Optionally, the step of acquiring a preset number of infrared images, preprocessing each infrared image to obtain an infrared image training set corresponding to the preprocessed infrared training images includes: A preset number of infrared images are collected, and each infrared image is enhanced and normalized to obtain preprocessed infrared training images. An infrared image training set is then generated based on each infrared training image.
[0011] Optionally, the step of pre-training the infrared image training set based on the masked autoencoder (MAE) architecture of the self-supervised learning model to obtain the pre-trained encoder weights includes: The core encoder of the self-supervised learning model's masked autoencoder (MAE) architecture randomly masks image regions within a preset range of the infrared image training set. The self-supervised learning model is controlled to reconstruct the masked infrared image to obtain the visual representation of the infrared image. The visual representation of the infrared image is stored in the encoder weights of the core encoder to obtain the pre-trained encoder weights.
[0012] Optionally, loading the encoder weights into the feature extraction network for target detection to construct an infrared target detection model includes: The encoder weights are loaded into the feature extraction network for object detection, while the detection head for object detection is retained. The thermal morphological characteristic parameters of infrared targets are obtained from the infrared image training set, and the preset anchor box for target detection is optimized based on the thermal morphological characteristic parameters. The optimized anchor box is integrated into the detection head, and an infrared target detection model is constructed by combining it with a preset target detection loss function. The bounding box and class probability of the infrared target are output through forward inference.
[0013] Optionally, optimizing the self-supervised learning model and the infrared target detection model, and using the optimized self-supervised learning model and the optimized infrared target detection model to extract features and detect targets in the input infrared image to obtain detection results, includes: By dynamically adjusting the masking strategy and adopting a multi-scale loss function, the encoder in the MAE framework of the self-supervised learning model learns the physical feature representation of the unlabeled infrared training data, thus obtaining an optimized self-supervised learning model. The encoder in the optimized self-supervised learning model is used as the backbone of the target detection. The backbone is fine-tuned using labeled infrared target detection training data. The infrared target detection model corresponding to the fine-tuned backbone is then trained end-to-end using a complete infrared dataset to obtain the optimized infrared target detection model. Based on the optimized self-supervised learning model and the optimized infrared target detection model, feature extraction and target detection are performed on the input infrared image to obtain the detection results.
[0014] Optionally, the step of performing feature extraction and target detection on the input infrared image based on the optimized self-supervised learning model and the optimized infrared target detection model to obtain detection results includes: The encoder of the optimized self-supervised learning model extracts the image features of the input infrared image; The target bounding box and target category probability are generated based on the detection head of the optimized infrared target detection model, and a detection box is constructed based on the target bounding box and the target category probability. Non-maximum suppression (NMS) is used to remove redundant detection boxes in the detection boxes to obtain the final detection boxes. The image features are then detected based on the final detection boxes to obtain the detection results.
[0015] Secondly, to achieve the above objectives, the present invention also proposes a self-supervised learning infrared target detection device, the self-supervised learning infrared target detection device comprising: The pre-training module is used to pre-train the infrared image training set based on the masked autoencoder (MAE) architecture of the self-supervised learning model to obtain the pre-trained encoder weights. The model building module is used to load the encoder weights into the feature extraction network for target detection to build an infrared target detection model. The optimization detection module is used to optimize the self-supervised learning model and the infrared target detection model. The optimized self-supervised learning model and the optimized infrared target detection model are used to extract features and detect targets in the input infrared image to obtain detection results.
[0016] Thirdly, to achieve the above objectives, the present invention also proposes a self-supervised learning infrared target detection device, the self-supervised learning infrared target detection device comprising: a memory, a processor, and a self-supervised learning infrared target detection program stored in the memory and executable on the processor, the self-supervised learning infrared target detection program being configured to implement the steps of the self-supervised learning infrared target detection method as described above.
[0017] Fourthly, to achieve the above objectives, the present invention also proposes a storage medium storing a self-supervised learning infrared target detection program, wherein when the self-supervised learning infrared target detection program is executed by a processor, it implements the steps of the self-supervised learning infrared target detection method described above.
[0018] This invention proposes a self-supervised learning-based infrared target detection method. It pre-trains an infrared image training set using a masked autoencoder (MAE) architecture of a self-supervised learning model to obtain pre-trained encoder weights. These encoder weights are then loaded into a feature extraction network for target detection, constructing an infrared target detection model. The self-supervised learning model and the infrared target detection model are optimized. The optimized self-supervised learning model and the optimized infrared target detection model are then used to extract features and detect targets from the input infrared image, obtaining detection results. This method utilizes unlabeled infrared images for pre-training through self-supervised representation learning, reducing the need for large-scale labeled data. Pre-trained models, replacing those trained on large-scale visible light, extract more robust infrared image features, enhancing the model's detection capabilities in complex infrared scenes. This addresses the data scarcity issue in the field of infrared target detection, integrating MAE's robust feature extraction capabilities into an efficient target detection framework. This significantly improves the accuracy and robustness of infrared target detection, especially performing exceptionally well in low-resolution, high-noise infrared images. It narrows the inter-domain differences between visible light and infrared images, enabling the model to better adapt to infrared scenes. In complex and varied infrared scenes (such as nighttime, fog, and smoke), the model exhibits stronger generalization ability and stability, improving the speed and efficiency of self-supervised infrared target detection. Attached Figure Description
[0019] Figure 1 This is a schematic diagram of the device structure of the hardware operating environment involved in the embodiments of the present invention; Figure 2 This is a flowchart illustrating the first embodiment of the self-supervised learning infrared target detection method of the present invention; Figure 3 This is a flowchart illustrating the second embodiment of the self-supervised learning infrared target detection method of the present invention; Figure 4 This is a flowchart illustrating the third embodiment of the self-supervised learning infrared target detection method of the present invention; Figure 5 This is a flowchart illustrating the fourth embodiment of the self-supervised learning infrared target detection method of the present invention; Figure 6 This is a flowchart illustrating the fifth embodiment of the self-supervised learning infrared target detection method of the present invention; Figure 7This is a schematic diagram of the self-supervised training and target detection structure in the self-supervised learning infrared target detection method of the present invention. Figure 8 This is a functional block diagram of the first embodiment of the self-supervised learning infrared target detection device of the present invention.
[0020] The realization of the objective, functional features and advantages of the present invention will be further explained in conjunction with the embodiments and with reference to the accompanying drawings. Detailed Implementation
[0021] It should be understood that the specific embodiments described herein are merely illustrative of the invention and are not intended to limit the invention.
[0022] The solution of this invention mainly involves: pre-training an infrared image training set using a masked autoencoder (MAE) architecture of a self-supervised learning model to obtain pre-trained encoder weights; loading these encoder weights into a feature extraction network for target detection to construct an infrared target detection model; optimizing the self-supervised learning model and the infrared target detection model; and using the optimized self-supervised learning model and the optimized infrared target detection model to extract features and detect targets in the input infrared image to obtain detection results. This approach leverages self-supervised representation learning to pre-train using unlabeled infrared images, reducing the need for large-scale labeled data. The infrared pre-trained model trained through self-supervised learning replaces the pre-trained model trained on large-scale visible light, extracting more robust infrared image features and enhancing the model's detection capability in complex infrared scenes, thus solving the infrared... Addressing the data scarcity issue in target detection, this paper integrates MAE's robust feature extraction capabilities into an efficient target detection framework, significantly improving the accuracy and robustness of infrared target detection. It excels particularly in low-resolution, high-noise infrared images, narrowing the inter-domain differences between visible and infrared images. This allows the model to better adapt to infrared scenes, exhibiting stronger generalization ability and stability in complex and varied infrared scenarios (such as nighttime, fog, and smoke). It also improves the speed and efficiency of self-supervised infrared target detection, resolving the limitations of existing technologies where models struggle to fully adapt to the characteristics of infrared images, thus restricting detection performance. Furthermore, the high cost and complex annotation process of infrared image acquisition, coupled with insufficient data and incomplete annotation, limit infrared target detection performance, making it difficult for the model to fully learn the feature representations of targets in external images, thus impacting detection accuracy and robustness.
[0023] Reference Figure 1 , Figure 1 This is a schematic diagram of the device structure of the hardware operating environment involved in the embodiments of the present invention.
[0024] like Figure 1As shown, the device may include: a processor 1001, such as a CPU; a communication bus 1002; a user interface 1003; a network interface 1004; and a memory 1005. The communication bus 1002 is used to enable communication between these components. The user interface 1003 may include a display screen or an input unit such as a keyboard; optionally, the user interface 1003 may also include a standard wired interface or a wireless interface. The network interface 1004 may optionally include a standard wired interface or a wireless interface (such as a Wi-Fi interface). The memory 1005 may be high-speed RAM or non-volatile memory, such as a disk drive. Optionally, the memory 1005 may also be a storage device independent of the aforementioned processor 1001.
[0025] Those skilled in the art will understand that Figure 1 The device structure shown does not constitute a limitation on the device and may include more or fewer components than shown, or combine certain components, or have different component arrangements.
[0026] like Figure 1 As shown, the memory 1005, which serves as a storage medium, may include an operating device, a network communication module, a user interface module, and a self-supervised learning infrared target detection program.
[0027] The device of the present invention calls the self-supervised learning infrared target detection program stored in the memory 1005 through the processor 1001, and performs the following operations: The infrared image training set is pre-trained based on the masked autoencoder (MAE) architecture of the self-supervised learning model to obtain the pre-trained encoder weights. The encoder weights are loaded into the feature extraction network for target detection to construct an infrared target detection model; The self-supervised learning model and the infrared target detection model are optimized. The optimized self-supervised learning model and the optimized infrared target detection model are used to extract features and detect targets in the input infrared image to obtain detection results.
[0028] The device of the present invention, through processor 1001 calling the self-supervised learning infrared target detection program stored in memory 1005, also performs the following operations: A preset number of infrared images are collected, and each infrared image is preprocessed to obtain an infrared image training set corresponding to the preprocessed infrared training images.
[0029] The device of the present invention, through processor 1001 calling the self-supervised learning infrared target detection program stored in memory 1005, also performs the following operations: A preset number of infrared images are collected, and each infrared image is enhanced and normalized to obtain preprocessed infrared training images. An infrared image training set is then generated based on each infrared training image.
[0030] The device of the present invention, through processor 1001 calling the self-supervised learning infrared target detection program stored in memory 1005, also performs the following operations: The core encoder of the self-supervised learning model's masked autoencoder (MAE) architecture randomly masks image regions within a preset range of the infrared image training set. The self-supervised learning model is controlled to reconstruct the masked infrared image to obtain the visual representation of the infrared image. The visual representation of the infrared image is stored in the encoder weights of the core encoder to obtain the pre-trained encoder weights.
[0031] The device of the present invention, through processor 1001 calling the self-supervised learning infrared target detection program stored in memory 1005, also performs the following operations: The encoder weights are loaded into the feature extraction network for object detection, while the detection head for object detection is retained. The thermal morphological characteristic parameters of infrared targets are obtained from the infrared image training set, and the preset anchor box for target detection is optimized based on the thermal morphological characteristic parameters. The optimized anchor box is integrated into the detection head, and an infrared target detection model is constructed by combining it with a preset target detection loss function. The bounding box and class probability of the infrared target are output through forward inference.
[0032] The device of the present invention, through processor 1001 calling the self-supervised learning infrared target detection program stored in memory 1005, also performs the following operations: By dynamically adjusting the masking strategy and adopting a multi-scale loss function, the encoder in the MAE framework of the self-supervised learning model learns the physical feature representation of the unlabeled infrared training data, thus obtaining an optimized self-supervised learning model. The encoder in the optimized self-supervised learning model is used as the backbone of the target detection. The backbone is fine-tuned using labeled infrared target detection training data. The infrared target detection model corresponding to the fine-tuned backbone is then trained end-to-end using a complete infrared dataset to obtain the optimized infrared target detection model. Based on the optimized self-supervised learning model and the optimized infrared target detection model, feature extraction and target detection are performed on the input infrared image to obtain the detection results.
[0033] The device of the present invention, through processor 1001 calling the self-supervised learning infrared target detection program stored in memory 1005, also performs the following operations: The encoder of the optimized self-supervised learning model extracts the image features of the input infrared image; The target bounding box and target category probability are generated based on the detection head of the optimized infrared target detection model, and a detection box is constructed based on the target bounding box and the target category probability. Non-maximum suppression (NMS) is used to remove redundant detection boxes in the detection boxes to obtain the final detection boxes. The image features are then detected based on the final detection boxes to obtain the detection results.
[0034] This embodiment, through the above scheme, pre-trains the infrared image training set using a masked autoencoder (MAE) architecture of a self-supervised learning model to obtain pre-trained encoder weights; loads these encoder weights into a feature extraction network for target detection to construct an infrared target detection model; optimizes both the self-supervised learning model and the infrared target detection model, and uses the optimized self-supervised learning model and optimized infrared target detection model to perform feature extraction and target detection on the input infrared image to obtain detection results. This approach enables pre-training using unlabeled infrared images through self-supervised representation learning, reducing the need for large-scale labeled data. The infrared pre-trained model trained through self-supervised learning... This approach replaces pre-trained models trained on large-scale visible light, extracts more robust infrared image features, enhances the model's detection capabilities in complex infrared scenes, solves the problem of data scarcity in the field of infrared target detection, integrates MAE's robust feature extraction capabilities into an efficient target detection framework, significantly improves the accuracy and robustness of infrared target detection, especially performing well in low-resolution, high-noise infrared images, narrows the inter-domain differences between visible light and infrared images, and enables the model to better adapt to infrared scenes. In complex and variable infrared scenes (such as nighttime, foggy days, and smoke), the model exhibits stronger generalization ability and stability, improving the speed and efficiency of self-supervised learning for infrared target detection.
[0035] Based on the above hardware structure, an embodiment of the self-supervised learning infrared target detection method of the present invention is proposed.
[0036] Reference Figure 2 , Figure 2 This is a flowchart illustrating the first embodiment of the self-supervised learning infrared target detection method of the present invention.
[0037] In the first embodiment, the self-supervised learning infrared target detection method includes the following steps: Step S10: Pre-train the infrared image training set according to the masked autoencoder (MAE) architecture of the self-supervised learning model to obtain the pre-trained encoder weights.
[0038] It should be noted that by pre-training the infrared image training set using the Masked Autoencoder (MAE) architecture of the self-supervised learning model, the pre-trained encoder weights can be obtained. The MAE consists of two parts: the encoder, which encodes the visible patch into a low-dimensional feature vector, and the decoder, which reconstructs the masked patch based on the encoded features. The patch is a fixed-size block or region into which the input image is segmented.
[0039] Step S20: Load the encoder weights into the feature extraction network for target detection to construct an infrared target detection model.
[0040] It should be understood that by loading the encoder weights obtained after pre-training into the feature extraction network for target detection, an infrared target detection model can be constructed.
[0041] Step S30: Optimize the self-supervised learning model and the infrared target detection model. Use the optimized self-supervised learning model and the optimized infrared target detection model to extract features and detect targets in the input infrared image to obtain detection results.
[0042] It is understood that by optimizing the self-supervised learning model and the infrared target detection model, an optimized self-supervised learning model and an optimized infrared target detection model can be obtained. Then, based on the optimized self-supervised learning model and the optimized infrared target detection model, feature extraction and target detection are performed on the input infrared image to obtain detection results and generate the detection results after network feature extraction and target detection.
[0043] This embodiment, through the above scheme, pre-trains the infrared image training set using a masked autoencoder (MAE) architecture of a self-supervised learning model to obtain pre-trained encoder weights; loads these encoder weights into a feature extraction network for target detection to construct an infrared target detection model; optimizes both the self-supervised learning model and the infrared target detection model, and uses the optimized self-supervised learning model and optimized infrared target detection model to perform feature extraction and target detection on the input infrared image to obtain detection results. This approach enables pre-training using unlabeled infrared images through self-supervised representation learning, reducing the need for large-scale labeled data. The infrared pre-trained model trained through self-supervised learning... This approach replaces pre-trained models trained on large-scale visible light, extracts more robust infrared image features, enhances the model's detection capabilities in complex infrared scenes, solves the problem of data scarcity in the field of infrared target detection, integrates MAE's robust feature extraction capabilities into an efficient target detection framework, significantly improves the accuracy and robustness of infrared target detection, especially performing well in low-resolution, high-noise infrared images, narrows the inter-domain differences between visible light and infrared images, and enables the model to better adapt to infrared scenes. In complex and variable infrared scenes (such as nighttime, foggy days, and smoke), the model exhibits stronger generalization ability and stability, improving the speed and efficiency of self-supervised learning for infrared target detection.
[0044] Furthermore, Figure 3 This is a flowchart illustrating the second embodiment of the self-supervised learning infrared target detection method of the present invention, as shown below. Figure 3 As shown, based on the first embodiment, a second embodiment of the self-supervised learning infrared target detection method of the present invention is proposed. In this embodiment, before step S10, the self-supervised learning infrared target detection method further includes the following steps: Step S01: Collect a preset number of infrared images, preprocess each infrared image, and obtain the infrared image training set corresponding to the preprocessed infrared training images.
[0045] It should be noted that by collecting a preset number of large-scale infrared images and performing preprocessing operations on the infrared images, preprocessed infrared training images can be obtained, and then a corresponding infrared image training set can be generated based on the infrared training images.
[0046] Furthermore, step S01 specifically includes the following steps: A preset number of infrared images are collected, and each infrared image is enhanced and normalized to obtain preprocessed infrared training images. An infrared image training set is then generated based on each infrared training image.
[0047] It should be understood that after image enhancement and normalization of each infrared image, preprocessed infrared training images can be obtained, which in turn generate an infrared image training set.
[0048] In the specific implementation, the infrared image data integrates mainstream public infrared datasets (such as FLIR ADAS, M3FD and other multispectral datasets) with data collected by autonomous equipment to ensure the diversity and richness of data samples. Image preprocessing uses data augmentation techniques such as random cropping, rotation, flipping, and brightness adjustment to enhance the diversity of data; normalizes the image pixel values to facilitate model training; and can generate random mask regions to provide input for self-supervised pre-training tasks.
[0049] This embodiment, through the above-described scheme, acquires a preset number of infrared images, preprocesses each infrared image, and obtains an infrared image training set corresponding to the preprocessed infrared training images. This reduces the need for large-scale labeled data and improves the speed and efficiency of self-supervised learning for infrared target detection.
[0050] Furthermore, Figure 4 This is a flowchart illustrating the third embodiment of the self-supervised learning infrared target detection method of the present invention, as shown below. Figure 4 As shown, based on the first embodiment, a third embodiment of the self-supervised learning infrared target detection method of the present invention is proposed. In this embodiment, step S10 specifically includes the following steps: Step S11: Randomly mask the image region of the infrared image training set within a preset range according to the core encoder of the mask autoencoder (MAE) architecture of the self-supervised learning model.
[0051] It should be noted that the core encoder of the mask autoencoder (MAE) architecture of the self-supervised learning model can randomly mask image regions of the infrared image training set within a preset range. Through the masking mechanism, the model is forced to learn the most essential and robust visual features in the image (such as object outlines, textures, and thermal radiation distribution), rather than surface details.
[0052] Step S12: Control the self-supervised learning model to reconstruct the masked infrared image, obtain the visual representation of the infrared image, store the visual representation of the infrared image in the encoder weights of the core encoder, and obtain the pre-trained encoder weights.
[0053] It is understood that controlling the self-supervised learning model to reconstruct the masked infrared image can obtain the visual representation of the infrared image, and then store the visual representation of the infrared image in the encoder weights of the core encoder to obtain the pre-trained encoder weights. The parameters (i.e., weights) learned by the encoder in the reconstruction task represent the model's understanding of the infrared image. These weights can be used as "knowledge capsules" and transferred to other tasks (such as object detection and classification) to avoid training the model from scratch. That is, the pre-trained encoder weights can be directly used to initialize the model for downstream tasks, so that the model does not need to learn basic visual concepts from scratch. The visual representation is the encoder's abstract representation of the image (such as feature vectors). It captures the semantic information of the image. These representations are stored in the encoder weights and can be directly transferred to other tasks, which greatly improves the model performance and training efficiency.
[0054] In a specific implementation, a masked autoencoder (MAE) architecture can be used, with Vision Transformer (ViT) as the core encoder. By randomly masking 60%-80% of the training set image regions and forcing the model to reconstruct the masked infrared image, the model is forced to learn a robust visual representation of the infrared image through this "mask-reconstruction" mechanism, thus achieving self-supervised representation learning. After pre-training is completed, the pre-trained encoder weights are saved.
[0055] This embodiment, through the above-described scheme, randomly masks image regions within a preset range of the infrared image training set using the core encoder of the self-supervised learning model's masked autoencoder (MAE) architecture. It then controls the self-supervised learning model to reconstruct the masked infrared image, obtaining a visual representation of the infrared image. This visual representation is stored in the encoder weights of the core encoder, resulting in pre-trained encoder weights. This allows the infrared pre-trained model trained through self-supervised learning to replace the pre-trained model trained on a large scale using visible light, extracting more robust infrared image features and enhancing the model's detection capabilities in complex infrared scenes. This solves the problem of data scarcity in the field of infrared target detection and improves the speed and efficiency of self-supervised infrared target detection.
[0056] Furthermore, Figure 5 This is a flowchart illustrating the fourth embodiment of the self-supervised learning infrared target detection method of the present invention, as shown below. Figure 5 As shown, based on the first embodiment, a fourth embodiment of the self-supervised learning infrared target detection method of the present invention is proposed. In this embodiment, step S20 specifically includes the following steps: Step S21: Load the encoder weights into the feature extraction network for target detection, and retain the target detection head.
[0057] It should be noted that the encoder weights include the ability to represent low-level image features (edges, textures, thermal radiation patterns) and high-level semantics (object contours, temperature distribution). By loading the encoder weights into the feature extraction network for object detection and retaining the detection head for object detection, the object detection model can obtain stronger object recognition and localization capabilities with less data and in a shorter time.
[0058] In the specific implementation, taking YOLOv5 as an example, after loading the MAE pre-trained CSPDarknet encoder (the first 5 layers are frozen, and the last 3 layers can be fine-tuned), the feature extraction network inserts a thermal feature attention module (located between the C3 modules). Through the neck network: a PAN structure is constructed to fuse features at different scales; a temperature activation layer (such as ReLU6, which limits the response in low-temperature regions) is added to each scale feature; the detection head retains the original YOLOv5 detection head structure, but the anchor box size is adjusted (based on infrared target statistics); the output dimension of the classification branch is changed to the number of infrared scene categories (such as 2 categories: pedestrians and vehicles).
[0059] Step S22: Obtain the thermal morphological characteristic parameters of the infrared target from the infrared image training set, and optimize the preset anchor frame for target detection based on the thermal morphological characteristic parameters.
[0060] It is understood that after obtaining the thermal morphological characteristic parameters of infrared targets from the infrared image training set, the preset anchor frame for target detection can be optimized based on the thermal morphological characteristic parameters.
[0061] Step S23: Integrate the optimized anchor box into the detection head, construct an infrared target detection model by combining it with a preset target detection loss function, and output the bounding box and class probability of the infrared target through forward inference.
[0062] It should be understood that loading the pre-trained encoder weights of MAE into the feature extraction network for object detection, replacing the traditional pre-trained object detection weights, is to better adapt to the characteristics of infrared images. Simultaneously, the object detection head is retained to predict the bounding box and class probability of the target, constructing a complete infrared object detection model. This involves integrating the optimized anchor boxes into the detection head, which can be combined with a preset object detection loss function to build the infrared object detection model, and outputting the bounding box and class probability of the infrared target through forward inference.
[0063] In practice, thermal morphological parameters of the target (such as the average size of the thermal radiation region, aspect ratio, and temperature gradient distribution) can be obtained from the infrared dataset. Based on these parameters, the preset anchor boxes for target detection can be optimized (e.g., adjusting the size ratio and density distribution of the anchor boxes, or introducing temperature-related anchor box generation rules). Subsequently, the optimized anchor boxes are integrated into the detection head, and a complete model is constructed by combining thermal target detection loss functions (such as temperature-weighted classification loss and thermal shape constraint regression loss). After the model is trained, the bounding box coordinates and class probabilities of the target are output through forward inference.
[0064] This embodiment, through the above-described scheme, loads the encoder weights into the feature extraction network for target detection while retaining the target detection head; obtains the thermal morphological characteristic parameters of infrared targets from the infrared image training set, optimizes the preset anchor boxes for target detection based on the thermal morphological characteristic parameters; integrates the optimized anchor boxes into the detection head, constructs an infrared target detection model by combining it with a preset target detection loss function, and outputs the bounding boxes and class probabilities of the infrared targets through forward inference; it can integrate the robust feature extraction capability of MAE into an efficient target detection framework, significantly improving the accuracy and robustness of infrared target detection, especially performing well in low-resolution, high-noise infrared images, reducing the inter-domain difference between visible light and infrared images, enabling the model to better adapt to infrared scenes, and exhibiting stronger generalization ability and stability in complex and variable infrared scenes (such as nighttime, foggy days, smoke, etc.), improving the speed and efficiency of self-supervised learning infrared target detection.
[0065] Furthermore, Figure 6 This is a flowchart illustrating the fifth embodiment of the self-supervised learning infrared target detection method of the present invention, as shown below. Figure 6 As shown, based on the first embodiment, a fifth embodiment of the self-supervised learning infrared target detection method of the present invention is proposed. In this embodiment, step S30 specifically includes the following steps: Step S31: By dynamically adjusting the masking strategy and adopting a multi-scale loss function, the encoder in the MAE framework of the self-supervised learning model learns the physical feature representation of the unlabeled infrared training data, thereby obtaining the optimized self-supervised learning model.
[0066] It should be noted that the masking strategy can be dynamically adjusted to adapt to the thermal characteristic masking mechanism of infrared images. For example, dynamic masking based on temperature gradient: by calculating the temperature gradient map of the infrared image (reflecting the rate of change of thermal radiation), the mask ratio is reduced (or the whole block is retained) in areas with high temperature gradient (such as the edge of a hot target, the boundary between hot and cold), and the mask ratio is increased in areas with low temperature gradient (such as a uniform hot area in the background).
[0067] Understandably, the shape and scale of the mask can be adaptive, that is, the shape of the mask block can be dynamically adjusted according to the thermal morphology of the infrared target (such as point heat source, area heat source). For example, a circular mask can be used for a circular thermal target (such as an engine heat vent), and a bar mask can be used for a linear thermal target (such as a pipe), so that the mask area is more in line with the distribution pattern of infrared physical features. By masking on demand (preserving areas with high temperature gradients and destroying uniform thermal areas), the encoder is forced to extract the most critical features for thermal target detection (such as hot edges and hot spot positions) in the non-masked area, avoiding the learning of visual textures (such as colors and textures in visible light images) that are unrelated to infrared physics.
[0068] It should be understood that the physical features of infrared images contain multi-scale information: macroscopic scale: the overall shape of the thermal target (such as the thermal profile of a vehicle); mesoscopic scale: the distribution of thermal components (such as the hot zones of an engine or exhaust pipe); microscopic scale: subtle changes in temperature gradient (such as the attenuation curve of thermal radiation). Through multi-level supervision, the encoder's feature representation can simultaneously satisfy "accurate microscopic temperature changes, reasonable mesoscopic hot zone distribution, and identifiable macroscopic thermal morphology".
[0069] In its implementation, the proportion, shape, and scale of the mask region are adaptively changed based on the infrared image temperature gradient and the morphology of the thermal target, enabling the encoder to focus on key physical features such as the edges and hot spots of the thermal target. The multi-scale loss function calculates edge, semantic, and pixel-level losses on different levels of feature maps of the encoder and decoder, and assigns higher weights to high-temperature regions, thereby constraining the model to learn full-scale physical features of micro-temperature changes, meso-scale hot zone distribution, and macro-scale thermal morphology. The two work together to enable the encoder to output feature representations with infrared physical interpretability, providing higher-quality semantic features for subsequent target detection tasks.
[0070] Step S32: Use the encoder in the optimized self-supervised learning model as the backbone of the target detection, fine-tune the backbone using labeled infrared target detection training data, and perform data augmentation end-to-end training on the infrared target detection model corresponding to the fine-tuned backbone using the complete infrared dataset to obtain the optimized infrared target detection model.
[0071] It is understandable that by using the encoder in the optimized self-supervised learning model as the backbone of object detection and fine-tuning it, the infrared object detection model corresponding to the fine-tuned backbone can be trained end-to-end using the complete infrared dataset to obtain the optimized infrared object detection model.
[0072] In the specific implementation, see Figure 7 , Figure 7This is a schematic diagram of the self-supervised training and target detection structure in the self-supervised learning infrared target detection method of the present invention, as shown below. Figure 7 As shown, this embodiment employs a two-stage progressive training strategy, organically combining self-supervised pre-training with fine-tuning of object detection. In the first stage, masked image modeling is performed on a large-scale, unlabeled infrared dataset using the MAE framework. By dynamically adjusting the masking strategy and multi-scale loss function, the encoder learns the physical feature representation of infrared images. In the second stage, the backbone of object detection is fine-tuned on a small amount of labeled infrared data, followed by end-to-end training on the complete dataset. During training, commonly used data augmentation techniques for object detection, such as mosaic, pixel enhancement, and geometric enhancement, are incorporated to optimize the model's performance in infrared object detection tasks.
[0073] Understandably, this embodiment utilizes unlabeled infrared images for pre-training through MAE's self-supervised learning, significantly reducing the need for large-scale labeled data and solving the problem of data scarcity in the field of infrared target detection. By integrating MAE's robust feature extraction capabilities into an efficient target detection framework, the accuracy and robustness of infrared target detection are significantly improved, especially in low-resolution, high-noise infrared images. By extracting features from infrared images through self-supervised learning, the inter-domain differences between visible light and infrared images are reduced, enabling the model to better adapt to infrared scenes. In complex and variable infrared scenes (such as nighttime, foggy days, and smoke), the model exhibits stronger generalization ability and stability.
[0074] It should be understood that the encoder is trained on a large-scale unlabeled infrared dataset using self-supervised methods such as masked autoencoders (MAEs) to learn the general feature representations of infrared images. Then, the trained encoder is directly transferred to the object detection model as the backbone network, replacing the initial backbone structure of the original model. Subsequently, some layer parameters of the backbone network are fine-tuned on a small amount of labeled infrared data to allow the model to initially adapt to the object detection task. Finally, the entire object detection model (including the backbone network, neck network, and detection head) is trained end-to-end using a complete labeled infrared dataset. By optimizing the object detection loss function and comprehensively adjusting the model parameters, the model can accurately extract infrared target features, predict the target's bounding box and class probability, and ultimately achieve efficient infrared target detection.
[0075] Step S33: Based on the optimized self-supervised learning model and the optimized infrared target detection model, perform feature extraction and target detection on the input infrared image to obtain the detection results.
[0076] It should be understood that by using the optimized self-supervised learning model and the optimized infrared target detection model to extract features and detect targets from the input infrared image, the corresponding detection results can be obtained.
[0077] Furthermore, step S33 specifically includes the following steps: The encoder of the optimized self-supervised learning model extracts the image features of the input infrared image; The target bounding box and target category probability are generated based on the detection head of the optimized infrared target detection model, and a detection box is constructed based on the target bounding box and the target category probability. Non-maximum suppression (NMS) is used to remove redundant detection boxes in the detection boxes to obtain the final detection boxes. The image features are then detected based on the final detection boxes to obtain the detection results.
[0078] Understandably, the input infrared image undergoes preprocessing, feature extraction, and target detection. Image features are extracted using a fine-tuned encoder, and the target detection head generates bounding boxes and class probabilities for the targets. Non-maximum suppression (NMS) is then used to remove redundant detection boxes, and the final detection result is output.
[0079] It should be noted that the detection head usually includes a classification branch (category probability) and a regression branch (bounding box coordinates), which need to be adjusted according to the infrared scene: Classification branch: If there are few target categories in the infrared scene (such as pedestrians and vehicles), the output dimension can be reduced; Regression branch: Add thermal target shape prior (such as pedestrian thermal contours being approximately elliptical, and vehicles being rectangular) to optimize the anchor design.
[0080] In the specific implementation, the preprocessed image can be feature extracted by the encoder after self-supervised pre-training and fine-tuning to obtain image features. Generally, feature representations containing information such as thermal target contours and temperature gradients can be mined from the image. Subsequently, the extracted features are input into the target detection head, and the probability of the target belonging to each category and the coordinate information of the bounding box are generated through classification and regression branches, respectively. Finally, the non-maximum suppression (NMS) algorithm is used to process the multiple candidate detection boxes output by the detection head. By comparing the category probabilities and overlap of each detection box, redundant and low-confidence detection boxes can be removed, and the detection boxes most likely to actually contain the target can be retained, thus obtaining the final accurate infrared target detection result.
[0081] This embodiment, through the above scheme, dynamically adjusts the masking strategy and employs a multi-scale loss function, enabling the encoder in the MAE framework of the self-supervised learning model to learn the physical feature representation of unlabeled infrared training data, thus obtaining an optimized self-supervised learning model. The encoder in the optimized self-supervised learning model is used as the backbone of target detection. This backbone is fine-tuned using labeled infrared target detection training data. The infrared target detection model corresponding to the fine-tuned backbone is then subjected to data augmentation end-to-end training using a complete infrared dataset, resulting in an optimized infrared target detection model. Based on the optimized self-supervised learning model and the optimized infrared target detection model, feature extraction and target detection are performed on the input infrared image to obtain detection results. This demonstrates how self-supervised representation learning can utilize unlabeled data to achieve target detection. Pre-training with infrared images reduces the need for large-scale labeled data. The infrared pre-trained model trained through self-supervised learning replaces the pre-trained model trained with large-scale visible light data, extracting more robust infrared image features and enhancing the model's detection capabilities in complex infrared scenes. This solves the problem of data scarcity in the field of infrared target detection. By integrating MAE's robust feature extraction capabilities into an efficient target detection framework, the accuracy and robustness of infrared target detection are significantly improved, especially in low-resolution, high-noise infrared images. It narrows the inter-domain differences between visible light and infrared images, enabling the model to better adapt to infrared scenes. In complex and variable infrared scenes (such as nighttime, fog, and smoke), the model exhibits stronger generalization ability and stability, improving the speed and efficiency of self-supervised learning for infrared target detection.
[0082] Accordingly, the present invention further provides a self-supervised learning infrared target detection device.
[0083] Reference Figure 8 , Figure 8 This is a functional block diagram of the first embodiment of the self-supervised learning infrared target detection device of the present invention.
[0084] In a first embodiment of the self-supervised learning infrared target detection device of the present invention, the self-supervised learning infrared target detection device includes: The pre-training module 10 is used to pre-train the infrared image training set according to the mask autoencoder (MAE) architecture of the self-supervised learning model to obtain the pre-trained encoder weights.
[0085] The model building module 20 is used to load the encoder weights into the feature extraction network for target detection to build an infrared target detection model.
[0086] The optimization detection module 30 is used to optimize the self-supervised learning model and the infrared target detection model. The optimized self-supervised learning model and the optimized infrared target detection model are used to extract features and detect targets in the input infrared image to obtain detection results.
[0087] The pre-training module 10 is also used to acquire a preset number of infrared images, preprocess each infrared image, and obtain an infrared image training set corresponding to the preprocessed infrared training images.
[0088] The pre-training module 10 is also used to acquire a preset number of infrared images, perform image enhancement and normalization processing on each infrared image to obtain pre-processed infrared training images, and generate an infrared image training set based on each infrared training image.
[0089] The model building module 20 is further configured to randomly mask image regions of the infrared image training set within a preset range according to the core encoder of the masked autoencoder (MAE) architecture of the self-supervised learning model; control the self-supervised learning model to reconstruct the masked infrared image, obtain the visual representation of the infrared image, and store the visual representation of the infrared image in the encoder weights of the core encoder to obtain the pre-trained encoder weights.
[0090] The optimized detection module 30 is further configured to load the encoder weights into the feature extraction network for target detection and retain the detection head for target detection; obtain the thermal morphological feature parameters of infrared targets from the infrared image training set, optimize the preset anchor boxes for target detection based on the thermal morphological feature parameters; integrate the optimized anchor boxes into the detection head, construct an infrared target detection model by combining it with a preset target detection loss function, and output the bounding boxes and class probabilities of the infrared targets through forward inference.
[0091] The optimized detection module 30 is further configured to dynamically adjust the masking strategy and employ a multi-scale loss function to enable the encoder in the MAE framework of the self-supervised learning model to learn the physical feature representation of the unlabeled infrared training data, thereby obtaining an optimized self-supervised learning model; using the encoder in the optimized self-supervised learning model as the backbone of target detection, fine-tuning the backbone using labeled infrared target detection training data, and performing data augmentation end-to-end training on the infrared target detection model corresponding to the fine-tuned backbone using a complete infrared dataset to obtain an optimized infrared target detection model; and performing feature extraction and target detection on the input infrared image based on the optimized self-supervised learning model and the optimized infrared target detection model to obtain detection results.
[0092] The optimized detection module 30 is further configured to extract image features of the input infrared image based on the encoder of the optimized self-supervised learning model; generate target bounding boxes and target class probabilities based on the detection head of the optimized infrared target detection model; construct detection boxes based on the target bounding boxes and target class probabilities; remove redundant detection boxes in the detection boxes using non-maximum suppression (NMS) to obtain final detection boxes; and perform detection on the image features based on the final detection boxes to obtain detection results.
[0093] The steps for implementing each functional module of the self-supervised learning infrared target detection device can be referred to in the various embodiments of the self-supervised learning infrared target detection method of the present invention, and will not be repeated here.
[0094] Furthermore, this embodiment of the invention also proposes a storage medium storing a self-supervised learning infrared target detection program, which, when executed by a processor, performs the following operations: The infrared image training set is pre-trained based on the masked autoencoder (MAE) architecture of the self-supervised learning model to obtain the pre-trained encoder weights. The encoder weights are loaded into the feature extraction network for target detection to construct an infrared target detection model; The self-supervised learning model and the infrared target detection model are optimized. The optimized self-supervised learning model and the optimized infrared target detection model are used to extract features and detect targets in the input infrared image to obtain detection results.
[0095] Furthermore, when the self-supervised learning infrared target detection program is executed by the processor, it also performs the following operations: A preset number of infrared images are collected, and each infrared image is preprocessed to obtain an infrared image training set corresponding to the preprocessed infrared training images.
[0096] Furthermore, when the self-supervised learning infrared target detection program is executed by the processor, it also performs the following operations: A preset number of infrared images are collected, and each infrared image is enhanced and normalized to obtain preprocessed infrared training images. An infrared image training set is then generated based on each infrared training image.
[0097] Furthermore, when the self-supervised learning infrared target detection program is executed by the processor, it also performs the following operations: The core encoder of the self-supervised learning model's masked autoencoder (MAE) architecture randomly masks image regions within a preset range of the infrared image training set. The self-supervised learning model is controlled to reconstruct the masked infrared image to obtain the visual representation of the infrared image. The visual representation of the infrared image is stored in the encoder weights of the core encoder to obtain the pre-trained encoder weights.
[0098] Furthermore, when the self-supervised learning infrared target detection program is executed by the processor, it also performs the following operations: The encoder weights are loaded into the feature extraction network for object detection, while the detection head for object detection is retained. The thermal morphological characteristic parameters of infrared targets are obtained from the infrared image training set, and the preset anchor box for target detection is optimized based on the thermal morphological characteristic parameters. The optimized anchor box is integrated into the detection head, and an infrared target detection model is constructed by combining it with a preset target detection loss function. The bounding box and class probability of the infrared target are output through forward inference.
[0099] Furthermore, when the self-supervised learning infrared target detection program is executed by the processor, it also performs the following operations: By dynamically adjusting the masking strategy and adopting a multi-scale loss function, the encoder in the MAE framework of the self-supervised learning model learns the physical feature representation of the unlabeled infrared training data, thus obtaining an optimized self-supervised learning model. The encoder in the optimized self-supervised learning model is used as the backbone of the target detection. The backbone is fine-tuned using labeled infrared target detection training data. The infrared target detection model corresponding to the fine-tuned backbone is then trained end-to-end using a complete infrared dataset to obtain the optimized infrared target detection model. Based on the optimized self-supervised learning model and the optimized infrared target detection model, feature extraction and target detection are performed on the input infrared image to obtain the detection results.
[0100] Furthermore, when the self-supervised learning infrared target detection program is executed by the processor, it also performs the following operations: The encoder of the optimized self-supervised learning model extracts the image features of the input infrared image; The target bounding box and target category probability are generated based on the detection head of the optimized infrared target detection model, and a detection box is constructed based on the target bounding box and the target category probability. Non-maximum suppression (NMS) is used to remove redundant detection boxes in the detection boxes to obtain the final detection boxes. The image features are then detected based on the final detection boxes to obtain the detection results.
[0101] Those skilled in the art will understand that all or part of the steps in the methods described above can be implemented by a program instructing related hardware. The program is stored in a storage medium and includes several instructions to cause a device (which may be a microcontroller, chip, etc.) or processor to execute all or part of the steps of the methods described in the various embodiments of this application. The aforementioned storage medium is a computer-readable storage medium, including: USB flash drive, mobile hard drive, read-only memory (ROM), random access memory (RAM), magnetic disk or optical disk, and other media that can store program code.
[0102] It should be noted that, in this document, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Unless otherwise specified, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes that element.
[0103] The sequence numbers of the above embodiments of the present invention are for descriptive purposes only and do not represent the superiority or inferiority of the embodiments.
[0104] The above are merely preferred embodiments of the present invention and do not limit the scope of the patent. Any equivalent structural or procedural transformations made based on the description and drawings of the present invention, or direct or indirect applications in other related technical fields, are similarly included within the scope of patent protection of the present invention.
Claims
1. A self-supervised learning method for infrared target detection, characterized in that, The self-supervised learning infrared target detection method includes: The infrared image training set is pre-trained based on the masked autoencoder (MAE) architecture of the self-supervised learning model to obtain the pre-trained encoder weights. The encoder weights are loaded into the feature extraction network for target detection to construct an infrared target detection model; The self-supervised learning model and the infrared target detection model are optimized. The optimized self-supervised learning model and the optimized infrared target detection model are used to extract features and detect targets in the input infrared image to obtain the detection results.
2. The self-supervised learning infrared target detection method as described in claim 1, characterized in that, Before pre-training the infrared images in the training set according to the masked autoencoder (MAE) architecture of the self-supervised learning model to obtain the pre-trained encoder weights, the self-supervised learning infrared target detection method further includes: A preset number of infrared images are collected, and each infrared image is preprocessed to obtain an infrared image training set corresponding to the preprocessed infrared training images.
3. The self-supervised learning infrared target detection method as described in claim 2, characterized in that, The process involves acquiring a preset number of infrared images, preprocessing each infrared image to obtain a training set of infrared images corresponding to the preprocessed infrared training images, including: A preset number of infrared images are collected, and each infrared image is enhanced and normalized to obtain preprocessed infrared training images. An infrared image training set is then generated based on each infrared training image.
4. The self-supervised learning infrared target detection method as described in claim 1, characterized in that, The step of pre-training the infrared image training set based on the masked autoencoder (MAE) architecture of the self-supervised learning model to obtain the pre-trained encoder weights includes: The core encoder of the self-supervised learning model's masked autoencoder (MAE) architecture randomly masks image regions within a preset range of the infrared image training set. The self-supervised learning model is controlled to reconstruct the masked infrared image, obtain the visual representation of the infrared image, and store the visual representation of the infrared image in the encoder weights of the core encoder to obtain the pre-trained encoder weights.
5. The self-supervised learning infrared target detection method as described in claim 1, characterized in that, The step of loading the encoder weights into the feature extraction network for target detection to construct an infrared target detection model includes: The encoder weights are loaded into the feature extraction network for object detection, while the detection head for object detection is retained. The thermal morphological characteristic parameters of infrared targets are obtained from the infrared image training set, and the preset anchor box for target detection is optimized based on the thermal morphological characteristic parameters. The optimized anchor box is integrated into the detection head, and an infrared target detection model is constructed by combining it with a preset target detection loss function. The bounding box and class probability of the infrared target are output through forward inference.
6. The self-supervised learning infrared target detection method as described in claim 1, characterized in that, The optimization of the self-supervised learning model and the infrared target detection model, and the subsequent feature extraction and target detection of the input infrared image using the optimized self-supervised learning model and optimized infrared target detection model to obtain detection results, includes: By dynamically adjusting the masking strategy and adopting a multi-scale loss function, the encoder in the MAE framework of the self-supervised learning model learns the physical feature representation of the unlabeled infrared training data, thus obtaining an optimized self-supervised learning model. The encoder in the optimized self-supervised learning model is used as the backbone of the target detection. The backbone is fine-tuned using labeled infrared target detection training data. The infrared target detection model corresponding to the fine-tuned backbone is then trained end-to-end using a complete infrared dataset to obtain the optimized infrared target detection model. Based on the optimized self-supervised learning model and the optimized infrared target detection model, feature extraction and target detection are performed on the input infrared image to obtain the detection results.
7. The self-supervised learning infrared target detection method as described in claim 6, characterized in that, The step of extracting features and detecting targets from the input infrared image based on the optimized self-supervised learning model and the optimized infrared target detection model to obtain detection results includes: The encoder of the optimized self-supervised learning model extracts the image features of the input infrared image; The target bounding box and target category probability are generated based on the detection head of the optimized infrared target detection model, and a detection box is constructed based on the target bounding box and the target category probability. Non-maximum suppression (NMS) is used to remove redundant detection boxes in the detection boxes to obtain the final detection boxes. The image features are then detected based on the final detection boxes to obtain the detection results.
8. A self-supervised learning infrared target detection device, characterized in that, The self-supervised learning infrared target detection device includes: The pre-training module is used to pre-train the infrared image training set based on the masked autoencoder (MAE) architecture of the self-supervised learning model to obtain the pre-trained encoder weights. The model building module is used to load the encoder weights into the feature extraction network for target detection to build an infrared target detection model. The optimization detection module is used to optimize the self-supervised learning model and the infrared target detection model. The optimized self-supervised learning model and the optimized infrared target detection model are used to extract features and detect targets in the input infrared image to obtain detection results.
9. A self-supervised learning infrared target detection device, characterized in that, The self-supervised learning infrared target detection device includes: a memory, a processor, and a self-supervised learning infrared target detection program stored in the memory and executable on the processor, wherein the self-supervised learning infrared target detection program is configured to implement the steps of the self-supervised learning infrared target detection method as described in any one of claims 1 to 7.
10. A storage medium, characterized in that, The storage medium stores a self-supervised learning infrared target detection program, which, when executed by a processor, implements the steps of the self-supervised learning infrared target detection method as described in any one of claims 1 to 7.
Citation Information
Patent Citations
Hyperspectral image classification method based on spectrum Transform self-supervised learning algorithm model
CN118038269A
Semi-supervised infrared small target detection method based on thermodynamic heuristic data enhancement
CN118071989A
Infrared target detection method and device, computer equipment and storage medium
CN118229961A
Image edge detection method based on deep supervision mask self-encoding network
CN118334065A
Mask recovery enhancement-based low-resolution weak and small target detection method and system
CN118864826A