Target identification method and training method in infrared image
Self-supervised learning with contrastive learning and augmented puzzle images addresses the challenges of low contrast and noise in infrared ship classification, enhancing feature extraction and reducing data reliance for improved maritime classification accuracy.
Patent Information
- Application Number
- CN202510484473.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-17
- Publication Date
- 2025-07-15
AI Technical Summary
The existing infrared ship classification methods rely on a large amount of manual labeling data, and the classification accuracy is low in low visibility environment, and the generalization ability is insufficient, making it difficult to meet the actual application needs.
Using self-supervised learning combined with contrast learning, a sample pool is constructed through image enhancement and puzzle tasks, optimized feature extraction, reduced annotation dependence, and enhanced feature extraction ability and classification robustness.
With a small amount of labeled data, it improves the accuracy and generalization ability of infrared image classification, is suitable for complex marine environments, and reduces the cost of data labeling.
Smart Images

Figure CN120318587A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of infrared image processing technology, and in particular to a method for target recognition and a training method in infrared images. Background Art
[0002] Infrared ship images have important values in applications such as maritime surveillance, target recognition, maritime search and rescue, and military defense. Traditional ship recognition methods mainly rely on visible light images. However, in low visibility environments such as at night, in bad weather (such as haze, heavy rain), and for long-distance target detection, the visible light imaging ability drops significantly, resulting in unstable classification effects. Therefore, infrared ship classification has become an important research direction, which can provide stable target detection and recognition capabilities in insufficient lighting conditions, ensuring the safety of the sea and the reliability of intelligent ship navigation. For example, in a maritime search and rescue mission, infrared images can help quickly locate the distressed ship; in military defense, infrared images can effectively identify enemy ships and provide important intelligence support. Therefore, the research on infrared ship classification technology is of great significance for improving the efficiency of maritime surveillance and target recognition.
[0003] However, infrared ship classification faces many challenges in practical applications. First, infrared images usually have problems such as low contrast, high noise, and unclear details. Due to the working principle of infrared sensors, the contrast between the target and the background in the image is relatively low, resulting in blurred target edges and making it difficult to accurately extract features. In addition, infrared images are easily affected by environmental noise (such as sea surface reflection, weather interference, etc.), further increasing the difficulty of classification. Second, the annotation of infrared ship images requires professional personnel and is time-consuming and laborious. Due to the complexity and diversity of infrared images, the acquisition cost of annotated data is relatively high, restricting the construction of large-scale annotated data sets. Finally, the maritime environment is complex and changeable, and factors such as weather, lighting, and waves will all affect the quality of infrared images, resulting in insufficient generalization ability of the classification model.
[0004] Current mainstream infrared ship classification methods mainly rely on deep neural networks for supervised learning, that is, use convolutional neural networks (CNNs) and deep residual networks (ResNets) for classification training. However, these methods have obvious limitations. First, they highly rely on manually labeled data. Traditional deep learning methods require a large amount of labeled data for training, and the acquisition and annotation costs of infrared images are relatively high. Especially in military and ocean navigation scenarios, it is difficult to construct a large-scale and high-quality dataset. Second, the generalization ability of the model is limited. Due to the significant influence of the environment on infrared images, existing supervised learning methods perform unstably under different weather and lighting conditions, and are prone to overfitting to the training data, resulting in a decrease in classification accuracy. Finally, the feature extraction ability is limited. The target contours in infrared images are blurred and there are few details, and it is difficult for traditional CNN structures to extract robust ship features, leading to poor adaptability of the model in the case of small samples.
[0005] Therefore, existing infrared ship classification methods are difficult to meet the actual application requirements, and a solution that can reduce annotation dependence, enhance feature extraction ability, and improve classification robustness is needed. Summary of the Invention
[0006] This application proposes a method for target recognition and training method in infrared images, which can solve one of the problems existing in the background technology.
[0007] To achieve the above object, this application adopts the following technical solutions:
[0008] In the first aspect, a training method for a target recognition model in infrared images is provided. The training method includes:
[0009] Constructing a training sample pool; and
[0010] Using the sample pool to train an initialized target recognition model, the target recognition model includes: a feature extraction sub-model and a classifier.
[0011] Specifically, constructing the training sample pool includes:
[0012] Obtaining the original infrared image;
[0013] Performing image enhancement processing on the original infrared image to obtain an enhanced image;
[0014] Obtaining a jigsaw image from the enhanced image; and
[0015] Inputting the enhanced image features corresponding to the enhanced image and the jigsaw image features corresponding to the jigsaw image into the sample pool.
[0016] The feature extraction sub-model aims to optimize the contrast loss between the enhanced image features and the jigsaw image features, and iteratively updates the sample pool.
[0017] The classifier is placed on the output side of the feature extraction sub-model.
[0018] Based on the above technical solution, a sample pool is constructed using the enhanced image features and the jigsaw image features, and the model is trained by optimizing the contrast loss between the enhanced image features and the jigsaw image features, enabling the model to reduce the dependence on annotation through contrastive learning and bringing closer the distances in the feature space between different versions of the same image, namely the enhanced image and the jigsaw image. In this way, the target recognition and classification reduce the dependence on annotation, and the contrastive learning using the enhanced image features and the jigsaw image features enhances the feature extraction ability and improves the classification robustness.
[0019] In a possible design of the first aspect, the contrast loss adopts a noise estimation contrast loss function, and the noise estimation contrast loss function includes: a first noise estimation contrast loss sub-function for measuring the approximation between the historical features and the jigsaw image features, and a second noise estimation contrast loss sub-function for measuring the approximation between the historical features and the enhanced image features, where the historical features are determined by the enhanced image features of the historical iteration rounds.
[0020] In a possible design of the first aspect, the historical features are updated by an exponential moving average method.
[0021] In a possible design of the first aspect, the sample pool further includes hard negative samples. The first noise estimation contrast loss sub-function includes: a first part for measuring the approximation between the historical features and the jigsaw image features, and a second part for measuring the approximation between the jigsaw image features and the hard negative samples. The second noise estimation contrast loss sub-function includes: a third part for measuring the approximation between the historical features and the enhanced image features, and a fourth part for measuring the approximation between the enhanced image features and the hard negative samples.
[0022] In a possible design of the first aspect, specifically constructing the training sample pool further includes:
[0023] Determining hard negative samples from the sample pool based on the cosine similarity of features.
[0024] In a possible design of the first aspect, obtaining the jigsaw image from the enhanced image specifically includes:
[0025] Dividing the enhanced image to obtain a number of image patches; and
[0026] Randomly shuffle the image blocks to obtain the jigsaw image.
[0027] In a possible design of the first aspect, the training method further includes:
[0028] Extract features and reduce the dimensionality of the enhanced image to obtain the enhanced image features; and
[0029] Extract features and reduce the dimensionality of the image blocks to obtain image block features, and splice the image block features to obtain the jigsaw image features, where the enhanced image features and the jigsaw image features have the same feature dimension.
[0030] In a possible design of the first aspect, the original infrared image is an infrared image captured by an infrared sensor in different environments. The image enhancement process for the original infrared image is specifically:
[0031] Denoise, enhance the contrast, and normalize the size of the original infrared image to obtain the enhanced image.
[0032] In a possible design of the first aspect, the training method further includes:
[0033] Fine-tune the target recognition model using a small amount of class-labeled data, and calculate the error between the predicted class and the true class using the cross-entropy loss for the fine-tuning.
[0034] In a second aspect, an object recognition method for infrared images is provided. The recognition method includes:
[0035] Obtain the infrared image to be processed; and
[0036] Process the infrared image to be processed using the trained target recognition model as described above to obtain the recognition result.
[0037] In a third aspect, an electronic device is provided. The electronic device includes: a processor, and a memory coupled to the processor. The memory is used to store a computer program; the processor is used to execute the computer program stored in the memory so that the electronic device executes the training method in any possible implementation manner of the first aspect, or executes the recognition method described in the second aspect.
[0038] In a fourth aspect, a computer-readable storage medium is provided, including a computer program or instruction. When the computer program or instruction runs on a computer, the computer is caused to execute the training method in any possible implementation manner of the first aspect, or execute the recognition method described in the second aspect.
[0039] Fifth aspect, there is provided a computer program product, including: a computer program or instructions, when the computer program or instructions run on a computer, enabling the computer to execute the training method according to any possible implementation manner in the first aspect, or execute the recognition method according to the second aspect. Description of the Drawings
[0040] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the drawings required for use in the embodiments or related technical descriptions. Obviously, the drawings in the following description are only some embodiments of the embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0041] Figure 1 It is a framework diagram of self-supervised feature extraction based on contrastive learning provided by an embodiment of the present application;
[0042] Figure 2 It is a jigsaw puzzle schematic diagram provided by an embodiment of the present application for splitting an original image into image patches and randomly shuffling them;
[0043] Figure 3 It is a schematic diagram of hard negative sample screening provided by an embodiment of the present application;
[0044] Figure 4 It is a schematic diagram of classification prediction provided by an embodiment of the present application. Detailed Embodiments
[0045] In order to make the objectives, technical solutions and advantages of the present application clearer, the following further details the present application in conjunction with the drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.
[0046] It should be noted that although the functional modules are divided in the device schematic diagram and the logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in a different order from the module division in the device or the flowchart in the flowchart. Terms such as "first" and "second" in the specification, claims and the above drawings are used to distinguish similar objects and do not necessarily need to be used to describe a specific order or sequence.
[0047] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those of ordinary skill in the technical field to which the present application belongs. The terms used herein are only for the purpose of describing the embodiments of the present application and are not intended to limit the present application.
[0048] The following is an exemplary description of the target recognition method and training method in the infrared image of the embodiment of the present application in combination with the infrared ship classification method based on self-supervised feature extraction.
[0049] This embodiment provides an infrared ship classification method based on self-supervised feature extraction, aiming to improve the classification accuracy and robustness of infrared images in complex marine environments. Aiming at problems such as low contrast, high noise, and scarce labeled data in infrared images, this embodiment combines self-supervised learning with contrastive learning to achieve efficient feature extraction and reduce the dependence on labeled data. First, use an infrared sensor to collect diverse ship images, and improve the image quality through adaptive denoising and contrast enhancement. Then, introduce a jigsaw task as the pre-training objective to enable the model to learn global structural information, and combine a dynamic negative sample selection strategy to optimize the contrastive learning process, improving the inter-class discrimination ability and intra-class consistency. Finally, with the support of a small amount of labeled data, fine-tuning is performed to optimize the classification performance. This embodiment effectively improves the accuracy and generalization ability of infrared ship classification, and is applicable to application scenarios such as maritime monitoring, military defense, and intelligent navigation.
[0050] This embodiment adopts the method of self-supervised feature extraction and fine-tuning with a small amount of labeled data, effectively reducing the dependence on large-scale labeled data. In the early stage of model training, this embodiment uses self-supervised contrastive learning to learn the feature representation of infrared ships on unlabeled data, enabling the model to automatically extract stable features and possess strong generalization ability. In the fine-tuning stage, this embodiment only uses a small amount of labeled data to optimize the classification model, without having to train the entire model from scratch, thereby reducing the data annotation cost and improving the classification accuracy and robustness of the model in complex environments. Therefore, the present invention has research significance and practical significance.
[0051] The specific implementation method of this embodiment is as follows:
[0052] Step 1: Use an infrared sensor to collect marine ship images to ensure data diversity and extensiveness. The infrared sensor can capture targets under low light conditions such as at night and in fog, improving the adaptability in complex environments. The collected images cover different weather conditions, such as sunny days, rainy days, foggy days, etc., to enhance the generalization ability of the model. The obtained infrared images will be used as the input data for the subsequent preprocessing (Step 2).
[0053] Step 2: In order to optimize the effect of subsequent feature extraction, the collected infrared images are denoised, contrast enhanced and resized to improve the model's ability to identify targets. First, the non-local mean filter (NLM) is used to remove the noise caused by the infrared sensor while retaining the edge information to improve the image quality. Then, the low-contrast area is enhanced by local histogram equalization to make the ship target more visible against the complex background. The adaptive contrast enhancement method can also be used to dynamically adjust the local contrast of the image to ensure that the target in the low-contrast area is clearly visible and avoid the noise amplification problem that may be caused by global equalization. Subsequently, the adaptive non-local mean filter (Adaptive NLM) is used to adjust the filter parameters according to the local texture complexity of the image, so as to remove noise while retaining the image detail information. Finally, all images are resized to a uniform size to adapt to the input requirements of the neural network. The infrared images after this preprocessing step will be used for jigsaw task construction (step 3) and data enhancement (step 4) to provide high-quality input data for the model.
[0054] Step 3: This embodiment uses the jigsaw puzzle task as a pre-training task for self-supervised learning. The core idea of the jigsaw puzzle task is to disrupt the spatial arrangement of the image so that the model must rely on global features rather than local texture information to understand the image, thereby enhancing its spatial perception ability. First, for the enhanced infrared image I generated in step 2, divide it into n×n grids, with a total of n 2 Sub-block
[0055]
[0056] Among them, each sub-block P i Represents a local area of the original image.
[0057] Then, the order of the puzzle pieces is randomly disrupted to generate the puzzle image I t :
[0058] I t = shuffle(I)(2)
[0059] The goal of the shuffle operation is to rearrange these sub-blocks in a random order.
[0060] Next, the shuffled puzzle image block I t As one of the inputs of the network, it and the enhanced infrared image I together constitute the image pair required for contrastive learning. Through contrastive learning, the distance between different transformed versions of the same image (enhanced infrared image and puzzle image) in the feature space is shortened.
[0061] Step 4: In order to improve the robustness of the model, different data transformations are performed on the enhanced infrared image generated in step 3 and the puzzle image generated in step 4. For the enhanced infrared image, this embodiment applies global transformations, including random flipping, random cropping, color transformation, Gaussian blur and normalization, to increase the diversity of the data and prevent the model from over-relying on certain specific visual features. For each puzzle image block in the puzzle image, this embodiment applies local transformations, including random cropping, Gaussian noise addition and local contrast adjustment to simulate fine-grained changes in the real environment, so that the model learns a more stable feature representation. The enhanced infrared image and puzzle image after data enhancement are then input into the feature extraction network (step 5) for deep feature learning.
[0062] Step 5: This embodiment uses ResNet-50 as the feature extraction network to extract the enhanced infrared image I and the puzzle image block I after data enhancement. t The features are extracted and the dimensionality is reduced by linear projection to ensure the consistency and compactness of the feature representation. The enhanced infrared image is input into ResNet-50 for forward propagation to extract 2048-dimensional features, which are then mapped to 128-dimensional feature representation through a linear projection layer. For the transformed puzzle image, the features of each puzzle piece are calculated separately, and the 2048-dimensional features of each piece are transformed to 128 dimensions through linear projection, and then all n 2 The characteristic concatenation of the puzzle pieces is n 2 × 128 dimensions and is mapped to a 128-dimensional feature vector through a final linear projection. The extracted features are used for dynamic negative sample selection (step 6) and model updating (step 7).
[0063] Step 6: This embodiment constructs a sample pool to improve the model's ability to distinguish features and enhance the effectiveness of contrastive learning. The sample pool is used to store sample features extracted from previous training during training and provide stable and diverse negative sample references for contrastive loss calculation.
[0064] Before training begins, a forward propagation is performed on each image in the dataset, and its feature representation is extracted and stored in the sample pool as the initial value, that is, where f(·) is the feature projection head, is the original feature extracted by the backbone network. The feature representation in the sample pool is continuously updated as the training iteration proceeds.
[0065] During the training process, for each training sample, the present invention selects the most challenging negative sample from the sample pool, namely, the difficult negative sample. The difficult negative sample refers to those samples that have a high similarity with the current sample in the feature space, but are actually of different categories. In order to screen the difficult negative samples, the present embodiment uses formula (4) to calculate the cosine similarity between the features of the current original sample I and the features of all negative samples I' in the sample pool:
[0066]
[0067] Among them, f(v I ) is the current original sample feature representation, and m I' is the feature representation of the negative samples stored in the sample pool.
[0068] Then, select the N samples with the highest similarity as the hard negative samples of the current sample to ensure that the model focuses on learning how to distinguish difficult-to-distinguish sample pairs.
[0069] Step 7: After completing feature extraction (Step 5) and dynamic negative sample selection (Step 6), this embodiment combines the jigsaw task with contrastive learning to optimize the ResNet-50 network, enabling it to learn infrared image representations with strong inter-class discrimination and high intra-class consistency in an unsupervised manner.
[0070] To enable the model to learn more stable and discriminative features, based on the negative sample pool selected in Step 6, this invention introduces contrastive learning to make the enhanced infrared image I and the transformed version of the jigsaw task, the jigsaw image I t maintain a high similarity in the feature space. To this end, first, similar to formula (4), calculate the cosine similarity s(v I between the enhanced image feature v It and the transformed image feature v I , v I t).
[0071] Then, the model uses noise contrastive estimation to model the similarity of positive and negative sample pairs. Specifically, the noise contrastive estimation calculates the matching probability of positive sample pairs through the following formula (5):
[0072]
[0073] The matching probability of negative sample pairs is also the same.
[0074] Among them, is the hard negative samples selected in Step 8 as the negative sample pool for this batch, is the entire sample pool, and τ is the temperature parameter used to control the distribution of the contrastive loss, enabling the model to more effectively adjust the similarity between features.
[0075] Before calculating the cosine similarity score, different projection heads are used for the features to map the feature vectors into a low-dimensional representation space. Specifically, apply the projection head f(·) to the feature v I of the enhanced infrared image I, and to the feature v t of the transformed image I ItApply the projection head g(·) above. Then the noise contrast estimation loss is equal to minimizing the following loss:
[0076]
[0077] The first term is the positive sample contrast term, minimizing the distance between the enhanced infrared image I and its transformed version, the puzzle image I t in the representation space, guiding the model to learn a feature representation that is robust to image transformation. Even if the image undergoes different forms of perturbation, its semantic features should remain consistent, that is, different transformations of the same image should have similar semantic features. The second term is to maximize the distance between the puzzle image I t and the negative sample I’ in the representation space, ensuring that images of different semantic classes have significant distinguishability in the representation space.
[0078] In the original noise contrast estimation loss function formula (6), only the contrast relationships of "enhanced infrared image I - puzzle image I t " and "puzzle image I t - hard negative sample I'" are established, lacking the direct constraint of "enhanced infrared image I - hard negative sample I'. At the same time, the feature representation of the enhanced infrared image I is easily affected by the random perturbation of single-batch parameter updates, resulting in possible large fluctuations in the feature space, which in turn affects the stability of model optimization.
[0079] To solve the above problems and improve the contrast effect, this embodiment introduces the historical feature representation m I in the sample pool to enhance the learning stability of the model, and adopts the convex combination of two noise contrast estimation loss functions to balance the influence of different loss terms. Specifically, the two contrast loss functions are combined in a weighted manner, so as to enhance the model's ability to distinguish negative samples while maintaining the invariance of image representation. Therefore, the final noise contrast estimation loss function can be expressed as:
[0080]
[0081] where λ is a hyperparameter that controls the influence of the two losses on model optimization. Specifically, the first term of formula (7) is a direct application of formula (6), but here the feature representation m I of the current sample in the sample pool is used I to replace the feature representation f(v I ) of the original sample in the current training round. The historical feature representation m IRepresents the exponential moving average of the historical iteration round features of the enhanced image I, enabling the model to utilize more stable historical feature information for contrastive learning without expanding the training batch, effectively reducing the fluctuation amplitude of parameter updates. The second term has two functions: one is to make the representation f(v I ) of image I similar to its memory representation m I , which helps reduce the drift of feature distributions during training, minimize the impact of batch noise, and thus make the model updates more stable; the other is to directly promote the difference between the representation f(v I ) of image I and the representation of the hard negative sample image I', improving the model's ability to distinguish different images. At the same time, to further enhance the training stability, the feature representation m I' of the hard negative sample in the sample pool is used to replace the original f(v I' ), thereby leveraging more stable feature information from historical batches. These two work together to ensure that the model can learn feature representations invariant to transformations and enhance the model's accuracy in distinguishing different images. In short, the first loss makes the model invariant to image transformations by comparing the features of the jigsaw-transformed image with the historical features of the enhanced images in the memory bank; the second loss stabilizes the feature representation and enhances discriminability by comparing the enhanced image features with their own historical features. Both losses use the negative sample features in the sample pool for contrastive learning to jointly optimize the robust expression of infrared ship features by the model.
[0082] The above contrastive loss needs to consider both positive and negative sample pairs simultaneously and is trained through contrastive learning to minimize the feature distance (positive sample pair) between the enhanced image features and the jigsaw image features, while separately maximizing the feature distances (negative sample pairs) between the enhanced image features, the jigsaw image features, and the hard negative samples sampled from the sample pool. Through the contrastive loss, the enhanced image features are made similar to the jigsaw image features while dissimilar to other image features (negative samples).
[0083] During the training process, the model uses the stochastic gradient descent optimizer for parameter updates and uses momentum acceleration to improve the convergence speed. The entire training process is iterated until the loss of the model converges. Through iterative optimization, the network learns feature representations robust to interference factors such as jigsaw transformations.
[0084] Step 7: At different stages of training, the sample distribution in the sample pool changes over time, so the sample pool needs to be updated dynamically. Specifically, the implementation of the dynamic update strategy includes the following steps: First, after each training iteration, the sample features f(v I ) in the current batch are added to the sample pool, and the features of these samples update the feature representation m calculated in the historical training rounds in the sample pool through exponential moving averageI , as shown in the formula:
[0085] m I =αm I +(1-α)f(v I ) (8)
[0086] Where α is a hyperparameter between 0 and 1, which controls the weight of historical information and current information.
[0087] Secondly, in order to maintain the diversity and timeliness of the memory bank, the oldest stored samples will be removed. This removal strategy can prevent the memory bank from being occupied by outdated samples and ensure that the samples in the memory bank always reflect the feature distribution of the current training stage.
[0088] Step 8: After the self-supervised learning is completed, the classification model is fine-tuned using a small amount of labeled data. Specifically, in the fine-tuning stage, a linear classifier is used, which takes the 128-dimensional feature representation of the enhanced infrared image I extracted in step 5 as input and optimizes it in a supervised manner, so that the model can have good classification ability on limited labeled data.
[0089] First, the feature representation v after self-supervised learning is I Enter a fully connected classification layer to predict the final category:
[0090]
[0091] Among them, W c is the learnable weight matrix of the classifier, with a shape of C×128, where C is the number of categories and b c is the bias vector. The Softmax function is used to convert the network output into a class probability distribution.
[0092] Then, the cross entropy loss is used to calculate the error between the predicted category and the true category:
[0093]
[0094] Among them, y c is the one-hot encoding of the true category, is the class probability distribution predicted by the model.
[0095] Finally, the loss gradient is calculated by backpropagation and updated using a stochastic gradient descent optimizer to optimize the classifier until the training loss converges.
[0096] Step 9: Use the fine-tuned classification model to perform classification prediction on the new infrared ship images. Specifically, input the new infrared images into the feature extraction network to extract their feature representations, and then input the features into the classifier to output the categories of the ships. The classification results can be used in application scenarios such as maritime surveillance and target recognition.
[0097] The embodiment of the present application also provides an electronic device, including: a processor, and a memory coupled to the processor, where the memory is used to store a computer program; the processor is used to execute the computer program stored in the memory so that the electronic device executes the method described in any one of the above embodiments.
[0098] The electronic device can be a computing device such as a desktop computer, a notebook, a palm computer, and a cloud server. The electronic device may include, but is not limited to, a processor and a memory.
[0099] The so-called processor may be a central processing unit (CPU), or may also be other general-purpose processors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc. The processor is the control center of the electronic device, and uses various interfaces and lines to connect various parts of the entire device.
[0100] The memory may be used to store the computer program, and the processor realizes various functions of the electronic device by running or executing the computer program stored in the memory and calling the data stored in the memory.
[0101] The memory may mainly include a program storage area and a data storage area. Among them, the program storage area may store an operating system, application programs required for at least one function, etc.; the data storage area may store data created according to the use of the mobile phone, etc. In addition, the memory may include high-speed random access memory, and may also include non-volatile memory, such as a hard disk, a memory, a plug-in hard disk, a smart media card (SMC), a secure digital (SD) card, a flash card, at least one magnetic disk storage device, a flash memory device, or other volatile solid-state storage devices.
[0102] An embodiment of the present application further provides a storage medium, which is a computer-readable storage medium. The computer program is stored in the computer-readable storage medium. When the computer program is executed by a processor, the steps of the above method embodiments can be implemented. Among them, the computer program includes computer program code, and the computer program code can be in the form of source code, object code, executable file or some intermediate form, etc. The computer-readable medium may include: any entity or device capable of carrying the computer program code, recording medium, USB flash drive, mobile hard disk, magnetic disk, optical disk, computer memory, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), electrical carrier signal, telecommunication signal, and software distribution medium, etc.
[0103] An embodiment of the present application further provides a computer program product, including: a computer program or instruction. When the computer program or instruction runs on a computer, the computer is caused to execute the method of any of the above possible implementation manners.
[0104] The above is the preferred implementation manner of the present application. It should be noted that for those of ordinary skill in the art, without departing from the principle of the present application, several improvements and refinements can be made, and these improvements and refinements are also regarded as the protection scope of the present application.
Claims
1. A training method for a target recognition model in an infrared image, characterized in that, The training method includes: Constructing a sample pool for training; and Using the sample pool to train an initialized target recognition model, the target recognition model including: a feature extraction sub-model and a classifier, Specifically, constructing the sample pool for training includes: Obtaining an original infrared image; Performing image enhancement processing on the original infrared image to obtain an enhanced image; Obtaining a jigsaw image from the enhanced image; and Inputting the enhanced image features corresponding to the enhanced image and the jigsaw image features corresponding to the jigsaw image into the sample pool, The feature extraction sub-model aims to optimize the contrast loss between the enhanced image features and the jigsaw image features, and iteratively updates the sample pool, The classifier is placed on the output side of the feature extraction sub-model.
2. The training method according to claim 1, wherein The contrast loss adopts a noise estimation contrast loss function, and the noise estimation contrast loss function includes: a first noise estimation contrast loss sub-function for measuring the approximation degree between historical features and the jigsaw image features, and a second noise estimation contrast loss sub-function for measuring the approximation degree between historical features and the enhanced image features, where the historical features are determined by the enhanced image features of historical iteration rounds.
3. The training method according to claim 2, wherein Updating the historical features by an exponential moving average method.
4. The training method according to claim 2, wherein The sample pool further includes hard negative samples. The first noise estimation contrast loss sub-function includes: a first part for measuring the approximation degree between the historical features and the jigsaw image features, and a second part for measuring the approximation degree between the jigsaw image features and the hard negative samples. The second noise estimation contrast loss sub-function includes: a third part for measuring the approximation degree between the historical features and the enhanced image features, and a fourth part for measuring the approximation degree between the enhanced image features and the hard negative samples.
5. The training method according to claim 4, wherein Specifically, constructing the sample pool for training further includes: Determining hard negative samples from the sample pool based on the cosine similarity of features.
6. The training method according to claim 1, wherein Specifically, obtaining a jigsaw image from the enhanced image includes: dividing the enhanced image to obtain a plurality of image blocks; and Randomly shuffling the image blocks to obtain the jigsaw image.
7. The training method according to claim 6, wherein The training method further includes: Performing feature extraction and feature dimensionality reduction on the enhanced image to obtain the enhanced image features; and Performing feature extraction and feature dimensionality reduction on the image blocks to obtain image block features, and splicing the image block features to obtain the jigsaw image features, where the enhanced image features and the jigsaw image features have the same feature dimension.
8. The training method according to claim 1, wherein The original infrared image is an infrared image captured by an infrared sensor in different environments. Specifically, performing image enhancement processing on the original infrared image includes: Performing denoising, contrast enhancement, and size normalization processing on the original infrared image to obtain the enhanced image.
9. The training method according to claim 1, wherein The training method further includes: Fine-tuning the target recognition model with a small amount of class-labeled data, and calculating the error between the predicted class and the true class by using the cross-entropy loss for the fine-tuning.
10. A method for target recognition in infrared images, characterized in that, The recognition method includes: Obtaining an infrared image to be processed; and Process the infrared image to be processed by using the target recognition model trained according to any one of claims 1-9 to obtain a recognition result.
Citation Information
Cited By
Self-adaptive full-scale infrared target detection network based on YOLO
CN121725341A