Seeker infrared re-identification method based on DINOv2 model fine tuning

Through fine-tuning technology and preprocessing method based on DINOv2 model, the problem of low accuracy and efficiency of infrared image re-recognition under complex viewing angles and distance differences is solved, and efficient and accurate re-recognition of infrared images on sea surface targets is achieved.

CN119992288APending Publication Date: 2025-05-13NORTHWESTERN POLYTECHNICAL UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510201447.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-24
Publication Date
2025-05-13

AI Technical Summary

Technical Problem

The existing image re-recognition algorithms are not very accurate and efficient when processing the target poses observed by the seeker and the distance differences are significant.

Method used

The infrared re-identification method of seeker based on DINOv2 model fine-tuning is adopted. Through the fine-tuning training stage and query recognition stage, the pre-trained DINOv2 backbone model and the post-random initialization classifier are used for fine-tuning. Combined with pre-processing technologies such as automatic gamma correction, size specification and pixel standardization, the accuracy of feature extraction and re-identification is improved.

Benefits of technology

It significantly improves the accuracy and efficiency of infrared image target re-identification, and can efficiently and accurately re-identify infrared images of sea surface targets under complex viewing angles and distance differences.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119992288A_ABST
    Figure CN119992288A_ABST
Patent Text Reader

Abstract

The invention discloses a seeker infrared re-identification method based on DINOv2 model fine tuning. A fine tuning training stage comprises the following steps: preprocessing an infrared image of a cooperative target; setting a learning rate change strategy in the training process; then, a pre-trained DINOv2 backbone model is used, a randomly initialized classifier is arranged behind the pre-trained DINOv2 backbone model for fine tuning, and a fine-tuned DINOv2 model is obtained after training is finished; the query recognition stage comprises the following steps: preprocessing an infrared image of a non-cooperative target; performing feature extraction on the infrared image of the non-cooperative target by using a fine-tuned DINOv2 model to obtain discriminant features, which are easy to classify, of various targets; and calculating the similarity between the features of the specified target image and the features of the target image in the candidate library, and sorting the images in the candidate library according to the similarity between the images to obtain the most similar image, thereby determining the identity of the specified image. The problem that an existing image re-recognition algorithm is low in re-recognition accuracy and efficiency is solved.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical field]

[0001] The present invention belongs to the technical field of image recognition, and in particular relates to a seeker infrared re-identification method based on DINOv2 model fine-tuning. [Background technology]

[0002] Re-identification (ReID) is a technology that uses computer vision technology to determine whether a specific object exists in an image or video sequence. It aims to accurately identify a specific target individual on the sea surface, such as a red-painted cargo ship, from different perspectives, distances or lighting conditions.

[0003] There may be multiple targets in the seeker's field of view, so the seeker needs to accurately identify and determine the target to prevent the target from being lost due to incorrect target identification. During the seeker's flight, the seeker's attitude relative to the target changes rapidly, causing the attitude, shape and size of the target in the seeker's field of view to change dramatically, making the target's motion state uncertain.

[0004] Due to the difference in the shooting posture of the seeker (including changes in distance, pitch angle and azimuth), the same target presents large visual differences in images taken in different time and space, which has a negative impact on target re-identification. Existing image re-identification algorithms often find it difficult to effectively deal with such image differences caused by changes in shooting time and space, resulting in low re-identification accuracy and efficiency. [Summary of the invention]

[0005] The purpose of the present invention is to provide a seeker infrared re-identification method based on DINOv2 model fine-tuning to solve the problems of low re-identification accuracy and efficiency of existing image re-identification algorithms.

[0006] The present invention adopts the following technical scheme: a seeker infrared re-identification method based on DINOv2 model fine-tuning, including a fine-tuning training stage and a query recognition stage;

[0007] The fine-tuning training phase includes the following:

[0008] Preprocess the infrared image of the cooperative target; set the learning rate change strategy during the training process; then use the pre-trained DINOv2 backbone model and post-randomly initialized classifier for fine-tuning, and obtain the fine-tuned DINOv2 model after the training;

[0009] The query identification phase includes the following:

[0010] The infrared images of non-cooperative targets are subjected to feature extraction using the fine-tuned DINOv2 model, and discriminant features that are easy to classify for various types of targets are obtained;

[0011] Calculate the similarity between the features of the specified target image and the target image features in the candidate library, sort the images in the candidate library according to the similarity between the images, and obtain the most similar image, thereby determining the identity of the specified image.

[0012] Further, the preprocessing includes automatic gamma correction, size specification and pixel normalization, wherein the corresponding gamma value is obtained using the following gamma value calculation method, and the gamma value of the process is determined by converting the brightness average to the center of the histogram:

[0013]

[0014] Among them, X is the average brightness of the normalized image, X∈[0,1], 0.5 is the center of the brightness histogram, and after obtaining the corresponding gamma value, the image is gamma corrected using the following formula to obtain an image with reasonable brightness:

[0015] S=R γ ,

[0016] Among them, S is the rectified image, R is the original normalized image, and its range is R∈[0, 1].

[0017] Furthermore, the specific process of size normalization is: normalize the size to a specified size, namely (224, 224, 3), where the two 224s refer to the width and height of the image, respectively, and 3 is the number of channels of the image.

[0018] Furthermore, the specific process of pixel normalization in the fine-tuning training stage is: the image is normalized using the mean and standard deviation, and the normalized mean of each channel of the image is 0.5 and the standard deviation is 0.5.

[0019] Furthermore, the learning rate change strategy introduces learning rate warm-up and decay strategies to optimize the model training effect:

[0020] In the early stage of model training, a smaller learning rate is needed to make the model weight update more stable. The learning rate change rule of the warm-up process is shown in the following formula:

[0021]

[0022] Where lr refers to the learning rate, i refers to the current training round, a refers to the initial learning rate coefficient in the warm-up phase, a∈(0,1), lr i Refers to the learning rate of the current round, lr base Refers to the basic learning rate, epochs warmup Refers to the number of rounds required for the warm-up process;

[0023] By using the learning rate decay strategy, the model can adjust parameters more finely in the later stages of training, thereby converging to a better solution;

[0024] The specific cosine attenuation strategy is as follows:

[0025]

[0026] Among them, i refers to the current training round, epoch warmup Indicates the last round of warm-up, epochs max Represents the total number of training rounds.

[0027] Furthermore, the method for fine-tuning the DINOv2 model in the fine-tuning training phase is as follows:

[0028] The backbone of the DINOv2 training model is selected to extract features from the input target image and obtain feature representations containing category tokens;

[0029] Perform layer normalization on all tokens output by the backbone (including category tokens and tile tokens);

[0030] Add a classification layer and use a linear fully connected network to perform linear transformation on the normalized features;

[0031] Determine the optimization goal of re-ID model training, that is, to make the similarity of features between different images of the same target greater and the similarity of features between different target images smaller;

[0032] Using the total loss function, the parameters of the model are continuously adjusted through the back propagation algorithm to gradually reduce the total loss to a stable level; the total loss function is the sampling triplet loss function of difficult sample mining and the cross entropy loss function.

[0033] Furthermore, the loss function of the sampling triplet for difficult sample mining is as follows:

[0034]

[0035] Where d represents the distance between samples, a refers to anchor samples, p refers to positive samples, n refers to negative samples, margin is a constant greater than 0, P is the number of targets in each batch, K different images are randomly selected for each target, A is defined as the set of images with the same target as the anchor sample in a training batch, and B is defined as the set of images with different targets from the anchor sample in a training batch.

[0036] Furthermore, the specific method of the query identification stage is:

[0037] Preprocess the infrared images of non-cooperative targets, including gamma correction, size standardization and image normalization;

[0038] Use a fine-tuned DINOv2 model for feature extraction and perform layer normalization;

[0039] Select the category token features to query;

[0040] Calculate the Euclidean distance between the query image feature vector and the feature vector of each image in the database to obtain a feature distance list; the smaller the feature distance, the closer the distance between the two images in the feature space, that is, the more similar they are; return the identity of the image with the smallest distance as the query recognition result.

[0041] The beneficial effects of the present invention are as follows: the present invention focuses on the problem that the difficulty of re-identifying the target image increases sharply when the seeker observes a variety of target postures and significant distance differences. In response to this problem, the present invention proposes a seeker infrared re-identification method based on DINOv2 model fine-tuning, which significantly improves the accuracy and efficiency of re-identifying infrared images of sea surface targets captured from different distances and viewing angles.

[0042] The present invention introduces a fine-tuned DINOv2 pre-trained model in the field of image recognition technology, especially in the infrared image re-identification of sea surface strike targets. This innovation significantly improves the feature representation capability and robustness of infrared image target re-identification, enabling the model to efficiently and accurately re-identify infrared image targets under complex viewing angles and distance differences. When re-identifying infrared images of sea surface targets captured at different distances and viewing angles, the accuracy and efficiency can be significantly improved.

[0043] In both the fine-tuning training stage and the query recognition stage, the present invention uses automatic gamma correction preprocessing to adjust the brightness of the target image to a reasonable range, thereby solving the problem of image nonlinear effects caused by the limitations of infrared imaging equipment or inappropriate ambient light, and significantly enhancing the robustness and compatibility of the model. At the same time, the image is also size-standardized and pixel-normalized to meet the dimension and range of the model input data, further improving the recognition performance of the model.

[0044] In order to complete the training of the re-identification model, the present invention introduces the hard sample mining sampling triplet loss as a loss function to enhance the discriminability of features. At the same time, the cross entropy loss is also incorporated into the training framework to further constrain and optimize the feature space. The combined use of this loss function enables the model to better understand the similarities and differences between individuals at the feature level, thereby improving the accuracy of re-identification.

Brief Description of the Drawings

[0045] Figure 1 (a) is a flow chart of the fine-tuning training phase of the present invention;

[0046] Figure 1 (b) Figure 1 Model learning sub-flowchart in (a);

[0047] Figure 2 A flowchart of the query identification phase of the present invention;

[0048] Figure 3 This is a visualization result diagram of a 0-degree pitch angle image as a query set in an embodiment of the present invention;

[0049] Figure 4 This is a visualization result diagram of a 20-degree pitch angle image as a query set in an embodiment of the present invention. [Specific implementation method]

[0050] The present invention is described in detail below with reference to the accompanying drawings and specific embodiments.

[0051] The present invention provides a seeker infrared re-identification method based on fine-tuning of the DINOv2 model, and the technology belongs to the field of image recognition. Compared with the existing deep learning re-identification method, the present invention introduces a fine-tuned DINOv2 pre-trained model, which significantly improves the feature representation ability and robustness of infrared image target re-identification. Combined with automatic gamma correction preprocessing, this technology deeply explores the detailed features of the target in the infrared image, and relies on the powerful representation learning ability of the DINOv2 pre-trained model to achieve efficient and accurate re-identification of infrared image targets under complex viewing angles and distance differences.

[0052] The present invention is divided into two stages: a fine-tuning training stage and a query recognition stage. In the fine-tuning training stage, the infrared image of the cooperative target is preprocessed (including automatic gamma correction, size specification and pixel normalization), a learning rate change strategy is set during the training process, and then a pre-trained DINOv2 backbone model and a post-randomly initialized classifier are used for fine-tuning. After the training is completed, a fine-tuned DINOv2 model is obtained.

[0053] In the query recognition stage, the infrared images of non-cooperative strike targets are first preprocessed by automatic gamma correction, size standardization and pixel normalization. Then, the fine-tuned DINOv2 model is used to extract features from the infrared images of strike targets to obtain discriminant features that are easy to classify various types of targets. Subsequently, the similarity between the features of the specified target image and the target image features in the candidate library is calculated, and the images in the library are sorted according to the similarity between the images to obtain the most similar image, thereby determining the identity of the specified image.

[0054] The detailed process of the fine-tuning training phase is as follows Figure 1 (a) and Figure 1(b) As shown. Due to the limitations of infrared imaging equipment or inappropriate ambient light, image preprocessing is an essential stage. In order to make the input image match the reasonable input data range of the deep learning model and prevent the nonlinear effects of the image such as overexposure, darkness, and uneven brightness from causing the model training to fail, we use an automatic gamma correction method based on the brightness average to adjust the brightness of the target image to a reasonable range. Generally, the expected brightness mean for the normalized image is generally 0.5, so we use the following gamma value calculation method to obtain the corresponding gamma value. The gamma value of this process is determined by converting the brightness average to the center of the histogram:

[0055]

[0056] Where X is the average brightness of the normalized image, X∈[0,1]. In the equation, 0.5 is the center of the brightness histogram. After obtaining the corresponding gamma value, use the following formula to perform gamma correction on the image to obtain an image with reasonable brightness:

[0057] S=R γ ,

[0058] Among them, S is the rectified image, R is the original normalized image, and its range is R∈[0, 1].

[0059] Since the target shooting angle and distance are different, the target image size is different. In order to meet the dimension of the model input data, the size is standardized to the specified size, i.e. (224, 224, 3), where the two 224 refer to the width and height of the image respectively, and 3 is the number of channels of the image. When converting a single-channel image to an RGB image, the grayscale value of each pixel will be copied to the three RGB channels, i.e. the R, G, and B values ​​of each pixel will be the same. The image also needs to be pixel normalized, using the mean and standard deviation to normalize the image. The normalized mean of each channel of the image is 0.5, and the standard deviation is 0.5. This helps to reduce numerical instability and gradient vanishing and explosion problems during model training, thereby accelerating the convergence of the model.

[0060] During fine-tuning, the weights of the DINOv2 model are initialized based on pre-trained parameters. These weights do not fully match the data characteristics of the current task. In order to prevent training oscillations during fine-tuning and improve model performance, we optimize the model training effect by introducing learning rate warm-up and decay strategies. In the early stage of model training, due to the difference between the training weights and the current training data characteristics, the training process is oscillating and unstable. A smaller learning rate is needed to make the model weight update more stable. The learning rate change rule of the warm-up process is shown in the following formula:

[0061]

[0062] Where lr refers to the learning rate, i refers to the current training round, a refers to the initial learning rate coefficient in the warm-up phase, a∈(0,1), lr i Refers to the learning rate of the current round, lr base Refers to the basic learning rate, epochs warmup Refers to the number of rounds required for the warm-up process.

[0063] As the training process progresses, the model will gradually approach the global optimal solution. If a high learning rate is continued, the model may miss the optimal solution, making it difficult for the training to converge to the optimal solution. Therefore, by using a learning rate decay strategy, the model can adjust parameters more finely in the later stages of training, thereby converging to a better solution. We use the cosine decay strategy, as follows:

[0064]

[0065] Among them, i refers to the current training round, epoch warmup Indicates the last round of warm-up, epochs max Represents the total number of training rounds.

[0066] Conventional methods generally update the learning rate after training with a batch of images, while the present invention reduces the frequency of learning rate updates and updates the learning rate once in each round. Compared with the learning rate change rule in the general method: warm-up-maintain-decay, the present invention removes the learning rate maintenance stage. The learning rate adjustment strategy of the present invention makes the training process more stable and converges faster.

[0067] In order to complete the training of the re-ID model, in addition to using the backbone of the DINOv2 pre-trained model, we also need to add an additional classification layer to complete the discriminative feature constraint task during the training process. The DINOv2 model is a Vision Transformer (ViT) model, which outputs the class token as the discriminative feature of the target image to calculate the similarity between images. In order to better constrain the discriminative features of the image, all tokens (including class tokens and patch tokens) are first layer-normalized, as shown in the following formula:

[0068]

[0069] Among them, E(x) represents the mean of the input data, Var(x) represents the variance of the input data, γ and β are learnable affine transformation parameters, and ∈ is a number close to 0 to prevent the divisor from being 0.

[0070] The classification layer uses a linear fully connected network to input x i×nA linear transformation is performed, namely:

[0071] y n×o =x n×i W i×o + b,

[0072] Where W is a learning parameter with a dimension of n×o, b is an o-dimensional bias vector, n is the number of rows in the input vector, which refers to the number of samples in a batch, and i is the number of input neurons. In the present invention, i=dim(token class ), o is the number of output neurons, and in the present invention, o=dim(the number of training target categories).

[0073] The main optimization goal of re-ID model training is to extract highly discriminative features, that is, to make the similarity of features between different images of the same target larger and the similarity of features between different target images smaller, so that the model can accurately identify the target individual in complex scenes regardless of changes in posture, lighting or viewing angle.

[0074] To achieve this goal, we introduce triplet loss with batch hard mining as the loss function. The basic triplet loss optimizes the distance distribution in the feature space by constructing triplets containing anchor samples, positive samples (belonging to the same category as the anchor samples), and negative samples (belonging to different categories from the anchor samples). It encourages the distance between the anchor sample and the positive sample to be smaller than the distance between the anchor sample and the negative sample, thereby enhancing the discriminability of the features. The design of this loss function helps the model better understand the similarities and differences between individuals at the feature level, and is an important means to improve the accuracy of re-identification. The triplet loss is shown below:

[0075] L triplet =max(d(a,p)-d(a,n)+margin,0),

[0076] Where d represents the distance between samples, a refers to the anchor sample, p refers to the positive sample, n refers to the negative sample, and margin is a constant greater than 0. This formula achieves the similarity calculation between samples by optimizing the distance between the anchor sample and the positive sample to be smaller than the distance between the anchor sample and the negative sample.

[0077] In order to better optimize the above goals, the hard sample mining sampling triple loss randomly selects P targets for each training batch, and randomly selects K different images for each target, that is, a batch contains P×K images. Then for the anchor sample in the training batch, we can select a sample with the largest distance from it in the positive sample, and a sample with the closest distance from it in the negative sample and the anchor sample to form a triplet, which can be expressed as the following formula:

[0078]

[0079] Among them, A is defined as the set of pictures with the same target as the anchor sample in a training batch, and B is defined as the set of pictures with different targets from the anchor sample in a training batch.

[0080] At the same time, cross-entropy loss is also incorporated into the training framework as a classification loss to further constrain and optimize the feature space. Classification loss guides model learning through label information so that the extracted features can effectively distinguish different individual categories. During the training process, the classification layer maps features to the category space and calculates the difference between the predicted label and the actual label through functions such as cross-entropy loss, thereby back-propagating the gradient and adjusting the model parameters. The cross-entropy loss function is shown below:

[0081]

[0082] Among them, x is the input, y is the input category label, w is the parameter weight, C is the total number of categories, and N is the number of images in each batch. The total loss function during training is as follows:

[0083] L sum =L trihard +L CE ,

[0084] By calculating the total loss function L sum , and back propagation is performed to gradually reduce the loss to a stable level, thereby achieving the optimization goal of fine-tuning training.

[0085] The detailed process of the query identification phase is as follows: Figure 2 As shown in the figure, for the non-cooperative images of the specified query, image preprocessing (including automatic gamma correction, size specification and pixel normalization) is also required to make the image input conform to the reasonable data range of the model. Then the fine-tuned DINOv2 model is used for feature extraction and layer normalization, and the category token features are selected for query. This feature vector is a compact representation of the image in the high-level semantic space, which can accurately reflect the intrinsic discriminative features of the image.

[0086] In order to measure the similarity between the query image and the image in the database, the present invention adopts a similarity measurement method, the Euclidean distance. The Euclidean distance directly reflects the relative position relationship of the feature vectors in the multidimensional space. The closer the distance, the more similar the two feature vectors are, and vice versa. The calculation formula is as follows:

[0087]

[0088] Among them, d Euclid represents the Euclidean distance between two feature vectors, n represents the dimension of the feature vector, x i and i Respectively represent the values ​​of the two feature vectors in the i-th dimension. In the present invention, we calculate the cosine similarity between the query image feature vector and the feature vector of each image in the database to obtain a similarity score list. The higher the similarity score, the closer the distance between the two images in the feature space, that is, the more similar they are. After obtaining the similarity score list, the present invention uses a sorting algorithm to sort the list in ascending order to find the image that is most similar to the query image (with the smallest distance). After the sorting is completed, we return the identity of the image with the smallest distance (maximum similarity) as the query recognition result.

[0089] Example:

[0090] In order to verify the effect of the invention, we selected the IRShips dataset for verification. This dataset is a synthetic infrared image dataset. The images are generated using a virtual camera in the 8-14μm band in the CounterSim simulation program. There are 10 types of ships in total. This dataset comes from the paper "Deeplearning for automatic target recognition with real and synthetic infraredmaritime imagery" published by Westlake, Samuel T. et al. in 2020. The main parameters of the ships captured in this dataset are:

[0091] 1. The distance of the camera relative to the ship is 1000, 1111, 1250, 1429, 1667, 2000, 2500, 3333,

[0092] 5000,10000m;

[0093] 2. The camera’s azimuth angle relative to the ship is 0-359 degrees;

[0094] 3. The camera's pitch angle relative to the ship is 0, 10 and 20 degrees.

[0095] In order to demonstrate the effect of the re-identification method of the present invention on the dataset, we selected ship images at distances of 1000, 1111, 1250, 1429, 1667, 2000, 2500, 3333, and 5000 meters; in order to better demonstrate the recognition performance of the present invention between different viewing angles, we selected the pitch angles of 0 and 20 degrees with the largest difference to verify the recognition effect of the present invention; for a target ship, 180 images were uniformly sampled at the above two pitch angles, for a total of 360 images. During the verification, we cropped the synthesized image, that is, using the coordinates of the ship in the image provided by the dataset, we cropped the main body of the ship in the image for recognition.

[0096] We selected 6 types of ships (Ada, Independence, Visby, Akizuki, Sejong Daewang, Zumwalt) as cooperative targets and 4 types of ships as non-cooperative targets (Alvaro De Bazan, Jiangkai II, Oliver Hazard Perry, Armourique). The 6 types of cooperative target images are applied to the fine-tuning training stage of the first stage of the present invention to obtain a fine-tuned trained model, and the non-cooperative targets do not participate in the fine-tuning training process; in the second stage, the 4 types of non-cooperative target images are divided into a query set and a candidate library according to the pitch angle, and the query set and the candidate library are respectively input into the above model to obtain the corresponding features, and the distance between the query set features and the candidate library features is calculated and sorted to obtain the corresponding identity of the query and the image.

[0097] The effect demonstration of the present invention pays special attention to the impact of pitch angle changes on the re-identification performance of infrared ship images. The visualization results and test performance data of the embodiment are based on the dinov2_vits14_pretrain sub-model of the DINOv2 pre-trained model. The test process uses 0-degree pitch angle and 20-degree pitch angle images of 4 types of non-cooperative targets as query sets, and the corresponding remaining data as candidate libraries. The query visualization results are shown in Figure 2. Figure 3 and Figure 4 The results shown are randomly selected. Figure 3 The visual query results are as follows: the image with a pitch angle of 0 degrees is used as the query set, and the image with a pitch angle of 20 degrees is used as the candidate library; Figure 4 The 20-degree pitch angle images are used as the query set, and the 0-degree pitch angle images are used as the visualization query results of the candidate library.

[0098] The image in the leftmost column of the visualization is the anchor sample, and the images in the 5 columns on the right are candidate images that are increasingly dissimilar to the anchor sample from left to right. The green frame of the candidate image is the sample that is correctly identified, and the red frame is the sample that is incorrectly identified. The image identification information in the visualization result is: the first line is the target category; the second line is the pitch angle, azimuth angle, and distance; the third line of the candidate image is the feature distance between the sample and the anchor sample (the smaller the distance, the higher the similarity).

[0099] pass Figure 3 and Figure 4 It can be observed that the present invention can correctly identify the identity of the target when the target perspective in the image is greatly different, and the first five candidate images are basically correct, indicating that the retrieval efficiency and performance of the present invention are very good. Even in the face of complex and changeable perspective changes, the present invention can still maintain a high degree of accuracy and stability, which is difficult to match with traditional methods. The results shown in this visualization cannot represent the overall performance of the present invention under all data.

[0100] The re-identification performance data of the present invention under the above data parameter settings are as follows:

[0101] Table 1 Test performance data

[0102]

[0103] Rank-n in the table indicates the probability of having a correct result in the first n pictures with the highest confidence in the search library. Rank-n can reflect the first hit rate and the first n hit rates of the model in the retrieval task. The higher the value, the better the accuracy of the re-identification method. When n=1, Rank-1 is the first hit rate, which reflects whether the model can accurately find the target in the first retrieval. The Rank-1 of the present invention in both cases reaches more than 80%, indicating that the first hit rate is high, that is, the model has a strong retrieval ability under complex perspective differences and can quickly retrieve the target. When n>1, Rank-n indicates whether the target can be found in the first n retrievals. The first 5 and 10 hit rates reach more than 90%, indicating that the model can still maintain a high accuracy in multiple retrievals, which is crucial to user experience in practical applications.

[0104] mAP is the mean average precision, which evaluates the model performance by calculating the average precision of each category and taking the average. In the re-identification task, mAP can reflect the average performance of the model on the entire data set. The higher the mAP value, the better the accuracy and robustness of the re-identification method. The mAP values ​​of the present invention in both cases are about 65%, indicating that the present invention has excellent retrieval accuracy under complex perspective changes, and can still maintain good performance and high retrieval efficiency when facing complex and changeable data.

[0105] In order to demonstrate the advanced performance of the present invention, a comparative experiment was conducted using an existing general model, the ResNet50 network structure pre-trained model on the ImageNet dataset. The comparative experiment followed the first and second stage processes proposed by the present invention and used the above-specified dataset, but the automatic gamma correction technology was not applied in either stage. The comparative experiment performance data is as follows:

[0106] Table 2 Comparative experimental test performance data

[0107]

[0108] By comparing the test performance of the present invention with that of the comparative experiment, the present invention has achieved significant improvements in multiple key performance indicators. When the pitch angle of 0 degrees is used as the query library, the present invention improves the Rank-1 accuracy by about 14% compared with the comparative method; there is a certain improvement in the Rank-5 and Rank-10 accuracy. When the pitch angle of 20 degrees is used as the query library, the present invention improves the Rank-1 performance by about 23% compared with the comparative method, and the performance of Rank-5 and Rank-10 is also significantly improved.

[0109] What is particularly important is that in terms of the comprehensive measurement indicator of mean average precision (mAP), the method of the present invention achieved a significant improvement of more than 20% in both cases compared to the comparison method. This result verifies the effectiveness and advancement of the present invention. Compared with the comparison method, the performance results shown by the method proposed in the present invention are almost consistent when using images of different pitch angles as the query set. However, there are significant differences in the performance of the comparison method in these two scenarios. This phenomenon strongly proves that the method of the present invention can effectively reduce the differences in image features caused by changes in viewing angles. In other words, the image features extracted by the re-identification method of the present invention are highly robust.

Claims

1. A seeker infrared re-identification method based on DINOv2 model fine-tuning, characterized in that: It includes a fine-tuning training phase and a query recognition phase; The fine-tuning training phase includes the following: Preprocess the infrared image of the cooperative target; set the learning rate change strategy during the training process; then use the pre-trained DINOv2 backbone model and post-randomly initialized classifier for fine-tuning, and obtain the fine-tuned DINOv2 model after the training; The query identification stage includes the following contents: Preprocessing the infrared image of the non-cooperative target; extracting features from the infrared image of the non-cooperative target using the fine-tuned DINOv2 model to obtain discriminant features that are easy to classify for various types of targets; Calculate the similarity between the features of the specified target image and the target image features in the candidate library, sort the images in the candidate library according to the similarity between the images, and obtain the most similar image, thereby determining the identity of the specified image.

2. A seeker infrared re-identification method based on DINOv2 model fine-tuning as described in claim 1, characterized in that: The preprocessing includes automatic gamma correction, size specification and pixel normalization, wherein the corresponding gamma value is obtained using the following gamma value calculation method, and the gamma value of the process is determined by converting the brightness average value to the center of the histogram: Among them, X is the average brightness of the normalized image, X∈[0,1], 0.5 is the center of the brightness histogram, and after obtaining the corresponding gamma value, the image is gamma corrected using the following formula to obtain an image with reasonable brightness: S=R γ , Among them, S is the rectified image, R is the original normalized image, and its range is R∈[0, 1].

3. A seeker infrared re-identification method based on DINOv2 model fine-tuning as described in claim 2, characterized in that: The specific process of the size standardization is: standardize the size to a specified size, namely (224, 224, 3), wherein the two 224s refer to the width and height of the image respectively, and 3 is the number of channels of the image.

4. A seeker infrared re-identification method based on DINOv2 model fine-tuning as described in claim 2 or 3, characterized in that: The specific process of pixel normalization is: using mean and standard deviation to normalize the image, the normalized mean of each channel of the image is 0.5, and the standard deviation is 0.

5.

5. A seeker infrared re-identification method based on DINOv2 model fine-tuning as described in claim 1, characterized in that: The learning rate change strategy in the fine-tuning training phase introduces a learning rate warm-up and decay strategy to optimize the model training effect: In the early stage of model training, a smaller learning rate is needed to make the model weight update more stable. The learning rate change rule of the warm-up process is shown in the following formula: Where lr refers to the learning rate, i refers to the current training round, a refers to the initial learning rate coefficient in the warm-up phase, a∈(0,1), lr i Refers to the learning rate of the current round, lr base Refers to the basic learning rate, epochs warmup Refers to the number of rounds required for the warm-up process; By using the learning rate decay strategy, the model can adjust parameters more finely in the later stages of training, thereby converging to a better solution; The specific cosine attenuation strategy is as follows: Among them, i refers to the current training round, epoch warmup Indicates the last round of warm-up, epochs max Represents the total number of training rounds.

6. A seeker infrared re-identification method based on DINOv2 model fine-tuning as described in claim 1, characterized in that: The method for fine-tuning the DINOv2 model in the fine-tuning training phase is: The backbone of the DINOv2 training model is selected to extract features from the input target image and obtain feature representations containing category tokens; Perform layer normalization on all tokens output by the backbone (including category tokens and tile tokens); Add a classification layer and use a linear fully connected network to perform linear transformation on the normalized features; Determine the optimization goal of re-ID model training, that is, to make the similarity of features between different images of the same target greater and the similarity of features between different target images smaller; Using the total loss function, the parameters of the model are continuously adjusted through the back propagation algorithm to gradually reduce the total loss to a stable level; the total loss function is the sampling triplet loss function of difficult sample mining and the cross entropy loss function.

7. A seeker infrared re-identification method based on DINOv2 model fine-tuning as described in claim 6, characterized in that: The loss function of the difficult sample mining sampling triplet is as follows: Where d represents the distance between samples, a refers to anchor samples, p refers to positive samples, n refers to negative samples, margin is a constant greater than 0, P is the number of targets in each batch, K different images are randomly selected for each target, A is defined as the set of images with the same target as the anchor sample in a training batch, and B is defined as the set of images with different targets from the anchor sample in a training batch.

8. A seeker infrared re-identification method based on DINOv2 model fine-tuning as described in claim 1, characterized in that: The specific method of the query identification stage is: Preprocessing the infrared image of the non-cooperative target, wherein the preprocessing includes gamma correction, size standardization and image normalization; Use a fine-tuned DINOv2 model for feature extraction and perform layer normalization; Select the category token features to query; Calculate the Euclidean distance between the query image feature vector and each image feature vector in the database to obtain a feature distance list; The smaller the feature distance is, the closer the distance between the two images is in the feature space, that is, the more similar they are; the identity of the image with the smallest distance is returned as the query recognition result.