Self-supervision target detection method based on rid

Through the object detection method based on reid self-supervised, the multi-attribute reid matching algorithm and loss function adjustment are used to solve the problem of training set limitations and noise interference, and the feature extraction ability and training efficiency of the object detection model are improved.

CN120299047APending Publication Date: 2025-07-11LINKER
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510317096.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-18
Publication Date
2025-07-11

AI Technical Summary

Technical Problem

The training set production of existing object detection models has limitations, high manual calibration costs, and a lot of invalid information in the data set, resulting in weak model feature extraction capabilities and affecting the overall detection effect.

Method used

The object detection method based on reid self-supervised is adopted, and the multi-attribute reid matching algorithm is used to filter effective pseudo-labels to ensure that the model reduces noise interference during training, increases difficult sample learning, and prevents overfitting.

Benefits of technology

It realizes reducing noise interference in self-supervised training, improving the effectiveness of pseudo-labels, preventing the model from overfitting simple samples, and ensuring model training efficiency and accuracy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure BDA0005316250940000041
    Figure BDA0005316250940000041
  • Figure BDA0005316250940000071
    Figure BDA0005316250940000071
  • Figure FDA0005316250930000031
    Figure FDA0005316250930000031
Patent Text Reader

Abstract

The invention discloses a self-supervision target detection method based on rid. Firstly collecting and processing data and classifying; reducing the confidence of the target detection model to obtain related information; inputting the target-containing picture into a large oracle model to obtain coordinates and description; and performing IOU matching and recording on prediction results of the two models. And performing similarity matching with a database, and processing according to a result. Carrying out matting on unmatched category data, inputting the matted data into the clip model and recording the matted data; and related unmatched data matting is input into a rid network record. And matting the low-similarity data, inputting the rid, and recording according to the similarity. Then, data sampling is carried out to ensure that specific targets and the number are consistent, a loss training model is calculated, easily-detected samples are removed, and difficult cases and background negative samples are mined until the model is stable for detection; according to the method, through multi-step cooperation, data interaction processing among models is utilized, effective training data is mined, the adaptability of the model to complex conditions is enhanced, the accuracy and stability of the target detection model are improved, and the overall detection performance is optimized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to an object detection method, and more specifically to a reid self-supervised object detection method. Background Art

[0002] In object detection algorithms, an effective training set is one of the important factors for improving the object detection model. However, collecting, calibrating, and evaluating an effective data set requires a large amount of manpower and resources. Currently, there are the following problems in training set production: ① The data set has limitations, resulting in limitations of the trained object detection model; ② Manual calibration increases the training cost and time consumption of the model; ③ The collected training set may contain invalid information, increasing the training time consumption. Although the current object detection model adds text features to achieve zero-shot detection, the visual module is restricted by data, resulting in weak feature extraction ability, thus affecting the overall object detection ability. Summary of the Invention

[0003] Aiming at the deficiencies of the existing technology, the purpose of the present invention is to provide a reid self-supervised object detection method to solve the technical problems proposed in the background art.

[0004] To achieve the above purpose, the present invention provides the following technical solution: A reid self-supervised object detection method includes the following steps:

[0005] Step 1, collect data, and then process the collected data, classify the data, and obtain different groups of the same category;

[0006] Step 2, reduce the confidence of the object detection model to obtain the coordinates, categories, and scores of a large number of objects;

[0007] Step 3, input the image with an object into a large prediction model to obtain the coordinates and description information of each object in the image;

[0008] Step 4, perform IOU matching on the object coordinates predicted by the large language model in Step 3 and the prediction results of the object detection model. If no match is found, record it as no_match_iou_data1. If a match is found but the category does not match, record it as no_match_label_data1. If both the category and the coordinates match, record it as match_data1;

[0009] Step 5, perform similarity matching on the match_data1 recorded in Step 4 with the database. If the similarity > threshold, do not add it to the database, denoted as match_data_low_2. If ≤ threshold, add it to the database, denoted as match_data_high_2;

[0010] Step 6: Crop the no_match_label_data1 data recorded in Step 4 to obtain an image with only one target, and then input it into the CLIP multi-modal classification model to extract visual features and perform scale conversion operations, which are recorded as no_match_label_data_reid_2 and no_match_label_data_reid_unmatch_2;

[0011] Step 7: Crop no_match_label_data_reid_unmatch_2 and no_match_label_data1 to obtain the target of an image, input it into the ReID network, calculate the similarity between the obtained feature vectors and the database, and record the results according to the similarity calculation as no_match_reid_3 and no_match_unreid_3;

[0012] Step 8: Crop match_data_low_2 and input it into ReID. As in Step 7, obtain the average similarity. If it is greater than threshold 2, record it as match_data_low_hard_4; if it is less than or equal to threshold 2, record it as match_data_low_unreid_4;

[0013] Step 9: Perform data sampling to ensure that each iteration of the data must contain the targets of match_data_high_2, no_match_reid_3, no_match_label_data_reid_2, match_data_low_hard_4, and match_data_low_unreid_4. Ensure that the amount of target data remains the same for each iteration. If the quantity is insufficient, use the copy operation to copy and paste the images trained in the previous rounds to ensure the quantity is consistent;

[0014] Step 10: Calculate the loss and train the model simultaneously, and during the model training process, eliminate easy-to-detect sample targets, mine difficult example samples, and background negative samples;

[0015] Step 11: Continuously execute the above steps until the model becomes stable, and finally use the stable model for object detection.

[0016] As a further improvement of the present invention, the specific method for classifying the data in Step 1 is as follows: For samples of the same category, use the image-image ReID method to extract feature extraction methods to obtain the similarity values between each image and other images. Then, according to the clustering method based on the similarity values, classify the images of the same category to obtain different groups of the same category.

[0017] As a further improvement of the present invention, the specific steps for performing IOU matching in Step 4 are as follows:

[0018] Step 4-1: Input the category predicted by the detection network into Bert to obtain the category encoding;

[0019] Step 4-2: Input the description information predicted by the large model language into Bert to obtain the description encoding;

[0020] Step 4-3: Calculate the cosine similarity between the category encoding and the description encoding, as shown in formula (1)

[0021] sim_val 文本-文本 = seq det_文本 · seq 大语言模型文本 (1)

[0022] In formula (1), it represents the text-text reid comparison method.

[0023] As a further improvement of the present invention, the specific steps for performing scale conversion operation on the extracted visual features in Step 6 are as follows:

[0024] Step 6-1: Input the obtained N*1 feature vector and the no_match_label_data1 category into the clip multi-modal classification model, extract the text features for scale conversion operation, and obtain an N*1 feature vector;

[0025] Step 6-2: Then calculate the cosine similarity between the two vectors in Step 6-1, as shown in formula (2):

[0026] sim_val 文本-视觉 = seq 文本 · seq 视觉 (2)

[0027] 1. Formula 2 represents the similarity calculation between text and visual features. If sim_val 文本-视觉 > threshold 2, record no_match_label_data_reid_2, otherwise record no_match_label_data_reid_unmatch_2.

[0028] As a further improvement of the present invention, the specific method for calculating the similarity between the feature vector and the database in Step 7 is as follows:

[0029] The obtained feature vectors are used to calculate the similarity with the database, and then the similarities of the same category and the same group are averaged to obtain the average similarity. If the similarity is greater than threshold 2, it is recorded as no_match_reid_3; if it is less than or equal to threshold 2, it is recorded as no_match_unreid_3, and the recorded label is ignored, as shown in formula (3).

[0030]

[0031] As a further improvement of the present invention, the specific steps for calculating the loss in step ten are as follows:

[0032] Step 1, if it matches match_data_high_2, calculate the match_data_high_2 loss. The method is to set a relatively high weight during the early stage of model training and reduce the weight in the later stage of model training, as shown in formula (4).

[0033] loss1 = weight1 × (min(loss list_obj ) + loss label ) (4)

[0034] Formula (4) indicates that the results predicted by the object detection model and the large language model are consistent, and the objects with high confidence are used for pseudo-label loss calculation. Weight1 represents the weight of the object. In the early stage, in order to accelerate model convergence, the value of weight1 is set relatively high. When the model accuracy reaches threshold 3, the value of weight1 is set to be relatively high to let the object detection model learn to reduce the learning of easily detectable samples. loss list_obj represents the regression loss value between the coordinates predicted by the object detection model in the previous round, the coordinates of the prediction result of the large model, and the coordinates of the prediction result in this round; loss label represents the class loss value;

[0035] Step 2, if it matches match_data_low_hard_4, calculate the match_data_low_hard_4 loss, specifically as shown in formula (5):

[0036] loss2 = weight2 × (min(loss list_obj ) + loss label ) (5)

[0037] Formula (5) represents the calculation of the hard example sample loss, and weight2 represents the weight of the object, which increases the model's learning of this object;

[0038] Step 3, if no_match_label_data_reid_2 and no_match_reid_3 are matched, calculate the losses of no_match_label_data_reid_2 and no_match_reid_3. Adopt multi-class loss calculation, select the loss of the smallest class and set a lower weight to guide the model learning, as shown in formula (6) specifically:

[0039] loss4 = weight4 × (min(loss list_obj ) + min(loss label_list )) (6)

[0040] Step 4, if no_match_unreid_3 is matched, it means that the model cannot determine and no loss is calculated. If the model does not match the above records, no loss is calculated either.

[0041] As a further improvement of the present invention, the specific methods for removing easily detectable sample targets, mining difficult example samples and background negative samples in the tenth step are as follows:

[0042] For easily detectable samples: If the target appears in match_data_high_2 continuously for N times, remove the target label and ignore it without calculating the loss;

[0043] For mining difficult example samples: If the target appears in match_data_low_hard_4 or match_data_low_unreid_4 datasets continuously for N times, add its target box to the match_data_low_hard_4 dataset;

[0044] For background negative samples: If the target appears in no_match_unreid_3 continuously for N times, it indicates background noise.

[0045] Advantages of the present invention:

[0046] 1. Online calibration of the multi-attribute reid matching algorithm, and correction of the text-visual reid matching algorithm, visual-visual reid matching algorithm and text-text reid matching algorithm in the target detection algorithm during the training process, realizing the reduction of noise interference in the self-supervised training process of target detection and ensuring the effectiveness of pseudo-labels;

[0047] 2. Combining the online calibration of the multi-attribute reid matching algorithm and the loss change of the model during the training process, realizing the interference of noise and the training of effective data, so as to reduce the training time-consuming and prevent the target detection model from overfitting to simple samples and underfitting to difficult example samples;

[0048] 3. Online calibration is performed according to the multi-attribute reid matching algorithm to obtain the number of easy and difficult samples. Based on the easy and difficult samples, the dataset sampling is modified to ensure the difficulty of the samples in each iteration, preventing oscillations and local minima during model training.

[0049] 4. Weights are added to the loss function. During model training, different weight values are assigned to each target according to the loss value and multi-template supervision, increasing the learning of difficult example samples by the detection model and reducing the learning of noisy samples. Detailed implementation method

[0050] The following examples will be used to further elaborate on the present invention.

[0051] A reid self-supervised object detection method in this embodiment is mainly completed by the following eleven steps.

[0052] Step 1: Collect a database and try to collect samples in different scenarios. For samples of the same category, the image-image reid method (swin-reid is used in this embodiment) is adopted for feature extraction, obtaining the similarity value between each image and other images. Then, according to the clustering method (k-mean is used in this embodiment), the images of the same category are classified based on the similarity value, obtaining different groups of the same category.

[0053] Step 2: The object detection model reduces the confidence level to obtain the coordinates, categories, and scores of a large number of objects.

[0054] Step 3: Input the images with objects into the large prediction model to obtain the coordinates and description information of each object in the images.

[0055] Step 4: The object coordinates predicted by the large language model are matched with the prediction results of the object detection model using IOU (IOU is set to 0.5). If no match is found, it is recorded as no_match_iou_data1. If a match is found but the category does not match, it is recorded as no_match_label_data1. If both the category and coordinates match, it is recorded as match_data1; category matching method: 1) Input the category predicted by the detection network into Bert to obtain the category encoding; 2) Input the description information predicted by the large model language into Bert to obtain the description encoding; 3) Calculate the cosine similarity between the category encoding and the description encoding, as shown in formula (1)

[0056] sim_val 文本-文本 =seq det_文本 ·seq 大语言模型文本 (1)

[0057] In formula (1), it represents the text-text reid comparison method.

[0058] Step 5: Perform similarity matching between match_data1 and the database. If the similarity > threshold, do not add it to the database, denoted as match_data_low_2. If ≤ threshold, add it to the database, denoted as match_data_high_2;

[0059] Step 6: Crop the no_match_label_data1 data to obtain a picture with only one target, input it into the clip multi-modal classification model, extract visual features and perform scale conversion operations to obtain an N*1 feature vector, and input the no_match_label_data1 category into the clip multi-modal classification model, extract text features and perform scale conversion operations to obtain an N*1 feature vector. Then calculate the cosine similarity between the two vectors, as shown in formula (2):

[0060] sim_val 文本-视觉 = seq 文本 ·seq 视觉 (2)

[0061] Formula 2 represents the similarity calculation between text and visual features. If sim_val 文本-视觉 > threshold 2 (set to 0.5 in this embodiment), record no_match_label_data_reid_2, otherwise record

[0062] no_match_label_data_reid_unmatch_2

[0063] Step 7: Crop no_match_label_data_reid_unmatch_2 and no_match_label_data1 to obtain the target of a picture, input it into the reid network (set to swin_reid in this embodiment), calculate the similarity between the obtained feature vector and the database, and then calculate the average of the similarities of the same category and the same group to obtain the average similarity. If it is greater than threshold 2, record it as no_match_reid_3. If it is less than or equal to threshold 2, record it as no_match_unreid_3, and the recorded label is ignored, as shown in formula (3)

[0064]

[0065] Step 8: Crop match_data_low_2 and input it into reid. As in step 7, obtain the average similarity. If it is greater than threshold 2, record it as match_data_low_hard_4. If it is less than or equal to threshold 2, record it as match_data_low_unreid_4

[0066] Step 9: Data Sampling. In each iteration, the data must contain the targets of match_data_high_2, no_match_reid_3, no_match_label_data_reid_2, match_data_low_hard_4, and match_data_low_unreid_4. Ensure that the amount of target data remains consistent in each iteration. If the quantity is insufficient, perform a copy operation by copying and pasting the images trained in the previous rounds to ensure the quantity is consistent.

[0067] Step 9: Loss Calculation:

[0068] Compare the confidence levels of the match_data_high_2 dataset. The probability of noise is low. During the early stage of model training, set a relatively high weight. In the later stage of model training, reduce the weight. The reason is to accelerate model convergence in the early stage and let the model learn more difficult example samples and reduce learning of simple samples in the later stage, as shown in formula (4).

[0069] loss1 = weight1 × (min(loss list_obj ) + loss label ) (4)

[0070] Formula (4) indicates that the results predicted by the object detection model and the large language model are consistent. Calculate the pseudo-label loss for the targets with relatively high confidence. Weight1 represents the weight of the target. To accelerate model convergence in the early stage, the value of weight1 is set relatively high (set to 1.2 in this embodiment). When the model accuracy reaches the threshold 3 (set to 0.5 in this embodiment), set the value of weight1 relatively high (set to 0.8 in this embodiment) to let the object detection model learn to reduce the learning of easily detectable samples. loss list_obj represents the regression loss value between the coordinates predicted by the object detection model in the previous round, the coordinates of the prediction result of the large model, and the coordinates of the prediction result in this round; loss label represents the class loss value

[0071] Calculate the loss of match_data_low_hard_4. This dataset is a difficult example dataset. The large language model, the object detection model, and the vision-vision reid are all correct. Therefore, it is necessary to enhance the learning ability of this target throughout the process, as shown in formula (5).

[0072] loss2 = weight2 × (min(loss list_obj ) + loss label ) (5)

[0073] Formula 5 represents the loss calculation of difficult sample, weight2 represents the weight of the target, and in this embodiment, it is set to 1.2 to increase the weight of the model to learn the target;

[0074] match_data_low_unreid_4 indicates that data may not exist in the database, but is predicted by both the large language model and the target detection model, indicating that there is a small probability of noise. Therefore, when calculating the loss, a relatively small weight is set to guide the model to learn the target on the one hand, and prevent the model from learning too much and preventing the target from being noise on the other hand. This embodiment sets 0.5

[0075] loss3=weight3×(min(loss list_obj )+loss label ) (6)

[0076] no_match_label_data_reid_2 and no_match_reid_3 indicate that there is much uncertainty in the categories and coordinates. Therefore, multi-category loss calculation is adopted, the loss of the smallest category is selected and a lower weight is set to guide model learning. In this embodiment, it is set to 0.2.

[0077] loss4=weight4×(min(loss list_obj )+min(loss label_list )) (6)

[0078] no_match_unreid_3 means that the model cannot determine. If this area is matched, no loss is calculated. If it is not matched, no loss is calculated either.

[0079] Step 10: During the model training process, remove easy-to-detect sample targets and mine difficult-to-detect samples and background negative samples.

[0080] Easy-to-detect samples: If a target appears in match_data_high_2 N times in a row (5 times in this embodiment), the target is removed and marked as ignored, and no loss calculation is performed.

[0081] Mining difficult examples: If the target appears in the match_data_low_hard_4 or match_data_low_unreid_4 dataset N times (5 times in this embodiment), its target box is added to the match_data_low_hard_4 dataset;

[0082] Background negative samples: If the target appears in no_match_unreid_3 N times (5 times in this embodiment) in a row, it means it is background noise.

[0083] Step Eleven: Continuously perform the above operations until the model converges to stability.

[0084] In summary, the reid self-supervised object detection method of this embodiment;

[0085] 1. Online calibration of the multi-attribute reid matching algorithm: It is corrected by the text-visual reid matching algorithm, visual-visual reid matching algorithm, and text-text reid matching algorithm respectively. (1) Text-visual reid matching algorithm: The object coordinates and the description of the object are obtained through the object detection model (in this patent, a multi-modal object detection network such as (Grounding DINO) is used). Based on the object coordinates, image cropping is performed to obtain the image of the predicted object. The image of the predicted object is input into the large model (in this patent, qwen2-72b is adopted) to obtain the description of the image. Then, the text description obtained by the large model and the object description (category) output by the object detection model are input into the text model (BERT) for encoding to obtain text features, and similarity comparison is performed. (2) Visual-visual reid matching algorithm: Image cropping is performed on the prediction result, feature extraction is performed to obtain image features, and the image features are compared with the database features to observe whether they are consistent; (3) Text-text reid matching algorithm: Image cropping of the prediction result is input into the visual part of clip, and the predicted category is input into the text part of clip to match the similarity of visual features and text features.

[0086] 2. Screening effective pseudo-label objects. Based on the change of the loss value during the training process of the object detection model and the multi-attribute reid matching verification, more effective pseudo-label data is obtained, realizing two functions. ① Reduce the noise interference in model training, improve the accuracy of pseudo-labels, and thus improve the accuracy of the model; ② Distinguish between difficult example samples and simple samples in the training set, and prevent too many simple samples from participating in model training, resulting in overfitting of the detection network model to simple samples and difficult example samples, thus deteriorating the effect of the object detection model.

[0087] 3. Modify the dataset sampling, ensuring the difficulty of samples in each iteration to prevent oscillations and local minima during model training.

[0088] 4. Modify the loss function, add weights for samples with low confidence to prevent noise and reduce interference to the model.

[0089] The above is only the preferred embodiment of the present invention. The protection scope of the present invention is not limited to the above embodiments. All technical solutions within the idea of the present invention belong to the protection scope of the present invention. It should be noted that for those of ordinary skill in the art in this technical field, several improvements and refinements made without departing from the principle of the present invention should also be regarded as within the protection scope of the present invention.

Claims

1. A self-supervised object detection method based on reid, characterized in that: It includes the following steps: Step 1: Collect data, then process the collected data, classify the data, and obtain different groups of the same category; Step 2: Reduce the confidence of the object detection model to obtain the coordinates, categories, and scores of a large number of objects; Step 3: Input the images with objects into the large prediction model to obtain the coordinates and description information of each object in the images; Step 4: Perform IOU matching between the object coordinates predicted by the large language model in Step 3 and the prediction results of the object detection model. If no match is found, record it as no_match_iou_data1. If a match is found but the category does not match, record it as no_match_label_data1. If both the category and coordinates match, record it as match_data1; Step 5: Perform similarity matching between the match_data1 recorded in Step 4 and the database. If the similarity > threshold, do not add it to the database and record it as match_data_low_2. If ≤ threshold, add it to the database and record it as match_data_high_2; Step 6: Crop the no_match_label_data1 data recorded in Step 4 to obtain images with only one object, then input them into the clip multi-modal classification model, extract visual features and perform scale conversion operations, and record them as no_match_label_data_reid_2 and no_match_label_data_reid_unmatch_2; Step 7: Crop the no_match_label_data_reid_unmatch_2 and no_match_label_data1 to obtain the objects of one image, input them into the reid network, calculate the similarity between the obtained feature vectors and the database, and record according to the similarity calculation results as no_match_reid_3 and no_match_unreid_3; Step 8: Crop the match_data_low_2 and input it into the reid. As in Step 7, obtain the average similarity. If it is greater than threshold 2, record it as match_data_low_hard_4. If it is less than or equal to threshold 2, record it as match_data_low_unreid_4; Step 9: Perform data sampling to ensure that each iteration of the data must contain match_data_high_2, no_match_reid_3, no_match_label_data_reid_2, match_data_low_hard_4, and match_data_low_unreid_4 objects. Ensure that the amount of object data remains consistent in each iteration. If the quantity is insufficient, use the copy operation to copy and paste the images trained in each previous round to ensure the quantity is consistent; Step 10: Calculate the loss and train the model at the same time, and remove easy-to-detect sample targets and mine difficult-to-detect samples and background negative samples during the model training process; Step 11: Continuously execute the above steps until the model becomes stable, and finally use the stable model for target detection.

2. The Reid self-supervised object detection method according to claim 1, characterized in that: The specific method of classifying the data in step 1 is: using the image-image reid method to extract features from samples of the same category to obtain similarity values ​​between each image and other images, and then classifying the images of the same category according to the similarity values ​​using a clustering method to obtain different groups of the same category.

3. The Reid self-supervised object detection method according to claim 1 or 2, characterized in that: The specific steps for performing IOU matching in step 4 are as follows: Step 41: Input the category predicted by the detection network into Bert to obtain the category code; Step 42, input the description information predicted by the large model language into Bert to obtain the description code; Step 43: Calculate the cosine similarity of the category code and the description code, as shown in formula (1) sim_val 文本-文本 = seq det_文本 · seq 大语言模型文本 (1) Formula (1) represents the text-to-text re-ID comparison method.

4. The Reid self-supervised object detection method according to claim 1 or 2, characterized in that: The specific steps of extracting visual features and performing scale conversion operation in step 6 are as follows: step 61, inputting the obtained N*1 feature vector and the no_match_label_data1 category into the clip multimodal classification model, extracting text features and performing scale conversion operation to obtain an N*1 feature vector; step 62, then performing cosine similarity calculation on the two vectors in step 61, as shown in formula (2): sim_val 文本-视觉 = seq 文本 · seq 视觉 (2) Formula 2 represents the similarity calculation of text and visual features. If sim_val 文本-视觉 > the threshold 2, record no_match_label_data_reid_2; otherwise, record no_match_label_data_reid_unmatch_2.

5. The Reid self-supervised object detection method according to claim 4, characterized in that: The specific method of calculating the similarity between the feature vector and the database in step 7 is as follows: The obtained feature vector is used to calculate the similarity with the database, and then the similarity of the same category and the same group is averaged to obtain the average similarity. If it is greater than the threshold 2, it is recorded as no_match_reid_3, and if it is less than or equal to the threshold 2, it is recorded as no_match_unreid_3. The label of the record is ignored, as shown in formula (3).

6. The Reid self-supervised object detection method according to claim 1 or 2, characterized in that: The specific steps for calculating the loss in step 10 are as follows: Step 1: If the match is match_data_high_2, the match_data_high_2 loss is calculated by setting the weight to be high in the early model training and reducing the weight in the later stage of model training, as shown in formula (4): loss1 = weight1 × (min(loss list_obj ) + loss label ) (4) Formula (4) indicates that the results predicted by the object detection model and the large language model are consistent. For objects with relatively high confidence, pseudo-label loss calculation is performed. Weight1 represents the weight of the object. In the early stage, to accelerate model convergence, the value of weight1 is set relatively high. When the model accuracy reaches threshold 3, the value of weight1 is set relatively high to enable the object detection model to learn to reduce the learning of easily detectable samples, loss list_obj represents the regression loss value between the coordinates predicted by the object detection model in the previous round, the coordinates of the prediction result of the large model, and the coordinates of the prediction result in this round; loss label represents the class loss value; Step 2: If the match is match_data_low_hard_4, the match_data_low_hard_4 loss is calculated, as shown in formula (5): loss2 = weight2 × (min(loss list_obj ) + loss label ) (5) Formula (5) represents the loss calculation of difficult sample, weight2 represents the weight of the target, which increases the model's learning of the target; Step 3: If no_match_label_data_reid_2 and no_match_reid_3 are matched, the loss calculation of no_match_label_data_reid_2 and no_match_reid_3 is performed. Multi-category loss calculation is adopted, the loss of the smallest category is selected and a lower weight is set to guide model learning, as shown in formula (6): loss4 = weight4 × (min(loss list_obj )) + min(loss label_list )) (6) Step 4, if no_match_unreid_3 is matched, it means that the model cannot determine and no loss is calculated. If the model does not match any of the above records, no loss is calculated either.

7. The method for reid self-supervised object detection according to claim 1 or 2, characterized in that: The specific methods for eliminating easily detectable sample targets, mining difficult example samples, and background negative samples in Step 10 are as follows: For easily detectable samples: If the target appears in match_data_high_2 continuously for N times, the target label is eliminated and ignored, and no loss calculation is performed. For mining difficult example samples: If the target appears in the match_data_low_hard_4 or match_data_low_unreid_4 datasets continuously for N times, its target box is added to the match_data_low_hard_4 dataset. For background negative samples: If the target appears in no_match_unreid_3 continuously for N times, it indicates background noise.