An Unsupervised Ship Detection Method and System Based on Multimodal Model

Through the unsupervised ship detection method based on multimodal model, the training set is automatically calibrated and feature matching supervision is added, which solves the problem of manual calibration in the existing technology that consumes a lot of manpower and the unstable unsupervised model, and achieves efficient and accurate ship detection.

CN119559384BActive Publication Date: 2025-05-30ZHEJIANG WHYIS TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510131746.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-02-06
Publication Date
2025-05-30
Estimated Expiration
2045-02-06

AI Technical Summary

Technical Problem

The existing ship detection model requires manual calibration, which consumes a lot of manpower, and the unsupervised model lacks effective binding, resulting in instability, and some scenarios cannot collect ship data sets, resulting in limitations of the training set.

Method used

The unsupervised ship detection method based on multimodal model is adopted, and the training set is automatically calibrated by building target ship matching models and multimodal models, adding feature matching supervision, and improving the loss structure to improve detection accuracy and robustness.

Benefits of technology

It can improve the accuracy and efficiency of the ship detection model without manual calibration, enhance the robustness and detection capabilities of the model, and reduce labor costs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119559384B_ABST
    Figure CN119559384B_ABST
Patent Text Reader

Abstract

The present invention discloses an unsupervised ship detection method and system based on a multimodal model. Among them, the method modifies the original ship detection model to add a matching module to the original ship detection model, so that the ship detection features are added with feature matching supervision, improving the accuracy of ship detection; uses the multimodal model to automatically calibrate the training set without consuming manpower for manual calibration, greatly improving the efficiency; the ship detection model adds multimodal model supervision to the loss structure, further improving the accuracy of ship detection.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of ship detection, and in particular, to an unsupervised ship detection method and system based on a multi-modal model. Background Art

[0002] With the rapid development of water transportation technology, water traffic safety has become increasingly important. Among them, the position of ships is one of the important factors in water traffic. However, due to the limited quantity and variety of training set data, the detection rate and robustness of current ship detection models have limitations, and it is impossible to achieve self-training of ship detection and real-time update of model weights, thereby improving the detection rate and robustness of ship detection models. Currently, ship detection models have the following defects:

[0003] (1) Currently, supervised detection models require manual calibration, which consumes a large amount of manpower;

[0004] (2) Due to the inability to collect ship data sets in some scenarios, the ship training set has limitations;

[0005] (3) Due to the lack of manual calibration, unsupervised models lack effective constraints, resulting in instability of unsupervised models.

[0006] In response to the above problems, no effective solutions have been proposed yet. Summary of the Invention

[0007] An embodiment of the present invention provides an unsupervised ship detection method based on a multi-modal model to solve the problems in the prior art that supervised detection models require manual calibration and consume a large amount of manpower, and unsupervised models are unstable due to the lack of manual calibration, lack of effective constraints, and the inability to collect ship data sets in some scenarios.

[0008] To achieve the above object, on the one hand, the present invention provides an unsupervised ship detection method based on a multimodal model, and the method includes: S1. Construct a target ship matching model according to the original ship detection model and the matching module; S2. Extract each target in the manually calibrated training set to obtain a first picture set; input the first picture set into the target ship matching model to obtain a first feature sequence of each target; input the uncalibrated training set into the SAM segmentation model to obtain the first coordinates of each target; and extract each target to obtain a second picture set; input the second picture set into the clip multimodal model to obtain the first category and the first confidence of each target; input the second picture set into the target ship matching model to obtain a second feature sequence of each target; according to the first feature sequence of each target in the first picture set and the second feature sequence of each target in the second picture set, obtain the second category and the second confidence of each target in the second picture set; S3. Input the uncalibrated training set into the first ship detection model for prediction to obtain the third category, the second coordinates and the third confidence of each target in each picture; S4. Input the uncalibrated training set into the second ship detection model for prediction to obtain the fourth category, the third coordinates and the fourth confidence of each target in each picture; calculate the first loss value, the second loss value and the third loss value of the current target in the current picture of the uncalibrated training set according to the fourth category, the third category, the second category, the first category, the third coordinates, the second coordinates, the first coordinates, the fourth confidence, the third confidence, the second confidence and the first confidence; S5. Extract each target according to the second coordinates of each target in each picture of the uncalibrated training set to obtain a third picture set; input the third picture set into the target ship matching model to obtain the fifth category and the fifth confidence of each target, and input the third picture set into the clip multimodal model to obtain the sixth category and the sixth confidence of each target; calculate the fourth loss value of the current target in the current picture of the uncalibrated training set according to the second coordinates, the first coordinates, the third category, the fourth category, the fifth category, the sixth category, the third confidence, the fourth confidence, the fifth confidence and the sixth confidence; S6. Update the second ship detection model in reverse according to the sum of the first loss value, the second loss value, the third loss value and the fourth loss value of all targets in the uncalibrated training set; S7. Repeat S2 to S6, stop training if the sum value fluctuates within a first preset range, and use the finally updated second ship detection model as the target ship detection model; input the ship picture to be detected into the target ship detection model for detection to obtain the target category and position.

[0009] Optionally, S1 includes: S11, copying the manually calibrated training set and keeping only the targets of one category in each picture of the copied manually calibrated training set to obtain an updated training set; S12, adding a matching module to the backbone network of the original ship detection model to obtain an original ship matching model, and freezing the parameters of the backbone network of the original ship detection model; S13, inputting the updated training set into the original ship detection model for prediction to obtain the category, coordinates, and confidence of each picture; S14, screening the updated training set according to the category, coordinates, and confidence predicted for each picture in the updated training set and the manually calibrated category and coordinates, and training the screened updated training set through the original ship matching model to obtain the feature vector of each picture; S15, calculating the total matching loss value based on the feature vectors of all pictures in the screened updated training set, and reversely updating the matching module according to the total matching loss value; S16, repeating S13 - S15 until the finally calculated total matching loss value fluctuates within a second preset range and stops training to obtain the target ship matching model.

[0010] Optionally, the step of obtaining the second category and the second confidence of each target in the second picture set according to the first feature sequence of each target in the first picture set and the second feature sequence of each target in the second picture set includes: calculating the similarity values respectively according to the second feature sequence of the current target in the second picture set and the first feature sequence of each target in the first picture set to obtain the similarity values between the current target in the second picture set and each target in the first picture set; classifying the targets in the first picture set corresponding to all the similarity values by category to obtain the actual target quantity of each category; screening all the similarity values according to a preset similarity threshold to obtain the screened similarity values, and counting the number as the predicted quantity; classifying the targets in the first picture set corresponding to all the screened similarity values by category to obtain the accurate target quantity of each category, and calculating the accuracy rate of the current target in the second picture set and each category according to the accurate target quantity and the predicted quantity of each category; calculating the recall rate of the current target in the second picture set and each category according to the accurate target quantity and the actual target quantity of each category; screening the accuracy rates of the current target in the second picture set and each category according to an accuracy rate threshold to obtain the screened accuracy rates; taking each category corresponding to the screened accuracy rates as a pre-target category; sorting the recall rates of the current target in the second picture set and each pre-target category, and selecting the pre-target category corresponding to the highest recall rate as the target category of the current target in the second picture set; the target category is the second category; calculating the second confidence of the current target in the second picture set according to the accuracy rate of the current target in the second picture set and the target category and the recall rate of the current target in the second picture set and the target category.

[0011] Optionally, S4 includes: if the intersection over union (IoU) between the first coordinate of the current target in the second image set and the second coordinate of the current target in the uncalibrated training set is greater than or equal to the first preset IoU threshold, and the first category, second category of the current target in the second image set are the same as the third category of the current target in the uncalibrated training set, then calculate the first coordinate loss value based on the first coordinate of the current target in the second image set, the second coordinate of the current target in the uncalibrated training set, and the third coordinate of the current target in the uncalibrated training set; calculate the first category loss value based on the first category or second category of the current target in the second image set, or the third category of the current target in the uncalibrated training set, and the fourth category of the current target in the uncalibrated training set; calculate the first weight based on the first confidence and second confidence of the current target in the second image set; calculate the first loss value based on the first coordinate loss value, the first category loss value, and the first weight; if the IoU between the first coordinate of the current target in the second image set and the second coordinate of the current target in the uncalibrated training set is greater than or equal to the first preset IoU threshold, and the first category, second category of the current target in the second image set are completely different from the third category of the current target in the uncalibrated training set, or only two of them are the same, then calculate the second category loss value based on the first category, first confidence, second category, second confidence of the current target in the second image set, the third category, third confidence, fourth category, fourth confidence of the current target in the uncalibrated training set; calculate the second loss value based on the second category loss value and the first coordinate loss value; if the IoU between the first coordinate of the current target in the second image set and the second coordinate of the current target in the uncalibrated training set is less than the first preset IoU threshold, and the first confidence or second confidence of the current target in the second image set is greater than the first preset confidence, then calculate the third loss value based on the first category, first confidence, second category, second confidence of the current target in the second image set, and the fourth category, fourth confidence of the current target in the uncalibrated training set.

[0012] Optionally, calculating the fourth loss value of the current target in the current image in the uncalibrated training set according to the second coordinate, the first coordinate, the third category, the fourth category, the fifth category, the sixth category, the third confidence, the fourth confidence, the fifth confidence, and the sixth confidence includes: if the IoU between the second coordinate of the current target in the third image set and the first coordinate of the current target in the second image set is less than the second preset IoU threshold, then calculate the fourth loss value based on the third category, third confidence, fifth category, fifth confidence, sixth category, sixth confidence of the current target in the third image set, and the fourth category, fourth confidence of the current target in the uncalibrated training set.

[0013] Optionally, the first loss value is calculated according to the following formula:

[0014] ;

[0015] Wherein, is the first loss value, is the first weight, is the first coordinate loss value, is the first category loss value, is the coordinate loss value between the third coordinate of the current target in the uncalibrated training set and the first coordinate of the current target in the second image set, and the coordinate loss value between the third coordinate of the current target in the uncalibrated training set and the second coordinate of the current target in the uncalibrated training set. min is to find the minimum coordinate loss value, is the second confidence of the current target in the second image set, is the first confidence of the current target in the second image set, is the accuracy rate of the current target in the second image set and the target category, is the recall rate of the current target in the second image set and the target category.

[0016] Optionally, the second loss value is calculated according to the following formula:

[0017] ;

[0018] Wherein, is the category loss value a between the fourth category of the current target in the uncalibrated training set and the first category of the current target in the second image set, or the category loss value b between the fourth category of the current target in the uncalibrated training set and the second category of the current target in the second image set, or the category loss value c between the fourth category of the current target in the uncalibrated training set and the third category of the current target in the uncalibrated training set; when is the category loss value a, is the first confidence of the current target in the second image set, y is 0 or 1. If the fourth category of the current target in the uncalibrated training set matches the first category of the current target in the second image set, y is 1, otherwise, y is 0; when is the category loss value b, is the second confidence of the current target in the second image set, y is 0 or 1. If the fourth category of the current target in the uncalibrated training set matches the second category of the current target in the second image set, y is 1, otherwise, y is 0; when is the category loss value c, is the third confidence of the current target in the uncalibrated training set, y is 0 or 1. If the fourth category of the current target in the uncalibrated training set matches the third category of the current target in the uncalibrated training set, y is 1, otherwise, y is 0; is the fourth confidence of the current target in the uncalibrated training set, is the average value of class loss value a, class loss value b, and class loss value c, that is, the second class loss value. is the first coordinate loss value. is the second loss value.

[0019] Optionally, the third loss value is calculated according to the following formula:

[0020] ;

[0021] ;

[0022] where, is the third loss value, is the class loss value a between the fourth class of the current target in the unlabeled training set and the first class of the current target in the second image set, or the class loss value b between the fourth class of the current target in the unlabeled training set and the second class of the current target in the second image set. When is the class loss value a, is the first confidence of the current target in the second image set, y is 0 or 1. If the fourth class of the current target in the unlabeled training set matches the first class of the current target in the second image set, y is 1; otherwise, y is 0. When is the class loss value b, is the second confidence of the current target in the second image set, y is 0 or 1. If the fourth class of the current target in the unlabeled training set matches the second class of the current target in the second image set, y is 1; otherwise, y is 0. is the fourth confidence of the current target in the unlabeled training set. When both the first confidence and the second confidence of the current target in the second image set are greater than the first preset confidence, k is 2; when only one confidence is greater than the first preset confidence, k is 1.

[0023] Optionally, the fourth loss value is calculated according to the following formula:

[0024] ;

[0025] where, is the fourth loss value, is the class loss value d between the fourth class of the current target in the unlabeled training set and the third class of the current target in the third image set, or the class loss value e between the fourth class of the current target in the unlabeled training set and the fifth class of the current target in the third image set, or the class loss value f between the fourth class of the current target in the unlabeled training set and the sixth class of the current target in the third image set. is the average value of class loss value d, class loss value e, and class loss value f. When is the class loss value d, is the third confidence level of the current target in the third picture set, y is 0 or 1. If the fourth category of the current target in the calibrated training set does not match the third category of the current target in the third picture set, y is 1; otherwise, y is 0. When is the class loss value e, is the fifth confidence level of the current target in the third picture set, y is 0 or 1. If the fourth category of the current target in the calibrated training set does not match the fifth category of the current target in the third picture set, y is 1; otherwise, y is 0. When is the class loss value f, is the sixth confidence level of the current target in the third picture set, y is 0 or 1. If the fourth category of the current target in the calibrated training set does not match the sixth category of the current target in the third picture set, y is 1; otherwise, y is 0.

[0026] On the other hand, the present invention provides an unsupervised ship detection system based on a multimodal model, which system includes: a construction unit for constructing a target ship matching model according to an original ship detection model and a matching module; a first prediction unit for extracting each target in a manually calibrated training set to obtain a first picture set; inputting the first picture set into the target ship matching model to obtain a first feature sequence of each target; inputting an uncalibrated training set into a SAM segmentation model to obtain a first coordinate of each target; and extracting each target to obtain a second picture set; inputting the second picture set into a clip multimodal model to obtain a first category and a first confidence level of each target; inputting the second picture set into the target ship matching model to obtain a second feature sequence of each target; obtaining a second category and a second confidence level of each target in the second picture set according to the first feature sequence of each target in the first picture set and the second feature sequence of each target in the second preprocessed picture set; a second prediction unit for inputting the uncalibrated training set into a first ship detection model for prediction to obtain a third category, a second coordinate, and a third confidence level of each target in each picture; a first calculation unit for inputting the uncalibrated training set into a second ship detection model for prediction to obtain a fourth category, a third coordinate, and a fourth confidence level of each target in each picture; calculating a first loss value, a second loss value, and a third loss value of the current target in the current picture in the uncalibrated training set according to the fourth category, the third category, the second category, the first category, the third coordinate, the second coordinate, the first coordinate, the fourth confidence level, the third confidence level, the second confidence level, and the first confidence level; a second calculation unit for extracting each target according to the second coordinate of each target in each picture in the uncalibrated training set to obtain a third picture set; inputting the third picture set into the target ship matching model to obtain a fifth category and a fifth confidence level of each target, and inputting the third picture set into the clip multimodal model to obtain a sixth category and a sixth confidence level of each target; calculating a fourth loss value of the current target in the current picture in the uncalibrated training set according to the second coordinate, the first coordinate, the third category, the fourth category, the fifth category, the sixth category, the third confidence level, the fourth confidence level, the fifth confidence level, and the sixth confidence level; a reverse update unit for reversely updating the second ship detection model according to the sum value of the first loss value, the second loss value, the third loss value, and the fourth loss value of all targets in the uncalibrated training set; a detection unit for repeating the first prediction unit, the second prediction unit, the first calculation unit, the second calculation unit, and the reverse update unit, stopping training if the sum value fluctuates within a first preset range, and using the finally updated second ship detection model as the target ship detection model; and inputting the ship picture to be detected into the target ship detection model for detection to obtain the target category and position.

[0027] Advantages of the present invention:

[0028] The present invention provides an unsupervised ship detection method and system based on a multimodal model. Among them, the method uses the multimodal model to automatically calibrate the training set; modifies the original ship detection model to add a matching module to the original ship detection model, so that the ship detection features are increased with feature matching supervision; and the ship detection model adds multimodal model supervision in the loss structure. The method of the present invention can improve the accuracy and efficiency of ship detection. BRIEF DESCRIPTION OF THE DRAWINGS

[0029] Figure 1 is a flowchart of an unsupervised ship detection method based on a multimodal model provided by an embodiment of the present invention;

[0030] Figure 2 is a schematic structural diagram of an unsupervised ship detection system based on a multimodal model provided by an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0031] In order to make the objectives, technical solutions and advantages of the present invention clearer, the present invention will be further described in detail below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the scope of protection of the present invention.

[0032] Figure 1 is an unsupervised ship detection method based on a multimodal model provided by an embodiment of the present invention. As Figure 1 shown, the method includes:

[0033] S1. Construct a target ship matching model according to the original ship detection model and the matching structure;

[0034] Specifically, the S1 includes:

[0035] S11. Copy the manually calibrated training set and keep only the targets of one category in each picture of the copied manually calibrated training set to obtain an updated training set;

[0036] Taking a picture (the current picture) in the manually calibrated training set as an example:

[0037] If there are two targets in the current picture, and the two targets are of different categories, namely: A and B, then the current picture is copied into two pictures. In one picture, only the targets of category A are kept, and the targets of category B are blacked out or the background information is filled into the targets of category B; in the other picture, only the targets of category B are kept, and the targets of category A are blacked out or the background information is filled into the targets of category A.

[0038] If there are two targets in the current image, and the categories of the two targets are the same, both being A, then the current image is copied into a new image, and both targets are retained in this new image.

[0039] Perform the above processing on all images in the manually calibrated training set to obtain an updated training set.

[0040] S12. Add a matching module to the backbone network of the original ship detection model to obtain the original ship matching model, and freeze the parameters of the backbone network of the original ship detection model;

[0041] Specifically, the backbone network of the original ship matching model uses the backbone network of the original ship detection model (ResNet50). The original ship matching model adds a matching module compared to the original ship detection model. The specific steps of the matching module are as follows:

[0042] Perform self-attention operation on the feature map output by the backbone network to obtain a self-attention feature map;

[0043] Add the self-attention feature map to the feature map output by the backbone network to obtain a fused feature map;

[0044] Perform convolution operation on the fused feature map to obtain features with more valid information;

[0045] Perform global max pooling and fully convolutional operation on the fused feature map after convolution operation to obtain a 1*512-dimensional feature vector.

[0046] During the training process of the original ship matching model, the backbone network is frozen and does not perform backpropagation. Only the matching module performs backpropagation.

[0047] S13. Input the updated training set into the original ship detection model for prediction to obtain the category, coordinates, and confidence of each image;

[0048] S14. Screen the updated training set according to the category, coordinates, confidence predicted for each image in the updated training set and the manually calibrated category and coordinates, and train the screened updated training set through the original ship matching model to obtain the feature vector of each image;

[0049] Take an image (the current image) in the updated training set as an example:

[0050] If the category of the current image in the updated training set is consistent with the manually calibrated category, the confidence of the current image in the updated training set is greater than the second preset confidence, and the intersection over union of the coordinates of the current image in the updated training set and the manually calibrated coordinates is greater than the third preset intersection over union threshold, then the current image in the updated training set is retained.

[0051] Specifically, if there is only one target in the current image, and the category of the target is consistent with the manually calibrated category, the confidence of the target is greater than the second preset confidence level (set to 0.1 in the present invention), and the intersection over union (IoU) between the coordinates of the target and the manually calibrated coordinates is greater than the third preset IoU threshold (set to 0.25 in the present invention), then the current image is retained for subsequent matching model training.

[0052] If there are at least two targets in the current image, then as long as one target meets the above three conditions, the current image can be retained for subsequent matching model training.

[0053] All the retained images, that is, the filtered updated training set, are trained through the original ship matching model. Taking an image (the current image) in the filtered updated training set as an example:

[0054] The current image in the filtered updated training set is sequentially passed through the backbone network and the matching module to obtain the feature vector of the current image; that is, the feature map output by the current image through the backbone network undergoes self-attention operation, and the feature map after self-attention operation and the feature map output by the backbone network are sequentially subjected to fusion operation, convolution operation, global max pooling, and fully convolutional operation to obtain the feature vector of the current image.

[0055] S15. Calculate the total matching loss value based on the feature vectors of all images in the filtered updated training set, and update the matching module in the reverse direction according to the total matching loss value;

[0056] Specifically, the filtered updated training set is divided into multiple iterations. Taking one iteration (the current iteration) as an example:

[0057] The matching loss value of the current iteration is calculated according to the following formula:

[0058] ;

[0059] where, represents the matching loss value of the current iteration, N represents the number of images in the current iteration, i represents the i-th image in the current iteration, j represents the j-th image among the N images in the current iteration other than itself i, represents the feature vector obtained by the i-th image in the current iteration through the original ship matching model, represents the feature vector obtained by the j-th image in the current iteration through the original ship matching model, y is 0 or 1. If the category of the i-th image in the current iteration matches the category of the j-th image in the current iteration, then y is 1, otherwise, y is 0.

[0060] Sum and average the matching loss values of all iterations to obtain the total matching loss value.

[0061] S16. Repeat S13 - S15 until the finally calculated total matching loss value fluctuates within the second preset range, then stop training to obtain the target ship matching model.

[0062] Repeat the above process to update the original ship matching model multiple times until the finally calculated total matching loss value fluctuates within the second preset range ( 0.1%) and then stop training to obtain the target ship matching model. Among them, only the matching module in the original ship matching model is updated.

[0063] S2. Crop each target in the manually calibrated training set to obtain the first picture set; input the first picture set into the target ship matching model to obtain the first feature sequence of each target; input the uncalibrated training set into the SAM segmentation model to obtain the first coordinates of each target; and crop each target to obtain the second picture set; input the second picture set into the clip multi-modal model to obtain the first category and the first confidence of each target; input the second picture set into the target ship matching model to obtain the second feature sequence of each target; according to the first feature sequence of each target in the first picture set and the second feature sequence of each target in the second picture set, obtain the second category and the second confidence of each target in the second picture set;

[0064] Specifically, crop each target in the manually calibrated training set according to the coordinates of each target manually calibrated to obtain the first picture set; each picture in the first picture set has only one target.

[0065] Input the first picture set into the target ship matching model to obtain the first feature sequence of each target;

[0066] Input the uncalibrated training set into the SAM segmentation model to obtain the feature information of each target in each picture in the uncalibrated training set; then use the connected component method to frame each target to obtain the target box and the first coordinates of each target in each picture in the uncalibrated training set; crop each target according to the target box of each target in each picture in the uncalibrated training set to obtain the second picture set; each picture in the second picture set has only one target.

[0067] Input the second picture set into the clip multi-modal model to obtain the first category and the first confidence of each target; input the second picture set into the target ship matching model to obtain the second feature sequence of each target;

[0068] According to the first feature sequence of each target in the first picture set and the second feature sequence of each target in the second picture set, obtain the second category and the second confidence of each target in the second picture set, including:

[0069] Calculate the similarity values respectively according to the second feature sequence of the current target in the second picture set and the first feature sequence of each target in the first picture set, and obtain the similarity values between the current target in the second picture set and each target in the first picture set;

[0070] In a preferred embodiment, if there are 50 targets in the first picture set, calculate the similarity values with the current target in the second picture set to obtain 50 similarity values.

[0071] Classify the targets in the first picture set corresponding to all the similarity values by category to obtain the true target quantity of each category; Screen all the similarity values according to a preset similarity threshold to obtain all the screened similarity values, and count the number as the predicted quantity; Classify the targets in the first picture set corresponding to all the screened similarity values by category to obtain the accurate target quantity of each category, and calculate the accuracy rate of the current target in the second picture set and each category according to the accurate target quantity of each category and the predicted quantity; Calculate the recall rate of the current target in the second picture set and each category according to the accurate target quantity of each category and the true target quantity of each category;

[0072] If there are three categories in total for the targets in the first picture set, namely category A, category B, and category C; The true target quantity of category A is 20, the true target quantity of category B is 20, and the true target quantity of category C is 10.

[0073] Screen the 50 similarity values according to a preset similarity threshold (set to 0.35 in the present invention) to obtain 40 similarity values, then the predicted quantity is 40. Count the quantity of each category in the first picture set corresponding to the 40 similarity values to obtain the accurate target quantity of category A as 18, the accurate target quantity of category B as 15, and the accurate target quantity of category C as 7.

[0074] The accuracy rate of category A is: 18 / 40, the accuracy rate of category B is: 15 / 40, and the accuracy rate of category C is 7 / 40;

[0075] The recall rate of category A is: 18 / 20, the recall rate of category B is: 15 / 20, and the recall rate of category C is 7 / 10;

[0076] Screen the accuracy rates of the current target in the second picture set and each category according to an accuracy rate threshold to obtain the screened accuracy rates; Take each category corresponding to the screened accuracy rates as the pre-target category; Sort the recall rates of the current target in the second picture set and each pre-target category, and select the pre-target category corresponding to the highest recall rate as the target category of the current target in the second picture set; The target category is the second category;

[0077] In an alternative embodiment, if the accuracy threshold is 1 / 2, filter the accuracies greater than the accuracy threshold, and use the corresponding A-class and B-class after filtering as the pre-target classes;

[0078] If the recall rate of the A-class is greater than that of the B-class, then the A-class is used as the target class; therefore, the class of the current target in the second image set is the A-class, and it is used as the second class.

[0079] Calculate the second confidence of the current target in the second image set based on the accuracy of the current target and the target class in the second image set and the recall rate of the current target and the target class in the second image set.

[0080] In an alternative embodiment, the second confidence of the current target in the second image set is (18 / 40 + 18 / 20) / 2 = 0.675.

[0081] S3. Input the uncalibrated training set into the first ship detection model for prediction to obtain the third class, second coordinates, and third confidence of each target in each image;

[0082] S4. Input the uncalibrated training set into the second ship detection model for prediction to obtain the fourth class, third coordinates, and fourth confidence of each target in each image; calculate the first loss value, second loss value, and third loss value of the current target in the current image of the uncalibrated training set based on the fourth class, third class, second class, first class, third coordinates, second coordinates, first coordinates, fourth confidence, third confidence, second confidence, and first confidence.

[0083] Specifically, use the first coordinates, first class, first confidence, second class, second confidence of the current target in the second image set, the third class, second coordinates, and third confidence of the current target in the uncalibrated training set as pseudo-labels, and compare them with the fourth class, third coordinates, and fourth confidence of the current target in the uncalibrated training set to calculate the loss value.

[0084] The said S4 includes:

[0085] ① If the intersection over union (IoU) between the first coordinate of the current target in the second image set and the second coordinate of the current target in the uncalibrated training set is greater than or equal to the first preset IoU threshold, and the first category and the second category of the current target in the second image set are the same as the third category of the current target in the uncalibrated training set, then calculate the first coordinate loss value based on the first coordinate of the current target in the second image set, the second coordinate of the current target in the uncalibrated training set, and the third coordinate of the current target in the uncalibrated training set; calculate the first category loss value based on the first category or the second category of the current target in the second image set, or the third category of the current target in the uncalibrated training set, and the fourth category of the current target in the uncalibrated training set; calculate the first weight based on the first confidence and the second confidence of the current target in the second image set; calculate the first loss value based on the first coordinate loss value, the first category loss value, and the first weight;

[0086] The first loss value is calculated according to the following formula:

[0087] ;

[0088] where, is the first loss value, is the first weight, is the first coordinate loss value, is the first category loss value, is the coordinate loss value between the third coordinate of the current target in the uncalibrated training set and the first coordinate of the current target in the second image set, and the coordinate loss value between the third coordinate of the current target in the uncalibrated training set and the second coordinate of the current target in the uncalibrated training set. min is used to find the minimum coordinate loss value, is the second confidence of the current target in the second image set, is the first confidence of the current target in the second image set, is the accuracy of the current target in the second image set with respect to the target category, is the recall rate of the current target in the second image set with respect to the target category.

[0089] Since the first category of the current target in the second image set, the second category of the current target in the second image set, and the third category of the current target in the uncalibrated training set are the same, only one first category loss value needs to be calculated. That is, use the category loss value between the first category of the current target in the second image set and the fourth category of the current target in the uncalibrated training set as the first category loss value, or use the category loss value between the second category of the current target in the second image set and the fourth category of the current target in the uncalibrated training set as the first category loss value, or use the category loss value between the third category of the current target in the uncalibrated training set and the fourth category of the current target in the uncalibrated training set as the first category loss value.

[0090] ② If the intersection over union of the first coordinate of the current target in the second picture set and the second coordinate of the current target in the uncalibrated training set is greater than or equal to the first preset intersection over union threshold, and the first category, second category of the current target in the second picture set and the third category of the current target in the uncalibrated training set are completely inconsistent, or only two of them are the same, then calculate the second category loss value according to the first category, first confidence level, second category, second confidence level of the current target in the second picture set, the third category, third confidence level of the current target in the uncalibrated training set, the fourth category, fourth confidence level of the current target in the uncalibrated training set; calculate the second loss value according to the second category loss value and the first coordinate loss value;

[0091] The second loss value is calculated according to the following formula:

[0092] ;

[0093] Wherein, is the category loss value a between the fourth category of the current target in the uncalibrated training set and the first category of the current target in the second picture set, or the category loss value b between the fourth category of the current target in the uncalibrated training set and the second category of the current target in the second picture set, or the category loss value c between the fourth category of the current target in the uncalibrated training set and the third category of the current target in the uncalibrated training set; when is the category loss value a, is the first confidence level of the current target in the second picture set, y is 0 or 1, if the fourth category of the current target in the uncalibrated training set matches the first category of the current target in the second picture set, y is 1, otherwise, y is 0; when is the category loss value b, is the second confidence level of the current target in the second picture set, y is 0 or 1, if the fourth category of the current target in the uncalibrated training set matches the second category of the current target in the second picture set, y is 1, otherwise, y is 0; when is the category loss value c, is the third confidence level of the current target in the uncalibrated training set, y is 0 or 1, if the fourth category of the current target in the uncalibrated training set matches the third category of the current target in the uncalibrated training set, y is 1, otherwise, y is 0; is the fourth confidence level of the current target in the uncalibrated training set, is the average value of the category loss value a, category loss value b, category loss value c, that is, the second category loss value, is the first coordinate loss value, is the second loss value.

[0094] ③ If the intersection over union (IoU) between the first coordinate of the current target in the second image set and the second coordinate of the current target in the uncalibrated training set is less than the first preset IoU threshold, and the first confidence or the second confidence of the current target in the second image set is greater than the first preset confidence, then the third loss value is calculated based on the first category, the first confidence, the second category, the second confidence of the current target in the second image set, the fourth category, and the fourth confidence of the current target in the uncalibrated training set.

[0095] The third loss value is calculated according to the following formula:

[0096] ;

[0097] Where, is the third loss value, is the class loss value a between the fourth category of the current target in the uncalibrated training set and the first category of the current target in the second image set, or the class loss value b between the fourth category of the current target in the uncalibrated training set and the second category of the current target in the second image set. When is the class loss value a, is the first confidence of the current target in the second image set, y is 0 or 1. If the fourth category of the current target in the uncalibrated training set matches the first category of the current target in the second image set, y is 1; otherwise, y is 0. When is the class loss value b, is the second confidence of the current target in the second image set, y is 0 or 1. If the fourth category of the current target in the uncalibrated training set matches the second category of the current target in the second image set, y is 1; otherwise, y is 0. is the fourth confidence of the current target in the uncalibrated training set. When both the first confidence and the second confidence of the current target in the second image set are greater than the first preset confidence, k is 2; when only one confidence is greater than the first preset confidence, k is 1.

[0098] Specifically, if the intersection over union (IoU) between the first coordinate of the current target in the second image set and the second coordinate of the current target in the uncalibrated training set is less than the first preset IoU threshold, and the first confidence of the current target in the second image set is greater than the first preset confidence, but the second confidence of the current target in the second image set is less than the first preset confidence, then the class loss value b between the fourth category of the current target in the uncalibrated training set and the second category of the current target in the second image set is 0, and the third loss value is the class loss value a / 2.

[0099] If the intersection over union of the first coordinate of the current target in the second image set and the second coordinate of the current target in the uncalibrated training set is less than the first preset intersection over union threshold, and the first confidence of the current target in the second image set is less than the first preset confidence, but the second confidence of the current target in the second image set is greater than the first preset confidence, then the class loss value a of the fourth class of the current target in the uncalibrated training set and the second class of the current target in the second image set is 0, and the third loss value is the class loss value b / 2.

[0100] If the intersection over union of the first coordinate of the current target in the second image set and the second coordinate of the current target in the uncalibrated training set is less than the first preset intersection over union threshold, and the first confidence of the current target in the second image set is greater than the first preset confidence, and the second confidence of the current target in the second image set is also greater than the first preset confidence, then the third loss value is (class loss value a + class loss value b) / 2.

[0101] S5. Crop each target according to the second coordinate of each target in each image in the uncalibrated training set to obtain a third image set; input the third image set into the target ship matching model to obtain the fifth class and the fifth confidence of each target, and input the third image set into the clip multimodal model to obtain the sixth class and the sixth confidence of each target; calculate the fourth loss value of the current target in the current image in the uncalibrated training set according to the second coordinate, the first coordinate, the third class, the fourth class, the fifth class, the sixth class, the third confidence, the fourth confidence, the fifth confidence, and the sixth confidence;

[0102] Specifically, use the third class, the third confidence, the fifth class, the fifth confidence, the sixth class, and the sixth confidence of the current target in the third image set as pseudo-labels, compare them with the fourth class and the fourth confidence of the current target in the uncalibrated training set, and calculate the loss value.

[0103] The S5 includes:

[0104] ① If the intersection over union of the second coordinate of the current target in the third image set and the first coordinate of the current target in the second image set is greater than or equal to the second preset intersection over union threshold, and the third class, the fifth class, and the sixth class of the current target in the third image set are the same, no processing is performed.

[0105] ② If the intersection over union of the second coordinate of the current target in the third image set and the first coordinate of the current target in the second image set is greater than or equal to the second preset intersection over union threshold, and the third class, the fifth class, and the sixth class of the current target in the third image set are completely different, or only two of them are the same, no processing is performed.

[0106] ③ If the intersection over union of the second coordinate of the current target in the third picture set and the first coordinate of the current target in the second picture set is less than the second preset intersection over union threshold, then according to the third category, third confidence level, fifth category, fifth confidence level, sixth category, and sixth confidence level of the current target in the third picture set, and the fourth category and fourth confidence level of the current target in the unlabeled training set, the fourth loss value is calculated.

[0107] The fourth loss value is calculated according to the following formula:

[0108] ;

[0109] Where is the fourth loss value, is the category loss value d between the fourth category of the current target in the unlabeled training set and the third category of the current target in the third picture set, or the category loss value e between the fourth category of the current target in the unlabeled training set and the fifth category of the current target in the third picture set, or the category loss value f between the fourth category of the current target in the unlabeled training set and the sixth category of the current target in the third picture set, is the average value of the category loss value d, the category loss value e, and the category loss value f. When is the category loss value d, is the third confidence level of the current target in the third picture set, y is 0 or 1. If the fourth category of the current target in the unlabeled training set matches the third category of the current target in the third picture set, y is 1; otherwise, y is 0. When is the category loss value e, is the fifth confidence level of the current target in the third picture set, y is 0 or 1. If the fourth category of the current target in the unlabeled training set matches the fifth category of the current target in the third picture set, y is 1; otherwise, y is 0. When is the category loss value f, is the sixth confidence level of the current target in the third picture set, y is 0 or 1. If the fourth category of the current target in the unlabeled training set matches the sixth category of the current target in the third picture set, y is 1; otherwise, y is 0.

[0110] S6. Reverse update the second ship detection model according to the sum of the first loss value, second loss value, third loss value, and fourth loss value of all targets in the unlabeled training set;

[0111] Sum the first loss value, second loss value, third loss value, and fourth loss value of the current target in the current picture in the unlabeled training set to obtain the total loss value of the current target;

[0112] Reverse update the second ship detection model with the sum obtained by summing the total loss values of all targets in the unlabeled training set.

[0113] S7. Repeat the steps S2 - S6. If the sum value fluctuates within the first preset range ( ), stop the training, and use the last updated second ship detection model as the target ship detection model; input the ship picture to be detected into the target ship detection model for detection to obtain the target category and location.

[0114] Figure 2 This is an unsupervised ship detection system based on a multi - modal model provided by an embodiment of the present invention. As Figure 2 shown, the system includes:

[0115] A construction unit 201, configured to construct a target ship matching model according to the original ship detection model and the matching module;

[0116] A first prediction unit 202, configured to extract each target in the manually calibrated training set to obtain a first picture set; input the first picture set into the target ship matching model to obtain a first feature sequence of each target; input the uncalibrated training set into the SAM segmentation model to obtain the first coordinates of each target; and extract each target to obtain a second picture set; input the second picture set into the clip multi - modal model to obtain the first category and the first confidence of each target; input the second picture set into the target ship matching model to obtain a second feature sequence of each target; according to the first feature sequence of each target in the first picture set and the second feature sequence of each target in the second pre - processed picture set, obtain the second category and the second confidence of each target in the second picture set;

[0117] A second prediction unit 203, configured to input the uncalibrated training set into the first ship detection model for prediction to obtain the third category, the second coordinates and the third confidence of each target in each picture;

[0118] A first calculation unit 204, configured to input the uncalibrated training set into the second ship detection model for prediction to obtain the fourth category, the third coordinates and the fourth confidence of each target in each picture; calculate the first loss value, the second loss value and the third loss value of the current target in the current picture in the uncalibrated training set according to the fourth category, the third category, the second category, the first category, the third coordinates, the second coordinates, the first coordinates, the fourth confidence, the third confidence, the second confidence and the first confidence;

[0119] A second calculation unit 205, configured to crop each target according to the second coordinates of each target in each picture in the uncalibrated training set to obtain a third picture set; input the third picture set into the target ship matching model to obtain the fifth category and the fifth confidence of each target, input the third picture set into the clip multi-modal model to obtain the sixth category and the sixth confidence of each target; calculate the fourth loss value of the current target in the current picture in the uncalibrated training set according to the second coordinates, the first coordinates, the third category, the fourth category, the fifth category, the sixth category, the third confidence, the fourth confidence, the fifth confidence, and the sixth confidence;

[0120] A reverse update unit 206, configured to reversely update the second ship detection model according to the sum of the first loss value, the second loss value, the third loss value, and the fourth loss value of all targets in the uncalibrated training set;

[0121] A detection unit 207, configured to repeat the first prediction unit, the second prediction unit, the first calculation unit, the second calculation unit, and the reverse update unit, stop training if the sum value fluctuates within a first preset range, and use the last updated second ship detection model as the target ship detection model; input the ship picture to be detected into the target ship detection model for detection to obtain the target category and position.

[0122] The system of the present invention corresponds to the above method, and the specific embodiments of the system will not be repeated here.

[0123] Advantages of the present invention:

[0124] The method of the present invention uses a multi-modal model to automatically calibrate the training set, eliminating the need for manual calibration, which greatly improves the efficiency; modifies the original ship detection model to add a matching structure to the original ship detection model, thereby increasing feature matching supervision for ship detection features and improving the accuracy of ship detection; the ship detection model adds multi-modal model supervision to the loss structure, further improving the accuracy of ship detection.

[0125] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements for some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. An unsupervised ship detection method based on a multimodal model, characterized in that: include: S1, constructing a target ship matching model based on the original ship detection model and the matching module; S2, extracting each target in the manual calibration training set to obtain a first picture set; inputting the first picture set into the target ship matching model to obtain a first feature sequence of each target; Input the uncalibrated training set into the SAM segmentation model to obtain the first coordinate of each target; And extract each target to obtain a second picture set; Input the second picture set into the clip multimodal model to obtain the first category and the first confidence of each target; Inputting the second picture set into the target ship matching model to obtain a second feature sequence of each target; According to the first feature sequence of each target in the first picture set and the second feature sequence of each target in the second picture set, a second category and a second confidence level of each target in the second picture set are obtained; S3, inputting the uncalibrated training set into the first ship detection model for prediction, and obtaining the third category, second coordinates, and third confidence of each target in each picture; S4, inputting the uncalibrated training set into the second ship detection model for prediction, and obtaining the fourth category, third coordinate and fourth confidence of each target in each picture; calculating the first loss value, second loss value and third loss value of the current target in the current picture in the uncalibrated training set according to the fourth category, third category, second category, first category, third coordinate, second coordinate, first coordinate, fourth confidence, third confidence, second confidence and first confidence; S5, clipping each target according to the second coordinate of each target in each picture in the uncalibrated training set to obtain a third picture set; inputting the third picture set into the target ship matching model to obtain a fifth category and a fifth confidence level of each target; inputting the third picture set into the clip multimodal model to obtain a sixth category and a sixth confidence level of each target; A fourth loss value of the current target in the current picture in the uncalibrated training set is calculated according to the second coordinate, the first coordinate, the third category, the fourth category, the fifth category, the sixth category, the third confidence level, the fourth confidence level, the fifth confidence level, and the sixth confidence level; S6, reversely updating the second ship detection model according to the sum of the first loss value, the second loss value, the third loss value, and the fourth loss value of all targets in the uncalibrated training set; S7, repeating S2 to S6, if the sum value fluctuates within the first preset range, stop training, and use the last updated second ship detection model as the target ship detection model; Input the image of the ship to be detected into the target ship detection model for detection to obtain the target category and position; The S1 includes: S11, copying the manually calibrated training set and storing only one category of targets in each picture in the copied manually calibrated training set to obtain an updated training set; S12, adding a matching module to the backbone network of the original ship detection model to obtain the original ship matching model, and freezing the parameters of the backbone network of the original ship detection model; S13, inputting the updated training set into the original ship detection model for prediction, and obtaining the category, coordinates, and confidence of each image; S14, screening the updated training set according to the predicted category, coordinates, confidence level of each image in the updated training set and the manually calibrated category and coordinates, and training the screened updated training set through the original ship matching model to obtain a feature vector of each image; S15, calculating a total matching loss value according to the feature vectors of all images in the screened update training set, and reversely updating the matching module according to the total matching loss value; S16, repeating S13 to S15 until the final calculated total matching loss value fluctuates within the second preset range and the training is stopped to obtain a matching model of the target ship; Among them, the matching module is used to perform a self-attention operation on the feature map output by the backbone network to obtain a self-attention feature map; add the self-attention feature map to the feature map output by the backbone network to obtain a fused feature map; perform a convolution operation on the fused feature map; perform global maximum pooling and full convolution operations on the fused feature map after the convolution operation to obtain a 1*512-dimensional feature vector.

2. The method according to claim 1, characterized in that The step of obtaining a second category and a second confidence level of each target in the second picture set according to the first feature sequence of each target in the first picture set and the second feature sequence of each target in the second picture set comprises: Calculate similarity values ​​according to the second feature sequence of the current target in the second picture set and the first feature sequence of each target in the first picture set, and obtain similarity values ​​between the current target in the second picture set and each target in the first picture set; Classify the targets in the first picture set corresponding to all similarity values ​​by category to obtain the number of real targets in each category; filter all similarity values ​​according to a preset similarity threshold to obtain all filtered similarity values, and count the number as the predicted number; classify the targets in the first picture set corresponding to all filtered similarity values ​​by category to obtain the number of accurate targets in each category, and calculate the accuracy of the current target in the second picture set and each category based on the number of accurate targets in each category and the predicted number; calculate the recall rate of the current target in the second picture set and each category based on the number of accurate targets in each category and the number of real targets in each category; The accuracy of the current target and each category in the second picture set is screened according to the accuracy threshold to obtain the screened accuracy; each category corresponding to the screened accuracy is used as the pre-target category; the recall rate of the current target and each pre-target category in the second picture set is sorted, and the pre-target category corresponding to the highest recall rate is selected as the target category of the current target in the second picture set; the target category is the second category; A second confidence of the current target in the second picture set is obtained by calculating the accuracy of the current target and the target category in the second picture set and the recall rate of the current target and the target category in the second picture set.

3. The method according to claim 1, characterized in that The S4 includes: If the intersection-and-union ratio of the first coordinate of the current target in the second picture set and the second coordinate of the current target in the uncalibrated training set is greater than or equal to the first preset intersection-and-union ratio threshold, and the first category, the second category of the current target in the second picture set, and the third category of the current target in the uncalibrated training set are consistent, then the first coordinate loss value is calculated according to the first coordinate of the current target in the second picture set, the second coordinate of the current target in the uncalibrated training set, and the third coordinate of the current target in the uncalibrated training set; the first category loss value is calculated according to the first category or the second category of the current target in the second picture set, or the third category of the current target in the uncalibrated training set, and the fourth category of the current target in the uncalibrated training set; the first weight is calculated according to the first confidence and the second confidence of the current target in the second picture set; the first loss value is calculated according to the first coordinate loss value, the first category loss value, and the first weight; If the intersection-and-union ratio of the first coordinate of the current target in the second picture set and the second coordinate of the current target in the uncalibrated training set is greater than or equal to the first preset intersection-and-union ratio threshold, and the first category, the second category of the current target in the second picture set, and the third category of the current target in the uncalibrated training set are completely inconsistent, or only two of the categories are consistent, then the second category loss value is calculated according to the first category, the first confidence, the second category, the second confidence of the current target in the second picture set, the third category and the third confidence of the current target in the uncalibrated training set, and the fourth category and the fourth confidence of the current target in the uncalibrated training set; the second loss value is calculated according to the second category loss value and the first coordinate loss value; If the intersection and union of the first coordinate of the current target in the second picture set and the second coordinate of the current target in the uncalibrated training set is less than the first preset intersection and union threshold, and the first confidence or the second confidence of the current target in the second picture set is greater than the first preset confidence, then the third loss value is calculated based on the first category, first confidence, second category, and second confidence of the current target in the second picture set, and the fourth category and fourth confidence of the current target in the uncalibrated training set.

4. The method according to claim 1, characterized in that The fourth loss value of the current target in the current picture in the uncalibrated training set is calculated according to the second coordinate, the first coordinate, the third category, the fourth category, the fifth category, the sixth category, the third confidence level, the fourth confidence level, the fifth confidence level, and the sixth confidence level, including: If the intersection-and-union ratio of the second coordinate of the current target in the third picture set and the first coordinate of the current target in the second picture set is less than the second preset intersection-and-union ratio threshold, the fourth loss value is calculated according to the third category, third confidence level, fifth category, fifth confidence level, sixth category, and sixth confidence level of the current target in the third picture set and the fourth category and fourth confidence level of the current target in the uncalibrated training set.

5. The method according to claim 3, characterized in that: The first loss value is calculated according to the following formula: conf_pre=((sim_pre+sim_recall) / 2+conf_clip) / 2 in, is the first loss value, e 1-conf_pre is the first weight, is the first coordinate loss value, loss cls is the first category loss value, is the coordinate loss value of the third coordinate of the current target in the uncalibrated training set and the first coordinate of the current target in the second picture set, the coordinate loss value of the third coordinate of the current target in the uncalibrated training set and the second coordinate of the current target in the uncalibrated training set, min is the minimum coordinate loss value, (sim_pre+sim_recall) / 2 is the second confidence of the current target in the second picture set, conf_clip is the first confidence of the current target in the second picture set, sim_pre is the accuracy of the current target and the target category in the second picture set, and sim_recall is the recall rate of the current target and the target category in the second picture set.

6. The method according to claim 3, characterized in that: The second loss value is calculated according to the following formula: in, is the category loss value a of the fourth category of the current target in the uncalibrated training set and the first category of the current target in the second picture set, or the category loss value b of the fourth category of the current target in the uncalibrated training set and the second category of the current target in the second picture set, or the category loss value c of the fourth category of the current target in the uncalibrated training set and the third category of the current target in the uncalibrated training set; when When is the category loss value a, P1 is the first confidence of the current target in the second picture set, y is 0 or 1, if the fourth category of the current target in the uncalibrated training set matches the first category of the current target in the second picture set, y is 1, otherwise, y is 0; when When is the category loss value b, P1 is the second confidence of the current target in the second picture set, y is 0 or 1, if the fourth category of the current target in the uncalibrated training set matches the second category of the current target in the second picture set, y is 1, otherwise, y is 0; when When the category loss value is c, P1 is the third confidence of the current target in the uncalibrated training set, y is 0 or 1. If the fourth category of the current target in the uncalibrated training set matches the third category of the current target in the uncalibrated training set, y is 1, otherwise, y is 0; pre_conf is the fourth confidence of the current target in the uncalibrated training set, is the average of the category loss values ​​a, b, and c, which is the second category loss value. is the first coordinate loss value, is the second loss value.

7. The method according to claim 3, characterized in that The third loss value is calculated according to the following formula: in, is the third loss value, is the category loss value a of the fourth category of the current target in the uncalibrated training set and the first category of the current target in the second picture set, or the category loss value b of the fourth category of the current target in the uncalibrated training set and the second category of the current target in the second picture set. When is the category loss value a, P2 is the first confidence of the current target in the second picture set, y is 0 or 1, if the fourth category of the current target in the uncalibrated training set matches the first category of the current target in the second picture set, y is 1, otherwise, y is 0; when When is the category loss value b, P2 is the second confidence of the current target in the second picture set, y is 0 or 1, if the fourth category of the current target in the uncalibrated training set matches the second category of the current target in the second picture set, y is 1, otherwise, y is 0; pre_conf is the fourth confidence of the current target in the uncalibrated training set, when the first confidence and the second confidence of the current target in the second picture set are both greater than the first preset confidence, k is 2, when only one confidence is greater than the first preset confidence, k is 1.

8. The method according to claim 4, characterized in that The fourth loss value is calculated according to the following formula: in, is the fourth loss value, is the category loss value d of the fourth category of the current target in the uncalibrated training set and the third category of the current target in the third picture set, or the category loss value e of the fourth category of the current target in the uncalibrated training set and the fifth category of the current target in the third picture set, or the category loss value f of the fourth category of the current target in the uncalibrated training set and the sixth category of the current target in the third picture set, is the average value of category loss value d, category loss value e, and category loss value f. When is the category loss value d, P3 is the third confidence of the current target in the third picture set, y is 0 or 1. If the fourth category of the current target in the uncalibrated training set matches the third category of the current target in the third picture set, y is 1, otherwise, y is 0; when When is the category loss value e, P3 is the fifth confidence of the current target in the third picture set, y is 0 or 1. If the fourth category of the current target in the uncalibrated training set matches the fifth category of the current target in the third picture set, y is 1, otherwise, y is 0; when When it is the category loss value f, P3 is the sixth confidence of the current target in the third picture set, y is 0 or 1. If the fourth category of the current target in the uncalibrated training set matches the sixth category of the current target in the third picture set, y is 1, otherwise, y is 0.

9. An unsupervised ship detection system based on a multimodal model, characterized in that: include: A construction unit, used for constructing a target ship matching model according to the original ship detection model and the matching module; The first prediction unit is used to extract each target in the manual calibration training set to obtain a first picture set; the first picture set is input into the target ship matching model to obtain a first feature sequence of each target; Input the uncalibrated training set into the SAM segmentation model to obtain the first coordinate of each target; And extract each target to obtain a second picture set; Input the second picture set into the clip multimodal model to obtain the first category and the first confidence of each target; Inputting the second picture set into the target ship matching model to obtain a second feature sequence of each target; Obtaining a second category and a second confidence of each target in the second picture set according to a first feature sequence of each target in the first picture set and a second feature sequence of each target in the second preprocessed picture set; a second prediction unit, used for inputting the uncalibrated training set into the first ship detection model for prediction, and obtaining a third category, a second coordinate, and a third confidence level of each target in each picture; The first calculation unit is used to input the uncalibrated training set into the second ship detection model for prediction, and obtain the fourth category, third coordinate and fourth confidence of each target in each picture; and calculate the first loss value, second loss value and third loss value of the current target in the current picture in the uncalibrated training set according to the fourth category, third category, second category, first category, third coordinate, second coordinate, first coordinate, fourth confidence, third confidence, second confidence and first confidence; The second calculation unit is used to clip each target according to the second coordinate of each target in each picture in the uncalibrated training set to obtain a third picture set; input the third picture set into the target ship matching model to obtain a fifth category and a fifth confidence of each target; input the third picture set into the clip multimodal model to obtain a sixth category and a sixth confidence of each target; A fourth loss value of the current target in the current picture in the uncalibrated training set is calculated according to the second coordinate, the first coordinate, the third category, the fourth category, the fifth category, the sixth category, the third confidence level, the fourth confidence level, the fifth confidence level, and the sixth confidence level; A reverse updating unit, configured to reversely update the second ship detection model according to the sum of the first loss value, the second loss value, the third loss value, and the fourth loss value of all targets in the uncalibrated training set; a detection unit, configured to repeatedly perform its functions using the first prediction unit, the second prediction unit, the first calculation unit, the second calculation unit, and the reverse update unit, and to stop training if the sum value fluctuates within a first preset range, and to use the last updated second ship detection model as the target ship detection model; Input the image of the ship to be detected into the target ship detection model for detection to obtain the target category and position; The building block comprises: The saving subunit is used to copy the manually calibrated training set and save only one category of targets for each picture in the copied manually calibrated training set to obtain an updated training set; Adding a subunit, used for adding a matching module to the backbone network of the original ship detection model, obtaining the original ship matching model, and freezing the parameters of the backbone network of the original ship detection model; The prediction subunit is used to input the updated training set into the original ship detection model for prediction, and obtain the category, coordinates, and confidence of each image; A screening subunit is used to screen the updated training set according to the predicted category, coordinates, confidence level of each picture in the updated training set and the manually calibrated category and coordinates, and train the screened updated training set through the original ship matching model to obtain a feature vector of each picture; A calculation subunit, configured to calculate a total matching loss value according to the feature vectors of all images in the filtered update training set, and reversely update the matching module according to the total matching loss value; The training subunit is used to make the prediction subunit, the screening subunit, and the calculation subunit repeatedly perform their functions until the final calculated total matching loss value fluctuates within a second preset range and the training is stopped to obtain a target ship matching model; Among them, the matching module is used to perform a self-attention operation on the feature map output by the backbone network to obtain a self-attention feature map; add the self-attention feature map to the feature map output by the backbone network to obtain a fused feature map; perform a convolution operation on the fused feature map; perform global maximum pooling and full convolution operations on the fused feature map after the convolution operation to obtain a 1*512-dimensional feature vector.

Citation Information

Patent Citations

  • Multi-level unsupervised field adaptive target detection and identification method

    CN114972920A

  • Unsupervised target detection model training method based on metric learning

    CN117218404A