Adaptive Ship Detection Method Based on Contrastive Learning

By comparing the learning of pre-trained backbone network model and the adaptive area suggestion network model, the problem of existing ship detection methods increasing computing costs when improving the detection capabilities of targets of different sizes is solved, and efficient and low-cost ship detection effects are achieved.

CN115147611BActive Publication Date: 2025-06-13SHANGHAI MARITIME UNIVERSITY
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210900090.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-07-28
Publication Date
2025-06-13
Estimated Expiration
2042-07-28

AI Technical Summary

Technical Problem

While improving the detection capabilities of different size targets, existing ship detection methods increase computing costs. In the process of pre-training Backbone, the performance of unsupervised methods is weak, requiring a large number of manual marking samples, resulting in high costs.

Method used

Adaptive ship detection method based on contrast learning is adopted to pre-train the backbone network model through comparative learning, and the training and execution efficiency is reduced, and the recommendation box is directly predicted through the adaptive area suggestion network model to avoid pre-defining anchor boxes of different scales.

Benefits of technology

It realizes efficient ship detection, reduces the computational cost and manual marking cost, can effectively detect targets of different sizes, and improves the execution efficiency of the detection model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115147611B_ABST
    Figure CN115147611B_ABST
Patent Text Reader

Abstract

The present invention discloses an adaptive ship detection method based on contrastive learning. The method includes: pre-training a backbone network model; obtaining a ship image dataset to be detected, and inputting the ship image dataset to be detected into the backbone network model. Through the backbone network model and a feature pyramid model connected to the backbone network model, a plurality of feature maps are processed, wherein the plurality of feature maps are used to detect ship images of different scales; inputting the plurality of feature maps into an adaptive region proposal network model connected to the feature pyramid model, and obtaining proposal boxes that meet preset requirements through the processing of the adaptive region proposal network model, so as to detect the ship image to be detected through the proposal boxes. The present invention can improve the training and execution efficiency of the ship detection model and can achieve object detection of different scales.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of ship detection, and in particular, to an adaptive ship detection method based on contrast learning. Background Art

[0002] In recent years, the security of territorial waters has received extensive attention from various countries. Ship detection, as the core function of a ship target monitoring system, has important research significance.

[0003] Currently, object detection based on deep learning is the mainstream ship detection method. Such methods can extract high-level semantic features through a deep neural network to achieve end-to-end object detection, and have higher detection accuracy compared to traditional ship detection methods. Ship detection methods based on deep learning can be roughly divided into two categories: one-stage and two-stage. Faster R-CNN (an object detection method) is the most popular two-stage ship detection method, and the Faster R-CNN model is as Figure 1 shown. This method innovatively uses a Region Proposal Network (RPN) to solve the problem of generating proposal boxes. The RPN will pre-define a large number of anchor boxes (Anchors), and generate corresponding scores and regression parameters for them through a convolutional neural network. Finally, the Anchors are screened according to the scores generated by the network and the Non-Maximum Suppression (NMS) algorithm, and then high-quality proposal boxes are output. Faster R-CNN calculates the foreground probability of the Anchor and the regression parameters of the proposal box using the feature map output by the Backbone, and has end-to-end detection capabilities.

[0004] To improve the detection ability of ship detection methods for targets of different sizes, a Feature Pyramid Networks (FPN) is introduced into the Backbone of ship detection methods. The FPN consists of a series of parallel convolutional layers, and these convolutional layers are located at different positions in the Backbone, enabling the FPN to extract feature information of different scales, and thus realizing the detection of ship targets of different sizes. However, in the RPN executed after the FPN, a series of Anchor boxes of different sizes need to be constructed, which greatly increases its computational cost.

[0005] For a ship detection model, it is generally necessary to pre-train the Backbone, and then transfer the pre-trained Backbone to the object detection model for fine-tuning. In order to obtain an effective ship feature representation, the pre-training is generally performed on a large-scale image dataset. When pre-training the Backbone, the performance of the unsupervised method is relatively weak and cannot achieve the expected goal, while the supervised method requires manual labeling of a large number of samples, thus incurring a large cost in terms of manpower and material resources. Summary of the Invention

[0006] The present invention aims to solve at least one of the technical problems in the related art to some extent. For this reason, an object of the present invention is to provide an adaptive ship detection method based on contrast learning, which can meet the performance requirements of its corresponding ship detection model, reduce the training and execution efficiency of the ship detection model, and can achieve object detection at different scales.

[0007] To achieve the above object, the present invention is realized through the following technical solutions:

[0008] An adaptive ship detection method based on contrast learning, comprising:

[0009] Step S1: Pre-train the backbone network model;

[0010] Step S2: Obtain the ship image dataset to be detected, and input the ship image dataset to be detected into the backbone network model. Through the backbone network model and the feature pyramid model connected to the backbone network model, a plurality of feature maps are processed, wherein the plurality of feature maps are used to detect ship images of different scales;

[0011] Step S3: Input the plurality of feature maps into the adaptive region proposal network model connected to the feature pyramid model, and obtain proposal boxes that meet the preset requirements through the adaptive region proposal network model, so as to detect the ship image to be detected through the proposal boxes.

[0012] Optionally, the step S1 includes:

[0013] Step S11: Obtain the ship image dataset;

[0014] Step S12: Perform image enhancement processing and image encoding processing on each image in the ship image dataset, and obtain a plurality of positive sample pair features and a plurality of negative sample pair features according to the image data after the image encoding processing;

[0015] Step S13: Pre-train the backbone network model according to the plurality of positive sample pair features and the plurality of negative sample pair features.

[0016] Optionally, step S12 includes:

[0017] S121: Perform two image enhancement processes on the same image to obtain first image enhancement result data and second image enhancement result data respectively;

[0018] S122: Input the first image enhancement result data as anchor sample data into the image encoder for processing to obtain first image encoded data, input the second image enhancement result data into the image momentum encoder for processing to obtain second image encoded data, and form the positive sample pair features with the first image encoded data and the second image encoded data;

[0019] S123: Construct a negative sample feature queue, obtain negative sample data from the negative sample feature queue, and form the negative sample pair features with the first image encoded data and the negative sample data, where the negative sample data in the negative sample feature queue is the image encoded data after processing the anchor sample data of other images by the image encoder.

[0020] Optionally, step S13 includes:

[0021] Step S131: Construct a loss function model;

[0022] Step S132: Input several of the positive sample pair features and several of the negative sample pair features into the loss function model to calculate the contrast loss between samples, and when the contrast loss value reaches the minimum and no longer changes, use the current image encoder as the backbone network model.

[0023] Optionally, step S3 includes:

[0024] Step S31: Obtain the anchor coordinates corresponding to each feature map;

[0025] Step S32: Construct a multi-head perceptron model, input the anchor coordinates corresponding to each feature map into the multi-head perceptron model for operation to obtain corresponding predicted proposal box parameters and predicted foreground probabilities;

[0026] Step S33: Determine the corresponding initial proposal boxes according to the predicted proposal box parameters and the anchor coordinates.

[0027] Optionally, step S3 further includes: performing a screening operation on the initial proposal boxes determined by each anchor.

[0028] Optionally, the step of performing a screening operation on the initial proposal boxes determined by each anchor includes:

[0029] Filter out the proposal boxes with an overlapping area greater than the overlap threshold in the initial proposal boxes;

[0030] Set hyperparameters, and filter out the proposal boxes with the predicted foreground probability greater than the hyperparameters from the filtered proposal boxes as the final proposal boxes, so as to detect the ship to-be-detected image through the final proposal boxes.

[0031] Optionally, before using the current image encoder as the backbone network model, the method further includes: updating the negative sample feature queue.

[0032] Optionally, the step of updating the negative sample feature queue includes:

[0033] When the negative sample data in the negative sample feature queue dequeues, enqueue the image encoding data of the image corresponding to the negative sample data after image enhancement processing and image momentum encoder processing, so as to update the negative sample feature queue.

[0034] Optionally, the adaptive ship detection method is applied to an adaptive ship detection model. After filtering out the final proposal boxes, the method further includes: determining the multi-head perceptron loss and the adaptive RPN loss of all the final proposal boxes, and training the adaptive ship detection model according to the multi-head perceptron loss and the adaptive RPN loss.

[0035] The present invention has at least the following technical effects:

[0036] The present invention pre-trains the backbone network model through contrastive learning, which can save a large amount of manpower and material resources spent on sample labeling, and can enable the backbone network model after pre-training to output higher-quality feature maps; in addition, the adaptive region proposal network model in the present invention can directly predict proposal boxes through anchor points, without the need to pre-define different-scale anchor boxes for ships, and without setting hyperparameters related to anchor boxes, thereby improving the execution efficiency of the adaptive ship detection model, and because the present invention directly predicts proposal boxes through anchor points, the present invention can predict targets with large size changes, so as to avoid the problem that it is difficult for anchor boxes to match the real boxes when the target size changes greatly; the number of proposal boxes generated by the adaptive region proposal network model in the present invention is much smaller than the number of proposal boxes generated by the traditional region proposal network model, so the present invention can save a large amount of computing resources when calculating the coordinates of proposal boxes and filtering proposal boxes subsequently.

[0037] Additional aspects and advantages of the present invention will be given in part in the following description, become apparent in part from the following description, or be understood through the practice of the present invention. Description of the Drawings

[0038] Figure 1The working principle diagram of the ship target detection model based on Faster R-CNN;

[0039] Figure 2 The flowchart of the adaptive ship detection method based on contrast learning provided by an embodiment of the present invention;

[0040] Figure 3 The working principle diagram of the contrast learning pre-training model provided by an embodiment of the present invention;

[0041] Figure 4 The flowchart of the adaptive ship detection method based on contrast learning provided by a specific example of the present invention. Detailed implementation manners

[0042] The following details this embodiment. The examples of the embodiment are shown in the accompanying drawings, where the same or similar reference numerals represent the same or similar elements or elements with the same or similar functions throughout. The embodiments described below with reference to the accompanying drawings are exemplary and are intended to explain the present invention and should not be construed as limiting the present invention.

[0043] The following describes the adaptive ship detection method based on contrast learning of this embodiment with reference to the accompanying drawings.

[0044] Figure 2 The flowchart of the adaptive ship detection method based on contrast learning provided by an embodiment of the present invention. As Figure 2 shown, the method includes:

[0045] Step S1: Pre-train the backbone network model.

[0046] Among them, step S1 includes:

[0047] Step S11: Obtain the ship image dataset.

[0048] Step S12: Perform image enhancement processing and image encoding processing on each image in the ship image dataset, and obtain a number of positive sample pair features and a number of negative sample pair features according to the image data after the image encoding processing.

[0049] The step S12 includes:

[0050] S121: Perform two image enhancement processes on the same image to obtain the first image enhancement processing result data and the second image enhancement processing result data respectively;

[0051] S122: Input the first image enhancement processing result data as anchor sample data into the image encoder to obtain the first image encoding data, input the second image enhancement processing result data into the image momentum encoder to obtain the second image encoding data, and form a positive sample pair feature by combining the first image encoding data and the second image encoding data.

[0052] S123: Construct a negative sample feature queue, obtain negative sample data from the negative sample feature queue, and form a negative sample pair feature by combining the first image encoding data and the negative sample data. The negative sample data in the negative sample feature queue is the image encoding data obtained by processing the anchor sample data of other images through the image encoder.

[0053] Specifically, a ship image dataset prepared for pre-training can be obtained. For example, a large-scale ship image dataset ImageNet (large-scale image database) can be obtained. Then, as Figure 3 shown, perform image enhancement processing on each image twice. The first image enhancement processing result data and the second image enhancement processing result data obtained after image enhancement processing of the same image form a positive sample pair. Then, use the first image enhancement processing result data as anchor sample data and form a negative sample pair with the anchor sample data generated by other images.

[0054] Among them, the image enhancement processing methods are random cropping, Gaussian blur, and mosaic processing. Specifically, different-sized image patches can be randomly cropped on the image, the image patches can be blurred through a Gaussian blur filter, and multiple cropped image patches can be stitched together. For example, randomly select image patches from different images and crop them to a suitable size for the second time, and then stitch multiple image patches into a target image.

[0055] Furthermore, as Figure 3-4 shown, the image encoder and the image momentum encoder can be used to encode the positive sample data in the positive sample pair, that is, the first image enhancement processing result data and the second image enhancement processing result data, to obtain the first image encoding data and the second image encoding data, that is, the anchor sample feature and the positive sample feature. The anchor sample feature and the positive sample feature form the positive sample pair feature.

[0056] In this embodiment, a queue with a length of N can be randomly initialized. The queue elements represent initial negative sample features, that is, negative sample data. Among them, N is a hyperparameter, which can be set to 4096 here to generate enough negative sample pair features. In this embodiment, as Figure 4 shown, negative sample features, that is, negative sample data, can be selected from the negative sample feature queue, and then a negative sample pair feature can be formed by combining it with the anchor sample feature, that is, the first image encoding data.

[0057] It should be noted that the image encoder and the image momentum encoder are networks with the same structure. Among them, the parameters of the image momentum encoder can be updated through the following formula:

[0058] θ ME ←mθ ME +(1 - m)θ E (1)

[0059] Among them, θ ME is the model parameter of the image momentum encoder, θ E is the model parameter of the image encoder, and m ∈ [0, 1] is a hyperparameter used to determine the update rate of the model parameters of the image momentum encoder. Among them, the parameter m can be set to 0.999, so as to slowly update the image momentum encoder model to ensure that the model parameters are as consistent as possible when calculating the feature of different samples.

[0060] Step S13: Pre-train the backbone network model based on the features of a number of positive sample pairs and the features of a number of negative sample pairs.

[0061] Among them, step S13 includes:

[0062] Step S131: Construct a loss function model;

[0063] Step S132: Input the features of a number of positive sample pairs and the features of a number of negative sample pairs into the loss function model to calculate the contrast loss between samples. When the contrast loss value reaches the minimum and basically no longer changes, use the current image encoder as the backbone network model, that is, when the image encoder converges, use the current image encoder as the backbone network model.

[0064] Among them, the loss function model can be expressed by the following formula:

[0065]

[0066] Among them, L n is the loss function, x p is the anchor sample feature, x i is the positive sample feature, x p and x i constitute the positive sample pair feature, x j is the negative sample feature, x p and x j constitute the negative sample pair feature. x p and x i are respectively generated by the image encoder and the image momentum encoder, x jProvided by the negative sample feature queue X, where f represents the similarity between different samples. In this embodiment, by minimizing this loss function, the goal of maximizing the similarity of positive sample pairs and minimizing the similarity of negative sample pairs can be achieved, so as to construct a high-quality semantic representation space for downstream tasks. That is, when the value of the loss function is the smallest, as Figure 4 shown, the current image encoder can be used as the backbone network model in the ship detection model used in the downstream task, thereby completing the training process of the backbone network model.

[0067] In this embodiment, before using the current image encoder as the backbone network model, the negative sample feature queue needs to be updated.

[0068] Among them, the steps for updating the negative sample feature queue include:

[0069] When the negative sample data in the negative sample feature queue dequeues, the image encoding data of the image corresponding to the negative sample data after image enhancement processing and image momentum encoder processing is enqueued to update the negative sample feature queue.

[0070] For example, when the negative sample data at the head of the negative sample feature queue dequeues, the image encoding data processed by the image momentum encoder, that is, the positive sample features, can be enqueued to update the queue.

[0071] In this embodiment, using contrastive learning for the pre-training of the backbone network model can save a large amount of manpower and material resources spent on sample labeling, and the saved resources can be used to construct a larger-scale pre-training dataset. The pre-training method based on supervised learning is prone to the situation of model capacity saturation and insignificant improvement in detection accuracy on large-scale datasets. For contrastive learning pre-training, the more training samples are used, the more prior knowledge the model can extract, and the better its performance in downstream tasks.

[0072] In addition, contrastive learning can maximize the mutual information between positive sample pairs and project them to adjacent positions in the feature space. For negative samples, they can be evenly projected into the feature space. In this embodiment, by maximizing the similarity between positive samples and minimizing the similarity between negative samples, a high-quality feature space can be constructed, so that contrastive learning can obtain better results than supervised learning in downstream tasks.

[0073] Step S2: Obtain a ship image dataset to be detected, and input the ship image dataset to be detected into the backbone network model. Through the backbone network model and the feature pyramid model connected to the backbone network model, several feature maps are obtained, where the several feature maps are used to detect ship images of different scales.

[0074] In this embodiment, the backbone network model obtained through contrastive learning and the feature pyramid model connected thereto can obtain feature maps with higher quality.

[0075] It should be noted that, as Figure 4 shown, the generation of the proposal boxes is completed by the ship detection model. In this embodiment, the image encoder pre-trained through contrastive learning, that is, the backbone network model, can be migrated to the ship detection model, and then the obtained ship image dataset to be detected is input into the backbone network model, and a plurality of feature maps are output through the feature pyramid model.

[0076] As an example, ship images to be detected can be selected in batches, scaled to a unified size, and then input into the backbone network model. The backbone network model is connected to the feature pyramid model, that is, a convolutional network is connected in parallel after the convolutional layers conv2_x, conv3_x, conv4_x, and conv5_x of ResNet50 (a deep residual network with 50 convolutional layers). The ship image to be detected undergoes a MaxPooling (maximum pooling) operation based on the convolutional layer conv5_x, and finally five feature maps F 1 -F 5 are obtained, which are used to detect ship images of different scales. The size of the feature map is h i *w i *128, i = 1, 2,..., 5.

[0077] Step S3: Input the plurality of feature maps into the adaptive region proposal network model connected to the feature pyramid model, and process them through the adaptive region proposal network model to obtain proposal boxes that meet the preset requirements, so as to detect the ship image to be detected through the proposal boxes.

[0078] Among them, step S3 includes:

[0079] Step S31: Obtain the anchor point coordinates corresponding to each feature map;

[0080] Step S32: Construct a multi-head perceptron model, input the anchor point coordinates corresponding to each feature map into the multi-head perceptron model for calculation, and obtain the corresponding predicted proposal box parameters and predicted foreground probabilities;

[0081] Step S33: Determine the corresponding initial proposal boxes according to the predicted proposal box parameters and the anchor point coordinates.

[0082] The adaptive region proposal network model in this embodiment can implement proposal box prediction and loss calculation. In this embodiment, all feature maps can be input into the adaptive region proposal network model for proposal box prediction.

[0083] Specifically, assuming the original image size is h0 *w 0 *3, for all That is, the anchor point, and the predicted proposal box parameters (l i , r i , t i , b i ) and the predicted foreground probability p i can be obtained through the multi-head perceptron model, where α = 0, 1,..., h i -1, β = 0, 1,..., w i -1, Both are rounded down to obtain integer coordinates. Thus, the vertex coordinates of the proposal box can be calculated as (x i -l i , y i -t i ), (x i +r i , y i -t i ), (x i -l i , y i +b i ), (x i +r i , y i +b i ), and the initial proposal box is predicted.

[0084] The step S3 further includes: screening the initial proposal boxes determined by each anchor point. Among them, the step of screening the initial proposal boxes determined by each anchor point includes:

[0085] Filter out the proposal boxes with overlapping areas greater than the overlap threshold in the initial proposal boxes;

[0086] Set hyperparameters, and select the proposal boxes with predicted foreground probabilities greater than the hyperparameters from the filtered proposal boxes as the final proposal boxes, so as to detect the ship image to be detected through the final proposal boxes.

[0087] Specifically, the proposal boxes with too large overlapping areas can be filtered out through the NMS algorithm, then set the hyperparameter P = 0.5, and output all p i > P as the final proposal boxes, so as to detect the ship image to be detected through the final proposal boxes. It should be noted that there may be more than one final proposal box.

[0088] In one embodiment of the present invention, the adaptive ship detection method is applied to the adaptive ship detection model, i.e., the ship detection model described above. After the final proposal boxes are screened, the method further includes: determining the multi-head perceptron loss and the adaptive RPN loss of all the final proposal boxes, and training the adaptive ship detection model according to the multi-head perceptron loss and the adaptive RPN loss.

[0089] Specifically, the positive and negative samples used to calculate the adaptive RPN loss can be defined first: Assume that t is a true box in the original image. If the anchor point (x i , y i ) falls within t, that is, when the corresponding proposal box successfully matches the ship target, this proposal box is set as a positive sample, otherwise it is a negative sample. If the anchor point (x i , y i ) falls within multiple true boxes, the corresponding proposal box is set as a negative sample.

[0090] For all positive samples, the proposal box loss is calculated through the Smooth L1 loss function, and the formula is as follows:

[0091]

[0092]

[0093] Where L box represents the proposal box loss, d represents the difference between the proposal box parameter d i and the true box parameter , {l * , r * , t * , b *} is calculated through the true box.

[0094] The foreground probability loss is calculated through the cross-entropy loss function, and the formula is as follows:

[0095] L c =-∑q i logp i +(1 - q i )log(1 - p i ) (5)

[0096] Where L c is the foreground probability loss, q i is the true foreground probability, and p i is the predicted foreground probability.

[0097] Scale the suggestion box from the original image scale to the feature map scale, that is, adjust it to the size of the Region of Interest (ROI). Since 5 feature maps F 1 -F 5 are obtained through the feature pyramid model, different scaling ratios can be set for suggestion boxes of different areas here: for example, set the suggestion box area threshold to {32 2 , 64 2 , 128 2 , 256 2}. Divide the suggestion boxes into five parts according to their areas, and then select the corresponding scaling ratio i ∈ 1, …, 5 for scaling. Finally, obtain the two-dimensional image of the ROI region.

[0098] Furthermore, the two-dimensional image of the ROI region can be flattened to obtain one-dimensional serialized data, and the one-dimensional serialized data is input into the multi-head perceptron model to obtain the target box class probability and the target box regression parameters through the multi-head perceptron model. Specifically, the target box class probability loss can be calculated through the cross-entropy loss function:

[0099] L′ C = -∑q′ i logp′ i + (1 - q′ i )log(1 - p′ i ) (6)

[0100] where L′ C is the target box class probability loss, q′ i is the true class probability, p′ i is the predicted class probability. Then, the target box regression parameter loss is calculated through the Smooth L1 loss function:

[0101]

[0102] where, L′ box is the target box regression parameter loss, and d′ is the difference between the target box regression parameter and the true box regression parameter.

[0103] Finally, the total loss is obtained as:

[0104] L total = αL box + βL c + γL′ box + δL′ C (8)

[0105] where, α, β, γ, δ are for L box , L c , L′ box , L′C The weights are all set to 0.25 here.

[0106] In this embodiment, the sum of the adaptive RPN loss and the multi-head perceptron loss of all proposed boxes can be calculated to obtain the final total loss. Through multiple iterations of the total loss, the adaptive ship detection model can be trained. When the total loss reaches the minimum value and basically no longer changes, the target ship detection model can be trained.

[0107] In summary, the present invention pre-trains the backbone network model through contrastive learning, which can save a large amount of manpower and material resources spent on sample labeling, and can enable the backbone network model after pre-training to output higher-quality feature maps. In addition, the adaptive region proposal network model in the present invention can directly predict proposed boxes through anchor points, without the need to pre-define different-scale anchor boxes for ships, and without setting hyperparameters related to anchor boxes, thereby improving the execution efficiency of the adaptive ship detection model. And because the present invention directly predicts proposed boxes through anchor points, the present invention can predict targets with large size changes, thereby avoiding the problem that it is difficult for anchor boxes to match the true boxes when the target size changes greatly. The number of proposed boxes generated by the adaptive region proposal network model in the present invention is much smaller than the number of proposed boxes generated by the traditional region proposal network model. Therefore, the present invention can save a large amount of computing resources when calculating the coordinates of proposed boxes and filtering proposed boxes in the follow-up.

[0108] It should be noted that in this article, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements not only includes those elements, but also includes other elements not expressly listed, or also includes elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "including a..." does not exclude the existence of additional identical elements in the process, method, article or device including the element.

[0109] Although the content of the present invention has been introduced in detail through the above preferred embodiments, it should be recognized that the above description should not be considered as a limitation of the present invention. After those skilled in the art read the above content, various modifications and substitutions to the present invention will be obvious. Therefore, the protection scope of the present invention should be defined by the appended claims.

Claims

1. An adaptive ship detection method based on contrast learning, characterized in that, it includes: Step S1: Pre-train the backbone network model, which specifically includes: Step S11: Obtain the ship image data set; Step S12: Perform image enhancement processing and image encoding processing on each image in the ship image data set, and obtain a number of positive sample pair features and a number of negative sample pair features according to the image data after image encoding processing; Step S13: Pre-train the backbone network model according to a number of the positive sample pair features and a number of the negative sample pair features; Step S131: Construct a loss function model; Step S132: Input a number of the positive sample pair features and a number of the negative sample pair features into the loss function model to calculate the contrast loss between samples, and when the contrast loss value reaches the minimum and no longer changes, use the current image encoder as the backbone network model; Step S2: Obtain the ship image data set to be detected, and input the ship image data set to be detected into the backbone network model. A number of feature maps are obtained through processing by the backbone network model and a feature pyramid model connected to the backbone network model, wherein the number of the feature maps is used to detect ship images of different scales; Step S3: Input a number of the feature maps into an adaptive region proposal network model connected to the feature pyramid model, and obtain proposal boxes that meet preset requirements through processing by the adaptive region proposal network model, so as to detect the ship image to be detected through the proposal boxes; Step S31: Obtain the anchor point coordinates corresponding to each feature map; Step S32: Construct a multi-head perceptron model, the multi-head perceptron model is located in the adaptive region proposal network model, input the anchor point coordinates corresponding to each feature map into the multi-head perceptron model for operation, and obtain corresponding predicted proposal box parameters and predicted foreground probabilities; Step S33: Determine the corresponding initial proposal boxes according to the predicted proposal box parameters and the anchor point coordinates.

2. The adaptive ship detection method based on contrast learning according to claim 1, characterized in that, the step S12 includes: S121: Perform two image enhancement processes on the same image to obtain first image enhancement result data and second image enhancement result data respectively; S122: Input the first image enhancement result data as anchor sample data into the image encoder for processing to obtain first image encoding data, input the second image enhancement result data into the image momentum encoder for processing to obtain second image encoding data, and form the positive sample pair features by combining the first image encoding data and the second image encoding data; S123: Construct a negative sample feature queue, obtain negative sample data from the negative sample feature queue, and form the negative sample pair features by combining the first image encoding data and the negative sample data, wherein the negative sample data in the negative sample feature queue is the image encoding data obtained by processing the anchor sample data of other images through the image encoder.

3. The adaptive ship detection method based on contrast learning according to claim 2, wherein, the step S3 further includes: performing a screening operation on the initial proposal boxes determined by each anchor point.

4. The adaptive ship detection method based on contrast learning according to claim 3, wherein, the step of performing a screening operation on the initial proposal boxes determined by each anchor point includes: filtering out the proposal boxes with overlapping areas greater than the overlapping threshold in the initial proposal boxes; setting hyperparameters, and screening out the proposal boxes with a predicted foreground probability greater than the hyperparameters from the filtered proposal boxes as the final proposal boxes, so as to detect the ship image to be detected through the final proposal boxes.

5. The adaptive ship detection method based on contrast learning according to claim 1, wherein, before using the current image encoder as the backbone network model, the method further includes: updating the negative sample feature queue.

6. The adaptive ship detection method based on contrast learning according to claim 5, wherein, the step of updating the negative sample feature queue includes: when the negative sample data in the negative sample feature queue dequeues, enqueueing the image encoding data of the image corresponding to the negative sample data after image enhancement processing and image momentum encoder processing, so as to update the negative sample feature queue.

7. The adaptive ship detection method based on contrast learning according to claim 4, wherein, the adaptive ship detection method is applied to an adaptive ship detection model. After screening out the final proposal boxes, the method further includes: determining the multi-head perceptron loss and the adaptive RPN loss of all the final proposal boxes, and training the adaptive ship detection model according to the multi-head perceptron loss and the adaptive RPN loss.

Citation Information

Patent Citations

  • Fine-grained ship identification method based on comparative learning

    CN113255793A

  • Self-supervised visual model pre-training method based on dense semantic comparison

    CN113989582A