Ship re-identification method based on sample expansion

By using a sample expansion-based approach, leveraging the dynamic clustering of HDBSCAN and DINO-ViT models and the multi-scale fuzzy clustering of the CMS-FCM model, combined with adversarial learning, the problems of limited ship datasets and pseudo-label noise were solved, thereby improving the accuracy of ship re-identification and the model's generalization ability.

CN120808250APending Publication Date: 2025-10-17NANJING TECH UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510674974.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-23
Publication Date
2025-10-17

AI Technical Summary

Technical Problem

Existing unsupervised ship re-identification methods are not ideal under the influence of limited ship datasets and pseudo-label noise. In particular, the differences between different viewpoints of the same ship are large (large intra-class differences), while the differences between different ships at the same viewpoint are small (small inter-class differences), resulting in poor model generalization ability and inaccurate clustering results.

Method used

A sample expansion-based approach is adopted, using the HDBSCAN algorithm for dynamic clustering, the DINO-ViT model to capture global and local features, and the CMS-FCM model for multi-scale fuzzy clustering. Adversarial learning and random combination are combined to generate training samples, optimize the weight parameters of the ship re-identification clustering model, and reduce the influence of label noise and image style.

Benefits of technology

It improves the accuracy of ship re-identification and the model's generalization ability, enhances the network's robustness, and ensures accurate identification and differentiation under different camera perspectives.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120808250A_ABST
    Figure CN120808250A_ABST
Patent Text Reader

Abstract

The invention discloses a ship re-identification method based on sample expansion. A construction process of a ship re-identification clustering model comprises the following steps: obtaining a ship original image set and performing data expansion to obtain a training sample; training a ship re-identification clustering model by using the training sample, wherein the ship re-identification clustering model comprises a dynamic clustering module, a clustering intra-class recombination module and a clustering inter-class recombination module; acquiring ship monitoring data in a set sea area, and inputting the ship monitoring data into a preset ship re-identification clustering model to obtain a ship unsupervised identification result; under the condition of few ship sample data sets, the negative influence of label noise and image styles on the ship re-identification effect is reduced, and the accuracy of ship re-identification is improved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application belongs to the technical field of image recognition, and particularly relates to an unsupervised ship re-identification method. BACKGROUND

[0002] The ocean is the main channel for global trade and transportation. Maritime transportation not only supports international trade, but also promotes the stability and development of the global economy. The safety of maritime traffic directly affects the smoothness of global supply chains and the stability of the economy. However, as the process of globalization accelerates, the busy maritime traffic also makes the ocean a breeding ground for illegal activities. It has become difficult to rely solely on traditional monitoring methods to cope with the challenges of the new situation, and ship re-identification technology has begun to be used in maritime management.

[0003] Ship re-identification, as an important auxiliary means for regulating ships, accurately identifies and distinguishes the identities of different ships through the analysis of ship images or video streams, effectively preventing maritime criminal activities, and ensuring the safety of international maritime trade. The main challenges of existing unsupervised ship re-identification methods are the lack of ship data sets and the fact that pseudo-labels are easily affected by noise such as different camera styles, lighting, and viewing angles, resulting in unsatisfactory results.

[0004] For the problem of lack of ship data, existing methods mainly use transfer learning. However, the generalization ability of the model trained by transfer learning is often not ideal, which is due to the fact that ship pictures have a large difference between different angles of the same ship (large intra-class difference) and a small difference between different ships at the same angle (small inter-class difference).

[0005] For the problem that ship clustering is easily affected by pseudo-label noise, resulting in poor clustering results, the current main method to solve the pseudo-label noise problem is to use soft labels. However, soft labels require a large data set to support, and in essence, they assign a set of probabilities to a given data point rather than explicit hard labels, which can cause label inconsistency and ambiguity, leading the model to learn some incorrect associations and produce uncertain label assignments, thereby affecting the model's effectiveness. In addition, the features of ships in the same category are less different, and it is difficult to produce reliable pseudo-label results relying on a single clustering algorithm. SUMMARY

[0006] The present application provides a ship re-identification method based on sample expansion, which reduces the negative impact of label noise and image style on unsupervised ship re-identification effectiveness and improves the accuracy of unsupervised ship re-identification in the case of a small number of ship sample data sets.

[0007] To achieve the above purpose, the technical solution adopted by the present application is:

[0008] The first aspect of the present application provides a ship re-identification method based on sample expansion, comprising:

[0009] obtain ship unsupervised identification results by inputting the ship monitoring data into a preset ship re-identification clustering model;

[0010] The training process of the ship re-identification clustering model comprises:

[0011] obtain training samples by performing data expansion on the ship original image set;

[0012] The ship re-identification clustering model comprises a dynamic clustering module, a clustering intra-class reorganization module, and a clustering inter-class reorganization module;

[0013] The training samples are input into the dynamic clustering module, the training samples are clustered by an HDBSCAN algorithm, and the training samples collected by the same camera are divided into sample clusters; the training samples in the sample clusters are grouped by cosine similarity to obtain training sample groups; and the training samples are re-divided after adjusting the center positions of the training sample groups according to the Euclidean distances between the training sample groups to obtain first sample sub-clusters;

[0014] The first sample sub-clusters are input into the clustering intra-class reorganization module, the reliable probabilities of the first sample sub-clusters are calculated, and the first sample sub-clusters are split and reorganized according to the reliable probabilities to obtain second sample sub-clusters;

[0015] The second sample sub-clusters are input into the clustering inter-class reorganization module, the sharing probabilities between the second sample sub-clusters are calculated, and the second sample sub-clusters are merged and reorganized according to the sharing probabilities to obtain third sample sub-clusters; the correlation contrast loss of the third sample sub-clusters is calculated according to the third sample sub-clusters, the weight parameters of the ship re-identification clustering model are optimized according to the correlation contrast loss, the training process of the ship re-identification clustering model is repeated and iterated until a preset termination condition is reached, and the trained ship re-identification clustering model is output.

[0016] Further, the ship original image set is expanded to obtain training samples, specifically comprising:

[0017] The global features and local features in the ship original image are captured by using the self-attention mechanism in the DINO-ViT model, the global features and local features in the ship original image are fused to obtain comprehensive image features;

[0018] The ship mask is obtained by performing multi-scale fuzzy clustering on the comprehensive image features by using the CMS-FCM model; and the ship mask is divided into ship region features and background region features after being smoothed;

[0019] The ship region feature and the background region feature are randomly combined to form a fusion feature map; and the fusion feature map is subjected to adversarial learning to obtain a training sample.

[0020] Further, the ship mask is obtained by performing multi-scale fuzzy clustering on the comprehensive image feature through the CMS-FCM model, and specifically includes:

[0021] The image fuzzy clustering is obtained by performing multi-scale fuzzy clustering on the comprehensive image feature through the fuzzy clustering algorithm .

[0022] The image fuzzy clustering is added with the center of gravity constraint and the buoyancy center constraint, and the ship mask of the comprehensive image feature is obtained according to the image fuzzy clustering . , and the expression formula is:

[0023]

[0024] In the formula, is the number of clusters of the image fuzzy clustering, and r is the serial number of the cluster of the image fuzzy clustering; is the comprehensive image feature vector, is the cluster center of the image fuzzy clustering ; is the weight set for the control center of gravity term, is the weight set for the control buoyancy center term; is the center of gravity of the ship in the comprehensive image feature; is the buoyancy center of the ship in the comprehensive image feature.

[0025] Further, the ship mask is divided into a ship region feature and a background region feature after being smoothed; specifically including:

[0026] The first segmentation mask image is obtained by adding Gaussian noise to the ship mask to smooth the edge of the ship mask, and the expression formula is:

[0027]

[0028] In the formula, is the ship mask of the comprehensive image feature generated according to the image fuzzy clustering , is the Gaussian noise;

[0029] After removing isolated pixels in the first segmentation mask image through an opening operation, the first segmentation mask image is further optimized using a conditional random field to obtain a second segmentation mask image , and the expression formula is:

[0030]

[0031] In the formula, is the pixel with the serial number in the first segmentation mask image; is the pixel set in the first segmentation mask image; represents the segmentation label of the pixel with the serial number in the first segmentation mask image and the feature vector of the pixel with the serial number in the first segmentation mask image ; is the segmentation label of the pixel with the serial number in the first segmentation mask image, is the feature vector of the pixel with the serial number in the first segmentation mask image; is the pixel with the serial number in the first segmentation mask image; represents the segmentation label of the pixel with the serial number in the first segmentation mask image and the association feature of the segmentation label of the pixel with the serial number in the first segmentation mask image ; is the segmentation label of the pixel with the serial number in the first segmentation mask image;

[0032] The similarity between the segmentation label of the pixel with the serial number in the first segmentation mask image and the feature vector of the pixel with the serial number in the first segmentation mask image is calculated, and the expression formula is:

[0033]

[0034] In the formula, represents the feature variance of the segmentation label of the pixel with the serial number in the first segmentation mask image ; is the L1 norm; represents the feature mean of the segmentation label of the pixel with the serial number in the first segmentation mask image ;

[0035] The association feature between the segmentation label of the pixel with the serial number in the first segmentation mask image and the segmentation label of the pixel with the serial number in the first segmentation mask image is calculated ​, the expression formula is:

[0036]

[0037] wherein, is the weight of the smoothing term, is the feature vector of the pixel with the serial number in the first segmentation mask map, is an indicator function, which takes the value 1 when and are not equal, otherwise 0;

[0038] The ship original image is segmented by using the second segmentation mask map to obtain ship region features and background region features , the expression formula is:

[0039]

[0040]

[0041] In the formula, is the ship original image; is the second segmentation mask map.

[0042] Further, the ship region features and the background region features are randomly combined to form a fusion feature map, specifically including:

[0043] The ship region features are randomly rotated at an angle to simulate the situation on the sea, randomly flipped to increase multi-view information, and randomly translated and scaled to simulate the far and near situation to obtain ship simulation features ;

[0044] The ship simulation features and the background region features are randomly combined using an Alpha fusion strategy to form a fusion feature map , the expression formula is:

[0045]

[0046] wherein, is the weight coefficient of feature fusion.

[0047] Further, the fusion feature map is subjected to adversarial learning to obtain a training sample, specifically including:

[0048] A target ship image is selected from the ship original image set; the fusion feature map is input into a generator of adversarial learning to obtain a target style image, and a cycle consistency loss and an adversarial loss of the generated target style image and the target image are calculated by using a discriminator of adversarial learning, the expression formula is:

[0049]

[0050]

[0051] In the formula, is a generator of the adversarial learning from the target domain to the source domain is a generator of the adversarial learning from the source domain to the target domain is a fused feature map is an L1 norm is a target ship image is a mathematical expectation function is a discriminator of the adversarial learning represents a probability that belongs to represents an image generated according to belongs to a probability

[0052] The generator parameters are optimized according to the cycle consistency loss and the adversarial loss, and the adversarial learning process is repeated and iterated until the cycle consistency loss and the adversarial loss converge to obtain an output image of the adversarial learning; and the training sample is generated from the output image of the adversarial learning

[0053] An anchor point image is selected from the ship original image set; the output image is divided into positive samples and negative samples according to the anchor point image, the total loss of data expansion of the original image set is calculated according to the output image of the adversarial learning and the target anchor point image, and the cross-entropy loss , the triplet loss and the perception loss are utilized, and the expression formula is as follows:

[0054]

[0055] In the formula, is the total loss of data expansion of the original image set is a hyperparameter of the cross-entropy loss is a hyperparameter of the triplet loss is a hyperparameter of the perception loss

[0056] The parameters of the CMS-FCM model and the generator are optimized according to the total loss, and the ship original image set expansion process is repeated and iterated until the total loss converges to output a ship expanded image, and the training sample is obtained.

[0057] ​​​​​Furthermore, after adjusting the center positions of the training sample groups according to the Euclidean distances between the training sample groups, the training sample groups are re-divided to obtain the first sample sub-clusters, specifically including:

[0058] If the Euclidean distance between two training sample groups is less than the preset minimum distance threshold , then by introducing the perturbation vector Push the Euclidean distance between the two training sample groups to a greater value if the Euclidean distance between the two training sample groups is greater than the preset maximum distance threshold. , then by introducing the contraction vector To shorten the Euclidean distance between two training sample groups, the expression formula is:

[0059]

[0060]

[0061] In the formula, is the sequence number of the training sample group, For the The center point of the training sample group;

[0062] Recalculate the similarity between the training sample and the center point of the training sample group, remove the training samples whose similarity is lower than the similarity threshold, and re-divide the training samples according to the similarity and the adjusted center point of the training sample group to obtain the first sample sub-cluster.

[0063] Furthermore, the reliability probability of the first sample sub-cluster is calculated; and the first sample sub-cluster is split and reorganized according to the reliability probability to obtain a second sample sub-cluster. Specifically, the method includes:

[0064] Calculate the probability that the training sample belongs to the first sample subcluster , the expression formula is:

[0065]

[0066]

[0067] In the formula, For serial number Training samples of Belongs to the serial number The first sample subcluster The probability of membership; The sequence number in the first sample subcluster is training samples; For serial number The first sample subcluster of For serial number The first sample subcluster The number of training samples in ; is a training sample with a serial number of is similar to a first sample sub-cluster with a serial number of is a training sample with a serial number of in the first sample sub-cluster with a serial number of is a training sample with a serial number of is a training sample with a serial number of is a training sample with a serial number of is a training sample with a serial number of is a training sample with a serial number of has a total number of features that is not repeated;

[0068] averaging the membership probability of the training sample belonging to the first sample sub-cluster obtains a reliable probability of the first sample sub-cluster; the first sample sub-cluster with a reliable probability lower than a first threshold is split and reorganized, and a contrastive loss of the first sample sub-cluster is calculated through a contrastive learning algorithm; after the distance between the training samples is optimized using the contrastive loss, the second sample sub-cluster is obtained by re-dividing.

[0069] Further, the second sample sub-cluster is input into the clustering inter-class reorganization module, the shared probability between the second sample sub-clusters is calculated, and the third sample sub-cluster is obtained by merging and reorganizing the second sample sub-clusters according to the shared probability, specifically including:

[0070] the training samples in the second sample sub-cluster are denoted as training samples , and the membership probability of the training sample belonging to the second sample sub-cluster is calculated ; when the membership probability is greater than a second threshold , the training sample is determined as a similar training sample of the second sample sub-cluster and the second sample sub-cluster ;

[0071] The shared probability of the second sample sub-cluster and the second sample sub-cluster sharing the same ship instance is calculated according to the number of similar training samples of the second sample sub-cluster and the second sample sub-cluster , and the expression formula is:

[0072]

[0073] wherein, is a shared probability of the second sample sub-cluster with the same ship instance as the second sample sub-cluster with the sequence number is a shared probability of the second sample sub-cluster with the same ship instance as the second sample sub-cluster with the sequence number is a number of similar training samples within the second sample sub-cluster with the sequence number is a number of similar training samples within the second sample sub-cluster with the sequence number is a number of training samples within the second sample sub-cluster with the sequence number is a number of training samples within the second sample sub-cluster with the sequence number is a number of training samples within the second sample sub-cluster with the sequence number is a number of training samples within the second sample sub-cluster with the sequence number

[0074] the second sample sub-cluster with the shared probability greater than a set shared probability threshold is merged and reorganized to obtain a third sample sub-cluster.

[0075] Further, a correlation contrast loss of the third sample sub-cluster is calculated according to the third sample sub-cluster, a weight parameter of the ship re-identification clustering model is optimized according to the correlation contrast loss, and a training process of the ship re-identification clustering model is repeated and iterated until a set termination condition is reached to output the trained ship re-identification clustering model; specifically including:

[0076] calculating a membership probability of a training sample belonging to the third sample sub-cluster ; calculating a correlation contrast loss and a difficulty index of clustering and reorganization of the third sample sub-cluster according to the membership probability , and the expression formula is:

[0077]

[0078]

[0079] wherein, and are adjustment parameters of the correlation contrast loss ; is a difficulty index of clustering and reorganization of the third sample sub-cluster with the sequence number is a number of training samples within the third sample sub-cluster with the sequence number is a number of training samples within the third sample sub-cluster with the sequence number ​​​​​​​​​​​​​ For serial number The third sample subcluster of The serial number is training samples;

[0080] The weight parameters of the ship re-identification clustering model are optimized based on the association contrast loss, and the difficulty index is used to dynamically adjust the degree of attention the ship re-identification clustering model pays to the training samples; the association contrast loss function is used to optimize the parameters of the ship re-identification clustering model;

[0081] Repeat the iteration to train the ship re-identification clustering model. When the silhouette coefficient of the clustering model reaches the set silhouette threshold in 5 consecutive iterations, Or the Calinski-Harabasz index of the model clustering result reaches the set index threshold in 5 consecutive iterations When , the training is ended and the final clustering result of the model is obtained, that is, the result of ship re-identification, and the trained ship re-identification clustering model is output.

[0082] Compared with the prior art, the present invention has the following beneficial effects:

[0083] The present invention re-identifies ship images based on intra-class and inter-class reorganization of training samples, effectively reducing the negative impact of label noise and image style on the unsupervised ship re-identification effect, improving the accuracy of label recognition, and enhancing the generalization ability and robustness of the network, while also improving the accuracy of ship re-identification.

[0084] The present invention inputs training samples into a dynamic clustering module and clusters the training samples using the HDBSCAN algorithm. The HDBSCAN algorithm can automatically determine the number of clusters and effectively identify outliers. By discarding outliers, it can effectively remove noise points and accurately divide training samples collected by the same camera into sample clusters. The cosine similarity is used to group the training samples in the sample cluster to obtain training sample groups, which can ensure that images captured by the same camera can be grouped more accurately, thereby reducing confusion caused by similar appearance. After adjusting the center position of the training sample group according to the Euclidean distance between the training sample groups, the training samples are re-divided to obtain a first sample sub-cluster. A second pseudo-label that identifies the ship instance is added to the first sample sub-cluster. This can improve the quality of clustering and ensure that ship images under different camera perspectives can be accurately identified and distinguished.

[0085] The first sample sub-cluster is input into the intra-cluster and inter-cluster reorganization module, the first sample sub-cluster is split and reorganized according to the reliable probability to obtain a second sample sub-cluster; the second sample sub-cluster is input into the inter-cluster reorganization module, the sharing probability between the second sample sub-clusters is calculated, and the second sample sub-cluster is merged and reorganized according to the sharing probability to obtain a third sample sub-cluster; the first sample sub-cluster is reorganized in the intra-cluster, the quality of the cluster is improved, the second sample sub-cluster is reorganized in the inter-cluster, and each ship instance obtains a more accurate label in different camera sub-clusters, so that the performance of the model in the re-identification task is more accurate and reliable. BRIEF DESCRIPTION OF DRAWINGS

[0086] Figure 1 A schematic diagram of the ship re-identification method based on sample expansion disclosed in Embodiment 1;

[0087] Figure 2 A flowchart of the expanded ship image disclosed in Embodiment 2;

[0088] Figure 3 A flowchart of the sample dynamic clustering disclosed in Embodiment 2;

[0089] Figure 4 A flowchart of the intra-cluster reorganization and inter-cluster reorganization of the sample clustering disclosed in Embodiment 2;

[0090] Figure 5 A flowchart of the ship re-identification method based on sample expansion disclosed in Embodiment 2. DETAILED DESCRIPTION

[0091] The application will be further described below with reference to the accompanying drawings. The following examples are only used to more clearly illustrate the technical solutions of the application, and cannot be used to limit the protection scope of the application.

[0092] Embodiment 1

[0093] The embodiment provides a ship re-identification method based on sample expansion, which comprises the following steps:

[0094] Ship monitoring data in a set sea area is acquired, and the ship monitoring data is input into a preset ship re-identification clustering model to obtain a ship supervision identification result;

[0095] The training process of the ship re-identification clustering model comprises the following steps:

[0096] I. A ship original image set is acquired, and the ship original image set is expanded to obtain training samples, specifically comprising the following steps:

[0097] Global features and local features in the ship original image are fused to obtain a fused image feature by using a self-attention mechanism in a self-supervised visual learning model (referred to as a DINO-ViT model) to capture the global features and the local features in the ship original image.

[0098] The embodiment adds physical constraints to the fuzzy clustering algorithm to obtain more accurate clustering results. The new fuzzy clustering algorithm is referred to as a constrained multi-scale fuzzy clustering model (Constrained Multi-Scale Fuzzy C-Means), referred to as a CMS-FCM model. The CMS-FCM model is used to perform multi-scale fuzzy clustering on the fused image feature to obtain a ship mask. The ship mask is smoothed and divided into a ship region feature and a background region feature.

[0099] The ship region feature and the background region feature are randomly combined to form a comprehensive feature map. The comprehensive feature map is subjected to adversarial learning to obtain a training sample. The adversarial learning makes the generated training sample consistent with the feature style of the ship original image. The ship original image set is expanded to solve the problem of a small number of ship data sets and uneven distribution.

[0100] II. Training a ship re-identification clustering model using the training sample, specifically including:

[0101] The ship re-identification clustering model includes a dynamic clustering module, a clustering intra-class reorganization module, and a clustering inter-class reorganization module.

[0102] The training sample is input into the dynamic clustering module, and the training sample is clustered by an HDBSCAN algorithm (Hierarchical Density-Based Spatial Clustering of Applications with Noise). Training samples collected by the same camera are divided into sample clusters, and a first pseudo-label identifying the camera identity is added to the sample cluster. The training samples in the sample cluster are grouped by cosine similarity to obtain a training sample group. After adjusting the center position of the training sample group according to the Euclidean distance between the training sample groups, the training sample is re-divided to obtain a first sample sub-cluster. A second pseudo-label identifying the ship instance is added to the first sample sub-cluster.

[0103] The dynamic clustering module clusters the ship original image dataset according to the camera style to obtain different sample clusters. The camera style refers to the fact that most cameras are fixed, and the feature distribution of ship pictures under the same camera (picture background, angle, distance, etc.) has great commonality. The commonality is used to divide ship original images collected by different cameras. Then, different ship identities under the same camera style are divided to obtain a first sample sub-cluster. The style division method greatly reduces the noise influence of different camera styles.

[0104] The first sample sub-cluster is input into the clustering intra-reorganization module, and the reliable probability of the first sample sub-cluster is calculated. The first sample sub-cluster is split and reorganized according to the reliable probability, the contrast loss of the first sample sub-cluster is calculated through a contrast learning algorithm, and the second sample sub-cluster is obtained by re-dividing after optimizing the distance between training samples using the contrast loss. The reliability and contrast learning are used to ensure the quality of the cluster;

[0105] The second sample sub-cluster is input into the clustering inter-reorganization module, and the shared probability between the second sample sub-cluster is calculated. The third sample sub-cluster is obtained by merging and reorganizing the second sample sub-cluster according to the shared probability;

[0106] The membership probability of the training sample belonging to the third sample sub-cluster is calculated. The associated contrast loss is calculated according to the membership probability of the training sample belonging to the third sample sub-cluster and the difficulty index of the third sample sub-cluster clustering and reorganization. The difficulty index is used to dynamically adjust the attention degree of the ship re-identification clustering model to the training sample. The associated contrast loss function is used to optimize the parameters of the ship re-identification clustering model, thereby solving the problem of dividing the same ship instance into several different cross-camera clusters;

[0107] The third sample sub-cluster is obtained, that is, the training sample recognition result is obtained. The silhouette coefficient and Calinski-Harabasz index of the model clustering result are calculated. When the silhouette coefficient of the model clustering result reaches the set silhouette threshold value or the Calinski-Harabasz index of the model clustering result reaches the set index threshold value in the last 5 iterations, the training is ended, and the final clustering result of the model is obtained, that is, the ship re-identification result, and the trained ship re-identification clustering model is output.

[0108] As Figure 1As shown, the data expansion is performed on the original image set of the ship to solve the problem of small sample data set. In order to prevent the number of samples in different clusters from being too large, which leads to poor learning effect, a dynamic clustering module is designed. The training samples are divided into different camera clusters according to the camera style by using HDBSCAN, and then the number distribution of each cluster is adjusted to be similar by using the dynamic center method, so that the first sample sub-cluster divided according to the camera style is obtained. Then, the quality of the second sample sub-cluster formed by clustering is ensured by using intra-class reorganization of clustering, and then the third sample sub-cluster is obtained by merging the clusters of the same ship identity divided into different camera clusters due to the camera style by using inter-class reorganization of clustering. The third sample sub-cluster is iteratively optimized by using correlation contrast learning, and finally the reliable and accurate clustering result is obtained.

[0109] Embodiment 2

[0110] As Figures 2 to 5 shown, the present embodiment provides a ship re-identification method based on sample expansion, comprising:

[0111] Obtaining ship monitoring data in a set sea area, inputting the ship monitoring data into a preset ship re-identification clustering model to obtain a ship supervision identification result;

[0112] The construction process of the ship re-identification clustering model comprises:

[0113] Obtaining a set of original ship images, expanding the set of original ship images to obtain training samples, specifically comprising:

[0114] Capturing global features and local features in the original ship images by using the self-attention mechanism in the DINO-ViT model, fusing the global features and local features in the original ship images to obtain fused image features , the formula is:

[0115]

[0116] The formula is, is a multi-scale fusion function; is the feature representation of each layer.

[0117] Obtaining a ship mask by performing multi-scale fuzzy clustering on the comprehensive image features by using the CMS-FCM model, specifically comprising:

[0118] Performing multi-scale fuzzy clustering on the comprehensive image features by using a fuzzy clustering algorithm to obtain image fuzzy clustering ; it can better capture the details of the picture, which is conducive to increasing the generalization ability of the ship re-identification clustering model.

[0119] The image fuzzy clustering Add ship's center of gravity constraint and buoyancy center constraint, and perform fuzzy clustering based on the image. Generate ship masks with comprehensive image features , the expression formula is:

[0120]

[0121] In the formula, is the number of clusters of image fuzzy clustering, r is the serial number of the cluster of image fuzzy clustering; s is the feature vector of the comprehensive image, Fuzzy clustering for images The cluster center of The weight set for the control center of gravity term, The weight set for controlling the center of buoyancy; is the center of gravity of the ship in the comprehensive image features; is the buoyancy center of the ship in the comprehensive image features.

[0122] The ship mask is smoothed and divided into ship area features and background area features; specifically including:

[0123] By adding Gaussian noise to the ship mask and smoothing the edge of the ship mask, the first segmentation mask map is obtained. , the expression formula is;

[0124]

[0125] In the formula, To perform fuzzy clustering based on images Generate a ship mask that synthesizes image features, is Gaussian noise; Gaussian noise can effectively smooth edges and reduce the appearance of jagged edges.

[0126] After removing isolated pixels in the first segmentation mask through opening operation, the first segmentation mask is further optimized using conditional random field to obtain the second segmentation mask. , the expression formula is:

[0127]

[0128] In the formula, The sequence number in the first segmentation mask is pixels; is a set of pixels in the first segmentation mask image; Indicates that the sequence number in the first segmentation mask image is The segmentation labels of the pixels and the sequence number in the first segmentation mask is The feature vector of the pixel The similarity of The sequence number in the first segmentation mask is The segmentation labels of the pixels, The sequence number in the first segmentation mask is The feature vector of the pixel; The sequence number in the first segmentation mask is pixels; Indicates that the sequence number in the first segmentation mask image is The segmentation labels of the pixels The sequence number in the first segmentation mask is The segmentation labels of the pixels The associated features of The sequence number in the first segmentation mask is The segmentation labels of the pixels;

[0129] Calculate the sequence number in the first segmentation mask image The segmentation labels of the pixels and the sequence number in the first segmentation mask is The feature vector of the pixel Similarity , the expression formula is:

[0130]

[0131] In the formula, Indicates that the sequence number in the first segmentation mask image is The segmentation labels of the pixels The characteristic variance of is the L1 norm, Indicates that the sequence number in the first segmentation mask image is The segmentation labels of the pixels The feature mean of the first segmentation mask is The feature vector of the pixel and serial number The segmentation label The mean distance of is proportional to the energy of the data item. The smaller the distance, the lower the energy, indicating that the sequence number in the first segmentation mask is The feature vector of the pixel and serial number The segmentation labels of the pixels The higher the degree of match.

[0132] Calculate the sequence number in the first segmentation mask image The segmentation labels of the pixels The sequence number in the first segmentation mask is The segmentation labels of the pixels The associated features , the expression formula is:

[0133]

[0134] wherein, is a weight of a smoothing term, is an indicator function, when and is 1 if they are not equal, otherwise 0; is a feature vector of a pixel with a serial number of in the first segmentation mask map;

[0135] The embodiment utilizes a conditional random field (CRF) to combine context information to further optimize the first segmentation mask map, enhance continuity and make the image boundary more accurate.

[0136] The ship original image is segmented using the second segmentation mask map to obtain ship region features and background region features , and the expression formula is:

[0137]

[0138]

[0139] In the formula, is a ship original image, is a second segmentation mask map of the ship.

[0140] The ship region features and the background region features are randomly combined to form a fusion feature map, specifically including:

[0141] The ship region features are randomly rotated at an angle to simulate the situation on the sea, randomly flipped to increase multi-view information, and randomly translated and scaled to simulate the far and near situation to obtain ship simulation features .

[0142] The ship simulation features and the background region features are randomly combined using an Alpha fusion strategy to form a fusion feature map , and the expression formula is:

[0143]

[0144] wherein, is a weight coefficient of feature fusion;

[0145] The fusion feature map is subjected to adversarial learning to obtain a training sample, specifically including:

[0146] Select a target ship image from the ship original image set; input the fusion feature map into the generator of the adversarial learning to obtain a target style image, and calculate the cycle consistency loss of the generated target style image and the target image by using the discriminator of the adversarial learning and the adversarial loss , and the expression formula is

[0147]

[0148]

[0149] In the formula, is the L1 norm; is the generator of the adversarial learning of the target domain converted into the source domain is the generator of the adversarial learning of the source domain converted into the target domain; is the fusion feature map; is the target ship image; is the mathematical expectation function; is the discriminator of the adversarial learning, represents the probability of belonging to , represents the probability that the image generated according to belongs to ,

[0150] According to the cycle consistency loss and the adversarial loss, the generator parameters are optimized, and the adversarial learning process is repeated until the cycle consistency loss and the adversarial loss converge to obtain the output image of the adversarial learning.

[0151] The training sample is generated from the output image of the adversarial learning, specifically including:

[0152] Select an anchor image from the ship original image set; divide the output image into positive samples and negative samples according to the anchor image, calculate the total loss of data expansion on the original image set according to the output image of the adversarial learning and the target anchor image, and use the cross-entropy loss , the triplet loss and the perceptual loss , and the expression formula is

[0153]

[0154] In the formula, is the total loss of data expansion on the original image set; is the hyperparameter of the cross-entropy loss , is the hyperparameter of the triplet loss , is the hyperparameter of the perceptual loss ​hyperparameters of the CMS-FCM model;

[0155] The performance of the CMS-FCM model in segmenting the ship image is evaluated by using a cross-entropy loss, and a dynamic weight adjustment is combined to balance its influence. A triplet loss is used to ensure the distinguishability of the ship and the background in the feature space, and an improved perceptual loss is introduced to measure the perceptual similarity between the output image and the original ship image, so as to ensure the visual quality of the generated image.

[0156] The parameters of the CMS-FCM model and the generator are optimized according to the total loss, and the iteration of the ship original image set expansion process is repeated until the total loss converges to output the ship expansion image, and the training sample is obtained.

[0157] The ship region features and the background region features are randomly combined to form a fusion feature map; and the training sample is obtained by performing adversarial learning on the fusion feature map.

[0158] The ship re-identification clustering model is trained using the training sample, and the specific process includes:

[0159] The ship re-identification clustering model includes a dynamic clustering module, a clustering intra-class reorganization module, and a clustering inter-class reorganization module.

[0160] The training sample is input into the dynamic clustering module, and the training sample is clustered by using the HDBSCAN algorithm. There will be two types of training samples in the clustering result: individual outlier training samples and training samples in cross-camera clusters. The individual outlier training samples are directly discarded; the training samples collected by the same camera are divided into sample clusters; the first pseudo-label identifying the camera identity is added to the sample cluster; the training samples in the sample cluster are grouped by using the cosine similarity to obtain a training sample group; the cosine similarity can ensure that the images captured by the same camera can be more accurately grouped, thereby reducing the confusion caused by similar appearance.

[0161] After adjusting the center positions of the training sample groups according to the Euclidean distances between the training sample groups, the training samples are re-divided to obtain first sample sub-clusters, and the Euclidean distances between two training sample groups are calculated.

[0162] If the Euclidean distance is less than a set minimum distance threshold , the Euclidean distance between the two training sample groups is pushed away by introducing a disturbance vector , and if the Euclidean distance is greater than a set maximum distance threshold , the Euclidean distance between the two training sample groups is pulled closer by introducing a contraction vector , and the expression formula is:

[0163]

[0164]

[0165] In the formula, is the sequence number of the training sample group, For the The center point of the training sample group;

[0166] The similarity between the training sample and the center point of the training sample group is recalculated, and training samples with similarities below a similarity threshold are removed. The training samples are then re-divided based on the similarity and the adjusted center point of the training sample group to obtain a first sample subcluster; a second pseudo-label identifying the ship instance is added to the first sample subcluster. This embodiment first uses the mean to obtain the cluster center point within each training sample group. The distances between the center points of different training sample groups are then compared, pushing those that are too close together further away and bringing those that are too far together closer. After multiple iterations of cluster center points, each training sample group is forced to have a uniform distribution, avoiding the problem of poor learning results caused by large differences in the number of samples in different training sample groups.

[0167] Calculating the reliability probability of the first sample sub-cluster; splitting and recombining the first sample sub-cluster according to the reliability probability to obtain a second sample sub-cluster; specifically including:

[0168] Calculate the probability that the training sample belongs to the first sample subcluster , the expression formula is:

[0169]

[0170]

[0171] In the formula, For serial number Training samples of Belongs to the serial number The first sample subcluster The probability of membership; The sequence number in the first sample subcluster is training samples; For serial number The first sample subcluster of For serial number The first sample subcluster The number of training samples in ; For serial number Training samples of With serial number The first sample subcluster similarity; For serial number The first sample subcluster The serial number is training samples; Indicates the serial number is Training samples of and serial number Training samples of Same number of features; Indicates the serial number is Training samples of and serial number Training samples of The total number of features with no duplication;

[0172] The probability that the training sample belongs to the first sample subcluster Take the average value to get the reliability probability of the first sample sub-cluster; set the reliability probability below the first threshold The first sample sub-cluster is split and reorganized, the contrast loss of the first sample sub-cluster is calculated by the contrastive learning algorithm, and the contrast loss is used to optimize the distance between training samples and then the second sample sub-cluster is obtained by re-dividing.

[0173] Inputting the second sample sub-clusters into the clustering inter-cluster recombination module, calculating the sharing probability between the second sample sub-clusters, and merging and recombining the second sample sub-clusters according to the sharing probability to obtain a third sample sub-cluster, specifically including:

[0174] The second sample subcluster The training samples in , calculate the training samples Belongs to the second sample subcluster The probability of membership ; When the subordinate probability Greater than the second threshold When the training sample The second sample subcluster With the second sample subcluster Similar training samples within

[0175] According to the second sample subcluster With the second sample subcluster The number of similar training samples in the second sample sub-cluster is calculated With the second sample subcluster The probability of sharing the same ship instance is expressed as:

[0176]

[0177] In the formula, For serial number The second sample subcluster With serial number The second sample subcluster the probability of sharing the same ship instance; For serial number second sample sub-cluster second sample sub-cluster second sample sub-cluster similar training samples within the second sample sub-cluster; second sample sub-cluster second sample sub-cluster training samples within the second sample sub-cluster; second sample sub-cluster second sample sub-cluster training samples within the second sample sub-cluster;

[0178] merging and reorganizing the second sample sub-clusters with a shared probability greater than a set shared probability threshold to obtain a third sample sub-cluster.

[0179] calculating a correlation contrast loss of the third sample sub-cluster according to the third sample sub-cluster, optimizing weight parameters of the ship re-identification clustering model according to the correlation contrast loss, repeating iteration to reach a training process of the ship re-identification clustering model, and outputting the trained ship re-identification clustering model when a set termination condition is reached; specifically including:

[0180] calculating a membership probability of the training sample belonging to the third sample sub-cluster ; calculating a correlation contrast loss and a difficulty index of the third sample sub-cluster clustering reorganization according to the membership probability , and the expression formula is:

[0181]

[0182]

[0183] In the formula, and are adjustment parameters of the correlation contrast loss ; the difficulty index of the third sample sub-cluster clustering reorganization is ; the number of training samples within the third sample sub-cluster is ; the number of training samples within the third sample sub-cluster is ; the number of training samples within the third sample sub-cluster is ; and the number of training samples within the third sample sub-cluster is training samples. The difficulty index is used to dynamically adjust the attention degree of the ship re-identification clustering model to the training samples; thereby improving the learning effect and robustness of the ship re-identification clustering model in contrast learning. This method ensures that the ship re-identification clustering model can reduce the influence when dealing with clustering samples with high similarity, while maintaining effective learning of dissimilar clusters. The most difficult sample sub-cluster of the instance is selected as the representative to participate in contrast learning, which ensures that the ship re-identification clustering model can focus on the most challenging negative samples when reorganizing between clusters, thereby improving the learning effect.

[0184] The ship re-identification clustering model parameters are optimized using the associated contrast loss function; when the silhouette coefficient of the model clustering result reaches the set silhouette threshold value in the last 5 iterations or the Calinski-Harabasz index of the model clustering result reaches the set index threshold value in the last 5 iterations , the training is ended, and the final clustering result of the model is obtained, that is, the ship re-identification result, and the trained ship re-identification clustering model is output.

[0185] The embodiment proposes a new unsupervised ship re-identification method based on intra-class and inter-class reorganization, effectively reducing the negative impact of label noise and image style on unsupervised ship re-identification effect, improving the accuracy of label identification, and enhancing the generalization ability and robustness of the ship re-identification clustering model. At the same time, the accuracy of ship re-identification is improved.

[0186] Those skilled in the art will appreciate that embodiments of the present application can be provided as methods, systems, or computer program products. Therefore, the present application can take the form of an entirely hardware embodiment, an entirely software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present application can take the form of a computer program product implemented on one or more computer-usable storage media (including, but not limited to, disk storage, CD-ROMs, optical storage, etc.) containing computer-usable program code.

[0187] The present application is described with reference to flowcharts and / or block diagrams according to the methods, devices (systems), and computer program products of the embodiments of the present application. It should be understood that each flow and / or block in the flowcharts and / or block diagrams, as well as the combination of flows and / or blocks in the flowcharts and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing devices produce a device that implements the functions specified in the flow Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.

[0188] These computer program instructions can also be stored in a computer readable memory that can direct a computer or other programmable data processing apparatus to function in a particular manner, such that the instructions stored in the computer readable memory produce an article of manufacture including instructions which implement the flow Figure 1 flow or flows and / or blocks Figure 1 of the functions specified in the flow or flows and / or blocks in the block or blocks.

[0189] These computer program instructions can also be loaded onto a computer or other programmable data processing apparatus to cause a series of operational steps to be performed on the computer or other programmable apparatus to produce a computer implemented process such that the instructions which execute on the computer or other programmable apparatus provide steps for implementing the flow Figure 1 flow or flows and / or blocks Figure 1 of the functions specified in the flow or flows and / or blocks in the block or blocks.

[0190] The above description is only preferred embodiments of the present application, it should be pointed out that for those skilled in the art, without departing from the technical principles of the present application, a number of improvements and modifications can be made, these improvements and modifications should also be considered as the protection scope of the present application.

Claims

1. A ship re-identification method based on sample expansion, characterized in that: include: Obtain ship monitoring data within the specified sea area and input the ship monitoring data into the preset ship re-identification clustering model to obtain the unsupervised ship recognition results; The training process of the ship re-identification clustering model includes: Obtain the original ship image set and perform data expansion to obtain training samples; The ship re-identification clustering model is trained using the training samples, wherein the ship re-identification clustering module includes a dynamic clustering module, an intra-clustering reorganization module, and an inter-clustering reorganization module; The training samples are input into the dynamic clustering module and clustered using the HDBSCAN algorithm. The training samples collected by the same camera are divided into sample clusters. The training samples within the sample clusters are grouped using cosine similarity to obtain training sample groups. After adjusting the center position of the training sample group according to the Euclidean distance between the training sample groups, the training samples are re-divided to obtain the first sample sub-cluster. Inputting the first sample subcluster into the intra-cluster recombination module to calculate the reliability probability of the first sample subcluster; splitting and recombining the first sample subcluster according to the reliability probability to obtain a second sample subcluster; Inputting the second sample sub-clusters into the clustering inter-cluster recombination module, calculating the sharing probability between the second sample sub-clusters, and merging and recombining the second sample sub-clusters according to the sharing probability to obtain a third sample sub-cluster; Based on the obtained third sample subcluster, the association contrast loss of the third sample subcluster is calculated, and the weight parameters of the ship re-identification clustering model are optimized according to the association contrast loss. The training process of the ship re-identification clustering model is repeated iteratively until the set termination condition is met. The final clustering result, i.e., the ship re-identification result, is obtained, and the trained ship re-identification clustering model is output.

2. The ship re-identification method based on sample expansion according to claim 1, characterized in that: The original ship image set is expanded to obtain training samples, including: The self-attention mechanism in the DINO-ViT model is used to capture the global and local features of the original ship image, and the global and local features in the original ship image are fused to obtain the comprehensive image features. The ship mask is obtained by performing multi-scale fuzzy clustering on the comprehensive image features using the CMS-FCM model; the ship mask is smoothed and divided into ship area features and background area features; The ship area features and background area features are randomly combined to form a fusion feature map; adversarial learning is performed on the fusion feature map to obtain training samples.

3. The ship re-identification method based on sample expansion according to claim 2, characterized in that: The ship mask is obtained by performing multi-scale fuzzy clustering on the comprehensive image features through the CMS-FCM model, including: The image fuzzy clustering is obtained by performing multi-scale fuzzy clustering on the comprehensive image features through the fuzzy clustering algorithm. ; Fuzzy clustering of images Add ship's center of gravity constraint and buoyancy center constraint, and perform fuzzy clustering based on the image Get the ship mask of comprehensive image features , the expression formula is: ; In the formula, is the number of clusters of image fuzzy clustering, r is the serial number of the cluster of image fuzzy clustering; s is the comprehensive image feature vector, Fuzzy clustering for images The cluster center of The weight set for the control center of gravity term, The weight set for controlling the center of buoyancy; is the center of gravity of the ship in the comprehensive image features; is the buoyancy center of the ship in the comprehensive image features.

4. The ship re-identification method based on sample expansion according to claim 3 is characterized in that: The ship mask is smoothed and divided into ship area features and background area features; specifically including: The first segmentation mask is obtained by adding Gaussian noise to the ship mask and smoothing the edge of the ship mask. , the expression formula is; ; In the formula, To perform fuzzy clustering based on images Generate a ship mask that integrates image features, is Gaussian noise; After removing isolated pixels in the first segmentation mask through opening operation, the first segmentation mask is further optimized using conditional random field to obtain the second segmentation mask. , the expression formula is: ; In the formula, The sequence number in the first segmentation mask is pixels; is a set of pixels in the first segmentation mask image; Indicates that the sequence number in the first segmentation mask image is The segmentation labels of the pixels and the sequence number in the first segmentation mask is The feature vector of the pixel The similarity of The sequence number in the first segmentation mask is The segmentation labels of the pixels, The sequence number in the first segmentation mask is The feature vector of the pixel; The sequence number in the first segmentation mask is pixels; Indicates that the sequence number in the first segmentation mask image is The segmentation labels of the pixels The sequence number in the first segmentation mask is The segmentation labels of the pixels The associated features of The sequence number in the first segmentation mask is The segmentation labels of the pixels; Calculate the sequence number in the first segmentation mask image The segmentation labels of the pixels and the sequence number in the first segmentation mask is The feature vector of the pixel Similarity , the expression formula is: ; In the formula, Indicates that the sequence number in the first segmentation mask image is The segmentation labels of the pixels The characteristic variance of is the L1 norm; Indicates that the sequence number in the first segmentation mask image is The segmentation labels of the pixels The characteristic mean of Calculate the sequence number in the first segmentation mask image The segmentation labels of the pixels The sequence number in the first segmentation mask is The segmentation labels of the pixels The associated features , the expression formula is: ; in, is the weight of the smoothing term, The sequence number in the first segmentation mask is The feature vector of the pixel, is the indicator function, when and The indicator function takes the value 1 when they are not equal, otherwise it is 0; Use the second segmentation mask to segment the original ship image to obtain the ship area features and background area features , the expression formula is: ; ; In the formula, is the original image of the ship; is the second segmentation mask image.

5. The ship re-identification method based on sample expansion according to claim 4 is characterized in that: The ship area features and background area features are randomly combined to form a fusion feature map, specifically including: The ship area features are rotated at random angles to simulate the situation at sea, randomly flipped to add multi-view information, and randomly translated and scaled to simulate the distance to obtain ship simulation features. ; Use Alpha fusion strategy to simulate ship features Randomly combine with background area features to form a fusion feature map , the expression formula is: ; in, is the weight coefficient of feature fusion.

6. The ship re-identification method based on sample expansion according to claim 2, characterized in that: Perform adversarial learning on the fused feature map to obtain training samples, including: The target ship image is selected from the original ship image set; the fused feature map is input into the adversarial learning generator to obtain the target style image, and the discriminator of adversarial learning is used to calculate the cycle consistency loss of the generated target style image and the target image. and combat losses , the expression formula is: ; ; In the formula, Generator for adversarial learning of target domain to source domain Generator for adversarial learning to convert source domain into target domain; is the fusion feature map; is the L1 norm; is the target ship image; is the mathematical expectation function; For adversarial learning discriminator, Indicates belonging The probability of Indicates based on Generated image belong probability; The generator parameters are optimized based on the cycle consistency loss and the adversarial loss, and the adversarial learning process is repeated until the cycle consistency loss and the adversarial loss converge to obtain the adversarial learning output image; training samples are generated from the adversarial learning output image; Anchor images are selected from the original image set of the ship; the output image is divided into positive samples and negative samples according to the anchor image, and the total loss of data expansion of the original image set is calculated based on the output image of the adversarial learning and the target anchor image. The cross entropy loss is used , triplet loss and perceptual loss , the expression formula is: ; In the formula, is the total loss of data expansion on the original image set; is the cross entropy loss The hyperparameters of is the triplet loss The hyperparameters of Perceptual loss Hyperparameters of The parameters of the CMS-FCM model and generator are optimized according to the total loss, and the process of expanding the original ship image set is repeated until the total loss converges and the expanded ship image is output to obtain the training samples.

7. The ship re-identification method based on sample expansion according to claim 1, characterized in that: After adjusting the center position of the training sample group according to the Euclidean distance between the training sample groups, the training sample group is re-divided to obtain the first sample sub-cluster, specifically including: If the Euclidean distance between two training sample groups is less than the preset minimum distance threshold , then by introducing the perturbation vector Push the Euclidean distance between the two training sample groups to a greater value if the Euclidean distance between the two training sample groups is greater than the preset maximum distance threshold. , then by introducing the contraction vector To shorten the Euclidean distance between two training sample groups, the expression formula is: ; ; In the formula, is the sequence number of the training sample group, For the The center point of the training sample group; Recalculate the similarity between the training sample and the center point of the training sample group, remove the training samples whose similarity is lower than the similarity threshold, and re-divide the training samples according to the similarity and the adjusted center point of the training sample group to obtain the first sample sub-cluster.

8. The ship re-identification method based on sample expansion according to claim 1, characterized in that: Calculating the reliability probability of the first sample sub-cluster; splitting and recombining the first sample sub-cluster according to the reliability probability to obtain a second sample sub-cluster; specifically including: Calculate the probability that the training sample belongs to the first sample subcluster , the expression formula is: ; ; In the formula, For serial number Training samples of Belongs to the serial number The first sample subcluster The probability of membership; The sequence number in the first sample subcluster is training samples; For serial number The first sample subcluster of For serial number The first sample subcluster The number of training samples in ; For serial number Training samples of With serial number The first sample subcluster similarity; For serial number The first sample subcluster The serial number is training samples; Indicates the serial number is Training samples of and serial number Training samples of Same number of features; Indicates the serial number is Training samples of and serial number Training samples of The total number of features with no duplication; The probability that the training sample belongs to the first sample subcluster Take the average value to get the reliability probability of the first sample sub-cluster; set the reliability probability below the first threshold The first sample sub-cluster is split and reorganized, the contrast loss of the first sample sub-cluster is calculated by the contrastive learning algorithm, and the contrast loss is used to optimize the distance between training samples and then the second sample sub-cluster is obtained by re-dividing.

9. The ship re-identification method based on sample expansion according to claim 1, characterized in that: Inputting the second sample sub-clusters into the clustering inter-cluster recombination module, calculating the sharing probability between the second sample sub-clusters, and merging and recombining the second sample sub-clusters according to the sharing probability to obtain a third sample sub-cluster, specifically including: The second sample subcluster The training samples in , calculate the training samples Belongs to the second sample subcluster The probability of membership ; When the subordinate probability Greater than the second threshold When the training sample The second sample subcluster With the second sample subcluster Similar training samples within According to the second sample subcluster With the second sample subcluster The number of similar training samples in the second sample sub-cluster is calculated With the second sample subcluster The probability of sharing the same ship instance is expressed as: ; In the formula, For serial number The second sample subcluster With serial number The second sample subcluster the probability of sharing the same ship instance; For serial number The second sample subcluster With serial number The second sample subcluster The number of similar training samples within ; For serial number The second sample subcluster The number of internal training samples; For serial number The second sample subcluster The number of internal training samples; The sharing probability is greater than the set sharing probability threshold The second sample sub-clusters are merged and reorganized to obtain a third sample sub-cluster.

10. The ship re-identification method based on sample expansion according to claim 1, characterized in that: Calculating the association contrast loss of the third sample subcluster based on the third sample subcluster, optimizing the weight parameters of the ship re-identification clustering model based on the association contrast loss, and repeatedly iterating to achieve the training process of the ship re-identification clustering model until the set termination condition is met and the trained ship re-identification clustering model is output; specifically, including: Calculate the probability that the training sample belongs to the third sample subcluster ; According to the subordinate probability Calculate the association contrast loss And the difficulty index of clustering and reorganization of the third sample sub-cluster is expressed as: ; ; In the formula, and is the association contrast loss The adjustment parameters of For serial number The third sample subcluster of Difficulty index of cluster recombination; For serial number The third sample subcluster of The number of internal training samples; For serial number The third sample subcluster of The serial number is training samples; The weight parameters of the ship re-identification clustering model are optimized based on the association contrast loss, and the difficulty index is used to dynamically adjust the degree of attention the ship re-identification clustering model pays to the training samples; the association contrast loss function is used to optimize the parameters of the ship re-identification clustering model; Repeat the iteration to train the ship re-identification clustering model. When the silhouette coefficient of the clustering model reaches the set silhouette threshold in 5 consecutive iterations, Or the Calinski-Harabasz index of the model clustering result reaches the set index threshold in 5 consecutive iterations When , the training is ended and the final clustering result of the model is obtained, that is, the result of ship re-identification, and the trained ship re-identification clustering model is output.