Semi-supervised long-term ship re-identification method based on correlation forgetting learning
Through a semi-supervised approach of associative forgetting learning and semantic embedding, the problems of poor generalization ability and noise influence in ship re-identification are solved, high robustness and reliability in long-term recognition are achieved, and feature matching is dynamically adjusted to adapt to appearance changes.
Patent Information
- Application Number
- CN202510953897.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-11
- Publication Date
- 2025-09-23
AI Technical Summary
Existing ship re-identification methods have poor generalization ability in long-term recognition and are difficult to adapt to changing appearance changes. In addition, the model easily loses old knowledge, resulting in reduced domain adaptability and robustness, and single modal data is easily affected by noise.
A semi-supervised method based on associative forgetting learning is adopted. The memory process is simulated through multidimensional feature clustering and associative learning. The visual and text features are fused with the semantic embedding module. The forgetting learning is used to reduce the attention to irrelevant features. The old knowledge is retained through retrospective learning, and the feature matching strategy is dynamically adjusted.
The model's domain adaptability and generalization capabilities in long-term ship re-identification are improved, the robustness and reliability of the model are enhanced, the noise problem in visual recognition is solved, and the recognition performance is improved.
Smart Images

Figure CN120689826A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the technical field of ship re-identification, and in particular to a semi-supervised long-term ship re-identification method based on associative forgetting learning. Background Art
[0002] Ship re-identification (Ship Re-ID) is of great significance in maritime transportation and port management, especially in the context of growing global trade, the role of ships in international logistics has become increasingly prominent. Accurately identifying and tracking ships not only helps to improve maritime safety and prevent illegal activities, but also optimizes port resource utilization and improves shipping efficiency. However, traditional ship identification methods often rely on single image matching and cannot cope with changes in ships at different times, perspectives and environmental conditions. In the context of long-term operation, ships are likely to change in appearance due to cargo changes and equipment replacements. Therefore, the demand for long-term ship re-identification (LTS-ReID) is becoming increasingly apparent. Long-term ship re-identification requires the ability to accurately identify and track ships at multiple time points and in complex environments, greatly enhancing the level of intelligent ship management.
[0003] Most existing ship re-identification methods focus on short-term ship re-identification, resulting in poor model generalization and difficulty adapting to changing real-world conditions. Compared to short-term ship re-identification, long-term ship re-identification places higher demands on feature learning and must cope with larger differences in appearance. Furthermore, the multi-perspective problem encountered in short-term ship re-identification is further amplified in long-term ship re-identification. Furthermore, ship re-identification presents a significant problem: once the model adapts to a new domain, it easily loses the knowledge acquired in previously observed domains, which leads to a decrease in the model's domain adaptability and generalization capabilities. Furthermore, most existing ship re-identification methods process data solely at the visual level, and single-modal data is easily affected by noise or omissions, resulting in a decrease in the robustness and reliability of the model.
[0004] It should be noted that the information disclosed in the above background technology section is only used to enhance the understanding of the background of this application, and therefore may include information that does not constitute prior art known to ordinary technicians in this field. Summary of the Invention
[0005] In order to provide a basic understanding of some aspects of the disclosed embodiments, a brief summary is given below. The summary is not an extensive review, nor is it intended to identify key / critical elements or delineate the scope of protection of these embodiments, but rather serves as a prelude to the detailed description that follows.
[0006] The disclosed embodiments provide a semi-supervised long-term ship re-identification method based on associative forgetting learning to improve the robustness and reliability of the model.
[0007] In some embodiments, the method includes: collecting and clustering the multidimensional features of each ship to generate a feature cluster for each ship; performing feature matching on the multidimensional features of the ship image with the features in the feature cluster and calculating the similarity, and outputting the ship identification result based on the similarity; performing association learning on the multidimensional features of the ship image and the features of the feature cluster; performing forgetting learning on the features of the feature cluster based on the frequency of occurrence of the features and the feature cluster; and updating the feature cluster through association learning and forgetting learning during each ship identification training.
[0008] In some embodiments, collecting multi-dimensional features of each vessel includes: collecting tag information and visual features of each vessel; wherein the tag information includes vessel type, color, angle, and environment;
[0009] The label information is rewritten and expanded to generate high-level semantic information for each ship; the high-level semantic information is converted into text features and fused with visual features, and the fused features are converted into feature sequences; the self-attention mechanism is used to calculate the attention weight relationship between visual features and text features based on the feature sequence; the visual features and text features are fused according to the attention weight relationship to generate multi-dimensional features for each ship.
[0010] In some embodiments, a self-attention mechanism is used to calculate the attention weight relationship between visual features and text features based on a feature sequence, including: calculating a query matrix Q, a key matrix K, and a value matrix V based on the feature sequence; wherein the query matrix Q comes from the visual feature projection, and the key matrix K and the value matrix V come from the text features; using the self-attention mechanism to calculate the attention score of the visual feature to the semantic feature; and using the Softmax function to convert the attention score into the attention weight of the visual feature to the semantic feature.
[0011] In some embodiments, all features are converted into Hamming code representations; the multidimensional features of the ship image and the features of the feature clustering cluster are associated with learning, including: calculating the Hamming code distance between the multidimensional features of the ship image and the features of the feature clustering cluster to obtain the first σ1 closest feature samples; randomly selecting σ2 random feature samples from the feature clustering cluster; clustering the multidimensional features of the ship, σ1 closest feature samples and σ2 random feature samples, and using cosine similarity for metric learning; adjusting the features in the feature clustering cluster according to the cosine similarity.
[0012] In some embodiments, adjusting features in a feature cluster according to cosine similarity includes:
[0013] When the cosine similarity of the multidimensional features of the ship image and the features in the feature cluster is higher than the first threshold θ1, the multidimensional features of the ship image and the features in the feature cluster are weighted averaged and the features in the feature cluster are adjusted; when the cosine similarity of the multidimensional features of the ship image and the features in the feature cluster are both lower than the first threshold θ1, a new feature cluster is generated based on the multidimensional features of the ship image; the update loss function L is calculated based on the feature clusters before and after the feature update. update , when the loss function is greater than or equal to the second threshold θ2, the update is determined, and when the loss function is less than the second threshold θ2, the update is canceled.
[0014] In some embodiments, a new feature cluster is generated based on the multidimensional features of the ship image, including: determining the feature cluster with the highest similarity to the new feature cluster based on the clustering result of the sample; assigning the label of the feature cluster with the highest similarity to the new feature cluster to generate a pseudo label of the new feature cluster; calculating the adaptive learning loss function based on the new feature cluster and the feature cluster with the highest similarity; in the adaptive learning loss function L adaptive When it is greater than the third threshold θ3, the features of the new feature cluster are aligned with the features of the feature cluster with the highest similarity, and the new feature cluster after feature alignment is aligned with the feature cluster with the highest similarity.
[0015] In some embodiments, forgetting learning is performed on the features of the feature clusters according to the frequencies of occurrence of the features and the feature clusters, including: updating the frequencies of occurrence of the features and the feature clusters each time the features are matched; constructing a forgetting loss function L forget When the difference between the occurrence frequency of a feature and the occurrence frequency of the corresponding feature cluster is less than the fourth threshold θ4, the feature is forgotten; the occurrence frequency of the corresponding feature cluster and the occurrence frequency of all its features are updated to the frequency of the least used feature in the cluster.
[0016] In some embodiments, the feature clusters are updated by association learning and forgetting learning each time a ship is identified, including: updating the loss function L update , adaptive learning loss function L adaptive And the forgetting loss function L forget Perform weighted fusion to generate a comprehensive loss function L all ; Using the comprehensive loss function L all Balanced update loss function L update , adaptive learning loss function L adaptive And the forgetting loss function L forget Updates.
[0017] In some embodiments, after updating the feature clusters by association learning and forgetting learning during each ship recognition training, the method further includes: clustering the updated features using the previous clustering parameters and the current clustering parameters after each clustering is completed; calculating the image-to-prototype similarity consistency loss L according to the clustering results. image1 and image-to-image similarity consistency loss L image2 ; Calculate the gap loss function L between the updated sample and the sample of the original feature cluster reg ; Consistency loss L based on image-to-prototype similarity image1 , image-to-image similarity consistency loss L image2 and the gap loss function L reg Adjust the current clustering parameters.
[0018] In some embodiments, the method further comprises: performing a similarity consistency loss L on the image to the prototype. image1 , image-to-image similarity consistency loss L image2 and the gap loss function L reg Perform weighted fusion to generate the total loss function L full ; Using the total loss function L full Balanced image-to-prototype similarity consistency loss L image1 , image-to-image similarity consistency loss L image2 and the gap loss function L reg Updates.
[0019] The semi-supervised long-term ship re-identification method based on associative forgetting learning provided by the embodiments of the present disclosure can achieve the following technical effects:
[0020] 1. After extracting the multi-dimensional features of a ship, association learning is used to simulate the human memory association process to identify the same ship even though it has been modified by external equipment or cargo. Forgetting learning is used to reduce attention to features irrelevant to the identity, thereby improving the model's domain adaptability and generalization capabilities in long-term ship re-identification, and enhancing the model's robustness and reliability.
[0021] 2. The semantic embedding auxiliary module integrates the visual features of the image with the textual features of the image semantics, using semantic information to provide additional context and category relevance, addressing noise issues such as ship perspectives and the environment, and effectively improving visual recognition performance.
[0022] 3. Using the idea of "review", the clustering results of the new and old models are compared, and the old model is used to guide the update of the clustering parameters of the new model, so as to solve the problem of the model's reduced domain adaptability caused by discarding old knowledge after learning new knowledge, that is, the problem of the model's poor domain adaptability and generalization ability.
[0023] The above general description and the following description are exemplary and explanatory only and are not intended to limit the present application. BRIEF DESCRIPTION OF THE DRAWINGS
[0024] Figure 1 is a schematic diagram of a long-term ship re-identification method based on associative forgetting learning provided by an embodiment of the present disclosure;
[0025] Figure 2 2 is a schematic diagram of the structure of a long-term ship re-identification task based on associative forgetting learning provided by an embodiment of the present disclosure;
[0026] Figure 3 This is a flowchart of semi-supervised long-term ship re-identification based on associative forgetting learning provided by an embodiment of the present disclosure;
[0027] Figure 4 This is a multi-branch classification framework diagram of the pre-trained ResNet-50 provided in an embodiment of the present disclosure;
[0028] Figure 5 is a schematic diagram of a semantic embedding auxiliary module provided by an embodiment of the present disclosure;
[0029] Figure 6 is a schematic diagram of an associative forgetting learning module provided by an embodiment of the present disclosure;
[0030] Figure 7 It is a schematic diagram of a review learning diagram provided by an embodiment of the present disclosure. DETAILED DESCRIPTION
[0031] In order to make the purpose, technical solutions and advantages of this application more clear, the following further describes this application in detail with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain this application and are not intended to limit this application.
[0032] The present application provides a semi-supervised long-term ship re-identification model based on association and forgetting learning, which includes a semantic embedding auxiliary module, an association and forgetting learning module, and a review learning module.
[0033] Reference Figure 1 、 Figure 2 and Figure 3 Based on the above model, the embodiment of the present disclosure provides a semi-supervised long-term ship re-identification method based on associative forgetting learning, such as Figure 2 Shown, including:
[0034] S101, collecting and clustering the multidimensional features of each ship to generate a feature cluster for each ship;
[0035] S201, matching the multi-dimensional features of the ship image with the features in the feature cluster and calculating the similarity, and outputting the ship recognition result according to the similarity;
[0036] S301, performing association learning between the multi-dimensional features of the ship image and the features of the feature cluster;
[0037] S401, performing forgetting learning on the features of the feature clusters according to the frequencies of occurrence of the features and the feature clusters;
[0038] S501, updating the feature clusters through association learning and forgetting learning during each ship recognition training.
[0039] After extracting the multi-dimensional features of the ship, the human memory association process is simulated through association learning to identify the same ship that has changed due to external equipment or cargo. Forgetting learning is used to reduce attention to features unrelated to the identity, thereby improving the model's adaptability and generalization capabilities in long-term ship re-identification, and enhancing the model's robustness and reliability.
[0040] In some embodiments, collecting multi-dimensional features of each vessel includes: collecting tag information and visual features of each vessel; wherein the tag information includes vessel type, color, angle, and environment;
[0041] The label information is rewritten and expanded to generate high-level semantic information for each ship; the high-level semantic information is converted into text features and fused with visual features, and the fused features are converted into feature sequences; the self-attention mechanism is used to calculate the attention weight relationship between visual features and text features based on the feature sequence; the visual features and text features are fused according to the attention weight relationship to generate multi-dimensional features for each ship.
[0042] In some embodiments, a self-attention mechanism is used to calculate the attention weight relationship between visual features and text features based on the feature sequence, including:
[0043] Calculate the query matrix Q, key matrix K and value matrix V according to the feature sequence; where the query matrix Q comes from the visual feature projection, and the key matrix K and value matrix V come from the text features;
[0044] Use the self-attention mechanism to calculate the attention score of visual features to semantic features;
[0045] The Softmax function is used to convert the attention score into the attention weight of the visual feature to the semantic feature.
[0046] Specifically, refer to Figure 2 、 Figure 3 and Figure 4, using semantic embedding auxiliary modules to solve the problem of missing detailed features and noise influence of a single modality, and provide more refined feature expression for subsequent feature learning. First, the different attributes of the ship are regarded as independent classification tasks, and a multi-task learning framework is constructed; wherein the attributes include color, type, angle and environment. Multi-task learning learns the features of multiple related tasks by sharing network layers, and finally outputs their respective results through different branches to obtain the semantic information of the image and output the semantic information in the form of text. Here, the present invention selects ResNet-50 as the backbone of the pre-training model to extract the attribute feature representation F of the ship image shared Training method: Multi-task learning training is performed using a labeled ship dataset, with ResNet-50 as the backbone network to extract ship features. Multiple branch outputs are then constructed, with each branch sharing network layer data. The ship dataset includes ship type, color, angle, and environment labels, and the multiple branch outputs include ship type, color, angle, and environment classifications.
[0047] In order to measure the difference between the model prediction results and the actual labels, the cross entropy function L is used. CE As a loss function:
[0048]
[0049] Among them, N1 is the number of real labels, m is the label sequence number, and y m is the true label of sequence number m, For the prediction of y m probability.
[0050] In order to train the multi-task learning model, a multi-branch loss function L is set total , to optimize the output of multiple tasks simultaneously:
[0051]
[0052] Among them, N2 is the number of model branches, a is the branch number, ω a The weight corresponding to the branch numbered a is represented. This paper selects the four most representative branches: ship type, color, shooting angle, and environment. This paper uses a validation set to verify the model training results. Training is completed when the validation set accuracy does not improve for five consecutive epochs.
[0053] The label of the ship and the specific environmental description of the ship are input into DeepSeek-v1, and DeepSeek-v1 is allowed to imitate the extended information of the label according to the example, and set a regular constraint to limit the output format and number of words, and at the same time add abnormal labels and abnormal items to improve the robustness of the model to avoid inconsistency between the output of the model and the set format, and then perform batch processing optimization to obtain a pre-trained DeepSeek-v1. The present invention uses the pre-trained DeepSeek-v1 model to rewrite and expand the low-level semantic information obtained in the first stage, that is, the ship label information obtained in step one (such as container ship, red, left front side, sea, daytime), to generate more expressive high-level semantic information (for example, this is a red container ship sailing on the sea during the day, and the picture shows the left front side of the ship).
[0054] DeepSeek-v1's capabilities in natural language generation and information abstraction help expand the semantic representation of ship images, enabling the model to capture more details. DeepSeek-v1 is also superior to other mainstream language models in processing Chinese-like language information.
[0055] To more efficiently extract image features and fuse multimodal information, this paper uses PVTv2 (PyramidVision Transformer V2) as its core backbone model. PVTv2 combines the Transformer's self-attention mechanism with the multi-scale nature of its pyramid structure, enabling it to capture detailed visual feature representations at different scales. This paper inputs ship images into PVTv2 to obtain their feature representations.
[0056] Since Transformer does not depend on the specific type of input data, it can convert visual information and text information into a unified token representation and process the differences between different modal data (visual features and text features). The present invention converts text information into text features through Transformer and combines it with the visual features obtained in the previous step to obtain a comprehensive image feature representation.
[0057] The self-attention mechanism processes the integrated features. Based on the resulting sequence, the query matrix Q, key matrix K, and value matrix V are calculated. Q is derived from the visual feature projection, while K and V are derived from the text features. The self-attention mechanism can learn the complex dependencies between image and semantic features, further improving ship recognition.
[0058] Because multimodal fusion can lead to the loss of some detailed features, token upsampling is used to restore this detail, significantly improving the recognition accuracy of small components such as masts and portholes. This involves converting feature maps into higher-resolution representations. Through a reverse T2T (Tokens to Token) conversion, each token is expanded into multiple sub-tokens, while global features are also converted into tokens. This allows for the recovery of more detailed information and mitigates the impact of noise such as illumination variations. Visual and semantic textual features are fused according to the attention weight matrix A to generate fine-grained ship image features.
[0059] The first stage mentioned above is completed by the semantic embedding auxiliary module, which is used to solve the problems of missing detailed features and noise influence of a single modality, and provide more refined feature expression for subsequent feature learning.
[0060] In some embodiments, all features are converted into Hamming code representations; the multidimensional features of the ship image and the features of the feature clustering cluster are associated with learning, including: calculating the Hamming code distance between the multidimensional features of the ship image and the features of the feature clustering cluster to obtain the first σ1 closest feature samples; randomly selecting σ2 random feature samples from the feature clustering cluster; clustering the multidimensional features of the ship, σ1 closest feature samples and σ2 random feature samples, and using cosine similarity for metric learning; adjusting the features in the feature clustering cluster according to the cosine similarity.
[0061] Specifically, the multi-dimensional features of the ship image are defined as new features.
[0062] Reference Figure 6 ,First, the corresponding Hamming code is generated based on the ,feature expression obtained in the first stage, ,to facilitate the subsequent rapid association learning.
[0063] The HDBSCAN (Hierarchical Density-Based Spatial Clustering of Applications with Noise) clustering algorithm is used to cluster the extracted ship image features (fine-grained multimodal feature representation F att ) for clustering. The HDBSCAN algorithm can automatically identify and remove noise while generating an appropriate number of clusters. The present invention groups and stores the characteristic representations of each ship according to the clustering results, with each cluster representing a ship identity.
[0064] Based on the clustering results, a long-term memory module was designed. Each cluster represents a ship feature storage module, storing the ship's multidimensional features (such as shape, color, type, angle, etc.). During initialization, each cluster records the initial feature representation of the ship. The long-term memory module stores each ship feature in the long-term memory area, along with its cluster ID and frequency of occurrence, for subsequent associative learning and forgetting mechanisms.
[0065] M i ={H i ,C i ,F i ,R i ...}
[0066] Among them, M i is the identity of the ship with serial number i in the long-term memory area, i is the serial number of the ship identity, H i is the Hamming code of the ship identity with serial number i, C i is the cluster with ship identity number i, F i is the fine-grained ship image feature of the ship identity with sequence number i, R i is the frequency of occurrence of the ship identity with serial number i.
[0067] For each training sample of ship images, the Hamming code distance between its features and those stored in the memory module is calculated. The first σ1 closest feature samples and an additional σ2 random feature samples are selected for clustering. Using Hamming codes can speed up retrieval, while using partial samples can improve model efficiency. To avoid overfitting, the present invention also adds random samples. The new and selected samples are then clustered, and cosine similarity is used for metric learning. The cosine similarity metric formula is as follows:
[0068]
[0069] Among them, F j is the fine-grained ship image feature of the ship identity with sequence number j, and |.| is the modulo operation.
[0070] In some embodiments, adjusting features in a feature cluster according to cosine similarity includes: when the cosine similarity between the multidimensional features of the ship image and the features in the feature cluster is higher than a first threshold value θ1, adjusting the features in the feature cluster after weighted averaging the multidimensional features of the ship image and the features in the feature cluster; when the cosine similarity between the multidimensional features of the ship image and the features in the feature cluster is lower than the first threshold value θ1, generating a new feature cluster according to the multidimensional features of the ship image; and calculating an updated loss function L according to the feature cluster before and after the feature update. update , when the loss function is greater than or equal to the second threshold θ2, the update is determined, and when the loss function is less than the second threshold θ2, the update is canceled.
[0071] Specifically, refer to Figure 6 If the similarity between the fine-grained ship image features of the ship image training sample and a feature in the long-term memory area exceeds the preset first similarity threshold θ1, the feature representation is updated using a weighted average method. The weighting coefficient α is adjusted according to the similarity. A higher similarity will give the new feature a greater weight when updating. Here, taking the ship identity with serial number i as an example, the formula is as follows:
[0072]
[0073] in, is the memory feature of the ship identity with serial number g in the updated long-term memory area, g is the serial number of the memory feature, α is the weighting coefficient, is the old memory feature of the ship identity with sequence number g in the long-term memory area, F i is the fine-grained ship image feature of the ship identity with sequence number i, is the memory feature of the ship identity with sequence number g in the long-term memory area The corresponding new features.
[0074] If the similarity between the fine-grained ship image features of the training sample of the ship image and all the features in the long-term memory area is lower than the first threshold θ1, the fine-grained ship image features of the ship image are regarded as new instances and added to the long-term memory module to form a new memory cluster.
[0075] In order to avoid excessive memory updates, which may cause the model to pay too much attention to features such as external devices, the present invention introduces an update loss function L update To limit the amplitude of feature updates and increase the distance between features before and after the update, the model can avoid unnecessary changes and maintain stability. Here, taking the identity of the ship with serial number i as an example, the formula is as follows:
[0076]
[0077] Among them, N f is the number of memory features, g is the sequence number of the memory feature, ||.|| is the L2 norm, is the old memory feature of the ship identity with sequence number g in the long-term memory area, R′ g is the frequency of occurrence of the memory feature with sequence number g in the ship identity with sequence number i in the long-term memory area, R iis the frequency of occurrence of ship identity number i, and is the maximum frequency of occurrence of features in the long-term memory for ship identity number i, where i is the ship identity number. If the cluster corresponding to the ship identity is stable, the denominator will be large, resulting in a smaller value. If the cluster is unstable, the denominator will be small, resulting in a larger value. A second threshold θ2 is set. If the updated loss function is greater than or equal to the second threshold θ2, the update is performed; if the updated loss function is less than the second threshold θ2, the update is not performed.
[0078] In some embodiments, a new feature cluster is generated based on the multidimensional features of the ship image, including: determining the feature cluster with the highest similarity to the new feature cluster based on the clustering result of the sample; assigning the label of the feature cluster with the highest similarity to the new feature cluster to generate a pseudo label of the new feature cluster; calculating the adaptive learning loss function based on the new feature cluster and the feature cluster with the highest similarity; in the adaptive learning loss function L adaptive When it is greater than the third threshold θ3, the features of the new feature cluster are aligned with the features of the feature cluster with the highest similarity, and the new feature cluster after feature alignment is aligned with the feature cluster with the highest similarity.
[0079] The clustering results are used to assign pseudo labels to the unlabeled data, i.e., the training samples of new ship images. Pseudo labels can be generated by finding the most similar cluster in the long-term memory module and assigning the label of the cluster to the training sample. In order to avoid the negative impact of pseudo labels on the model, the present invention introduces an adaptive learning loss function L adaptive , ensuring that the learning of new data of new training samples does not interfere with the feature representation of old data, and guiding the model to update adaptively through pseudo labels. Here, taking the identity of the ship with serial number i as an example, the formula is as follows:
[0080]
[0081] Among them, N3 is the number of features of the training sample of the new ship image, g is the sequence number of the memory feature, is the memory feature of the ship identity with serial number g in the updated long-term memory area, is the pseudo label of the memory feature with sequence number g, For serial number The characteristic mean of is the cluster center, For serial number The mean of the old cluster features. η is the hyperparameter that controls the learning setting of new data. For serial number The first half of this loss function aligns the new features with the old ones, while the second half prevents the original cluster center from shifting too much due to the addition or modification of new features. A third threshold θ3 is set. If the loss function exceeds this third threshold θ3, the new features are aligned with the old ones, and the adjusted new clusters are aligned with the old ones.
[0082] In some embodiments, as the environment changes (such as the ship's cargo, equipment replacement, changes in lighting conditions, etc.), the model needs to be able to dynamically adjust its feature matching strategy to enhance adaptability to new features. This invention introduces an adaptive matching strategy. Specifically, when new external equipment or cargo features become important, the model will increase its attention to them based on their matching frequency and stability, thereby dynamically adjusting the matching strategy. Here, taking the ship identity with serial number i as an example, the formula is as follows:
[0083]
[0084] where w g R is the dynamic matching weight of the new feature corresponding to the memory feature with the serial number g in the long-term memory area, and g is the serial number of the memory feature. ′ g is the frequency of occurrence of the memory feature with sequence number g in the ship identity with sequence number i in the long-term memory area, β is the adjustment coefficient, which controls the flexibility of the matching strategy, and R i is the frequency of occurrence of the ship identity with serial number i.
[0085] In some embodiments, forgetting learning is performed on the features of the feature clusters according to the frequencies of occurrence of the features and the feature clusters, including: updating the frequencies of occurrence of the features and the feature clusters each time the features are matched; constructing a forgetting loss function L forget When the difference between the occurrence frequency of a feature and the occurrence frequency of the corresponding feature cluster is less than the fourth threshold θ4, the feature is forgotten; the occurrence frequency of the corresponding feature cluster and the occurrence frequency of all its features are updated to the frequency of the least used feature in the cluster.
[0086] Reference Figure 6, each time a feature is matched, the frequency of occurrence of the feature is updated. Each ship identity has a corresponding memory cluster, and each memory cluster has a corresponding frequency of occurrence, which indicates the number of times the cluster (ship identity) has been successfully matched. When the number of matches of a cluster is different from the number of matches of certain features of the cluster, or the degree of match is low, these features are removed from the memory module. The present invention sets a fourth threshold value θ4. If the frequency of occurrence of a certain feature is lower than the fourth threshold value θ4 compared with the cluster to which it belongs, it is regarded as an outdated feature and is ready to be forgotten. The frequency of occurrence of the cluster and the frequency of occurrence of all its features are updated to the frequency of occurrence of the feature with the lowest frequency in the cluster. This process takes the memory cluster as the unit, and will not forget the features of other clusters. Here, taking the ship identity with serial number i as an example, the formula is as follows:
[0087]
[0088] Among them, N4 is the number of memory features in the memory cluster corresponding to the ship identity with serial number i in the long-term memory area, I[R ′ g <R i +θ f ] is the indicator function, which is used to select low-frequency features for forgetting, θ f To control the frequency threshold, set the hyperparameters. ||.|| is the L2 norm. It is the old memory feature of the ship identity with sequence number g in the long-term memory area with sequence number i.
[0089] In some embodiments, the feature clusters are updated by association learning and forgetting learning during each ship recognition training, including: updating the loss function L update , adaptive learning loss function L adaptive And the forgetting loss function L forget Perform weighted fusion to generate a comprehensive loss function L all ; Using the comprehensive loss function L all Balanced update loss function L update , adaptive learning loss function L adaptive And the forgetting loss function L forget Updates.
[0090] All loss functions are combined into a comprehensive loss function L all , by balancing the memory smooth update, forgetting and adaptive learning functions in a weighted sum manner, the formula is as follows:
[0091] L all =λ1L update +λ2L forget +λ3L adaptive
[0092] Among them, λ1 is the updated loss function L updateThe weight coefficient, λ2 is the adaptive learning loss function L adaptive The weight coefficient, λ3 is the forgetting loss function L forget During training, the gradient of each loss function is determined, and weights with large gradient changes are reduced, while those with small gradients are increased. The training is then verified using validation sets from both the new and old data domains. This means that mAP (Mean Average Precision) is used to further adjust each weight. The training ends when mAP is greater than or equal to the fifth threshold θ5.
[0093] The second stage mentioned above is completed by the association and forgetting learning module, which is used to reduce the focus on the ship's external equipment and cargo, so that the model can focus more on the ship's identity information.
[0094] In some embodiments, after updating the feature clusters by association learning and forgetting learning each time a ship is identified, the method further includes: clustering the updated features using the previous clustering parameters and the current clustering parameters after each clustering is completed; calculating the image-to-prototype similarity consistency loss L according to the clustering results; image1 and image-to-image similarity consistency loss L image2 ; Calculate the gap loss function L between the updated sample and the sample of the original feature cluster reg ; Consistency loss L based on image-to-prototype similarity image1 , image-to-image similarity consistency loss L image2 and the gap loss function L reg Adjust the current clustering parameters.
[0095] Reference Figure 7 The third stage is completed by the review learning module. The idea of "review" is used to solve the problem of poor domain adaptability and generalization ability of the model. After the model clustering in the second stage is completed, further optimization is carried out in this stage. First, before adapting to the new domain, the momentum encoder of the old model (the model that was fully trained before, if it is the first training, it will be set as both the new model and the old model) is frozen as the "old knowledge expert model θ old The purpose of freezing the momentum encoder is to avoid forgetting the old domain knowledge during the new domain training process and to ensure that the clustering parameters of the old domain model remain unchanged. Freezing means fixing the parameters of the model and not updating the gradient. The momentum encoder is based on the parameters saved during the training process of the old domain model and has an effective memory of the old domain information.
[0096] The historical memory area stores key samples from each model training session. This method uses samples from the historical memory area to "review" the clustering results of the old model under the old domain model. Simultaneously, the newly obtained model is used to process samples from the historical memory area to obtain clustering results under the new model. The clustering results of the new and old models are compared, and the old model is used to guide the update of the clustering parameters of the new model.
[0097] In order to ensure that the knowledge of the old domain is not forgotten when training in the new domain, and to enable the model to effectively adapt to the characteristics of the new domain, this paper introduces two similarity consistency losses: image to prototype similarity consistency loss L image1 and image-to-image similarity consistency loss L image2 The image-to-prototype similarity consistency loss is used to ensure that during the new domain training process, the similarity calculated by the model from the old domain (momentum encoder) and the new domain model is consistent, that is, the clustering results of the two for samples in the historical memory area are similar, to ensure that the new model can adapt to the data domain of the old model. The sixth threshold θ6 is set. If the value of the loss function exceeds the sixth threshold θ6, it means that the gap is large. The parameters of the new model are adjusted with reference to the old model to reduce the gap. The formula is as follows:
[0098]
[0099] Among them, N5 is the number of samples in the historical memory area, t is the serial number of the sample in the historical memory area, ||.|| is the L2 norm, P t is the prototype similarity vector, The prototype similarity vector calculated by the new domain model for the sample with sequence number t, The prototype similarity vector calculated by the old domain model for the sample with sequence number t in the historical memory area. The formula of the prototype similarity vector is as follows:
[0100]
[0101] Among them F t Represents the fine-grained ship image features of the sample with sequence number t in the historical memory area. t Represents the category prototype matrix, which is obtained by calculating the sample feature value of each category. Represents the category prototype matrix π t The transpose of τ is the temperature parameter used to control the smoothness of the distribution. The image-to-image similarity consistency loss L image2, is to learn the difference between the samples in the historical memory area and the samples in the new data domain, so as to improve the generalization ability in different fields. The seventh threshold θ7 is set. If the value of the loss function exceeds the seventh threshold θ7, it means that the data distribution gap between the two domains is large and similarity cannot be found. It is transformed into finding the distribution of the correlation of the parts in the domain, that is, finding the correlation between the two data domains. Here, the samples in the historical memory area are called old domain samples, and the new ship image training samples are called new domain samples. The image-to-image similarity consistency loss L image2 The formula is as follows:
[0102]
[0103] Among them, N v is the total number of samples participating in the comparison, N6 is the number of new domain samples participating in the comparison, N7 is the number of historical memory area samples participating in the comparison, e represents the serial number of the new domain sample, o represents the serial number of the historical memory area sample, sim() is the cosine similarity measurement formula, is the feature representation obtained by the new domain model based on the sample with sequence number e in the new domain sample, It is the feature representation obtained by the new domain model based on the sample with sequence number o in the historical memory area. is the feature representation obtained by the old domain model based on the sample with sequence number e in the new domain sample, It is the feature representation obtained by the old domain model based on the sample with sequence number o in the historical memory area.
[0104] The number of samples in each cluster is used to rank them, and the top U clusters containing the most samples are selected. To improve robustness, the present invention uses two strategies to select samples to be retained in the historical memory buffer. First, the sample closest to each cluster center is selected. Then, any sample within the cluster, excluding previously selected samples, is randomly selected. These samples and prototypes are used to regularize training in the new domain, ensuring that the model can adapt to the new domain while retaining important information from the old domain.
[0105] The present invention uses regularization to avoid excessive overlap between samples newly added to the historical memory area and samples in the original historical memory area, while ensuring the reliability of the selected samples, and sets the gap loss function L reg , the formula is as follows:
[0106]
[0107] Among them, N8 is the number of samples in the historical memory area, N9 is the number of samples newly added to the historical memory area, o is the serial number of the sample in the historical memory area, z is the serial number of the sample newly added to the historical memory area, ||.|| is the L2 norm, It is the feature representation obtained by the new domain model based on the sample with sequence number o in the historical memory area. Indicates the feature representation of the sample with sequence number z added to the historical memory area. The old domain model obtains the feature representation of the sample with serial number o in the historical memory area. In order to select more representative samples and avoid duplication, the eighth threshold θ8 is set. If the loss function is greater than the eighth threshold θ8, it means that the gap between the new sample and the existing sample is large and it can be included. Otherwise, if it is less than the threshold, it means that the gap between the new sample and the existing sample is small and it will not be included.
[0108] In some embodiments, the consistency loss L for image-to-prototype similarity image1 , image-to-image similarity consistency loss L image2 and the gap loss function L reg Perform weighted fusion to generate the total loss function L full ;
[0109] Using the total loss function L full Balanced image-to-prototype similarity consistency loss L image1 , image-to-image similarity consistency loss L image2 and the gap loss function L reg Updates.
[0110] During the training process, in order to balance the relationship between new domain adaptation and old domain knowledge retention, the total loss function L is designed. full , the formula is as follows:
[0111] L full =γ1L image1 +γ2L image2 +γ3L reg
[0112] Among them, γ1 is the similarity consistency loss L from image to prototype image1 The weight parameter, γ2 is the image-to-image similarity consistency loss L image2 The weight parameter, γ3 is the gap loss function L reg The weight parameters are used to balance the contributions of different loss terms to the total loss. During training, the gradient of each loss function is determined, and weights with large gradient changes are reduced, while those with small gradients are increased. The training is then verified using validation sets from both the new and old data domains. This means that the weights are further adjusted using mAP (Mean Average Precision). Model training is completed when mAP exceeds the ninth threshold, θ9, resulting in the final semi-supervised long-term ship re-identification model based on associative forgetting learning.
[0113] The above description and the accompanying drawings fully illustrate the embodiments of the present disclosure so that those skilled in the art can practice them. Other embodiments may include structural, logical, electrical, process and other changes. The embodiments represent only possible variations. Unless explicitly required, individual components and functions are optional, and the order of operations may vary. Parts and features of some embodiments may be included in or replace parts and features of other embodiments. Moreover, the words used in this application are only used to describe the embodiments and are not used to limit the claims. As used in the description of the embodiments and claims, unless the context clearly indicates otherwise, the singular forms "a", "an" and "the" are intended to also include plural forms. Similarly, the term "and / or" as used in this application refers to any and all possible combinations of one or more associated listings. In addition, when used in this application, the term "comprise" and its variations "comprises" and / or comprising refer to the presence of stated features, wholes, steps, operations, elements, and / or components, but do not exclude the presence or addition of one or more other features, wholes, steps, operations, elements, components and / or groups of these. In the absence of further restrictions, an element defined by the sentence "comprising a..." does not exclude the presence of other identical elements in the process, method or device that includes the element. In this article, each embodiment may focus on the differences from other embodiments, and the same and similar parts between the various embodiments can be referenced to each other. For the methods, products, etc. disclosed in the embodiments, if they correspond to the method part disclosed in the embodiments, then the relevant parts can be found in the description of the method part.
[0114] Those skilled in the art will appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are performed in hardware or software may depend on the specific application and design constraints of the technical solution. The technicians may use different methods to implement the described functions for each specific application, but such implementation should not be considered to be beyond the scope of the embodiments of the present disclosure. The technicians will clearly understand that, for the convenience and brevity of description, the specific working processes of the systems, devices and units described above can refer to the corresponding processes in the aforementioned method embodiments and will not be repeated here.
Claims
1. A semi-supervised long-term ship re-identification method based on associative forgetting learning, characterized by: include: Collect and cluster the multidimensional features of each ship to generate a feature cluster for each ship; Match the multidimensional features of the ship image with the features in the feature cluster and calculate the similarity, and output the ship recognition result based on the similarity; Perform association learning between the multi-dimensional features of ship images and the features of feature clusters; Perform forgetting learning on the features of feature clusters according to the frequencies of occurrence of features and feature clusters; The feature clusters are updated through association learning and forgetting learning during each ship recognition training.
2. The method according to claim 1, characterized in that Collect multi-dimensional characteristics of each vessel, including: Collect tag information and visual features of each ship; tag information includes ship type, color, angle and environment; Rewrite and expand the label information to generate high-level semantic information for each ship; Convert high-level semantic information into text features and fuse them with visual features, and convert the fused features into feature sequences; The self-attention mechanism is used to calculate the attention weight relationship between visual features and text features based on the feature sequence; Visual features and text features are fused according to the attention weight relationship to generate multi-dimensional features for each ship.
3. The method according to claim 2, characterized in that The self-attention mechanism is used to calculate the attention weight relationship between visual features and text features based on the feature sequence, including: Calculate the query matrix Q, key matrix K and value matrix V according to the feature sequence; where the query matrix Q comes from the visual feature projection, and the key matrix K and value matrix V come from the text features; Use the self-attention mechanism to calculate the attention score of visual features to semantic features; The Softmax function is used to convert the attention score into the attention weight of the visual feature to the semantic feature.
4. The method according to claim 1, wherein Convert all features into Hamming code representation; perform association learning between the multi-dimensional features of ship images and the features of feature clusters, including: Calculate the Hamming code distance between the multidimensional features of the ship image and the features of the feature cluster, and obtain the first σ1 closest feature samples; Randomly select σ2 random feature samples from the feature cluster; The multidimensional features of the ship, σ1 closest feature samples and σ2 random feature samples are clustered, and the cosine similarity is used for metric learning; Adjust features in feature clustering clusters based on cosine similarity.
5. The method according to claim 4, characterized in that Adjust the features in the feature clustering cluster based on cosine similarity, including: When the cosine similarity between the multidimensional features of the ship image and the features in the feature cluster is higher than a first threshold θ1, the multidimensional features of the ship image and the features in the feature cluster are weighted averaged to adjust the features in the feature cluster; When the cosine similarity of the multidimensional features of the ship image and the features in the feature cluster are both lower than the first threshold θ1, a new feature cluster is generated based on the multidimensional features of the ship image; Calculate the updated loss function L based on the feature clustering before and after feature update update , when the loss function is greater than or equal to the second threshold θ2, the update is determined, and when the loss function is less than the second threshold θ2, the update is canceled.
6. The method according to claim 5, characterized in that Generate new feature clusters based on the multidimensional features of ship images, including: Determine the feature cluster with the highest similarity to the new feature cluster based on the clustering results of the samples; Assign the label of the feature cluster with the highest similarity to the new feature cluster to generate a pseudo label for the new feature cluster; Calculate the adaptive learning loss function based on the new feature clustering cluster and the feature clustering cluster with the highest similarity; In the adaptive learning loss function L adaptive When it is greater than the third threshold θ3, the features of the new feature cluster are aligned with the features of the feature cluster with the highest similarity, and the new feature cluster after feature alignment is aligned with the feature cluster with the highest similarity.
7. The method according to claim 6, characterized in that The features of the feature clusters are forgotten based on the frequency of occurrence of the features and feature clusters, including: Each time a feature is matched, the frequency of occurrence of the feature and the feature cluster is updated; Construct the forgetting loss function L forget , when the difference between the occurrence frequency of a feature and the occurrence frequency of the corresponding feature cluster is less than the fourth threshold θ4, the feature is forgotten; Update the occurrence frequency of the corresponding feature cluster and the occurrence frequency of all its features to the frequency of the feature that appears least in the cluster.
8. The method according to claim 7, characterized in that The feature clusters are updated through association learning and forgetting learning at each ship recognition, including: Update the loss function L update , adaptive learning loss function L adaptive And the forgetting loss function L forget Perform weighted fusion to generate a comprehensive loss function L all ; Using the comprehensive loss function L all Balanced update loss function L update , adaptive learning loss function L adaptive And the forgetting loss function L forget Updates.
9. The method according to claim 1, characterized in that After updating the feature clusters through association learning and forgetting learning during each training, it also includes: After each clustering is completed, the updated features are clustered using the previous clustering parameters and the current clustering parameters; Calculate the similarity consistency loss L between the image and the prototype based on the clustering results image1 and image-to-image similarity consistency loss L image2 ; Calculate the gap loss function L between the updated sample and the sample of the original feature cluster cluster reg ; According to the image-to-prototype similarity consistency loss L image1 , image-to-image similarity consistency loss L image2 and the gap loss function L reg Adjust the current clustering parameters.
10. The method according to claim 9, characterized in that Also includes: Similarity consistency loss L for image to prototype image1 , image-to-image similarity consistency loss L image2 and the gap loss function L reg Perform weighted fusion to generate the total loss function L full ; Using the total loss function L full Balanced image-to-prototype similarity consistency loss L image1 , image-to-image similarity consistency loss L image2 and the gap loss function L reg Updates.
Citation Information
Cited By
Auxiliary memory method and system based on machine learning
CN121808050A