A Ship Image Retrieval Method Based on Deep Hashing

The deep hashing method with ViT and dynamic hashing addresses the limitations of existing ship image retrieval by enhancing accuracy and adaptability in complex marine environments through incremental learning and dynamic code updates.

CN119441532BActive Publication Date: 2025-07-15GUANGZHOU SHANGSAI ELECTRONIC TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411606604.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-11-12
Publication Date
2025-07-15
Estimated Expiration
2044-11-12

AI Technical Summary

Technical Problem

Existing ship image retrieval technology is difficult to achieve efficient and accurate image recognition and retrieval in complex and dynamic maritime environments. Traditional methods have problems such as strong subjectivity of feature extraction, low computational efficiency, poor adaptability, and catastrophic forgetting. Traditional hash encoding cannot effectively capture complex features and has high computational cost.

Method used

The deep hash network based on the ViT feature extraction module and the dynamic hash coding module is adopted, combined with the incremental learning mechanism, the deep features of the ship image are extracted through the ViT feature extraction module, and the dynamic hash coding module is used to generate and update the hash code, which solves the catastrophic forgetting problem in incremental learning and reduces the calculation cost.

Benefits of technology

It improves the accuracy and robustness of ship image retrieval, reduces hash conflicts, improves retrieval efficiency, reduces computing power and time costs, and adapts to complex and large-scale data changes.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119441532B_ABST
    Figure CN119441532B_ABST
Patent Text Reader

Abstract

The present invention discloses a ship image retrieval method based on deep hashing, which includes establishing an initial ship image database D0; constructing a deep hashing network based on a feature extraction module and a dynamic hashing encoding module; training the deep hashing network with D0 to obtain a ship image retrieval model M0, generating binary hash codes for the samples in D0 with M0, and storing them in the retrieval database; updating the ship image retrieval model based on incremental learning, generating binary hash codes for the newly added samples with the updated model and incorporating them into the retrieval database for ship image retrieval. The present invention proposes a scheme combining the ViT network with a dynamic hash code generation mechanism, which can dynamically adjust the hash codes when the data stream changes to ensure the flexibility and adaptability of the system in the face of large-scale and complex environmental data, thereby reducing hash conflicts, improving retrieval efficiency, and reducing training costs. It is applicable to application scenarios such as small-scale ship images and complex backgrounds in complex marine environments.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of image data processing, and particularly relates to a ship image retrieval method based on deep hashing. Background Art

[0002] Ships can be classified into cargo ships, oil tankers, fishing boats, etc. Their diversity in the marine environment increases the complexity of image recognition and retrieval. The existing image retrieval technologies mainly include the following. Method 1: Methods based on manual feature extraction, such as feature detection algorithms like SIFT (Scale-Invariant Feature Transform) and SURF (Speeded-Up Robust Features), which perform image matching by extracting key point features; Method 2: Convolutional neural network (CNN) methods based on deep learning, which use deep learning models to automatically extract deep features of images to achieve higher-precision image retrieval; Method 3: Image retrieval technologies based on hash coding, which compress image features into low-dimensional hash codes to achieve fast and efficient retrieval. However, these methods all have some defects.

[0003] Regarding Method 1: The subjectivity of feature extraction is strong, relying on the experience of designers, and it is difficult to extract features with universal adaptability. Manually extracting features is difficult to cope with complex backgrounds and environmental interferences, resulting in low retrieval accuracy. The computational efficiency is low. The process of manually extracting features is cumbersome and time-consuming, and it is difficult to meet the real-time retrieval requirements of large-scale data. Traditional computer vision feature detection algorithms such as SIFT and SURF are used for ship image retrieval, but their adaptability in dealing with scale, rotation, and illumination changes is limited, especially in complex marine environments.

[0004] Regarding Method 2, when extracting image features using convolutional neural network methods based on deep learning, they are restricted by local receptive fields. The reason is that CNN uses convolutional operations, which makes the network only learn from local features and relies on the receptive fields accumulated layer by layer to obtain global information; in addition, convolutional neural networks are sensitive to spatial changes in features and perform poorly in the face of spatial changes such as rotation, scaling, or deformation of objects. The above defects result in poor performance of CNN when extracting ship image features in complex marine backgrounds.

[0005] Regarding Method 3, the image retrieval technology based on traditional hash coding has deficiencies in multiple aspects. First, traditional hash coding methods rely on manual features or shallow feature transformations and cannot effectively capture complex features in images, resulting in low accuracy of image retrieval. Second, traditional hash coding is static, and the generated hash codes cannot be adjusted according to data changes, lacking adaptability and being difficult to cope with changes in different scenarios and data. In addition, traditional hash coding has a serious problem of information loss when facing high-dimensional data, leading to low retrieval accuracy. These deficiencies reflect the limitations of traditional hash coding technology when facing the diversity of ships and dynamic environments in complex maritime environments. Finally, due to the catastrophic forgetting problem, traditional deep hash networks need to retrain all historical data every time new data arrives. When the amount of data is large, there will be a serious computational burden.

[0006] Vision Transformer (ViT) is an image classification model based on Transformer, mainly including three parts: an image feature embedding module, a Transformer encoder module, and an MLP classification module. The image feature embedding module divides the input image into fixed-size patches, and then flattens these patches into a series of vectors. The Transformer encoder module contains multiple encoder layers, and each encoder layer contains a multi-head self-attention sub-layer and a feed-forward neural network sub-layer. In the ViT model, the image feature embedding module and the Transformer encoder module are mainly used for feature extraction, so they are collectively called the ViT feature extraction module. Summary of the Invention

[0007] The object of the present invention is to provide a method for retrieving ship images based on deep hashing that can solve the above problems, efficiently and accurately complete the retrieval task in complex and dynamic environments, overcome the catastrophic forgetting problem in the incremental learning process, enable the new network to maintain the hash code expression of the old network for old data with a small amount of historical data, and at the same time add the ability to express new data, thereby reducing the computing power cost and time cost.

[0008] To achieve the above object, the technical solution adopted by the present invention is as follows: A method for retrieving ship images based on deep hashing includes the following steps;

[0009] S1, establish an initial ship image database D0;

[0010] Obtain multiple ship images of different categories. For each ship image, preprocess and label the category to obtain samples, and all samples constitute the ship image database D0;

[0011] S2, construct a deep hash network, including a ViT feature extraction module M connected in sequenceViT and the dynamic hashing encoding module M hash , where M ViT is used to extract features from the sample x in D0 to obtain the corresponding image feature f(x), and M hash is used to perform dynamic hashing encoding on f(x) to output a real-valued hash code , and then convert it into a binary hash code h(x);

[0012] S3. Train the deep hashing network with D0 to obtain the ship image retrieval model M0, generate binary hash codes for all samples in D0 using M0, and store them in the retrieval database;

[0013] S4. Update the ship image retrieval model based on incremental learning. Each update results in an updated model. The update method for the t-th incremental learning includes steps S41 to S45;

[0014] S41. Obtain the ship image database D for the t-th incremental learning t , marked as the new database D new , t≥1, and when t = 1, D t = D0; Select q samples from each category in D0 to D t-1 to form the old database D old , mark the updated model after the (t - 1)-th time as M old ;

[0015] S42. Construct the total loss L hash for updating M total , including the cost function L stability and the hashing contrast loss L hash ;

[0016] ,

[0017] ,

[0018] = ,

[0019] In the formula, x i is a sample in D old , x j is a sample in D new , , are the real-valued hash codes of the sample x hash before and after the t-th update of M i respectively, is the L2 norm, is the real-valued hash code of the sample x hash after the t-th update of M j , Sij For sample x i and x j The similarity value of x i and x j If the categories are the same, then S ij =1, otherwise S ij =0,max is the maximum value operation, m is the boundary distance for distinguishing samples of different categories, and γ is the adjustment parameter of Lhash;

[0020] S43, constructed to update M ViT The loss function L ViT ;

[0021] ,

[0022] In the formula, , M ViT The corresponding samples x before and after the tth update i The image features of M ViT Adjustment parameters for the ability to remember old data, M ViT In D new The cross entropy loss calculated above;

[0023] S44, construct the total loss of the model L, L = L total +L ViT ;

[0024] S45, D old , D new After merging, enter M old , to minimize the total model loss L training M ViT and M hash , get the tth updated model M t ;

[0025] S5, use M t For D new Generate binary hash codes for the samples and store them in the retrieval database;

[0026] S6, based on M t Conduct ship image retrieval;

[0027] Get the ship image x to be queried query , by M t Generate a binary hash code, search in the search database, and output the search results.

[0028] As a preferred embodiment: S1 specifically comprises: unifying the resolution of the collected ship images to 256 256, classify and label the samples by category, then rotate each sample by several angles. Each rotation forms a new sample. The samples and the new samples together constitute the initial ship image database D0.

[0029] Preferably: In S2, M hash Obtain the real-valued hash code according to the following formula and the binary hash code h(x);

[0030] ,

[0031] ,

[0032] In the formula, W h , b h are respectively the initial bias matrix and the initial weight matrix of M hash , is the hyperbolic tangent function, is the sign function. If ≥0, then , otherwise , and K is the length of the hash code.

[0033] Preferably: The ViT feature extraction module is composed of the image feature embedding module and the Transformer encoder module of the ViT model;

[0034] The image feature embedding module is used to divide the sample x into several image patches, then flatten each image patch into a one-dimensional vector, map it to a D-dimensional vector through a linear mapping layer and perform position encoding to obtain the position encoding vector;

[0035] The Transformer encoder module includes multiple cascaded Transformer encoders, which are used to input the position encoding vector and output the image feature f(x) after multiple feature extractions.

[0036] Preferably: In S45, each time M hash is trained, based on L total update the weight matrix and bias matrix of the hash function during backpropagation;

[0037] ,

[0038] In the formula, , are respectively the updated and the pre-updated weight matrices, and when t = 1, , , are respectively the updated and the pre-updated bias matrices, and when t = 1, , is the learning rate.

[0039] Preferably, S6 specifically is;

[0040] S61, preset distance threshold τ and preset number of retrieval results M;

[0041] S62, use M t to generate the binary hash code h(x query ) of x, and calculate the Hamming distance between h(x query ) and all binary hash codes in the retrieval database; query

[0042] S63, sort the samples in the retrieval database in ascending order of Hamming distance to generate a sequence list;

[0043] S64, select the first M samples in the sequence list whose Hamming distance is less than τ as the retrieval results.

[0044] Preferably, the categories include seagoing ships, container ships, three-five ships, civilian self-use ships and law enforcement ships.

[0045] The present invention is applicable to the identification of an increasing number of ship samples at a maritime checkpoint. The idea of the present invention is:

[0046] First, construct a deep hash network based on a ViT feature extraction module and a dynamic hash encoding module, train the deep hash network with an initial ship image database D0 to obtain a ship image retrieval model M0, generate binary hash codes for all samples in D0 with M0, and store them in the retrieval database.

[0047] Subsequently, continuously collect ship images, regularly form ship image databases. For example, the ship image databases collected for the first time and the second time later are marked as D1 and D2, and the ship image databases constructed multiple times form a data stream D = {D1, D2,..., D t}, and then propose a dynamic hash code generation and update mechanism based on the data stream combined with sample incremental learning. When there is a new database / data stream input, re-encode the new samples by updating the weight matrix and bias matrix during dynamic hash encoding. Since the above two matrices are updated every time there is a database input, the encoding methods of the samples in each database are different, forming a dynamic encoding method. Specifically:

[0048] (1) First, define a hash function to generate a real-valued hash code according to the formula, which is to make the gradient transmittable during the training process, so the continuous activation function tanh is used to approximate the sign function; subsequently, in order to make the output of M be a binary code, it is necessary to perform binary encoding on hash . The present invention performs binary encoding according to the formula , convert the real-valued hash code into a binary hash code h(x), where is the sign function, otherwise .

[0049] (2) When a new database arrives, the original hash function may no longer be applicable to the new samples, which may lead to problems such as an increase in hash conflicts and a decrease in retrieval accuracy. To solve this problem, the present invention introduces a dynamic hash code update mechanism. While updating the hash function, it tries to keep the hash codes of the old data unchanged as much as possible, reducing the impact on the existing index structure, and at the same time enabling the new hash function to effectively distinguish between similar and dissimilar samples in the new data. The dynamic hash function update method is divided into the following steps: First, to keep the hash codes of the old data stable, for the sample x in the old database i , it is desired that the updated hash code be as consistent as possible with the original hash code. Therefore, a cost function is designed. At the same time, to adapt to the hash code learning of the sample x j in the new database, a hash code contrast loss is used to ensure that the similarity relationship is maintained in the hash space. A hash contrast loss L hash is designed, which aims to minimize the hash code distance between similar samples and maximize the distance between dissimilar samples. The γ in L hash is used to balance the adaptability to new data and the stability of old data. In summary, for L stability , it means ensuring that the hash codes of the sampled old data samples are as consistent as possible between the model outputs before and after; for L hash , it means ensuring that the hash codes of similar images are pulled closer and the hash codes of dissimilar images are pushed farther apart, thereby generating hash codes with higher discrimination.

[0050] (3) To further improve the adaptability and stability of the overall model when the data changes, the present invention also combines the dynamic hash code update mechanism with the incremental learning mechanism to jointly ensure the adaptability and stability of the system when the data changes. An L ViT is designed, and its formula fine-tunes M ViT during each training, updates the parameters of M ViT to adapt to the new data while maintaining the performance on the old data.

[0051] Compared with the prior art, the advantages of the present invention are as follows:

[0052] (1) Combining the powerful feature extraction ability of the ViT network: In the present invention, the deep features of ship images are extracted through the ViT feature extraction module. Compared with traditional convolutional neural networks, ViT can better capture the global features of images, thereby improving the accuracy and robustness of ship image retrieval.

[0053] (2) Dynamic hash code generation and update mechanism: The present invention proposes a dynamic hash code generation and update mechanism that can dynamically adjust the hash code when the data stream changes, ensuring the flexibility and adaptability of the system in the face of large-scale and complex environmental data, thereby reducing hash conflicts and improving retrieval efficiency.

[0054] (3) Incremental learning mechanism to support real-time updates: By introducing an incremental learning mechanism, the present invention can update the model when new data is input and preserve the expression ability on the old data as much as possible. The method is to randomly select a small amount of data from the old dataset and mix it with the new dataset for training, avoiding the high computational and time costs of training the deep dynamic hash model using all the old data, and ensuring that the training cost is reduced in the scenario of expanding data. BRIEF DESCRIPTION OF THE DRAWINGS

[0055] Figure 1 Structural diagram of the deep hash network of the present invention;

[0056] Figure 2 It is a flowchart of the present invention. DETAILED DESCRIPTION OF THE INVENTION

[0057] The present invention will be further described below in conjunction with the accompanying drawings.

[0058] Example 1: Refer to Figure 1 and Figure 2 , a ship image retrieval method based on deep hashing, comprising the following steps;

[0059] S1, establish an initial ship image database D0;

[0060] Obtain multiple ship images of different categories. For each ship image, preprocess and label the category to obtain samples. All samples constitute the ship image database D0, and the categories include seagoing ships, container ships, sampans, civilian self-use ships, and law enforcement ships;

[0061] S2, construct a deep hash network, including a ViT feature extraction module M ViT and a dynamic hash encoding module M hash , where M ViT is used to extract features from the sample x in D0 to obtain the corresponding image feature f(x), and M hash is used to perform dynamic hash encoding on f(x) and output a real-valued hash code , and then convert it into a binary hash code h(x);

[0062] S3, train the deep hash network with D0 to obtain a ship image retrieval model M0, generate binary hash codes for all samples in D0 with M0, and store them in the retrieval database;

[0063] S4. Update the ship image retrieval model based on incremental learning. Each update results in an updated model. The update method for the t-th incremental learning includes steps S41 to S45;

[0064] S41. Obtain the ship image database D for the t-th incremental learning t , marked as the new database D new , where t ≥ 1, and when t = 1, D t = D0; Select q samples from each category in D0 to D t-1 to form the old database D old , and mark the updated model of the (t - 1)-th time as M old ; The value of q can be selected as needed, such as 30, 50, 80, etc.;

[0065] S42. Construct the total loss L hash for updating M total , including the cost function L stability and the hash comparison loss L hash ;

[0066] ,

[0067] ,

[0068] = ,

[0069] where x i is a sample in D old , x j is a sample in D new , , are the real-valued hash codes of the sample x hash before and after the t-th update of M i , is the L2 norm, is the real-valued hash code of the sample x hash after the t-th update of M j , S ij is the similarity value of the samples x i and x j . If the categories of x i and x j are the same, then S ij = 1, otherwise S ij = 0, max is the maximum operation, m is the boundary distance for distinguishing samples of different categories, and γ is the adjustment parameter of Lhash;

[0070] S43. Construct for updating MViT Loss function L ViT ;

[0071] ,

[0072] wherein, and are the image features of the corresponding sample x before and after the t-th update respectively, β is a regulation parameter for controlling the ability of M to remember old data, ViT is the cross-entropy loss calculated by M i on D ViT ; is M ViT on D new ;

[0073] S44. Construct the total loss L of the model, L = L total + L ViT ;

[0074] S45. After merging D old and D new , input them into M old to train M ViT and M hash to minimize the total loss L of the model, and obtain the updated model M t at the t-th time;

[0075] S5. Use M t to generate binary hash codes for the samples in D new , and store them in the retrieval database;

[0076] S6. Perform ship image retrieval based on M t ;

[0077] Obtain the ship image x query to be queried, generate a binary hash code through M t , retrieve in the retrieval database, and output the retrieval result.

[0078] Example 2: Refer to Figure 1 and Figure 2 . On the basis of Example 1, S1 is specifically to unify the resolution of the collected ship images to 256 × 256, classify and label them by category to obtain samples, then rotate each sample by several angles, and form a new sample every time it rotates. The samples and the new samples together constitute the initial ship image database D0. For example, there are 2000 ship images, which are divided into a training set and a test set of 1800 and 200 respectively. The 1800 ship images are classified to obtain 1800 samples, and then the samples are rotated by three angles of 90°, 180°, and 270° respectively to obtain 5400 new samples. Therefore, the initial ship image database D0 has a total of 7200 samples.

[0079] Example 3: Refer to Figure 1 and Figure 2 , on the basis of Example 1, in S2, M hash obtains the real-valued hash code according to the following formula and the binary hash code h(x);

[0080] ,

[0081] ,

[0082] In the formula, W h and b h are the initial bias matrix and the initial weight matrix of M hash respectively, is the hyperbolic tangent function, is the sign function. If ≥0, then , otherwise , and K is the length of the hash code.

[0083] Example 4: Refer to Figure 1 and Figure 2 , on the basis of Example 1, the ViT feature extraction module is composed of the image feature embedding module and the Transformer encoder module of the ViT model; the image feature embedding module is used to divide the sample x into several image patches, then flatten each image patch into a one-dimensional vector, map it to a D-dimensional vector through a linear mapping layer, and perform position encoding to obtain the position encoding vector; the Transformer encoder module includes multiple layers of cascaded Transformer encoders, which are used to input the position encoding vector and output the image feature f(x) after multiple feature extractions.

[0084] Regarding the image feature embedding module: The image feature embedding module includes an image block layer, a linear mapping layer, and a position embedding layer. The image block layer is used to perform the block operation, divide the sample x into several image patches and flatten each image patch into a one-dimensional vector. The linear mapping layer is used to map each one-dimensional vector to a D-dimensional vector. The position embedding layer is used to perform position encoding on each D-dimensional vector to obtain the position encoding vector.

[0085] Regarding the Transformer encoder module: In the Transformer encoder module, each layer of the Transformer encoder includes a multi-head self-attention mechanism and a feed-forward neural network. The CLS token vector of the last layer of the Transformer encoder is used as the image feature f(x).

[0086] Example 5: Refer to Figure 1 and Figure 2, based on Example 1, in S45, each time when training M hash , the weight matrix and bias matrix of the hash function are updated during backpropagation; total When backpropagating, update the weight matrix and bias matrix of the hash function;

[0087] ,

[0088] In the formula, , are the updated and pre-updated weight matrices respectively, and when t = 1, , , are the updated and pre-updated bias matrices respectively, and when t = 1, , is the learning rate.

[0089] Example 6: Refer to Figure 1 and Figure 2 , based on Example 1, S6 specifically includes steps S61 to S64;

[0090] S61, preset a distance threshold τ and a preset number of retrieval results M;

[0091] S62, use M t to generate the binary hash code h(x query ) of x, and calculate the Hamming distance between h(x query ) and all binary hash codes in the retrieval database; query ) Calculate the Hamming distance between h(x

[0092] S63, sort the samples in the retrieval database in ascending order of Hamming distance to generate a sequence list;

[0093] S64, select the first M samples in the sequence list whose Hamming distance is less than τ as the retrieval results.

[0094] Example 7: Refer to Figure 1 and Figure 2 , based on Example 1, we present a comparative experiment example for the ship data retrieval scenario to verify the advantages of the present invention over the prior art in terms of accuracy and efficiency. In the experiment, a real ship image dataset is selected, which contains a large number of ship images of different types, such as seagoing ships, container ships, sampans, civilian self-use ships, law enforcement ships, etc., with 22 classifications, and the dataset is divided into an 85% training set and a 15% test set.

[0095] In the experiment, we selected the general retrieval accuracy in the industry as the evaluation index and compared the performance of the present invention with three other existing methods on these indexes. The experimental results of each model are shown in Table 1.

[0096] Table 1. Comparison table of experimental results of image retrieval models

[0097] Method \ Hash Code Length (number of bits) 64 bits 128 bits 256 bits Traditional SIFT Feature Method 73.28% 74.2% 75.5% Traditional SURT Feature Method 75.3% 77.1% 78.4% Deep Hashing Network (based on CNN) 84.1% 86.5% 88.2% Method of the Present Invention 88.9% 91.2% 93.1%

[0098] As can be seen from Table 1, with the increase in the length of the hash code, the retrieval accuracy of each method has been improved. However, the deep hash coding of the present invention always performs better than other methods under different hash code lengths. Especially when the hash code length reaches 256 bits, the present invention can significantly improve the retrieval accuracy and achieve a better effect, while there is still a large gap in accuracy between other methods and the present invention.

[0099] The above are only the preferred embodiments of the present invention and are not intended to limit the present invention. Any modifications, equivalent replacements, and improvements made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.

Claims

1. A ship image retrieval method based on deep hashing, characterized in that: Including the following steps; S1. Establish an initial ship image database D0; Obtain multiple ship images of different categories. For each ship image, preprocess and label the category to obtain samples. All the samples constitute the ship image database D0; S2. Construct a deep hash network, including a ViT feature extraction module M ViT and a dynamic hash coding module M hash , where M ViT is used to extract features from the sample x in D0 to obtain the corresponding image feature f(x), and M hash is used to perform dynamic hash coding on f(x) to output a real-valued hash code , and then convert it into a binary hash code h(x); S3. Use D0 to train a deep hashing network to obtain a ship image retrieval model M0. Use M0 to generate binary hash codes for all samples in D0 and store them in the retrieval database; S4. Update the ship image retrieval model based on incremental learning. Each update results in an updated model. The update method for the t-th incremental learning includes steps S41 to S45; S41. Obtain the ship image database D for the t-th incremental learning t , and label it as the new database D new , where t ≥ 1, and when t = 1, D t = D0; Select q samples from each category in D0~D t-1 to form the old database D old , and label the updated model of the (t - 1)-th time as M old ; S42, construct for updating M hash total loss L total , including cost function L stability and hash comparison loss L hash ; , , = , where x i is a sample in D old and x j is a sample in D new . and are the real-valued hash codes of sample x hash before and after the t-th update by M i respectively. is the L2 norm, is the real-valued hash code of sample x hash after the t-th update by M j . S ij is the similarity value of samples x i and x j . If x i and x j are in the same class, then S ij = 1; otherwise, S ij = 0. max is the operation of taking the maximum value, m is the boundary distance for distinguishing samples of different classes, and γ is the adjustment parameter of Lhash; S43, construct a loss function L for updating M ViT ; ViT ; , In the formula, and are the image features of the corresponding sample x ViT before and after the t-th update of M, respectively, β is the adjustment parameter for controlling the ability of M i to remember old data, ViT and is the cross-entropy loss calculated by M ViT on D new . S44, construct the total loss L of the model, L = L total + L ViT ; S45, input D old and D new After merging, input M old and train M to minimize the total model loss L ViT and M hash to obtain the updated model M at the t-th time t ; S5, using M t to generate binary hash codes for the samples in D new and store them in the retrieval database; S6, based on M t Perform ship image retrieval; Obtain the image x of the ship to be queried query , through M t Generate a binary hash code, retrieve it in the retrieval database, and output the retrieval result, specifically including S61~S64; S61. Preset a distance threshold τ and a preset number M of retrieval results; S62, using M t Generate x query The binary hash code h(x query ) Calculate h(x query ) and the Hamming distance from all binary hash codes in the retrieval database; S63. Sort the samples in the retrieval database in ascending order of Hamming distance to generate a sequence list; S64. Select the first M samples in the sequence list with a Hamming distance less than τ as the retrieval results.

2. The method for retrieving ship images based on deep hashing according to claim 1, wherein: Specifically, S1 is to unify the resolution of the collected ship images to 256 256, classify and label the samples by category, then rotate each sample by several angles, and form a new sample each time it rotates. The samples and the new samples together constitute the initial ship image database D0.

3. A method for retrieving ship images based on deep hashing according to claim 1, characterized in that: In S2, M hash Obtain a real-valued hash code according to the following formula and a binary hash code h(x); , , where, W h , b h are respectively the initial bias matrix and the initial weight matrix of M hash , is the hyperbolic tangent function, is the sign function. If ≥0, then , otherwise , and K is the length of the hash code.

4. A method for retrieving ship images based on deep hashing according to claim 1, characterized in that: The ViT feature extraction module consists of an image feature embedding module and a Transformer encoder module of the ViT model; The image feature embedding module is used to divide the sample x into several image patches, then flatten each image patch into a one-dimensional vector, map it to a D-dimensional vector through a linear mapping layer, and perform position encoding to obtain a position encoding vector; The Transformer encoder module includes multiple cascaded Transformer encoders, which are used to input the position encoding vector and output the image feature f(x) after multiple feature extractions.

5. A method for retrieving ship images based on deep hashing according to claim 1, characterized in that: In S45, each time M is trained hash when, based on L total update the weight matrix and bias matrix of the hash function during backpropagation; , In the formula, and are the weight matrices after and before the update respectively. When t = 1, , and are the bias matrices after and before the update respectively. When t = 1, , is the learning rate.

6. The method for retrieving ship images based on deep hashing according to claim 1, wherein: The categories include seagoing ships, container ships, sampans, civilian self-use ships, and law enforcement ships.

Citation Information

Patent Citations

  • Unsupervised image deep hash retrieval method and system

    CN113722529A

  • Ocean remote sensing ship image retrieval method based on self-attention hashing

    CN118312636A