Retrieval Method, Device, Electronic Device, Storage Medium and Product for Ship Images
By using visual language models based on ship full-color and multispectral image datasets for multimodal contrast learning and semantic fusion, hash codes are generated to improve the cross-modal retrieval accuracy of ship images, solving the problem of degradation of search accuracy in the prior art.
Patent Information
- Application Number
- CN202510325169.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-19
- Publication Date
- 2025-06-13
- Estimated Expiration
- 2045-03-19
AI Technical Summary
In the prior art, simply searching text to image or searching image to image leads to a decrease in cross-modal retrieval accuracy of remote sensing images.
A ship visual language model trained based on ship full-color image dataset, ship multi-spectral image dataset and label text is used to generate a hash code to determine the target search results through multimodal contrast learning, cross-modal semantic fusion and hash learning processing.
The accuracy of cross-modal search of ship images to be retrieved in the ship sample database is improved, and the problem of reduced search accuracy in the prior art is solved.
Smart Images

Figure CN119848287B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of image processing, and in particular, to a method, device, electronic device, storage medium and product for retrieving ship images. Background Art
[0002] With the rapid development of earth observation technology, a large amount of remote sensing image data has been generated, and there is an urgent need for an efficient and convenient data management strategy.
[0003] In the prior art, for the retrieval of cross-modal remote sensing text images, there is a text-to-image retrieval method, which mainly uses the inherent high-level semantic information in the text to overcome the limitations of image-to-image retrieval, and realizes the retrieval of cross-modal remote sensing ship images through text features. There is also an image-to-image retrieval method, which mainly uses the detailed visual information from the image itself for retrieval. The above text-to-image retrieval method cannot capture the rich and detailed image information existing in the remote sensing image, while the image-to-image retrieval method ignores the abstract semantic information provided by the text label. Therefore, the accuracy of cross-modal retrieval of remote sensing images is reduced. Summary of the Invention
[0004] The present invention provides a method, device, electronic device, storage medium and product for retrieving ship images, which is used to solve the defect that the accuracy of cross-modal retrieval of remote sensing images is reduced due to simply text-to-image retrieval or image-to-image retrieval in the prior art, and realizes multi-modal contrast learning, cross-modal semantic fusion and hashing learning processing on the ship image to be retrieved and the ship sample database by using a ship visual language model combined with images and texts trained based on a ship panchromatic image dataset, a ship multi-spectral image dataset and label texts, and finally determines the target retrieval result based on the obtained hash code and the sample hash code, so as to improve the accuracy of cross-modal retrieval of the ship image to be retrieved in the ship sample database.
[0005] The present invention provides a method for retrieving ship images, including the following steps.
[0006] Obtain a ship image to be retrieved and a ship sample database; wherein, the ship sample database contains at least two ship sample images, and the ship sample images are ship images of different modalities from the ship image to be retrieved.
[0007] Input the ship image to be retrieved and all ship sample images in the ship sample database into the ship vision-language model respectively, to obtain the hash code corresponding to the ship image to be retrieved and all sample hash codes corresponding to all ship sample images in the ship sample database output by the ship vision-language model; wherein, the ship vision-language model is trained based on the obtained multi-modal ship sample pattern dataset including the ship panchromatic image dataset, the ship multi-spectral image dataset and the label text, and the ship vision-language model is a model that processes the ship image to be retrieved and the ship sample database based on multi-modal contrast learning, cross-modal semantic fusion and hashing learning to obtain the corresponding hash code and sample hash code.
[0008] Determine all Hamming similarities respectively according to the hash code and all sample hash codes, and determine a similarity ranking list according to all Hamming similarities.
[0009] Determine the target retrieval result according to the ranking list of all Hamming similarities; wherein, the target retrieval result is a set of ship sample images similar to the ship image to be retrieved selected based on all similarities in the ranking list.
[0010] According to a ship image retrieval method provided by the present invention, the ship vision-language model is trained based on the following steps: obtain a multi-modal ship sample pattern dataset; wherein, the multi-modal ship sample pattern dataset includes a ship panchromatic image dataset, a ship multi-spectral image dataset and label text, and the label text is the paired label text corresponding to the paired samples determined by pairing the ship panchromatic image dataset and the ship multi-spectral image dataset; input the ship panchromatic image dataset, the ship multi-spectral image dataset and the label text into a basic ship vision-language model, to obtain the dataset hash code and the target loss function output by the basic ship vision-language model; wherein, the basic ship vision-language model is a pre-determined basic model for training; optimize the basic ship vision-language model according to the dataset hash code and the target loss function to obtain the ship vision-language model.
[0011] A method for retrieving ship images provided by the present invention inputs a ship panchromatic image dataset, a ship multispectral image dataset, and label texts into a basic ship vision-language model to obtain a dataset hash code and an objective loss function output by the basic ship vision-language model, including: inputting the ship panchromatic image dataset and the ship multispectral image dataset into a multimodal contrastive learning network in the basic ship vision-language model to obtain image features and a contrastive loss output by the multimodal contrastive learning network; the image features are obtained by the multimodal contrastive learning network performing image enhancement processing and similarity measurement on the ship panchromatic image dataset and the ship multispectral image dataset; inputting the image features and the label texts into a cross-modal semantic fusion network in the basic ship vision-language model to obtain fusion features and a fusion loss output by the cross-modal semantic fusion network; wherein, the fusion features are obtained by the cross-modal semantic fusion network performing alignment and fusion on the image features and the label texts; inputting the fusion features into a hash learning network in the basic ship vision-language model to obtain a dataset hash code and a hash loss output by the hash learning network; wherein, the dataset hash code is obtained by the hash learning network performing hash learning on the fusion features; determining the objective loss function according to the contrastive loss, the fusion loss, and the hash loss.
[0012] A method for retrieving ship images provided by the present invention inputs the fusion features into a hash learning network in the basic ship vision-language model to obtain a dataset hash code output by the hash learning network, including: inputting the fusion features into a hash learning layer in the hash learning network to obtain hash features output by the hash learning layer; wherein, the hash learning layer is a layer that processes the fusion features based on a non-linear activation function; inputting the hash features into a binarization layer in the hash learning network to obtain a dataset hash code output by the binarization layer; wherein, the binarization layer is a layer that performs binary processing on the hash features based on a sign function.
[0013] A method for retrieving ship images provided by the present invention determines all Hamming similarities according to the hash code and all sample hash codes respectively, including: multiplying the hash code and all sample hash codes respectively to obtain all Hamming similarities.
[0014] A method for retrieving ship images provided by the present invention determines a target retrieval result according to a ranking list of all Hamming similarities, including: determining all candidate similarities according to all Hamming similarities in the ranking list and a similarity threshold; wherein, all candidate similarities are Hamming similarities in the ranking list that are greater than the similarity threshold, and the similarity threshold is a preset threshold; determining all ship sample images corresponding to all candidate similarities as the target retrieval result.
[0015] The present invention also provides a ship image retrieval device, including the following modules.
[0016] An image acquisition module for acquiring the ship image to be retrieved and a ship sample database; wherein, the ship sample database contains at least two ship sample images, and the ship sample images are ship images of different modalities from the ship image to be retrieved.
[0017] A hash code output module for respectively inputting the ship image to be retrieved and all ship sample images in the ship sample database into a ship visual language model to obtain the hash code corresponding to the ship image to be retrieved output by the ship visual language model and all sample hash codes corresponding to all ship sample images in the ship sample database; wherein, the ship visual language model is trained based on the obtained multi-modal ship sample pattern dataset including a ship panchromatic image dataset, a ship multi-spectral image dataset, and label texts, and the ship visual language model is a model that processes the ship image to be retrieved and the ship sample database based on multi-modal contrast learning, cross-modal semantic fusion, and hash learning to obtain the corresponding hash code and sample hash code.
[0018] A list determination module for respectively determining all Hamming similarities based on the hash code and all sample hash codes, and determining a similarity ranking list based on all Hamming similarities.
[0019] A result determination module for determining a target retrieval result based on the ranking list of all Hamming similarities; wherein, the target retrieval result is a set of ship sample images similar to the ship image to be retrieved selected based on all similarities in the ranking list.
[0020] The present invention also provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the computer program, the retrieval method of any one of the above ship images is implemented.
[0021] The present invention also provides a non-transitory computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the retrieval method of any one of the above ship images is implemented.
[0022] The present invention also provides a computer program product, including a computer program. When the computer program is executed by a processor, the retrieval method of any one of the above ship images is implemented.
[0023] The present invention provides a method, device, electronic device, storage medium and product for searching ship images. The method comprises the following steps: obtaining a ship image to be searched and a ship sample database; wherein the ship sample database contains at least two ship sample images, and the ship sample images are ship images of different modalities from the ship image to be searched; the ship image to be searched and all the ship sample images in the ship sample database are respectively input into a ship visual language model, and a hash code corresponding to the ship image to be searched output by the ship visual language model and all sample hash codes corresponding to all the ship sample images in the ship sample database are obtained; wherein the ship visual language model is based on the obtained full-color image of the ship. The ship visual language model is trained on a multimodal ship sample pattern data set, a ship multispectral image data set and a multimodal ship sample pattern data set with labeled texts. The ship visual language model is a model that processes the ship images to be retrieved and the ship sample database based on multimodal contrastive learning, cross-modal semantic fusion and hash learning to obtain corresponding hash codes and sample hash codes; all Hamming similarities are determined according to the hash codes and all sample hash codes, respectively, and a similarity ranking list is determined according to all Hamming similarities; a target retrieval result is determined according to the ranking list of all Hamming similarities; wherein the target retrieval result is a set of ship sample images similar to the ship image to be retrieved, selected based on all similarities in the ranking list. The technical solution of the present invention is used to solve the defect in the prior art that the accuracy of cross-modal retrieval of remote sensing images is reduced by simply performing text-to-image retrieval or image-to-image retrieval, and realizes a ship visual language model that combines images and texts obtained by training based on a ship panchromatic image dataset, a ship multispectral image dataset and label text, and performs multimodal comparative learning, cross-modal semantic fusion and hash learning processing on the ship images to be retrieved and the ship sample database. Finally, the target retrieval result is determined based on the obtained hash code and the sample hash code, thereby improving the accuracy of cross-modal retrieval of the ship images to be retrieved in the ship sample database. BRIEF DESCRIPTION OF THE DRAWINGS
[0024] In order to more clearly illustrate the technical solutions in the present invention or the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.
[0025] Figure 1 It is a flow chart of the ship image retrieval method provided by the present invention.
[0026] Figure 2 It is a structural schematic diagram of the ship visual language model provided by the present invention.
[0027] Figure 3It is a schematic diagram of the alignment and fusion method provided by the present invention.
[0028] Figure 4 It is a schematic diagram of the attention mechanism fusion method provided by the present invention.
[0029] Figure 5 It is a schematic diagram of the image sample provided by the present invention.
[0030] Figure 6 It is a schematic structural diagram of the retrieval device for ship images provided by the present invention.
[0031] Figure 7 It is a schematic structural diagram of the electronic device provided by the present invention. Detailed implementation manners
[0032] To make the objectives, technical solutions and advantages of the present invention clearer, the technical solutions in the present invention will be clearly and completely described below with reference to the accompanying drawings in the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments in the present invention without making creative efforts shall fall within the protection scope of the present invention.
[0033] The following combines Figure 1 Describe the retrieval method for ship images provided by the present invention. The retrieval method for ship images provided by the present invention is applicable to the retrieval of cross-modal remote sensing ship images. The execution subject of this method can be an electronic device or a retrieval device for ship images set in the electronic device. The retrieval device for ship images can be implemented by software, hardware, or a combination of both. Figure 1 It is a schematic flow diagram of the retrieval method for ship images provided by the present invention. As Figure 1 shown, the method includes the following steps 101, 102, 103, and 104.
[0034] Step 101: Obtain the ship image to be retrieved and the ship sample database.
[0035] In this step, the ship sample database contains at least two ship sample images, and the ship sample images are ship images of different modalities from the ship image to be retrieved.
[0036] Specifically, obtain the ship image to be retrieved that needs to be retrieved, and at the same time obtain the ship sample database containing at least two ship sample images of different modalities from the ship image to be retrieved.
[0037] Step 102: Input the ship image to be retrieved and all ship sample images in the ship sample database into the ship vision-language model respectively, to obtain the hash code corresponding to the ship image to be retrieved output by the ship vision-language model and all sample hash codes corresponding to all ship sample images in the ship sample database.
[0038] In this step, the ship vision-language model is trained based on the obtained multi-modal ship sample pattern dataset including a ship panchromatic image dataset, a ship multi-spectral image dataset, and label texts. The ship vision-language model is a model that processes the ship image to be retrieved and the ship sample database based on multi-modal contrast learning, cross-modal semantic fusion, and hash learning to obtain the corresponding hash code and sample hash codes.
[0039] Specifically, based on the trained ship vision-language model, input the ship image to be retrieved and all ship sample images in the ship sample database into the ship vision-language model respectively, and process the ship image to be retrieved and all ship sample images in the ship sample database based on the trained ship vision-language model through multi-modal contrast learning, cross-modal semantic fusion, and hash learning, so as to obtain the hash code corresponding to the ship image to be retrieved and all sample hash codes corresponding to all ship sample images in the ship sample database.
[0040] In a specific embodiment, the ship vision-language model is trained based on the following steps: Obtain a multi-modal ship sample pattern dataset; wherein, the multi-modal ship sample pattern dataset includes a ship panchromatic image dataset, a ship multi-spectral image dataset, and label texts, and the label texts are the paired label texts corresponding to the paired samples determined by pairing the ship panchromatic image dataset and the ship multi-spectral image dataset; Input the ship panchromatic image dataset, the ship multi-spectral image dataset, and the label texts into the basic ship vision-language model to obtain the dataset hash code and the target loss function output by the basic ship vision-language model; wherein, the basic ship vision-language model is a pre-determined basic model for training; Optimize the basic ship vision-language model according to the dataset hash code and the target loss function to obtain the ship vision-language model.
[0041] In this step, the basic ship vision-language model is a language model pre-obtained by the user for image recognition and processing, and this embodiment does not limit this.
[0042] Specifically, Figure 2 is the structural schematic diagram of the ship vision-language model provided by the present invention, as Figure 2 shown. The obtained ship vision-language model includes a multi-modal contrast learning network, a cross-modal semantic fusion network, and a hash learning network.
[0043] In a specific embodiment, a ship full-color image dataset, a ship multispectral image dataset, and label texts are input into a basic ship vision-language model to obtain a dataset hash code and an objective loss function output by the basic ship vision-language model, including: inputting the ship full-color image dataset and the ship multispectral image dataset into a multimodal contrastive learning network in the basic ship vision-language model to obtain image features and a contrastive loss output by the multimodal contrastive learning network; the image features are obtained by the multimodal contrastive learning network performing image enhancement processing and similarity measurement on the ship full-color image dataset and the ship multispectral image dataset; inputting the image features and the label texts into a cross-modal semantic fusion network in the basic ship vision-language model to obtain fusion features and a fusion loss output by the cross-modal semantic fusion network; wherein, the fusion features are obtained by the cross-modal semantic fusion network performing alignment fusion on the image features and the label texts; inputting the fusion features into a hash learning network in the basic ship vision-language model to obtain a dataset hash code and a hash loss output by the hash learning network; wherein, the dataset hash code is obtained by the hash learning network performing hash learning on the fusion features; determining the objective loss function according to the contrastive loss, the fusion loss, and the hash loss.
[0044] In a specific embodiment, the ship full-color image dataset and the ship multispectral image dataset are input into a multimodal contrastive learning network in the basic ship vision-language model to obtain image features and a contrastive loss output by the multimodal contrastive learning network.
[0045] In this step, the image features are obtained by the multimodal contrastive learning network performing image enhancement processing and similarity measurement on the ship full-color image dataset and the ship multispectral image dataset.
[0046] The multimodal contrastive learning network includes a siamese image encoder for extracting high-dimensional image embeddings of the ship full-color image dataset and the ship multispectral image dataset, a similarity measurement module, and a contrastive loss applied to the image embeddings in the ship full-color image dataset and the ship multispectral image dataset. The siamese image encoder includes a full-color vision branch and a multispectral vision branch. The multimodal contrastive learning network aims to reduce the distance between similar image embeddings within the same modality and between different modalities, effectively align the feature representations, and at the same time bridge the heterogeneity gap between modalities.
[0047] Specifically, first, the multimodal image dataset is formally defined as , where N represents the total number of paired samples of label texts in the multimodal dataset, and each pair consists of a full-color image in the ship full-color image dataset and a multispectral image in the corresponding ship multispectral image dataset. Specifically, represents the Panchromatic image, represents the height of the panchromatic image, represents the width of the panchromatic image, represents the number of channels of the panchromatic image. Among them, represents a single-band grayscale image. Similarly, represents the th multispectral image in the ship multispectral image dataset, represents the number of spectral bands. represents the label text corresponding to the th paired sample. Among them, represents a set of predefined semantic categories for panchromatic images and multispectral images. Among them, the set of predefined language categories can be, for example, destroyers, tankers, etc. This embodiment does not limit this. To improve the robustness of feature extraction, the ship panchromatic image dataset and the ship multispectral image dataset are input into the multimodal contrast learning network in the basic ship vision language model. Through the multimodal contrast learning network, data augmentation is performed on all the panchromatic images in the ship panchromatic image dataset and all the multispectral images in the ship multispectral image dataset, and the enhanced ship panchromatic image enhanced dataset and the ship multispectral image enhanced dataset are obtained.
[0048] Exemplarily, the ship panchromatic image enhanced dataset , the ship multispectral image enhanced dataset . Among them, represents including common enhancement techniques such as random cropping and flipping. This embodiment does not limit this. After the enhanced ship panchromatic image enhanced dataset is processed through the panchromatic branch of the siamese image encoder, and the panchromatic branch uses the shared backbone network , so as to obtain the panchromatic image features in the image features . Among them, , represents the dimension of the feature space. Similarly, after the enhanced ship multispectral image enhanced dataset is processed through the multispectral branch of the siamese image encoder, and the multispectral branch uses the shared backbone network , so as to obtain the multispectral image features in the image features . Among them, , and during the processing, the weights of the panchromatic branch and the weights of the multispectral branch Initialize.
[0049] However, during the feature embedding process, it is necessary to determine the similarity of features for feature embedding. Specifically, the similarity measurement module is responsible for quantifying the similarity of feature embeddings within the same modality (intra-modal) or different modalities (inter-modal). The similarity measurement module supports using cosine similarity or Pearson correlation coefficient to measure the relationship between feature vectors, which is not limited in this embodiment.
[0050] Specifically, the calculation process of similarity measurement is as follows. For example, given the feature embedding from the panchromatic branch and from the multispectral branch , then for the first method, use cosine similarity to calculate the similarity between the two feature embeddings and . The similarity is shown in formula (1).
[0051] (1)
[0052] In formula (1), represents the dot product of and , represents the Euclidean norm of , represents the Euclidean norm of .
[0053] For the second method, use the standardized Pearson correlation coefficient to calculate the similarity between the two feature embeddings and . The similarity is shown in formula (2).
[0054] (2)
[0055] In formula (2), represents the covariance between and , represents the standard deviation of , represents the standard deviation of . The Pearson correlation coefficient is used to capture the linear relationship between features, making it particularly suitable for aligning multi-modal representations, especially when there are modality-specific scale and variance differences, which is not limited in this embodiment.
[0056] Subsequently, the contrastive loss is applied to maximize the similarity of similar samples within and between modalities, thereby reducing the variance within the same class and the differences between cross-modal classes.
[0057] Exemplarily, the calculation of the contrastive loss is as follows. To learn discriminative and aligned high-dimensional embeddings within the same modality, a contrastive loss function is adopted. For a batch of paired samples, where represents the number of different categories in the ship panchromatic image dataset and the ship multispectral image dataset, let represent the panchromatic embedding, , represent the multispectral embedding, a batch of one modality contains two groups of samples: the first group contains images, each image representing a unique category, and the second group contains images, each image representing an image of the same category as the first group and paired in the same order. In other words, within the same modality, the th sample and the th sample belong to the same category.
[0058] The within-modality loss maximizes the similarity of samples of the same category within a single modality. The within-panchromatic-modality loss is calculated as shown in Equation (3).
[0059] (3)
[0060] In Equation (3), represents that it can be instantiated with image features that can use cosine similarity or image features of the standardized Pearson correlation coefficient, represents the temperature parameter, which is used to control the steepness of the similarity distribution.
[0061] Therefore, for the multispectral modality, the within-multispectral-modality loss is calculated as shown in Equation (4).
[0062] (4)
[0063] In summary, the total within-modality loss is obtained as the sum of the within-panchromatic-modality loss and the within-multispectral-modality loss, that is, the total within-modality loss .
[0064] After calculating the total within-modality loss, to learn high-dimensional embeddings aligned between the panchromatic modality and the multispectral modality, a cross-modal contrastive loss is further defined by utilizing four groups of multimodal features.
[0065] For the cross-modal contrastive loss, first, for a batch of paired samples, let , , and respectively represent two sets of panchromatic features, and let and , and respectively represent two sets of multispectral features. The features at the same index and belong to the same class. For example, and may belong to the same class. The cross-modal contrast loss is calculated based on the comparison of four pairs of cross-modal features. For example, the first contrast loss is calculated as shown in formula (5).
[0066] (5)
[0067] Similarly, the second contrast loss , the third contrast loss and the fourth contrast loss are also calculated in the same way as formula (5). Thus, according to the first contrast loss , the second contrast loss , the third contrast loss and the fourth contrast loss , the total cross-modal contrast loss is determined. The total cross-modal contrast loss promotes the embedding alignment of samples of the same class between the two modalities, while reducing the similarity between samples of different classes between the modalities.
[0068] Finally, as Figure 2 shown, the contrast loss between high-dimensional image embeddings is determined as the sum of the total intra-modal loss and the total cross-modal contrast loss , that is, the contrast loss .
[0069] In a specific embodiment, the image features and label text are input into the cross-modal semantic fusion network in the basic ship vision-language model to obtain the fusion features and fusion loss output by the cross-modal semantic fusion network.
[0070] In this step, the fusion features are obtained by the cross-modal semantic fusion network through aligning and fusing the image features and label text. The cross-modal semantic fusion network is a network for enhancing the fusion of high-level semantic information of label text.
[0071] Specifically, the image features and label texts are input into the cross-modal semantic fusion network in the basic ship vision-language model to obtain the fusion features and fusion loss output by the cross-modal semantic fusion network. First, a pre-trained text editor is used to extract text features offline. For example, for each ship label text , the ship label text is extended to a descriptive text prompt, such as {a remote sensing image of , a type of ship}, and the descriptive text prompt is encoded into a high-level semantic embedding . The high-level semantic embedding is then fused with the image features through a semantic embedding fusion module.
[0072] For fusion, it is usually based on alignment fusion and attention-based fusion methods, and this embodiment does not limit this.
[0073] For image and text feature fusion based on the alignment fusion method, Figure 3 is a schematic diagram of the alignment fusion method provided by the present invention. As Figure 3 shown, the frozen text features and the image (panchromatic or multispectral) features are aligned through a contrastive loss mechanism, and the accuracy is Ground-Trunth to enforce semantic correspondence between modalities. The alignment ensures that text and image pairs with similar semantics are closer in the embedding space, while dissimilar pairs are pushed further apart. As Figure 3 shown, each value in the similarity matrix . Among them, represents the similarity measure of the image features (for example, it can be determined by cosine similarity or normalized Pearson correlation coefficient). The alignment is achieved through the fusion loss . The calculation of the fusion loss is shown in formula (6).
[0074] (6)
[0075] The advantage of such a setting is to ensure the alignment of text and image features in the semantic space. Through alignment-based fusion, the image features of different modalities are not only enhanced with high-level semantic information but also encouraged to be consistent with their corresponding text descriptions. This effectively bridges the gap between modalities and enhances the model's ability to associate semantically related visual and text features, which is very important for interpretable cross-modal retrieval.
[0076] Specifically, for the fusion method of the attention mechanism, Figure 4It is a schematic diagram of the attention mechanism fusion method provided by the present invention, as shown in Figure 4 In the attention-based fusion mechanism, text features are integrated into the image features of the image. The image features can be panchromatic image features from the panchromatic modality or multispectral image features from the multispectral modality . To introduce the attention mechanism fusion, frozen text features are used as semantic anchors. The fusion process is described as follows: Let represent the panchromatic image features, represent the multispectral image features, where , represents the batch size, is the dimension of the features. Let represent the pre-computed text features, represent the number of categories in the dataset. The panchromatic image features or multispectral image features in the input image features are first normalized to obtain the normalized panchromatic normalized image features and the multispectral normalized image features . The calculation process of the panchromatic normalized image features is shown in Equation (7). The calculation process of the multispectral normalized image features is shown in Equation (8).
[0077] (7)
[0078] (8)
[0079] After obtaining the normalized panchromatic normalized image features and the multispectral normalized image features , the text features are also normalized to obtain the normalized text features . The calculation process of the normalized text features is shown in Equation (9).
[0080] (9)
[0081] Then, the text-guided attention mechanism is applied to calculate the correlation between the image features and the text features. Calculate the attention value between the panchromatic normalized image features or the multispectral normalized image features and the normalized text features . The calculation formula of the attention value A is shown in Equation (10).
[0082] (10)
[0083] In formula (10), the attention value is the attention value matrix, , denotes or .
[0084] Furthermore, an activation function is introduced, which can be, for example, softmax. The attention weights are calculated through the softmax operation, , . After determining the attention weights , matrix multiplication is performed according to the attention weights and the text features to obtain the semantic features by calculating the weighted sum of the text features and the attention weights . The semantic features , . Finally, the semantic features and the image features of the original input are input into the semantic embedding fusion module for feature splicing to obtain the fusion features , and the fusion features , , denote the panchromatic image features or the multispectral image features .
[0085] The advantage of such a setting is that the attention-based fusion mechanism not only promotes the integration of higher-level semantic information from the label text, but also enables the cross-modal semantic fusion network to adaptively adjust and focus on the ship area most relevant to the text. Combining with the alignment-based fusion helps reduce the background noise in the multimodal ship images. Therefore, the attention-based fusion not only enhances the semantic understanding of the images, but also reduces the influence of irrelevant background elements in some cases.
[0086] In a specific embodiment, the fusion features are input into the hash learning network in the basic ship vision-language model to obtain the dataset hash code and the hash loss output by the hash learning network. The target loss function is determined according to the contrast loss, the fusion loss, and the hash loss.
[0087] Among them, the dataset hash code is obtained by the hash learning network performing hash learning on the fusion features.
[0088] In a specific embodiment, the fused feature is input into the hash learning network in the basic ship vision - language model to obtain the dataset hash code output by the hash learning network, including: inputting the fused feature into the hash learning layer in the hash learning network to obtain the hash feature output by the hash learning layer; wherein, the hash learning layer is a layer that processes the fused feature based on a non - linear activation function; inputting the hash feature into the binarization layer in the hash learning network to obtain the dataset hash code output by the binarization layer; wherein, the binarization layer is a layer that performs binary processing on the hash feature based on the sign function.
[0089] Specifically, after obtaining the fused feature the fused feature is input into the hash learning network in the basic ship vision - language model, and hash learning is performed through the hash learning layer of the hash learning network, so as to obtain the compact hash feature output by the hash learning layer . The hash feature wherein, represents normalization processing. represents a fully - connected network with a non - linear activation function, ensuring that the value of the dataset hash code falls within the range [-1, 1], represents the length of the dataset hash code.
[0090] Furthermore, the hash feature is trained using the same contrast loss as in the multi - modal contrast learning network. This ensures that semantically similar samples are closer in the hash space, while semantically dissimilar samples are farther apart. This design maintains the semantic relationship between different modalities while compressing features into a low - dimensional space, making the retrieval task efficient.
[0091] Among them, the obtained hash loss consists of two parts, the hash loss wherein, represents the intra - modal alignment loss, represents the inter - modal alignment loss.
[0092] After determining the contrast loss , the fusion loss and the hash loss , the target loss function is determined according to the contrast loss , the fusion loss and the hash loss . The target loss function wherein, is a hyper - parameter, and the hyper - parameter is used to control the relative importance of the hash loss in the overall objective, and the default value is set to 1.
[0093] The advantage of such a setting is that by jointly optimizing the contrastive loss , the fusion loss and the hash loss , a ship visual language model is obtained, enabling the ship visual language model to learn semantically rich and aligned representations between the panchromatic modality and the multispectral modality, and at the same time encoding these representations into a discriminative and computationally efficient hash space, which is suitable for large-scale retrieval tasks.
[0094] Furthermore, during the hash retrieval process, the learned hash features are input into the binarization layer in the hash learning network, and the hash features are binarized based on the binarization layer using the sign function, thereby obtaining the dataset hash code output by the binarization layer. The dataset hash code is a binary hash code, and the dataset hash code , where the function maps each element in as shown in the following formula (11).
[0095] (11)
[0096] Finally, the basic ship visual language model is further optimized through the dataset hash code and the target loss function, thereby obtaining the ship visual language model.
[0097] After obtaining the ship visual language model, the ship image to be retrieved and all ship sample images in the ship sample database are respectively input into the ship visual language model. Image features are extracted through the multi-modal contrastive learning network of the ship visual language model, and then the image features and text features are fused through the cross-modal semantic fusion network to obtain fused features. Finally, the fused features are input into the hash learning network, and hash learning is performed on the fused features through the hash learning network, thereby obtaining the hash code corresponding to the ship image to be retrieved output by the ship visual language model and all sample hash codes corresponding to all ship sample images in the ship sample database.
[0098] The advantage of such a setting is that by performing feature fusion and introducing the calculation of hash features at the same time, compact and robust hash codes are generated for subsequent efficient retrieval, realizing efficient and effective cross-modal ship image retrieval.
[0099] Step 103: Determine all Hamming similarities respectively according to the hash code and all sample hash codes, and determine the similarity ranking list according to all Hamming similarities.
[0100] In a specific embodiment, determining all Hamming similarities according to the hash code and all sample hash codes respectively includes: multiplying the hash code and all sample hash codes respectively to obtain all Hamming similarities.
[0101] Specifically, after obtaining the hash code corresponding to the ship image to be retrieved output by the ship vision-language model and all sample hash codes corresponding to all ship sample images in the ship sample database, respectively according to the hash code and each sample hash code product, all Hamming similarities are obtained. The Hamming similarity , the Hamming similarity value range is to . Among them, represents the length, and a larger value indicates a higher similarity between the corresponding hash codes. After obtaining all Hamming similarities, all Hamming similarities are sorted in ascending order to determine the similarity ranking list.
[0102] Step 104, determining the target retrieval result according to the ranking list of all Hamming similarities.
[0103] In this step, the target retrieval result is a set of ship sample images similar to the ship image to be retrieved selected based on all similarities in the ranking list.
[0104] In a specific embodiment, determining the target retrieval result according to the ranking list of all Hamming similarities includes: determining all candidate similarities according to all Hamming similarities in the ranking list and the similarity threshold; where all candidate similarities are Hamming similarities in the ranking list that are greater than the similarity threshold, and the similarity threshold is a preset threshold; determining all ship sample images corresponding to all candidate similarities as the target retrieval result.
[0105] Specifically, after obtaining the similarity ranking list, comparing all Hamming similarities in the similarity ranking list with the similarity threshold, determining the Hamming similarities greater than the similarity threshold among all Hamming similarities as candidate similarities, and determining the ship sample images in the ship sample database corresponding to the candidate similarities as the target retrieval result.
[0106] Exemplarily, obtain the ship image to be retrieved and the ship sample database; input the ship image to be retrieved and all ship sample images in the ship sample database into the ship vision-language model respectively, to obtain the hash code corresponding to the ship image to be retrieved output by the ship vision-language model and all sample hash codes corresponding to all ship sample images in the ship sample database; determine all Hamming similarities according to the hash code and all sample hash codes respectively, and determine the similarity ranking list according to all Hamming similarities; determine the target retrieval result according to the ranking list of all Hamming similarities.
[0107] Table 1 Experimental environment configuration
[0108]
[0109] The retrieval of the target retrieval result through the ship vision-language model can be verified through the following embodiments. All experimental verification environments are carried out under the same environment. The detailed information of the experimental environment configuration is shown in Table 1, and this embodiment does not limit this.
[0110] Figure 5 is the schematic diagram of the image sample provided by the present invention, as Figure 5 shown, the ship image samples used in the experimental data set are some publicly available ship images, such as Destroyers image samples, Littoral combatships image samples, Combat boats image samples, Bulk carriers image samples, Container ships image samples, and Oil tankers image samples, etc. This embodiment does not limit this.
[0111] The experiment is carried out on the only publicly available cross-modal remote sensing ship image retrieval dataset MRSSID. The cross-modal remote sensing ship image retrieval dataset MRSSID includes 2632 pairs of panchromatic and multispectral ship images obtained by GF-2 multispectral sensors and GF-2 panchromatic sensors. Each pair of ship image slices represents a combination of panchromatic and multispectral images captured simultaneously in the same area. In the experiment, the spatial resolution of the multispectral image is 4 meters, and the number of spectral channels is 4. The spatial resolution of the panchromatic image is 1 meter, and the number of spectral channels is only 1. This embodiment does not limit this.
[0112] During the experimental setup process, for the data augmentation part, the data augmentation strategy includes random cropping and horizontal flipping. All images are first scaled to 256×256 pixels and randomly cropped to 224×224 pixels. In order to utilize the knowledge of multimodal data, the ship vision-language model is used as the backbone of the siamese image encoder.
[0113] Specifically, the Contrastive Language-Image Pre-Training (CLIP)-ResNet-50 is adopted as the siamese image encoder, and the pre-trained language model CLIP-Transformer in the ship visual language model is used as the text encoder. This model is pre-trained on 400M general image-text pairs. To address the problem of unbalanced classes, a class-balanced resampling method is adopted. Meanwhile, the default number of training epochs is set to 50, the default setting of γ = 1, and the learning parameters are initialized to 0.07. In the proposed ship visual language model, the Adam optimizer with decay rates β1 = 0.9 and β2 = 0.999 is used for model optimization. The initial learning rate of the siamese image encoder is set to , while the initial learning rate of the hash learning network is set to .
[0114] In the first training epoch, a learning rate warm-up strategy is adopted, where the learning rate linearly increases from 0 to the predefined initial value.
[0115] After the warm-up stage, a cosine learning rate strategy without restart is applied to the remaining training epochs, gradually reducing the learning rate to improve convergence and model performance.
[0116] The final experimental results are shown in Table 2 as follows.
[0117] Table 2
[0118]
[0119] In Table 2, FSISR represents the latest state-of-the-art (SOTA) cross-modal hashing method. To ensure a fair comparison with the latest SOTA method, the backbone of FSISR is also replaced, and the modified version is denoted as FSISR-r. The comparative experimental results of the proposed ship visual language model and the SOTA method are shown in Table 2. Table 2 provides the mean average precision (mAP) values for two cross-modal retrieval tasks: PAN→MS (panchromatic to multispectral) and MS→PAN (multispectral to panchromatic), which are the averages for various hash code lengths (16, 24, 32, 48, and 64 bits). Table 2 also shows the experimental results of three variants of the proposed ship visual language model: SFCH-ali, which uses alignment-based fusion in the semantic embedding fusion module; SFCH-att, which utilizes attention-based fusion; and SFCH-fus, which combines these two strategies.
[0120] As can be seen from Table 2, the proposed ship visual language model achieved better performance on both tasks. In particular, the ship visual language model outperformed FSISR, the latest SOTA method, and FSISR-r with the same backbone for a fair comparison. Specifically, SFCH-att improved by 3.45% on average compared to FSISR-r, and SFCH-fus improved by 3.13% on average compared to FSISR-r, further verifying the superiority of the proposed ship visual language model. The superior performance of the ship visual language model can be attributed to the integration of its multi-modal contrastive learning network, hash learning network, and cross-modal semantic fusion network. The multi-modal contrastive learning network and hash learning network effectively align high-dimensional image embeddings and hash features of different modalities by optimizing their representations in the shared latent space, ensuring that the features are compact and semantically consistent. In addition, the cross-modal semantic fusion network utilizes high-level semantic information to achieve a more comprehensive and adaptive integration of complementary features. These results indicate that the proposed ship visual language model effectively bridges the gap between modalities and achieves excellent retrieval performance at a series of hash code lengths using semantic fusion and contrastive learning strategies.
[0121] Table 2 also shows the experimental results of three variants of the proposed ship visual language model: SFCH-ali, SFCH-att, and SFCH-fus. The results indicate that SFCH-ali, which only uses alignment-based fusion, achieved relatively lower performance compared to the other two variants. In contrast, SFCH-att and SFCH-fus, which incorporate attention-based fusion mechanisms, showed better performance. SFCH-att and SFCH-fus were almost comparable in terms of accuracy. However, SFCH-att lacks supervised alignment with semantic information, which limits its interpretability at the semantic level. In contrast, SFCH-fus, which combines alignment-based and attention-based fusion strategies, achieved semantic-level interpretability without sacrificing accuracy.
[0122] A method for retrieving ship images provided by the present invention includes obtaining a ship image to be retrieved and a ship sample database. The ship sample database contains at least two ship sample images, and the ship sample images are ship images in different modalities from the ship image to be retrieved. The ship image to be retrieved and all ship sample images in the ship sample database are respectively input into a ship visual language model to obtain a hash code corresponding to the ship image to be retrieved and all sample hash codes corresponding to all ship sample images in the ship sample database output by the ship visual language model. The ship visual language model is trained based on a multi-modal ship sample pattern data set including a ship panchromatic image data set, a ship multi-spectral image data set, and label texts. The ship visual language model is a model that processes the ship image to be retrieved and the ship sample database based on multi-modal contrast learning, cross-modal semantic fusion, and hash learning to obtain corresponding hash codes and sample hash codes. All Hamming similarities are determined respectively according to the hash code and all sample hash codes, and a similarity ranking list is determined according to all Hamming similarities. A target retrieval result is determined according to the ranking list of all Hamming similarities. The target retrieval result is a set of ship sample images similar to the ship image to be retrieved selected based on all similarities in the ranking list. In the above embodiment, the technical solution of the present invention is used to solve the defect that the accuracy of cross-modal retrieval of remote sensing images decreases due to pure text-to-image retrieval or image-to-image retrieval in the prior art. It realizes multi-modal contrast learning, cross-modal semantic fusion, and hash learning processing of the ship image to be retrieved and the ship sample database by using a ship visual language model combined with images and texts trained based on a ship panchromatic image data set, a ship multi-spectral image data set, and label texts. Finally, the target retrieval result is determined based on the obtained hash code and sample hash code, improving the accuracy of cross-modal retrieval of the ship image to be retrieved in the ship sample database.
[0123] The ship image retrieval device provided by the present invention will be described below. The ship image retrieval device described below can be correspondingly referred to the ship image retrieval method described above.
[0124] Figure 6 is a schematic structural diagram of the ship image retrieval device provided by the present invention. Refer to Figure 6 As shown, the ship image retrieval device 600 includes an image acquisition module 601, a hash code output module 602, a list determination module 603, and a result determination module 604.
[0125] The image acquisition module 601 is used to obtain a ship image to be retrieved and a ship sample database. The ship sample database contains at least two ship sample images, and the ship sample images are ship images in different modalities from the ship image to be retrieved.
[0126] The hash code output module 602 is configured to input the ship image to be retrieved and all ship sample images in the ship sample database into the ship vision-language model respectively, so as to obtain the hash code corresponding to the ship image to be retrieved output by the ship vision-language model and all sample hash codes corresponding to all ship sample images in the ship sample database. The ship vision-language model is trained based on the obtained multi-modal ship sample pattern dataset including the ship panchromatic image dataset, the ship multi-spectral image dataset, and the label text. The ship vision-language model is a model that processes the ship image to be retrieved and the ship sample database based on multi-modal contrast learning, cross-modal semantic fusion, and hash learning to obtain the corresponding hash code and sample hash code.
[0127] The list determination module 603 is configured to determine all Hamming similarities respectively according to the hash code and all sample hash codes, and determine a similarity ranking list according to all Hamming similarities.
[0128] The result determination module 604 is configured to determine the target retrieval result according to the ranking list of all Hamming similarities. The target retrieval result is a set of ship sample images similar to the ship image to be retrieved selected based on all similarities in the ranking list.
[0129] In an exemplary embodiment, the device further includes a model training module. The model training module is configured to: obtain a multi-modal ship sample pattern dataset. The multi-modal ship sample pattern dataset includes a ship panchromatic image dataset, a ship multi-spectral image dataset, and label text. The label text is the paired label text corresponding to the paired samples determined by pairing the ship panchromatic image dataset and the ship multi-spectral image dataset. Input the ship panchromatic image dataset, the ship multi-spectral image dataset, and the label text into the basic ship vision-language model to obtain the dataset hash code and the target loss function output by the basic ship vision-language model. The basic ship vision-language model is a pre-determined basic model for training. Optimize the basic ship vision-language model according to the dataset hash code and the target loss function to obtain the ship vision-language model.
[0130] In an exemplary embodiment, the model training module inputs the full-color ship image dataset, the multi-spectral ship image dataset, and the label text into the basic ship vision-language model to obtain the dataset hash code and the target loss function output by the basic ship vision-language model. Specifically, it is used to: input the full-color ship image dataset and the multi-spectral ship image dataset into the multi-modal contrastive learning network in the basic ship vision-language model to obtain the image features and the contrastive loss output by the multi-modal contrastive learning network; the image features are obtained by the multi-modal contrastive learning network through image enhancement processing and similarity measurement of the full-color ship image dataset and the multi-spectral ship image dataset; input the image features and the label text into the cross-modal semantic fusion network in the basic ship vision-language model to obtain the fusion features and the fusion loss output by the cross-modal semantic fusion network; wherein, the fusion features are obtained by the cross-modal semantic fusion network through alignment and fusion of the image features and the label text; input the fusion features into the hash learning network in the basic ship vision-language model to obtain the dataset hash code and the hash loss output by the hash learning network; wherein, the dataset hash code is obtained by the hash learning network through hash learning of the fusion features; determine the target loss function according to the contrastive loss, the fusion loss, and the hash loss.
[0131] In an exemplary embodiment, the model training module inputs the fusion features into the hash learning network in the basic ship vision-language model to obtain the dataset hash code output by the hash learning network. Specifically, it is used to: input the fusion features into the hash learning layer in the hash learning network to obtain the hash features output by the hash learning layer; wherein, the hash learning layer is a layer that processes the fusion features based on a non-linear activation function; input the hash features into the binarization layer in the hash learning network to obtain the dataset hash code output by the binarization layer; wherein, the binarization layer is a layer that performs binary processing on the hash features based on the sign function.
[0132] In an exemplary embodiment, the list determination module 603 determines all Hamming similarities according to the hash code and all sample hash codes respectively. Specifically, it is used to: multiply the hash code and all sample hash codes respectively to obtain all Hamming similarities.
[0133] In an exemplary embodiment, the result determination module 604 is specifically used to: determine all candidate similarities according to all Hamming similarities and the similarity threshold in the ranking list; wherein, all candidate similarities are the Hamming similarities in the ranking list that are greater than the similarity threshold, and the similarity threshold is a preset threshold; determine the target retrieval results corresponding to all candidate similarities and all ship sample images.
[0134] The device of this embodiment can be used to execute the method of any embodiment in the method embodiments of ship image retrieval. The specific implementation process and technical effects are similar to those in the method embodiments of ship image retrieval. For details, please refer to the detailed introduction in the method embodiments of ship image retrieval, which will not be elaborated here.
[0135] Figure 7 is a schematic structural diagram of an electronic device provided by the present invention. As Figure 7 shown, the electronic device may include: a processor 710, a communication interface 720, a memory 730, and a communication bus 740. Among them, the processor 710, the communication interface 720, and the memory 730 communicate with each other through the communication bus 740. The processor 710 can call the logical instructions in the memory 730 to execute the ship image retrieval method, and the method includes: obtaining a ship image to be retrieved and a ship sample database; wherein, the ship sample database contains at least two ship sample images, and the ship sample images are ship images in different modalities from the ship image to be retrieved; inputting the ship image to be retrieved and all the ship sample images in the ship sample database into a ship visual language model respectively to obtain the hash code corresponding to the ship image to be retrieved and all the sample hash codes corresponding to all the ship sample images in the ship sample database output by the ship visual language model; wherein, the ship visual language model is trained based on the obtained multi-modal ship sample pattern data set including a ship panchromatic image data set, a ship multi-spectral image data set, and label texts, and the ship visual language model is a model that processes the ship image to be retrieved and the ship sample database based on multi-modal contrast learning, cross-modal semantic fusion, and hash learning to obtain the corresponding hash code and sample hash code; determining all Hamming similarities respectively according to the hash code and all the sample hash codes, and determining a similarity ranking list according to all the Hamming similarities; determining a target retrieval result according to the ranking list of all the Hamming similarities; wherein, the target retrieval result is a set of ship sample images similar to the ship image to be retrieved selected based on all the similarities in the ranking list.
[0136] In addition, when the logical instructions in the above-mentioned memory 730 are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or a part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods of various embodiments of the present invention. The aforementioned storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memories (ROM, Read-Only Memory), random access memories (RAM, Random Access Memory), magnetic disks, or optical discs that can store program codes.
[0137] On the other hand, the present invention also provides a computer program product. The computer program product includes a computer program. The computer program can be stored on a non-transitory computer-readable storage medium. When the computer program is executed by a processor, the computer can execute the method for retrieving ship images provided by the above-mentioned various methods. The method includes: obtaining a ship image to be retrieved and a ship sample database; wherein, the ship sample database contains at least two ship sample images, and the ship sample images are ship images in different modalities from the ship image to be retrieved; inputting the ship image to be retrieved and all the ship sample images in the ship sample database into a ship visual language model respectively to obtain the hash code corresponding to the ship image to be retrieved output by the ship visual language model and all the sample hash codes corresponding to all the ship sample images in the ship sample database; wherein, the ship visual language model is trained based on the obtained multi-modal ship sample pattern data set including a ship panchromatic image data set, a ship multi-spectral image data set, and label texts, and the ship visual language model is a model that processes the ship image to be retrieved and the ship sample database based on multi-modal contrast learning, cross-modal semantic fusion, and hash learning to obtain the corresponding hash code and sample hash code; determining all Hamming similarities respectively according to the hash code and all the sample hash codes, and determining a similarity ranking list according to all the Hamming similarities; determining a target retrieval result according to the ranking list of all the Hamming similarities; wherein, the target retrieval result is a set of ship sample images similar to the ship image to be retrieved selected based on all the similarities in the ranking list.
[0138] In another aspect, the present invention also provides a non-transitory computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it implements the retrieval method of ship images provided by the above-mentioned various methods. The method includes: obtaining a ship image to be retrieved and a ship sample database; wherein, the ship sample database contains at least two ship sample images, and the ship sample images are ship images in different modalities from the ship image to be retrieved; inputting the ship image to be retrieved and all the ship sample images in the ship sample database into a ship vision-language model respectively, to obtain a hash code corresponding to the ship image to be retrieved and all sample hash codes corresponding to all the ship sample images in the ship sample database output by the ship vision-language model; wherein, the ship vision-language model is trained based on a multi-modal ship sample pattern data set including a ship panchromatic image data set, a ship multi-spectral image data set and label texts, and the ship vision-language model is a model that processes the ship image to be retrieved and the ship sample database based on multi-modal contrast learning, cross-modal semantic fusion and hash learning to obtain the corresponding hash code and sample hash codes; determining all Hamming similarities respectively according to the hash code and all the sample hash codes, and determining a similarity ranking list according to all the Hamming similarities; determining a target retrieval result according to the ranking list of all the Hamming similarities; wherein, the target retrieval result is a set of ship sample images similar to the ship image to be retrieved selected based on all the similarities in the ranking list.
[0139] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place, or may be distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment. Those of ordinary skill in the art can understand and implement it without creative efforts.
[0140] Through the description of the above embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus a necessary general hardware platform, and of course, also by hardware. Based on such an understanding, the essence of the above technical solution, or the part that contributes to the prior art, can be embodied in the form of a software product. The computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in each embodiment or some parts of the embodiments.
[0141] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements for some of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the various embodiments of the present invention.
Claims
1. A ship image retrieval method, characterized in that: include: Acquire a ship image to be retrieved and a ship sample database; wherein the ship sample database contains at least two ship sample images, and the ship sample images are ship images of different modalities from the ship image to be retrieved; The ship image to be retrieved and all the ship sample images in the ship sample database are respectively input into the ship visual language model, and the hash code corresponding to the ship image to be retrieved and all the sample hash codes corresponding to all the ship sample images in the ship sample database output by the ship visual language model are obtained; wherein the ship visual language model is obtained by training a basic ship visual language model based on an acquired multimodal ship sample pattern data set including a ship panchromatic image data set, a ship multispectral image data set and a label text to obtain a data set hash code and a target loss function, and the basic ship visual language model is optimized according to the data set hash code and the target loss function, and the ship visual language model is a model that processes the ship image to be retrieved and the ship sample database based on multimodal contrast learning, cross-modal semantic fusion and hash learning to obtain the corresponding hash code and the sample hash code; the basic ship visual language model is trained based on the acquired multimodal ship sample pattern data set including a ship panchromatic image data set, a ship multispectral image data set and a label text Obtaining a data set hash code and a target loss function, including: inputting the ship panchromatic image data set and the ship multispectral image data set into a multimodal contrast learning network in a basic ship visual language model to obtain image features and contrast losses output by the multimodal contrast learning network; the image features are obtained by the multimodal contrast learning network performing image enhancement processing and similarity measurement on the ship panchromatic image data set and the ship multispectral image data set; inputting the image features and the label text into a cross-modal semantic fusion network in the basic ship visual language model to obtain fusion features and fusion losses output by the cross-modal semantic fusion network; the fusion features are obtained by the cross-modal semantic fusion network performing alignment and fusion on the image features and the label text; inputting the fusion features into a hash learning network in the basic ship visual language model to obtain the data set hash code and hash loss output by the hash learning network; the data set hash code is obtained by the hash learning network performing hash learning on the fusion features; determining the target loss function according to the contrast loss, the fusion loss and the hash loss; Determine all Hamming similarities according to the hash code and all the sample hash codes respectively, and determine a similarity ranking list according to all the Hamming similarities; The target retrieval result is determined according to the ranking list of all the Hamming similarities; wherein the target retrieval result is a set of the sample ship images similar to the ship image to be retrieved, selected based on all the similarities in the ranking list.
2. The ship image retrieval method according to claim 1, characterized in that: The ship visual language model is trained based on the following steps: Acquire the multimodal ship sample pattern dataset; wherein the multimodal ship sample pattern dataset includes a ship panchromatic image dataset, a ship multispectral image dataset and a label text, and the label text is a paired label text corresponding to a paired sample determined by pairing the ship panchromatic image dataset and the ship multispectral image dataset; Inputting the ship panchromatic image dataset, the ship multispectral image dataset and the label text into a basic ship visual language model to obtain a dataset hash code and a target loss function output by the basic ship visual language model; wherein the basic ship visual language model is a predetermined basic model for training; The basic ship visual language model is optimized according to the data set hash code and the target loss function to obtain the ship visual language model.
3. The ship image retrieval method according to claim 2, characterized in that: Inputting the fusion feature into the hash learning network in the basic ship visual language model to obtain the data set hash code output by the hash learning network includes: Inputting the fused features into the hash learning layer in the hash learning network to obtain the hash features output by the hash learning layer; wherein the hash learning layer is a layer that processes the fused features based on a nonlinear activation function; The hash feature is input into a binarization layer in the hash learning network to obtain a hash code of the data set output by the binarization layer; wherein the binarization layer is a layer that performs binary processing on the hash feature based on a sign function.
4. The ship image retrieval method according to any one of claims 1 to 2, characterized in that: The determining all Hamming similarities according to the hash code and all the sample hash codes respectively includes: The hash code and all the sample hash codes are multiplied respectively to obtain all the Hamming similarities.
5. The ship image retrieval method according to any one of claims 1 to 3, characterized in that: The step of determining the target search result according to the ranking list of all the Hamming similarities comprises: Determine all candidate similarities according to all the Hamming similarities in the ranking list and the similarity threshold; wherein all the candidate similarities are the Hamming similarities in the ranking list whose Hamming similarities are greater than the similarity threshold, and the similarity threshold is a preset threshold; All the candidate similarities corresponding to all the ship sample images are determined as the target retrieval results.
6. A ship image retrieval device, characterized in that: include: An image acquisition module, used to acquire a ship image to be retrieved and a ship sample database; wherein the ship sample database contains at least two ship sample images, and the ship sample images are ship images of different modalities from the ship image to be retrieved; A hash code output module is used to input the ship image to be retrieved and all the ship sample images in the ship sample database into a ship visual language model respectively, and obtain the hash code corresponding to the ship image to be retrieved output by the ship visual language model and all sample hash codes corresponding to all the ship sample images in the ship sample database; wherein the ship visual language model is obtained by training a basic ship visual language model based on an acquired multimodal ship sample pattern data set including a ship panchromatic image data set, a ship multispectral image data set and a label text to obtain a data set hash code and a target loss function, and optimizing the basic ship visual language model according to the data set hash code and the target loss function; the ship visual language model is a model that processes the ship image to be retrieved and the ship sample database based on multimodal contrast learning, cross-modal semantic fusion and hash learning to obtain the corresponding hash code and sample hash code; the basic ship visual language model is trained based on an acquired multimodal ship sample pattern data set including a ship panchromatic image data set, a ship multispectral image data set and a label text Obtaining a data set hash code and a target loss function, including: inputting the ship panchromatic image data set and the ship multispectral image data set into a multimodal contrastive learning network in a basic ship visual language model, obtaining image features and contrast losses output by the multimodal contrastive learning network; the image features are obtained by the multimodal contrastive learning network performing image enhancement processing and similarity measurement on the ship panchromatic image data set and the ship multispectral image data set; inputting the image features and the label text into a cross-modal semantic fusion network in the basic ship visual language model, obtaining fusion features and fusion losses output by the cross-modal semantic fusion network; wherein the fusion features are obtained by the cross-modal semantic fusion network aligning and fusing the image features and the label text; inputting the fusion features into a hash learning network in the basic ship visual language model, obtaining the data set hash code and hash loss output by the hash learning network; wherein the data set hash code is obtained by the hash learning network performing hash learning on the fusion features; determining the target loss function according to the contrast loss, the fusion loss and the hash loss; A list determination module, used to determine all Hamming similarities according to the hash code and all the sample hash codes respectively, and determine a similarity ranking list according to all the Hamming similarities; A result determination module is used to determine a target retrieval result according to a ranking list of all the Hamming similarities; wherein the target retrieval result is a set of sample ship images similar to the ship image to be retrieved, selected based on all the similarities in the ranking list.
7. An electronic device comprising a memory, a processor, and a computer program stored in the memory and running on the processor, characterized in that: When the processor executes the computer program, the method for retrieving the ship image according to any one of claims 1 to 5 is implemented.
8. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the method for retrieving a ship image according to any one of claims 1 to 5 is implemented.
9. A computer program product, comprising a computer program, characterized in that When the computer program is executed by a processor, the method for retrieving a ship image according to any one of claims 1 to 5 is implemented.
Citation Information
Patent Citations
Depth cross-modal hash retrieval method and device and electronic equipment
CN115757711A
Cross-modal hash retrieval method and device based on multiple comparison and two-way confrontation
CN116431847A