Information retrieval method and device, computer equipment, readable storage medium and program product
By establishing sample pairs of text and image data, and using the cross-modal information retrieval model to generate fusion features, optimizing the model to improve the search accuracy, the problem of difficult-to-bridge semantic alignment and feature gap between modals in the prior art is solved, and more efficient cross-modal retrieval is achieved.
Patent Information
- Application Number
- CN202510001751.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-02
- Publication Date
- 2025-05-16
AI Technical Summary
The prior art has shortcomings in semantic alignment between modals and the bridging of feature divides, resulting in unsatisfactory cross-modal data retrieval accuracy.
By establishing a sample pair of text data and image data, using the text information extraction model, image information extraction model and cross-attention module in the cross-modal information retrieval model, text fusion features and image fusion features are generated, and the trained cross-modal information retrieval model is obtained through the adaptive triple loss function and quantized loss function optimization model.
The accuracy of cross-modal retrieval is improved, and the gap in features is bridged by aligning the semantics between the modals across the attention fusion module and the adaptive triple loss function.
Smart Images

Figure CN120011581A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of deep learning technology, and in particular to an information retrieval method, apparatus, computer device, computer-readable storage medium, and computer program product. Background Art
[0002] With the rapid growth of multimedia data, massive multimodal data such as images, videos and texts are rapidly accumulating in various applications. Traditional information retrieval methods are usually limited to single-modal information queries, while practical applications often require retrieving relevant content from one modality (such as images) to another modality (such as text).
[0003] However, existing methods still have shortcomings in semantic alignment between modalities and bridging the feature gap. Especially when dealing with complex and diverse cross-modal data, the retrieval accuracy is often not ideal. Summary of the invention
[0004] Based on this, it is necessary to provide an information retrieval method, apparatus, computer device, computer-readable storage medium and computer program product that can improve retrieval accuracy in response to the above technical problems.
[0005] In a first aspect, the present application provides an information retrieval method, comprising:
[0006] Acquire text data and image data, and establish multiple sample pairs based on the text data and the image data;
[0007] For each sample pair, a text fusion feature and an image fusion feature are generated through a text information extraction model, an image information extraction model and a cross-attention module in a cross-modal information retrieval model to be trained; the text fusion feature is converted into a text hash code through the text information extraction model; the image fusion feature is converted into an image hash code through the image information extraction model;
[0008] Based on an adaptive triple loss function and a quantization loss function, the total loss of the text hash code and the image hash code is determined; based on the total loss, the cross-modal information retrieval model to be trained is optimized to obtain a trained cross-modal information retrieval model; the trained cross-modal information retrieval model is used to retrieve images according to input text, or to retrieve text according to input images.
[0009] In one embodiment, the text information extraction model includes: a word segmentation layer and a text feature extraction block; the text feature extraction block includes a preset number of text extraction layers; the image information extraction model includes: a detection layer, a fully connected layer and an image feature extraction block; the image feature extraction block includes a preset number of image extraction layers; the preset number of text extraction layers and the preset number of image extraction layers correspond to each other in sequence; the text information extraction model, the image information extraction model and the cross-attention module in the cross-modal information retrieval model to be trained are used to generate text fusion features and image fusion features, including:
[0010] Input the text in the sample pair into the word segmentation layer to obtain word segmentation features; use the word segmentation features as the input of the text feature extraction block; for any text extraction layer, convert the text features of the previous text extraction layer of the current text extraction layer into a text query vector, a text key vector and a text value vector; generate a first text feature based on the text query vector, the text key vector and the text value vector; input the image in the sample pair into the detection layer to obtain a detection frame, and input the side frame into the fully connected layer to obtain an image feature vector; use the image feature vector as the input of the image feature extraction block; for the image extraction layer corresponding to the current text extraction layer, convert the image features of the previous image extraction layer into an image query vector, an image ... generate a first text feature based on the text query vector, the text key vector and the text value vector; generate a first text feature based on the text query vector, the text key vector and the text value vector; generate a first text feature based on the text key vector and image value vector; based on the image query vector, the image key vector and the image value vector, generate a first image feature; through a cross-attention fusion module, fuse the text query vector with the image key vector and the image value vector to obtain a second text feature, and fuse the image query vector with the text key vector and the text value vector to obtain a second image feature; based on the first text feature and the second text feature, determine the text feature of the current text extraction layer; based on the first image feature and the second image feature, determine the image feature of the current image extraction layer; wherein the text feature of the last layer of text extraction layer is the text fusion feature; the image feature of the last layer of image extraction layer is the image fusion feature.
[0011] In one embodiment, the text information extraction model further includes: a text hash layer; the image information extraction model further includes: an image hash layer; the converting the text fusion feature into a text hash code by the text information extraction model includes:
[0012] The text fusion feature is input into the text hash layer to obtain a text hash code; the image fusion feature is input into the image hash layer to obtain an image hash code.
[0013] In one embodiment, determining the total loss of the text hash code and the image hash code based on the adaptive triple loss function and the quantization loss function includes:
[0014] Based on the correlation, boundary parameter and weight factor of the image hash code and the text hash code, an adaptive triple loss of the image hash code and the text hash code is determined by an adaptive triple loss function; a quantization loss of the image hash code and the text hash code is determined by a quantization loss function; and a total loss of the text hash code and the image hash code is determined based on the adaptive triple loss and the quantization loss.
[0015] In one embodiment, the image data includes N images, and the text data includes N texts describing each image; and establishing a plurality of sample pairs based on the text data and the image data includes:
[0016] For each image, the current image is combined with N texts respectively to obtain N sample pairs corresponding to the current image.
[0017] In one embodiment, the trained cross-modal information retrieval model is used to retrieve images based on input text, including:
[0018] The text is input into the text information extraction model of the trained cross-modal information retrieval model to obtain a text hash code; the image hash codes of all images are generated through the image information extraction model of the trained cross-modal information retrieval model; the similarity between the text hash code and the hash codes of all images is calculated respectively, and the image corresponding to the image hash code with the highest similarity is taken as the result of image retrieval.
[0019] In a second aspect, the present application also provides an information retrieval device, comprising:
[0020] An acquisition module, used for acquiring text data and image data, and establishing a plurality of sample pairs based on the text data and the image data;
[0021] An extraction module is used to generate text fusion features and image fusion features for each sample pair through a text information extraction model, an image information extraction model and a cross-attention module in a cross-modal information retrieval model to be trained; convert the text fusion features into text hash codes through the text information extraction model; and convert the image fusion features into image hash codes through the image information extraction model;
[0022] The optimization module is used to determine the total loss of the text hash code and the image hash code based on an adaptive triple loss function and a quantization loss function; based on the total loss, optimize the cross-modal information retrieval model to be trained to obtain a trained cross-modal information retrieval model; the trained cross-modal information retrieval model is used to retrieve images according to input text, or to retrieve text according to input images.
[0023] In a third aspect, the present application further provides a computer device, including a memory and a processor, wherein the memory stores a computer program, and when the processor executes the computer program, the following steps are implemented:
[0024] Acquire text data and image data, and establish multiple sample pairs based on the text data and the image data;
[0025] For each sample pair, a text fusion feature and an image fusion feature are generated through a text information extraction model, an image information extraction model and a cross-attention module in a cross-modal information retrieval model to be trained; the text fusion feature is converted into a text hash code through the text information extraction model; the image fusion feature is converted into an image hash code through the image information extraction model;
[0026] Based on an adaptive triple loss function and a quantization loss function, the total loss of the text hash code and the image hash code is determined; based on the total loss, the cross-modal information retrieval model to be trained is optimized to obtain a trained cross-modal information retrieval model; the trained cross-modal information retrieval model is used to retrieve images according to input text, or to retrieve text according to input images.
[0027] In a fourth aspect, the present application further provides a computer-readable storage medium having a computer program stored thereon, wherein when the computer program is executed by a processor, the following steps are implemented:
[0028] Acquire text data and image data, and establish multiple sample pairs based on the text data and the image data;
[0029] For each sample pair, a text fusion feature and an image fusion feature are generated through a text information extraction model, an image information extraction model and a cross-attention module in a cross-modal information retrieval model to be trained; the text fusion feature is converted into a text hash code through the text information extraction model; the image fusion feature is converted into an image hash code through the image information extraction model;
[0030] Based on an adaptive triple loss function and a quantization loss function, the total loss of the text hash code and the image hash code is determined; based on the total loss, the cross-modal information retrieval model to be trained is optimized to obtain a trained cross-modal information retrieval model; the trained cross-modal information retrieval model is used to retrieve images according to input text, or to retrieve text according to input images.
[0031] In a fifth aspect, the present application further provides a computer program product, including a computer program, which implements the following steps when executed by a processor:
[0032] Acquire text data and image data, and establish multiple sample pairs based on the text data and the image data;
[0033] For each sample pair, a text fusion feature and an image fusion feature are generated through a text information extraction model, an image information extraction model and a cross-attention module in a cross-modal information retrieval model to be trained; the text fusion feature is converted into a text hash code through the text information extraction model; the image fusion feature is converted into an image hash code through the image information extraction model;
[0034] Based on an adaptive triple loss function and a quantization loss function, the total loss of the text hash code and the image hash code is determined; based on the total loss, the cross-modal information retrieval model to be trained is optimized to obtain a trained cross-modal information retrieval model; the trained cross-modal information retrieval model is used to retrieve images according to input text, or to retrieve text according to input images.
[0035] The above-mentioned information retrieval method, apparatus, computer device, computer-readable storage medium and computer program product obtain text data and image data, and establish multiple sample pairs based on the text data and the image data; for each sample pair, generate text fusion features and image fusion features through the text information extraction model, image information extraction model and cross-attention module in the cross-modal information retrieval model to be trained; convert the text fusion features into text hash codes through the text information extraction model; convert the image fusion features into image hash codes through the image information extraction model; determine the total loss of the text hash codes and the image hash codes based on the adaptive triple loss function and the quantization loss function; optimize the cross-modal information retrieval model to be trained based on the total loss to obtain a trained cross-modal information retrieval model; the trained cross-modal information retrieval model is used to retrieve images according to the input text, or to retrieve text according to the input image. The semantics between modalities are aligned and the gap in features is bridged through the cross-attention fusion module and the adaptive triple loss, thereby improving the accuracy of cross-modal retrieval. BRIEF DESCRIPTION OF THE DRAWINGS
[0036] In order to more clearly illustrate the technical solutions in the embodiments of the present application or related technologies, the drawings required for use in the embodiments of the present application or related technical descriptions will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other related drawings can be obtained based on these drawings without paying creative work.
[0037] Figure 1 is a flow chart of an information retrieval method in one embodiment;
[0038] Figure 2 is a network structure diagram of a cross-modal information retrieval model in one embodiment;
[0039] Figure 3 A retrieval result diagram of a text retrieval based on an input image in one embodiment;
[0040] Figure 4 A retrieval result diagram of retrieving images according to input text in one embodiment;
[0041] Figure 5 is a structural block diagram of an information retrieval device in one embodiment;
[0042] Figure 6 FIG. 4 is a diagram showing the internal structure of a computer device in one embodiment. DETAILED DESCRIPTION
[0043] In order to make the purpose, technical solution and advantages of the present application more clearly understood, the present application is further described in detail below in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.
[0044] In one embodiment, Figure 1 As shown, an information retrieval method is provided. This embodiment takes the method applied to a terminal as an example. It can be understood that the method can also be applied to a server, and can also be applied to a system including a terminal and a server, and is implemented through the interaction between the terminal and the server. In this embodiment, the method includes the following steps:
[0045] Step 102: Acquire text data and image data, and create multiple sample pairs based on the text data and the image data.
[0046] Each text in the text data corresponds to each image in the image data, and the text is used to describe the image. The sample pair is a text-image pair, and all texts and all images are combined to create a sample pair.
[0047] Exemplarily, text data and image data are obtained, all texts in the text data and all images in the image data are combined respectively, and a plurality of sample pairs are established.
[0048] Step 104, for each sample pair, generate text fusion features and image fusion features through the text information extraction model, image information extraction model and cross-attention module in the cross-modal information retrieval model to be trained; convert the text fusion features into text hash codes through the text information extraction model; convert the image fusion features into image hash codes through the image information extraction model.
[0049] Among them, when the text information extraction model extracts the text information in the sample pair and the image information extraction model extracts the image information in the sample pair, the cross-attention fusion module fuses the information extracted by the above two models and returns the fusion results to the above two models.
[0050] Among them, hash code is an encoding method that converts input data of any length into an output value of fixed length through a specific algorithm. Hash code is often used to quickly find, compare and store data.
[0051] Step 106: determine the total loss of the text hash code and the image hash code based on the adaptive triple loss function and the quantization loss function; optimize the cross-modal information retrieval model to be trained based on the total loss to obtain a trained cross-modal information retrieval model; the trained cross-modal information retrieval model is used to retrieve images according to the input text, or to retrieve text according to the input image.
[0052] Among them, when retrieving images through text, after entering the text, the input text is converted into a text hash code through the cross-modal information retrieval model, and then all images are converted into image hash codes. The similarity of the text hash code and the image hash code is calculated in turn, and they are sorted according to the similarity. The image corresponding to the image hash code most similar to the text hash code is the retrieval result; the method of retrieving text through images is the same as above.
[0053] The above-mentioned information retrieval method, apparatus, computer device, computer-readable storage medium and computer program product obtain text data and image data, and establish multiple sample pairs based on the text data and the image data; for each sample pair, generate text fusion features and image fusion features through the text information extraction model, image information extraction model and cross-attention module in the cross-modal information retrieval model to be trained; convert the text fusion features into text hash codes through the text information extraction model; convert the image fusion features into image hash codes through the image information extraction model; determine the total loss of the text hash codes and the image hash codes based on the adaptive triple loss function and the quantization loss function; optimize the cross-modal information retrieval model to be trained based on the total loss to obtain a trained cross-modal information retrieval model; the trained cross-modal information retrieval model is used to retrieve images according to the input text, or to retrieve text according to the input image. The semantics between modalities are aligned and the gap in features is bridged through the cross-attention fusion module and the adaptive triple loss, thereby improving the accuracy of cross-modal retrieval.
[0054] In an exemplary embodiment, the text information extraction model includes: a word segmentation layer and a text feature extraction block; the text feature extraction block includes a preset number of text extraction layers; the image information extraction model includes: a detection layer, a fully connected layer and an image feature extraction block; the image feature extraction block includes a preset number of image extraction layers; the preset number of text extraction layers and the preset number of image extraction layers correspond to each other in sequence; the text information extraction model, the image information extraction model and the cross-attention module in the cross-modal information retrieval model to be trained are used to generate text fusion features and image fusion features, including:
[0055] Input the text in the sample pair into the word segmentation layer to obtain word segmentation features; use the word segmentation features as the input of the text feature extraction block; for any text extraction layer, convert the text features of the previous text extraction layer of the current text extraction layer into a text query vector, a text key vector and a text value vector; generate a first text feature based on the text query vector, the text key vector and the text value vector; input the image in the sample pair into the detection layer to obtain a detection frame, and input the side frame into the fully connected layer to obtain an image feature vector; use the image feature vector as the input of the image feature extraction block; for the image extraction layer corresponding to the current text extraction layer, convert the image features of the previous image extraction layer into an image query vector, an image ... generate a first text feature based on the text query vector, the text key vector and the text value vector; generate a first text feature based on the text query vector, the text key vector and the text value vector; generate a first text feature based on the text key vector and image value vector; based on the image query vector, the image key vector and the image value vector, generate a first image feature; through a cross-attention fusion module, fuse the text query vector with the image key vector and the image value vector to obtain a second text feature, and fuse the image query vector with the text key vector and the text value vector to obtain a second image feature; based on the first text feature and the second text feature, determine the text feature of the current text extraction layer; based on the first image feature and the second image feature, determine the image feature of the current image extraction layer; wherein the text feature of the last layer of text extraction layer is the text fusion feature; the image feature of the last layer of image extraction layer is the image fusion feature.
[0056] Optionally, the word segmentation layer is a word segmentation model such as Byte-Pair Encoding (BPE) or WordPiece. The text feature extraction block is a Transformer model. The text extraction layer is a Transformer block. The detection layer is Faster R-CNN. The image feature extraction block is a VIT model. The image extraction layer is a VIT block, including a normalization layer, a multi-head self-attention module, and a multi-layer perceptron module. The preset number is 12; 12 layers of text extraction layers and 12 layers of image extraction layers correspond to each other in order.
[0057] Exemplarily, firstly, the word segmentation model is used to perform word segmentation processing on the text in the sample pair to obtain word segmentation features, and the word segmentation features are input into the Transformer model; for any Transformer module in the Transformer model, the text features of the previous Transformer module of the current Transformer module are converted into a text query vector, a text key vector and a text value vector; based on the text query vector, the text key vector and the text value vector, a first text feature is generated. The image in the sample pair is input into Faster R-CNN to detect the specific content of the image, and then each detection box is linearly mapped through a fully connected layer to generate a feature vector as the input of the VIT model; for the VIT module corresponding to the current Transformer module, as the current VIT module, the image features of the previous VIT module of the current VIT module are normalized, and then the normalized image features are After three matrices , and Convert it into image query vector Q, image key vector K and image value vector V, and then generate the final feature vector based on Q, K and V. The formula is as follows:
[0058]
[0059]
[0060] Among them, Softmax is the activation function, Image features The dimension of is the transpose of K, and i is the serial number of the current attention head. The attention features of all heads are concatenated to generate the final feature vector. The formula is as follows:
[0061]
[0062] in, yes The convolution layer of n can be 8. Thus, the first image feature is generated.
[0063] Through the cross-attention fusion module, the text query vector is fused with the image key vector and the image value vector to obtain a second text feature, and the image query vector is fused with the text key vector and the text value vector to obtain a second image feature. The first text feature and the second text feature are added to obtain the text feature of the current text extraction layer; the first image feature and the second image feature are added to obtain the image feature of the current image extraction layer. The text feature of the last text extraction layer is used as the text fusion feature; the image feature of the last image extraction layer is used as the image fusion feature.
[0064] In this embodiment, a cross-attention fusion module is used to realize cross-modal information interaction between image features and text features, so that image information is used to supplement text information, and text information is used to supplement image information, thereby solving the problem of information difficulty in multimodal scenarios.
[0065] In an exemplary embodiment, the text information extraction model further includes: a text hash layer; the image information extraction model further includes: an image hash layer; the converting the text fusion feature into a text hash code by the text information extraction model includes:
[0066] The text fusion feature is input into the text hash layer to obtain a text hash code; the image fusion feature is input into the image hash layer to obtain an image hash code.
[0067] Exemplarily, the text fusion features are input into the text hash layer, and the text fusion features are converted into text hash codes; the image fusion features are input into the image hash layer, and the image fusion features are converted into image hash codes.
[0068] In this embodiment, by converting both the text fusion features and the image fusion features into hash codes, data can be easily retrieved.
[0069] In an exemplary embodiment, determining the total loss of the text hash code and the image hash code based on the adaptive triple loss function and the quantization loss function includes:
[0070] Based on the correlation, boundary parameter and weight factor of the image hash code and the text hash code, an adaptive triple loss of the image hash code and the text hash code is determined by an adaptive triple loss function; a quantization loss of the image hash code and the text hash code is determined by a quantization loss function; and a total loss of the text hash code and the image hash code is determined based on the adaptive triple loss and the quantization loss.
[0071] Among them, an improved adaptive triple loss is used, and a dynamic weight adjustment mechanism is introduced to adaptively adjust the influence of the loss term according to the semantic similarity of the samples, so that the hash codes of similar samples are closer and the hash codes of dissimilar samples are more dispersed, thereby improving the discriminative ability of the hash codes.
[0072] Exemplarily, the loss of text hash code and image hash code is calculated by an adaptive triple loss formula, and the specific formula of adaptive triple loss is as follows:
[0073]
[0074]
[0075]
[0076] in, , and Represents the index of different instances (images or texts). The examples and instances are semantically related, while The examples and The instances are not semantically related. Whether two instances are semantically related can be determined by whether they share at least one label. and Represent image hash code and text hash code respectively, is the boundary parameter. is the adaptive weight factor between the three instances, which is the key to achieve adaptive optimization. It makes the hash codes of similar samples closer and the hash codes of dissimilar samples more dispersed. Specifically, the weight factor The similarity between instances and the idea of focal loss can be combined for calculation, so as to dynamically adjust the influence of each sample on the loss. The specific formula is as follows:
[0077]
[0078] in, Represents anchor point samples and positive samples The similarity between This reflects the strong semantic correlation between positive samples and anchor points, making the loss function pay more attention to these important positive sample pairs. Represents anchor point samples and negative samples The similarity between When is larger (i.e., the anchor points and negative samples are more similar), higher weights need to be applied to these difficult negative samples to enhance the model’s ability to distinguish difficult samples. is a scaling factor that adjusts the influence of the overall loss. and It is a parameter used to control the influence of different similarities on the weight factor. In addition, a quantization loss is further applied to the model to reduce the information loss caused by the hashing process. The specific quantization loss is The definition is as follows:
[0079]
[0080] in, It is The adaptive triplet loss is used to unify the binary hash codes of the instances. and quantify the loss Add them together to get the total loss.
[0081] In this embodiment, an adaptive triple loss is designed to adaptively adjust the influence of the loss term, so that the hash codes of similar samples can be closer and the hash codes of dissimilar samples can be more dispersed, thereby improving the discrimination ability of the hash codes.
[0082] In an exemplary embodiment, the image data includes N images, and the text data includes N texts describing each image; and establishing a plurality of sample pairs based on the text data and the image data includes:
[0083] For each image, the current image is combined with N texts respectively to obtain N sample pairs corresponding to the current image.
[0084] Among them, N is the number determined when the data set is made, each text corresponds to each image one by one, and the text content describes the content of the corresponding image.
[0085] Exemplarily, N images are paired with N texts respectively to obtain N×N sample pairs.
[0086] In this embodiment, by creating sample pairs for all images and all texts, the model can learn which ones are truly corresponding and which ones are not corresponding, thereby improving the retrieval performance of the model.
[0087] In an exemplary embodiment, the trained cross-modal information retrieval model is used to retrieve images based on input text, including:
[0088] The text is input into the text information extraction model of the trained cross-modal information retrieval model to obtain a text hash code; the image hash codes of all images are generated through the image information extraction model of the trained cross-modal information retrieval model; the similarity between the text hash code and the hash codes of all images is calculated respectively, and the image corresponding to the image hash code with the highest similarity is taken as the result of image retrieval.
[0089] Exemplarily, when it is necessary to retrieve images through text, the text is input into a text information extraction model of a trained cross-modal information retrieval model to obtain a text hash code; feature extraction is performed on all images through the image information extraction model of the trained cross-modal information retrieval model to generate image hash codes for all images; the similarity between the text hash code and the hash codes of all the images is calculated respectively, and they are sorted according to the similarity, and the image corresponding to the image hash code with the highest similarity is selected as the result of image retrieval.
[0090] In this embodiment, by calculating the similarity between all image hash codes and the text hash code used for search, the image that best matches the text description can be retrieved.
[0091] In an exemplary embodiment, Figure 2 As shown, an information retrieval method comprises:
[0092] Obtain text data and image data, the image data includes multiple images, and the text data includes multiple texts describing each image; pair the multiple images with the multiple texts to obtain sample pairs of the number of images × the number of texts. First, use the word segmentation model to separate the texts in the sample pairs. Perform word segmentation processing to obtain word segmentation features, and input the word segmentation features into the Transformer model, which includes N=12 Transformer modules; for any Transformer module in the Transformer model, convert the text features of the previous Transformer module of the current Transformer module into a text query vector, a text key vector, and a text value vector; based on the text query vector, the text key vector, and the text value vector, generate the first text feature. Input Faster R-CNN to detect the specific content of the image, then linearly map each detection box through the fully connected layer FC to generate a feature vector as the input of the VIT model. The VIT model includes N=12 VIT modules. For the VIT module corresponding to the current Transformer module, as the current VIT module, the image features of the previous VIT module of the current VIT module are normalized, and then the normalized image features are After three matrices , and Convert it into image query vector Q, image key vector K and image value vector V, and then generate the final feature vector based on Q, K and V. The formula is as follows:
[0093]
[0094]
[0095] Among them, Softmax is the activation function, Image features The dimension of is the transpose of K, and i is the serial number of the current attention head. The attention features of all heads are concatenated to generate the final feature vector. The formula is as follows:
[0096]
[0097] in, yes The convolution layer of n can be 8. Thus, the first image feature is generated.
[0098] Through the cross-attention fusion module, the text query vector is fused with the image key vector and the image value vector to obtain a second text feature, and the image query vector is fused with the text key vector and the text value vector to obtain a second image feature. The first text feature and the second text feature are added to obtain the text feature of the current text extraction layer; the first image feature and the second image feature are added to obtain the image feature of the current image extraction layer. The text features of the last layer of text extraction layer are used as text fusion features; the image features of the last layer of image extraction layer are used as image fusion features. The text fusion features are input into the text hash layer to convert the text fusion features into text hash codes; the image fusion features are input into the image hash layer to convert the image fusion features into image hash codes. Based on the adaptive triple loss function and the quantization loss function, the total loss of the text hash code and the image hash code is calculated; the specific formula for the adaptive triple loss is as follows:
[0099]
[0100]
[0101]
[0102] in, , and Represents the index of different instances (images or texts). The examples and instances are semantically related, while The examples and The instances are semantically unrelated. Whether two instances are semantically related can be determined by whether they share at least one label. and Represent image hash code and text hash code respectively, is the boundary parameter. is the adaptive weight factor between the three instances, which is the key to achieve adaptive optimization. It makes the hash codes of similar samples closer and the hash codes of dissimilar samples more dispersed. Specifically, the weight factor The similarity between instances and the idea of focal loss can be combined for calculation, so as to dynamically adjust the influence of each sample on the loss. The specific formula is as follows:
[0103]
[0104] in, Represents anchor point samples and positive samples The similarity between This reflects the strong semantic correlation between positive samples and anchor points, making the loss function pay more attention to these important positive sample pairs. Represents anchor point samples and negative samples The similarity between When is larger (i.e., the anchor points and negative samples are more similar), higher weights need to be applied to these difficult negative samples to enhance the model’s ability to distinguish difficult samples. is a scaling factor that adjusts the influence of the overall loss. and It is a parameter used to control the influence of different similarities on the weight factor. In addition, a quantization loss is further applied to the model to reduce the information loss caused by the hashing process. The specific quantization loss is The definition is as follows:
[0105]
[0106] in, It is The adaptive triplet loss is used to unify the binary hash codes of the instances. and quantify the loss The total loss is obtained by adding them together. Based on the total loss, the cross-modal information retrieval model to be trained is optimized to obtain a trained cross-modal information retrieval model; the trained cross-modal information retrieval model is used to retrieve images based on input text, or to retrieve text based on input images.
[0107] In an exemplary embodiment, Figure 3 and Figure 4 As shown in Figure 1, the retrieval results of the trained cross-modal information retrieval model in actual application are shown in Figure 1. Figure 3 It is the result of searching text based on image. The first column represents the input image, and the second to fifth columns represent the text retrieved based on the input image, arranged from left to right according to similarity. Figure 4 The first column represents the input text, and the second to fifth columns represent the images retrieved based on the input text, which are arranged from left to right according to similarity.
[0108] It should be understood that, although the various steps in the flowcharts involved in the above-mentioned embodiments are displayed in sequence according to the indication of the arrows, these steps are not necessarily executed in sequence according to the order indicated by the arrows. Unless there is a clear explanation in this article, the execution of these steps does not have a strict order restriction, and these steps can be executed in other orders. Moreover, at least a part of the steps in the flowcharts involved in the above-mentioned embodiments can include multiple steps or multiple stages, and these steps or stages are not necessarily executed at the same time, but can be executed at different times, and the execution order of these steps or stages is not necessarily to be carried out in sequence, but can be executed in turn or alternately with other steps or at least a part of the steps or stages in other steps.
[0109] In an exemplary embodiment, Figure 5 As shown, an information retrieval device is provided, comprising: an acquisition module, an extraction module and an optimization module, wherein:
[0110] An acquisition module, used for acquiring text data and image data, and establishing a plurality of sample pairs based on the text data and the image data;
[0111] An extraction module is used to generate text fusion features and image fusion features for each sample pair through a text information extraction model, an image information extraction model and a cross-attention module in a cross-modal information retrieval model to be trained; convert the text fusion features into text hash codes through the text information extraction model; and convert the image fusion features into image hash codes through the image information extraction model;
[0112] The optimization module is used to determine the total loss of the text hash code and the image hash code based on an adaptive triple loss function and a quantization loss function; based on the total loss, optimize the cross-modal information retrieval model to be trained to obtain a trained cross-modal information retrieval model; the trained cross-modal information retrieval model is used to retrieve images according to input text, or to retrieve text according to input images.
[0113] In one embodiment, the text information extraction model includes: a word segmentation layer and a text feature extraction block; the text feature extraction block includes a preset number of text extraction layers; the image information extraction model includes: a detection layer, a fully connected layer and an image feature extraction block; the image feature extraction block includes a preset number of image extraction layers; the preset number of text extraction layers and the preset number of image extraction layers correspond to each other in sequence; the extraction module is further used to:
[0114] Input the text in the sample pair into the word segmentation layer to obtain word segmentation features; use the word segmentation features as the input of the text feature extraction block; for any text extraction layer, convert the text features of the previous text extraction layer of the current text extraction layer into a text query vector, a text key vector and a text value vector; generate a first text feature based on the text query vector, the text key vector and the text value vector; input the image in the sample pair into the detection layer to obtain a detection frame, and input the side frame into the fully connected layer to obtain an image feature vector; use the image feature vector as the input of the image feature extraction block; for the image extraction layer corresponding to the current text extraction layer, convert the image features of the previous image extraction layer into an image query vector, an image ... generate a first text feature based on the text query vector, the text key vector and the text value vector; generate a first text feature based on the text query vector, the text key vector and the text value vector; generate a first text feature based on the text key vector and image value vector; based on the image query vector, the image key vector and the image value vector, generate a first image feature; through a cross-attention fusion module, fuse the text query vector with the image key vector and the image value vector to obtain a second text feature, and fuse the image query vector with the text key vector and the text value vector to obtain a second image feature; based on the first text feature and the second text feature, determine the text feature of the current text extraction layer; based on the first image feature and the second image feature, determine the image feature of the current image extraction layer; wherein the text feature of the last layer of text extraction layer is the text fusion feature; the image feature of the last layer of image extraction layer is the image fusion feature.
[0115] In one embodiment, the text information extraction model further includes: a text hash layer; the image information extraction model further includes: an image hash layer; the extraction module is further used to:
[0116] The text fusion feature is input into the text hash layer to obtain a text hash code; the image fusion feature is input into the image hash layer to obtain an image hash code.
[0117] In one embodiment, the optimization module is further used to:
[0118] Based on the correlation, boundary parameter and weight factor of the image hash code and the text hash code, an adaptive triple loss of the image hash code and the text hash code is determined by an adaptive triple loss function; a quantization loss of the image hash code and the text hash code is determined by a quantization loss function; and a total loss of the text hash code and the image hash code is determined based on the adaptive triple loss and the quantization loss.
[0119] In one embodiment, the image data includes N images, and the text data includes N texts describing each image; the acquisition module is further used to:
[0120] For each image, the current image is combined with N texts respectively to obtain N sample pairs corresponding to the current image.
[0121] In one embodiment, a training module is further included for:
[0122] The text is input into the text information extraction model of the trained cross-modal information retrieval model to obtain a text hash code; the image hash codes of all images are generated through the image information extraction model of the trained cross-modal information retrieval model; the similarity between the text hash code and the hash codes of all images is calculated respectively, and the image corresponding to the image hash code with the highest similarity is taken as the result of image retrieval.
[0123] Each module in the above information retrieval device can be implemented in whole or in part by software, hardware or a combination thereof. Each module can be embedded in or independent of a processor in a computer device in the form of hardware, or can be stored in a memory in a computer device in the form of software, so that the processor can call and execute the operations corresponding to each module.
[0124] In an exemplary embodiment, a computer device is provided. The computer device may be a server, and its internal structure diagram may be as shown in FIG. Figure 6 As shown. The computer device includes a processor, a memory, an input / output interface (Input / Output, referred to as I / O) and a communication interface. The processor, the memory and the input / output interface are connected through a system bus, and the communication interface is connected to the system bus through the input / output interface. The processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The database of the computer device is used to store text data and image data. The input / output interface of the computer device is used to exchange information between the processor and an external device. The communication interface of the computer device is used to communicate with an external terminal through a network connection. When the computer program is executed by the processor, an information retrieval method is implemented.
[0125] Those skilled in the art will understand that Figure 6 The structure shown in the figure is only a block diagram of a part of the structure related to the solution of the present application, and does not constitute a limitation on the computer device to which the solution of the present application is applied. The specific computer device may include more or fewer components than those shown in the figure, or combine certain components, or have a different arrangement of components.
[0126] Acquire text data and image data, and establish multiple sample pairs based on the text data and the image data;
[0127] For each sample pair, a text fusion feature and an image fusion feature are generated through a text information extraction model, an image information extraction model and a cross-attention module in a cross-modal information retrieval model to be trained; the text fusion feature is converted into a text hash code through the text information extraction model; the image fusion feature is converted into an image hash code through the image information extraction model;
[0128] Based on an adaptive triple loss function and a quantization loss function, the total loss of the text hash code and the image hash code is determined; based on the total loss, the cross-modal information retrieval model to be trained is optimized to obtain a trained cross-modal information retrieval model; the trained cross-modal information retrieval model is used to retrieve images according to input text, or to retrieve text according to input images.
[0129] In one embodiment, the text information extraction model includes: a word segmentation layer and a text feature extraction block; the text feature extraction block includes a preset number of text extraction layers; the image information extraction model includes: a detection layer, a fully connected layer and an image feature extraction block; the image feature extraction block includes a preset number of image extraction layers; the preset number of text extraction layers and the preset number of image extraction layers correspond to each other in sequence; when the processor executes the computer program, the following steps are also implemented:
[0130] Input the text in the sample pair into the word segmentation layer to obtain word segmentation features; use the word segmentation features as the input of the text feature extraction block; for any text extraction layer, convert the text features of the previous text extraction layer of the current text extraction layer into a text query vector, a text key vector and a text value vector; generate a first text feature based on the text query vector, the text key vector and the text value vector; input the image in the sample pair into the detection layer to obtain a detection frame, and input the side frame into the fully connected layer to obtain an image feature vector; use the image feature vector as the input of the image feature extraction block; for the image extraction layer corresponding to the current text extraction layer, convert the image features of the previous image extraction layer into an image query vector, an image ... generate a first text feature based on the text query vector, the text key vector and the text value vector; generate a first text feature based on the text query vector, the text key vector and the text value vector; generate a first text feature based on the text key vector and image value vector; based on the image query vector, the image key vector and the image value vector, generate a first image feature; through a cross-attention fusion module, fuse the text query vector with the image key vector and the image value vector to obtain a second text feature, and fuse the image query vector with the text key vector and the text value vector to obtain a second image feature; based on the first text feature and the second text feature, determine the text feature of the current text extraction layer; based on the first image feature and the second image feature, determine the image feature of the current image extraction layer; wherein the text feature of the last layer of text extraction layer is the text fusion feature; the image feature of the last layer of image extraction layer is the image fusion feature.
[0131] In one embodiment, the text information extraction model further includes: a text hash layer; the image information extraction model further includes: an image hash layer; when the processor executes the computer program, the following steps are also implemented:
[0132] The text fusion feature is input into the text hash layer to obtain a text hash code; the image fusion feature is input into the image hash layer to obtain an image hash code.
[0133] In one embodiment, when the processor executes the computer program, the processor further implements the following steps:
[0134] Based on the correlation, boundary parameter and weight factor of the image hash code and the text hash code, an adaptive triple loss of the image hash code and the text hash code is determined by an adaptive triple loss function; a quantization loss of the image hash code and the text hash code is determined by a quantization loss function; and a total loss of the text hash code and the image hash code is determined based on the adaptive triple loss and the quantization loss.
[0135] In one embodiment, the image data includes N images, and the text data includes N texts describing each image; when the processor executes the computer program, the following steps are also implemented:
[0136] For each image, the current image is combined with N texts respectively to obtain N sample pairs corresponding to the current image.
[0137] In one embodiment, when the processor executes the computer program, the processor further implements the following steps:
[0138] The text is input into the text information extraction model of the trained cross-modal information retrieval model to obtain a text hash code; the image hash codes of all images are generated through the image information extraction model of the trained cross-modal information retrieval model; the similarity between the text hash code and the hash codes of all images is calculated respectively, and the image corresponding to the image hash code with the highest similarity is taken as the result of image retrieval.
[0139] In one embodiment, a computer readable storage medium is provided, on which a computer program is stored, and when the computer program is executed by a processor, the following steps are implemented:
[0140] Acquire text data and image data, and establish multiple sample pairs based on the text data and the image data;
[0141] For each sample pair, a text fusion feature and an image fusion feature are generated through a text information extraction model, an image information extraction model and a cross-attention module in a cross-modal information retrieval model to be trained; the text fusion feature is converted into a text hash code through the text information extraction model; the image fusion feature is converted into an image hash code through the image information extraction model;
[0142] Based on an adaptive triple loss function and a quantization loss function, the total loss of the text hash code and the image hash code is determined; based on the total loss, the cross-modal information retrieval model to be trained is optimized to obtain a trained cross-modal information retrieval model; the trained cross-modal information retrieval model is used to retrieve images according to input text, or to retrieve text according to input images.
[0143] In one embodiment, the text information extraction model includes: a word segmentation layer and a text feature extraction block; the text feature extraction block includes a preset number of text extraction layers; the image information extraction model includes: a detection layer, a fully connected layer and an image feature extraction block; the image feature extraction block includes a preset number of image extraction layers; the preset number of text extraction layers and the preset number of image extraction layers correspond to each other in sequence; when the computer program is executed by the processor, the following steps are also implemented:
[0144] Input the text in the sample pair into the word segmentation layer to obtain word segmentation features; use the word segmentation features as the input of the text feature extraction block; for any text extraction layer, convert the text features of the previous text extraction layer of the current text extraction layer into a text query vector, a text key vector and a text value vector; generate a first text feature based on the text query vector, the text key vector and the text value vector; input the image in the sample pair into the detection layer to obtain a detection frame, and input the side frame into the fully connected layer to obtain an image feature vector; use the image feature vector as the input of the image feature extraction block; for the image extraction layer corresponding to the current text extraction layer, convert the image features of the previous image extraction layer into an image query vector, an image ... generate a first text feature based on the text query vector, the text key vector and the text value vector; generate a first text feature based on the text query vector, the text key vector and the text value vector; generate a first text feature based on the text key vector and image value vector; based on the image query vector, the image key vector and the image value vector, generate a first image feature; through a cross-attention fusion module, fuse the text query vector with the image key vector and the image value vector to obtain a second text feature, and fuse the image query vector with the text key vector and the text value vector to obtain a second image feature; based on the first text feature and the second text feature, determine the text feature of the current text extraction layer; based on the first image feature and the second image feature, determine the image feature of the current image extraction layer; wherein the text feature of the last layer of text extraction layer is the text fusion feature; the image feature of the last layer of image extraction layer is the image fusion feature.
[0145] In one embodiment, the text information extraction model further includes: a text hash layer; the image information extraction model further includes: an image hash layer; when the computer program is executed by the processor, the following steps are also implemented:
[0146] The text fusion feature is input into the text hash layer to obtain a text hash code; the image fusion feature is input into the image hash layer to obtain an image hash code.
[0147] In one embodiment, when the computer program is executed by a processor, the following steps are also implemented:
[0148] Based on the correlation, boundary parameter and weight factor of the image hash code and the text hash code, an adaptive triple loss of the image hash code and the text hash code is determined by an adaptive triple loss function; a quantization loss of the image hash code and the text hash code is determined by a quantization loss function; and a total loss of the text hash code and the image hash code is determined based on the adaptive triple loss and the quantization loss.
[0149] In one embodiment, the image data includes N images, and the text data includes N texts describing each image; when the computer program is executed by the processor, the following steps are also implemented:
[0150] For each image, the current image is combined with N texts respectively to obtain N sample pairs corresponding to the current image.
[0151] In one embodiment, when the computer program is executed by a processor, the following steps are also implemented:
[0152] The text is input into the text information extraction model of the trained cross-modal information retrieval model to obtain a text hash code; the image hash codes of all images are generated through the image information extraction model of the trained cross-modal information retrieval model; the similarity between the text hash code and the hash codes of all images is calculated respectively, and the image corresponding to the image hash code with the highest similarity is taken as the result of image retrieval.
[0153] In one embodiment, a computer program product is provided, comprising a computer program, which, when executed by a processor, implements the following steps:
[0154] Acquire text data and image data, and establish multiple sample pairs based on the text data and the image data;
[0155] For each sample pair, a text fusion feature and an image fusion feature are generated through a text information extraction model, an image information extraction model and a cross-attention module in a cross-modal information retrieval model to be trained; the text fusion feature is converted into a text hash code through the text information extraction model; the image fusion feature is converted into an image hash code through the image information extraction model;
[0156] Based on an adaptive triple loss function and a quantization loss function, the total loss of the text hash code and the image hash code is determined; based on the total loss, the cross-modal information retrieval model to be trained is optimized to obtain a trained cross-modal information retrieval model; the trained cross-modal information retrieval model is used to retrieve images according to input text, or to retrieve text according to input images.
[0157] In one embodiment, the text information extraction model includes: a word segmentation layer and a text feature extraction block; the text feature extraction block includes a preset number of text extraction layers; the image information extraction model includes: a detection layer, a fully connected layer and an image feature extraction block; the image feature extraction block includes a preset number of image extraction layers; the preset number of text extraction layers and the preset number of image extraction layers correspond to each other in sequence; when the computer program is executed by the processor, the following steps are also implemented:
[0158] Input the text in the sample pair into the word segmentation layer to obtain word segmentation features; use the word segmentation features as the input of the text feature extraction block; for any text extraction layer, convert the text features of the previous text extraction layer of the current text extraction layer into a text query vector, a text key vector and a text value vector; generate a first text feature based on the text query vector, the text key vector and the text value vector; input the image in the sample pair into the detection layer to obtain a detection frame, and input the side frame into the fully connected layer to obtain an image feature vector; use the image feature vector as the input of the image feature extraction block; for the image extraction layer corresponding to the current text extraction layer, convert the image features of the previous image extraction layer into an image query vector, an image ... generate a first text feature based on the text query vector, the text key vector and the text value vector; generate a first text feature based on the text query vector, the text key vector and the text value vector; generate a first text feature based on the text key vector and image value vector; based on the image query vector, the image key vector and the image value vector, generate a first image feature; through a cross-attention fusion module, fuse the text query vector with the image key vector and the image value vector to obtain a second text feature, and fuse the image query vector with the text key vector and the text value vector to obtain a second image feature; based on the first text feature and the second text feature, determine the text feature of the current text extraction layer; based on the first image feature and the second image feature, determine the image feature of the current image extraction layer; wherein the text feature of the last layer of text extraction layer is the text fusion feature; the image feature of the last layer of image extraction layer is the image fusion feature.
[0159] In one embodiment, the text information extraction model further includes: a text hash layer; the image information extraction model further includes: an image hash layer; when the computer program is executed by the processor, the following steps are also implemented:
[0160] The text fusion feature is input into the text hash layer to obtain a text hash code; the image fusion feature is input into the image hash layer to obtain an image hash code.
[0161] In one embodiment, when the computer program is executed by a processor, the following steps are also implemented:
[0162] Based on the correlation, boundary parameter and weight factor of the image hash code and the text hash code, an adaptive triple loss of the image hash code and the text hash code is determined by an adaptive triple loss function; a quantization loss of the image hash code and the text hash code is determined by a quantization loss function; and a total loss of the text hash code and the image hash code is determined based on the adaptive triple loss and the quantization loss.
[0163] In one embodiment, the image data includes N images, and the text data includes N texts describing each image; when the computer program is executed by the processor, the following steps are also implemented:
[0164] For each image, the current image is combined with N texts respectively to obtain N sample pairs corresponding to the current image.
[0165] In one embodiment, when the computer program is executed by a processor, the following steps are also implemented:
[0166] The text is input into the text information extraction model of the trained cross-modal information retrieval model to obtain a text hash code; the image hash codes of all images are generated through the image information extraction model of the trained cross-modal information retrieval model; the similarity between the text hash code and the hash codes of all images is calculated respectively, and the image corresponding to the image hash code with the highest similarity is taken as the result of image retrieval.
[0167] A person of ordinary skill in the art can understand that all or part of the processes in the above-mentioned embodiment method can be completed by instructing the relevant hardware through a computer program, and the computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above-mentioned methods. Among them, any reference to the memory, database or other medium used in the embodiments provided in the present application can include at least one of non-volatile memory and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetic random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. As an illustration and not limitation, RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM). The database involved in each embodiment provided in this application may include at least one of a relational database and a non-relational database. Non-relational databases may include distributed databases based on blockchains, etc., but are not limited to this. The processor involved in each embodiment provided in this application may be a general-purpose processor, a central processing unit, a graphics processor, a digital signal processor, a programmable logic device, a data processing logic device based on quantum computing, an artificial intelligence (AI) processor, etc., but are not limited to this.
[0168] The technical features of the above embodiments may be combined arbitrarily. To make the description concise, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this application.
[0169] The above-described embodiments only express several implementation methods of the present application, and the descriptions thereof are relatively specific and detailed, but they cannot be understood as limiting the scope of the present application. It should be pointed out that, for a person of ordinary skill in the art, several variations and improvements can be made without departing from the concept of the present application, and these all belong to the protection scope of the present application. Therefore, the protection scope of the present application shall be subject to the attached claims.
Claims
1. An information retrieval method, characterized in that: The method comprises: Acquire text data and image data, and establish multiple sample pairs based on the text data and the image data; For each sample pair, a text fusion feature and an image fusion feature are generated through a text information extraction model, an image information extraction model and a cross-attention module in a cross-modal information retrieval model to be trained; the text fusion feature is converted into a text hash code through the text information extraction model; the image fusion feature is converted into an image hash code through the image information extraction model; Based on an adaptive triple loss function and a quantization loss function, the total loss of the text hash code and the image hash code is determined; based on the total loss, the cross-modal information retrieval model to be trained is optimized to obtain a trained cross-modal information retrieval model; the trained cross-modal information retrieval model is used to retrieve images according to input text, or to retrieve text according to input images.
2. The method according to claim 1, characterized in that The text information extraction model includes: a word segmentation layer and a text feature extraction block; the text feature extraction block includes a preset number of text extraction layers; the image information extraction model includes: a detection layer, a fully connected layer and an image feature extraction block; the image feature extraction block includes a preset number of image extraction layers; the preset number of text extraction layers and the preset number of image extraction layers correspond to each other in sequence; the text information extraction model, the image information extraction model and the cross-attention module in the cross-modal information retrieval model to be trained are used to generate text fusion features and image fusion features, including: Input the text in the sample pair into the word segmentation layer to obtain word segmentation features; use the word segmentation features as input to the text feature extraction module; For any text extraction layer, converting the text features of the previous text extraction layer of the current text extraction layer into a text query vector, a text key vector and a text value vector; generating a first text feature based on the text query vector, the text key vector and the text value vector; Input the image in the sample pair into the detection layer to obtain a detection frame, input the detection frame into the fully connected layer to obtain an image feature vector; and use the image feature vector as the input of the image feature extraction block; For the image extraction layer corresponding to the current text extraction layer, converting the image features of the previous image extraction layer into an image query vector, an image key vector and an image value vector; generating a first image feature based on the image query vector, the image key vector and the image value vector; The text query vector is fused with the image key vector and the image value vector through a cross-attention fusion module to obtain a second text feature, and the image query vector is fused with the text key vector and the text value vector to obtain a second image feature; Determine text features of a current text extraction layer based on the first text features and the second text features; determine image features of a current image extraction layer based on the first image features and the second image features; Among them, the text features of the last text extraction layer are text fusion features; the image features of the last image extraction layer are image fusion features.
3. The method according to claim 2, characterized in that The text information extraction model further includes: a text hash layer; the image information extraction model further includes: an image hash layer; The step of converting the text fusion feature into a text hash code by using the text information extraction model includes: Inputting the text fusion feature into the text hash layer to obtain a text hash code; The image fusion feature is input into the image hash layer to obtain an image hash code.
4. The method according to claim 1, characterized in that: The determining the total loss of the text hash code and the image hash code based on the adaptive triple loss function and the quantization loss function includes: Based on the correlation, boundary parameter and weight factor of the image hash code and the text hash code, an adaptive triple loss of the image hash code and the text hash code is determined through an adaptive triple loss function; Determine the quantization loss of the image hash code and the text hash code through a quantization loss function; Based on the adaptive triplet loss and the quantization loss, a total loss of a text hash code and an image hash code is determined.
5. The method according to claim 1, characterized in that: The image data includes N images, and the text data includes N texts describing each image; The establishing of a plurality of sample pairs based on the text data and the image data comprises: For each image, the current image is combined with N texts respectively to obtain N sample pairs corresponding to the current image.
6. The method according to claim 1, characterized in that The trained cross-modal information retrieval model is used to retrieve images according to input text, including: Input the text into the text information extraction model of the trained cross-modal information retrieval model to obtain the text hash code; Generate image hash codes for all images through the image information extraction model of the trained cross-modal information retrieval model; The similarity between the text hash code and the hash codes of all the images is calculated respectively, and the image corresponding to the image hash code with the highest similarity is taken as the result of image retrieval.
7. An information retrieval device, characterized in that: The device comprises: An acquisition module, used for acquiring text data and image data, and establishing a plurality of sample pairs based on the text data and the image data; An extraction module is used to generate text fusion features and image fusion features for each sample pair through a text information extraction model, an image information extraction model and a cross-attention module in a cross-modal information retrieval model to be trained; convert the text fusion features into text hash codes through the text information extraction model; and convert the image fusion features into image hash codes through the image information extraction model; The optimization module is used to determine the total loss of the text hash code and the image hash code based on an adaptive triple loss function and a quantization loss function; based on the total loss, optimize the cross-modal information retrieval model to be trained to obtain a trained cross-modal information retrieval model; the trained cross-modal information retrieval model is used to retrieve images according to input text, or to retrieve text according to input images.
8. A computer device comprising a memory and a processor, wherein the memory stores a computer program, wherein: When the processor executes the computer program, the steps of the method according to any one of claims 1 to 6 are implemented.
9. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 6 are implemented.
10. A computer program product, comprising a computer program, characterized in that When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 6 are implemented.