Cross-modal alignment image text matching method and device, equipment and medium
By using weight vectors to correct prompt word offset in the image text matching model, the problems of difficulty in feature extraction and weakened model expression ability in traditional methods are solved, and fast and accurate image text matching is achieved.
Patent Information
- Application Number
- CN202510184813.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-19
- Publication Date
- 2025-06-06
AI Technical Summary
Traditional image text matching methods are difficult to extract features quickly and accurately, and when pre-trained model migration application, setting prompt words will affect the original model and weaken expression capabilities.
By obtaining the image and text to be matched, the text encoder and image encoder of the preset image text matching model encode text and images respectively, generate text vectors and image vectors, and determine the weight vectors based on the converted vectors, which are used to correct the offset caused by prompt words and pay attention to important features.
It realizes rapid and accurate matching of images and text, eliminates prompt word offsets through weight vectors, and improves model expression ability and image text space consistency.
Smart Images

Figure CN120107735A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the technical field of multimodal data processing, and in particular to a method, device, equipment and medium for cross-modal aligned image-text matching. Background Art
[0002] With the development of digital information, the matching of images and texts can establish the association between image content and text description, and has wide applications in cross-modal retrieval, image annotation, and image classification.
[0003] Traditional image-text matching methods usually use deep learning models to output the features of images and texts, and use the features of images and texts to associate them. However, this method is difficult to fully and accurately extract the required features. In related technologies, pre-trained model migration is also applied to the task of image-text matching, and the required features are focused on through prompt words. However, setting prompt words, migration and training will affect the original model, weaken the model's expressive power, and make it difficult to quickly and accurately achieve image and text matching. Summary of the invention
[0004] The embodiments of the present application provide a method, apparatus, device and medium for cross-modal aligned image-text matching to quickly and accurately achieve image and text matching.
[0005] In a first aspect, an embodiment of the present application provides a cross-modal aligned image text matching method, comprising:
[0006] Obtaining the image to be matched, the text to be matched and the text prompt word;
[0007] Inputting the text to be matched and the text prompt word into a text encoder of a preset image-text matching model to obtain a text vector and a first conversion vector output by the text encoder;
[0008] Input the image to be matched into the image encoder of the image-text matching model to obtain an image vector and a second conversion vector output by the image encoder; wherein the image vector and the text vector have the same dimension;
[0009] Determine a weight vector according to the first conversion vector and the second conversion vector; wherein the dimension of the weight vector is the same as the dimension of the text vector, and the weight vector is used to determine the weight of each dimension in the image vector and the text vector;
[0010] Determining the similarity between the image to be matched and the text to be matched according to the image vector, the text vector and the weight vector;
[0011] Based on the similarity, a matching result between the image to be matched and the text to be matched is determined.
[0012] In a possible implementation manner, determining a weight vector according to the first conversion vector and the second conversion vector includes:
[0013] Generate a key vector and a value vector of a cross attention mechanism according to the first transformation vector;
[0014] Generate a query vector of a cross-attention mechanism according to the second transformation vector;
[0015] Performing cross attention calculation on the query vector, the key vector and the value vector to obtain a cross vector;
[0016] The cross vector is extracted and projected to obtain a weight vector.
[0017] In a possible implementation, performing cross attention calculation on the query vector, the key vector, and the value vector according to the cross attention mechanism to obtain a cross vector includes:
[0018] According to the expression: Get the cross vector;
[0019] Wherein, logits represents the cross vector; softmax represents the activation function in the cross attention mechanism; Q represents the query vector; K represents the key vector; V represents the value vector; d k represents the dimension of the key vector; T represents the transpose of the vector.
[0020] In a possible implementation, the text vector includes text elements of multiple dimensions, the image vector includes image elements of multiple dimensions, and the weight vector includes weight factors of multiple dimensions;
[0021] Determining the similarity between the image to be matched and the text to be matched according to the image vector, the text vector and the weight vector includes:
[0022] Calculate the similarity score of the text vector and the image vector in each dimension according to the text elements of each dimension in the text vector, the image elements of the corresponding dimension in the image vector and the weight factors of the corresponding dimension in the weight vector;
[0023] The similarity between the image to be matched and the text to be matched is determined according to the similarity scores of all dimensions.
[0024] In a possible implementation, the text encoder includes a plurality of text conversion layers, and each text conversion layer is correspondingly provided with a text prompt word;
[0025] The step of inputting the text to be matched and the text prompt word into a text encoder of a preset image-text matching model to obtain a text vector and a first conversion vector output by the text encoder includes:
[0026] Concatenate the text to be matched and the text prompt word corresponding to the first text conversion layer to obtain a concatenated vector;
[0027] Inputting the concatenated vector into a first text conversion layer in the text encoder to obtain a third conversion vector output by the first text conversion layer;
[0028] The text element at the first preset position in the third conversion vector is replaced with a text prompt word corresponding to the second text conversion layer, and the replaced vector is output to the second text conversion layer to obtain a fourth conversion vector output by the second text conversion layer, and the conversion process corresponding to one text conversion layer is completed; the conversion process is continuously performed until the first conversion vector output by the last text conversion layer is obtained;
[0029] Based on the first conversion vector, a text vector output by the text encoder is obtained.
[0030] In a possible implementation, the image encoder includes a plurality of image conversion layers;
[0031] The step of inputting the image to be matched into the image encoder of the image-text matching model to obtain the image vector and the second conversion vector output by the image encoder comprises:
[0032] Acquire multiple visual prompt words; wherein each image conversion layer except the first image conversion layer is correspondingly set with a visual prompt word;
[0033] Inputting the image to be matched into a first image conversion layer in the image encoder to obtain a fifth conversion vector;
[0034] The image element at the second preset position in the fifth conversion vector is replaced with a visual prompt word corresponding to the second image conversion layer, and the replaced vector is input into the second image conversion layer to obtain a sixth conversion vector, thereby completing the conversion process of one image conversion layer; the conversion process is continuously performed until a second conversion vector output by the last image conversion layer is obtained;
[0035] Based on the second conversion vector, an image vector output by the image encoder is obtained.
[0036] In a possible implementation, the to-be-matched text includes a plurality of image tags, and the text prompt word includes a tag prompt word corresponding to each image tag;
[0037] The step of inputting the text to be matched and the text prompt word into a text encoder of a preset image-text matching model to obtain a text vector and a first conversion vector output by the text encoder includes:
[0038] Input each image label and the corresponding label prompt word into the text encoder, and obtain the text vector and the first conversion vector corresponding to each image label output by the text encoder;
[0039] After inputting the image to be matched into the image encoder of the image-text matching model to obtain the image vector and the second conversion vector output by the image encoder, the method further includes:
[0040] Determining weight vectors of the image vector and each image label respectively according to the second conversion vector and the first conversion vector corresponding to each image label;
[0041] Determine the similarity between the image to be matched and each image label according to the image vector, the text vector corresponding to each image label and the corresponding weight vector;
[0042] The image to be matched is classified based on the similarity between the image to be matched and each image label.
[0043] In a second aspect, an embodiment of the present application provides a cross-modal aligned image-text matching device, comprising:
[0044] An acquisition module is used to acquire the image to be matched, the text to be matched and the text prompt word;
[0045] A text encoding module, used for inputting the text to be matched and the text prompt word into a text encoder of a preset image-text matching model, and obtaining a text vector and a first conversion vector output by the text encoder;
[0046] An image encoding module, used for inputting the image to be matched into the image encoder of the image-text matching model to obtain an image vector and a second conversion vector output by the image encoder; wherein the image vector and the text vector have the same dimension;
[0047] A weight determination module, configured to determine a weight vector according to the first conversion vector and the second conversion vector; wherein the dimension of the weight vector is the same as the dimension of the text vector, and the weight vector is used to determine the weight of each dimension in the image vector and the text vector;
[0048] A matching module is used to determine the similarity between the image to be matched and the text to be matched according to the image vector, the text vector and the weight vector; and to determine the matching result of the image to be matched and the text to be matched based on the similarity.
[0049] In a third aspect, an embodiment of the present application provides an electronic device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the computer program, the steps of the method described in the first aspect or any possible implementation method of the first aspect are implemented.
[0050] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium, which stores a computer program. When the computer program is executed by a processor, it implements the steps of the method described in the first aspect or any possible implementation method of the first aspect.
[0051] In a fifth aspect, an embodiment of the present application provides a computer program product. When the computer program product runs on an electronic device, the electronic device executes the steps of the method described in the first aspect or any possible implementation method of the first aspect.
[0052] Compared with the prior art, the embodiments of the present invention have the following beneficial effects:
[0053] In the embodiment of the present application, a text encoder is used to encode the text to be matched and the text prompt word to obtain a text vector, and an image encoder is used to encode the image to be matched to obtain an image vector; a first conversion vector and a second conversion vector are then used to determine a weight vector, which can fully consider various features in the text to be matched and the image to be matched, determine the correlation between the encoded text vector and the image vector, and consider the offset caused by the prompt word on each dimension in the vector; the similarity between the image to be matched and the text to be matched is determined through the image vector, the text vector and the weight vector, and the weight vector can be used to eliminate the offset caused by the prompt word. At the same time, attention can be paid to important features in the image to be matched and the text to be matched, the consistency of the image space and the text space can be improved, and the expression ability of the model can be guaranteed, so as to accurately obtain the similarity between the image to be matched and the text to be matched, so as to accurately match the image and the text according to the similarity. BRIEF DESCRIPTION OF THE DRAWINGS
[0054] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.
[0055] Figure 1 is a schematic diagram of the structure of the CLIP model provided in the embodiment of the present application;
[0056] Figure 2 is a flowchart of an implementation of a cross-modal aligned image-text matching method provided in an embodiment of the present application;
[0057] Figure 3 is a comparative schematic diagram before and after the weight factor adjustment provided in the embodiment of the present application;
[0058] Figure 4 is a structural diagram of an image-text matching model provided in an embodiment of the present application;
[0059] Figure 5 is a schematic diagram of the structure of a cross-modal aligned image-text matching device provided in an embodiment of the present application;
[0060] Figure 6 It is a schematic diagram of an electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION
[0061] In the following description, specific details such as specific system structures, technologies, etc. are provided for the purpose of illustration rather than limitation, so as to provide a thorough understanding of the embodiments of the present application. However, it should be clear to those skilled in the art that the present application may also be implemented in other embodiments without these specific details. In other cases, detailed descriptions of well-known systems, devices, circuits, and methods are omitted to prevent unnecessary details from obstructing the description of the present application.
[0062] The core of the Contrastive Language-Image Pre-Training (CLIP) model is contrastive learning, which learns representations by minimizing the distance between positive samples (matching image and text pairs) while maximizing the distance between negative samples (unmatched image and text pairs), enabling the model to match images with corresponding texts.
[0063] See also Figure 1 The structural diagram of the CLIP model shown in the figure shows that when the CLIP model is used for image-text matching, the text encoder of the CLIP model will encode the text to obtain a text vector, and the image encoder will encode the image to obtain an image vector. Finally, the matching result of the image and text is determined by calculating the cosine similarity of the two vectors.
[0064] The inventors found that when the CLIP model is applied to a specific downstream image-text matching task, a prompt word can be added to the input of the text encoder for guidance to adjust the processing method of the image and text, thereby improving the model's performance for image-text matching. However, when applied to different tasks, setting prompt words will also cause the text vector to shift in different dimensions, affecting the original model, weakening the model's expressive power, and destroying the alignment relationship between text and image in its feature space.
[0065] In order to achieve the idea of matching images and texts quickly and accurately, in the embodiments of the present application, it is considered to add weights to different dimensions of the text vector to correct the offset caused by the prompt word, and to strengthen the focus on important features, improve the consistency of the text and image space, thereby improving the expressiveness of the model, so as to achieve the matching of images and texts quickly and accurately. However, directly adding weights to each dimension is prone to overfitting problems during training, and it is difficult to achieve the above purpose quickly and accurately. Therefore, the inventors consider determining the weights of each dimension by encoding the text vector and image vector, associating the weights of each dimension with the corresponding text and image, achieving accurate adjustment of each dimension, reducing the risk of overfitting, and achieving the matching of images and texts quickly and accurately.
[0066] In order to make the purpose, technical solutions and advantages of the present application clearer, specific embodiments will be described below in conjunction with the accompanying drawings.
[0067] Figure 2 The implementation flow chart of the cross-modal aligned image text matching method provided in the embodiment of the present application is described in detail as follows:
[0068] Step 201, obtaining an image to be matched, a text to be matched, and a text prompt word.
[0069] In this embodiment, the image to be matched can be any image containing features such as objects, scenery or colors. The text to be matched can be text corresponding to the features contained in the image to be matched, or text unrelated to the image to be matched, which mainly depends on the specific matching task.
[0070] For example, in the target recognition task of industrial production, the image to be matched can be the image of the product produced, and the text to be matched can be the name of the product produced. In the field of remote sensing, the image to be matched can be the type of ground or buildings, etc., and the text to be matched can be the name of the ground type or the name of the building. In natural scenery, the image to be matched can be an image containing trees or animals, and the text to be matched can be the name of the tree or animal.
[0071] In addition, the text to be matched can be text that is unrelated to the image. Such text to be matched and the image to be matched are negative samples.
[0072] Here, the text prompt word can adopt a prompt word template similar to "a photo of a{class}", where "aphoto of a" can be a text prompt word, and {class} can be the corresponding text to be matched. For example, when matching the image and name of a cat, Chinese rural cat, ragdoll cat, Persian cat, Norwegian forest cat, Siberian cat, British shorthair cat and hairless cat can be filled in {class} as the text to be matched.
[0073] In addition, the text prompt words can also be colors, positions, actions, etc. For example, yellow, white, black, located on a tree, located on a roof, jumping or running, etc.
[0074] In addition to the above method of directly adding text prompt words, obtaining text prompt words can also be performed by the following random initialization method: obtaining the number of text prompt words; randomly initializing the text prompt words according to the number of text prompt words; and obtaining the final text prompt words by training the image-text matching model. Among them, the randomly initialized text prompt words can be initialized by using Gaussian distribution.
[0075] Step 202: input the text to be matched and the text prompt word into a text encoder of a preset image-text matching model to obtain a text vector and a first conversion vector output by the text encoder.
[0076] In this embodiment, the text encoder can encode the to-be-matched text and the text prompt words to obtain the corresponding text vector. In particular, by adding the text prompt words, the text prompt words can be used to guide the text vector to focus on the required features.
[0077] The text encoder may include multiple text conversion layers and a linear layer. The text conversion layer is used to encode the text to be matched and the text prompt word. The linear layer can perform linear transformation on the encoded vector and output a vector that meets actual needs. Among them, the vector output by the last text conversion layer is the first conversion vector output by the text encoder, and the vector output by the linear layer is the text vector finally output by the text encoder.
[0078] Since the first conversion vector has not been subjected to linearization processing and projection by the linear layer, the first conversion vector can contain more features of the text to be matched and the text prompt word.
[0079] Here, the preset image-text matching model is obtained through training based on specific tasks. For example, when performing image-text matching in the field of industrial production, images of different industrial products can be used as images to be matched, names of different industrial products can be used as texts to be matched, and text prompt words can be used to train CLIP to obtain an image-text matching model.
[0080] Step 203, input the image to be matched into the image encoder of the image-text matching model to obtain the image vector and the second conversion vector output by the image encoder; wherein the dimensions of the image vector and the text vector are the same.
[0081] In this embodiment, the image-text matching model may further include an image encoder, and the image encoder may be used to encode the image to be matched to obtain a corresponding image vector.
[0082] The image encoder may include multiple image conversion layers and a linear layer. The image conversion layer is used to encode the image to be matched to obtain a second conversion vector. Since the second conversion vector is not processed by the linear layer, more features of the image to be matched can be retained in the second conversion vector.
[0083] The linear layer can transform the encoded vector and output a vector that meets actual needs, namely the image vector. The linear transformation of the linear layer in the image encoder and the text encoder can make the dimensions of the image vector and the text vector consistent, so as to facilitate the subsequent similarity calculation.
[0084] Step 204, determining a weight vector according to the first conversion vector and the second conversion vector; wherein the dimension of the weight vector is the same as the dimension of the text vector, and the weight vector is used to determine the weight of each dimension in the image vector and the text vector.
[0085] In this embodiment, the inventor of the present application takes into account that the addition of text prompt words in the above-mentioned encoding process will cause the text vector to be offset in different dimensions. Therefore, in order to improve the spatial consistency between the image vector and the text vector, a second transformation vector with more features of the image to be matched and a first transformation vector with more features of the text to be matched are used to determine a weight vector for correcting the dimension of the vector.
[0086] Here, the dimensions of the weight vector, image vector, and text vector are the same. The elements of different dimensions in the weight vector are weight factors. The weight factor of each dimension can adjust the elements of the corresponding dimension in the image vector and text vector, correct the offset caused by the text prompt words, and increase the weights of important features.
[0087] Step 205: Determine the similarity between the image to be matched and the text to be matched according to the image vector, the text vector and the weight vector.
[0088] In this embodiment, in order to adjust the image vector and the text vector by the weight vector, the weight vector is also added when determining the similarity between the image to be matched and the text to be matched, thereby affecting the similarity between the image to be matched and the text to be matched.
[0089] Here, the similarity can be determined using cosine similarity.
[0090] Step 206: Determine the matching result of the image to be matched and the text to be matched based on the similarity.
[0091] In this embodiment, when the similarity is higher than a preset threshold, it can be determined that the image to be matched and the text to be matched are matched.
[0092] For example, in industrial production, by photographing the production line, the image to be matched is obtained, and the text to be matched is the product produced by the production line. If the image to be matched matches the text to be matched, it can be determined that the production line is performing a production task; if the image to be matched does not match the text to be matched, it can be further determined whether the production line is faulty or stopped, etc., which facilitates subsequent processing.
[0093] In the embodiment of the present application, a text encoder is used to encode the text to be matched and the text prompt word to obtain a text vector, and an image encoder is used to encode the image to be matched to obtain an image vector; a first conversion vector and a second conversion vector are then used to determine a weight vector, which can fully consider various features in the text to be matched and the image to be matched, determine the correlation between the encoded text vector and the image vector, and consider the offset caused by the prompt word on each dimension in the vector; the similarity between the image to be matched and the text to be matched is determined through the image vector, the text vector and the weight vector, and the weight vector can be used to eliminate the offset caused by the prompt word. At the same time, attention can be paid to important features in the image to be matched and the text to be matched, the consistency of the image space and the text space can be improved, and the expression ability of the model can be guaranteed, so as to accurately obtain the similarity between the image to be matched and the text to be matched, so as to accurately match the image and the text according to the similarity.
[0094] In some embodiments, the weight vector is determined according to the first transformation vector and the second transformation vector. The key vector and value vector of the cross-attention mechanism are generated according to the first transformation vector; the query vector of the cross-attention mechanism is generated according to the second transformation vector; the query vector, the key vector and the value vector are cross-attention calculated to obtain the cross vector; and the cross vector is extracted and projected to obtain the weight vector.
[0095] In this embodiment, the cross-attention mechanism can be used to determine the weight vector to better focus on the image vector and the text vector to ensure the consistency of the image and text spaces.
[0096] The core of the cross-attention mechanism is to calculate the attention weights between different vectors. According to the calculated attention weights, the first transformation vector and the second transformation vector can be weighted fused, so that the relevant areas in the image can be focused on according to the semantics of the text, or the semantics of the text can be focused on according to the relevant areas in the image, thereby achieving effective interaction and fusion of image and text.
[0097] Here, by performing a linear transformation on the first transformation vector, a key vector and a value vector can be generated, and by performing a linear transformation on the second transformation vector, a query vector can be generated; the cross vector is the attention weight obtained by cross-attention calculation using the query vector, key vector and value vector. By further extracting and projecting the cross vector, the attention weight can be scaled to an appropriate range to obtain a weight vector, so that the image vector and text vector can be subsequently corrected.
[0098] Since the dimensions of the key vector and the value vector are the same, but the dimensions of the key vector and the query vector are different, we can extract the first one of the sequence and then do a projection to project the cross vector onto the required dimension. For example, Linear(d, 512) can be used for projection to 512 dimensions.
[0099] In addition, in this embodiment, considering the addition of text prompt words, the text vector is adjusted using the text prompt words to increase the focus on important features. Therefore, it is considered to use the first splicing vector to generate the key vector and value vector in the cross-attention mechanism, and use the second splicing vector to generate the query vector in the cross-attention mechanism, and perform cross-attention calculations to improve the consistency of the image and text space.
[0100] Optionally, this embodiment calculates the query vector, the key vector, and the value vector according to the cross attention mechanism to obtain a cross vector, which can be based on the expression: Get the cross vector; where logits represents the cross vector; softmax represents the activation function in the cross attention mechanism; Q represents the query vector; K represents the key vector; V represents the value vector; d k represents the dimension of the key vector; T represents the transpose of the vector.
[0101] Here, see Figure 3 The comparison diagram before and after the weight factor adjustment is shown in the figure. In the image text matching model before the weight vector is set, no weight vector is set. In fact, the weight vector in the similarity calculation can be used as 1, which has no effect on the text vector and image vector. For reference, Figure 3In the left part, each rectangle corresponds to a weight factor of a dimension, and the length of the rectangle indicates the size of the weight factor. The lengths of the rectangles on the left are equal, that is, the weight factors are the same. After the weight vector is determined in the above way, the weight factors of each dimension will fluctuate around 1, thereby adjusting the elements of the corresponding dimensions in the text vector and the image vector, so that the text vector and the image vector can be more matched. For the corresponding reference, Figure 3 In the middle right part, each rectangle also corresponds to a weight factor of a dimension. The length of the rectangle represents the size of the weight factor. The lengths of the rectangles on the right are different. The dotted line on the right side of the rectangle represents the reduced part of the weight factor after adjustment, that is, the weight factor becomes smaller than the corresponding weight factor on the left; the solid line on the right side of the rectangle represents the increased part of the weight factor after adjustment, that is, the weight factor becomes larger than the corresponding weight factor on the left.
[0102] In some embodiments, the text vector includes text elements of multiple dimensions, the image vector includes image elements of multiple dimensions, and the weight vector includes weight factors of multiple dimensions. Here, the number of dimensions of the text vector, the image vector, and the weight vector is equal, for example, can be set to 128 dimensions or 512 dimensions, etc.
[0103] See also Figure 4 The structural diagram of the image-text matching model shown in the figure, this embodiment determines the similarity between the image to be matched and the text to be matched based on the image vector, the text vector and the weight vector, and can be based on the text elements of each dimension in the text vector, the image elements of the corresponding dimension in the image vector and the weight factors of the corresponding dimension in the weight vector, calculate the similarity score between the text vector and the image vector in each dimension; based on the similarity scores of all dimensions, determine the similarity between the image to be matched and the text to be matched.
[0104] In this embodiment, the text vector output by the text encoder of the image-text matching model consists of elements of n dimensions, namely Figure 4 The y in 1 ,y 2 ,y 3 ,……,y n The image vector output by the image encoder of the image-text matching model is also composed of elements of n dimensions, that is, Figure 4 x in 1 、x 2 、x 3 ,……,x n The weight vector obtained by cross-attention calculation of text vector and image vector is also composed of elements of n dimensions, that is, Figure 4 α in 1 , α 2 , α3 ,……,α n By adjusting and correcting the text vector and image vector of the corresponding dimension through the weight factor of each dimension, the text space and the image space can be aligned, so as to accurately calculate the similarity score between the image to be matched and the text to be matched in each dimension, thereby obtaining the similarity between the image to be matched and the text to be matched.
[0105] Optionally, this embodiment may be based on the expression: Calculate the similarity between the image to be matched and the text to be matched; where cosθ j Indicates the similarity between the image to be matched and the text to be matched, x i Represents the element of the i-th dimension in the image vector corresponding to the image to be matched, y i Represents the element of the i-th dimension in the text vector corresponding to the text to be matched, α i Represents the weight factor of the i-th dimension in the weight vector, i represents the i-th dimension, and n represents the set of all dimensions.
[0106] In this embodiment, the similarity between the image to be matched and the text to be matched can be directly calculated according to the above formula, and the weight factor, the element of the corresponding dimension in the image vector and the element of the corresponding dimension in the text vector are directly multiplied in the numerator to adjust the image vector and the text vector using the weight factor. Then, by calculating the cosine similarity, the similarity between the image to be matched and the text to be matched can be obtained.
[0107] In some embodiments, the image to be matched is input into the image encoder of the image-text matching model to obtain the image vector and the second conversion vector output by the image encoder. The image prompt word can be obtained first; then the image to be matched and the image prompt word are input into the image encoder of the image-text matching model to obtain the image vector and the second conversion vector output by the image encoder.
[0108] In this embodiment, considering that text prompt words can guide text vectors, image prompt words can also be added, and the image prompt words are used to guide the image vectors. The image prompt words and the image to be matched are input into the image encoder for encoding, so as to focus on the corresponding areas or features in the image to be matched, thereby improving the interaction and fusion of image and text.
[0109] Here, in order to make the text space and the image space more matched, the text prompt words can be mapped to the image space through a coupling function, and then the image prompt words can be obtained by training the image-text matching model; the value of the image to be matched can also be initialized by initializing it to 0 or Gaussian distribution, and the final image prompt words can be obtained by training the image-text matching model.
[0110] In some embodiments, see Figure 4 As shown in the structural diagram of the image-text matching model, the text encoder includes multiple text conversion layers, and a text prompt word can be correspondingly set in each text conversion layer to guide the text vector in each text conversion layer of the text encoder.
[0111] Similarly, the image encoder may also include multiple image conversion layers and multiple visual prompt words, wherein each image conversion layer except the first image conversion layer in the image encoder may be correspondingly provided with a visual prompt word.
[0112] In addition, the number of image conversion layers in the image encoder and the number of text conversion layers in the text encoder may be the same or different. If the number of image conversion layers is the same as the number of text conversion layers, the two may correspond. In this case, the image prompt words may be obtained by coupling mapping the text prompt words, and for a set of text prompt words and visual prompt words, the number of conversion layers they play a role in the text encoder and the image encoder may be the same.
[0113] Taking the text encoder as an example, before each text conversion layer, a text prompt word can be added to the vector corresponding to the text. Figure 4 The rectangle with dots under each text conversion layer represents the set text prompt words, and the rectangles on its left and right represent the text to be matched or the output of the previous text conversion layer. The rectangle under the first text conversion layer is specifically the vector corresponding to the text to be matched, which is obtained after the text vector mapping. The rectangles of the second layer and later are specifically the output of the previous text conversion layer of the rectangle (the text conversion layer under the rectangle). By replacing and adding text prompt words in the process of converting the text to be matched, a more accurate text vector can be obtained.
[0114] In addition, in the text encoder, the rectangle with the regular triangle is also part of the vector output by the text conversion layer, representing the EOT tag, which is used to indicate the end of the text. After the text to be matched is mapped through the text vector, the EOT tag can be directly obtained from the vector determined by the mapping.
[0115] In the image encoder, a rectangle with an inverted triangle represents a CLS tag, which is used to indicate a category or level tag and is generally set at the beginning of a sequence or vector. After the image to be matched is mapped to the image vector, a CLS tag can be added to the front of the vector determined by the mapping and input to the first image conversion layer as a complete vector. The CLS tag after the first image conversion layer is also part of the vector output by the image conversion layer, and the CLS tag processed by all image conversion layers can represent the global features of the entire image.
[0116] Here, the text prompt words corresponding to each text conversion layer of the text encoder and the image prompt words corresponding to each image conversion layer of the image encoder can be trained according to specific tasks to make the image-text matching model more suitable for the selected specific tasks. In addition, other parameters in the text encoder and image encoder of the image-text matching model can be frozen (see Figure 4 , where the snowflake symbol indicates parameter freezing).
[0117] Optionally, in this embodiment, the text to be matched and the text prompt word are input into a text encoder of a preset image-text matching model to obtain a text vector and a first conversion vector output by the text encoder, which can be:
[0118] Step 1: Concatenate the text to be matched and the text prompt word corresponding to the first text conversion layer to obtain a concatenated vector.
[0119] In this embodiment, the initial prompt word can be concatenated into the text to be matched as a vector input to the first text conversion layer, so that the text encoder can process the text to be matched and the text prompt word.
[0120] Step 2: Input the concatenated vector to the first text conversion layer in the text encoder to obtain a third conversion vector output by the first text conversion layer.
[0121] Here, the concatenated vector is processed using the first text transformation layer.
[0122] Step three, replace the text element at the first preset position in the third conversion vector with the text prompt word corresponding to the second text conversion layer, and output the replaced vector to the second text conversion layer to obtain the fourth conversion vector output by the second text conversion layer, and complete the conversion processing corresponding to one text conversion layer; continue the conversion processing until the first conversion vector output by the last text conversion layer is obtained.
[0123] Here, when adding the text prompt words corresponding to each text conversion layer, the text prompt words can be spliced into the vector input to the text conversion layer by replacement. Specifically, in the processing of the second text conversion layer, the text element at the first preset position in the third conversion vector input to the text conversion layer can be removed, and then the text prompt words corresponding to the text conversion layer can be added to the position of the removed text element, and then input to the second text conversion layer to obtain a fourth conversion vector. For example, the second position in the corresponding vector can be replaced with the text prompt words.
[0124] The processing is performed in each text conversion layer in turn, so that all the text prompt words are spliced into the text to be matched, thereby realizing the conversion processing of the text to be matched.
[0125] In addition, since the encoded length of the to-be-matched text is relatively small, the text prompt word may be directly concatenated to the vector input to the text conversion layer, for example, concatenated to the end of the vector.
[0126] Step three, based on the first conversion vector, obtain the text vector output by the text encoder.
[0127] In this embodiment, the first conversion vector output by the last text conversion layer can be used as the text vector output by the text encoder. In addition, the text encoder can also include a linear layer. After obtaining the first conversion vector, the EOT mark in the first conversion vector can be input into the linear layer for linear transformation to obtain a text vector of a required dimension, such as 256 dimensions, 512 dimensions, and 1024 dimensions.
[0128] Accordingly, when setting the visual cue words, the steps of obtaining the image vector using the image encoder are as follows.
[0129] Optionally, in this embodiment, the image to be matched is input into the image encoder of the image-text matching model to obtain the image vector and the second conversion vector output by the image encoder, which can be:
[0130] Step 1: Acquire multiple visual prompt words; wherein, each image conversion layer except the first image conversion layer is correspondingly set with a visual prompt word.
[0131] Here, the visual prompt words can be obtained through a pre-trained model, or can be obtained by mapping text prompt words to the image space through a coupling function and then performing model training.
[0132] Step 2: Input the image to be matched into the first image conversion layer in the image encoder to obtain a fifth conversion vector.
[0133] In this embodiment, the first image conversion layer may be used to perform encoding conversion on the image to be matched to obtain the fifth conversion vector.
[0134] Here, before the image to be matched is input into the first image conversion layer, the image to be matched can be encoded through image vector mapping, and a CLS marker is added to the front of the encoded vector.
[0135] Step three, replace the image element at the second preset position in the fifth conversion vector with the visual prompt word corresponding to the second image conversion layer, and input the replaced vector into the second image conversion layer to obtain the sixth conversion vector, completing the conversion processing corresponding to one image conversion layer; continue the conversion processing until the second conversion vector output by the last image conversion layer is obtained.
[0136] In this embodiment, since the length of the vector corresponding to the image is large, direct splicing will increase the length of the vector to a greater extent. Therefore, before the second image conversion layer and each subsequent image conversion layer, the image element at the end of the vector input to the image conversion layer can be replaced with the visual prompt word corresponding to the image conversion layer, thereby combining the image prompt word with the image to be matched to achieve conversion processing of the image to be matched.
[0137] Step 4: Based on the second conversion vector, obtain the image vector output by the image encoder.
[0138] In this embodiment, the second conversion vector output by the last image conversion layer can be used as the image vector output by the image encoder. Since the dimension of the second splicing vector is large and the dimension of the second splicing vector is different from that of the text vector, it is difficult to directly calculate the similarity. A linear layer can also be set in the image encoder to perform a linear transformation on the second conversion vector to obtain an image vector of the required dimension, such as 256 dimensions, 512 dimensions, and 1024 dimensions.
[0139] Optionally, since the CLS marker in the second conversion vector output by the last image conversion layer can represent the global features of the image, the CLS marker can be input into the linear layer for transformation.
[0140] In some embodiments, the cross entropy may be used as a loss function to train the model, thereby obtaining a trained image-text matching model to match images and texts.
[0141] In this embodiment, the minimum value of the loss function can be solved by gradient descent, and the model training can be considered completed when the loss function reaches the minimum value.
[0142] The loss function can be expressed as Where L represents the loss function of the image-text matching model, which can be calculated using cross entropy. N represents the total number of all sample images during the training of the image-text matching model. n represents the nth sample image among all sample images, and its value can be 0, 1, 2, ..., N-1. n represents the true similarity between the nth sample image and each sample text, p n represents the predicted similarity between the nth sample image and each sample text, F represents the image-text matching model, α represents the weight vector, I represents the input sample image of the image-text matching model, T represents the input sample text of the image-text matching model, and θ represents other parameters, such as hyperparameters or prompt words.
[0143] Here, the gradient descent method can adopt the stochastic gradient descent method, and the stochastic gradient descent method and the loss function are used to adjust the hyperparameters or prompt words in the image text matching model, which can be based on the expression: Adjust the hyperparameters or prompt words; where θ i+1 represents the hyperparameter or prompt word for the i+1th iteration in the stochastic gradient descent method, θ i represents the hyperparameter or prompt word of the i-th iteration in the stochastic gradient descent method, η represents the learning rate in the stochastic gradient descent method, represents the partial derivative of the loss function.
[0144] The above mainly introduces image-text matching. Image-text matching can be applied to specific tasks such as image classification and image retrieval. The following takes the image classification task as an example for a detailed explanation.
[0145] In some embodiments, the text to be matched includes multiple image tags, and the text prompt word includes a tag prompt word corresponding to each image tag. Here, the image text matching method provided in the above embodiments can be applied to the task of image classification. In this case, the text to be matched can be multiple image tags to divide the image into corresponding image tags. Among them, the image label can be determined according to the object in the image to be matched, that is, the image label can be the name of the object; it can also be determined according to the requirements of the specific classification task. For example, when the image to be matched needs to be classified by color, the image label can be different colors.
[0146] In this embodiment, the text to be matched and the text prompt word are input into the text encoder of the preset image-text matching model to obtain the text vector and the first conversion vector output by the text encoder. Alternatively, each image label and the corresponding label prompt word are input into the text encoder to obtain the text vector and the first conversion vector corresponding to each image label output by the text encoder.
[0147] Here, each image label and the corresponding label prompt word can be input into the text encoder to encode the image label, so as to match the image to be matched with each image label and find the most similar image label, thereby realizing image classification.
[0148] Correspondingly, in this embodiment, after inputting the image to be matched into the image encoder of the image-text matching model and obtaining the image vector and the second conversion vector output by the image encoder, the weight vectors of the image vector and each image label are respectively determined according to the second conversion vector and the first conversion vector corresponding to each image label; the similarity between the image to be matched and each image label is determined according to the image vector, the text vector corresponding to each image label and the corresponding weight vector; and the image to be matched is classified based on the similarity between the image to be matched and each image label.
[0149] In this embodiment, when classifying images, the obtained second transformation vector can be matched with the first transformation vector corresponding to each image label, that is, the weight vector and the similarity are calculated. The specific steps can refer to the steps in the above embodiments and will not be repeated here.
[0150] When classifying images, the image label with the highest similarity between the image to be matched and each image label may be determined as the image label to which the image to be matched is classified.
[0151] In the embodiment of the present application, a text encoder is used to encode the text to be matched and the text prompt word to obtain a text vector, and an image encoder is used to encode the image to be matched to obtain an image vector; then the first conversion vector and the second conversion vector are used to determine the weight vector, which can fully consider the various features in the text to be matched and the image to be matched, determine the correlation between the encoded text vector and the image vector, and consider the offset of each dimension in the vector caused by the prompt word; through the image vector, the text vector and the weight vector, the similarity between the image to be matched and the text to be matched is determined, and the weight vector can be used to eliminate the offset caused by the prompt word, and at the same time, the important features in the image to be matched and the text to be matched can be strengthened, the consistency of the image space and the text space can be improved, and the expression ability of the model can be guaranteed, so as to accurately obtain the similarity between the image to be matched and the text to be matched, so as to accurately match the image and the text according to the similarity. Among them, by using cross attention, the first conversion vector and the second conversion vector to determine the weight vector, the weight vector can be used to correct the influence of the prompt word on each dimension, and at the same time, the attention to important features can be strengthened, the consistency of the image and text space can be improved, and the image and text can be accurately matched. At the same time, the image-text matching method provided in this embodiment can also be applied to tasks such as image classification and image retrieval to improve the accuracy of image classification and image retrieval.
[0152] It should be understood that the size of the serial numbers of the steps in the above embodiments does not mean the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present application.
[0153] The following is an embodiment of the device of the present application. For details not described in detail, please refer to the corresponding method embodiment described above.
[0154] Figure 5 The schematic diagram of the structure of the cross-modal aligned image text matching device provided in the embodiment of the present application is shown. For the convenience of explanation, only the part related to the embodiment of the present application is shown, which is described in detail as follows:
[0155] like Figure 5 As shown, the cross-modal aligned image-text matching device 50 comprises:
[0156] An acquisition module 51 is used to acquire an image to be matched, a text to be matched, and a text prompt word;
[0157] A text encoding module 52, used for inputting the to-be-matched text and the text prompt word into a text encoder of a preset image-text matching model, and obtaining a text vector and a first conversion vector output by the text encoder;
[0158] An image encoding module 53 is used to input the image to be matched into the image encoder of the image-text matching model, and obtain an image vector and a second conversion vector output by the image encoder; wherein the image vector and the text vector have the same dimension;
[0159] A weight determination module 54, configured to determine a weight vector according to the first conversion vector and the second conversion vector; wherein the dimension of the weight vector is the same as the dimension of the text vector, and the weight vector is used to determine the weight of each dimension in the image vector and the text vector;
[0160] The matching module 55 is used to determine the similarity between the image to be matched and the text to be matched according to the image vector, the text vector and the weight vector; and to determine the matching result between the image to be matched and the text to be matched based on the similarity.
[0161] In a possible implementation, the weight determination module 54 is specifically configured to:
[0162] Generate a key vector and a value vector of a cross attention mechanism according to the first transformation vector;
[0163] Generate a query vector of a cross-attention mechanism according to the second transformation vector;
[0164] Perform cross attention calculation on the query vector, key vector and value vector to obtain the cross vector;
[0165] The cross vector is extracted and projected to obtain the weight vector.
[0166] In a possible implementation, the weight determination module 54 is specifically configured to:
[0167] According to the expression: Get the cross vector;
[0168] In the formula, logits represents the cross vector; softmax represents the activation function in the cross attention mechanism; Q represents the query vector; K represents the key vector; V represents the value vector; d k represents the dimension of the key vector; T represents the transpose of the vector.
[0169] In a possible implementation, the text vector includes text elements of multiple dimensions, the image vector includes image elements of multiple dimensions, and the weight vector includes weight factors of multiple dimensions;
[0170] The matching module 55 is specifically used for:
[0171] Calculate the similarity score of the text vector and the image vector in each dimension according to the text elements of each dimension in the text vector, the image elements of the corresponding dimension in the image vector and the weight factors of the corresponding dimension in the weight vector;
[0172] According to the similarity scores of all dimensions, the similarity between the image to be matched and the text to be matched is determined.
[0173] In a possible implementation, the text encoder includes a plurality of text conversion layers, each text conversion layer being provided with a corresponding text prompt word;
[0174] The text encoding module 52 is specifically used for:
[0175] Concatenate the text to be matched and the text prompt word corresponding to the first text conversion layer to obtain a concatenated vector;
[0176] Input the concatenated vector to the first text conversion layer in the text encoder to obtain a third conversion vector output by the first text conversion layer;
[0177] The text element at the first preset position in the third conversion vector is replaced with the text prompt word corresponding to the second text conversion layer, and the replaced vector is output to the second text conversion layer to obtain a fourth conversion vector output by the second text conversion layer, and the conversion process of one text conversion layer is completed; the conversion process is continuously performed until the first conversion vector output by the last text conversion layer is obtained;
[0178] Based on the first conversion vector, a text vector output by the text encoder is obtained.
[0179] In one possible implementation, the image encoder includes a plurality of image conversion layers;
[0180] The image encoding module 53 is specifically used for:
[0181] Acquire multiple visual prompt words; wherein each image conversion layer except the first image conversion layer is correspondingly set with a visual prompt word;
[0182] Inputting the image to be matched into the first image conversion layer in the image encoder to obtain a fifth conversion vector;
[0183] The image element at the second preset position in the fifth conversion vector is replaced with the visual prompt word corresponding to the second image conversion layer, and the replaced vector is input into the second image conversion layer to obtain the sixth conversion vector, thus completing the conversion process corresponding to one image conversion layer; the conversion process is continuously performed until the second conversion vector output by the last image conversion layer is obtained;
[0184] Based on the second conversion vector, an image vector output by the image encoder is obtained.
[0185] In a possible implementation, the text to be matched includes multiple image tags, and the text prompt word includes a tag prompt word corresponding to each image tag;
[0186] The text encoding module 52 is specifically used for:
[0187] Input each image label and the corresponding label prompt word into the text encoder, and obtain the text vector and the first conversion vector corresponding to each image label output by the text encoder;
[0188] The cross-modal aligned image-text matching apparatus 40 further includes a classification module, which is used to:
[0189] Determine the weight vectors of the image vector and each image label respectively according to the second transformation vector and the first transformation vector corresponding to each image label;
[0190] Determine the similarity between the image to be matched and each image label according to the image vector, the text vector corresponding to each image label and the corresponding weight vector;
[0191] Based on the similarity between the image to be matched and each image label, the image to be matched is classified.
[0192] Figure 6 Schematic diagram of an electronic device provided in an embodiment of the present application. Figure 6 As shown, the electronic device 60 of this embodiment includes: a processor 61, a memory 62, and a computer program 63 stored in the memory 62 and executable on the processor 61. When the processor 61 executes the computer program 63, the steps in the above-mentioned various cross-modal aligned image text matching method embodiments are implemented, for example Figure 2 Alternatively, when the processor 61 executes the computer program 63, the functions of each module in the above-mentioned device embodiments are realized, for example, Figure 5 The functions of modules 51 to 55 are shown.
[0193] Exemplarily, the computer program 63 may be divided into one or more modules / units, one or more modules / units are stored in the memory 62 and executed by the processor 61 to complete the present application. The one or more modules / units may be a series of computer program instruction segments that can complete specific functions, and the instruction segments are used to describe the execution process of the computer program 63 in the electronic device 60. For example, the computer program 63 may be divided into Figure 5 Modules 51 to 55 are shown.
[0194] The electronic device 60 may include, but is not limited to, a processor 61 and a memory 62. Those skilled in the art will appreciate that Figure 6 It is only an example of the electronic device 60 and does not constitute a limitation of the electronic device 60. It may include more or fewer components than shown in the figure, or a combination of certain components, or different components. For example, the electronic device may also include input and output devices, network access devices, buses, etc.
[0195] The processor 61 may be a central processing unit (CPU), or other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), field-programmable gate arrays (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. A general-purpose processor may be a microprocessor or the processor may be any conventional processor, etc.
[0196] The memory 62 may be an internal storage unit of the electronic device 60, such as a hard disk or memory of the electronic device 60. The memory 62 may also be an external storage device of the electronic device 60, such as a plug-in hard disk, a smart media card (SMC), a secure digital (SD) card, a flash card, etc. equipped on the electronic device 60. Further, the memory 62 may also include both an internal storage unit of the electronic device 60 and an external storage device. The memory 62 is used to store computer programs and other programs and data required by the electronic device. The memory 62 may also be used to temporarily store data that has been output or is to be output.
[0197] For the convenience and simplicity of description, only the division of the above functional units and modules is used as an example for illustration. In actual applications, the above functions can be assigned to different functional units / modules as needed. The above integrated units can be implemented in the form of hardware, software functional units, or a combination of hardware and software.
[0198] The embodiment of the present invention further provides a computer-readable storage medium, which stores a computer program. When the computer program is executed by a processor, the methods in the above method embodiments are implemented.
[0199] The embodiment of the present invention further provides a computer program product, including a computer program. When the computer program is executed by a processor, the methods in the above method embodiments are implemented.
[0200] The computer program includes computer program code, which may be in source code form, object code form, executable file or some intermediate form, etc. Computer readable media may include: any entity or device capable of carrying computer program code, recording medium, USB flash drive, mobile hard disk, magnetic disk, optical disk, computer memory, read-only memory (ROM), random access memory (RAM), electric carrier signal, telecommunication signal and software distribution medium, etc.
[0201] In the above embodiments, the description of each embodiment has its own emphasis. For parts that are not described or recorded in detail in a certain embodiment, reference can be made to the relevant descriptions of other embodiments. If there is no special explanation and logical conflict, the terms and / or descriptions between different embodiments are consistent and can be referenced to each other. The technical features in different embodiments can be combined to form a new embodiment according to their internal logical relationship.
[0202] The above embodiments are only used to illustrate the technical solutions of the present application, rather than to limit them. Although the present application has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. These modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present application, and should all be included in the protection scope of the present application.
Claims
1. A cross-modal aligned image-text matching method, characterized in that: include: Obtaining the image to be matched, the text to be matched and the text prompt word; Inputting the text to be matched and the text prompt word into a text encoder of a preset image-text matching model to obtain a text vector and a first conversion vector output by the text encoder; Input the image to be matched into the image encoder of the image-text matching model to obtain an image vector and a second conversion vector output by the image encoder; wherein the image vector and the text vector have the same dimension; Determine a weight vector according to the first conversion vector and the second conversion vector; wherein the dimension of the weight vector is the same as the dimension of the text vector, and the weight vector is used to determine the weight of each dimension in the image vector and the text vector; Determining the similarity between the image to be matched and the text to be matched according to the image vector, the text vector and the weight vector; Based on the similarity, a matching result between the image to be matched and the text to be matched is determined.
2. The cross-modal aligned image-text matching method according to claim 1, characterized in that: The determining of a weight vector according to the first conversion vector and the second conversion vector comprises: Generate a key vector and a value vector of a cross attention mechanism according to the first transformation vector; Generate a query vector of a cross-attention mechanism according to the second transformation vector; Performing cross attention calculation on the query vector, the key vector and the value vector to obtain a cross vector; The cross vector is extracted and projected to obtain a weight vector.
3. The cross-modal aligned image-text matching method according to claim 2, characterized in that: Performing cross attention calculation on the query vector, the key vector, and the value vector to obtain a cross vector, including: According to the expression: Get the cross vector; Wherein, logits represents the cross vector; softmax represents the activation function in the cross attention mechanism; Q represents the query vector; K represents the key vector; V represents the value vector; d k represents the dimension of the key vector; T represents the transpose of the vector.
4. The cross-modal aligned image-text matching method according to any one of claims 1 to 3, characterized in that: The text vector includes text elements of multiple dimensions, the image vector includes image elements of multiple dimensions, and the weight vector includes weight factors of multiple dimensions; Determining the similarity between the image to be matched and the text to be matched according to the image vector, the text vector and the weight vector includes: Calculate the similarity score of the text vector and the image vector in each dimension according to the text elements of each dimension in the text vector, the image elements of the corresponding dimension in the image vector and the weight factors of the corresponding dimension in the weight vector; The similarity between the image to be matched and the text to be matched is determined according to the similarity scores of all dimensions.
5. The cross-modal aligned image-text matching method according to any one of claims 1 to 3, characterized in that: The text encoder includes a plurality of text conversion layers, each of which is provided with a corresponding text prompt word; The step of inputting the text to be matched and the text prompt word into a text encoder of a preset image-text matching model to obtain a text vector and a first conversion vector output by the text encoder includes: Concatenate the text to be matched and the text prompt word corresponding to the first text conversion layer to obtain a concatenated vector; Inputting the concatenated vector into a first text conversion layer in the text encoder to obtain a third conversion vector output by the first text conversion layer; The text element at the first preset position in the third conversion vector is replaced with a text prompt word corresponding to the second text conversion layer, and the replaced vector is output to the second text conversion layer to obtain a fourth conversion vector output by the second text conversion layer, and the conversion process corresponding to one text conversion layer is completed; the conversion process is continuously performed until the first conversion vector output by the last text conversion layer is obtained; Based on the first conversion vector, a text vector output by the text encoder is obtained.
6. The cross-modal aligned image-text matching method according to any one of claims 1 to 3, characterized in that: The image encoder includes a plurality of image conversion layers; The step of inputting the image to be matched into the image encoder of the image-text matching model to obtain the image vector and the second conversion vector output by the image encoder comprises: Acquire multiple visual prompt words; wherein each image conversion layer except the first image conversion layer is correspondingly set with a visual prompt word; Inputting the image to be matched into a first image conversion layer in the image encoder to obtain a fifth conversion vector; The image element at the second preset position in the fifth conversion vector is replaced with a visual prompt word corresponding to the second image conversion layer, and the replaced vector is input to the second image conversion layer to obtain a sixth conversion vector, completing the conversion process corresponding to one image conversion layer; the conversion process is continuously performed until the second conversion vector output by the last image conversion layer is obtained; Based on the second conversion vector, an image vector output by the image encoder is obtained.
7. The cross-modal aligned image-text matching method according to any one of claims 1 to 3, characterized in that: The text to be matched includes a plurality of image tags, and the text prompt words include the tag prompt words corresponding to each image tag; The step of inputting the text to be matched and the text prompt word into a text encoder of a preset image-text matching model to obtain a text vector and a first conversion vector output by the text encoder includes: Input each image label and the corresponding label prompt word into the text encoder, and obtain the text vector and the first conversion vector corresponding to each image label output by the text encoder; After inputting the image to be matched into the image encoder of the image-text matching model to obtain the image vector and the second conversion vector output by the image encoder, the method further includes: Determining weight vectors of the image vector and each image label respectively according to the second conversion vector and the first conversion vector corresponding to each image label; Determine the similarity between the image to be matched and each image label according to the image vector, the text vector corresponding to each image label and the corresponding weight vector; The image to be matched is classified based on the similarity between the image to be matched and each image label.
8. A cross-modal aligned image-text matching device, characterized in that: include: An acquisition module is used to acquire the image to be matched, the text to be matched and the text prompt word; A text encoding module, used for inputting the text to be matched and the text prompt word into a text encoder of a preset image-text matching model, and obtaining a text vector and a first conversion vector output by the text encoder; An image encoding module, used for inputting the image to be matched into the image encoder of the image-text matching model to obtain an image vector and a second conversion vector output by the image encoder; wherein the image vector and the text vector have the same dimension; A weight determination module, configured to determine a weight vector according to the first conversion vector and the second conversion vector; wherein the dimension of the weight vector is the same as the dimension of the text vector, and the weight vector is used to determine the weight of each dimension in the image vector and the text vector; A matching module is used to determine the similarity between the image to be matched and the text to be matched according to the image vector, the text vector and the weight vector; and to determine the matching result of the image to be matched and the text to be matched based on the similarity.
9. An electronic device comprising a memory and a processor, wherein the memory is used to store a computer program, and the processor is used to call and run the computer program stored in the memory, characterized in that: When the processor executes the computer program, the steps of the method according to any one of claims 1 to 7 are implemented.
10. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 7 are implemented.