Method, device and equipment for establishing image-text retrieval model and storage medium

By generating the probability distribution of images and text and guiding cross-modal similarity learning, the problem of ignoring the large in-modal similarity and computational overhead in the prior art is solved, and more accurate and efficient picture-text matching is achieved.

CN120162471APending Publication Date: 2025-06-17GUANGZHOU UNIVERSITY
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510144864.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-10
Publication Date
2025-06-17

AI Technical Summary

Technical Problem

The prior art ignores similarity within the modality in graphic and text retrieval, resulting in limited application when understanding and matching complex visual and semantic patterns, and only focusing on positive sample pairs when aligning across modal features, there is error, and the feature information generated by external large models introduces a lot of computational overhead.

Method used

By using the teacher model to generate the image and text feature set of the training set, the aligned feature sets of images and text are encoded, and the distribution of images to text and text to images is calculated, and the probability distribution of images and text is fused together with external and internal distributions to generate images and text. These probability distributions are used to guide cross-modal similarity learning and construct loss functions to adjust depth model parameters.

Benefits of technology

The semantic relationships within the modality of the image and text are fully described, which enhances modal consistency, improves the accuracy of matching, and reduces the computational burden through a lightweight scheme, and improves the efficiency of picture and text matching.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120162471A_ABST
    Figure CN120162471A_ABST
Patent Text Reader

Abstract

The invention provides an image-text retrieval model establishment method and device, equipment and a storage medium, and the method comprises the steps: generating a teacher image feature set and a teacher text feature set; encoding to obtain an image encoding set and a text encoding set, and matching an image alignment feature set corresponding to the image encoding set and a text alignment feature set corresponding to the text encoding set by using probability distribution; generating an image-to-text distribution and a text-to-image distribution based on calculation between the image alignment feature set and the text alignment feature set; calculating and generating distribution among external images based on the teacher image feature set, and calculating and generating distribution among external texts based on the teacher text feature set; internal image distribution is generated based on calculation in the image coding set, and internal text distribution is generated based on calculation in the text coding set; fusing to generate image probability distribution and text probability distribution; the image probability distribution is used to guide the image-to-text distribution, and the text probability distribution is used to guide the text-to-image distribution.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of graphic and text retrieval, and particularly to a method, device, equipment and storage medium for establishing a graphic and text retrieval model. Background Art

[0002] Graphic and text retrieval is a cross-modal information retrieval technology aimed at realizing mutual retrieval between images and texts. Specifically, it can find the most relevant text description according to a picture, or query the most matching picture according to a piece of text.

[0003] In the related art, traditional graphic and text retrieval methods mainly focus on cross-modal alignment, that is, how to map image and text features to a common semantic space. These methods usually ignore the intra-modal similarity, that is, the inherent semantic relationship between images or between texts, which greatly limits their application in understanding and matching complex visual and semantic patterns. Some methods only focus on positive sample pairs during cross-modal feature alignment, resulting in errors in the matching process. In addition, using the feature information generated by an external large model as a soft label index for cross-modal learning introduces a large amount of unnecessary computational overhead in the feature fusion and knowledge distillation processes.

[0004] Based on the above analysis of the development status of this technical field, the existing technologies lack a solution that considers intra-modal similarity, sets specific feature enhancement steps to increase the difference between positive and negative samples, and guides cross-modal similarity learning through soft label distillation. Summary of the Invention

[0005] The purpose of the present invention is to provide a method, device, equipment and storage medium for establishing a graphic and text retrieval model, aiming to solve the above problems in the existing technology.

[0006] According to the first aspect of the embodiments of the present invention, a method for establishing a graphic and text retrieval model is provided, including:

[0007] Generating a teacher image feature set and a teacher text feature set corresponding to a training set by using a teacher model;

[0008] Encoding the training set to obtain an image encoding set and a text encoding set respectively, and using probability distribution to match the image alignment feature set corresponding to the image encoding set and the text alignment feature set corresponding to the text encoding set;

[0009] Calculating based on the image alignment feature set and the text alignment feature set to generate an image-to-text distribution and a text-to-image distribution;

[0010] Generate the distribution between external images based on the calculation within the teacher image feature set, and generate the distribution between external texts based on the calculation within the teacher text feature set; generate the distribution between internal images based on the calculation within the image encoding set, and generate the distribution between internal texts based on the calculation within the text encoding set;

[0011] Fuse the distribution between external images and the distribution between internal images to generate an image probability distribution, and fuse the distribution between external texts and the distribution between internal texts to generate a text probability distribution;

[0012] Use the image probability distribution to guide the image-to-text distribution, use the text probability distribution to guide the text-to-image distribution, construct a loss function to adjust the parameters of the deep model, and obtain a graph-text retrieval model after the iteration ends.

[0013] According to the second aspect of the embodiments of the present invention, there is provided an apparatus for establishing a graph-text retrieval model, including:

[0014] A teacher module for using a teacher model to generate a teacher image feature set and a teacher text feature set corresponding to a training set;

[0015] A probability distribution matching module for encoding the training set to obtain an image encoding set and a text encoding set respectively, and using the probability distribution to match the image alignment feature set corresponding to the image encoding set and the text alignment feature set corresponding to the text encoding set;

[0016] A cross-domain distribution module for generating an image-to-text distribution and a text-to-image distribution based on the calculation between the image alignment feature set and the text alignment feature set;

[0017] A same-modal distribution module for generating the distribution between external images based on the calculation within the teacher image feature set, and generating the distribution between external texts based on the calculation within the teacher text feature set; generating the distribution between internal images based on the calculation within the image encoding set, and generating the distribution between internal texts based on the calculation within the text encoding set;

[0018] A fusion module for fusing the distribution between external images and the distribution between internal images to generate an image probability distribution, and fusing the distribution between external texts and the distribution between internal texts to generate a text probability distribution;

[0019] A guidance training module for using the image probability distribution to guide the image-to-text distribution, using the text probability distribution to guide the text-to-image distribution, constructing a loss function to adjust the parameters of the deep model, and obtaining a graph-text retrieval model after the iteration ends.

[0020] According to the third aspect of the embodiments of the present invention, there is provided an electronic device, including: a memory, a processor, and a computer program stored on the memory and executable on the processor, and when the computer program is executed by the processor, it implements the steps of the method for establishing a graph-text retrieval model provided in the first aspect of the present disclosure.

[0021] According to a fourth aspect of an embodiment of the present invention, there is provided a computer-readable storage medium, on which an implementation program for information transmission is stored. When the program is executed by a processor, the steps of the method for establishing a graphic and text retrieval model provided in the first aspect of the present disclosure are implemented.

[0022] The technical solutions provided by the embodiments of the present invention include the following beneficial effects: For features of two types, namely images and texts, probability distributions that are independent of each other and non-cross-modal are generated, that is, an image probability distribution and a text probability distribution, which fully describe the semantic relationships within the modality of the image or text itself, adding modality itself consistency to improve the accuracy of matching; using the probability distribution to guide cross-modal similarity learning, and realizing soft label guidance through a lightweight scheme, reducing the computational burden to improve the efficiency of graphic and text matching.

[0023] It should be understood that the above general description and the following detailed description are only exemplary and explanatory, and cannot limit the present disclosure. BRIEF DESCRIPTION OF THE DRAWINGS

[0024] In order to more clearly illustrate the technical solutions in one or more embodiments of this specification or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are only some embodiments recorded in this specification. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0025] Figure 1 is a flowchart of the method for establishing a graphic and text retrieval model according to an embodiment of the present invention;

[0026] Figure 2 is a schematic diagram of the model architecture according to an embodiment of the present invention;

[0027] Figure 3 is a schematic diagram of the device for establishing a graphic and text retrieval model according to an embodiment of the present invention;

[0028] Figure 4 is a schematic diagram of an electronic device according to an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0029] In order to enable those skilled in the art to better understand the technical solutions in one or more embodiments of this specification, the following will clearly and completely describe the technical solutions in one or more embodiments of this specification with reference to the drawings in one or more embodiments of this specification. Obviously, the described embodiments are only a part of the embodiments of this specification, rather than all the embodiments. Based on one or more embodiments of this specification, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of this document.

[0030] Method Embodiment

[0031] According to an embodiment of the present invention, a method for establishing a graphic and text retrieval model is provided. Figure 1 It is a flowchart of the method for establishing a graphic and text retrieval model according to an embodiment of the present invention, as Figure 1 shown. The method for establishing a graphic and text retrieval model according to an embodiment of the present invention specifically includes:

[0032] In step S110, a teacher image feature set and a teacher text feature set corresponding to the training set are generated using a teacher model, which specifically includes:

[0033] By obtaining the public data sets MSCOCO and Flickr, processing and packing the data respectively and dividing them into three groups: a training set, a validation set, and a test set. The sources of the public data sets are https: / / cocodataset.org / #download and https: / / www.kaggle.com / datasets / hsankesara / flickr-image-dataset; in the embodiment of the present invention, the training process of the model is mainly described, and the training set includes positive and negative samples.

[0034] Preferably, after the public data set is obtained, it is cleaned, and the data format is sorted out as:

[0035]

[0036] Build a deep learning environment. Install anaconda on the server and create a virtual environment. Build Pytorch-GPU in the virtual environment, and then install the required library packages such as the pytotch library, scipy, numpy, sentence-transformers, tqdm, scikit-learn, and ftfy; obtain the required model and the soft label distillation ICSD method.

[0037] Use the forward_features function in https: / / github.com / deepglint / unicom to load CLIP ViT-B / 32 as the teacher image model, and use the teacher image model to offline generate the teacher image feature set and save it under train_unicom.npy. Among them, the internal format of the train_unicom.npy file is: {"<image_id>1":" <feature>"};

[0038] Load the text teacher model configuration file locally using https: / / huggingface.co / sentence-transformers / all-mpnet-basev2, and use the teacher text model to generate the teacher text feature set

[0039] In step S120, the training set is encoded to obtain an image encoding set and a text encoding set respectively. The image alignment feature set corresponding to the image encoding set and the text alignment feature set corresponding to the text encoding set are matched using probability distribution, specifically including:

[0040] Load the image encoder and the text encoder, encode the images in the training set through the image encoder to obtain the image encoding set, and encode the text in the training set through the text encoder to obtain the text encoding set;

[0041] Obtain the image encoding set and the text encoding set N represents the batch size, d represents the dimensional feature. Normalize the image encoding set into the set Normalize the text encoding set into the set

[0042] Based on the set Construct the first similarity matrix S1, and based on the set Construct the second similarity matrix S2. In the first similarity matrix, the element Element S1 i,j represents the similarity between images i and j, γ represents the temperature parameter. In the second similarity matrix, the element Element S2 i,j represents the similarity between texts i and j;

[0043] To enhance the bidirectional dependence relationship between similar samples, use formula 1 to construct the first symmetric matrix S1′ corresponding to the first similarity matrix, and use formula 2 to construct the second symmetric matrix S2′ corresponding to the second similarity matrix:

[0044] S1′ = softmax(S1, dim = 1) + softmax(S1, dim = 0) Formula 1;

[0045] S2′ = softmax(S2, dim = 1) + softmax(S2, dim = 0) Formula 2;

[0046] Among them, softmax(S, dim = 1) means that the sum of the activated elements in each row of matrix S is 1, and softmax(S, dim = 0) means that the sum of the activated elements in each column of matrix S is 1;

[0047] The weights in the symmetric matrix can enhance the negative impact of different samples on the target sample. Therefore, the impact of image negative samples can be calculated using Equation 3 Calculate the impact of the negative samples in this article using Equation 4

[0048] On this basis, a learnable scaling factor is used for dynamic adjustment to comprehensively consider the impacts of the original encoded features and the negative sample features. While retaining the original encoded information, the distance between positive and negative samples is increased at the same time. The set is adjusted by the scaling factor β using Equation 5 The adjusted image features are obtained, and the set is adjusted by the scaling factor β using Equation 6 The adjusted text features are obtained:

[0049]

[0050] The adjusted image features are smoothed and then input into the fully connected layer to output the image alignment feature set I′. The adjusted text features are smoothed and then input into the fully connected layer to output the text alignment feature set T′. In the embodiments of the present invention, feature smoothing is implemented using LayerNorm to ensure that the feature values between samples fall within a similar range. The fully connected layer is set as the FC layer to refine the feature representation. The image alignment feature set I′ is obtained using Equation 16, and the text alignment feature set T′ is obtained using Equation 17:

[0051]

[0052] The image alignment feature set and the text alignment feature set are the results obtained through probability distribution matching PDM.

[0053] In step S130, based on the calculation between the image alignment feature set and the text alignment feature set, an image-to-text distribution and a text-to-image distribution are generated, specifically including:

[0054] Calculate the similarity probability distribution between the i-th feature in the image alignment feature set and the j-th feature in the text alignment feature set using Equation 7 Calculate the similarity probability distribution between the i-th feature in the text alignment feature set and the j-th feature in the image alignment feature set using Equation 8

[0055]

[0056]

[0057] Among them, represents the cosine similarity between the i-th feature I of the image alignment feature set i ′ and the j-th feature T of the text alignment feature set j ′, represents the cosine similarity between the i-th feature T of the text alignment feature set i ′ and the j-th feature I of the image alignment feature set j ′, τ represents the learnable temperature parameter, N represents the number of features, and k represents the increment;

[0058] Generate the image-to-text distribution

[0059] Generate the text-to-image distribution

[0060] Both are probability distributions of cross-modal similarity.

[0061] In step S140, based on the teacher image feature set, calculate and generate the external image-to-image distribution, and based on the teacher text feature set, calculate and generate the external text-to-text distribution; based on the image encoding set, calculate and generate the internal image-to-image distribution, and based on the text encoding set, calculate and generate the internal text-to-text distribution, specifically including:

[0062] Use formula 9 to calculate the similarity probability distribution between the i-th feature and the j-th feature in the teacher image feature set Use formula 10 to calculate the similarity probability distribution between the i-th feature and the j-th feature in the teacher text feature set

[0063]

[0064] Among them, represents the i-th feature in the teacher image feature set and the j-th feature 's cosine similarity, represents the i-th feature in the teacher text feature set and the j-th feature 's cosine similarity, N represents the number of features, and k represents the increment;

[0065] Generate the external image-to-image distribution

[0066] Generate the external text-to-text distribution

[0067] are used for external intra-modal similarity to guide cross-modal similarity;

[0068] Calculate the similarity probability distribution between the $i$-th feature and the $j$-th feature in the image encoding set using Formula 11 Calculate the similarity probability distribution between the $i$-th feature and the $j$-th feature in the text encoding set using Formula 12

[0069]

[0070] where represents the $i$-th feature in the image encoding set and the $j$-th feature of the cosine similarity, represents the $i$-th feature in the text encoding set and the $j$-th feature of the cosine similarity, $\tau$ represents the learnable temperature parameter, $N$ represents the number of features, and $k$ represents the increment;

[0071] Generate the internal image - to - image distribution

[0072] Generate the internal text - to - text distribution

[0073] Both are internal intra - modality similarities used to guide cross - modality similarities.

[0074] In step S150, fuse the external image - to - image distribution and the internal image - to - image distribution to generate the image probability distribution, and fuse the external text - to - text distribution and the internal text - to - text distribution to generate the text probability distribution, specifically including:

[0075] Fuse the internal knowledge and the external knowledge, and use Formula 13 to fuse the external image - to - image distribution and the internal image - to - image distribution to generate the image probability distribution Use Formula 14 to fuse the external text - to - text distribution and the internal text - to - text distribution to generate the text probability distribution

[0076]

[0077]

[0078] where $\lambda$ represents the balance factor, is the final probability distribution of image - to - image, is the final probability distribution of text - to - text.

[0079] In step S160, use the image probability distribution to guide the image - to - text distribution, use the text probability distribution to guide the text - to - image distribution, construct the loss function to adjust the deep model parameters, and obtain the image - text retrieval model after the iteration ends, specifically including:

[0080] After integrating knowledge in the fusion modality, transfer the similarity in the modality to cross-modal similarity through knowledge distillation. Use the image probability distribution as the first target distribution and the KL divergence to guide the learnable image-to-text distribution. Use the text probability distribution as the second target distribution and the KL divergence to guide the learnable text-to-image distribution;

[0081] Use Equation 15 to construct the guiding loss function

[0082]

[0083] where KL represents the KL divergence, represents the guiding loss of the first target distribution, represents the guiding loss of the second target distribution. In this way, maintain the semantic consistency between the data in the modality and the cross-modal data, thereby enhancing the correct matching of cross-modal image-text or text-image pairs;

[0084] Combine the guiding loss function and the similarity loss function as the loss function. The similarity loss function uses a traditional loss to calculate the similarity between image-text and text-image. Use Equation 18 to calculate the similarity loss function, and use Equations 19 and 20 to represent the parameters in the similarity loss function:

[0085]

[0086] where, represents the similarity loss function, represents the image-text similarity, represents the text-image similarity, H(,) represents the cross-entropy operation, and y i =(y i1 ,...,y iN ) is a one-shot label, that is, for positive sample pairs, the value of y ii is 1, and in other cases, y ij is 0;

[0087] Use Equation 21 to combine the guiding loss function and the similarity loss function as the loss function

[0088]

[0089] where α represents a weight factor ranging from 0.1 to 1.0.

[0090] Introduce the in-modal consistency combined with soft-label distillation ICSD method in this embodiment of the present invention to train on the existing model, and adjust the parameters of the deep model through the loss function. The deep model is a model itself used to implement image-text retrieval;

[0091] Using the Recall@K metric, where K is typically set to 1, 5, 10, this metric reflects the proportion of correct matches retrieved among the top K results, providing an accurate measure of the model's retrieval accuracy. To comprehensively understand the model's performance, six recall values are further aggregated to calculate the RSUM score, providing a balanced perspective on the model's effectiveness and ensuring performance evaluation from multiple perspectives.

[0092] During the iteration process of the embodiment of the present invention, (1) Initialize the number of iterations. Initialize the number of iterations epoch = 0; the value of epoch ranges from 30 to 40 for different datasets; (2) The batch size batchsize ranges from 50 to 300 for different datasets; (3) Use the AdamW optimizer with a learning rate of 0.00001 and a weight decay of 0.0001; (4) Regarding parameter settings, set γ = 1 to adjust the sensitivity of the similarity score, and λ = 0.3 to balance the internal and external within-modal knowledge; (5) Evaluate the model through the built-in functions of pytorch. After each epoch of training, use the validation set to validate the currently trained model. If the RSUM is higher than the result of the previous epoch, save the parameter information of the current model, and perform cyclic validation until the optimal result is saved.

[0093] Save the optimal parameter information of the model or the parameter information after the iteration ends, and load the test set to calculate the accuracy rate and learnable parameters.

[0094] The method further includes:

[0095] In step S170, use the trained image-text retrieval model to perform image-text retrieval prediction.

[0096] Combined with the following drawings, the above technical solutions of the embodiments of the present invention are illustrated by examples.

[0097] Figure 2 is a schematic diagram of the model architecture of the embodiment of the present invention, as Figure 2 shown, showing the complete implementation framework of the present invention, including the Probability Matching Distribution (PDM) and the process of using within-modal probability distribution knowledge distillation to guide cross-modal probability distribution.

[0098] To sum up, in response to the existing problems, the method for establishing a graphic and text retrieval model of this invention generates independent and non-cross-modal probability distributions for two types of features, namely, image probability distribution and text probability distribution, which fully describe the semantic relationship within the image or text modality itself, and increase the consistency of the modality itself to improve the matching accuracy; the image probability distribution is formed by the fusion of the external image distribution and the internal image distribution, and the text probability distribution is formed by the fusion of the external text distribution and the internal text distribution, which fully considers both internal and external knowledge; in the process of generating the aligned feature set, the negative impact of different samples on the target sample is enhanced to assist in distinguishing similar and dissimilar samples and reduce errors in the matching process; probability distribution is used to guide cross-modal similarity learning, and soft label guidance is achieved through a lightweight solution to reduce the computational burden and improve the efficiency of graphic and text matching.

[0099] Device Embodiment

[0100] According to an embodiment of the present invention, a device for establishing a graph-text retrieval model is provided. Figure 3 Schematic diagram of a device for establishing a graphic and text retrieval model according to an embodiment of the present invention. Figure 3 As shown, the device for establishing a graphic-text retrieval model according to an embodiment of the present invention specifically includes:

[0101] The teacher module 30 is used to generate a teacher image feature set and a teacher text feature set corresponding to the training set using the teacher model, and is specifically used to:

[0102] Generate teacher image feature set using teacher image model

[0103] Generate teacher text feature set using teacher text model

[0104] The probability distribution matching module 32 is used to encode the training set to obtain an image code set and a text code set, and use the probability distribution to match the image alignment feature set corresponding to the image code set and the text alignment feature set corresponding to the text code set, specifically for:

[0105] Get image encoding set Text encoding set Encode the image set Normalize to a set Text encoding set Normalize to a set Collection-based Construct the first similarity matrix S1 based on the set Constructing a second similarity matrix S2;

[0106] Construct the first symmetric matrix S1′ corresponding to the first similarity matrix using Formula 1, and construct the second symmetric matrix S2′ corresponding to the second similarity matrix using Formula 2:

[0107] S1′ = softmax(S1, dim = 1) + softmax(S1, dim = 0) Formula 1;

[0108] S2′ = softmax(S2, dim = 1) + softmax(S2, dim = 0) Formula 2;

[0109] Among them, softmax(S, dim = 1) means making the sum of the activated elements in each row of matrix S equal to 1, and softmax(S, dim = 0) means making the sum of the activated elements in each column of matrix S equal to 1;

[0110] Calculate the influence of the image negative sample using Formula 3 Calculate the influence of the text negative sample in this article using Formula 4

[0111] Adjust the set using Formula 5 with the scaling factor β Obtain the adjusted image features, and adjust the set using Formula 6 with the scaling factor β Obtain the adjusted text features:

[0112]

[0113] After smoothing the adjusted image features, input them into the fully connected layer to output the image alignment feature set I′, and after smoothing the adjusted text features, input them into the fully connected layer to output the text alignment feature set T′.

[0114] The cross-modal distribution module 34 is used to calculate based on both the image alignment feature set and the text alignment feature set to generate the image-to-text distribution and the text-to-image distribution, specifically for:

[0115] Calculate the similarity probability distribution of the i-th feature in the image alignment feature set and the j-th feature in the text alignment feature set using Formula 7 Calculate the similarity probability distribution of the i-th feature in the text alignment feature set and the j-th feature in the image alignment feature set using Formula 8

[0116]

[0117] Among them, represents the cosine similarity between the i-th feature I′ of the image alignment feature set i and the j-th feature T′ of the text alignment feature set j and represents the i-th feature T′ of the text alignment feature set i and the cosine similarity of the j-th feature I' of the image alignment feature set j where τ represents the learnable temperature parameter, N represents the number of features, and k represents the increment;

[0118] Generate image-to-text distribution

[0119] Generate text-to-image distribution

[0120] The same-domain distribution module 36 is used to calculate and generate the external image-to-image distribution based on the teacher image feature set, and calculate and generate the external text-to-text distribution based on the teacher text feature set; calculate and generate the internal image-to-image distribution based on the image encoding set, and calculate and generate the internal text-to-text distribution based on the text encoding set. Specifically, it is used for:

[0121] Use Equation 9 to calculate the similarity probability distribution of the i-th feature and the j-th feature in the teacher image feature set Use Equation 10 to calculate the similarity probability distribution of the i-th feature and the j-th feature in the teacher text feature set

[0122]

[0123] where represents the i-th feature in the teacher image feature set and the j-th feature of the cosine similarity, represents the i-th feature in the teacher text feature set and the j-th feature of the cosine similarity, N represents the number of features, and k represents the increment;

[0124] Generate the external image-to-image distribution

[0125] Generate the external text-to-text distribution

[0126] Use Equation 11 to calculate the similarity probability distribution of the i-th feature and the j-th feature in the image encoding set Use Equation 12 to calculate the similarity probability distribution of the i-th feature and the j-th feature in the text encoding set

[0127]

[0128] where represents the i-th feature in the image encoding set and the j-th feature of the cosine similarity, Represents the i-th feature in the text encoding set and the j-th feature The cosine similarity, τ represents the learnable temperature parameter, N represents the number of features, and k represents the increment;

[0129] Generate the distribution among internal images

[0130] Generate the distribution among internal texts

[0131] The fusion module 38 is used to fuse the external image distribution and the internal image distribution to generate an image probability distribution, and fuse the external text distribution and the internal text distribution to generate a text probability distribution. Specifically, it is used for:

[0132] Use formula 13 to fuse the external image distribution and the internal image distribution to generate an image probability distribution Use formula 14 to fuse the external text distribution and the internal text distribution to generate a text probability distribution

[0133]

[0134] where λ represents the balance factor.

[0135] The guidance training module 310 is used to use the image probability distribution to guide the image-to-text distribution, use the text probability distribution to guide the text-to-image distribution, construct a loss function to adjust the parameters of the deep model, and obtain an image-text retrieval model after the iteration. Specifically, it is used for:

[0136] Perform knowledge distillation, use the image probability distribution as the first target distribution, use KL divergence to guide the image-to-text distribution, use the text probability distribution as the second target distribution, and use KL divergence to guide the text-to-image distribution;

[0137] Use formula 15 to construct a guidance loss function

[0138]

[0139] where KL represents KL divergence, represents the guidance loss of the first target distribution, represents the guidance loss of the second target distribution;

[0140] Comprehensively calculate the guidance loss function and the similarity loss function as the loss function, and adjust the parameters of the deep model through the loss function.

[0141] In summary, in view of the problems existing in the current situation, the apparatus for establishing the invention's graphic and text retrieval model generates independent and non-cross-modal probability distributions for the features of two types, namely images and texts, that is, the image probability distribution and the text probability distribution, which fully describe the semantic relationships within the modalities of the images or texts themselves, adding the consistency of the modalities themselves to improve the accuracy of matching; the image probability distribution is formed by fusing the external image distribution and the internal image distribution, and the text probability distribution is formed by fusing the external text distribution and the internal text distribution, fully considering both internal and external knowledge; during the process of generating the alignment feature set, the negative impacts of different samples on the target sample are enhanced to assist in distinguishing similar and dissimilar samples, reducing errors during the matching process; the probability distribution is used to guide cross-modal similarity learning, and soft label guidance is achieved through a lightweight scheme, reducing the computational burden to improve the efficiency of graphic and text matching.

[0142] Embodiment of the electronic device

[0143] Figure 4 It is a schematic diagram of the electronic device according to an embodiment of the present invention. The electronic device 400 may include at least one processor 410 and a memory 420. The processor 410 may execute instructions stored in the memory 420. The processor 410 is communicatively connected to the memory 420 via a data bus. In addition to the memory 420, the processor 410 may also be communicatively connected to an input device 430, an output device 440, and a communication device 450 via the data bus.

[0144] The processor 410 may be any conventional processor, such as a commercially available CPU. The processor may also include, for example, a Graphic Process Unit (GPU), a Field Programmable Gate Array (FPGA), a System on Chip (SOC), an Application Specific Integrated Circuit (ASIC), or a combination thereof.

[0145] The memory 420 may be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as Static Random Access Memory (SRAM), Electrically Erasable Programmable Read-Only Memory (EEPROM), Erasable Programmable Read-Only Memory (EPROM), Programmable Read-Only Memory (PROM), Read-Only Memory (ROM), magnetic memory, flash memory, a magnetic disk, or an optical disk.

[0146] In an embodiment of the present disclosure, executable instructions are stored in the memory 420. The processor 410 may read the executable instructions from the memory 420 and execute the instructions to implement all or part of the steps of the method for establishing a graphic and text retrieval model in any of the above exemplary embodiments.

[0147] Embodiment of computer-readable storage medium

[0148] In addition to the above methods and apparatuses, an exemplary embodiment of the present disclosure may also be a computer program product or a computer-readable storage medium storing the computer program product. The computer program product includes computer program instructions that can be executed by a processor to implement all or part of the steps described in the method for establishing a graphic and text retrieval model in any of the above exemplary embodiments.

[0149] The computer program product may be written in any combination of one or more programming languages for programming code to perform the operations of the embodiments of the present application. The programming languages include object-oriented programming languages such as Java, C++, etc., and also include conventional procedural programming languages such as the "C" language or similar programming languages and scripting languages (such as Python). The programming code may be executed entirely on the user computing device, partially on the user device, executed as an independent software package, partially on the user computing device and partially on a remote computing device, or entirely on a remote computing device or server.

[0150] The computer-readable storage medium may adopt any combination of one or more readable media. The readable media may be a readable signal medium or a readable storage medium. The readable storage medium may include, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples of the readable storage medium include: a static random access memory (SRAM) with one or more wire electrical connections, an electrically erasable programmable read-only memory (EEPROM), an erasable programmable read-only memory (EPROM), a programmable read-only memory (PROM), a read-only memory (ROM), a magnetic memory, a flash memory, a magnetic disk or an optical disk, or any suitable combination of the above.

[0151] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit them. Although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements on some or all of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention.< / feature>

Claims

1. A method for establishing a picture and text retrieval model, characterized in that: include: Use the teacher model to generate the teacher image feature set and teacher text feature set corresponding to the training set; Encode the training set to obtain an image encoding set and a text encoding set respectively, and use probability distribution to match an image alignment feature set corresponding to the image encoding set and a text alignment feature set corresponding to the text encoding set; Generate image-to-text distribution and text-to-image distribution based on calculation between the image alignment feature set and the text alignment feature set; Generate external image distribution based on the teacher image feature set, and generate external text distribution based on the teacher text feature set; generate internal image distribution based on the image code set, and generate internal text distribution based on the text code set; The external image distribution and the internal image distribution are integrated to generate an image probability distribution, and the external text distribution and the internal text distribution are integrated to generate a text probability distribution; The image probability distribution is used to guide the image-to-text distribution, and the text probability distribution is used to guide the text-to-image distribution. A loss function is constructed to adjust the deep model parameters, and an image-text retrieval model is obtained after the iteration.

2. The method according to claim 1, characterized in that The method of using the teacher model to generate the teacher image feature set and the teacher text feature set corresponding to the training set specifically includes: Generate teacher image feature set using teacher image model Generate teacher text feature set using teacher text model 3. The method according to claim 1, characterized in that The image alignment feature set corresponding to the image encoding set and the text alignment feature set corresponding to the text encoding set using probability distribution matching specifically include: Get image encoding set and text encoding set The image encoding set Normalize to a set The text encoding set Normalize to a set Based on the collection Construct a first similarity matrix S1 based on the set Constructing a second similarity matrix S2; Use formula 1 to construct the first symmetric matrix S1 corresponding to the first similarity matrix ′ , use formula 2 to construct the second symmetric matrix S2 corresponding to the second similarity matrix ′ : S1′=softmax(S1,dim=1)+softmax(S1,dim=0) Formula 1; S2′=softmax(S2, dim=1)+softmax(S2, dim=0) Formula 2; Among them, softmax(S, dim = 1) means that the sum of the activated elements in each row of the matrix S is 1, and softmax(S, dim = 0) means that the sum of the activated elements in each column of the matrix S is 1; Use Formula 3 to calculate the impact of negative samples in the image Use formula 4 to calculate the negative sample impact of this paper Use Equation 5 to adjust the set by the scaling factor β Get the image adjustment features and adjust the set by the scaling factor β using Formula 6 Get the text justification features: The image adjustment features are smoothed and input into the fully connected layer to output the image alignment feature set I ′ , the text adjustment feature is smoothed and input into the fully connected layer to output the text alignment feature set T ′ .

4. The method according to claim 1, characterized in that: The generating of image-to-text distribution and text-to-image distribution based on the calculation between the image alignment feature set and the text alignment feature set specifically includes: Use Formula 7 to calculate the similarity probability distribution between the i-th feature in the image alignment feature set and the j-th feature in the text alignment feature set Use Formula 8 to calculate the similarity probability distribution between the i-th feature in the text alignment feature set and the j-th feature in the image alignment feature set in, Represents the i-th feature I of the image alignment feature set i ′ And the jth feature T of the text alignment feature set j ′ The cosine similarity of Represents the i-th feature T of the text alignment feature set i ′ And the jth feature I of the image alignment feature set j ′ The cosine similarity of , τ represents the learnable temperature parameter, N represents the number of features, and k represents the increment; Generating image-to-text distribution Generating Text-to-Image Distribution 5. The method according to claim 4, characterized in that The method of generating the external image distribution based on the teacher image feature set and the external text distribution based on the teacher text feature set; generating the internal image distribution based on the image code set and the internal text distribution based on the text code set specifically includes: Use Formula 9 to calculate the similarity probability distribution between the i-th feature and the j-th feature in the teacher image feature set Formula 10 is used to calculate the similarity probability distribution between the i-th feature and the j-th feature in the teacher text feature set. in, Represents the i-th feature in the teacher image feature set With the jth feature The cosine similarity of Represents the i-th feature in the teacher text feature set With the jth feature The cosine similarity of , N represents the number of features, and k represents the increment; Generate external image distribution Generate external text distribution Use formula 11 to calculate the similarity probability distribution between the i-th feature and the j-th feature in the image encoding set Use formula 12 to calculate the similarity probability distribution between the i-th feature and the j-th feature in the text encoding set in, Represents the i-th feature in the image encoding set With the jth feature The cosine similarity of Represents the i-th feature in the text encoding set With the jth feature The cosine similarity of , τ represents the learnable temperature parameter, N represents the number of features, and k represents the increment; Generate internal inter-image distribution Generate internal text distribution 6. The method according to claim 5, characterized in that The fusing of the external image distribution and the internal image distribution to generate an image probability distribution, and the fusing of the external text distribution and the internal text distribution to generate a text probability distribution specifically include: The image probability distribution is generated by fusing the external image distribution and the internal image distribution using formula 13. Formula 14 is used to fuse the external text distribution and the internal text distribution to generate a text probability distribution. Here, λ represents the balance factor.

7. The method according to claim 6, characterized in that The using the image probability distribution to guide the image-to-text distribution, the using the text probability distribution to guide the text-to-image distribution, and constructing a loss function to adjust the deep model parameters specifically include: Perform knowledge distillation, take the image probability distribution as target distribution one, use KL divergence to guide the image-to-text distribution, take the text probability distribution as target distribution two, and use KL divergence to guide the text-to-image distribution; Use Formula 15 to construct the guided loss function Among them, KL represents KL divergence, represents the guidance loss of target distribution one, represents the guidance loss of target distribution 2; The guided loss function and the similarity loss function are combined as the loss function, and the deep model parameters are adjusted through the loss function.

8. A device for establishing a picture and text retrieval model, characterized in that: include: A teacher module, used to generate a teacher image feature set and a teacher text feature set corresponding to a training set using a teacher model; A probability distribution matching module is used to encode the training set to obtain an image encoding set and a text encoding set respectively, and use the probability distribution to match the image alignment feature set corresponding to the image encoding set and the text alignment feature set corresponding to the text encoding set; A cross-modal distribution module, configured to generate image-to-text distribution and text-to-image distribution based on calculations between the image alignment feature set and the text alignment feature set; A homomodal distribution module, used to generate external image distribution based on the teacher image feature set, and generate external text distribution based on the teacher text feature set; generate internal image distribution based on the image code set, and generate internal text distribution based on the text code set; A fusion module, used for fusing the external image distribution and the internal image distribution to generate an image probability distribution, and fusing the external text distribution and the internal text distribution to generate a text probability distribution; A guidance training module is used to use the image probability distribution to guide the image-to-text distribution, use the text probability distribution to guide the text-to-image distribution, construct a loss function to adjust the deep model parameters, and obtain a picture-text retrieval model after the iteration.

9. An electronic device, characterized in that: include: A memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the computer program, when executed by the processor, implements the steps of the method for establishing a graphic and text retrieval model as described in any one of claims 1 to 7.

10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores an implementation program for information transmission, and when the program is executed by the processor, the steps of the method for establishing a graphic and text retrieval model as described in any one of claims 1 to 7 are implemented.