Cross-language and cross-modal retrieval method based on semantic decoupling and dynamic parameter generation
By adopting semantic decoupling and dynamic parameter generation methods in cross-language cross-modal retrieval, the problem of knowledge forgetting and adapting to different language expression methods in cross-language migration is solved, and a more efficient and accurate cross-language cross-modal retrieval effect is achieved.
Patent Information
- Application Number
- CN202411874910.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-19
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2044-12-19
AI Technical Summary
Existing cross-language and cross-modal retrieval technologies are prone to knowledge forgetting when cross-language migration, and it is difficult to adapt to the differences in expression methods of different languages. Especially for low-resource languages, it is difficult to effectively capture their unique expressions and semantic information.
A cross-language cross-modal retrieval method based on semantic decoupling and dynamic parameter generation is adopted to decouple sentences into semantic related and semantic independent features through semantic decoupling branches, and these features are fused through a dynamic adapter module to generate dynamic parameters suitable for the target language.
It improves the accuracy and efficiency of cross-language and cross-modal retrieval, can better adapt to the expression methods of different languages, improves support for low-resource languages, and reduces the phenomenon of knowledge forgetting.
Smart Images

Figure CN119322859B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of cross-language and cross-modal retrieval, and in particular to a cross-language and cross-modal retrieval method based on semantic decoupling and dynamic parameter generation. Background Art
[0002] With the rapid emergence of images and videos on the Internet, users around the world have a huge demand for visual content of interest through natural language retrieval, i.e., cross-modal retrieval. Cross-modal retrieval is the mutual retrieval of data from at least two modalities, usually using one modality as a query to retrieve relevant data from another modality. By finding the potential associations between data from different modalities, relatively accurate cross-matching can be achieved. Cross-modal retrieval allows users to quickly find what they need in a vast ocean of information. The rapid development of deep learning technology, especially in recent years, has greatly promoted the research and application of cross-modal retrieval technology. Recent neural network-based cross-modal retrieval models tend to require a large amount of human-labeled text-image pair data for training, which is only available for a few languages in the world. Therefore, it is extremely challenging to build a cross-modal retrieval system for users with different language backgrounds, especially for low-resource languages (such as Czech). Therefore, cross-language cross-modal retrieval is very important. It uses visual-text pair data in high-resource languages to build retrieval models for new target languages to cope with the lack of a large amount of manually labeled data in the target language.
[0003] A simple and low-cost solution is to use machine translation tools to convert the source language labeled data into the target language. With these resources generated by machine translation, existing work tends to transfer the cross-modal alignment ability of the visual language model to the target language through cross-language alignment so as to train multimodal and multilingual models. However, when transferring across languages, the model often forgets knowledge and the performance on high-resource languages decreases. To alleviate this problem, adapter-based methods have emerged. These methods freeze the parameters of the visual language model and use lightweight adapters for cross-language transfer. However, due to the large number of languages and different expressions in different languages, even in the same language, the description of the same image is likely to be very different, making it impossible for the model to be efficiently and accurately extended to new languages, which poses a major challenge to training and improving cross-modal retrieval systems. Existing adapter methods are difficult to adapt to target language sentences with different expressions. Considering the large number of low-resource languages, how to capture the unique expressions and corresponding semantic information of these low-resource languages is a very challenging task. Summary of the invention
[0004] The present invention aims to address the deficiencies of the prior art and to propose a cross-language and cross-modal retrieval method based on semantic decoupling and dynamic parameter generation.
[0005] The object of the present invention is achieved through the following technical solution: a cross-language and cross-modal retrieval method based on semantic decoupling and dynamic parameter generation, the method comprising the following steps:
[0006] S1. Segment and encode the source language and target language sentences to obtain the embedding vectors of the corresponding language texts.
[0007] S2, input the embedding vector into the text encoders of the source language branch, the target language branch and the semantic decoupling branch to obtain the corresponding feature vector, wherein the text encoders of the source language branch and the target language branch are 12 layers, and the text encoder of the semantic decoupling branch is 1 layer;
[0008] The semantic decoupling branch is as follows: using the pre-trained text encoding layer output to two multi-layer perceptrons to obtain semantically relevant and semantically irrelevant feature outputs respectively; concatenating the two obtained features and then passing them through a multi-layer perceptron to obtain the parameters of the dynamic adapter;
[0009] The target language branch is output from the pre-trained text encoder layer to the dynamic adapter to obtain the final text information as the target language text output or the target language input of the next layer, wherein the parameters of the dynamic adapter are obtained by semantic decoupling;
[0010] S3, conduct adversarial learning on the output of semantically irrelevant features and source language text information and construct constraints;
[0011] S4. Based on the target language text features and source language text features output by the text encoder, the MSE loss is used to train the parameters of the dynamic adapter of the target language branch to achieve cross-language migration.
[0012] S5, input the target language text into the text encoder trained in the cross-language alignment stage to obtain the target language branch text features of cross-modal alignment;
[0013] S6. Use the pre-trained image encoder to obtain image features, and use InfoNCE loss to calculate the retrieval model that realizes cross-modal alignment of text features and image features;
[0014] S7. Input the text into the trained retrieval model to achieve cross-language and cross-modal retrieval.
[0015] Furthermore, the encoding after word segmentation of the source language and the target language sentences is specifically as follows:
[0016] Load the pre-trained mBERT and clip tokenizers, perform token segmentation on the source language text and the target language text to obtain a fixed-length tokenID vector after token segmentation;
[0017] Use a multilingual word embedding matrix and the word embedding matrix of clip to perform word embedding operation on the tokenID vector obtained previously, obtain the word embedding encoding vectors of the source language text and the target language text, and add position encoding to the two text sub-embedding vectors.
[0018] Furthermore, the text encoder layer is specifically:
[0019] The obtained embedding vectors of the source language text and the target language text are used to obtain the corresponding feature vectors through the text encoder of each layer of the pre-trained model CLIP. , where i represents the i-th layer of the encoder;
[0020]
[0021]
[0022] here It is at the token level and can be expressed as .
[0023] Furthermore, the semantic decoupling branch is specifically:
[0024] Just need to pass the features of a layer of text encoder After the semantically relevant and semantically irrelevant features are obtained through the semantically relevant adapter and the semantically irrelevant adapter respectively, they are fused through a feature fusion layer to obtain the parameters of the dynamic adapter. : ;
[0025] in , , where H represents the size of the hidden layer, I represents the size of the intermediate layer of semantically relevant adapters and semantically irrelevant adapters, Represents a direct splicing operation, and Represent the upper projection layer of semantically relevant adapter and semantically irrelevant adapter respectively, and Represent the lower projection layers of semantically relevant adapters and semantically irrelevant adapters, respectively. and Represents the upper and lower projection layers of the feature fusion layer, and By selecting according to the value of tokenID Vector Get and , represents semantically related features, Represents semantically irrelevant features. After concatenating the two, a feature fusion layer is used to obtain a sentence-level feature vector Z with a dimension of (B, I). Then, a Reshape operation is performed to change the dimension of the feature vector Z from I to ,and , thereby obtaining the parameter matrix of the dynamic adapter .
[0026] Furthermore, the target language branch dynamic adapter module is specifically: inputting the text feature vector into the dynamic adapter module Get the target language text features :
[0027]
[0028] =
[0029] Where i represents the number of encoder layers. , , where H represents the hidden layer size, represents the size of the intermediate layer of the feature fusion layer and , so the number of parameters has been reduced to a certain extent. Represents the matrix after the adapter; after 12 layers of encoder, the feature vector obtained is based on The value of [EOS] is used to select the [EOS] vector and obtain the feature vector at the target language sentence level .
[0030] Furthermore, the specific steps in S3 are: the semantically relevant and semantically irrelevant features of the semantic decoupling module need to be constrained; specifically, the features of the semantically irrelevant branches in the semantic decoupling module are kept away from the semantic features of the source language, and the semantically irrelevant feature constraints are:
[0031] )
[0032] Where F is the discriminator, which consists of a simple linear layer. When it becomes smaller, the model will not be able to distinguish The source language features corresponding to the positive sample pairs;
[0033] The semantically relevant feature constraints are:
[0034]
[0035] Where B is the batch size, is the source language feature, are semantically related features; A feature vector representing the sentence level of the source language.
[0036] Furthermore, the use of MSE loss to train the parameters of the dynamic adapter of the target language branch and the parameters of the semantic decoupling branch to achieve cross-language migration is specifically as follows:
[0037] +
[0038] Where B is the batch size, is the source language feature, is the target language feature. Cross-language alignment can be achieved by adding constraints.
[0039] Furthermore, the retrieval model that uses InfoNCE loss calculation to achieve cross-modal alignment of text features and image features is specifically:
[0040] NCE loss is calculated for image features and text features to train the semantic disentanglement branch and the dynamic adapter part:
[0041] ,
[0042] ,
[0043] ;
[0044] in is the temperature coefficient, B is the batch size, Representative image features To text features The similarity loss is Representative text features To image features The similarity loss is represents the cosine similarity; Under the constraint of this loss, the model will make relevant text features and image features close to each other, and irrelevant ones far away from each other.
[0045] According to another aspect of the specification, the present invention specification also provides a cross-language and cross-modal retrieval device based on semantic decoupling and dynamic parameter generation, including a memory and one or more processors, wherein the memory stores executable code, and when the processor executes the executable code, it implements the cross-language and cross-modal retrieval method based on semantic decoupling and dynamic parameter generation.
[0046] According to another aspect of the specification, the present invention specification also provides a computer-readable storage medium having a program stored thereon, and when the program is executed by a processor, the cross-language and cross-modal retrieval method based on semantic decoupling and dynamic parameter generation is implemented.
[0047] Beneficial effects of the present invention:
[0048] 1. The present invention freezes all parameters of the original text encoder and introduces a semantic decoupling module and a dynamic adapter module. The semantic decoupling module decouples a sentence into semantically relevant features and semantically irrelevant features, which are simply the specific meaning and expression form of a sentence. Based on these two pieces of information, the model can perform better in processing sentences with the same semantics but different expressions.
[0049] 2. The dynamic adapter module of the present invention fuses the output of the semantic decoupling module into the adapter by means of low-rank decomposition, and fuses semantically relevant information with semantically irrelevant information to improve cross-language and cross-modal retrieval capabilities. BRIEF DESCRIPTION OF THE DRAWINGS
[0050] Figure 1 A schematic flow chart of a cross-language and cross-modal retrieval method based on semantic decoupling and dynamic parameter generation provided by an embodiment of the present invention;
[0051] Figure 2 A schematic diagram of a specific process of word embedding provided in an embodiment of the present invention;
[0052] Figure 3 A schematic diagram of the structures of a semantically relevant adapter and a semantically irrelevant adapter provided in an embodiment of the present invention;
[0053] Figure 4 A schematic diagram of adversarial learning in the semantic decoupling process provided by an embodiment of the present invention;
[0054] Figure 5 An example diagram of qualitative analysis of semantically irrelevant features provided by an embodiment of the present invention;
[0055] Figure 6 A schematic diagram of a cross-language and cross-modal retrieval device based on semantic decoupling and dynamic parameter generation provided by an embodiment of the present invention. DETAILED DESCRIPTION
[0056] The specific implementation modes of the present invention are further described in detail below with reference to the accompanying drawings.
[0057] like Figure 1 As shown, the present invention provides a cross-language and cross-modal retrieval method based on semantic decoupling and dynamic parameter generation, and the specific steps include:
[0058] S1. In the cross-language transfer stage, feature encoding is performed on the source language text and the target language text to obtain the embedding vector of the corresponding language text.
[0059] Load the pre-trained mBERT and clip tokenizers, perform tokenization on the source language text and the target language text to obtain a fixed-length tokenID vector after tokenization. , , where B represents the batch size and L represents the length of the word segmentation. Figure 1 and Figure 2 The word segmentation operation in the target language sentence [A cat sitting on the cobblestone ground.] is segmented into [“one”, “only”, “cat”, …]; the source language sentence [A cat sitting on the cobblestone ground.] is segmented into [“A”, “cat”, “sitting”, …].
[0060] Because the large-scale pre-trained model used is the English visual model clip, we keep the original processing method of clip unchanged when processing the English and visual ends.
[0061] like Figure 2 As shown in the figure, a multilingual word embedding matrix and the word embedding matrix of clip are used to perform word embedding operations on the previously obtained tokenID matrix to obtain the word embedding encoding vectors of the source language text and the target language text. , , where H represents the hidden layer size, Represents the source language, Represents the target language.
[0062] Add position encoding to both text word embedding vectors .
[0063]
[0064]
[0065] S2. Input the embedding vector into the text encoder to obtain its corresponding feature vector. The text part has three branches: source language branch, target language branch and semantic decoupling branch. All three branches contain text encoders. The source language branch and target language branch text encoders have 12 layers, and the semantic decoupling branch text encoder has 1 layer.
[0066] The source language branch is a pre-trained text encoder layer, and the obtained text information is used as the source language text output or the next layer of source language text information input;
[0067] The semantic decoupling branch uses only one pre-trained text encoding layer and outputs it to two multi-layer perceptrons (MLPs) to obtain semantically relevant and semantically irrelevant feature outputs respectively. The two features are concatenated and then passed through an MLP to obtain a matrix as the parameters of the dynamic adapter.
[0068] The target language branch is output from the pre-trained text encoder layer to the dynamic adapter module to obtain the final text information as the target language text output or the target language input of the next layer, wherein the parameters of the dynamic adapter are obtained by semantic decoupling;
[0069] The text encoder layer is specifically:
[0070] The obtained embedding vectors of the source language text and the target language text are used to obtain the corresponding feature vectors through the text encoder of each layer of the pre-trained model CLIP. , where i represents the i-th layer of the encoder.
[0071]
[0072]
[0073] here It is at the token level and can be expressed as .
[0074] The semantic decoupling module is specifically:
[0075] like Figure 1 As shown, we only need to pass the features of a layer of text encoder After the semantically relevant and semantically irrelevant features are obtained through the semantically relevant adapter and the semantically irrelevant adapter respectively, they are fused through a feature fusion layer to obtain:
[0076] ;
[0077] in , , where H represents the size of the hidden layer, I represents the size of the intermediate layer of semantically relevant adapters and semantically irrelevant adapters, Represents a direct splicing operation, and Represent the upper projection layer of semantically relevant adapter and semantically irrelevant adapter respectively, and Represent the lower projection layers of semantically relevant adapters and semantically irrelevant adapters, respectively. and Represents the upper and lower projection layers of the feature fusion layer. and By selecting according to the value of tokenID Vector Get and , represents semantically related features, Represents semantically irrelevant features. After concatenating the two, a feature fusion layer is used to obtain a sentence-level feature vector Z with a dimension of (B, I). Then, a Reshape operation is performed to change the dimension of the feature vector Z from I to ,and , thereby obtaining the parameter matrix of the dynamic adapter .
[0078] Furthermore, the target language branch dynamic adapter module is specifically: inputting the text feature vector into the dynamic adapter module Get the target language text features :
[0079]
[0080] =
[0081] Where i represents the number of encoder layers. , , where H represents the hidden layer size, represents the size of the intermediate layer of the feature fusion layer and , so the number of parameters has been reduced to a certain extent. Represents the matrix after the adapter. After 12 layers of encoder, the feature vector is obtained according to The value of [EOS] is used to select the [EOS] vector and obtain the feature vector at the target language sentence level .
[0082] Furthermore, the source language branch text encoder is specifically:
[0083]
[0084] in The text encoder in the visual text encoder clip we use, because English is already trained so we keep the parameters frozen. Represents the source language sentence-level feature vector.
[0085] like Figure 3 As shown, the MLP used by the semantic decoupling module to calculate semantically relevant features and semantically irrelevant features is specifically:
[0086] Downward projection layer, nonlinear function layer and upward projection layer, first use the downward projection layer to reduce the dimension, then use the activation layer to make the model better adapt to nonlinear changes, and then use the upward projection layer to restore the dimension to the input dimension. Then the output is residually linked to the input. Here we simply use two adapters to calculate two features separately, and the number of parameters and calculations are not large.
[0087] The global semantic information corresponding to the feature vector is input into the selector to obtain the weight matrix and bias matrix related to the global semantics. At the same time, the global semantic information corresponding to the feature vector is projected downward, and after dimension reduction, the feature matrix related to the global semantics is calculated through the output of the selector and input into the nonlinear function layer.
[0088] S3, such as Figure 4 As shown in Figure 1, we need to constrain the semantically relevant and semantically irrelevant features of the semantic decoupling module to ensure that they are relevant or irrelevant to the semantics, specifically:
[0089] Since we input parallel sentence pairs, the semantics of the English and Chinese words in the corresponding positions of each batch are consistent, such as Figure 4 Shown and are parallel, that is, the semantics are consistent. In order to obtain semantically irrelevant features, we convert the output of the English text encoder and Conduct an adversarial training. Specifically, we first establish two feature sample pairs ( , )and( , ) where i represents the sequence number of the sentence in the batch, and n represents a random number other than i. So it is obvious that ( , ) is a semantically related positive sample pair, and ( , ) is a negative sample pair. Based on this, we use Constraining semantically irrelevant features:
[0090] )
[0091] Where F is the discriminator, which consists of a simple linear layer. When it becomes smaller, the model will not be able to distinguish The English features corresponding to the positive sample pairs, for example, there are three English features , , , a Chinese feature , after training, the discriminator cannot determine which of the three English features can constitute a positive sample pair. So far, the training of semantically irrelevant features has been completed. The main purpose is to make the features of the semantically irrelevant branch in the semantic decoupling module away from the English semantic features. As for how to judge whether it is really away from the English semantic features, it depends on the discriminator. If the discriminator cannot identify which English feature is the positive sample pair of the input Chinese feature, then it means that the Chinese feature is semantically irrelevant.
[0092] In terms of semantically related features, we use Semantically related features Make constraints
[0093] ;
[0094] Where B is the batch size, is the source language feature, are semantically related features.
[0095] S4. Based on the target language text features and source language text features output by the text encoder, the mean square error loss (MSE) is used to train the semantic decoupling module and dynamic adapter part of the model to achieve cross-language transfer.
[0096] ;
[0097] Where B is the batch size, is the source language feature, is the target language feature (e.g. Chinese, German).
[0098] S5. In the cross-modal alignment phase, load the model trained in the cross-language phase, input the target language text into the text encoder, and obtain the text features of the target language branch. ;
[0099] S6. Use the pre-trained image encoder to obtain image features. Our retrieval model is mainly composed of a text side and an image side. It mainly relies on the similarity between text and image to determine whether the text matches the image. For example, if a text is input and the five images with the highest similarity are found, then these five images are the retrieval answers given by the retrieval model in this article. Therefore, the NCE loss is used to calculate the distance between the corresponding text features and image features, thereby realizing a cross-modal aligned retrieval model.
[0100] Specifically, in the image branch, the image is first reduced and cropped from the original image size to 224×224, and then a 2D convolutional layer is used to convert the input image into a feature map, and then position encoding is added to obtain the initial features of the image. , B is the batch size, Represents the feature size. The initial features are input into the pre-trained CLIP image encoder to obtain the final image features. ;
[0101] NCE loss is calculated for image features and text features to train the semantic disentanglement module and dynamic adapter part:
[0102]
[0103]
[0104]
[0105] in is the temperature coefficient, which is set to 0.01 in the present invention, and B is the batch size. Representative image features To text features The similarity loss is Representative text features To image features The similarity loss is Represents cosine similarity. Under the constraint of this loss, the model will make relevant text features and image features close to each other, and irrelevant ones far away from each other.
[0106] S7. Input the target language text into the trained retrieval model, and first pass it through the text embedding layer to get the target language text embedding. Then input it into the semantic decoupling branch to get the dynamic parameter matrix, and insert the dynamic parameter matrix into the dynamic adapter of the target language branch to achieve dynamic changes in the adapter parameters with the input. Then input the text into the target language branch to get a target language feature vector, input the picture into the image encoder to get the image feature vector, compare the target language feature vector with the image feature vector for similarity, and select the pictures ranked TOP-1, TOP-5, and TOP-10 in similarity to achieve cross-language and cross-modal retrieval.
[0107] Referring to FIG5 , we have qualitatively analyzed the specific content of the semantically irrelevant features in the present invention. We randomly selected 200 sentences from the test set of the MSCOCO dataset and performed tsne clustering visualization on them. The results are as follows: Figure 5 As shown, sentences with different expressions can be grouped together well. Figure 5 The corresponding visualization sentence example is shown below:
[0108] Group 1:
[0109] NO.174: A group of giraffes stand near a big tree.
[0110] NO.33: A truck with graffiti all over it.
[0111] NO.1: A cat stands next to two stuffed animals.
[0112] Group 2:
[0113] NO.147: There is a piece of bread and salad on the plate.
[0114] NO.124: The yard has a stove and chairs.
[0115] NO.98: A metal rack holds several toothbrushes.
[0116] Group 3:
[0117] NO.193: A man in jeans lits his skateboard with his foot.
[0118] NO.24: A woman holding an umbrella walks on a path.
[0119] NO.170: A man wearing a tie is walking on the street.
[0120] Group 4:
[0121] NO.61: This is a painting, in which a boy and a girl are sitting on a chair on the beach.
[0122] NO.80: A player is playing a game on the tennis court, with referees and many spectators nearby.
[0123] NO.57: A blond little boy is sitting at the table, holding a grilled sausage and eating breakfast.
[0124] The sentences in the first group all start with a quantifier plus a noun, the sentences in the second group all start with a locative word followed by a noun, the sentences in the third group all start with a quantifier plus an adjective plus a noun, and compared to the first group, there is an additional adjective, and the fourth group is all longer sentences, which can be separated by commas. Based on this qualitative analysis, we can see that semantically irrelevant features can learn the expression form of sentences.
[0125] Corresponding to the aforementioned embodiment of a cross-language and cross-modal retrieval method based on semantic decoupling and dynamic parameter generation, the present invention also provides an embodiment of a cross-language and cross-modal retrieval device based on semantic decoupling and dynamic parameter generation.
[0126] See also Figure 6 A cross-language and cross-modal retrieval device based on semantic decoupling and dynamic parameter generation provided by an embodiment of the present invention includes a memory and one or more processors, wherein the memory stores executable code, and when the processor executes the executable code, it is used to implement a cross-language and cross-modal retrieval method based on semantic decoupling and dynamic parameter generation in the above embodiment.
[0127] An embodiment of a cross-language and cross-modal retrieval device based on semantic decoupling and dynamic parameter generation provided by the present invention can be applied to any device with data processing capabilities, and the device with data processing capabilities can be a device or apparatus such as a computer. The device embodiment can be implemented through software, or through hardware or a combination of software and hardware. Taking software implementation as an example, as a device in a logical sense, it is formed by the processor of any device with data processing capabilities in which it is located reading the corresponding computer program instructions in the non-volatile memory into the internal memory for execution. From a hardware perspective, if Figure 6As shown, it is a hardware structure diagram of any device with data processing capability where a cross-language and cross-modal retrieval device based on semantic decoupling and dynamic parameter generation provided by the present invention is located, except Figure 6 In addition to the processor, memory, network interface, and non-volatile memory shown, any device with data processing capabilities in which the apparatus in the embodiments is located may also include other hardware, generally based on the actual functions of the device with data processing capabilities, which will not be described in detail.
[0128] The implementation process of the functions and effects of each unit in the above-mentioned device is specifically described in the implementation process of the corresponding steps in the above-mentioned method, and will not be repeated here.
[0129] For the device embodiment, since it basically corresponds to the method embodiment, the relevant parts can refer to the partial description of the method embodiment. The device embodiment described above is only schematic, wherein the units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they may be located in one place, or they may be distributed on multiple network units. Some or all of the modules may be selected according to actual needs to achieve the purpose of the scheme of the present invention. Ordinary technicians in this field can understand and implement it without paying creative work.
[0130] An embodiment of the present invention further provides a computer-readable storage medium having a program stored thereon. When the program is executed by a processor, a cross-language and cross-modal retrieval method based on semantic decoupling and dynamic parameter generation in the above embodiment is implemented.
[0131] The computer-readable storage medium may be an internal storage unit of any device with data processing capability described in any of the aforementioned embodiments, such as a hard disk or a memory. The computer-readable storage medium may also be an external storage device of any device with data processing capability, such as a plug-in hard disk, a smart media card (SMC), an SD card, a flash card, etc. equipped on the device. Furthermore, the computer-readable storage medium may also include both an internal storage unit and an external storage device of any device with data processing capability. The computer-readable storage medium is used to store the computer program and other programs and data required by any device with data processing capability, and may also be used to temporarily store data that has been output or is to be output.
[0132] The present invention also provides a computer program product, including a computer program, which, when executed by a processor, implements the cross-language and cross-modal retrieval method based on semantic decoupling and dynamic parameter generation.
[0133] Those skilled in the art will readily appreciate other embodiments of the present application after considering the description and practicing the contents disclosed herein. The present application is intended to cover any modification, use or adaptation of the present application, which follows the general principles of the present application and includes common knowledge or customary techniques in the art that are not disclosed in the present application. The description and examples are intended to be exemplary only, and the true scope and spirit of the present application are indicated by the claims.
[0134] It should be understood that the above general description and the detailed description below are only exemplary and explanatory and cannot limit the present application. The present application is not limited to the precise structure described above and shown in the drawings, and various modifications and changes can be made without departing from the scope thereof. The scope of the present application is limited only by the attached claims.
Claims
1. A cross-language and cross-modal retrieval method based on semantic decoupling and dynamic parameter generation, characterized in that: The method comprises the following steps: S1. Segment and encode the source language and target language sentences to obtain the embedding vectors of the corresponding language texts. S2, input the embedding vector into the text encoders of the source language branch, the target language branch and the semantic decoupling branch to obtain the corresponding feature vector, wherein the source language branch and the target language branch text encoders are 12 layers, and the text encoder of the semantic decoupling branch is 1 layer; The semantic decoupling branch is as follows: using the pre-trained text encoding layer output to two multi-layer perceptrons to obtain semantically relevant and semantically irrelevant feature outputs respectively; concatenating the two obtained features and then passing them through a multi-layer perceptron to obtain the parameters of the dynamic adapter; The target language branch is output from the pre-trained text encoder layer to the dynamic adapter to obtain the final text information as the target language text output or the target language input of the next layer, wherein the parameters of the dynamic adapter are obtained by semantic decoupling; S3, conduct adversarial learning on the output of semantically irrelevant features and source language text information and construct constraints; S4. Based on the target language text features and source language text features output by the text encoder, the MSE loss is used to train the parameters of the dynamic adapter of the target language branch to achieve cross-language migration. S5, input the target language text into the text encoder trained in the cross-language alignment stage to obtain the target language branch text features of cross-modal alignment; S6. Use the pre-trained image encoder to obtain image features, and use InfoNCE loss to calculate the retrieval model that realizes cross-modal alignment of text features and image features; S7. Input the text into the trained retrieval model to achieve cross-language and cross-modal retrieval.
2. According to claim 1, a cross-language and cross-modal retrieval method based on semantic decoupling and dynamic parameter generation is characterized in that: The encoding after word segmentation of the source language and target language sentences is specifically as follows: Load the pre-trained mBERT and clip tokenizers, perform token segmentation on the source language text and the target language text to obtain a fixed-length tokenID vector after token segmentation; Use a multilingual word embedding matrix and the word embedding matrix of clip to perform word embedding operation on the tokenID vector obtained previously, obtain the word embedding encoding vectors of the source language text and the target language text, and add position encoding to the two text sub-embedding vectors.
3. According to claim 1, a cross-language and cross-modal retrieval method based on semantic decoupling and dynamic parameter generation is characterized in that: The text encoder layer is specifically: The obtained embedding vectors of the source language text and the target language text are used to obtain the corresponding feature vectors through the text encoder of each layer of the pre-trained model CLIP. , where i represents the i-th layer of the encoder; ; Here It is at the token level, expressed as .
4. According to claim 1, a cross-language and cross-modal retrieval method based on semantic decoupling and dynamic parameter generation is characterized in that: The semantic decoupling branch is specifically: Only the features that have passed through a layer of text encoder need to be After the semantically relevant and semantically irrelevant features are obtained through the semantically relevant adapter and the semantically irrelevant adapter respectively, they are fused through a feature fusion layer to obtain the parameters of the dynamic adapter. : ; in , , where H represents the size of the hidden layer, I represents the size of the intermediate layer of semantically relevant adapters and semantically irrelevant adapters, Represents a direct splicing operation, and Represent the upper projection layer of semantically relevant adapter and semantically irrelevant adapter respectively, and Represent the lower projection layers of semantically relevant adapters and semantically irrelevant adapters, respectively. and Represents the upper and lower projection layers of the feature fusion layer, and By selecting according to the value of tokenID Vector Get and , represents semantically related features, Represents semantically irrelevant features. After concatenating the two, a feature fusion layer is used to obtain a sentence-level feature vector Z with a dimension of (B, I). Then, a Reshape operation is performed to change the dimension of the feature vector Z from I to ,and , thereby obtaining the parameter matrix of the dynamic adapter .
5. According to claim 1, a cross-language and cross-modal retrieval method based on semantic decoupling and dynamic parameter generation is characterized in that: The target language branch dynamic adapter module specifically includes: inputting the text feature vector into the dynamic adapter module Get the target language text features : , = ; Where i represents the number of encoder layers. , , where H represents the hidden layer size, represents the size of the intermediate layer of the feature fusion layer and , which reduces the number of parameters. Represents the matrix after the adapter; after 12 layers of encoder, the feature vector obtained is based on The value of [EOS] is used to select the [EOS] vector and obtain the feature vector at the target language sentence level .
6. The cross-language and cross-modal retrieval method based on semantic decoupling and dynamic parameter generation according to claim 1, characterized in that: The specific steps in S3 are: the semantically relevant and semantically irrelevant features of the semantic decoupling module need to be constrained; specifically, the features of the semantically irrelevant branches in the semantic decoupling module are kept away from the semantic features of the source language, and the semantically irrelevant feature constraints are: ); Where F is the discriminator, which consists of a simple linear layer. When it becomes smaller, the model will not be able to distinguish The source language features corresponding to the positive sample pairs; The semantically relevant feature constraints are: ; Where B is the batch size, is the source language feature, are semantically related features; Represents the source language sentence-level feature vector.
7. The cross-language and cross-modal retrieval method based on semantic decoupling and dynamic parameter generation according to claim 1, characterized in that: The method of using MSE loss to train the parameters of the dynamic adapter of the target language branch and the parameters of the semantic decoupling branch to achieve cross-language migration is as follows: ; Where B is the batch size, is the source language feature, The target language features.
8. The cross-language and cross-modal retrieval method based on semantic decoupling and dynamic parameter generation according to claim 1, characterized in that: The retrieval model that uses InfoNCE loss calculation to achieve cross-modal alignment of text features and image features is specifically: NCE loss is calculated for image features and text features to train the semantic disentanglement branch and the dynamic adapter part: , , ; in is the temperature coefficient, B is the batch size, Representative image features To text features The similarity loss is Representative text features To image features The similarity loss is represents cosine similarity; Under the constraint of this loss, the model will make relevant text features and image features close to each other, and irrelevant ones far away from each other.
9. A cross-language and cross-modal retrieval device based on semantic decoupling and dynamic parameter generation, comprising a memory and one or more processors, wherein the memory stores executable code, characterized in that: When the processor executes the executable code, a cross-language and cross-modal retrieval method based on semantic decoupling and dynamic parameter generation is implemented as described in any one of claims 1 to 8.
10. A computer-readable storage medium having a program stored thereon, characterized in that: When the program is executed by a processor, a cross-language and cross-modal retrieval method based on semantic decoupling and dynamic parameter generation as described in any one of claims 1 to 8 is implemented.
Citation Information
Patent Citations
Multi-modal neural machine translation method guided by multi-granularity visual pivot based on text perception cross-modal comparison decoupling
CN118313388A
Fused acoustic and text encoding for multimodal bilingual pretraining and speech translation
US20230169281A1