Cross-lingual Text Representation Method, Apparatus, Device, and Storage Medium of Fusion Word Alignment Adapter Module

By constructing a word alignment matrix and integrating an adapter module within a Transformer model for joint masked language and word alignment modeling, the method addresses the challenge of word alignment in cross-language representation, enhancing performance for low-resource languages.

CN115774998BActive Publication Date: 2025-07-15XINJIANG TECH INST OF PHYSICS & CHEM CHINESE ACAD OF SCI
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202211657931.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-12-22
Publication Date
2025-07-15
Estimated Expiration
2042-12-22

AI Technical Summary

Technical Problem

Existing cross-language pre-trained models are difficult to form word alignment mapping on low-resource small languages, resulting in a degradation of cross-language migration performance.

Method used

The source language-target language parallel corpus dataset is constructed, and the word alignment matrix is generated through the unsupervised word alignment algorithm, and a word alignment adapter module is inserted between each sublayer of the translingual pre-trained model of the Transformer structure is performed to conduct joint training of mask language modeling and word alignment modeling to generate a cross-language text representation with word alignment features.

Benefits of technology

The performance of cross-language text representation in low-resource small languages and multiple cross-language downstream tasks is improved, especially in tasks such as cross-language part-of-speech annotation, cross-language syntax analysis, and cross-language naming entity recognition.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115774998B_ABST
    Figure CN115774998B_ABST
Patent Text Reader

Abstract

The present invention discloses a cross - language text representation method, apparatus, device and storage medium integrating a word alignment adapter module, which relates to technical fields such as artificial intelligence, natural language processing, text modeling, etc. The specific implementation solution is as follows: constructing a source - language - target - language parallel corpus dataset, and constructing a word alignment matrix for parallel sentences through an unsupervised word alignment algorithm; inserting a word alignment adapter between each sub - layer of a cross - language pre - trained model with a Transformer structure, and jointly training through masked language modeling and word alignment modeling to achieve semantic alignment of cross - language representation features; inputting the cross - language representation features generated by the word alignment adapter module into a task adapter, so as to implement various cross - language downstream tasks. It solves the problem in the prior art that the cross - language representation effect is poor due to the difficulty of forming a word alignment mapping for low - resource minority languages. According to the technology of the present application, the performance of cross - language text representation for low - resource minority languages and various cross - language downstream tasks is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Embodiments of the present application relate to the field of artificial intelligence in the direction of computer technology, and particularly to technical fields such as natural language processing, deep learning, and text modeling. Specifically, embodiments of the present application provide a cross-lingual text representation method, apparatus, device, and storage medium that integrate a word alignment adapter module. Background Art

[0002] In recent years, under the dual background of globalization and the rapid development of information technology, the digital divide caused by resource differences between different languages has received extensive attention in the academic community. Cross-Lingual Representation maps texts in different languages to the same semantic representation space, thereby extracting unified semantic representation features and serving cross-lingual downstream tasks. Through cross-lingual representation learning technology, it is possible to achieve unified processing of multilingual texts and also realize knowledge transfer from high-resource languages to low-resource languages, which is an important method to narrow the digital divide between different languages.

[0003] Early cross-lingual representation learning mainly focused on static word vector research. Scholars such as Mikolov found that different language word vector spaces obtained by training word embedding models have certain isomorphic characteristics. Therefore, early methods used bilingual dictionaries, etc. as supervision and adopted methods based on linear mapping to learn the word mapping relationship between different languages. Subsequently, some scholars proposed unsupervised cross-lingual word vector training methods and achieved good results. With the pre-trained model showing strong performance in multiple natural language processing tasks, cross-lingual representation learning based on pre-trained models has become the mainstream method. Cross-lingual pre-trained models can better extract multilingual context features and adapt to various downstream tasks.

[0004] The inventors of the present application found in long-term research and development that due to the uneven distribution of training corpora in different languages and the lack of word alignment supervision in masked language modeling training, it is difficult for cross-lingual text representation methods based on pre-trained models to form word-level alignment between the source language and the target language during the training process, resulting in a decline in cross-lingual transfer performance. Summary of the Invention

[0005] The present invention provides a cross - language text representation method, device, equipment and storage medium integrating a word alignment adapter module. This method constructs a source - language - target - language parallel corpus dataset, and constructs a word alignment matrix for parallel sentences through an unsupervised word alignment algorithm; inserts word alignment adapters between each sub - layer of a cross - language pre - trained model with a Transformer structure, and realizes semantic alignment of cross - language representation features through joint training of masked language modeling and word alignment modeling; inputs the cross - language representation features generated by the word alignment adapter module into a task adapter, thereby realizing various cross - language downstream tasks. It solves the problem in the prior art that it is difficult to form a word alignment mapping for low - resource small languages, resulting in poor cross - language representation effects. According to the technology of this application, the performance of cross - language text representation for low - resource small languages and various cross - language downstream tasks is improved.

[0006] The cross - language text representation method integrating a word alignment adapter module according to the present invention is carried out according to the following steps:

[0007] a. Construct a word alignment matrix for the source - language and target - language texts in a bilingual parallel corpus. The bilingual parallel corpus is input into an unsupervised word alignment algorithm model, relevant parameters are set, and word alignment matrices in two directions, including from the source language to the target language and from the target language to the source language, are obtained;

[0008] b. Construct a word alignment adapter module for each source - language - target - language pair and initialize the corresponding adapter module parameters. The word alignment adapter module includes:

[0009] Two feed - forward neural network linear layers, a residual connection layer and a normalization network;

[0010] c. Insert the adapter module between each sub - layer of the encoder of the cross - language pre - trained model with a Transformer structure;

[0011] d. Use the bilingual parallel corpus as the input of the encoder of the cross - language pre - trained model, and conduct joint training of masked language modeling and word alignment modeling on the model, thereby generating cross - language text representations with word alignment features for each language pair;

[0012] e. Connect a task adapter module after the word alignment adapter module to realize specific cross - language downstream tasks.

[0013] Step d takes the bilingual parallel corpus as the input, specifically:

[0014] Concatenate the source - language and target - language parallel sentences, add a separator [SEP] at the concatenation point, and encode them through the tokenizer of the cross - language pre - trained model.

[0015] The joint training of the masked language modeling and word alignment modeling described in step d is specifically as follows: Words in the input text are masked and replaced according to a certain probability, and context modeling training is achieved by inferring the original words at the replacement positions; the word alignment modeling calculates the similarity of the word vectors of the aligned words in the bilingual parallel sentences according to the word alignment matrix to achieve the alignment of synonymous representations; the joint training refers to simultaneously performing the two trainings of masked language modeling and the word alignment modeling; the cross-language text representation with word alignment features is specifically: by using an adapter module, injecting the word alignment information on the basis of the cross-language pre-trained model, so that the synonymous features of different languages are aligned in the semantic space to generate cross-language representation features for serving downstream tasks; the downstream tasks include: cross-language part-of-speech tagging, cross-language syntactic analysis, cross-language named entity recognition, and other natural language processing tasks that rely on cross-language text representation features; the task adapter module includes: a two-layer feed-forward neural network, a residual connection, and a normalization network.

[0016] A cross-language text representation device integrating a word alignment adapter module, the device includes:

[0017] Word alignment matrix construction module: used to obtain a bilingual parallel corpus dataset, perform word alignment training on the dataset through an unsupervised word alignment algorithm, calculate word-level alignment scores for each pair of bilingual texts, and generate a word alignment matrix;

[0018] Word alignment adapter module: The word alignment adapter module is composed of a feed-forward neural network, a residual connection, and a normalization network, and is inserted between each sub-layer of the Transformer encoder. It is used to jointly train the masked language modeling and word alignment modeling of the word alignment adapter model, calculate the masked language loss and the word alignment modeling loss for each input of the source language-target language parallel sentence pair, where the word alignment modeling loss calculates the mean square error loss according to the aligned word pair feature vectors;

[0019] Task adapter module: The word alignment adapter module is composed of a feed-forward neural network, a residual connection, and a normalization network, and uses the cross-language representation features output by the word alignment adapter module as the model input to train specific cross-language downstream tasks.

[0020] Furthermore, the word alignment matrix construction module includes:

[0021] Source language-target language parallel corpus dataset construction unit: used to construct a certain-scale bilingual parallel corpus dataset composed of the source language and the target language;

[0022] A word alignment matrix generation unit, configured to perform word-level alignment on the parallel corpus dataset through an unsupervised word alignment algorithm, and generate a word alignment matrix for each group of parallel sentence pairs based on alignment scores.

[0023] Further, the word alignment adapter module includes:

[0024] A masked language modeling unit, configured to perform masked substitution on the concatenated parallel sentences, and enhance the context feature extraction ability by inferring the original words at the masked positions;

[0025] A word alignment modeling unit, configured to perform word alignment modeling on the concatenated parallel sentences, obtain aligned word pairs through the word alignment matrix, and achieve synonymous semantic alignment by calculating the similarity of the feature vectors of the aligned word pairs.

[0026] Further, the task adapter module includes:

[0027] A cross-lingual downstream task training unit, configured to use the output features of the word alignment adapter module as the input of the task adapter module to implement specific cross-lingual downstream tasks.

[0028] An electronic device, including:

[0029] At least one processor;

[0030] At least one GPU computing card; and

[0031] A memory communicatively connected to the at least one processor; wherein,

[0032] The memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor or the at least one GPU computing card, so that the at least one processor or the at least one GPU computing card can execute the method according to any one of claims 1-8.

[0033] A non-transitory computer-readable storage medium storing computer instructions, wherein the computer instructions are used to cause the computer to execute the method according to any one of claims 1-3.

[0034] The present invention provides a cross-lingual text representation method integrating a word alignment adapter module, and the method includes:

[0035] Construct a word alignment matrix according to the source language and target language parallel texts in the bilingual parallel corpus;

[0036] Insert a word alignment adapter module between each sub-layer of the cross-lingual pre-training model encoder of the Transformer structure;

[0037] Jointly train the model for masked language modeling and word alignment modeling to generate cross - language text representations with word alignment features;

[0038] Connect a task adapter module after the word alignment adapter module to implement specific cross - language downstream tasks.

[0039] According to another aspect of the present invention, there is provided a cross - language text representation device, which includes:

[0040] A word alignment matrix construction module: used to generate a word alignment matrix for bilingual parallel corpora through an unsupervised word alignment algorithm;

[0041] A word alignment adapter module: used for jointly training masked language modeling and word alignment modeling, injecting cross - language word alignment information into the model, and generating cross - language representation features;

[0042] A task adapter module: used to implement specific cross - language downstream tasks, taking the cross - language representation features output by the word alignment adapter module as input.

[0043] According to yet another aspect of the present invention, there is provided an electronic device, which includes:

[0044] At least one processor;

[0045] At least one GPU computing card; and

[0046] A memory communicatively connected to the at least one processor; wherein,

[0047] The memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor or the at least one GPU computing card, so that the at least one processor or the at least one GPU computing card can execute the method described in any one of the embodiments of the present application.

[0048] According to yet another aspect of the present invention, there is provided a non - transitory computer - readable storage medium storing computer instructions, and the computer instructions are used to cause the computer to execute the method described in any one of the embodiments of the present application.

[0049] According to the technology of the present application, cross - language text representation tasks can be completed and the performance of cross - language downstream tasks can be improved.

[0050] The beneficial effects of the present invention are as follows: Aiming at the problem that traditional cross - language pre - trained models have poor representation performance on low - resource minority languages, the present invention first constructs a word alignment matrix for source - language - target - language parallel corpora, and then inserts an adapter module with a small number of parameters into the cross - language pre - trained model of the Transformer structure. By means of the prior knowledge of the pre - trained model, joint training of masked language modeling and word alignment modeling is carried out on the adapter module, so as to form a word - level cross - language semantic mapping. The cross - language text representation features obtained by this method have certain word alignment information and can improve the performance of cross - language downstream tasks, especially related tasks for low - resource minority languages. Brief Description of the Drawings

[0051] Figure 1 It is a flowchart of a cross - language text representation method integrating a word alignment adapter module provided by an embodiment of the present application;

[0052] Figure 2 It is a structural diagram of a cross - language text representation model integrating a word alignment adapter module provided by an embodiment of the present application;

[0053] Figure 3 It is a flowchart of a word alignment matrix construction provided by an embodiment of the present application;

[0054] Figure 4 It is a flowchart of a masked language modeling training method provided by an embodiment of the present application;

[0055] Figure 5 It is a flowchart of a word alignment modeling training method provided by an embodiment of the present application;

[0056] Figure 6 It is a flowchart of a joint training method for a word alignment adapter module provided by an embodiment of the present application;

[0057] Figure 7 It is a structural diagram of a word alignment adapter module provided by an embodiment of the present application;

[0058] Figure 8 It is a flowchart of a task adapter module training method provided by an embodiment of the present application;

[0059] Figure 9 It is a schematic structural diagram of a cross - language text representation device integrating a word alignment adapter module provided by an embodiment of the present application;

[0060] Figure 10 It is a block diagram of an electronic device for cross - language text representation integrating a word alignment adapter module according to an embodiment of the present application. Detailed Description of the Embodiments

[0061] The following further elaborates on the specific embodiments of the present invention in conjunction with the accompanying drawings. It includes various details of the embodiments of the present application to facilitate understanding, and they should be considered merely exemplary. Therefore, those of ordinary skill in the art should recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope and spirit of the present application. Similarly, for the sake of clarity and conciseness, the descriptions of well-known functions and structures are omitted in the following description.

[0062] Embodiment

[0063] A cross - language text representation method integrating a word - alignment adapter module according to the present invention is carried out according to the following steps:

[0064] a. Construct a word - alignment matrix for the source - language and target - language texts in the bilingual parallel corpus. The bilingual parallel corpus is input into an unsupervised word - alignment algorithm model, relevant parameters are set, and word - alignment matrices in two directions, including from the source language to the target language and from the target language to the source language, are obtained;

[0065] b. Construct a word - alignment adapter module for each source - language - target - language pair and initialize the corresponding adapter - module parameters. The word - alignment adapter module includes:

[0066] Two feed - forward neural - network linear layers, a residual - connection layer, and a normalization network;

[0067] c. Insert the adapter module between each sub - layer of the encoder of the cross - language pre - trained model of the Transformer structure;

[0068] d. Use the bilingual parallel corpus as the input to the encoder of the cross - language pre - trained model, and jointly train the model for masked - language modeling and word - alignment modeling, so as to generate cross - language text representations with word - alignment features for each language pair;

[0069] e. Connect a task adapter module after the word - alignment adapter module to implement specific cross - language downstream tasks.

[0070] Step d uses the bilingual parallel corpus as the input, specifically:

[0071] Concatenate the source - language and target - language parallel sentences, add a separator [SEP] at the concatenation point, and encode them through the tokenizer of the cross - language pre - trained model.

[0072] The joint training of the masked language modeling and word alignment modeling described in step d is specifically as follows: The words in the input text are masked and replaced with a certain probability, and context modeling training is achieved by inferring the original words at the replacement positions; the word alignment modeling calculates the similarity of the word vectors of the aligned words in the bilingual parallel sentences according to the word alignment matrix to achieve the alignment of synonym representations; the joint training means that the masked language modeling and the word alignment modeling are performed simultaneously; the cross-language text representation with word alignment features is specifically as follows: By using the adapter module, the word alignment information is injected on the basis of the cross-language pre-trained model, so that the synonym feature representations of different languages are aligned in the semantic space to generate cross-language representation features for downstream tasks; the downstream tasks include cross-language part-of-speech tagging, cross-language syntactic analysis, cross-language named entity recognition, and other natural language processing tasks that rely on cross-language text representation features; the task adapter module includes a two-layer feed-forward neural network, a residual connection, and a normalization network.

[0073] A cross-language text representation device integrating a word alignment adapter module, the device includes:

[0074] Word alignment matrix construction module: used to obtain a bilingual parallel corpus dataset, perform word alignment training on the dataset through an unsupervised word alignment algorithm, calculate word-level alignment scores for each pair of bilingual texts, and generate a word alignment matrix;

[0075] Word alignment adapter module: The word alignment adapter module consists of a feed-forward neural network, a residual connection, and a normalization network, and is inserted between each sub-layer of the Transformer encoder. It is used to jointly train the masked language modeling and word alignment modeling of the word alignment adapter model, calculate the masked language loss and the word alignment modeling loss for each input of the source language-target language parallel sentence pair, and the word alignment modeling loss calculates the mean square error loss according to the feature vectors of the aligned word pairs;

[0076] Task adapter module: The word alignment adapter module consists of a feed-forward neural network, a residual connection, and a normalization network. Using the cross-language representation features output by the word alignment adapter module as the model input, it is used to train specific cross-language downstream tasks.

[0077] Furthermore, the word alignment matrix construction module includes:

[0078] Source language-target language parallel corpus dataset construction unit: used to construct a certain-scale bilingual parallel corpus dataset composed of the source language and the target language;

[0079] A word alignment matrix generation unit, configured to perform word-level alignment on the parallel corpus dataset through an unsupervised word alignment algorithm, and generate a word alignment matrix for each group of parallel sentence pairs based on alignment scores.

[0080] Further, the word alignment adapter module includes:

[0081] A masked language modeling unit, configured to perform masked substitution on the concatenated parallel sentences, and enhance the context feature extraction ability by inferring the original words at the masked positions;

[0082] A word alignment modeling unit, configured to perform word alignment modeling on the concatenated parallel sentences, obtain aligned word pairs through the word alignment matrix, and achieve synonymous semantic alignment by calculating the similarity of the feature vectors of the aligned word pairs.

[0083] Further, the task adapter module includes:

[0084] A cross-lingual downstream task training unit, configured to use the output features of the word alignment adapter module as the input of the task adapter module to implement specific cross-lingual downstream tasks.

[0085] An electronic device, including:

[0086] At least one processor;

[0087] At least one GPU computing card; and

[0088] A memory communicatively connected to the at least one processor; wherein,

[0089] The memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor or the at least one GPU computing card, so that the at least one processor or the at least one GPU computing card can execute the method according to any one of claims 1-8.

[0090] A non-transitory computer-readable storage medium storing computer instructions, wherein the computer instructions are used to cause the computer to execute the method according to any one of claims 1-3;

[0091] Figure 1 It is a flowchart of a cross-lingual text representation method provided by an embodiment of the present application; this embodiment is applicable to the situation of generating unified feature representations for multiple language texts. This method can be executed by a cross-lingual text representation device, and the device can be implemented in software and / or hardware. See Figure 1 As shown in, the cross-lingual text representation method provided by the embodiment of the present application includes:

[0092] S110. Construct a source language-target language parallel corpus dataset:

[0093] In one embodiment, the parallel corpus dataset can be a parallel corpus between any two languages.

[0094] S120. Construct a word alignment matrix for the source language-target language parallel sentences through an unsupervised word alignment algorithm:

[0095] In one embodiment, the unsupervised word alignment algorithm can be understood as finding synonym pairs in mutually parallel bilingual texts;

[0096] S130. Insert a word alignment adapter module between each sub-layer of the cross-lingual pre-trained model with a Transformer structure:

[0097] In one embodiment, the cross-lingual pre-trained model has multiple sub-layers. Inserting a word alignment adapter module between the sub-layers can be understood as using the output hidden layer features of each sub-layer as the input features of the word alignment adapter. Such a structure enables the word alignment adapter to make full use of the prior knowledge of each sub-layer of the pre-trained model;

[0098] S140. Perform joint training of masked language modeling and word alignment modeling on the word alignment adapter:

[0099] In one embodiment, while keeping the parameters of the pre-trained model fixed and the parameters of the word alignment adapter module trainable, by simultaneously calculating the masked language modeling and word alignment modeling losses, the model enhances the word-level semantic alignment between the two languages;

[0100] S150. Insert a task adapter module to implement cross-lingual downstream tasks:

[0101] The task adapter module is connected after the word alignment adapter module, and the output features of the word alignment adapter are used for specific cross-lingual downstream tasks;

[0102] Exemplarily, the cross-lingual downstream tasks can be cross-lingual part-of-speech tagging, cross-lingual syntactic analysis, cross-lingual named entity recognition, and other natural language processing tasks that rely on cross-lingual text representation features;

[0103] Figure 2 It is a structural diagram of a cross-lingual text representation model provided by an embodiment of the present application. The illustrated cross-lingual text representation model mainly consists of four parts, namely a cross-lingual pre-trained model, a word alignment matrix construction, a word alignment adapter, and a task adapter;

[0104] 201. Pre-trained model:

[0105] The pre-trained model is used to provide prior knowledge to the word alignment adapter module and the task adapter module. The pre-trained model consists of the encoder part of the Transformer structure. Each sub-layer is composed of a multi-head attention layer, a residual connection, and a feed-forward neural network, which can capture context features well and is used for a variety of natural language processing tasks. In the pre-training stage, the pre-trained model often adopts modeling paradigms such as masked language modeling and translation language modeling.

[0106] 202. Word Alignment Matrix Construction Module:

[0107] The word alignment matrix construction module performs word alignment on the input bilingual parallel text through an unsupervised word alignment algorithm, that is, finds synonyms in synonymous sentences, and outputs a word alignment matrix for each sample. The word alignment matrix is a binary matrix with a fixed dimension, and its dimension is the unified length after padding the input text. The aligned word pair positions are marked as 1, and the remaining positions are marked as 0.

[0108] 203. Word Alignment Adapter Module:

[0109] The word alignment adapter consists of two feed-forward neural networks, a residual connection, and a normalization network, and is used for joint training of masked language modeling and word alignment modeling. The word alignment adapter module is located between each sub-layer of the cross-lingual pre-trained model, and its number of layers is consistent with that of the pre-trained model. The hidden layer output of the pre-trained model is used as the input of the word alignment adapter, and up and down projection and normalization calculations are performed respectively within the word alignment adapter module.

[0110] Exemplarily, given the input h i , the output after passing through the word alignment adapter module can be obtained through the following formula:

[0111]

[0112] g(h i ) = max(0, h i W i b1)W2 + b2

[0113] In the above formula, are the training parameters of the feed-forward neural network, d represents the hidden layer dimension of the pre-trained model, d f represents the hidden layer dimension of the adapter module, and LN(·) represents the normalization network.

[0114] 204. Task Adapter Module:

[0115] The task adapter consists of two feedforward neural networks, residual connections, and a normalization network, and is used to implement specific cross - language downstream tasks. The task adapter module is located between each sub - layer of the cross - language pre - trained model, following the word alignment adapter, and its number of layers is consistent with that of the pre - trained model; the hidden layer output of the word alignment adapter module serves as the input of the task adapter, and upper and lower projections and normalization calculations are respectively performed within the task adapter module;

[0116] Figure 3 FIG. 4 is a flowchart of constructing a word alignment matrix provided by an embodiment of the present application. The word alignment matrix construction generates a word alignment matrix according to a bilingual parallel corpus for word alignment modeling training. For a schematic diagram of the word alignment matrix, see Figure 2-2 03. The word alignment matrix construction provided by this solution includes:

[0117] S310. Construct a bilingual parallel corpus dataset of a certain scale:

[0118] The parallel corpus dataset must be the same as the parallel corpus dataset in S110;

[0119] S320. Install an unsupervised word alignment algorithm tool and configure the running environment:

[0120] It is necessary to install a tool with unsupervised word alignment function and its necessary running environment in an eligible electronic device and configure relevant parameters;

[0121] S330. Perform unsupervised word alignment on the parallel corpus through the unsupervised word alignment algorithm to obtain word alignment scores:

[0122] The unsupervised word alignment algorithm scores the words in each pair of parallel texts in the parallel corpus dataset to generate word alignment scores;

[0123] S340. Generate a word alignment matrix according to the word alignment scores:

[0124] For each pair of parallel sentence pairs in the parallel corpus dataset, first generate a zero matrix with a fixed dimension. According to the word alignment scores between words, select the word with the highest score as the aligned word and mark it as 1 in the word alignment matrix, thereby generating a word alignment matrix;

[0125] Exemplarily, for the input bilingual text (s, t), constructing a cross - language word alignment matrix, then the aligned word pairs can be expressed as a(s, t) = {(i1, j1),..., (i m , j m )}, for example:

[0126]

[0127] a(s, t) = {(0, 0), (1, 1), (2, 2)}

[0128] Figure 4 is a flowchart of a masked language modeling training method provided by an embodiment of the present application. The masked language modeling training enables the model to infer the original words at the masked positions based on the context, thereby enabling the model to have certain context feature extraction and semantic representation capabilities. The masked language modeling training method provided by this solution includes:

[0129] S410. Use the source language - target language concatenated text as the input:

[0130] The source language text and the corresponding target language text are concatenated, separated by the separator [SEP] inherent in the pre - trained model in the middle, and the concatenated sentence is used as the input of the model;

[0131] Exemplarily, the input text can be: I eat apples [SEP] I eat apples;

[0132] S420. Tokenize and encode through the Tokenizer:

[0133] The Tokenizer inherent in the cross - language pre - trained model will tokenize the input text and map each word to the corresponding input encoding;

[0134] S430. Replace words with a certain ratio and input them into the model:

[0135] A certain ratio of the words in the input text will be replaced with [MASK], which is used to infer the original words at this position according to the context during the masked language modeling process;

[0136] Optionally, the replacement ratio can be adjusted according to the usage scenario, and 15% is a common practice. In the embodiment of the present application, a replacement ratio of 15% is adopted;

[0137] S440. Use the output of the last layer of the word alignment adapter as the semantic feature to infer the original words at the masked positions:

[0138] Use the output features of the last layer of the word alignment adapter module as the input, and the inference layer predicts the original words at the masked positions;

[0139] S450. Calculate the masked language modeling loss according to the true label and perform backpropagation;

[0140] Calculate the loss according to the predicted words and the true words, and perform backpropagation to update the parameters of the word alignment adapter. In the embodiment of the present application, the cross - entropy loss function is used to calculate the loss;

[0141] Figure 5 It is a flowchart of a word alignment modeling training method provided by an embodiment of the present application; according to the word alignment matrix, word vectors of each group of aligned words are obtained, and the similarity of the aligned word vectors is calculated to narrow the distance between the word vectors of synonyms in the semantic space. The word alignment modeling training method provided by this solution includes:

[0142] S510: Use the source language-target language concatenated text as input;

[0143] The source language text and the corresponding target language text are concatenated, separated by the separator [SEP] owned by the pre-trained model in the middle, and the concatenated sentence is used as the input of the model;

[0144] S520: Perform word segmentation and encoding through the Tokenizer;

[0145] The Tokenizer owned by the cross-lingual pre-trained model will perform word segmentation on the input text and map each word to the corresponding input encoding;

[0146] S530: Use the output of the last layer of the word alignment adapter as the semantic feature, and extract the word vectors of the aligned word pairs according to the word alignment matrix;

[0147] The encoded text is input into the model, and the output of the last layer of the word alignment adapter module is extracted as the semantic feature, and the word vectors of the aligned word pairs are extracted according to the word alignment matrix generated in the previous step;

[0148] S540: Calculate the word alignment similarity loss according to each pair of word vectors and perform backpropagation:

[0149] Calculate the similarity of the word vectors of each group of aligned word pairs, use the similarity score as the word alignment loss, and perform backpropagation based on this;

[0150] Optionally, the similarity calculation algorithm can be Euclidean distance, Chebyshev distance, cosine similarity, Pearson correlation coefficient, etc. In the embodiment of the present application, the MSE mean square loss is adopted, that is, the sum of the squares of the Euclidean distances;

[0151] Exemplarily, let f be a cross-lingual pre-trained model with a word alignment adapter module, sim(·) be the similarity calculation function, Ali(·) be the word alignment adapter, and (i, t) represent the t-th word of the i-th sample. Then the word alignment modeling training loss L ALIGN Can be expressed as:

[0152] L ALIGN (f; C) = -∑ (s,t)∈C ∑ (i,j)∈a(s,t) sim(Ali(i; s), Ali(i; t))

[0153] Figure 6 It is a flowchart of a joint training method for a word alignment adapter module provided by an embodiment of the present application. In order to enable the model to fully extract context features and perform better semantic alignment when representing multilingual texts, the embodiment of the present application adopts a joint training method. The joint training method for the word alignment adapter module provided includes:

[0154] S610. Keep the pre-trained model parameters fixed and the word alignment adapter parameters trainable:

[0155] In order to make full use of the prior knowledge of the pre-trained model and prevent the problem of catastrophic forgetting, a scheme of keeping the original parameters of the cross-lingual pre-trained model fixed and the adapter parameters trainable is adopted;

[0156] S620. Use the output of the last layer of the word alignment adapter as the semantic feature:

[0157] When the joint training method realizes masked language modeling and word alignment modeling, the output of the last layer of the word alignment adapter will be used as the semantic feature vector, that is, the cross-lingual text representation feature;

[0158] S630. Calculate the masked language modeling loss and the word alignment modeling loss respectively according to the semantic feature:

[0159] The semantic feature obtained according to S620 will be used to calculate the masked language modeling loss and the word alignment modeling loss simultaneously;

[0160] S640. Calculate the sum of the masked language modeling loss and the word alignment modeling loss as the joint loss:

[0161] The joint training method can be understood as combining the losses of multiple tasks. In the embodiment of the present application, the sum of the masked language modeling loss and the word alignment modeling loss is used as the joint loss:

[0162] Exemplarily, let L MLM be the masked language modeling loss, and L ALIGN be the word alignment loss. Then the joint loss Loss can be expressed as:

[0163] Loss = L MLM + L ALIGN

[0164] S650. Perform backpropagation according to the joint loss:

[0165] Figure 7 It is a structural diagram of a word alignment adapter module provided by an embodiment of the present application. Among them, the adapter module is composed of two linear transformation layers, a residual link and a normalization network;

[0166] The linear transformation layers are an upper projection layer and a lower projection layer respectively, and there is an activation function operation between the two layers;

[0167] Figure 8 It is a flowchart of a method for training a task adapter module provided by an embodiment of the present application. The overall structure of the task adapter is the same as that of the word alignment adapter, but it uses the output features of the word alignment adapter module as input to implement specific cross - language downstream tasks. The method for training the task adapter module provided by this solution includes:

[0168] S810. Keep the parameters of the word alignment adapter fixed and connect the task adapter module behind it;

[0169] After the word alignment adapter module is jointly trained by masked language modeling and word alignment modeling, it has cross - language representation ability. Therefore, keep its parameters fixed and connect the task adapter behind it;

[0170] S820. The output of the word alignment adapter is used as semantic features and input to the task adapter;

[0171] S830. The task adapter module is trained for specific cross - language tasks;

[0172] The task adapter is used to train specific cross - language downstream tasks, and different tasks have different task adapter modules. Among them, the same word alignment adapter can be connected to multiple task adapters to achieve multi - task training, but generally the same task adapter module is only used for the same task;

[0173] Exemplarily, the cross - language downstream tasks include but are not limited to cross - language part - of - speech tagging, cross - language named entity recognition, cross - language syntactic analysis, etc.;

[0174] Figure 9 It is a schematic structural diagram of a cross - language text representation device integrating a word alignment adapter module provided by an embodiment of the present application. Refer to Figure 9 This cross - language text representation device integrating a word alignment adapter module provided in this embodiment includes: a word alignment matrix construction module, a word alignment adapter module, and a task adapter module. Among them,

[0175] The word alignment matrix construction module: is used to obtain a bilingual parallel corpus dataset, perform word alignment training on the dataset through an unsupervised word alignment algorithm, calculate word - level alignment scores for each pair of bilingual texts, and generate a word alignment matrix;

[0176] Word alignment adapter module: The word alignment adapter module consists of a feed-forward neural network, a residual connection, and a normalization network, and is inserted between each sub-layer of the Transformer encoder for jointly training the word alignment adapter model for masked language modeling and word alignment modeling. For each set of source language-target language parallel sentence inputs, it calculates the masked language loss and the word alignment modeling loss, where the word alignment modeling loss calculates the mean square error loss based on the aligned word pair feature vectors;

[0177] Task adapter module: The word alignment adapter module consists of a feed-forward neural network, a residual connection, and a normalization network, and uses the cross-language representation features output by the word alignment adapter module as the model input for training specific cross-language downstream tasks.

[0178] Furthermore, the word alignment matrix construction module includes:

[0179] Source language-target language parallel corpus dataset construction unit, used to construct a certain-scale bilingual parallel corpus dataset composed of source language and target language;

[0180] Word alignment matrix generation unit, used to achieve word-level alignment for the parallel corpus dataset through an unsupervised word alignment algorithm, and generate a word alignment matrix for each set of parallel sentence pairs through alignment scores.

[0181] Furthermore, the word alignment adapter module includes:

[0182] Masked language modeling unit, used to perform mask replacement on the concatenated parallel sentences and enhance the context feature extraction ability by inferring the original words at the masked positions;

[0183] Word alignment modeling unit, used to perform word alignment modeling on the concatenated parallel sentences, obtain aligned word pairs through the word alignment matrix, and achieve synonymous semantic alignment by calculating the similarity of the aligned word pair feature vectors.

[0184] Furthermore, the task adapter module includes:

[0185] Cross-language downstream task training unit, used to use the output features of the word alignment adapter module as the input of the task adapter module to implement specific cross-language downstream tasks.

[0186] According to the embodiments of the present application, the present application also provides an electronic device and a readable storage medium.

[0187] Such as Figure 10As shown, it is a block diagram of an electronic device for a small-sample intent recognition method according to an embodiment of the present application. The electronic device refers to various modern electronic digital computers, including, for example: personal computers, portable computers, and various server devices. The components, their interconnections, and functions shown herein are only examples;

[0188] As Figure 10 shown, the electronic device includes: one or more multi-core processors, one or more GPU computing cards, and a memory. To enable the electronic device to interact, it should also include: an input device and an output device. Various devices are interconnected and communicate through a bus.

[0189] The memory is the non-transitory computer-readable storage medium provided by the present application. Among them, the memory stores instructions executable by at least one multi-core processor or at least one GPU computing card, so that the cross-language text representation method of the fusion word alignment adapter module provided by the present application can be executed. The non-transitory computer-readable storage medium of the present application stores computer instructions, and these computer instructions are used to cause a computer to execute the cross-language text representation method of the fusion word alignment adapter module provided by the present application.

[0190] The input device provides and accepts control signals input by the user into the electronic device, including a keyboard that generates digital or character information and a mouse that is used to control the device to generate other key signals. The output device provides feedback information to the user about the electronic device, including a display that prints the execution result or process.

[0191] It should be understood that the present invention is not limited to the exact structure already described and shown in the drawings, and various modifications and changes can be made without departing from its scope. The scope of the present invention is only limited by the appended claims.

Claims

1. A cross - language text representation method integrating a word alignment adapter module, characterized in that Proceed as follows: a. Construct a word alignment matrix for the source language and target language texts in the bilingual parallel corpus. The bilingual parallel corpus is input into an unsupervised word alignment algorithm model, relevant parameters are set, and word alignment matrices in both directions from the source language to the target language and from the target language to the source language are obtained; b. Construct a word alignment adapter module for each source language - target language pair and initialize the corresponding adapter module parameters. The word alignment adapter module includes: Two feed - forward neural network linear layers, a residual connection layer, and a normalization network; c. Insert the adapter module between each sub - layer of the encoder of the cross - language pre - trained model of the Transformer structure; d. Use the bilingual parallel corpus as the input to the encoder of the cross - language pre - trained model, and jointly train the model for masked language modeling and word alignment modeling, so as to generate cross - language text representations with word alignment features for each language pair; e. Connect a task adapter module after the word alignment adapter module to implement specific cross - language downstream tasks.

2. The cross-language text representation method of a fusion word alignment adapter module according to claim 1, characterized in that Step d uses the bilingual parallel corpus as the input, specifically: Concatenate the parallel sentences of the source language and the target language, add a separator at the concatenation point, and encode them through the tokenizer of the cross - language pre - trained model.

3. A cross - language text representation method for a fusion word alignment adapter module according to claim 1, characterized in that, The joint training of masked language modeling and word alignment modeling in step d is specifically: randomly mask some words in the input text with a certain probability, and infer the original words at the masked positions to achieve context modeling training; for word alignment modeling, calculate the similarity of the word vectors of the aligned words in the bilingual parallel sentences according to the word alignment matrix to achieve synonym representation alignment; the joint training means performing both masked language modeling and word alignment modeling at the same time; the cross - language text representation with word alignment features is specifically: by using the adapter module, inject the word alignment information on the basis of the cross - language pre - trained model, so that the synonym feature representations of different languages are aligned in the semantic space, generate cross - language representation features, and serve downstream tasks; The downstream tasks include: cross - language part - of - speech tagging, cross - language syntactic analysis, cross - language named entity recognition, and other natural language processing tasks that rely on cross - language text representation features; the task adapter module includes: two - layer feed - forward neural network, residual connection, and normalization network.

4. A cross - language text representation device integrating a word alignment adapter module, characterized in that The device includes: A word alignment matrix construction module: used to obtain a bilingual parallel corpus dataset, perform word alignment training on the dataset through an unsupervised word alignment algorithm, calculate word - level alignment scores for each pair of bilingual texts, and generate a word alignment matrix; Word alignment adapter module: The word alignment adapter module consists of a feed-forward neural network, a residual connection, and a normalization network, which is inserted between each sub-layer of the Transformer encoder for joint training of masked language modeling and word alignment modeling of the word alignment adapter model. For each set of source language-target language parallel sentence inputs, the masked language loss and the word alignment modeling loss are calculated, where the word alignment modeling loss calculates the mean squared error loss based on the aligned word pair feature vectors; Task adapter module: The word alignment adapter module consists of a feed-forward neural network, a residual connection, and a normalization network. Using the cross-language representation features output by the word alignment adapter module as the model input, it is used to train specific cross-language downstream tasks; Further, the word alignment matrix construction module includes: Source language-target language parallel corpus dataset construction unit, which is used to construct a certain-scale bilingual parallel corpus dataset composed of source language and target language; Word alignment matrix generation unit, which is used to achieve word-level alignment of the parallel corpus dataset through an unsupervised word alignment algorithm. For each pair of parallel sentences, a word alignment matrix is generated through alignment scores; Further, the word alignment adapter module includes: Masked language modeling unit, which is used to perform mask substitution on the concatenated parallel sentences and enhance the context feature extraction ability by inferring the original words at the masked positions; Word alignment modeling unit, which is used to perform word alignment modeling on the concatenated parallel sentences, obtain aligned word pairs through the word alignment matrix, and achieve synonymous semantic alignment by calculating the similarity of the aligned word pair feature vectors; Further, the task adapter module includes: Cross-language downstream task training unit, which is used to use the output features of the word alignment adapter module as the input of the task adapter module to implement specific cross-language downstream tasks.

5. An electronic device, wherein, Includes: At least one processor; At least one GPU computing card; And A memory communicatively connected to the at least one processor; wherein, The memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor or the at least one GPU computing card, so that the at least one processor or the at least one GPU computing card can execute the method according to any one of claims 1-3.

6. A non-transitory computer-readable storage medium storing computer instructions, characterized in that, The computer instructions are used to cause the computer to execute the method according to any one of claims 1-3.

Citation Information

Patent Citations

  • Text processing method and device

    CN111563381A

  • Semi-supervised adversarial learning cross-language abstract generation method based on word alignment

    CN112541343A