Cross-modal non-autoregressive decoding method and system combined with pre-training language model

By introducing the idea of ​​Random Drop and the CTC-enhanced soft copy method in the pre-training task, the non-autoregressive pre-trained language model cannot effectively model the vocabulary dependence and modal difference of the target end, and a more efficient cross-modal non-autoregressive translation effect is achieved.

CN119962548AActive Publication Date: 2025-05-09XIAMEN UNIV
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202510093157.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-21
Publication Date
2025-05-09
Estimated Expiration
2045-01-21

AI Technical Summary

Technical Problem

The existing non-autoregressive pretrained language models cannot effectively model the dependence between target-side vocabulary, and the problem of modal differences in decoder input during the pretraining and fine-tuning stages exists.

Method used

The idea of ​​Random Drop is introduced to model the dependence of target-side vocabulary in pre-training tasks, and the target-side information is integrated into the decoder input of the downstream cross-modal non-autoregressive model through a CTC-enhanced soft copy method to alleviate the problem of modal differences.

Benefits of technology

Effectively modeling the dependence between target-side vocabulary, alleviates the problem of modal differences in decoder inputs in pre-training and fine-tuning stages, and improves the performance of cross-modal non-autoregressive translation models.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119962548A_ABST
    Figure CN119962548A_ABST
Patent Text Reader

Abstract

The invention discloses a cross-modal non-autoregressive decoding method and system combined with a pre-training language model, and relates to cross-modal translation. The decoding system sequentially comprises a target arrangement sampling module, a pre-training text coding module, a pre-training text decoding module, a pre-training loss calculation module, a source end feature extraction module, a downstream coding module, a CTC enhancement module, a downstream decoding module, a downstream task loss calculation module, an output module and the like. The decoding method comprises the following steps: sampling a pre-training sample and a target arrangement; encoding the source end text into an implicit vector; decoding a target end text; carrying out loss calculation in a pre-training stage; parameters of a downstream cross-modal non-autoregression translation model are initialized; extracting source end features; source end feature coding; generating downstream decoding module input; the decoding module generates a translation result; the loss of the downstream cross-modal non-autoregression translation model is calculated; and outputting a downstream cross-modal non-autoregressive translation result. Random Drop and CTC are utilized to relieve the modal difference of decoder input in the pre-training stage and the fine tuning stage.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the fields of cross-modal translation, non-autoregressive text generation, and pre-trained language models, and in particular to a cross-modal non-autoregressive decoding method and system combined with a pre-trained language model. Background Art

[0002] In recent years, non-autoregressive (NAR) decoding mechanisms have been applied to cross-modal translation tasks such as sign language translation, effectively reducing the inference delay of the model. Currently, cross-modal non-autoregressive translation models using curriculum-based non-autoregressive decoders (CND) outperform autoregressive baseline models in both speed and performance.

[0003] The research paradigm of combining pre-training with fine-tuning has attracted widespread attention in academia. Existing work in academia has shown that the ability of cross-modal autoregressive translation models can be significantly improved after combining with autoregressive pre-trained language models. However, there are two problems in directly combining existing non-autoregressive pre-trained language models with cross-modal non-autoregressive translation models: First, existing non-autoregressive pre-trained language models cannot explicitly model the dependencies between target-side words, which is very important for cross-modal non-autoregressive translation models. Second, there is a modality difference between the representation space of the decoder input of the existing non-autoregressive pre-trained language model and the decoder input of the cross-modal non-autoregressive translation model. This is because the input of the non-autoregressive decoder is usually copied from the word embedding of the encoder input. For the non-autoregressive pre-trained language model, the input of the decoder is in the same text representation space as the encoder input. However, for the downstream cross-modal non-autoregressive translation model, the input of the encoder is in the non-text representation space, which causes the input of the decoder to be in the same non-text representation space as the encoder during the fine-tuning stage, resulting in the modality difference problem in the two stages. Summary of the invention

[0004] One object of the present invention is to better integrate a pre-trained language model into a non-autoregressive cross-modal translation model, and to provide a cross-modal non-autoregressive decoding method combined with a pre-trained language model by using the ideas of Random Drop and CTC.

[0005] A second object of the present invention is to realize a cross-modal non-autoregressive translation system combined with a pre-trained language model through a cross-modal non-autoregressive decoding method.

[0006] To achieve the above objectives, the present invention provides the following technical solutions.

[0007] The present invention provides a cross-modal non-autoregressive decoding method combined with a pre-trained language model, comprising the following steps:

[0008] 1) Sample pre-training training samples and target permutations: Sample a data pair from the corpus used for pre-training and sample a permutation of the target sentence ,in, Each element in is the subscript of the target word in the target sequence;

[0009] 2) Encode the source text into a latent vector: Use the text encoding module of the pre-trained model to encode the source text into a latent vector. Encoded as a latent vector;

[0010] 3) Target-side text decoding: The word embedding of the pre-trained source-side text is converted into the input of the non-autoregressive decoder through a copy mechanism, and then input into the text decoding module; after receiving the output vector of the pre-trained text encoding module, the text decoding module generates the translation result through the feature vector of the encoding module;

[0011] 4) Pre-training loss calculation: Use cross entropy loss to calculate the translation loss of each decoder layer, and add up the losses of all layers to get the total loss;

[0012] 5) Initialize the parameters of the downstream cross-modal non-autoregressive translation model:

[0013] After the pre-trained language model is trained, its parameters can be used to initialize the parameters of the encoder and decoder of the downstream cross-modal non-autoregressive translation model;

[0014] 6) Source feature extraction: The feature extraction module extracts features from each frame of source data to obtain a feature sequence containing source information;

[0015] 7) Source feature encoding: The source feature vector extracted in step 6) is sent to the downstream encoding module to obtain hidden representation;

[0016] 8) Downstream decoding module input generation: Combine the predicted probability of the CTC layer and the weight matrix of the target word embedding layer, calculate the attention score matrix through position encoding and learnable weights, and use it to generate the input of the decoding module to integrate the target context information;

[0017] 9) The decoding module generates the translation result: The downstream decoding module generates the translation result by simultaneously obtaining the input of the decoding module and the feature vector output of the encoding module;

[0018] 10) Calculate the loss of the downstream cross-modal non-autoregressive translation model: In the downstream loss calculation module, the CTC loss function is used to calculate the loss of the downstream encoding module, and the cross entropy loss function is used to calculate the translation loss of the generated words in the downstream decoding module;

[0019] 11) Downstream cross-modal non-autoregressive translation result output: The spoken sentence obtained by the downstream cross-modal non-autoregressive translation model is output through the output module.

[0020] In step 3), the specific steps of target-side text decoding may be:

[0021] The source word embedding is converted into the input of the non-autoregressive decoder through a copy mechanism, requiring each layer of the decoder to make predictions, and each layer only retains the subscripts belonging to Before The decoder first extracts the words in the elements and replaces the remaining words with a special symbol [mask]; concatenates the word embedding predicted by this layer with the hidden state of this layer, and uses an additional linear layer to reduce the dimension to get the output of the decoder at the current layer; in this way, the high-level decoder can receive the predictions of the low-level decoder, thereby predicting more words, and the entire sentence will be gradually predicted by the decoder from the bottom up.

[0022] In step 4), the specific steps of the pre-training stage loss calculation can be: in the loss calculation module of the pre-training stage, the translation loss of the non-autoregressive decoding module is calculated using the cross entropy loss; since the decoder performs predictions at each layer, the loss function of the pre-training stage This is the sum of the cross entropy losses of each layer of the decoder:

[0023]

[0024] in, Indicates the decoder Layer The probability distribution of words; Indicates the number of decoder layers; represents the number of words predicted in the i-th layer, , Indicates the total number of words, represents the tth target word;

[0025] In step 8), the specific steps of generating the input of the downstream decoding module can be: first obtain the predicted probability of the top CTC layer of the encoder, and then multiply it with the weight matrix of the pre-trained target end word embedding layer to obtain a representation containing the target end context In order to convert its length to the length of the target sequence, position encoding and a learnable weight are used to calculate the attention score matrix. The input of the CTC-enhanced decoding module is obtained by multiplying the attention score matrix and the representation. In this way, the target context information can be incorporated into the decoder input in the fine-tuning stage, thereby mapping its representation space from the source to the target, alleviating the modality difference problem.

[0026] In step 9), the specific steps of generating the translation result by the decoding module may be: the downstream decoding module generates the translation result by simultaneously acquiring the feature vectors of the input of the decoding module and the output of the encoding module; specifically, the present invention requires that each layer of the decoder performs prediction, and each layer only retains the top prediction with the highest confidence. words, and then replace the remaining words with a special symbol [mask]; then, the word embedding predicted by this layer is concatenated with the hidden state of this layer, and an additional linear layer is used for dimensionality reduction to obtain the output of the decoder at the current layer; in this way, the lower layers of the decoder can provide the most confident prediction results for the higher layers, thereby promoting the prediction of complex words by the higher layers; in addition, the number of words to be predicted at each layer increases linearly with the increase in the number of layers, and the top layer will predict the entire target sentence, so the decoder can predict the entire sentence from bottom to top and from easy to difficult.

[0027] In step 10), the downstream cross-modal non-autoregressive translation model loss is calculated, and during the fine-tuning experiment of the downstream task, the model is optimized by jointly minimizing the encoder loss and the decoder loss:

[0028]

[0029] in, represents the downstream cross-modal non-autoregressive translation model loss; Represents the translation loss of a word, represents the loss of the downstream encoding module; α is a hyperparameter used to balance the two training losses.

[0030] The present invention provides a cross-modal non-autoregressive decoding system combined with a pre-trained language model, which comprises: a target arrangement sampling module, a pre-trained text encoding module, a pre-trained text decoding module, a pre-trained loss calculation module, a source feature extraction module, a downstream encoding module, a CTC enhancement module, a downstream decoding module, a downstream task loss calculation module, and an output module;

[0031] The target permutation sampling module is used to randomly select a source text and a corresponding target text from the pre-training corpus, and perform permutation transformation on the target text;

[0032] The pre-trained text encoding module is used to convert the source text into a series of latent vector representations;

[0033] The pre-trained text decoding module is used to convert the word embedding of the source text into the input of the non-autoregressive decoder through a copy mechanism, and input it into the text decoding module; the decoding module generates a translation result using the output of the encoding module;

[0034] The pre-training loss calculation module is used to calculate the cross entropy loss of each decoding process, and accumulate the losses of all steps to obtain the total loss, which is used to guide model optimization;

[0035] The source feature extraction module is used to extract features from source data to obtain a feature sequence containing source information, and the extracted source feature vector is sent to a downstream encoding module;

[0036] The downstream encoding module is used to further encode the feature vector obtained by the source feature extraction module to obtain a higher level of abstract representation;

[0037] The CTC enhancement module is used to improve the model's ability to process sequence data. It combines the prediction probability of the CTC (Connectionist Temporal Classification) layer and the weight matrix of the target-side word embedding layer, calculates the attention score matrix through position encoding and learnable weights, and uses it to generate the input of the decoding module to fuse the target-side context information.

[0038] The downstream decoding module generates a translation result by simultaneously acquiring the input of the decoding module and the feature vector output of the encoding module. The parallel processing method is a typical non-autoregressive method.

[0039] The downstream task loss calculation module uses the CTC loss function to calculate the loss of the encoding module and uses the cross entropy loss function to calculate the translation loss of the generated words in the decoding module. These losses are used to guide the optimization of the entire downstream model.

[0040] The output module is used to output the translation results obtained by the downstream cross-modal non-autoregressive translation model, such as spoken sentences.

[0041] The advantages of the present invention are:

[0042] 1. The present invention introduces the idea of ​​Random Drop into the pre-training task of the non-autoregressive language model, effectively modeling the dependencies between target-side words.

[0043] 2. The present invention introduces the idea of ​​CTC into the downstream cross-modal non-autoregressive translation task. By incorporating text modality information into the decoder input, the modality difference problem of the decoder input in the pre-training and fine-tuning stages is effectively alleviated.

[0044] 3. The performance of the cross-modal non-autoregressive decoding method combined with a pre-trained language model proposed in the present invention is always better than all strong baselines, and can be applied to different types of cross-modal translation tasks. It has a broad application prospect and good application value. BRIEF DESCRIPTION OF THE DRAWINGS

[0045] Figure 1 A framework diagram of a cross-modal non-autoregressive decoding system combined with a pre-trained language model.

[0046] Figure 2 Flowchart of a cross-modal non-autoregressive decoding method combined with a pre-trained language model.

[0047] Figure 3 A cross-modal non-autoregressive translation system combined with a pre-trained language model. DETAILED DESCRIPTION

[0048] In order to better understand the above technical solution, the above technical solution will be described in detail below in conjunction with the drawings and specific implementation methods of the specification. The invention content of the present invention is described in detail below, and examples of the invention content are shown in the drawings, wherein the same or similar reference numerals throughout represent the same or similar elements or elements with the same or similar functions. The embodiments described below with reference to the drawings are exemplary and are intended to be used to explain the present invention, and should not be understood as limiting the present invention.

[0049] In order to solve the two problems existing in the above-mentioned prior art, firstly, the present invention is inspired by Random DropTraining and believes that each layer of the decoder can predict different parts of the entire target sequence during the pre-training process, and at the same time, the high-level layer of the decoder can perceive the prediction of the bottom layer, so as to learn the inter-layer dependencies of the target end vocabulary; and which words are predicted by each layer can be determined by random arrangement sampling. In this way, the decoder generates the entire sequence from bottom to top in a different order in each training iteration, so as to learn a variety of dependencies, thereby alleviating the first problem.

[0050] In order to alleviate the modality difference problem of decoder input in the pre-training and fine-tuning stages, the present invention is inspired by the speech translation and automatic speech recognition models based on CTC (Connectionist Temporal Classification). These methods usually add a CTC prediction layer on the top layer of the encoder and use the target text as a supervisory signal. In this way, a rough translation of the target text can be obtained. The present invention considers that if this rough translation result can be integrated into the decoder input, on the one hand, the decoder input can be mapped from the source end to the target end, thereby alleviating the modality difference problem; on the other hand, it can provide the decoder with certain target end information, making the model easier to learn.

[0051] Based on the above considerations, the present invention first uses the random permutation progressive prediction pre-training task to learn the diverse dependencies in the non-autoregressive decoder on large-scale pre-trained text pairs; then uses the CTC-enhanced soft copy method to integrate the target-end information into the decoder of the downstream cross-modal non-autoregressive model to alleviate the modality difference problem.

[0052] By introducing the ideas of Random Drop and CTC, the present invention first uses the random permutation progressive prediction pre-training task to model the diverse dependencies in the non-autoregressive decoder on large-scale pre-trained text pairs; then uses the CTC-enhanced soft copy method to integrate the target-side information into the decoder of the downstream cross-modal non-autoregressive model to alleviate the modality difference problem.

[0053] like Figure 1 As shown, an embodiment of the present invention provides a cross-modal non-autoregressive decoding system combined with a pre-trained language model, the system sequentially includes a target arrangement sampling module, a pre-trained text encoding module, a pre-trained text decoding module, a pre-trained loss calculation module, a source feature extraction module, a downstream encoding module, a CTC enhancement module, a downstream decoding module, a downstream task loss calculation module, and an output module;

[0054] Target permutation sampling module: used to randomly select a source text and a corresponding target text from the pre-training corpus, and perform permutation transformation on the target text;

[0055] Pre-trained text encoding module: used to convert source text into a series of latent vector representations;

[0056] Pre-trained text decoding module: used to convert the word embedding of the source text into the input of the non-autoregressive decoder through a copy mechanism, and input it into the text decoding module; the decoding module uses the output of the encoding module to generate the translation result;

[0057] Pre-training loss calculation module: used to calculate the cross entropy loss of each decoding process, accumulate the losses of all steps to get the total loss, and guide model optimization;

[0058] Source feature extraction module: used to extract features from source data to obtain a feature sequence containing source information. The extracted source feature vector is sent to the downstream encoding module.

[0059] Downstream encoding module: used to further encode the feature vector obtained by the source feature extraction module to obtain a higher level of abstract representation;

[0060] CTC enhancement module: used to improve the model's ability to process sequence data. It combines the prediction probability of the CTC (Connectionist Temporal Classification) layer and the weight matrix of the target-side word embedding layer, calculates the attention score matrix through position encoding and learnable weights, and uses it to generate the input of the decoding module to fuse the target-side context information.

[0061] Downstream decoding module: The downstream decoding module generates the translation result by simultaneously obtaining the input of the decoding module and the feature vector output of the encoding module. The parallel processing method is a typical non-autoregressive method.

[0062] Downstream task loss calculation module: In the downstream loss calculation module, the CTC loss function is used to calculate the loss of the encoding module, and the cross entropy loss function is used to calculate the translation loss of the generated words in the decoding module. These losses are used to guide the optimization of the entire downstream model;

[0063] Output module: The output module outputs the translation results, such as spoken sentences, obtained by the downstream cross-modal non-autoregressive translation model.

[0064] like Figure 2 , an embodiment of the present invention provides a cross-modal non-autoregressive decoding method combined with a pre-trained language model, comprising the following steps:

[0065] 1) Sample pre-training samples and target permutations: Sample text data pairs used for pre-training and sample a permutation of target sentences, with each element using the subscript of the target vocabulary in the target sequence. Specifically:

[0066] First, sample a piece of data from the corpus used for pre-training , and use the target sequence sampling module to sample a permutation of the target sequence ,Notice Each element in is the index of the target word in the target sequence rather than the word itself. For example, {1, 2, 4, 0, 3} can be a permutation of a target sequence of length 5.

[0067] 2) Input to the pre-trained text encoding module: The pre-trained text encoding module using the pre-trained model will The code is converted into a hidden vector and then into a hidden state. The pre-trained text encoding module includes a word embedding layer, a self-attention layer, a residual connection, a layer normalization layer, and a feedforward neural network layer stacked multiple times.

[0068] 3) Input to the pre-trained text decoding module: First, use the copy mechanism to The word embedding is converted into the input of the non-autoregressive decoder and input into the pre-trained text decoding module. After receiving the output vector of the text encoding module, the decoding module generates the translation result through the feature vector of the encoding module.

[0069] Specifically, consider the total The first layer of the pre-trained text decoding module Layer, first get the hidden vector of this layer , and pass it to the output layer to obtain the prediction of this layer . These operations are expressed as:

[0070]

[0071]

[0072] in, Represents the weight matrix of the linear output layer.

[0073] At the same time, only keeping The subscript belongs to Before The words in the elements are then replaced with the special symbol [mask] to get the updated .

[0074] Finally, splicing The word embedding and , and use an additional linear layer to reduce the dimension, we can get the decoder Output of the layer :

[0075]

[0076] in Represents the weight matrix of the linear layer. Since the decoder is The output of the layer is also In this way, the high-level decoder can receive the predictions of the low-level decoder, thereby predicting more words, and the entire sentence will be gradually predicted by the decoder from bottom to top.

[0077] 4) Calculate the loss in the pre-training stage: In the loss calculation module in the pre-training stage, the cross entropy loss function is used to calculate the loss of the non-autoregressive decoding module. Since the decoder makes predictions at each layer, the loss function in the pre-training stage is It is the sum of the cross entropy losses of each layer of the decoder. Specifically:

[0078] In the loss calculation module of the pre-training stage, the cross entropy loss is used to calculate the translation loss of the decoding module. Since the decoder makes predictions at each layer, the loss function of the pre-training stage is This is the sum of the cross entropy losses of each layer of the decoder:

[0079]

[0080] in Indicates the decoder Layer The probability distribution of words, and . Indicates the number of decoder layers; represents the number of words predicted in the i-th layer, , Indicates the total number of words, represents the tth target word.

[0081] 5) Initialize the parameters of the downstream cross-modal translation model: After training the non-autoregressive pre-trained language model, use the pre-trained model and use its parameters to initialize the parameters of the encoder and decoder of the downstream cross-modal non-autoregressive translation model.

[0082] 6) Input source data: Prepare source data (such as source language text and corresponding multimodal data such as images or audio), and extract features from each frame of the source data (sign language or speech, etc.) in the feature extraction module to obtain features containing source information.

[0083] 7) Source feature extraction: In the feature extraction module, feature extraction is performed on each frame of the source non-text data to obtain its frame-level features and obtain a feature sequence containing source information.

[0084] 8) Input to the downstream encoding module: The frame-level features extracted in step 7) are sent to the downstream encoding module, encoded by the encoder to obtain the hidden representation, and the target prediction probability after the CTC layer is obtained; specifically:

[0085] The extracted source feature vector is sent to the downstream encoding module, and the hidden representation is obtained through the encoder-decoder. Here, CTC (Connectionist Temporal Classification) is used as the loss function of the encoding module, and the supervision signal is the target text. The purpose is to enable the encoder to provide a probability distribution containing the target text information during the training phase. The encoding module has an attention mechanism, which can get rid of the limitation of the input sequence length. It is a technology that enables the model to focus on important information and fully learn and absorb it.

[0086] 9) Input to the CTC enhancement module: Combine the target-side prediction probability of the CTC layer and the weight matrix of the target-side word embedding layer, calculate the attention score matrix through position encoding and learnable weights, and use it to generate the input of the decoding module to fuse the target-side context information.

[0087] In the CTC enhancement module, first obtain the predicted probability of the top CTC layer of the encoder , and then the weight matrix of the pre-trained target word embedding layer and Multiply to get a representation containing the target context However, at this time Still the length of the source sequence, in order to convert its length into the length of the target sequence, the present invention uses position encoding and a learnable weight To calculate the attention score matrix ,Finally, the decoding module input From the attention matrix and Multiply them together to get .

[0088] These operations are expressed as:

[0089]

[0090]

[0091]

[0092]

[0093]

[0094] in Represents the weight matrix of the CTC prediction layer. In this way, the target-side context information can be incorporated into the decoder input of the fine-tuning stage, thereby mapping its representation space from the source to the target, alleviating the modality difference problem.

[0095] 10) Input to the downstream decoding module: The downstream decoding module generates the translation result by simultaneously obtaining the input of the decoding module and the feature vector output of the encoding module; specifically:

[0096] The downstream decoding module obtains the input of the decoding module at the same time And the feature vector output by the encoding module is used to generate the translation result.

[0097] The decoding module here adopts the non-autoregressive decoder (Curriculum-basedNon-autoregressive Decoder, CND) based on curriculum learning. Specifically, considering a total of The non-autoregressive decoder of the layer Layer, first get the hidden vector of this layer , and pass it to the output layer to obtain the prediction of this layer And the corresponding confidence . These operations are expressed as:

[0098]

[0099]

[0100]

[0101] in, Represents the weight matrix of the linear output layer. It is worth noting that the weight matrix of the prediction layer is shared among all layers of the decoder.

[0102] At the same time, only keeping The most confident elements, and then replace the remaining words with the special symbol [mask] to get the updated Finally, splicing The word embedding and , and use an additional linear layer to reduce the dimension, we can get the decoder Output of the layer :

[0103]

[0104] In this way, the lower layers of the decoder can provide the most confident predictions to the higher layers, thereby facilitating the prediction of complex words by the higher layers.

[0105] In addition, the number of words to be predicted at each layer increases linearly with the number of layers, and the top layer will predict the entire target sentence, so the decoder can predict the entire sentence from bottom to top and from easy to difficult.

[0106] 11) Calculate the loss of the downstream fine-tuning stage: In the downstream loss calculation module, use the CTC loss function to calculate the loss of the downstream encoding module , the cross entropy loss function is used to calculate the translation loss of the generated words in the downstream decoding module .

[0107] During fine-tuning experiments on downstream tasks, the model is optimized by jointly minimizing the encoder loss and the decoder loss.

[0108]

[0109] Here, α is a hyperparameter used to balance the two training losses.

[0110] 12) End: Finally, the spoken sentence obtained by the downstream cross-modal non-autoregressive translation model is output through the output module.

[0111] Our proposed method consistently outperforms all strong baselines and achieves competitive results with an autoregressive cross-modal translation model equipped with an autoregressive pre-trained language model.

[0112] The overall method flow is as follows:

[0113] Step 1: Sample pre-training training samples and target permutations.

[0114] Sample a data pair from the corpus used for pre-training and sample the target sentence.

[0115] Step 2: Encode the source text into a latent vector.

[0116] The text encoding module using the pre-trained model will Encoded as a latent vector.

[0117] Step 3: Target-side text decoding.

[0118] First, the copy mechanism is used to obtain the input of the non-autoregressive decoder and input it into the text decoding module. The text decoding module generates the translation result through the feature vector of the encoding module.

[0119] Step 4: Pre-training phase loss calculation.

[0120] The translation loss of the non-autoregressive decoding module is calculated using cross entropy loss during the pre-training stage.

[0121] Step 5: Initialize the parameters of the downstream cross-modal non-autoregressive translation model.

[0122] The trained non-autoregressive pre-trained language model is used to initialize the encoder and decoder parameters of the downstream cross-modal non-autoregressive translation model.

[0123] Step 6: Source-side feature extraction.

[0124] The feature extraction module is used to extract features from each frame of source data to obtain a feature sequence containing source information.

[0125] Step 7: Source-side feature encoding.

[0126] The source feature sequence is sent to the downstream encoding module to obtain the hidden representation.

[0127] Step 8: Downstream decoding module input generation.

[0128] Get the predicted probability of the top CTC layer of the encoder and multiply it with the weight matrix of the pre-trained target word embedding layer to obtain a representation containing the target context.

[0129] Step 9: The decoding module generates the translation result.

[0130] The downstream decoding module generates the translation result by simultaneously obtaining the feature vectors of the decoding module input and the encoding module output.

[0131] Step 10: Downstream cross-modal non-autoregressive translation model loss calculation.

[0132] In the downstream loss calculation module, the CTC loss function is used to calculate the loss of the downstream encoding module, and the cross entropy loss function is used to calculate the translation loss of the generated words in the downstream decoding module. At the same time, a hyperparameter is used to balance the two losses.

[0133] Step 11: Output of downstream cross-modal non-autoregressive translation results.

[0134] Finally, the spoken sentence obtained by the downstream cross-modal non-autoregressive translation model is output through the output module.

[0135] The translation system is mainly composed of an encoding module and a decoding module. The present invention is directed to an improved solution of a decoding method, which needs to be applied in combination with an overall translation system. The core of the present invention is to construct and optimize a decoding framework, while the translation system provides an application carrier for the decoding method, and such a structural design ensures the integrity of the entire system.

[0136] An embodiment of the present invention provides a cross-modal non-autoregressive translation system combined with a pre-trained language model. Figure 3 The specific instructions are as follows:

[0137] The system comprises, in sequence:

[0138] Video input unit: used for cross-modal translation data selection modeling, inputting source data into the system in the form of frames; the video input unit serves as the input end of the entire system;

[0139] CTC enhancement module: It is used to obtain the CTC enhanced decoding module input using the probability distribution of the encoder CTC layer output and the pre-trained decoding end text embedding layer. The input contains the target end sequence information. The CTC enhancement module can enhance the ability of the decoding module so that it can better generate the target sequence;

[0140] Video feature extraction module: used to extract features from each frame of source data to obtain sequence features containing source information.

[0141] Downstream encoding module: The extracted visual feature vector is sent to the downstream encoding module to obtain the hidden representation. The encoder parameters are initialized from the encoder parameters in the non-autoregressive pre-trained language model.

[0142] Downstream decoding module: After receiving the output vector of the downstream encoding module, the downstream decoding module generates the translation result through the feature vector of the encoding module. The decoder adopts a non-autoregressive decoder based on curriculum learning, which requires the bottom layer to predict simple words and the upper layer to predict more complex words based on the output of the bottom layer, so as to predict the entire sequence from bottom to top and from easy to difficult. The parameters of the decoder are initialized from the encoder parameters in the non-autoregressive pre-trained language model.

[0143] Storage unit: used to store cross-modal translation data.

[0144] Loss calculation module: used to calculate the loss of the system; in the loss calculation module, the CTC loss function is used to calculate the loss of the downstream encoding module, and the cross entropy loss function is used to calculate the translation loss of the generated words in the downstream decoding module.

[0145] Text output unit: outputs the spoken sentence output obtained through the cross-modal translation model.

[0146] The design of the cross-modal non-autoregressive translation system combined with the pre-trained language model fully considers the characteristics of video data and the advantages of the pre-trained language model. Through the above module design and function allocation, a complete translation process from video input to text output can be realized.

[0147] In summary, the present invention proposes a cross-modal non-autoregressive decoding method combined with a pre-trained language model, introduces the ideas of Random Drop and CTC, designs a non-autoregressive pre-training task and a CTC-enhanced soft copy method, and respectively solves two problems in the combination of non-autoregressive pre-trained language model and non-autoregressive cross-modal translation model. A cross-modal non-autoregressive decoding algorithm combined with a pre-trained language model provided in one of the invention contents of the present invention includes: first sampling data pairs for pre-training, and sampling a permutation of the target sentence; then using a text encoder to convert the pre-trained source text into a hidden state, and using a decoding module to predict the entire sequence from bottom to top according to the sampled permutation, and then training the non-autoregressive pre-trained language model, in the pre-training loss calculation module, using a cross entropy loss function to calculate the loss of the non-autoregressive decoding module; then using the pre-trained model to initialize the parameters of the encoder and decoder of the downstream cross-modal non-autoregressive translation model; after the parameters are initialized, the source data input module performs translation data selection modeling, and the source data is input into the model in the form of frames. ; Then, the feature extraction module at the source end extracts features from each frame of the source end data to obtain sequence features containing source end information; The extracted feature vector is sent to the downstream encoding module, and the hidden representation is obtained through the encoder, and the probability predicted by the encoder CTC layer and the pre-trained text-end word embedding are used to generate the input of the downstream decoding module; After receiving the output vector of the downstream encoding module and its input, the downstream decoding module generates the translation result through the feature vector of the encoding module; In the downstream task loss calculation module, the CTC loss function is used to calculate the loss of the downstream encoding module, and the cross entropy loss function is used to calculate the translation loss of the generated words in the downstream decoding module; Finally, the spoken sentence output obtained by the cross-modal translation model of the source end data is output through the output module. The method proposed by the present invention is always better than all strong baselines, and achieves competitive results with the cross-modal autoregressive translation model equipped with an autoregressive pre-trained language model, as shown in Table 1.

[0148] Table 1

[0149]

[0150] As shown in Table 1, the present invention (ROPre) was experimentally verified on the classic cross-modal translation task of sign language translation. The effectiveness of the present invention was evaluated on the PHOENIX-2014T German sign language data using ROUGE and BLEU4 as evaluation indicators and decoding speed. First, the running speed of the present invention is 7.92 times that of the autoregressive model (AR) combined with ARPLM as the speed baseline, and both evaluation indicators are better. Compared with the two iterative NAR models (i.e., CMLM and CMLMC), although its running speed is 2.47 times that of the baseline, it lags behind the present invention in terms of speed and translation quality. Secondly, compared with the three completely non-autoregressive (NAR) models (i.e., GLEAT, DSLP and CND), although the present invention lags behind in running speed, the generation quality is better than these baseline methods.

[0151] Table 2

[0152]

[0153] As shown in Table 2, the effectiveness of the present invention (ROPre) is also evaluated by ROUGE and BLEU4 and decoding speed on the CSL-Daily Chinese Sign Language data. First, compared with the two iterative NAR models (i.e., CMLM and CMLMC), the present invention achieved improvements of 2.33 and 1.79 BLEU4 scores, respectively, and the decoding latency was much lower on the test set of CSL-Daily. Second, compared with the three full NAR SLT models (i.e., GLEAT, DSLP, and CND), the present invention achieved an average improvement of 2.34 BLEU4 points and comparable decoding speed. In addition, the present invention is slightly better than CND equipped with a mutual learning framework. Third, compared with the AR base model, the present invention achieved a performance improvement of 0.92 BLEU4 points with an acceleration of about 8.02 times. The present invention even exceeds the AR STL model equipped with AR-PLM by 0.94 BLEU4 points, ultimately proving the effectiveness and efficiency of the present invention.

[0154] The above embodiments are only preferred embodiments of the present invention and cannot be considered to limit the scope of the present invention. All equivalent changes and improvements made within the scope of the present invention should still fall within the scope of the present invention.

Claims

1. A cross-modal non-autoregressive decoding method combined with a pre-trained language model, characterized by The following steps are involved: 1) Sample pre-training training samples and target permutations: Sample a data pair from the corpus used for pre-training and sample a permutation of the target sentence ,in, Each element in is the subscript of the target word in the target sequence; 2) Encode the source text into a latent vector: Use the text encoding module of the pre-trained model to encode the source text into a latent vector. Encoded as a latent vector; 3) Target-side text decoding: The word embedding of the pre-trained source-side text is converted into the input of the non-autoregressive decoder through a copy mechanism, and then input into the text decoding module; after receiving the output vector of the pre-trained text encoding module, the text decoding module generates the translation result through the feature vector of the encoding module; 4) Pre-training loss calculation: Use cross entropy loss to calculate the translation loss of each decoder layer, and add up the losses of all layers to get the total loss; 5) Initialize the parameters of the downstream cross-modal non-autoregressive translation model: After the pre-trained language model is trained, its parameters can be used to initialize the parameters of the encoder and decoder of the downstream cross-modal non-autoregressive translation model; 6) Source feature extraction: The feature extraction module extracts features from each frame of source data to obtain a feature sequence containing source information; 7) Source feature encoding: The source feature vector extracted in step 6) is sent to the downstream encoding module to obtain hidden representation; 8) Downstream decoding module input generation: Combine the predicted probability of the CTC layer and the weight matrix of the target word embedding layer, calculate the attention score matrix through position encoding and learnable weights, and use it to generate the input of the decoding module to integrate the target context information; 9) The decoding module generates the translation result: The downstream decoding module generates the translation result by simultaneously obtaining the input of the decoding module and the feature vector output of the encoding module; 10) Calculate the loss of the downstream cross-modal non-autoregressive translation model: In the downstream loss calculation module, the CTC loss function is used to calculate the loss of the downstream encoding module, and the cross entropy loss function is used to calculate the translation loss of the generated words in the downstream decoding module; 11) Downstream cross-modal non-autoregressive translation result output: The spoken sentence obtained by the downstream cross-modal non-autoregressive translation model is output through the output module.

2. The cross-modal non-autoregressive decoding method in combination with a pre-trained language model as claimed in claim 1, characterized in that In step 3), the specific steps of target-side text decoding are: Consider the total The first layer of the pre-trained text decoding module Layer, first get the hidden vector of this layer , and pass it to the output layer to obtain the prediction of this layer ; These operations are expressed as: in, Represents the weight matrix of the linear output layer; At the same time, only keeping The subscript belongs to Before The words in the elements are then replaced with the special symbol [mask] to get the updated ; Finally, splicing The word embedding and , and use an additional linear layer to reduce the dimension, we can get the decoder Output of the layer : in Represents the weight matrix of the linear layer; since the decoder is The output of the layer is also In this way, the high-level decoder can receive the predictions of the low-level decoder, thereby predicting more words, and the entire sentence will be gradually predicted by the decoder from bottom to top.

3. The cross-modal non-autoregressive decoding method in combination with a pre-trained language model as claimed in claim 1, characterized in that In step 4), the specific steps of the pre-training stage loss calculation are: in the loss calculation module of the pre-training stage, the translation loss of the non-autoregressive decoding module is calculated using the cross entropy loss; since the decoder performs predictions at each layer, the loss function of the pre-training stage This is the sum of the cross entropy losses of each layer of the decoder: in, Indicates the decoder Layer The probability distribution of words; Indicates the number of decoder layers; represents the number of words predicted by the i-th layer, represents the tth target word.

4. The cross-modal non-autoregressive decoding method in combination with a pre-trained language model as claimed in claim 1, characterized in that In step 8), the specific steps of generating the input of the downstream decoding module are: first, obtain the prediction probability of the top CTC layer of the encoder , and then the weight matrix of the pre-trained target word embedding layer and Multiply to get a representation containing the target context However, at this time It is still the length of the source sequence. In order to convert its length to the length of the target sequence, position encoding is used and a learnable weight To calculate the attention score matrix ,Finally, the decoding module input From the attention matrix and Multiply them together to get; These operations are expressed as: in, Represents the weight matrix of the CTC prediction layer; in this way, the target-side context information can be incorporated into the decoder input of the fine-tuning stage, thereby mapping its representation space from the source to the target, alleviating the modality difference problem.

5. The cross-modal non-autoregressive decoding method in combination with a pre-trained language model as claimed in claim 1, characterized in that In step 9), the specific steps of generating the translation result by the decoding module are as follows: the downstream decoding module generates the translation result by simultaneously acquiring the feature vectors of the input of the decoding module and the output of the encoding module; specifically, the present invention requires that each layer of the decoder performs prediction, and each layer only retains the top prediction with the highest confidence. words, and then replace the remaining words with the special symbol [mask]; Then, the predicted word embedding of this layer is concatenated with the hidden state of this layer, and an additional linear layer is used for dimensionality reduction, that is, the output of the decoder at the current layer is obtained; In this way, the lower layers of the decoder can provide the most confident predictions to the higher layers, thereby facilitating the prediction of complex words by the higher layers. In addition, the number of words to be predicted at each layer increases linearly with the number of layers, and the top layer will predict the entire target sentence, so the decoder can predict the entire sentence from bottom to top and from easy to difficult.

6. The cross-modal non-autoregressive decoding method in combination with a pre-trained language model as claimed in claim 5, characterized in that: The downstream decoding module adopts a non-autoregressive decoder CND based on curriculum learning. Specifically, considering a total of The non-autoregressive decoder of the layer Layer, first get the hidden vector of this layer , and pass it to the output layer to obtain the prediction of this layer And the corresponding confidence ; These operations are expressed as: in, Represents the weight matrix of the linear output layer; the weight matrix of the prediction layer is shared among all layers of the decoder; At the same time, only keeping The most confident elements, and then replace the remaining words with the special symbol [mask] to get the updated ; Finally, splicing The word embedding and , and use an additional linear layer to reduce the dimension, we can get the decoder Output of the layer : In this way, the lower layers of the decoder can provide the most confident predictions to the higher layers, thus facilitating the prediction of complex words by the higher layers. In addition, the number of words to be predicted at each layer increases linearly with the number of layers, and the top layer will predict the entire target sentence, so the decoder can predict the entire sentence from bottom to top and from easy to difficult.

7. The cross-modal non-autoregressive decoding method in combination with a pre-trained language model as claimed in claim 1, characterized in that In step 10), the downstream cross-modal non-autoregressive translation model loss is calculated, and during the fine-tuning experiment of the downstream task, the model is optimized by jointly minimizing the encoder loss and the decoder loss: in, represents the downstream cross-modal non-autoregressive translation model loss; Represents the translation loss of a word, represents the loss of the downstream encoding module; α is a hyperparameter used to balance the two training losses.

8. A cross-modal non-autoregressive decoding system combined with a pre-trained language model, characterized by Including in order: Target arrangement sampling module, pre-trained text encoding module, pre-trained text decoding module, pre-trained loss calculation module, source feature extraction module, downstream encoding module, CTC enhancement module, downstream decoding module, downstream task loss calculation module, output module; The target permutation sampling module is used to randomly select a source text and a corresponding target text from the pre-training corpus, and perform permutation transformation on the target text; The pre-trained text encoding module is used to convert the source text into a series of latent vector representations; The pre-trained text decoding module is used to convert the word embedding of the source text into the input of the non-autoregressive decoder through a copy mechanism, and input it into the text decoding module; the decoding module generates a translation result using the output of the encoding module; The pre-training loss calculation module is used to calculate the cross entropy loss of each decoding process, and accumulate the losses of all steps to obtain the total loss, which is used to guide model optimization; The source feature extraction module is used to extract features from source data to obtain a feature sequence containing source information, and the extracted source feature vector is sent to a downstream encoding module; The downstream encoding module is used to further encode the feature vector obtained by the source feature extraction module to obtain a higher level of abstract representation; The CTC enhancement module is used to improve the model's ability to process sequence data. It combines the prediction probability of the CTC layer and the weight matrix of the target-side word embedding layer, calculates the attention score matrix through position encoding and learnable weights, and uses it to generate the input of the decoding module to fuse the target-side context information. The downstream decoding module generates a translation result by simultaneously acquiring the input of the decoding module and the feature vector output of the encoding module. The parallel processing method is a typical non-autoregressive method. The downstream task loss calculation module uses the CTC loss function to calculate the loss of the encoding module and uses the cross entropy loss function to calculate the translation loss of the generated words in the decoding module. These losses are used to guide the optimization of the entire downstream model. The output module is used to output the translation result obtained by the downstream cross-modal non-autoregressive translation model.

Citation Information

Patent Citations

  • End-to-end sign language translation method and system

    CN113657257A

  • CTC-based non-autoregressive end-to-end speech translation method

    CN116227503A

  • Visual language translation method and system based on comparative learning and word granularity weight

    CN116484885A

  • Systems and methods for using neural codec language model for zero-shot cross-lingual text-to-speech synthesis

    WO2024178710A1