A translation method applied to low-resource machine translation

By performing multi-scale sampling and fusion attention weights in the multi-head self-attention model, the problem of insufficient representation learning and over-parameterization of the Transformer model under low resource conditions is solved, and the performance and stability of machine translation are improved.

CN114580441BActive Publication Date: 2025-07-01XIAONIU FANYI
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210172910.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-02-24
Publication Date
2025-07-01
Estimated Expiration
2042-02-24

AI Technical Summary

Technical Problem

The Transformer model faces the problems of insufficient representation learning and over-parameterization of the model under low resource conditions, resulting in instability in training and inefficiency. The top-level representation tends to be similar, making it difficult to improve translation performance.

Method used

Based on the multi-head self-attention model, multi-scale sampling is performed on the source language input text, data representations of multiple scales are constructed, and these data representations are used as Query and Key to perform dot product operations, multiple attention weight matrices are obtained, and then fused, and finally dot product operations are performed with the original input data to obtain the fused output vector.

Benefits of technology

Without adding additional parameters, the representation ability and translation performance of the model are improved, and the problem of over-parameterization and insufficient representation learning of the model is alleviated. It is suitable for machine translation tasks under low resource conditions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114580441B_ABST
    Figure CN114580441B_ABST
Patent Text Reader

Abstract

The present invention discloses a translation method applied to low-resource machine translation. The steps are as follows: preprocess the source language input text and perform word embedding operation to obtain the source language input text x ∈ R<supgt;L×D< / supgt; of the model; in the machine translation model, sample the processed source language input text based on the multi-head self-attention model to obtain data representations at multiple scales; respectively use the sampled data representations as Query and Key to perform dot product operations to obtain n attention weight matrices; fuse the n attention weight matrices obtained in the above process, and then perform a dot product operation on the fused attention weight matrix and the source language input text x to obtain a fused output vector, which is used as the decoder input to implement the subsequent translation process. The method of the present invention has good generality, can improve the model performance without increasing additional parameters, and can also enhance the representation ability of the model in the case of scarce data.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention discloses a method for representation learning of a model, specifically a translation method applied to low-resource machine translation. Background Art

[0002] With the continuous improvement of computer computing power, in the context of massive Internet data, computational models based on neural networks have been pushed onto the historical stage, gradually showing huge advantages in various fields of artificial intelligence, replacing traditional methods, and becoming a research hotspot in academia and industry. Currently, they have been widely applied to people's daily lives, such as speech recognition, machine translation, language models, etc. At the same time, language, as a bridge for human communication, has always been widely concerned by the academic community. Therefore, how to enable people speaking different languages to communicate with each other is a very challenging but extremely meaningful problem.

[0003] Neural Machine Translation (NMT) is a simple new architecture compared to traditional statistical machine translation for translating text from one language to another. Neural machine translation has now achieved remarkable results, significantly improving the fluency and accuracy of machine translation. However, the accuracy of translation depends on the representation ability of the model. Experiments show that deeper networks can have stronger learning abilities, can better fit the training data, and achieve better effects. On the other hand, in the context of the big data era, the acquisition of massive data is no longer a bottleneck for model training, but a larger amount of data requires a more robust network to obtain more excellent performance improvement. Therefore, deeper and more expressive networks have become the focus of people's attention. However, as the network capacity increases, the model also faces problems such as difficulty in training, difficulty in convergence, and over-parameterization of the model. Therefore, how to effectively train neural translation models has also become a popular research topic for machine translation researchers.

[0004] It can be seen from the development of machine translation that the models are tending to become more and more complex. Since shallow models, such as those studied by Mhaskar et al. from the perspective of network learning, have shown that deeper and more complex convolutional neural networks have stronger learning capabilities; Telgarsky et al. introduced the potential advantages of deep networks; Eldan and Shamir also verified on feedforward neural networks that increasing the network's representation ability can improve the effectiveness of the model. On the other hand, in the industrial field, models with simple structures cannot fully utilize the advantages of massive data, resulting in no obvious performance improvement when increasing the data volume, but instead increasing the training burden. These phenomena also indicate that only networks with more robust representation learning can better meet people's needs. However, neural network-based machine translation models, such as the Transformer model, face problems such as insufficient representation learning and over-parameterization of the model: for example, simply increasing the number of layers of the translation model does not lead to performance improvement, but instead brings problems such as unstable training and low training efficiency; at the same time, recent research has found that when the network model is continuously deepened, its top-layer representations tend to be similar, indicating that the top-layer network has not been effectively trained.

[0005] Multi-scale information refers to sampling the input data or representations at different granularities to form different data views, and different features of the data can generally be observed at different granularities. Multi-scale feature fusion is a method of fusing different-scale information to obtain a more robust data representation. Thanks to the continuous representation of images, multi-scale methods have been widely applied and studied in the field of computer vision. For example, in object detection tasks, multiple feature maps are usually constructed by upsampling or downsampling methods, and each feature map corresponds to information with different receptive fields. The model needs to comprehensively identify and detect the objects in the image based on the information with different receptive fields.

[0006] However, in the field of natural language processing, multi-scale methods are still rarely studied. Especially for machine translation models, there are many challenges in constructing multi-scale models. Firstly, the data learned by the model in the machine translation task is a symbol string of varying lengths composed of symbols, which makes the construction of scale information vulnerable to the influence of the sentence length distribution of the dataset; on the other hand, the relationship between different symbols is discrete, and its semantic information fusion mechanism needs further study, which makes the fusion of multi-scale information difficult.

[0007] The current Transformer model faces problems such as insufficient representation learning and over-parameterization of the model: for example, simply increasing the number of layers of the translation model does not lead to performance improvement, but instead brings problems such as unstable training and low training efficiency; at the same time, recent research has found that when the network model is continuously deepened, its top-layer representations tend to be similar, indicating that the top-layer network has not been effectively trained. Summary of the Invention

[0008] In view of the disadvantages of insufficient representation learning and over-parameterization faced by the Transformer model in the prior art, the technical problem to be solved by the present invention is to provide a translation method for low-resource machine translation, through which a more robust data representation can be obtained for the machine translation model.

[0009] To solve the above technical problems, the technical solution adopted by the present invention is as follows:

[0010] The present invention provides a translation method for low-resource machine translation, including the following steps:

[0011] 1) Preprocess the source language input text and perform word embedding operation to obtain the source language input text x ∈ R of the model L×D ;

[0012] 2) In the machine translation model, sample the processed source language input text x ∈ R L×D on the basis of the multi-head self-attention model to obtain data representations at multiple scales

[0013] 3) Respectively use the sampled data representations as Query and Key to perform dot product operations to obtain n attention weight matrices;

[0014] 4) Fuse the n attention weight matrices obtained in the above process, and then perform dot product operation on the fused attention weight matrix and the source language input text x to obtain the fused output vector, which is used as the decoder input to implement the subsequent translation process.

[0015] In step 2), in the machine translation model, sample the processed source language input text x ∈ R L×D on the basis of the multi-head self-attention model to obtain source language input texts at multiple scales Specifically:

[0016] Q, K, V = x

[0017]

[0018]

[0019]

[0020] where x is the input source language input text, Q, K, and V are the query value, key value, and source language input text respectively, and the data dimensions of Q, K, and V are W i Q , is the i-th learnable parameter matrix with a matrix dimension of are respectively the i-th query, key-value, and source language input text after linear mapping, sample is the sampling strategy, are respectively the query value and key-value after sampling.

[0021] Step 3) Use the sampled data representations as Query and Key respectively to perform dot product operations to obtain n attention weight matrices, specifically:

[0022]

[0023] where, is the multi-scale model input, is the word embedding dimension of the model, softmax(·) represents the softmax activation function, and its specific expression is:

[0024]

[0025] where y is the function output, x is the function input, x m 、x n are respectively the m-th and n-th dimensions of the input data, K is the input vector dimension, and e is the natural constant.

[0026] Step 4) Fuse the n attention weight matrices obtained in the above process, and perform dot product operations on the fused attention weight matrix and the original input data x to obtain the final output, specifically:

[0027]

[0028] where Concat(·) is the concatenation method, which concatenates two feature vectors along a certain dimension, W o is the learnable parameter matrix, atten_weight i is the attention weight matrix, output is the model output, is the i-th source language input text after linear mapping.

[0029] The present invention has the following beneficial effects and advantages:

[0030] 1. The present invention provides a translation method applied to low-resource machine translation and acts on the multi-head attention method, which can be widely applied to various sequence generation tasks and has good versatility; and is orthogonal to other methods for stable model training, such as some initialization methods of deep models and structural adjustments of neural networks, etc.

[0031] 2. Since the method of the present invention adopts a parameter-free vector decomposition method, it can improve the model performance without increasing the number of additional parameters, and can also enhance the representation ability of the model in the case of scarce data.

[0032] 3. The method of the present invention can alleviate problems such as over-parameterization of the model and insufficient representation learning, and is significantly helpful for low-resource machine translation. BRIEF DESCRIPTION OF THE DRAWINGS

[0033] Figure 1 It is a schematic diagram of the Transformer model structure involved in the translation method of the present invention applied to low-resource machine translation;

[0034] Figure 2 It is an example diagram of multi-scale feature fusion involved in the method of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0035] The present invention will be further described below with reference to the accompanying drawings of the specification.

[0036] As Figure 2 shown, the present invention provides a translation method applied to low-resource machine translation, including the following steps:

[0037] 1) Preprocess the source language input text and perform word embedding operation to obtain the source language input text x ∈ R of the model L×D ;

[0038] 2) In the machine translation model, sample the processed source language input text x ∈ R L×D on the basis of the multi-head self-attention model, so as to obtain data representations of multiple scales

[0039] 3) Respectively use the sampled data representations as Query and Key to perform dot product operations to obtain n attention weight matrices;

[0040] 4) Fuse the n attention weight matrices obtained in the above process, and then perform dot product operation on the fused attention weight matrix and the source language input text x to obtain a fused output vector, which is used as the decoder input to implement the subsequent translation process.

[0041] On the basis of multi-head self-attention, the present invention further enriches the representation ability of the model. Specifically, the present invention further proposes a multi-scale attention weight fusion model on the basis of the self-attention model, that is, sample the input data x ∈ R L×D on the basis of the multi-head self-attention model, so as to obtain data representations of multiple scales Perform dot product operations on these data representations as Query and Key respectively to obtain n attention weight matrices, then fuse these n attention weight matrices, and finally perform dot product operations on the fused attention weight matrix and the original input data x to obtain the final output.

[0042] The specific process in step 2) is as follows:

[0043] Q, K, V = x

[0044]

[0045]

[0046]

[0047] where x is the input source language input text, Q, K, and V are the query value, key value, and source language input text respectively, and the data dimensions of Q, K, and V are W i Q , is the i-th learnable parameter matrix, and its matrix dimension is are the query, key value, and source language input text after the i-th linear mapping respectively, sample is the sampling strategy, are the query value and key value after sampling respectively.

[0048] x ∈ R L×D is the input training data, and sample is the sampling strategy. Its core idea is to construct a multi-scale model on the hidden layer dimension. The multi-scale structure is reflected in that for the same data, this method can construct multiple different data representations according to its hidden layer dimension, which can avoid sampling operations on the input data, so that the model can adapt to data inputs of various lengths and reduce the impact of data changes on the model structure.

[0049] In step 3), perform dot product operations on the sampled data representations as Query and Key respectively to obtain n attention weight matrices, specifically:

[0050]

[0051] where, is the input of the multi-scale model, is the word embedding dimension of the model, softmax(·) represents the softmax activation function, and its specific expression is:

[0052]

[0053] where y is the function output, x is the function input, and x m , x n are the m-th and n-th dimensions of the input data respectively, K is the dimension of the input vector, and e is the natural constant.

[0054] Step 4) Fuse the n attention weight matrices obtained from the above process, and perform a dot product operation on the fused attention weight matrix and the original input data x to obtain the final output, specifically:

[0055]

[0056] where Concat(·) is the concatenation method, which concatenates two feature vectors along a certain dimension, W o is a learnable parameter matrix, is the attention weight matrix, output is the model output, is the i-th source language input text after linear mapping.

[0057] The present invention focuses on improving the representation ability of the model on low-resource datasets. That is, in this embodiment, text datasets in three directions of IWSLT14 are selected for experiments. For each direction of the dataset, after preprocessing such as tokenizing the source language, data cleaning, lowercasing letters, and byte pair encoding, the number of byte pair encoding merges is set to 10,000. The specific language directions and data scales are shown in Table 1.

[0058] Table 1

[0059] Number of De->En corpora Number of En->De corpora Number of Fa->en corpora 160000 160000 89000

[0060] The present invention can achieve an improvement in the BLEU score compared to the basic Transformer model (as Figure 1 shown) on the three datasets of IWSLT14. The experimental results are shown in Table 2.

[0061] Table 2

[0062] Model De->En En->De Fa->En Baseline model 35.96 29.5 24 The present invention 36.32 29.57 24.24

[0063] Using the multi-scale feature method in the translation method applied to low-resource machine translation can improve the performance of the model, and this method can effectively alleviate problems such as insufficient model representation learning and over-parameterization of the model.

Claims

1. A translation method applied to low-resource machine translation, characterized in that It includes the following steps: 1) Preprocess the source language input text and perform word embedding operations to obtain the source language input text x ∈ R of the model L×D ; 2) In the machine translation model, based on the multi-head self-attention model, sample the processed source language input text x ∈ R L×D to obtain data representations at multiple scales 3) Taking the sampled data representations as Query and Key respectively to perform dot product operations to obtain n attention weight matrices; 4) Fusing the n attention weight matrices obtained in the above process, and then performing dot product operations on the fused attention weight matrix and the source language input text x to obtain a fused output vector, which is used as the decoder input to implement the subsequent translation process; Step 2) In the machine translation model, based on the multi-head self-attention model, sample the processed source language input text x ∈ R L×D to obtain source language input texts at multiple scales Specifically: Where x is the input source language input text, and Q, K, and V are the query value, key value, and source language input text respectively. The data dimensions of Q, K, and V are is the i-th learnable parameter matrix, and its matrix dimension is are the i-th query, key value, and source language input text after linear mapping respectively. sample is the sampling strategy, are the query value and key value after sampling respectively.

2. The translation method applied to low-resource machine translation according to claim 1, characterized in that: Step 3) Taking the sampled data representations as Query and Key respectively to perform dot product operations to obtain n attention weight matrices, specifically: Among them, is the input of the multi-scale model, is the word embedding dimension of the model, and softmax(·) represents the softmax activation function, and its specific expression is: where y is the function output, x is the function input, x m , x n are the m-th and n-th dimensions of the input data respectively, K is the input vector dimension, and e is the natural constant.

3. The translation method applied to low-resource machine translation according to claim 1, characterized in that: Step 4) Fusing the n attention weight matrices obtained in the above process, and performing dot product operations on the fused attention weight matrix and the original input data x to obtain the final output, specifically: where Concat(·) is a concatenation method that concatenates two feature vectors along a certain dimension, and W o is a learnable parameter matrix, and atten_weight i is an attention weight matrix, output is the model output, is the i-th source language input text after linear mapping.

Citation Information

Patent Citations

  • Non-autoregressive neural machine translation method based on decoder input enhancement

    CN113468895A

  • Low-resource neural machine translation method based on bidirectional dependency self-attention mechanism

    CN113901845A