A Chinese-Vietnamese Cross-Language Dependency Parsing Method Incorporating a Language Adversarial Network
By integrating language adversarial networks and Chinese syntactic information, using Vietnamese label-free data to fine-tune the XLM-RoBERTa model, the performance degradation of resource-scarce languages in cross-language dependency syntactic analysis was solved, and the improvement of Vietnamese dependency syntactic analysis was achieved, especially in the case of limited resources, which significantly improved the accuracy rate.
Patent Information
- Application Number
- CN202310142778.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-02-21
- Publication Date
- 2025-07-22
- Estimated Expiration
- 2043-02-21
AI Technical Summary
In cross-language dependency syntax analysis, it is difficult for the existing technology to effectively utilize the resource-rich language training model to improve the dependency syntax analysis performance of languages with fewer resources, especially when the test data language changes, the performance drops sharply.
By integrating into the language adversarial network, the XLM-RoBERTa model is fine-tuned using Vietnamese label-free data to obtain Vietnamese language-related information, and combine it with Chinese syntactic information, encode it through BiLSTM and multi-layer perceptron. Finally, the dependency syntax tree is decoded through the maximum spanning tree algorithm to enhance the Vietnamese dependency syntax analysis ability.
In the case of small corpus size, the accuracy of Vietnamese dependent syntax analysis is significantly improved, the performance of Chinese-Jiangsu cross-language dependent syntax analysis is enhanced, common knowledge between languages is captured and interference with language-specific characteristics is reduced.
Smart Images

Figure CN116167357B_ABST
Abstract
Description
[0001] Technical language
[0002] The present invention relates to a Chinese-Vietnamese cross-lingual dependency syntax analysis method incorporating a language adversarial network, belonging to the technical field of natural language processing. Background Art
[0003] As the carrier of thinking, language is the most natural and convenient tool for humans to communicate ideas and express emotions. Deeply understanding the languages of Southeast Asian countries plays a crucial role in promoting the exchanges and development between our country and neighboring countries. Improving the quality of Vietnamese dependency syntax analysis not only helps to promote the development of Vietnamese natural language processing, such as word segmentation, part-of-speech tagging, and named entity recognition, etc.; but also provides support for upper-level artificial intelligence tasks, such as Chinese-Vietnamese machine translation, intelligent question answering for Vietnamese, and sentiment analysis, etc.
[0004] With the development of deep neural networks, supervised dependency syntax analysis has achieved remarkable performance improvement. However, training a high-quality syntax analysis model often depends on large-scale in-language labeled data. Once the language of the test data changes, the performance of dependency syntax analysis will drop sharply, which is the cross-lingual dependency syntax analysis problem. Currently, cross-lingual dependency syntax analysis is still a key challenge in the research of dependency syntax analysis. The goal of cross-lingual dependency syntax analysis is to use existing data resources to improve the dependency syntax analysis performance of texts different from the training data, especially to use the language with rich resources to train the model and use the trained model to process those languages with less resources. The present invention designs a language embedding model integrating a language adversarial network to improve the ability of Chinese-Vietnamese cross-lingual dependency syntax analysis. First, by researching the grammars of Chinese and Vietnamese, we found that there are many similarities between Chinese and Vietnamese, such as: the names of sentence components are the same, namely the subject predicate object adverbial attributive complement The positions of the subject, predicate, and object in a sentence are the same, all following the subject-predicate-object structure. However, there are also certain differences between these two languages. For example, in Chinese sentences, the attributive is placed before the noun; while in Vietnamese, on the contrary, the noun is in the front and the attributive is behind. In addition, the emergence of new words and phrases and the change of expression structures may all lead to changes in the distribution of syntactic trees, which is the language difference mentioned in this article. And there are also common points between the Vietnamese and Chinese grammar structures. The key to cross-lingual dependency syntax analysis is how to model the differences and commonalities between different languages. Summary of the Invention
[0005] The present invention provides a Chinese-Vietnamese cross-lingual dependency parsing method incorporating a language adversarial network to solve the problem of Chinese-Vietnamese cross-lingual dependency parsing, and the present invention has achieved good experimental results in Chinese-Vietnamese cross-lingual dependency parsing.
[0006] The technical solution of the present invention is: a Chinese-Vietnamese cross-lingual dependency parsing method based on incorporating a language adversarial network, and the specific steps of the method are as follows:
[0007] Step1. Collection and preprocessing of unannotated Vietnamese data; the specific steps of Step1 are as follows:
[0008] Step1.1. The crawled Vietnamese web pages are processed through rule extraction, deduplication, machine annotation, and manual proofreading to form a Vietnamese text corpus, which is used as the construction of unlabeled data. The labeled data of Chinese and Vietnamese are downloaded from the general dataset UD as laboratory data.
[0009] Step2. Use the unannotated Vietnamese data to fine-tune the XLM-RoBERTa model by minimizing the language model loss, so as to obtain information related to the Vietnamese language.
[0010] Step3. The output of the last four layers of the hidden layer of the fine-tuned XLM-RoBERTa model is used as an additional input to the dependency parsing model based on language embeddings.
[0011] Step4. Concatenate the XLM-RoBERTa representation obtained in the previous step and the language embedding representation, and use it as the input vector of the language embedding model. Then, it is encoded through BiLSTM, followed by dimensionality reduction through a multi-layer perceptron (MLPs), and finally, the scores of the dependency arcs are obtained through the bi-affine mechanism (Biaffines). Finally, the dependency syntax tree is obtained through the maximum spanning tree algorithm for decoding.
[0012] Step5. Incorporate the language adversarial network into the language embedding model enhanced by RoBERTa, and attempt to use the syntactic information of Chinese to improve the performance of Vietnamese dependency parsing.
[0013] As a further solution of the present invention, Step2 includes the following:
[0014] Download the XLM-RoBERTa model from the website, train the XLM-RoBERTa model using the unlabeled Vietnamese data, and fine-tune some parameters of the model using the language model loss to improve the training performance of XLM-RoBERTa. The specific steps are as follows: Step2.1. Download the XLM-RoBERTa model from the website and select the original XLM-RoBERTa model as the base model.
[0015] Step 2.2: Use the parameters in the original XLM-RoBERTa model as the starting point, and fine-tune the model parameters of XML-RoBERTa using Vietnamese unlabeled data. Here, the loss of the next sentence and the orthogonal loss are removed, and only the language model loss is used to adjust the XLM-RoBERTa model parameters.
[0016] As a further solution of the present invention, the specific steps of Step 3 are as follows:
[0017] Step 3.1: Obtain the outputs of the last four layers of the hidden layer in the fine-tuned XLM-RoBERTa.
[0018] Step 3.2: Calculate the average value of the outputs of the last four layers of the XLM-RoBERTa output, and then use linear mapping to transform the high-dimensional output into a low-dimensional output vector.
[0019] Step 3.3: Vectorize the labeled Chinese and Vietnamese data to obtain word embedding representations. Randomly initialize the language embedding representation.
[0020] As a further solution of the present invention, Step 4 specifically includes the following:
[0021] Step 4.1: Obtain the input vector x of each word by concatenating the fine-tuned XLM-RoBERTa representation. Word embedding representation. And language embedding representation. Obtain the input vector x of each word. i ;
[0022] Step 4.2: Use x i As the input of the private bidirectional long short-term memory network BiLSTM, and obtain the private context-related word representation through the encoding of the private bidirectional long short-term memory network BiLSTM.
[0023] Step 4.3 And the context-related representation obtained by the shared BiLSTM. After concatenation, perform dimensionality reduction through a multi-layer perceptron to obtain the vector representation of each word as the central word representation. And the vector representation of the modifier.
[0024] Step 4.4: Use the vector representation of each word as the central word representation. And the vector representation of the modifier. Pass through a bi-affine layer to obtain the score S of the dependency arc.
[0025] Step 4.5 Decode the score S of the dependency arc using the maximum spanning tree algorithm to obtain the final dependency syntax tree.
[0026] As a further solution of the present invention, the Step 5 specifically includes the following:
[0027] Step 5.1 Obtain the input vector corresponding to each word by splicing the representation of the fine-tuned XLM-RoBERTa and the word embedding representation
[0028] Step 5.2 Use as the input of the additional shared BiLSTM to obtain the common shared context-related word representation
[0029] Step 5.3 Splice the shared context-related word representation and the private context-related word representation into the MLPs for dimensionality reduction;
[0030] Step 5.4 Add a gradient reversal layer and a language classifier to the shared BiLSTM. By minimizing the language classification loss, let the shared BiLSTM capture more commonalities between Chinese and Vietnamese, thereby enhancing the model's ability to perform Vietnamese dependency syntax analysis, so that the common context word representation enters the gradient reversal layer and the language classifier;
[0031] Step 5.5 To further enhance the extraction of common features between Chinese and Vietnamese, an orthogonal constraint is added here to reduce the interference of language-specific features on language common features and increase the difference between specific language representations and target language invariant representations.
[0032] The beneficial effects of the present invention are:
[0033] 1. The present invention splices Chinese and Vietnamese to construct Chinese-Vietnamese bilingual word vectors and uses XLM-RoBERTa for pre-training, enhancing the cross-language transplantation performance of Chinese and Vietnamese dependency syntax analysis;
[0034] 2. The present invention uses the method of orthogonal constraint and an adversarial learning model, and obtains a more pure and effective language invariant representation.
[0035] 3. The method for constructing dependency syntax analysis proposed by the present invention significantly improves the accuracy in few-shot cross-language dependency syntax analysis compared with the case of a small corpus size. Description of the Drawings
[0036] Figure 1This is the flowchart in the present invention; Specific implementation manner
[0037] Example 1: As Figure 1 shown, a Chinese-Vietnamese cross-lingual dependency parsing method incorporating a language adversarial network, and the specific steps of the method are as follows:
[0038] Step1. Collection and preprocessing of unannotated Vietnamese data; The crawled Vietnamese web pages are processed through rule extraction, deduplication, machine annotation, and manual proofreading to form a Vietnamese text corpus, which is used as the construction of unlabeled data. Download Chinese and Vietnamese labeled data from the Universal Dependencies (UD) dataset as laboratory data.
[0039] Step2. Use the unannotated Vietnamese data to fine-tune the XLM-RoBERTa model by minimizing the language model loss, so as to obtain information related to the Vietnamese language; Download the XLM-RoBERTa model from the website, and use the unlabeled Vietnamese data to train the XLM-RoBERTa model, and use the language model loss to fine-tune some parameters of the model, so as to improve the training performance of XLM-RoBERTa. The specific steps are as follows:
[0040] Step2.1. Download the XLM-RoBERTa model from the website, and select the original XLM-RoBERTa model as the base model;
[0041] Step2.2. Use the parameters in the original XLM-RoBERTa model as the starting point, and fine-tune the model parameters of XML-RoBERTa using the unlabeled Vietnamese data; Here, the loss of the next sentence and the orthogonal loss are removed, and only the language model loss is used to adjust the XLM-RoBERTa model parameters.
[0042] Step3. Use the output of the last four layers of the hidden layer of the fine-tuned XLM-RoBERTa model as an additional input to the dependency parsing model based on language embeddings; The specific steps of Step3 are as follows:
[0043] Step3.1 Obtain the output of the last four layers of the hidden layer in the fine-tuned XLM-RoBERTa;
[0044] Step3.2 Calculate the average value of the last four layers of the XLM-RoBERTa output, and then use linear mapping to transform the high-dimensional output into a low-dimensional output vector
[0045] Step3.3 Vectorize the labeled Chinese and Vietnamese data to obtain word embedding representations Randomly initialize the language embedding representation
[0046] Step 4. Concatenate the XLM-RoBERTa representation obtained in the previous step and the language embedding representation, and use it as the input vector of the language embedding model. Then, encode it through a BiLSTM, followed by dimensionality reduction through a multi-layer perceptron (MLP). Finally, obtain the scores of the dependency arcs through a bi-affine mechanism (Biaffines), and ultimately decode to obtain the dependency syntax tree through the maximum spanning tree algorithm. The specific steps of Step 4 are as follows:
[0047] Step 4.1 Obtain the input vector x for each word by concatenating the fine-tuned XLM-RoBERTa representation the word embedding representation and the language embedding representation ; i
[0048] Step 4.2 Use x i as the input of the private bidirectional long short-term memory network (BiLSTM), and obtain the private context-related word representation through encoding by the private BiLSTM
[0049] Step 4.3 and the context-related representation obtained by the shared BiLSTM After concatenation, perform dimensionality reduction through a multi-layer perceptron to obtain the vector representation of each word as the head word representation and the modifier
[0050] Step 4.4 Use the vector representation of each word as the head word representation and the modifier Pass through a bi-affine layer to obtain the scores S of the dependency arcs;
[0051] Step 4.5 Decode the scores S of the dependency arcs using the maximum spanning tree algorithm to obtain the final dependency syntax tree.
[0052] Step 5. Incorporate the language adversarial network into the language embedding model enhanced by RoBERTa, and attempt to improve the performance of Vietnamese dependency syntax analysis by leveraging the syntactic information of Chinese. Step 5 specifically includes the following:
[0053] Step 5.1 Obtain the input vector corresponding to each word by concatenating the representation of the fine-tuned XLM-RoBERTa and the word embedding representation ;
[0054] Step 5.2 Use as the input of an additional shared BiLSTM to obtain the common shared context-related word representation
[0055] Step5.3 Shared context-related word representations And private context-related word representations Concatenate them into MLPs for dimensionality reduction;
[0056] Step5.4 Add a gradient reversal layer and a language classifier to the shared BiLSTM. By minimizing the language classification loss, let the shared BiLSTM capture more commonalities between Chinese and Vietnamese, thereby enhancing the model's ability to perform Vietnamese dependency parsing, so that the common context word representations Enter the gradient reversal layer and the language classifier;
[0057] Step5.5 To further enhance the extraction of common features between Chinese and Vietnamese, an orthogonality constraint is added here to reduce the interference of language-specific features on language common features and increase the difference between language-specific representations and target language invariant representations.
[0058] Example 2: As Figure 1 shown, a Chinese-Vietnamese cross-language dependency parsing method incorporating a language adversarial network, the specific steps of the method are as follows:
[0059] a1. Crawl corpus from various Vietnamese news websites. The crawled Vietnamese web pages are processed through rule extraction, deduplication, machine annotation, and manual proofreading to form a Vietnamese text corpus, which is used as an unlabeled Vietnamese corpus. Obtain a large amount of Chinese and Vietnamese labeled corpus from UD. Then preprocess the collected corpus, such as word segmentation, deduplication, and marking.
[0060] a2. Use the unlabeled Vietnamese data to fine-tune the XLM-RoBERTa model by minimizing the language model loss, so as to obtain information related to the Vietnamese language.
[0061] 1) Download the original XLM-RoBERTa model.
[0062] 2) Use the parameters in the original XLM-RoBERTa model as a starting point, and use the unlabeled Vietnamese data to fine-tune the original XLM-RoBERTa model. Here, the loss of the next sentence and the orthogonality loss are removed, and only the language model loss is used to adjust the parameters of the XLM-RoBERTa model.
[0063] a3. Calculate the average value of the outputs of the last four layers of the hidden layer in the fine-tuned XLM-RoBERTa to obtain the XLM-RoBERTa model representation; then use a linear mapping to transform the high-dimensional output into a low-dimensional output vector Vectorize the tagged Chinese and Vietnamese data to obtain word embedding representations Randomly initialize the language embedding representations
[0064] a4. The word embedding representations, language embedding representations, and XLM-RoBERTa representations are concatenated as the input vectors for the shared and private long short-term memory networks.
[0065] 1) Concatenate the word embedding representations and language embedding representations as the input to the private bidirectional long short-term memory network to obtain language-invariant representations.
[0066] 2) Concatenate the word embedding representations, language embedding representations, and XLM-RoBERTa representations as the input to the shared bidirectional long short-term memory network to obtain language-specific representations.
[0067] a5. The context representation words output by the shared long short-term memory network and the context representations output by the private long short-term memory network are concatenated and then enter a multi-layer perceptron, while the hidden state output by the private bidirectional long short-term memory network separately enters a gradient reversal layer.
[0068] 1) Due to the lack of tagged data in the target language, the parameters of the private bidirectional long short-term memory network corresponding to the target language may not be fully optimized. Therefore, when the input word is from the target language, this paper uses the output of the hybrid bidirectional long short-term memory network as the final language-specific representation. Otherwise, this paper uses the output of the original private bidirectional long short-term memory network as the language-specific representation of this language. The definition of the hybrid target language word representation is as follows:
[0069]
[0070] where and are the outputs of the source language private bidirectional long short-term memory network and the target language private bidirectional long short-term memory network respectively; γ is the weight used to balance and . In this work, the present invention selects a suitable γ value from 0 to 1 through experiments and finally sets γ to 0.1.
[0071] 2) Both the shared and private long short-term memory networks first use a three-layer bidirectional long short-term memory network to serially encode the input sentences. The long short-term memory network mainly consists of four parts, namely the forget gate F i , the input gate I i , the output gate O i , and the memory cell state C i . The forget gate is used to determine how much of the previous cell state is discarded by the current cell. The input gate Ii It is used to determine how much of the input at the current moment will be updated into the current cell. Is the cell state of the current input. The memory cell state C at the current moment i Is composed of the content of the cell state C at the previous moment i-1 And the content of the cell state of the current input i To determine. Output gate O i Is used to control how much output there is in the cell state at the current moment. h i Is the final output of the long short-term memory network and is passed to the next layer of the neural network. The values of the three gates range between 0 and 1, where 0 means complete discard and 1 means complete retention.
[0072]
[0073] Among them, W f , W i , W c , W o , b f , b i , b c And b o Are all weight matrices; σ is the sigmoid activation function; x i Is the input at the i-th moment, which here represents the word vector corresponding to the i-th position.
[0074] 3) Use the language-specific representation obtained from the private long short-term memory network as the final context word representation, and input the final context word representation into a multi-layer perceptron for dimensionality reduction, and finally use it for dependency parsing.
[0075] The multi-layer perceptron uses the explicit word representation h i As the input, and uses two independent multi-layer perceptrons to obtain two low-dimensional vector representations for each position 0 ≤ i ≤ n. Among them Is the representation vector of the word W i As the core word, Is the representation vector of the word W i As the modifier word. The multi-layer perceptron not only reduces the dimension of the context-related representation h i , but more importantly, it retains the semantically related information. The definition of the multi-layer perceptron is shown as follows.
[0076]
[0077]
[0078] 4) Input the language-invariant representation obtained from the shared long short-term memory network into the gradient reversal layer, and perform forward and backward propagation to prevent the language classifier from accurately predicting the language type corresponding to each word, thereby encouraging the shared bidirectional long short-term memory network to obtain more language-invariant representations. The definitions of the forward and backward propagation of the gradient reversal layer are as follows:
[0079] GRL λ (h i )
[0080]
[0081] a6. After the language-invariant representation obtained from the shared long short-term memory network passes through the gradient reversal layer, it enters the language classifier, and finally, the entire adversarial network is optimized by minimizing the standard cross-entropy loss.
[0082] 1) The language classification layer uses a multi-layer perceptron to calculate the scores of the language distribution and uses the softmax operation to obtain the probability values of the language type distribution corresponding to each word.
[0083] z i = softmax(W2ReLU(W1h1 + b1)+b2) (5)
[0084] where θ d = {W1, W2, b1, b2} represents the parameters of the language classifier. The entire adversarial network is trained by minimizing the standard cross-entropy loss, and the specific calculation process is as follows:
[0085]
[0086] where represents the actual language distribution vector. Here, only the element corresponding to the correct language of word W i is 1 and the rest are 0; z i,j represents the predicted probability value that word W i belongs to language j; n represents the number of words in a sentence; m represents the number of source language types.
[0087] a7. When the language-specific representation obtained from the private long short-term memory network passes through the multi-layer perceptron, an orthogonal constraint is used to ensure that different bidirectional long short-term memory networks separate the language-invariant representation and the language-specific representation.
[0088] 1) Although the present invention can separate the language-invariant representation and the language-specific representation through different bidirectional long short-term memory networks, it is difficult to ensure that there is no mutual interference between the two. To solve this problem, the present invention uses orthogonal constraints to make the language-specific representation and the language-invariant representation mutually exclusive, and the orthogonal loss function is calculated as follows:
[0089]
[0090] Where and are the outputs of the shared bidirectional long short-term memory network and the private bidirectional long short-term memory network, respectively.
[0091] a8. The language-specific representation enters the double-affine mechanism for scoring after passing through a multi-layer perceptron, and the maximum spanning tree is used for decoding.
[0092] 1) Calculate all dependency arc scores through double-affine operations. Where score(i←j) is the dependency arc score from the head word W i to the modifier word W i , and the matrix U 1 is the double-affine parameter. It should be noted that all dependency arc scores can be calculated simultaneously in a matrix form.
[0093]
[0094] The calculation method of the dependency relation label score is similar to that of the dependency arc, as follows:
[0095]
[0096] Where, U 2 and U 3 are weight matrices, and b is a bias vector. It should be noted that the dependency relation label prediction and the dependency arc prediction are calculated using independent multi-layer perceptrons and double-affine scoring layers.
[0097] After the scores are calculated, the best dependency syntax tree can be found through the maximum spanning tree algorithm for decoding. The calculation process is as follows:
[0098]
[0099] The double-affine syntax analysis model defines a local cross-entropy loss for each position i. Assume that the word w j is the correct head word of the word w i , and the dependency relation type between them is l. The definition of the related dependency syntax analysis loss function is as follows:
[0100]
[0101] a9. The total loss function of the adversarial-based language embedding model is shown in the following formula:
[0102]
[0103] Here, the present invention selects the tagged data of Chinese and Vietnamese in the Universal Dependencies (UD) as the experimental data of the present invention, and at the same time uses the processed Vietnamese corpus crawled from the Internet as the unlabeled data of the present invention. The commonly used UAS and LAS are used as the evaluation indicators of the dependency parsing model. The specific calculation method is shown as follows:
[0104]
[0105]
[0106] To verify the effectiveness of the method proposed by the present invention, the present invention selects three classic few-shot cross-lingual dependency parsing methods as the baseline models. 1) The basic model of directly concatenating corpora (Concat): directly concatenating the training corpora of Chinese and Vietnamese as a large training set to train the basic model. 2) The language embedding model (Language Embedding): setting additional language embedding representations to distinguish differences in different languages. 3) The feature augmentation model (Feature Augmentation): naturally separating the language-specific representation and the language-shared representation by using shared and private encoders. The present invention first uses XLM-RoBERTa for pre-training to enhance the performance of the cross-lingual dependency parsing basic model; secondly, by integrating a language adversarial network, it further improves the ability of the cross-lingual embedding model to obtain common knowledge between different languages; finally, setting orthogonal constraints to expand the difference between the language-invariant representation and the language-common representation.
[0107] The final experimental results are shown in Table 1 below. First, it can be seen that even though the ability of the basic model is enhanced by using XLM-RoBERTa, the method proposed by the present invention still achieves the best experimental results, which further verifies the effectiveness of the method; secondly, it is found that the performance of the language embedding model can be effectively improved by adding a language adversarial network, proving that the language adversarial network helps the model capture more common knowledge between languages, so as to help resource-scarce languages perform dependency parsing with the help of the rich dependency syntactic information of language resources. Finally, the performance of the method proposed by the present invention can be further improved by adding orthogonal constraints, proving that language embedding helps to obtain language-specific representations, while language adversarial helps to obtain language-common representations, and there is a certain complementarity between these two representations.
[0108] Table 1 Comparison between Other Methods and the Method of the Present Invention on Different Models
[0109]
[0110] The specific embodiments of the present invention have been described in detail above in conjunction with the accompanying drawings. However, the present invention is not limited to the above embodiments, and various changes can be made without departing from the spirit of the present invention within the scope of knowledge possessed by those of ordinary skill in the art.
Claims
1. A Chinese-Vietnamese cross-lingual dependency parsing method incorporating a language adversarial network, characterized in that: The specific steps of the method are as follows: Step1. Collection and preprocessing of unannotated Vietnamese data; Step2. Using the unannotated Vietnamese data to fine-tune the XLM-RoBERTa model by minimizing the language model loss, so as to obtain information related to the Vietnamese language; Step3. Taking the outputs of the last four layers of the hidden layer of the fine-tuned XLM-RoBERTa model as additional inputs for the dependency parsing model based on language embeddings; Step4. Concatenating the XLM-RoBERTa representation and the language embedding representation obtained in the previous step as the input vector of the language embedding model, then encoding through BiLSTM, then reducing the dimension through multi-layer perceptrons (MLPs), and finally obtaining the scores of dependency arcs through the biaffine mechanism Biaffines, and finally decoding through the maximum spanning tree algorithm to obtain the dependency parsing tree; Step5. Incorporating the language adversarial network into the language embedding model enhanced by RoBERTa, and attempting to use the syntactic information of Chinese to improve the performance of Vietnamese dependency parsing.
2. The Chinese-Vietnamese cross-lingual dependency parsing method integrating a language confrontation network according to claim 1, characterized in that: The specific steps of Step1 are as follows: Step1.
1. The crawled Vietnamese web pages are processed through rule extraction, deduplication, machine annotation, and manual proofreading to form a Vietnamese text corpus as the unlabeled data for construction; download the labeled data of Chinese and Vietnamese from the general dataset UD as the laboratory data.
3. The Chinese-Vietnamese cross-lingual dependency parsing method integrating a language adversarial network according to claim 1, characterized in that: The specific steps of Step2 are as follows: Step2.
1. Download the XLM-RoBERTa model from the website and select the original XLM-RoBERTa model as the base model; Step2.
2. Use the parameters in the original XLM-RoBERTa model as the starting point and fine-tune the model parameters of XML-RoBERTa using the unlabeled Vietnamese data; here, the loss of the next sentence and the orthogonal loss are removed, and only the language model loss is used to adjust the XLM-RoBERTa model parameters.
4. The Chinese-Vietnamese cross-lingual dependency parsing method incorporating a language adversarial network according to claim 1, characterized in that: The specific steps of Step3 are as follows: Step3.
1. Obtain the outputs of the last four layers of the hidden layer in the fine-tuned XLM-RoBERTa; Step 3.2 Calculate and obtain the average value of the last four layers of the BERT output, and then use linear mapping to transform the high-dimensional output into a low-dimensional output vector Step 3.3 Vectorize the labeled Chinese and Vietnamese data to obtain word embedding representations Randomly initialize the language embedding representation 5. The Chinese-Vietnamese cross-lingual dependency parsing method incorporating a language adversarial network according to claim 1, characterized in that: Step4 specifically includes the following: Step 4.1 Obtain the input vector x for each word by concatenating the fine-tuned XLM-RoBERTa representations Word embedding representation and the language embedding representation i ; Step4.2 Take x i as the input of the private bidirectional long short-term memory network BiLSTM, and obtain the private context-related word representation through the encoding of the private bidirectional long short-term memory network BiLSTM Step4.3 and the context-related representation obtained by the shared bidirectional long short-term memory network BiLSTM After concatenation, dimensionality reduction is performed through a multi-layer perceptron to obtain the representation of each word as the central word and the vector representation of the modifier Step4.4 Represent each word as the central word and the vector representation of the modifier Through the bi-affine layer, obtain the score S of the dependency arc; Step4.
5. Decode the scores S of the dependency arcs using the maximum spanning tree algorithm to obtain the final dependency parsing tree.
6. The Chinese-Vietnamese cross-lingual dependency parsing method incorporating a language adversarial network according to claim 1, characterized in that: Step5 specifically includes the following: Step 5.1 Obtain the input vector corresponding to each word by splicing the representations of the fine-tuned XLM-RoBERTa and the word embedding representations Step 5.2 Use as the input of the additional shared bidirectional long short-term memory network BiLSTM to obtain the common and shared context-related word representations Step5.3 Shared context-related word representations and the private context-related word representations are concatenated into MLPs for dimensionality reduction; Step 5.4 Add a gradient reversal layer and a language classifier to the shared bidirectional long short-term memory network (BiLSTM). By minimizing the language classification loss, the shared BiLSTM can capture more commonalities between Chinese and Vietnamese, thereby enhancing the model's ability to perform Vietnamese dependency parsing and enabling the common context word representations to enter the gradient reversal layer and the language classifier; Step5.
5. In order to further enhance the extraction of common features between Chinese and Vietnamese, the orthogonal loss is added here to reduce the interference of language-specific features on language common features.
Citation Information
Patent Citations
Language model fine tuning method for low-resource adhesive language text classification
CN113032559A
Vietnamese dependency syntactic analysis method based on parameter migration
CN114757167A