Vietnamese dependency syntax analysis method and system based on multiple annotators, and electronic equipment
By combining the advantages of the traditional dependency syntax analysis model and large language model, the multi-notator method and data enhancement technology are used to solve the problem of insufficient data of low-resource language dependency syntax analysis in Vietnamese, significantly improving the model performance and pseudo-data quality, and providing a more accurate syntax structure representation.
Patent Information
- Application Number
- CN202510161292.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-13
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2045-02-13
AI Technical Summary
Dependency syntactic analysis of low-resource languages such as Vietnamese faces the problem of insufficient data. Traditional data augmentation methods are difficult to generate high-quality and diverse data, which limits the application effect of deep learning models.
The Vietnamese dependency syntax analysis method based on multi-notchers is adopted. By downloading the relevant Vietnamese dependency syntax UD tree library and VnDT tree library, the XLM-RoBERTa model is used to fine-tune the label-free data, generate an embedded representation, and encode it through the bidirectional long and short-term memory network and KAN, and finally decode it through the maximum spanning tree algorithm to generate a dependent syntax tree. At the same time, the large language model DeepSeek is used to quadratic annotation of pseudo-data to generate high-quality training corpus.
It significantly improves the performance of the Vietnamese dependent syntax analysis model, improves the quality of pseudo-data, enhances the model's ability to capture complex syntactic features, and provides a more accurate syntactic structure representation.
Smart Images

Figure CN119990102A_ABST
Abstract
Description
Technical Field
[0001] The invention relates to a Vietnamese dependency syntax analysis method, system and electronic equipment based on multiple annotators, belonging to the technical field of natural language processing. Background Art
[0002] The cultural and economic interactions between my country and Southeast Asian countries such as Vietnam are becoming increasingly frequent and close. However, it cannot be ignored that language differences are very likely to become a stumbling block to smooth communication between the two sides. Language, as a key element of carrying thinking, is the most natural and convenient way for people to communicate ideas and convey emotions. Therefore, in-depth exploration of the languages of Southeast Asian countries has irreplaceable significance for promoting the communication process and common development between my country and neighboring countries. In view of this, it is increasingly important to carry out research on dependency syntactic analysis methods for low-resource languages in Southeast Asia, especially Vietnamese, Thai, Lao, etc. With the help of in-depth analysis of the syntactic structure hidden in the text, we can grasp the internal relationship of the language more accurately. Therefore, devoting ourselves to the study of dependency syntactic analysis methods for low-resource languages in Southeast Asia not only occupies an important position in the academic field of artificial intelligence, but also has extremely critical practical significance for enhancing friendly cooperation between my country and Southeast Asian countries and strengthening regional exchanges and interactions.
[0003] In recent years, the rapid development of deep learning technology has promoted the progress of syntactic analysis tasks. However, deep learning methods rely on large-scale annotated data, which limits their application effect in low-resource languages such as Vietnamese. In order to overcome these challenges, researchers have proposed a variety of solutions, among which data enhancement is considered to be an effective means. By expanding the training data set, data enhancement can help the model better learn syntactic knowledge and improve its performance on low-resource languages. Traditional data enhancement methods usually rely on rules or simple replacement strategies, which are difficult to generate high-quality and diverse data. With the rise of large language models, using their powerful language understanding and generation capabilities for data enhancement has become a new research direction. The present invention designs a Vietnamese dependency syntactic analysis method based on multiple annotators to improve the performance of Vietnamese dependency syntactic analysis. As one of the main languages in Southeast Asia, Vietnamese has unique grammatical structure and lexical features. It has a complex tone system, rich word form changes, and unique grammatical structures, such as postposition of verbs and frequent use of quantifiers. Research on dependency syntactic analysis for low-resource languages such as Vietnamese is still relatively lagging, mainly due to the lack of large-scale annotated data. Therefore, how to combine the advantages of traditional parsing models with large language models is the key to the problem. Summary of the invention
[0004] The present invention provides a Vietnamese dependency syntax analysis method, system, and electronic device based on multiple annotators to solve the problem of Vietnamese dependency syntax tree bank conversion. The present invention improves the performance of Vietnamese dependency syntax analysis and achieves good experimental results in Vietnamese dependency syntax tree bank conversion.
[0005] The technical solution of the present invention is: a Vietnamese dependency syntactic analysis method based on multiple annotators, and the specific steps of the method are as follows:
[0006] Step 1. Download the relevant Vietnamese dependency syntax UD treebank as training corpus, VnDT treebank as the original corpus for constructing pseudo data, and collect Vietnamese unlabeled data at the same time;
[0007] Step 2: Fine-tune the XLM-RoBERTa model using the collected Vietnamese unlabeled data to enhance the representation capability of the multilingual pre-trained model.
[0008] Step 3: After fine-tuning the XLM-RoBERTa model of the UD treebank, the relevant embedding representation, character embedding representation and word-level embedding representation are concatenated as the final input vector;
[0009] Step 4: Encode the input vector through a three-layer bidirectional long short-term memory network to obtain its contextual representation, which is used as the input of KAN.
[0010] Step 5, obtain the scores of dependency arcs through bi-affine layering, and finally decode them through the maximum spanning tree algorithm to obtain the dependency syntax tree;
[0011] Step 6. Use the dependency parsing model built by the UD treebank to parse 2,000 sentences randomly selected from the VnDT treebank to obtain pseudo data containing noise.
[0012] Step 7: Input the pseudo data into the large model DeepSeek for secondary annotation according to the pre-designed prompt template;
[0013] Step 8. Use the pseudo data annotated by the second annotation as additional training corpus and merge it with the UD treebank to retrain a new Vietnamese dependency parsing model.
[0014] As a further solution of the present invention, the specific steps of Step 1 are as follows:
[0015] Step 1.1, use the downloaded UD (Universal Dependencies) tree bank as the training corpus and the VnDT tree bank as the original corpus for constructing data.
[0016] Step 1.2: The crawled Vietnamese web pages are subjected to rule extraction, deduplication, machine annotation, and manual proofreading to form a Vietnamese text corpus to build unlabeled data;
[0017] As a further solution of the present invention, the specific steps of Step 2 are as follows:
[0018] Step 2.1. Download and select the original XLM-RoBERTa model as the base model from the huggingface website;
[0019] Step 2.2: Use the parameters in the original XLM-RoBERTa model as a starting point and fine-tune the XLM-RoBERTa model parameters using Vietnamese unlabeled data.
[0020] As a further solution of the present invention, Step 3 includes:
[0021] Step 3.1. Each word w in the input layer i The corresponding final input vector x i It consists of three parts: the first part is the character embedding representation It can capture the character-level information of words; the second part is the representation generated by the XLM-RoBERTa model It provides higher-level contextual semantic information; the third part is the word-level embedding representation It provides basic semantic information of each word.
[0022] As a further solution of the present invention, the specific steps of Step 4 are as follows:
[0023] Step 4.1. Input vector x i After encoding through a three-layer bidirectional long short-term memory network, the final context word representation h' is obtained i ;
[0024] Step 4.2, input context word representation h i After KAN processing, a low-dimensional core word representation is generated and modifiers
[0025] As a further solution of the present invention, Step 5 includes:
[0026] Step 5.1: Take each word as the vector representation of the core word and vector representation of modifiers After the double affine layer, the score S of the dependency arc is obtained;
[0027] Step 5.2: Use the maximum spanning tree algorithm to decode and find the dependency syntax tree with the highest score as the final result of syntax analysis.
[0028] As a further solution of the present invention, Step 6 includes:
[0029] Step 6.1. Input 2000 sentences randomly selected from the VnDT treebank into the Vietnamese dependency parsing model trained by the UD treebank to obtain pseudo data containing noise.
[0030] As a further solution of the present invention, Step 7 includes:
[0031] Step 7.1. Design the prompt template of the large model DeepSeek as required. The prompt template should include specific instructions and constraints on the output.
[0032] Step 7.2, input 2000 pseudo data into the large model according to the prompt template to obtain the pseudo data after secondary annotation;
[0033] The present invention also provides a Vietnamese dependency syntactic analysis system based on multiple annotators, the system comprising: a module for executing the above-mentioned Vietnamese dependency syntactic analysis method based on multiple annotators.
[0034] The present invention also provides an electronic device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the above-mentioned Vietnamese dependency syntactic analysis method based on multiple annotators when executing the program.
[0035] The beneficial effects of the present invention are:
[0036] 1. The present invention combines the advantages of traditional models and large language models, and uses large language models to perform secondary annotation on pseudo data, thereby obtaining more accurate syntactic parsing results, significantly improving the quality of pseudo data, and thus effectively improving the performance of the Vietnamese syntactic analysis model.
[0037] 2. This paper fine-tunes the pre-trained language model XLM-RoBERTa by using Vietnamese unlabeled data, improving the performance of Vietnamese dependency parsing. Replacing the traditional multi-layer perceptron with KAN helps the model better capture complex syntactic features and provide a more accurate representation of grammatical structure.
[0038] 3. The Vietnamese dependency syntactic analysis method based on multiple annotators proposed in this invention has significantly improved the accuracy on the benchmark dataset. BRIEF DESCRIPTION OF THE DRAWINGS
[0039] Figure 1 It is a flow chart of the present invention; DETAILED DESCRIPTION
[0040] Example 1: Figure 1 As shown, a Vietnamese dependency syntactic analysis method based on multiple annotators, the specific steps of the method are as follows:
[0041] a1. Crawling Vietnamese web pages from the website, through rule extraction, deduplication, machine annotation, and manual proofreading, to form a Vietnamese text corpus for building unlabeled data; downloading the Vietnamese tree bank UD from the universal dependencies dataset as training data; downloading the VnDT tree bank as the original pseudo corpus.
[0042] a2. Use the crawled Vietnamese unlabeled data to fine-tune the XLM-RoBERTa model parameters to enhance the representation ability of the pre-trained model.
[0043] 1) Download and select the original XLM-RoBERTa model as the base model from the huggingface website;
[0044] 2) Use Vietnamese unlabeled data to fine-tune the model parameters of XML-RoBERTa, and the parameter settings remain unchanged.
[0045] a3. The XLM-RoBERTa model after fine-tuning the UD treebank obtains the relevant embedding representation, character embedding representation and word-level embedding representation, which are concatenated as the final input vector;
[0046] 1) XLM-RoBERTa representation after splicing and fine-tuning Word Embedding Representation and word embedding representation Get each word w i The corresponding input vector x i ;
[0047]
[0048] a4. Encode the input vector through a three-layer bidirectional long short-term memory network to obtain the context word representation, which is used as the input of KAN;
[0049] 1) Input vector x i After encoding through a three-layer bidirectional long short-term memory network, the Vietnamese-related context word representation h′ is obtained i ;
[0050] h 0 h 1 …h n =BiLSTM(x 0 x 1 …xn ,θ BiLSTM ) (2)
[0051] 2) The core of the long short-term memory network layer is the cell state, denoted by C t The long short-term memory network deletes or adds information to the cell state through a gating mechanism. BiLSTM implements three gate calculations, namely the forget gate F t , Input Gate I t and output gate O t , used to protect and control cell state.
[0052] The forget gate is responsible for deciding how much of the cell state at the previous moment to retain in the cell state at the current moment.
[0053] F t =σ(W f ·[h t-1 ,x t ]+b f ) (3)
[0054] The input gate is responsible for deciding how much of the current input to keep in the current cell state. The sigmoid layer decides what value to update; the tanh layer creates a new cell state value vector will be added to the state. t-1 Updated to C t , the old state and F t Multiply, discard the information that needs to be discarded, and add This completes the update of the cell state.
[0055] I t =σ(W i ·[h t-1 ,x t ]+b i )
[0056]
[0057] The output gate is responsible for determining how many outputs the cell state has at the current moment, and a sigmoid layer is used to determine which part of the cell state will be output. The cell state is processed through tanh to obtain a value between -1 and 1, and it is multiplied by the output of the sigmoid gate, and finally only the part that determines the output is output.
[0058]
[0059] Where W f , W i , W c , W o 、b f、b i 、b c and b o are all weight matrices; σ is the sigmoid activation function; x t is the input at time t, which represents the word vector corresponding to the t-th position.
[0060] 5) KAN uses context words to represent h' i As input. It is the word w i As the representation vector of the core word, It is the word w i As the representation vector of the modifier. KAN not only reduces the context-dependent representation w i More importantly, it retains semantically relevant information.
[0061]
[0062] a5. Use the dual affine attention mechanism for scoring and use the maximum spanning tree for decoding to find the dependency syntactic tree with the highest score as the final result of syntactic analysis.
[0063] 1) The biaffine scoring layer is used to calculate the dependency arc score between words in a sentence, which represents the dependency arc score from the core word w j to modifier w i In this layer, the score of the dependency arc is determined by the weight matrix U 1 OK. The scoring formula for the biaffine layer is: Note that all dependency arc scores can be calculated simultaneously in a matrix form.
[0064]
[0065] 2) The loss function is used to optimize the dependency parsing model to ensure that the model can accurately predict the dependency relationship between words and their dependency labels. Assume that word w i The correct core word is w j , the dependency label is l, then the loss function of the model is defined as follows:
[0066]
[0067] a6. Use the dependency parsing model built by the UD treebank to parse sentences randomly extracted from the VnDT treebank to obtain pseudo data containing noise;
[0068] a7. Input the pseudo data into the large model DeepSeek for secondary annotation according to the pre-designed prompt template;
[0069] a8. The pseudo data annotated by the secondary annotation is used as additional training corpus and merged with the UD treebank to retrain a new Vietnamese dependency parsing model.
[0070] This paper uses the Vietnamese corpus crawled from web pages and preprocessed as unlabeled data, and selects the Vietnamese dependency treebank and Vietnamese VnDT treebank in the Universal Dependencies (UD) as experimental data. This paper mainly uses the Unlabeled Attachment Score (UAS) and the Labeled Attachment Score (LAS) as evaluation indicators of dependency syntax analysis performance. The detailed calculation methods of the two evaluation indicators are as follows:
[0071]
[0072] Among them, H refers to the total number of words with correct core words, HL refers to the total number of words with correct core words and dependency types, and T refers to the total number of words in the corpus.
[0073] In order to verify the effect of the method proposed in the present invention, two classic dependency syntactic analysis methods are selected as benchmark models. 1) Language Embedding Model: Set additional language embedding representation to improve the model's understanding and representation ability of Vietnamese. 2) Multi-task Learning Model: Treat the syntactic analysis of treebanks with inconsistent annotation specifications as different tasks, and use the idea of multi-task learning to simultaneously learn the syntactic information between heterogeneous treebanks. The corresponding model of the method of the present invention: Data Augmentation Model, which is to use the pseudo data generated by the traditional model to be annotated twice by a large model, and then use it as additional training corpus to train a new syntactic analysis model. The present invention adds the representation of XLM-RoBERTa to the input layer, which enhances the performance of the dependency syntactic analysis basic model; secondly, using KAN to replace the original MLP helps the model better capture complex contextual information, especially when processing dependency syntactic relations, it can provide a more accurate grammatical structure representation. Finally, the pseudo data after secondary annotation is merged with the UD treebank into a large training set to train the Vietnamese dependency syntactic analysis model.
[0074] The final experimental results are shown in Table 2 below. After adding pseudo data that has been re-annotated after iterative optimization of the large model to the training data, the parsing performance of the basic model has been significantly improved. This result proves the effectiveness of the method proposed in the present invention, especially after the Vietnamese pseudo data generated by the traditional model is re-annotated through the powerful reasoning ability of the large model, the quality of the pseudo data has been improved, thereby further enhancing the performance of the model. Combining the reasoning ability of the large language model with the traditional dependency syntactic analysis model can provide effective help in the dependency syntactic parsing task.
[0075] Table 2 Comparison of other methods and the method of the present invention on different models
[0076]
[0077] The specific implementation modes of the present invention are described in detail above in conjunction with the accompanying drawings, but the present invention is not limited to the above implementation modes, and various changes can be made within the knowledge scope of ordinary technicians in this field without departing from the purpose of the present invention.
Claims
1. A Vietnamese dependency parsing method based on multiple annotators, characterized by: The method comprises the following steps: Step 1. Download the relevant Vietnamese dependency syntax UD treebank as training corpus, VnDT treebank as the original corpus for constructing pseudo data, and collect Vietnamese unlabeled data at the same time; Step 2: Fine-tune the XLM-RoBERTa model using the collected Vietnamese unlabeled data to enhance the representation capability of the multilingual pre-trained model. Step 3: After fine-tuning the XLM-RoBERTa model of the UD treebank, the relevant embedding representation, character embedding representation and word-level embedding representation are concatenated as the final input vector; Step 4: Encode the input vector through a three-layer bidirectional long short-term memory network to obtain its contextual representation, which is used as the input of KAN. Step 5, obtain the scores of dependency arcs through bi-affine layering, and finally decode them through the maximum spanning tree algorithm to obtain the dependency syntax tree; Step 6: Use the dependency parsing model built by the UD treebank to parse the sentences randomly extracted from the VnDT treebank to obtain pseudo data containing noise; Step 7: Input the pseudo data into the large model DeepSeek for secondary annotation according to the pre-designed prompt template; Step 8. Use the pseudo data annotated by the second annotation as additional training corpus and merge it with the UD treebank to retrain a new Vietnamese dependency parsing model.
2. The Vietnamese dependency parsing method based on multiple annotators according to claim 1, characterized in that: The specific steps of Step 1 are as follows: Step 1.1, use the downloaded UD tree bank as training corpus, and use the VnDT tree bank as the original corpus for constructing data; Step 1.2: The crawled Vietnamese web pages are subjected to rule extraction, deduplication, machine annotation, and manual proofreading to form a Vietnamese text corpus for building unlabeled data.
3. The Vietnamese dependency parsing method based on multiple annotators according to claim 1, characterized in that: The specific steps of Step 2 are as follows: Step 2.
1. Download and select the original XLM-RoBERTa model as the base model from the huggingface website; Step 2.2: Use the parameters in the original XLM-RoBERTa model as a starting point and fine-tune the XLM-RoBERTa model parameters using Vietnamese unlabeled data.
4. The Vietnamese dependency parsing method based on multiple annotators according to claim 1, characterized in that: The Step 3 includes: Step 3.
1. Each word w in the input layer i The corresponding final input vector x i It consists of three parts: the first part is the character embedding representation It can capture the character-level information of words; the second part is the representation generated by the XLM-RoBERTa model It provides higher-level contextual semantic information; the third part is the word-level embedding representation It provides basic semantic information of each word.
5. The Vietnamese dependency parsing method based on multiple annotators according to claim 1, characterized in that: The specific steps of Step 4 are as follows: Step 4.
1. Input vector x i After encoding through a three-layer bidirectional long short-term memory network, the final context word representation h' is obtained i ; Step 4.2, input context word representation h i After KAN processing, a low-dimensional core word representation is generated and modifiers 6. The Vietnamese dependency parsing method based on multiple annotators according to claim 1, characterized in that: The Step 5 includes: Step 5.1: Take each word as the vector representation of the core word and vector representation of modifiers After the double affine layer, the score S of the dependency arc is obtained; Step 5.2: Use the maximum spanning tree algorithm to decode and find the dependency syntax tree with the highest score as the final result of syntax analysis.
7. The Vietnamese dependency parsing method based on multiple annotators according to claim 1, characterized in that: The Step 6 includes: Step 6.
1. Input several sentences randomly selected from the VnDT tree bank into the Vietnamese dependency parsing model trained by the UD tree bank to obtain pseudo data containing noise.
8. The Vietnamese dependency parsing method based on multiple annotators according to claim 1, characterized in that: The Step 7 includes: Step 7.
1. Design the prompt template of the large model DeepSeek as required. The prompt template should include specific instructions and constraints on the output. Step 7.2: Input several pieces of pseudo data into the large model according to the prompt template to obtain the pseudo data after secondary annotation.
9. A Vietnamese dependency parsing system based on multiple annotators, characterized in that: The system comprises: a module for executing the Vietnamese dependency parsing method based on multiple annotators according to any one of claims 1 to 8.
10. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the program, the method for Vietnamese dependency parsing based on multiple annotators as claimed in any one of claims 1 to 8 is implemented.
Citation Information
Patent Citations
Dependence mapping method and system
CN102760121A
Method for detecting Vietnamese dependency treebank error based on treebank conversion
CN108280060A
Vietnamese dependency syntax tree bank conversion method based on double syntactic constraints
CN118278387A