Vietnamese dependency syntax tree bank construction method and system based on step-by-step optimization training

Through step-by-step optimization training and multi-stage iteration, using multi-language pre-training models and prompt word technology, a high-quality Vietnamese-dependent syntax tree library was built, which solved the problem of scarcity of Vietnamese-dependent syntax data, improved analysis accuracy and efficiency, and provided a new direction for NLP tasks in low-resource languages.

CN120297263APending Publication Date: 2025-07-11KUNMING UNIV OF SCI & TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510357557.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-25
Publication Date
2025-07-11

AI Technical Summary

Technical Problem

In the natural language processing of Vietnamese, especially in the field of dependent syntactic analysis, it faces the problems of scarcity of resources and limited model performance. Traditional methods are difficult to effectively deal with the unique grammatical characteristics and language structure of Vietnamese in cross-language transfer learning.

Method used

Using a step-by-step optimization training method, by collecting publicly dependent syntax tree library and high-quality Vietnamese corpus, the multi-language pre-trained language model XLM-ROBERTa and the dual affine dependent syntax analysis model BiAffine Parser are trained. Combining block prompt words, main clause recognition prompt words and thinking chain strategies, a high-quality dependent syntax tree library is gradually optimized to generate.

Benefits of technology

It significantly improves the accuracy and efficiency of Vietnamese-dependent syntax analysis, reduces dependence on manual annotation data, improves the processing ability of long sentences and complex syntactic structures, and provides innovative solutions for natural language processing tasks in low-resource languages.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120297263A_ABST
    Figure CN120297263A_ABST
Patent Text Reader

Abstract

The invention relates to a Vietnamese dependency syntax tree bank construction method and system based on step-by-step optimization training, and belongs to the field of natural language processing. Comprising the following steps: designing block cue words and master and slave sentence recognition cue words to respectively guide a large model to carry out block segmentation and master and slave sentence recognition tasks on a high-quality Vietnamese corpus to obtain a block data set and a master and slave sentence data set; thinking chain cues are designed, a high-quality Vietnamese corpus data set, a pseudo-dependency syntactic tree of a traditional model, block data and master and slave sentence data are input into a large model, and a high-quality dependency syntactic tree bank data set is obtained through iterative optimization and fused with a public dependency syntactic tree bank; loading to a multi-language pre-training language model and a double-affine dependency syntax analysis model for re-training to obtain a model with better syntax analysis performance for constructing a Vietnamese dependency syntax tree bank. According to the method, the problem of Vietnamese dependency syntax data scarcity is effectively relieved, and the performance of Vietnamese dependency syntax analysis is remarkably improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a method and system for constructing a Vietnamese dependency syntax tree bank based on step - by - step optimization training, belonging to the technical field of natural language processing. Background Art

[0002] Language, as an important bridge for communication, plays an irreplaceable role. Vietnamese is not only the official language of Vietnam but also an important tool for deeply understanding Vietnamese social culture and promoting multilateral cooperation. However, research on natural language processing (NLP) of Vietnamese, especially in the field of dependency syntax analysis, faces major challenges such as scarce resources and limited model performance.

[0003] Dependency syntax analysis is one of the core technologies in the NLP field, aiming to analyze the grammatical structure of sentences and reveal the dependency relationships between words. Dependency syntax analysis technology is of great significance for advanced tasks such as machine translation, sentiment analysis, and information extraction. In Chinese - Vietnamese machine translation, the accuracy of Vietnamese dependency syntax analysis directly affects the quality of translation results. However, compared with high - resource languages such as English and Chinese, the Vietnamese corpus is small in scale and the annotated data is limited. Traditional methods perform poorly in cross - language transfer learning and are difficult to effectively handle the unique grammatical features and language structures of Vietnamese.

[0004] To address these problems, the present invention proposes a method for constructing a Vietnamese dependency syntax tree bank based on step - by - step optimization training. Summary of the Invention

[0005] The technical problem solved by the present invention is: The present invention provides a method and system for constructing a Vietnamese dependency syntax tree bank based on step - by - step optimization training, effectively alleviating the problem of scarce Vietnamese dependency syntax data and effectively improving the accuracy and efficiency of Vietnamese dependency syntax analysis.

[0006] The technical solution of the present invention is: A method for constructing a Vietnamese dependency syntax tree bank based on step - by - step optimization training, the method comprising:

[0007] Step 1, collect publicly available dependency syntax tree banks and high - quality Vietnamese corpora as experimental data, and pre - process the high - quality Vietnamese corpora;

[0008] Step 2, input the publicly available dependency syntax tree bank into the multilingual pre - trained language model XLM - ROBERTa and the bi - affine dependency syntax analysis model BiAffine Parser for training, and save the bi - affine dependency syntax analysis model with the best performance;

[0009] Step 3, use the bi - affine dependency syntax analysis model with the best performance to parse the dependency syntax tree to be processed, and obtain a traditional model pseudo - dependency syntax tree;

[0010] Step 4. Design chunking prompt words and main-clause recognition prompt words to respectively guide the large model to perform chunking and main-clause recognition tasks on high-quality Vietnamese corpora, and save the obtained chunked dataset and main-clause dataset;

[0011] Step 5. Design a chain-of-thought prompt word for guiding the large model to construct a dependency syntax tree. Utilize the high-quality Vietnamese corpus dataset, the pseudo-dependency syntax tree of the traditional model, the chunked data, and the main-clause data to help the large model deeply understand the syntactic structure information, thereby generating a correct dependency syntax tree; finally, obtain a high-quality dependency syntax tree library dataset through iterative optimization;

[0012] Step 6. After fusing the constructed high-quality dependency syntax tree library dataset with the public dependency syntax tree library, load it into the multilingual pre-trained language model XLM-ROBERTa and the bi-affine dependency parsing model BiAffine Parser for retraining to obtain a model with better syntactic parsing performance for constructing the Vietnamese dependency syntax tree library.

[0013] Furthermore, the said Step 1 includes:

[0014] Step 1.1. Download a labeled Vietnamese dataset from the general data as the public dependency syntax tree library dataset;

[0015] Step 1.2. Download an unlabeled Vietnamese dataset from the general dataset as the high-quality Vietnamese corpus;

[0016] Step 1.3. Use a tokenization tool to tokenize the high-quality Vietnamese corpus, and then process it into a format that conforms to the general dependency syntax tree standard according to the tokenization results, and use the tokenization results that conform to the general dependency syntax tree standard as the dependency syntax tree to be processed.

[0017] Furthermore, the said Step 2 includes:

[0018] Step 2.1. Input the public dependency syntax tree library into the multilingual pre-trained language model XLM-ROBERTa to obtain the corresponding word vectors, and then input the word vectors into the bi-affine dependency parsing BiAffine Parser model for training;

[0019] Step 2.2. After the training is completed, save the bi-affine dependency parsing model with the best performance.

[0020] Furthermore, the said Step 3 includes:

[0021] Step 3.1. Use the bi-affine dependency parsing model with the best performance to parse the dependency syntax tree to be processed;

[0022] Step 3.2. Save the dependency syntactic tree dataset obtained by parsing and name it as the pseudo-dependency syntactic tree of the traditional model.

[0023] Further, the said Step 4 includes:

[0024] Step 4.1. Design the chunking prompt words for guiding the large model to perform chunking.

[0025] Step 4.2. Input the chunking prompt words and the high-quality Vietnamese corpus into the large model and save the obtained chunking dataset.

[0026] Step 4.3. Design the main-clause and subordinate-clause recognition prompt words for guiding the large model to perform the main-clause and subordinate-clause recognition task.

[0027] Step 4.4. Input the main-clause and subordinate-clause recognition prompt words and the high-quality Vietnamese corpus into the large model and save the obtained main-clause and subordinate-clause dataset.

[0028] Further, the said Step 5 includes:

[0029] Step 5.1. Design the chain-of-thought prompt words for guiding the large model to perform the task of constructing the syntactic tree.

[0030] Step 5.2. Use the chain-of-thought strategy to input the chain-of-thought prompt words for the task of constructing the dependency syntactic tree, the high-quality Vietnamese corpus dataset, the chunking dataset, the main-clause and subordinate-clause dataset, and the pseudo-dependency syntactic tree of the traditional model into the large model to help the large model perform in-depth parsing, make more efficient use of all information, and generate a more accurate syntactic tree through repeated iteration and optimization.

[0031] Step 5.3. Save the output result of the large model and name it as the high-quality dependency syntactic tree dataset.

[0032] Further, in the said Step 5.2, the chain-of-thought strategy is as follows:

[0033] For each internal node N in the parse tree, first extract the child nodes and their ranges:

[0034] {C1, C2, …, C M} = Children(N)

[0035] {s1, s2, …, s M} = TextSpan(C1, C2, …, C M )

[0036] where C1, C2, …, C M are the child nodes of node N, and s1, s2, …, s MFor the text ranges corresponding to the child nodes C1, C2, …, C M That is, the content in the chunks, the child nodes of Children(.), and the text ranges of TextSpan(,);

[0037] Filter valid chunks for the extracted child nodes and their ranges;

[0038] If s1, s2, …, s M All belong to the chunk set FilteredChunks, then concatenate the text ranges of s1, s2, …, s M To generate a new range s, and construct a prompt Prompt(N) for the node N, then repeat the above steps to generate prompts for each node N, and the final set of prompts is Prompts:

[0039]

[0040] s = Concat(s1, s2, …, s M )

[0041] Prompt(N) = {s1, s2, …, s M are combined into a new text span s}

[0042]

[0043] Among them, Concat() represents concatenation, ∪ N∈InternalNodes Prompt(N) represents generating a prompt for each node N and combining them into the final set of prompts.

[0044] Furthermore, the said Step 6 includes:

[0045] Step 6.1: Integrate the constructed high-quality dependency syntactic tree dataset with the publicly available dependency syntactic tree bank;

[0046] Step 6.2: Load the integrated data into the multilingual pre-trained language model XLM-ROBERTa and the bi-affine dependency syntactic analysis model BiAffine Parser for retraining, and save the obtained new model.

[0047] The present invention also provides a Vietnamese dependency syntactic tree bank construction system based on step-by-step optimization training, and the system includes: a module for executing the above-mentioned Vietnamese dependency syntactic tree bank construction method based on step-by-step optimization training.

[0048] The present invention also provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, it implements the method for constructing a Vietnamese dependency syntax tree bank based on step-by-step optimization training.

[0049] The present invention also provides a non-transitory computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it implements the method for constructing a Vietnamese dependency syntax tree bank based on step-by-step optimization training.

[0050] The method of the present invention loads the data of the Universal Dependencies of the dependency syntax tree bank into the multilingual pre-trained language model XLM-RoBERTa and the bi-affine dependency syntax analysis model (BiAffine Parser) for training. After the training is completed, the bi-affine dependency syntax analysis model with the best performance is saved. Then, the trained model is used to perform dependency syntax parsing on the collected high-quality unannotated Vietnamese corpus to generate a dependency syntax tree bank of the traditional model. Next, the chain of thought strategy is used to guide the large model to perform corpus chunking and main-subordinate clause recognition tasks, so as to learn the main components and structural boundary information of the sentence, and improve the deep syntax parsing ability of the large model. Secondly, the pseudo-dependency syntax tree bank of the traditional model is combined to constrain the preference of the large model for generating syntax labels, and a more correct dependency syntax tree is simulated and generated. Finally, a high-quality Vietnamese dependency syntax tree bank is obtained through multiple iterations. Finally, the high-quality dependency syntax tree data set is mixed with the publicly available dependency syntax tree bank data set for re-optimizing and training the BiAffine Parser model to generate a Vietnamese dependency syntax analysis model with significantly improved performance. This method effectively improves the accuracy and stability of Vietnamese dependency syntax analysis through step-by-step optimization and multi-stage iteration, and at the same time provides a feasible solution for natural language processing tasks of resource-scarce languages.

[0051] The beneficial effects of the present invention are as follows:

[0052] 1. By designing a multi-step optimization training process, the present invention makes full use of the publicly available dependency syntax tree bank and high-quality Vietnamese corpus, and combines the chain of thought technology to achieve the efficient construction of the Vietnamese dependency syntax tree, significantly reducing the dependence on manually annotated data and lowering the development cost and time;

[0053] 2. By using the large model and prompt design technology, the present invention gradually optimizes the chunking and main-subordinate clause recognition tasks for Vietnamese, and the generated dependency syntax tree has higher accuracy and robustness, especially showing excellent performance in processing long sentences and complex syntactic structures;

[0054] 3. The method for constructing a Vietnamese dependency syntax treebank based on step-by-step optimization training proposed by the present invention not only effectively solves the problem of scarce Vietnamese resources, but also provides innovative ideas for the dependency syntax analysis of other low-resource languages, and has important promotion value;

[0055] 4. The present invention improves the performance of Vietnamese dependency syntax analysis by constructing a high-quality Vietnamese dependency syntax treebank, provides technical support for Sino-Vietnamese language and cultural exchanges and regional economic cooperation, and at the same time promotes the development of multi-language natural language processing technology;

[0056] 5. By combining the publicly available dependency syntax treebank and high-quality unannotated corpus, utilizing the powerful representation ability of multi-language pre-trained models, and using the chain-of-thought strategy to design prompting words to help the large model deeply understand syntactic structure information, the present invention effectively improves the accuracy and efficiency of Vietnamese dependency syntax analysis; the method of the present invention not only alleviates the problem of data scarcity in Vietnamese dependency syntax analysis, but also opens up a new direction for the research of natural language processing of low-resource languages;

[0057] 6. By integrating multi-language pre-trained models, prompting word optimization, chain-of-thought strategy, and large model technology, the present invention proposes an efficient and innovative step-by-step optimization strategy; the method of the present invention not only provides a new solution for Vietnamese dependency syntax analysis, but also accumulates valuable experience for the research of NLP tasks of other low-resource languages; in the future, the method of the present invention is expected to be further promoted to more fields of cross-language natural language processing, promote the progress of multi-language technology, and facilitate in-depth exchanges between different languages and cultures. BRIEF DESCRIPTION OF THE DRAWINGS

[0058] Figure 1 is an example of a chunking prompting word in an embodiment of the present invention;

[0059] Figure 2 is an example of a main-clause and subordinate-clause recognition prompting word in an embodiment of the present invention;

[0060] Figure 3 is an example of a prompting word for constructing a dependency syntax tree in an embodiment of the present invention;

[0061] Figure 4 is a flowchart of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0062] Example 1: As Figures 1-4 shown, a method for constructing a Vietnamese dependency syntax treebank based on step-by-step optimization training, the method includes:

[0063] Step 1. Collect the publicly available dependency syntax treebank and high-quality Vietnamese corpus as experimental data, and preprocess the high-quality Vietnamese corpus;

[0064] Further, the Step1 includes:

[0065] Step 1.1. Download the Vietnamese dataset with tags from Universal Dependencies as the public dependency syntax treebank dataset;

[0066] Step1.2. Download the untagged Vietnamese dataset from Asian Language Treebank as the high-quality Vietnamese corpus;

[0067] Step1.3. Use the underthesea word segmentation tool to segment the high-quality Vietnamese corpus, and then process it into a format that conforms to the standard of the general dependency syntax tree according to the segmentation results. The segmentation results that conform to the standard of the general dependency syntax tree are used as the dependency syntax tree to be processed.

[0068] Step 2. Input the public dependency syntax treebank into the multilingual pre-trained language model XLM-ROBERTa and the bi-affine dependency parsing model (BiAffine Parser) for training, and save the bi-affine dependency parsing model with the best performance;

[0069] Further, the Step2 includes:

[0070] Step 2.1. Input the public dependency syntax treebank into the multilingual pre-trained language model XLM-ROBERTa to obtain the corresponding word vectors, and then input the word vectors into the bi-affine dependency parsing BiAffine Parser model for training;

[0071] Step 2.2. After training, save the bi-affine dependency parsing model with the best performance.

[0072] Step 3. Use the bi-affine dependency parsing model with the best performance to parse the dependency syntax tree to be processed, and obtain the pseudo-dependency syntax tree of the traditional model;

[0073] Further, the Step3 includes:

[0074] Step 3.1. Use the bi-affine dependency parsing model with the best performance to parse the dependency syntax tree to be processed;

[0075] Step 3.2. Save the parsed dependency syntax tree dataset and name it the pseudo-dependency syntax tree of the traditional model.

[0076] Step 4. Design chunking prompts and main-subordinate clause recognition prompts to respectively guide the large model to perform chunking and main-subordinate clause recognition tasks on high-quality Vietnamese corpora, and save the obtained chunked dataset and main-subordinate clause dataset;

[0077] Further, the said Step 4 includes:

[0078] Step 4.1. Design chunking prompts to guide the large model to perform chunking; the chunking prompts are as Figure 1 shown:

[0079] Step 4.2. Input the chunking prompts and high-quality Vietnamese corpora into the large model, and save the obtained chunked dataset;

[0080] Step 4.3. Design main-subordinate clause recognition prompts to guide the training of the large model to perform main-subordinate clause recognition tasks; the main-subordinate clause recognition prompts are as Figure 2 shown:

[0081] Step 4.4. Input the main-subordinate clause recognition prompts and high-quality Vietnamese corpora into the large model, and save the obtained main-subordinate clause dataset.

[0082] Step 5. Design thought chain prompts for guiding the large model to perform the task of constructing dependency syntax trees. Utilize the high-quality Vietnamese corpus dataset, pseudo-dependency syntax trees of traditional models, chunked data, and main-subordinate clause data to help the large model deeply understand syntactic structure information, thereby generating correct dependency syntax trees; finally, obtain a high-quality dependency syntax tree library dataset through iterative optimization;

[0083] Further, the said Step 5 includes:

[0084] Step 5.1. Design thought chain prompts for guiding the large model to perform the task of constructing syntax trees; the prompts are as Figure 3 shown:

[0085] Step 5.2. Use the thought chain strategy to input the thought chain prompts for constructing dependency syntax trees, high-quality Vietnamese corpus dataset, chunked dataset, main-subordinate clause dataset, and pseudo-dependency syntax trees of traditional models into the large model to help the large model perform deep parsing, make more efficient use of all information, and generate more accurate syntax trees through repeated iterative optimization;

[0086] Step 5.3. Save the output results of the large model and name it as a high-quality dependency syntax tree dataset.

[0087] Further, in the said Step 5.2, the thought chain strategy is as follows:

[0088] For each internal node N in the parse tree, first extract the child nodes and their ranges:

[0089] {C1, C2, …, C M} = Children(N)

[0090] {s1, s2, …, s M} = TextSpan(C1, C2, …, C M )

[0091] where C1, C2, …, C M are the child nodes of node N, and s1, s2, …, s M are the text ranges corresponding to the child nodes C1, C2, …, C M , that is, the content in the chunks, Children(.) for child nodes, and TextSpan(,) for text ranges;

[0092] For example: The parse tree of the sentence "small birds fly" is like (S(NP(JJ Small)(NN Birds))(VP fly)). The analysis shows that:

[0093] The root node N = S: Children(S) = {NP, VP}

[0094] The child nodes of the noun phrase NP: Children(NP) = {JJ, NN}

[0095] The child nodes of the verb phrase VP: Children(VP) = {V}

[0096] The content of the child nodes:

[0097] TextSpan(JJ) = "Small"

[0098] TextSpan(NN) = "Birds"

[0099] TextSpan(V) = "fly"

[0100] Filter the valid chunks from the extracted child nodes and their ranges;

[0101] If s1, s2, …, s M all belong to the chunk set FilteredChunks, then concatenate the text ranges of s1, s2, …, s M to generate a new range s, and construct a prompt Prompt(N) for node N. Then repeat the above steps to generate prompts for each node N. The final set of prompts is Prompts:

[0102]

[0103] s = Concat(s1, s2, …, s M )

[0104] Prompt(N) = {s1, s2, …, s M are combined into a new text span s}

[0105]

[0106] Among them, Concat( ) represents concatenation, ∪ N∈InternalNodes Prompt(N) represents generating prompts for each node N and combining them into the final prompt set.

[0107] Step 6. After fusing the constructed high-quality dependency syntactic tree bank dataset with the public dependency syntactic tree bank, load it into the multilingual pre-trained language model XLM-ROBERTa and the bi-affine dependency parsing model BiAffine Parser for retraining to obtain a model with better syntactic parsing performance for constructing the Vietnamese dependency syntactic tree bank.

[0108] Furthermore, the said Step 6 includes:

[0109] Step 6.1. Fuse the constructed high-quality dependency syntactic tree dataset with the public dependency syntactic tree bank;

[0110] Step 6.2. Load the fused data into the multilingual pre-trained language model XLM-ROBERTa and the bi-affine dependency parsing model BiAffine Parser for retraining, and save the obtained new model.

[0111] The present invention also provides a Vietnamese dependency syntactic tree bank construction system based on step-by-step optimization training, and the said system includes:

[0112] A data collection and preprocessing module, which is used to collect the public dependency syntactic tree bank and high-quality Vietnamese corpus as experimental data, and preprocess the high-quality Vietnamese corpus;

[0113] A model training module, which is used to input the public dependency syntactic tree bank into the multilingual pre-trained language model XLM-ROBERTa and the bi-affine dependency parsing model for training, and save the best-performing bi-affine dependency parsing model;

[0114] A traditional model pseudo-dependency syntactic tree generation module, which is used to parse the dependency syntactic tree to be processed using the best-performing bi-affine dependency parsing model to obtain a traditional model pseudo-dependency syntactic tree;

[0115] The chunked dataset and main-clause dataset generation module is used to design chunking prompts and main-clause recognition prompts to guide the large model to perform chunking and main-clause recognition tasks on high-quality Vietnamese corpora, and save the obtained chunked dataset and main-clause dataset;

[0116] The high-quality dependency syntactic treebank dataset generation module is used to design a chain of thought prompt for guiding the large model to perform the task of constructing a dependency syntactic tree. Using the high-quality Vietnamese corpus dataset, the pseudo-dependency syntactic tree of the traditional model, the chunked data, and the main-clause data, it helps the large model deeply understand the syntactic structure information, so as to generate the correct dependency syntactic tree; finally, a high-quality dependency syntactic treebank dataset is obtained through iterative optimization;

[0117] The Vietnamese dependency syntactic treebank construction module is used to fuse the constructed high-quality dependency syntactic treebank dataset with the public dependency syntactic treebank, and then load it into the multilingual pre-trained language model XLM-ROBERTa and the bi-affine dependency parsing model BiAffine Parser for retraining, so as to obtain a model with better syntactic parsing performance for constructing the Vietnamese dependency syntactic treebank.

[0118] The present invention also provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, it implements the method for constructing a Vietnamese dependency syntactic treebank based on step-by-step optimization training.

[0119] The present invention also provides a non-transitory computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it implements the method for constructing a Vietnamese dependency syntactic treebank based on step-by-step optimization training.

[0120] Based on the public dependency syntactic treebank, the method of the present invention combines the unlabeled high-quality Vietnamese corpus, and through the multilingual pre-trained model, the large model, and the chain of thought strategy, gradually iteratively constructs a high-quality Vietnamese dependency syntactic treebank. Compared with the traditional method, the method of the present invention can make full use of the existing labeled data and unlabeled data, improve the generation accuracy of the dependency syntactic tree through step-by-step optimization, and reduce the dependence on manual annotation at the same time.

[0121] The core idea of the method of the present invention is to use a multi-step training strategy as a framework to improve the quality and parsing performance of dependency syntax trees in stages. First, a multilingual pre-trained model (XLM-RoBERTa) is fine-tuned through a publicly available dependency syntax tree bank to enhance its syntactic representation ability in Vietnamese, and combined with a bi-affine dependency syntax analysis model (BiAffine Parser) to train the best-performing bi-affine dependency syntax analysis model. Then, the best-performing bi-affine dependency syntax analysis model is used to perform syntactic parsing on high-quality Vietnamese corpora to generate traditional model dependency syntax trees. On this basis, specific prompt words are designed to guide the large model to perform chunking and main-clause identification on high-quality Vietnamese corpora, and the chain-of-thought strategy is used to further optimize the prompting process to help the large model utilize all information more efficiently in deep parsing. Finally, more accurate syntax trees are generated through repeated iterative optimization. Ultimately, the constructed high-quality dependency syntax tree dataset is fused with the publicly available dependency syntax tree bank for retraining the BiAffine Parser model, significantly improving the performance of the model in the dependency syntax analysis task. Experimental results show that the method of the present invention has achieved high LAS (Labelled Attachment Score) and UAS (Unlabelled Attachment Score) scores in the Vietnamese dependency syntax analysis task, especially performing excellently in the parsing of long sentences and complex grammar structures.

[0122] To verify the effectiveness of the method of the present invention, the present invention conducts experiments for verification. The present invention uses the test set to evaluate the new model and analyzes the LAS (Labelled Attachment Score) and UAS (Unlabelled Attachment Score) scores. The specific formula for obtaining a model with better syntactic parsing performance is as follows:

[0123] (1) LAS predicts the accuracy of the model in correctly predicting dependency relations and dependency labels;

[0124]

[0125] (2) UAS predicts the accuracy of dependency relations (without considering dependency labels);

[0126]

[0127] To verify the effectiveness of the method proposed in the present invention, the present invention selected data from the classic database (UniversalDependencies, UD) as the publicly available dependency syntactic treebank to train the model, and used the trained model as the baseline model parser. Then, the baseline model parser was used to perform syntactic parsing on high-quality Vietnamese corpora to generate a preliminary pseudo-dependency syntactic tree dataset. On this basis, specific prompting words were designed to guide the large model to perform chunking and main-clause identification on high-quality Vietnamese corpora, and the chain-of-thought technique was combined to further optimize the prompting process, thereby generating a more accurate syntactic tree dataset.

[0128] The purpose of this operation is to guide the large model to perform chunking and main-clause identification on high-quality Vietnamese corpora by designing targeted prompting words, and at the same time combine the chain-of-thought technique to gradually optimize the prompting process, thereby generating a more accurate chunking dataset and main-clause dataset. Subsequently, these data were jointly input into the large model for training with the Vietnamese pseudo-corpus treebank to significantly improve the efficiency and generalization ability of the dependency syntactic analysis model. The method of the present invention effectively alleviates the problem of scarce Vietnamese dependency syntactic data, significantly improves the processing performance of Vietnamese, and provides new ideas for natural language processing tasks of low-resource languages.

[0129] The experimental results of the present invention are shown in Table 1. The table shows that the method proposed in the present invention outperforms the baseline model in all evaluation metrics, which indicates that after introducing the multi-language collaborative training strategy, the model can not only effectively capture the unique grammatical structures of Vietnamese, but also achieve cross-linguistic representation transfer through Chinese-English pseudo-data, and the joint training of Chinese-English-Vietnamese further constructs a general syntactic analysis paradigm. By stepwise optimizing training and cross-lingual transfer learning, we have successfully improved the effect of Vietnamese dependency syntactic analysis.

[0130] Finally, the comprehensive application of this series of strategies effectively improves the overall performance of the method of the present invention and significantly enhances the processing ability of Vietnamese natural language processing tasks. This achievement not only provides new ideas and methods for the research on constructing Vietnamese treebanks, but also provides strong support for the development of multi-language technologies and Sino-Vietnamese language and cultural exchanges. Through cross-lingual collaboration and optimization, not only the analysis accuracy of the Vietnamese ontology is consolidated, but also an expandable technical path is provided for the language processing of the Southeast Asian language family, paving the way for the further application of Vietnamese in the field of natural language processing.

[0131] Table 1 is a comparison of the results of the original treebank and the syntactic treebank constructed by the present method

[0132]

[0133]

[0134] The specific embodiments of the present invention have been described in detail above in conjunction with the accompanying drawings. However, the present invention is not limited to the above embodiments, and various changes can be made without departing from the spirit of the present invention within the scope of knowledge possessed by those of ordinary skill in the art.

Claims

1. A method for constructing a Vietnamese dependency syntax treebank based on step-by-step optimization training, characterized in that: The method includes: Step 1. Collect the publicly available dependency syntactic treebank and high-quality Vietnamese corpus as experimental data, and preprocess the high-quality Vietnamese corpus; Step 2. Input the publicly available dependency syntactic treebank into the multilingual pre-trained language model XLM-ROBERTa and the bi-affine dependency parsing model BiAffine Parser for training, and save the bi-affine dependency parsing model with the best performance; Step 3. Use the bi-affine dependency parsing model with the best performance to parse the dependency syntactic tree to be processed, and obtain the pseudo-dependency syntactic tree of the traditional model; Step 4. Design chunking prompt words and main-clause recognition prompt words to guide the large model to perform chunking and main-clause recognition tasks on the high-quality Vietnamese corpus, and save the obtained chunked dataset and main-clause dataset; Step 5. Design a chain-of-thought prompt word for guiding the large model to construct a dependency syntactic tree. Utilize the high-quality Vietnamese corpus dataset, the pseudo-dependency syntactic tree of the traditional model, the chunked data, and the main-clause data to help the large model deeply understand the syntactic structure information, thereby generating the correct dependency syntactic tree; finally, obtain a high-quality dependency syntactic treebank dataset through iterative optimization; Step 6. After fusing the constructed high-quality dependency syntactic treebank dataset with the publicly available dependency syntactic treebank, load it into the multilingual pre-trained language model XLM-ROBERTa and the bi-affine dependency parsing model BiAffine Parser for retraining to obtain a model with better syntactic parsing performance for constructing the Vietnamese dependency syntactic treebank.

2. A method for constructing a Vietnamese dependency syntax tree bank based on step-by-step optimization training according to claim 1, characterized in that: The said Step 1 includes: Step 1.

1. Download the Vietnamese dataset with labels from the general data as the publicly available dependency syntactic treebank dataset; Step 1.

2. Download the Vietnamese dataset without labels from the general dataset as the high-quality Vietnamese corpus; Step 1.

3. Use a tokenization tool to tokenize the high-quality Vietnamese corpus, and then process it into a format that conforms to the general dependency syntactic tree standard according to the tokenization results. Take the tokenization results that conform to the general dependency syntactic tree standard format as the dependency syntactic tree to be processed.

3. A method for constructing a Vietnamese dependency syntax tree bank based on step-by-step optimization training according to claim 1, characterized in that: The said Step 2 includes: Step 2.

1. Input the publicly available dependency syntactic treebank into the multilingual pre-trained language model XLM-ROBERTa to obtain the corresponding word vectors, and then input the word vectors into the bi-affine dependency parsing BiAffine Parser model for training; Step 2.

2. After the training is completed, save the bi-affine dependency parsing model with the best performance obtained.

4. A method for constructing a Vietnamese dependency syntax tree bank based on step-by-step optimization training according to claim 1, characterized in that: The said Step 3 includes: Step 3.

1. Use the bi-affine dependency parsing model with the best performance to parse the dependency syntactic tree to be processed; Step 3.

2. Save the parsed dependency syntactic tree dataset and name it the pseudo-dependency syntactic tree of the traditional model.

5. A method for constructing a Vietnamese dependency syntax tree bank based on step-by-step optimization training according to claim 1, characterized in that: The said Step 4 includes: Step 4.

1. Design chunking prompt words for guiding the large model to perform chunking; Step 4.

2. Input the chunking prompt words and high-quality Vietnamese corpus into the large model, and save the obtained chunked dataset; Step 4.

3. Design the main-clause and subordinate-clause recognition prompt words for guiding the large model to perform the main-clause and subordinate-clause recognition task; Step 4.

4. Input the main-clause and subordinate-clause recognition prompt words and high-quality Vietnamese corpus into the large model, and save the obtained main-clause and subordinate-clause dataset.

6. A method for constructing a Vietnamese dependency syntax tree bank based on step-by-step optimization training according to claim 1, characterized in that: The said Step 5 includes: Step 5.

1. Design the thought-chain prompt words for guiding the large model to perform the syntactic tree construction task; Step 5.

2. Use the thought-chain strategy to input the thought-chain prompt words for the dependent syntactic tree construction task, the high-quality Vietnamese corpus dataset, the chunked dataset, the main-clause and subordinate-clause dataset, and the traditional model's pseudo-dependent syntactic tree into the large model to help the large model perform in-depth parsing, make more efficient use of all information, and generate a more accurate syntactic tree through repeated iteration and optimization; Step 5.

3. Save the output result of the large model and name it the high-quality dependent syntactic tree dataset; The said Step 6 includes: Step 6.

1. Integrate the constructed high-quality dependent syntactic tree dataset with the public dependent syntactic tree library; Step 6.

2. Load the integrated data into the multilingual pre-trained language model XLM-ROBERTa and the bi-affine dependent syntactic analysis model BiAffine Parser for retraining, and save the obtained new model.

7. A method for constructing a Vietnamese dependency syntactic tree bank based on step-by-step optimization training according to claim 6, characterized in that: In the said Step 5.2, the thought-chain strategy is as follows: For each internal node N in the parse tree, first extract the child nodes and their ranges: {C1, C2, …, C M} = Children(N) {s 1 ,s 2 ,…,s M} = TextSpan(C 1 ,C 2 ,…,C M ) Among them, C1, C2, …, C M are the child nodes of node N, and s1, s2, …, s M are the corresponding text ranges of the child nodes C1, C2, …, C M That is, the content in the block, Children(.) represents child nodes, and TextSpan(,) represents text ranges; Filter the valid chunks from the extracted child nodes and their ranges; If s1, s2, …, s M all belong to the chunk set FilteredChunks, then concatenate the text ranges of s1, s2, …, s M to generate a new range s, and construct a prompt Prompt(N) for the node N. Then repeat the above steps to generate prompts for each node N. The final set of prompts is Prompts: s = Concat(s1, s2, …, s m ) Prompt(N) = {s1, s2, …, s M are combined into a new text spans} Among them, Concat() represents concatenation, which means generating prompts for each node N and merging them into the final set of prompts.

8. A Vietnamese dependency syntax tree bank construction system based on step-by-step optimization training, characterized in that, The said system includes: a module for executing a method for constructing a Vietnamese dependent syntactic tree library based on step-by-step optimization training as described in any one of claims 1 to 7.

9. An electronic device, comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, characterized in that, When the said processor executes the said program, it implements a method for constructing a Vietnamese dependent syntactic tree library based on step-by-step optimization training as described in any one of claims 1 to 7.

10. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that, When the said computer program is executed by the processor, it implements a method for constructing a Vietnamese dependent syntactic tree library based on step-by-step optimization training as described in any one of claims 1 to 7.