Cross-language dependency syntactic analysis method based on large model migration and self-optimization synthetic data enhancement

Through large-scale model migration and self-optimized synthetic data enhancement methods, high-quality synthetic data is generated, which solves the problem of data scarcity in low-resource languages in cross-language dependency syntax analysis, improves the ability to dependency syntax parsing, and promotes data quality and language interoperability of cross-language tasks.

CN120297264APending Publication Date: 2025-07-11KUNMING UNIV OF SCI & TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510379408.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-28
Publication Date
2025-07-11

AI Technical Summary

Technical Problem

The lack of data in low-resource languages in cross-language dependency syntax analysis is particularly reduced in the accurate analysis of syntactic structures and dependencies with large language differences.

Method used

Through the method of large-model migration and self-optimization synthetic data enhancement, pseudo-data is generated using large-model and cross-language syntax parser, combined with fine-grained syntax parsing instructions and iterative self-optimization algorithms, high-quality synthetic data is constructed, language commonalities and differences are identified, and high-quality synthetic data is generated.

Benefits of technology

It significantly improves the accuracy of dependency syntax analysis of low-resource languages, constructs a high-quality synthetic syntax tree library, improves the data quality and reliability of cross-language tasks, breaks language barriers, and promotes regional economic integration and language interoperability.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120297264A_ABST
    Figure CN120297264A_ABST
Patent Text Reader

Abstract

The invention relates to a cross-language dependency syntactic analysis method based on large model migration and self-optimization synthetic data enhancement, and belongs to the field of natural language processing. Firstly, on one hand, a fine-grained syntactic analysis instruction is designed to guide a large model to recognize cross-language syntactic generality and difference information so as to generate pseudo data based on the large model; on the other hand, a cross-language syntax parser is trained and is also used for generating pseudo data based on the cross-language syntax parser; secondly, designing an iterative self-optimization algorithm to enable the large model to fuse the advantages of the two types of pseudo data, so as to obtain high-quality synthetic data; according to the method, organic fusion of a large-model semantic understanding advantage and a traditional model structure analysis advantage is effectively realized, syntactic generality and difference of different languages are deeply mined, the accuracy of target low-resource language dependency analysis is remarkably improved, a high-quality synthetic syntactic tree bank is constructed, and the method is suitable for large-scale and large-scale analysis. And more convenience is provided for cross-language tasks.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Technical language

[0002] The present invention relates to a cross - language dependency syntax analysis method based on large - model migration and self - optimizing synthetic data augmentation, belonging to the technical field of natural language processing. Background Art

[0003] As the core carrier of cross - cultural communication, the accurate parsing ability of language is directly related to the depth and efficiency of international cooperation, and can also, to a certain extent, avoid misunderstandings and conflicts. Dependency syntax analysis aims to intuitively display the syntactic structure and grammatical relationships of sentences using a hierarchical tree - like structure, which can effectively alleviate the above - mentioned problems. In addition, dependency syntax analysis can also be directly used to enhance other practical natural language processing tasks, such as neural machine translation, named entity recognition, and sentiment analysis, etc.

[0004] With the rapid development of natural language processing technology, the ability of large language models in processing rich - resource languages such as Chinese and English has been significantly improved. However, their performance in cross - language or low - resource language tasks will be significantly reduced, especially in the accurate parsing of syntactic structures and dependency relationships with large language differences. Building an efficient syntax analysis model often relies on a large amount of accurate labeled data, which poses a huge challenge to those languages with scarce resources. Facing this problem, this research proposes an innovative cross - language dependency syntax analysis method that uses large models and traditional syntax parsing models to label data for low - resource languages, and then iteratively optimizes by combining the preferences of two types of data generation to obtain higher - quality synthetic data. Through a detailed comparative analysis of the syntactic structures of multiple language pairs, we found that there are similarities and differences in the structures of the two languages in each pair. For example, in the Chinese - Vietnamese language pair, they both adopt the subject - verb - object syntactic pattern (e.g., "I love you" in Chinese corresponds to ) in Vietnamese), but in Chinese, there are unique structures such as the "ba" and "bei" characters, and most attributives are placed before the headword, while in Vietnamese, most attributives are placed after the headword. Fully identifying and utilizing the common structures while reducing the interference of different structures is crucial for cross - language dependency syntax analysis. In addition, this not only helps to break down language barriers and promote the process of regional economic integration, but also further promotes language interoperability and cultural exchanges between different countries. Summary of the Invention

[0005] The technical problem to be solved by the present invention is: The present invention provides a cross - language dependency syntax analysis method based on large - model migration and self - optimizing synthetic data augmentation, which is used to alleviate the problem of scarce data of low - resource languages in cross - language tasks, deeply explore language commonalities and differences, and at the same time achieve good experimental results.

[0006] The technical solution of the present invention is: a cross-lingual dependency parsing method based on large model migration and self-optimizing synthetic data augmentation. The specific steps of the method are as follows:

[0007] Step1: For synthetic data construction, collect annotated data in multiple languages and perform preprocessing, and pair the source language and the target language at the same time; for cross-lingual training of traditional models, collect a labeled common data set;

[0008] Step2: Use a large model and a cross-lingual syntactic parser to transfer useful syntactic knowledge from the source language, and perform syntactic parsing on the unannotated data of the target language to obtain two annotated versions of pseudo data; on the one hand, a fine-grained syntactic parsing instruction is designed to guide the large model to compare and extract syntactic commonalities and differences to generate large model-based pseudo data; on the other hand, a cross-lingual syntactic parser is trained to also generate cross-lingual syntactic parser-based pseudo data;

[0009] Step3: Design an iterative self-optimization algorithm to allow the large model to integrate the advantages of the two types of pseudo data, so as to obtain high-quality synthetic data.

[0010] Further, the specific steps of Step1 are as follows:

[0011] Step1.1: Download data sets in multiple languages with labels from the Universal Dependencies (UD) as traditional model training data, and download unannotated raw sentences in multiple languages from three parallel translation corpora ALT, WMT, and FLORES-200;

[0012] Step1.2: Remove duplicates and perform word segmentation on the translation corpus data, and at the same time process it into an unannotated syntactic corpus format.

[0013] Further, in Step2, a fine-grained syntactic parsing instruction is designed to guide the large model to compare and extract syntactic commonalities and differences to generate large model-based pseudo data; the specific strategy includes three parts: sentence splitting, multi-lingual chunking, and contrast alignment;

[0014] First, split the input sentence into a main clause and a subordinate clause, and stipulate that the root node in the main clause is the only root node of the whole sentence; second, continue to decompose the sentence into fine-grained semantic unit chunks with specific meanings, and this chunk-level processing achieves precise cross-lingual matching through position alignment; finally, shuffle the semantic chunks of the target language so that they are aligned with the semantic chunks of the source language; by analyzing the semantic chunks whose positions are unchanged and those whose positions are changed during the alignment process, the large model automatically learns language-invariant information and language-specific difference information, and uses this knowledge to help the large model perform syntactic parsing on the target language to obtain large model-based pseudo data.

[0015] Furthermore, in the Step 2, training a cross-lingual syntactic parser is also used to generate pseudo-data based on the cross-lingual syntactic parser. Specifically, the cross-lingual syntactic parser is divided into four parts: an input layer, an encoding layer, a multi-layer perceptron layer (MLPs), and a bi-affine decoding layer (BiAffine). First, the input layer converts the input sentences w1, w2, …, w n into dense vector representations x1, x2, …, x n , and each word vector x i is the average value rep i XLM-R of the last four-layer word vector representations extracted from the pre-training of XLM-RoBERTa and the randomly initialized word embedding representation emb i word added together, and then concatenated with the character representation word i char corresponding to each word. The formula is as follows,

[0016]

[0017] Then, a three-layer bidirectional long short-term memory network is used to encode x i into an intermediate vector h i with context information; then h i undergoes feature dimensionality reduction through the multi-layer perceptron (MLPs) to obtain the vector representation of each pair of words (w i , w j respectively as the central word and the vector representation of each pair of words (w i , w j as the modifier. Secondly, the bi-affine operation area is used to calculate the dependency arc score s(i←j) and the dependency label (Label, l) score of each neutral word and modifier. The specific formula is as follows,

[0018]

[0019] where U 1 , U 2 , U 3 represent the weight matrices of different bi-affine layers respectively, b represents the bias. In addition, the arc confidence score s arc with the highest probability and the label confidence score s label with the highest probability are obtained to provide confidence indicators for subsequent optimization operations. The formula is as follows:

[0020]

[0021] Finally, the cross - language syntactic parser defines a local cross - entropy loss function for each pair of words (W i , W j ) to optimize the model parameters. The loss function is defined as follows:

[0022]

[0023] where k is used to traverse any k - th word among the total of n words, distinguished from the j - th word, and l′ is used to traverse any l′ - th label among the total of L label types, distinguished from the l - th label predicted by the model;

[0024] In each training epoch, the source language and the target language are alternately trained so that the model contains rich common syntactic knowledge. When the model training ends or converges in advance to obtain a cross - language syntactic parser with optimal performance, this model is used to perform syntactic annotation on the unlabeled target - language data to obtain pseudo - data based on the cross - language syntactic parser.

[0025] Furthermore, the specific steps of Step 3 are as follows:

[0026] For the two types of generated pseudo - data: pseudo - data based on the large model and pseudo - data based on the cross - language syntactic parser, first, if the head - node prediction results of the two are the same, or the arc confidence score s arc of the cross - language syntactic parser exceeds the set threshold, then the label consistency is checked; if the labels are consistent or the label confidence score s label exceeds the threshold, the result of the cross - language syntactic parser is directly adopted; otherwise, the large model combines the results of the two to regenerate new labels and update the confidence scores; if the head - nodes are inconsistent or the arc confidence is lower than the threshold, the large model generates new head - nodes and labels and updates the confidence scores; this process continues to iterate until the confidence scores of the newly generated arcs and labels both exceed the threshold or reach the preset maximum number of iterations, thereby obtaining optimized high - quality synthetic data.

[0027] The present invention also provides a cross - language dependency syntactic analysis system based on large - model migration and self - optimizing synthetic data augmentation, and the system includes: a module for executing the above - mentioned cross - language dependency syntactic analysis method based on large - model migration and self - optimizing synthetic data augmentation.

[0028] The present invention also provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, it implements the above - mentioned cross - language dependency syntactic analysis method based on large - model migration and self - optimizing synthetic data augmentation.

[0029] The present invention also provides a non - transitory computer - readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it implements the above - mentioned cross - language dependency syntax analysis method based on large - model migration and self - optimizing synthetic data enhancement.

[0030] The present invention also provides a computer program product, including a computer program. When the computer program is executed by a processor, it implements the above - mentioned cross - language dependency syntax analysis method based on large - model migration and self - optimizing synthetic data enhancement.

[0031] The beneficial effects of the present invention are as follows:

[0032] 1. The present invention innovatively proposes a fine - grained cross - language syntactic knowledge migration strategy based on large - language models, aiming to guide large - language models to identify and extract deeper commonalities and differences among multiple languages, thereby significantly improving their dependency syntax parsing ability and significantly enhancing the accuracy of dependency parsing of target low - resource languages.

[0033] 2. The present invention innovatively designs a self - optimization method for aligning the preferences of large - language models and cross - language syntax parsers in syntactic structure generation. By extracting consistent syntactic knowledge from the pseudo - data generated by large - language models and the pseudo - data generated by cross - language syntax parsers and eliminating syntactic biases, our method significantly improves the quality and reliability of synthetic data, constructs a high - quality synthetic syntax tree bank, and provides more convenience for cross - language tasks.

[0034] 3. The present invention constructs a large number of synthetic syntax tree banks that can be used for cross - language tasks, and verifies the quality of the data and the correctness of the labels from two levels: model evaluation and manual verification. BRIEF DESCRIPTION OF THE DRAWINGS

[0035] Figure 1 It is a flowchart in the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0036] Embodiment 1: As Figure 1 shown, a cross - language dependency syntax analysis method based on large - model migration and self - optimizing synthetic data enhancement uses large - models and cross - language syntax parsers for experiments, which is mainly divided into three parts: migration - based pseudo - data generation, pseudo - data self - optimization, and synthetic data evaluation. The specific steps of the method are as follows:

[0037] Step1: For synthetic data construction, collect labeled data in multiple languages and perform pre - processing, and pair the source language and the target language at the same time; for cross - language training of traditional models, collect a labeled public data set.

[0038] The specific steps of Step1 are as follows:

[0039] Step1.1: Download datasets of multiple languages with tags from the Universal Dependencies (UD) treebank as training data for traditional models, and download raw sentences of multiple languages without annotation from three parallel translation corpora, namely ALT, WMT, and FLORES-200;

[0040] Step1.2: Remove duplicates and perform word segmentation on the translation corpus data, and at the same time process it into an unannotated syntactic corpus format.

[0041] Step2: Use a large model and a cross-lingual syntactic parser to transfer useful syntactic knowledge from the source language, and perform syntactic parsing on the unannotated data of the target language respectively to obtain two annotated versions of pseudo-data; on the one hand, fine-grained syntactic parsing instructions are designed to guide the large model to compare and extract syntactic commonalities and differences information to generate pseudo-data based on the large model; on the other hand, a cross-lingual syntactic parser is trained to also generate pseudo-data based on the cross-lingual syntactic parser;

[0042] In the above Step2, fine-grained syntactic parsing instructions are designed to guide the large model to compare and extract syntactic commonalities and differences information to generate pseudo-data based on the large model; the specific strategy includes three parts: sentence splitting, multi-lingual chunking, and contrastive alignment;

[0043] First, to solve the problem of multiple root node ambiguities in complex sentences, the input sentence is split into a main clause and a subordinate clause, and it is stipulated that the root node in the main clause is the only root node of the whole sentence; second, to explore the structural differences between languages, the sentence is further decomposed into fine-grained semantic unit chunks with specific meanings (e.g., noun / verb phrases), and this chunk-level processing achieves precise cross-lingual matching through position alignment; finally, the semantic chunks of the target language are scrambled so that they are aligned with the semantic chunks of the source language; by analyzing the semantic chunks that have not changed positions and those that have changed positions during the alignment process, the large model automatically learns language-invariant information (semantic chunks that have not changed positions) and language-specific difference information (semantic chunks that have changed positions), and uses this knowledge to help the large model perform syntactic parsing on the target language to obtain pseudo-data based on the large model.

[0044] Furthermore, in the above Step2, a cross-lingual syntactic parser is trained to also generate pseudo-data based on the cross-lingual syntactic parser. Specifically, this cross-lingual syntactic parser is divided into four parts: an input layer, an encoding layer, a multi-layer perceptron layer (MLPs), and a bi-affine decoding layer (BiAffine); first, the input layer converts the input sentences w1, w2, …, w n into dense vector representations x1, x2, …, x n , and each word vector x i is the average value rep of the word vector representations extracted from the last four layers of pre-training of XLM-RoBERTa iXLM-R and the randomly initialized word embedding representation emb i word are added together, and then combined with the character representation word corresponding to each word i char The formula is as follows

[0045]

[0046] Then, a three-layer bidirectional long short-term memory network is used to encode x i into an intermediate vector h with context information i ; Then h i undergoes dimensionality reduction of features through multi-layer perceptrons MLPs to obtain the vector representation of each pair of words (w i , w j ) as the vector representation of the central word respectively and the vector representation of each pair of words (w i , w j ) as the vector representation of the modifier Secondly, a double-affine operation area is used to calculate the dependency arc score s(i←j) and the dependency label (Label, l) score of each neutral word and modifier The specific formula is as follows

[0047]

[0048] Among them, U 1 , U 2 , U 3 represent the weight matrices of different double-affine layers respectively, b represents the bias. In addition, the arc confidence score s arc with the highest probability and the label confidence score s label with the highest probability are obtained to provide confidence indicators for subsequent optimization operations. The formula is as follows

[0049]

[0050] Finally, the cross-lingual syntactic parser defines a local cross-entropy loss function for each pair of words (w i , w j ) to optimize the model parameters. The loss function is defined as follows

[0051]

[0052] Among them, k is used to traverse any k-th word among a total of n words, distinguished from the j-th word, and l′ is used to traverse any l′-th label among a total of L label types, distinguished from the l-th label predicted by the model;

[0053] In each training cycle, the source language and the target language are alternately trained so that the model contains rich common syntactic knowledge. When the model training is completed or converges in advance to obtain a cross-lingual syntactic parser with optimal performance, this model is used to perform syntactic annotation on the unannotated target language data to obtain pseudo-data based on the cross-lingual syntactic parser.

[0054] Step3: Design an iterative self-optimization algorithm to enable the large model to integrate the advantages of the two types of pseudo-data, thereby obtaining high-quality synthetic data.

[0055] Furthermore, the specific steps of Step3 are as follows:

[0056] For the two types of generated pseudo-data: pseudo-data based on the large model and pseudo-data based on the cross-lingual syntactic parser. First, if the head node prediction results of the two are the same, or the arc confidence score s arc of the cross-lingual syntactic parser exceeds the set threshold, then check the label consistency; if the labels are consistent or the label confidence score s label exceeds the threshold, directly adopt the results of the cross-lingual syntactic parser; otherwise, the large model combines the results of the two to regenerate new labels and update the confidence scores; if the head nodes are inconsistent or the arc confidence is lower than the threshold, the large model generates new head nodes and labels and updates the confidence scores; this process is continuously iterated until the confidence scores of the newly generated arcs and labels both exceed the threshold or reach the preset maximum number of iterations, thereby obtaining optimized high-quality synthetic data.

[0057] The present invention also provides a cross-lingual dependency syntactic analysis system based on large model migration and self-optimized synthetic data enhancement. The system includes: a module for executing the above-mentioned cross-lingual dependency syntactic analysis method based on large model migration and self-optimized synthetic data enhancement.

[0058] The present invention also provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, it implements the above-mentioned cross-lingual dependency syntactic analysis method based on large model migration and self-optimized synthetic data enhancement.

[0059] The present invention also provides a non-transitory computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it implements the above-mentioned cross-lingual dependency syntactic analysis method based on large model migration and self-optimized synthetic data enhancement.

[0060] The present invention also provides a computer program product, including a computer program. When the computer program is executed by a processor, it implements the above-mentioned cross-lingual dependency syntactic analysis method based on large model migration and self-optimized synthetic data enhancement.

[0061] The present invention uses three classic cross - language models and three popular large models to verify the effectiveness of the method of the present invention.

[0062] Specifically, it includes the following:

[0063] Three classic cross - language syntactic parsing models are used to verify the effectiveness of the method, and finally the enhancement effect of the synthetic dataset on the model performance. The three classic cross - language syntactic parsing models are as follows:

[0064] 1) Fully shared model: We also share all the parameters of the model and alternately use the source - language and target - language training datasets to train the XLM - RoBERTa - enhanced bi - affine syntactic parser.

[0065] 2) Language embedding model: On the basis of the previous model, we also concatenate an 8 - dimensional language - type embedding vector at the last dimension of each word to distinguish language types.

[0066] 3) Multi - task learning model: We use rich - resource source - language syntactic parsing as an auxiliary task to enhance the target - language syntactic parsing task.

[0067] Three popular large models are used to verify the effectiveness of our method. The three large models are as follows:

[0068] 1) GPT - 4o - mini(8B): It is a distilled closed - source multilingual model. Its compact architecture retains the basic cross - language transfer ability, making it particularly effective in parsing languages mediated by English.

[0069] 2) Qwen2.5 - 7B - Instruct: It is customized for Chinese languages, uses word embeddings optimized for the CJK (Chinese, Japanese, and Korean) language family, and enhances its ability to handle complex grammar and semantic conversions in these languages.

[0070] 3) Llama3.1 - 8B - Instruct: It uses a longer context window to facilitate cross - language morphological analysis. Its open - source nature ensures transparency and reproducibility, enabling researchers to easily replicate experiments.

[0071] Finally, the quality of our synthetic data is verified from two aspects: model evaluation and human evaluation.

[0072] Specifically, it includes the following:

[0073] Use large models to evaluate the correctness of the final synthetic data;

[0074] Invite language experts to evaluate the correctness of our synthetic data, especially the syntactic labels.

[0075] The present invention uses the untagged dependency score UAS and the tagged dependency score LAS as evaluation metrics for the performance of the bi-affine syntactic analysis model. The calculation formulas are as follows:

[0076]

[0077] To prove the effectiveness of the present invention, the above-mentioned three cross-lingual dependency syntactic analysis models and three large models are selected as benchmark models. In the present invention, first, the large model and the cross-lingual syntactic parser are used to annotate the target language to obtain two annotated versions; then, a self-optimization process is designed to iteratively optimize the two annotated data to obtain high-quality synthetic data; finally, through strict multiple evaluation methods to verify our high-quality synthetic data.

[0078] Table 1 compares the model performance of three classic traditional models before and after adding our synthetic data. By adding our synthetic data, the performance of the three traditional models has been significantly improved, which fully proves that our synthetic data can provide more and effective syntactic knowledge to the traditional models, improve the ability of the models to extract syntactic representations of specific languages, and proves the effectiveness of our method and the reliability of the quality of the synthetic data.

[0079] Table 1 Scores of three traditional models on the test sets of Vietnamese and Maltese

[0080]

[0081] Table 2 Scores of three large models on the test sets of Vietnamese and Maltese

[0082]

[0083]

[0084] Table 2 compares the performance of different methods of three large models, and evaluates their performance under zero-shot, one-shot, few-shot, and our method. First, Llama-3.1-8B-Instruct shows the best performance, which highlights its stronger syntactic reasoning and language generalization ability; then, when a certain number of example prompts are provided, all large models show significant improvements, proving that example prompts can provide some internal grammar rules of the language to help large models better understand the grammar paradigms of different languages; finally, our transfer and self-optimization methods significantly improve the parsing performance of all languages, proving its effectiveness. By comparing the commonalities and differences of the grammar structures of different languages, and by combining the advantages of pseudo-data based on large models and cross-lingual syntactic parsers, optimization is carried out to generate higher-quality synthetic data, enabling large models to learn more fine-grained useful grammar knowledge.

[0085] Model and Manual Evaluation Results of the Correctness of Partial Dependency Labels in Synthetic Data in Table 3

[0086]

[0087] As shown in Table 3, it can be seen from the table that through the verification of both model evaluation and manual evaluation, the correctness of the partial dependency labels of our synthetic data is very high, which proves the effectiveness of our method and the quality reliability of the synthetic data.

[0088] The above table and attached drawings fully prove the effectiveness and rationality of the present invention. However, the present invention is not limited to the above embodiments, and various changes can be made without departing from the gist of the present invention within the scope of knowledge possessed by those of ordinary skill in the art.

Claims

1. A cross-lingual dependency syntax analysis method based on large model migration and self-optimizing synthetic data augmentation, characterized in that: The specific steps of the method are as follows: Step1: For synthetic data construction, collect annotated data in multiple languages and perform preprocessing, while pairing the source language and the target language; for cross-language training of traditional models, collect a labeled common dataset; Step2: Use a large model and a cross-language syntactic parser to transfer useful syntactic knowledge from the source language, and perform syntactic parsing on the unannotated data in the target language to obtain two annotated versions of pseudo-data. On the one hand, fine-grained syntactic parsing instructions are designed to guide the large model to compare and extract syntactic commonalities and differences to generate large model-based pseudo-data; on the other hand, a cross-language syntactic parser is trained to also generate cross-language syntactic parser-based pseudo-data; Step3: Design an iterative self-optimization algorithm to allow the large model to integrate the advantages of the two types of pseudo-data, thereby obtaining high-quality synthetic data.

2. The cross-lingual dependency syntactic analysis method based on large model migration and self-optimized synthetic data augmentation according to claim 1, characterized in that: The specific steps of Step1 are as follows: Step1.1: Download datasets in multiple languages with labels from the Universal Dependencies (UD) as traditional model training data, and download unannotated original sentences in multiple languages from three parallel translation corpora, ALT, WMT, and FLORES-200; Step1.2: Remove duplicates and perform word segmentation on the translation corpus data, and at the same time process it into an unannotated syntactic corpus format.

3. The cross-lingual dependency syntactic analysis method based on large model migration and self-optimizing synthetic data augmentation according to claim 1, characterized in that: In Step2, fine-grained syntactic parsing instructions are designed to guide the large model to compare and extract syntactic commonalities and differences to generate large model-based pseudo-data; the specific strategy includes three parts: sentence splitting, multi-language chunking, and contrastive alignment; First, split the input sentence into a main clause and a subordinate clause, and stipulate that the root node in the main clause is the only root node of the entire sentence; second, continue to decompose the sentence into fine-grained semantic unit chunks with specific meanings, and this chunk-level processing achieves precise cross-language matching through position alignment; finally, shuffle the semantic chunks in the target language so that they are aligned with the semantic chunks in the source language; By analyzing the semantic chunks that have not changed positions and those that have changed positions during the alignment process, the large model automatically learns language-invariant information and language-specific difference information, and uses this knowledge to help the large model perform syntactic parsing on the target language to obtain large model-based pseudo-data.

4. The cross-lingual dependency syntactic analysis method based on large model migration and self-optimizing synthetic data augmentation according to claim 1, wherein: In the above Step 2, a cross - language syntactic parser is trained and also used to generate pseudo - data based on the cross - language syntactic parser. Specifically, the cross - language syntactic parser is divided into four parts: an input layer, an encoding layer, a multi - layer perceptron layer (MLPs), and a bi - affine decoding layer (BiAffine). First, the input layer converts the input sentences \(w_1, w_2,\cdots, w\) n into dense vector representations \(x_1, x_2,\cdots, x\) n , where each word vector \(x\) i is the average value \(rep\) i XLM-R of the last four - layer word vector representations extracted from the pre - trained XLM - RoBERTa and the randomly initialized word embedding representation \(emb\) i word added together, and then concatenated with the character representation \(word\) i char corresponding to each word. The formula is as follows: Then, a three-layer bidirectional long short-term memory network is used to encode x i into an intermediate vector h with context information i ; then h i undergoes feature dimensionality reduction through multi-layer perceptrons (MLPs) to obtain the vector representations of each pair of words (w i , w j ) as the vector representation of the central word and the vector representations of each pair of words (w i , w j ) as the vector representation of the modifier Secondly, a bi-affine operation area is used to calculate the dependency arc score s(i←j) and the dependency label (Label, l) score for each neutral word and modifier The specific formula is as follows Among them, U 1 , U 2 , U 3 respectively represent the weight matrices of different bi - affine layers, b represents the bias. In addition, the arc confidence score s arc with the highest probability and the label confidence score s label with the highest probability are obtained to provide confidence metrics for subsequent optimization operations. The formula is as follows: Finally, the cross-lingual syntactic parser defines a local cross-entropy loss function for each pair of words (w i , w j ) to optimize the model parameters. The loss function is defined as follows: Among them, k is used to traverse any k-th word among a total of n words, distinguished from the j-th word, and l′ is used to traverse any l′-th label among a total of L label types, distinguished from the l-th label predicted by the model; In each training cycle, alternately train the source language and the target language so that the model contains rich common syntactic knowledge. When the model training ends or converges in advance to obtain a cross-language syntactic parser with optimal performance, use this model to perform syntactic annotation on the unannotated target language data to obtain cross-language syntactic parser-based pseudo-data.

5. The cross-lingual dependency syntactic analysis method based on large model migration and self-optimizing synthetic data augmentation according to claim 1, characterized in that: The specific steps of Step3 are as follows: For the two types of generated pseudo-data: pseudo-data based on large models and pseudo-data based on cross-lingual syntactic parsers. First, if the head node prediction results of both are the same, or the arc confidence score s of the cross-lingual syntactic parser arc exceeds the set threshold, then check the label consistency; if the labels are the same or the label confidence score s label exceeds the threshold, directly adopt the results of the cross-lingual syntactic parser; otherwise, the large model combines the results of both to regenerate new labels and update the confidence scores; If the head nodes are inconsistent or the arc confidence is lower than the threshold, the large model generates new head nodes and labels and updates the confidence scores; This process continues to iterate until both the newly generated arc and label confidence scores exceed the threshold or the preset maximum number of iterations is reached, thus obtaining optimized high-quality synthetic data.

6. A cross-lingual dependency parsing system based on large model migration and self-optimizing synthetic data augmentation, characterized in that, The system includes: a module for executing a cross-lingual dependency syntactic analysis method based on large model migration and self-optimizing synthetic data enhancement as described in any one of claims 1 to 5.

7. An electronic device, comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, characterized in that, When the processor executes the program, it implements a cross-lingual dependency syntactic analysis method based on large model migration and self-optimizing synthetic data enhancement as described in any one of claims 1 to 5.

8. A non-transitory computer-readable storage medium storing a computer program thereon, characterized in that, When the computer program is executed by the processor, it implements a cross-lingual dependency syntactic analysis method based on large model migration and self-optimizing synthetic data enhancement as described in any one of claims 1 to 5.

9. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by the processor, it implements a cross-lingual dependency syntactic analysis method based on large model migration and self-optimizing synthetic data enhancement as described in any one of claims 1 to 5.