Large model Mian machine translation method based on language knowledge retrieval

By constructing a tree-shaped search structure and cross-language text representation tool, the parallel corpus organization and utilization of low-resource languages is optimized, and the problems of lack of corpus and low quality in low-resource language translation are solved, and the translation accuracy and fluency are improved. It is suitable for translation tasks in a variety of low-resource languages.

CN120409500APending Publication Date: 2025-08-01KUNMING UNIV OF SCI & TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510555098.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-29
Publication Date
2025-08-01

AI Technical Summary

Technical Problem

In the translation tasks of low-resource languages, the existing technology faces the problems of lack of parallel corpus, high corpus noise, low translation quality and limited applicability in low-resource environments, resulting in increased model learning difficulty and insufficient translation performance.

Method used

The large-scale Chinese-Myanmar machine translation method based on language knowledge retrieval is adopted. By constructing a tree-shaped search structure, combining TF-IDF, BM25, LASER embedding and Frobenius norm regularization technology, the organization and utilization of parallel corpus are optimized, and the cross-language text representation tools are used for semantic matching and sorting, and the translation accuracy is improved.

Benefits of technology

It improves the accuracy and fluency of translation results of low-resource languages, alleviates the problem of insufficient corpus, provides new solutions for low-resource language translation, and is suitable for translation tasks of multiple low-resource languages.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120409500A_ABST
    Figure CN120409500A_ABST
Patent Text Reader

Abstract

The invention relates to a large model Mian machine translation method based on language knowledge retrieval, and belongs to the field of natural language processing. Comprising the following steps: constructing tree root nodes of a tree-shaped retrieval structure for parallel corpora from Chinese to Muranju language; dividing the parallel corpora from Chinese to Myanforn into long sentence pairs and short sentence pairs according to sentence lengths; constructing a preliminary tree-shaped retrieval structure through the short sentence pair; inserting the long sentence pair into a tree-shaped retrieval structure; in the retrieval stage, firstly, a to-be-translated text is coded through an LASER model and then matched with related words and sentence pairs in a tree retrieval structure, candidate sentence pairs are selected through a BM25 algorithm, and the cosine similarity between the candidate sentence pairs and the to-be-translated text is further calculated; before the large language model is used for final translation, the text reordering model is used for reordering the context prompt sentence pairs obtained through retrieval, and the context prompt sentence pairs and the to-be-translated text are input into the large language model to obtain a final translation result. The accuracy and smoothness of the translation result can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a Chinese-Myanmar machine translation method for large models based on language knowledge retrieval, and belongs to the technical field of natural language processing. Background Art

[0002] In the current era of globalization, the market integration, resource sharing, and cultural blending among countries along the line are deepening day by day. As an important medium for communication, language plays a key role in the political, economic, cultural, and technological cooperation among countries. However, for the translation needs of low-resource languages, language barriers are still a major bottleneck restricting international communication and cooperation.

[0003] The translation problem of low-resource languages is particularly prominent in the bilateral exchanges between medium- and low-resource countries. For example, the digital resources in Myanmar have been in a state of long-term scarcity. The existing Myanmar parallel corpus is small in scale and uneven in quality, lacking high-quality annotated corpus, which seriously restricts the research and application of neural machine translation (NMT) models. Although deep learning and NMT technologies have made remarkable progress in multilingual translation tasks in recent years, the improvement of their performance highly depends on large-scale high-quality parallel corpora, and such resources are still scarce for low-resource languages.

[0004] Facing this challenge, exploring effective low-resource translation methods has become a key research direction. However, the difficulties of low-resource translation tasks are far more than the scarcity of parallel corpora, including problems such as large corpus noise, low translation quality, and limited applicability of existing technologies in low-resource environments. On the one hand, low-resource languages often lack unified language norms, resulting in a large number of inconsistent expression forms and semantic ambiguities in text data, further increasing the difficulty of model learning. On the other hand, due to the relatively small number of users of low-resource languages, it is difficult to quickly expand high-quality corpora through traditional manual annotation methods. Even through technical means such as cross-lingual transfer learning or data augmentation, it is difficult to completely make up for the performance loss caused by insufficient corpora. In this context, how to fully tap the potential of existing corpora under limited resource conditions and build a more efficient corpus organization and utilization mechanism has become a key issue in low-resource translation research. Summary of the Invention

[0005] Aiming at the problems existing in the low-resource language translation task, the present invention provides a Chinese-Myanmar machine translation method for large models based on language knowledge retrieval. By introducing a structured parallel corpus retrieval mechanism and combining neural network representation and retrieval enhancement technologies, the present invention organizes and utilizes existing resources more efficiently; the method of the present invention improves the accuracy and fluency of the translation results of low-resource languages.

[0006] The technical solution of the present invention is: a Chinese-Myanmar machine translation method for large models based on language knowledge retrieval, the method comprising:

[0007] Step 1: Collect the Chinese-Burmese parallel corpus dataset;

[0008] Step 2: Segment the Chinese part of the Chinese-Burmese parallel corpus and perform word frequency statistics, remove stop words and words with low word frequencies, and construct the root node of the tree retrieval structure according to the occurrence frequency of Chinese words and phrases;

[0009] Step 3: Divide the Chinese-Burmese parallel corpus into long sentence pairs and short sentence pairs according to the sentence length;

[0010] Among them, the short sentence pairs are matched with the words and phrases of the root node referring to the TF-IDF algorithm and the BM25 algorithm, establish the connection between the short sentences and the words, and use the cross-lingual text representation tool to perform embedding representation on the inserted short sentence pairs and save them to complete the construction of the preliminary tree retrieval structure;

[0011] Step 4: For the processing of long sentence pairs, use the LASER cross-lingual representation tool to perform embedding representation on the sentences, screen the short sentences related to the long sentences in the retrieval tree through the TF-IDF algorithm and the BM25 algorithm, and then calculate the cosine similarity between the short sentences and the long sentences in the source language and the target language to determine the insertion position of the long sentence pairs in the tree retrieval structure;

[0012] Step 5: In the retrieval stage, segment the text to be translated and sequentially match the root nodes of the tree retrieval structure according to the segmentation results, screen the relevant candidate trees for the words and phrases appearing in the text to be translated; secondly, use the cross-lingual text representation tool to obtain the embedding representation of the text to be translated, and calculate the cosine similarity between the embedding representation of the text to be translated and the Chinese and Burmese embedding representations of the long and short sentence pairs in the candidate trees in turn, and take the average value of the cosine similarity scores of each pair of candidate sentence pairs as the final relevance score, and sort them from high to low according to the relevance score as the context hint sentence pairs for the translation task;

[0013] Step 6: Before using the large language model for the final translation, use the text reordering model to reorder the context hint sentence pairs retrieved, and input the context hint sentence pairs and the text to be translated into the large language model to obtain the final translation result.

[0014] Further, the Step 1 includes: collecting the Chinese-Burmese parallel corpus from the OPUS public parallel corpus website.

[0015] Further, the Step 2 includes:

[0016] Step 2.1: Use the jieba word segmentation tool to segment the Chinese part of the Chinese-Burmese parallel corpus and count the word frequencies. Sort the counted words according to the word frequencies to construct a word list. Clean the general Chinese words in the word list and remove the noise content. Remove the last 5% of the low-frequency words to avoid the influence of the long-tail effect of the word list on the retrieval efficiency of the final tree retrieval structure. The noise content includes blank lines and special symbols.

[0017] Step 2.2: Use the MongoDB non-relational database to save the word list information as the root node of the tree retrieval structure. The word list information includes words, word frequencies, and word indexes.

[0018] Further, the said Step 3 includes:

[0019] Step 3.1: Divide the collected Chinese-Burmese parallel corpus into long sentence pairs and short sentence pairs according to the sentence length. The ratio of long sentence pairs to short sentence pairs is 1:3.

[0020] Step 3.2: Filter and score the Chinese part of the short sentence pairs with the words of the root node of the tree retrieval structure in the MongoDB non-relational database in turn with reference to the TF-IDF algorithm and the BM25 algorithm, determine the insertion position of the short sentence pairs in the retrieval tree and insert them into the MongoDB non-relational database.

[0021] Step 3.3: Use the LASER cross-lingual text representation model to perform cross-lingual embedding representation on the Chinese and corresponding Burmese translations of the inserted short sentence pairs, and update the content of the MongoDB non-relational database.

[0022] Further, the said Step 3.3 includes:

[0023] Use the LASER cross-lingual text representation model to perform text representation on the short sentence pairs whose insertion positions are determined. Input the text into the LASER cross-lingual text representation model. First, enter the 5-layer bidirectional long short-term memory network BiLSTM to convert Chinese and Burmese into x1, x2....x i1 , (i1 = 1024) embedding representations. Then input the Decoder layer. After determining the generated language, the Decoder layer converts the embedding representation input from the upper layer into a [1, 1024]-dimensional sentence embedding representation vector through the loss function. After obtaining the sentence embedding representation of the bilingual text, save the sentence embedding representation of the bilingual text, the original text information, and the word information determined after screening in the MongoDB non-relational database.

[0024] Further, the Step4 includes: processing the divided long sentence pairs, using a cross-lingual text representation tool to perform embedding representation on the Chinese and corresponding Burmese translations in each group of long sentence pairs, screening the root nodes of the tree retrieval structure and their associated short sentence pairs through the TF-IDF algorithm and the BM25 algorithm, calculating the cross-cosine similarity between the current long sentence pair and the screened short sentence pairs to generate a similarity matrix, and regularizing the similarity matrix through the Frobenius norm to obtain the final similarity score; determining the insertion positions of the long sentence pairs in the tree retrieval structure according to the order of the similarity scores from high to low, thereby completing the insertion of the long sentence pairs in the tree retrieval structure.

[0025] Further, the specific steps of the Step4 include:

[0026] Step4.1: Use the LASER cross-lingual text representation model to perform cross-lingual representation on the source language and target language of the long sentence pairs;

[0027] Step4.2: Use the TF-IDF algorithm and the BM25 algorithm in sequence to filter and score the words, and retain the top 20 words with the highest scores and their affiliated short sentence pairs;

[0028] Step4.3: Use the cosine similarity to calculate the similarity scores of the source language and target language in the short sentence pairs and long sentence pairs in sequence to obtain a similarity matrix, take the Frobenius norm of the similarity matrix as the final similarity score of the current long sentence pair to be inserted, determine the position of the long sentence pair in the retrieval tree according to the similarity score from high to low, and insert it into the MongoDB non-relational database to complete the insertion of the long sentence pairs in the tree retrieval structure.

[0029] Further, in the Step4, the cosine similarity can calculate the semantic similarity between two texts after high-dimensional representation; if the semantics between the two texts are similar, the final result approaches 1, otherwise it approaches 0; the cosine similarity algorithm is defined as:

[0030]

[0031] where Sim(X,Y) represents the final score of the text similarity calculation, X and Y respectively represent the high-dimensional embedding vectors of the texts participating in the cosine similarity calculation, X i and Y i respectively represent an element in the vector, and n represents the total number of elements;

[0032] In the stage of confirming the insertion position of bilingual long sentence pairs in the tree-shaped retrieval structure, the LASER cross-lingual text representation model is used to obtain the embedded representation of the bilingual text of the long sentence pairs. The embedded representation of the bilingual text of the long sentence pairs is cross-calculated with the bilingual short sentence pairs and their embedded representations screened from the database to form a similarity matrix S. The similarity matrix S is defined as:

[0033]

[0034] where x s and y s represent the Chinese and Burmese bilingual short sentence pairs already inserted into the database, and x l and y l respectively represent the Chinese and Burmese bilingual long sentence pairs for which the insertion position in the tree-shaped retrieval structure is to be determined currently;

[0035] The definition of taking the Frobenius norm of the similarity matrix S is:

[0036]

[0037] where Score(S) represents the final score of the sentence pair to be inserted, ||S|| F represents the calculation of taking the Frobenius norm of the matrix, and s ab is the corresponding sentence pair score after cross-cosine similarity calculation for the element in the a-th row and b-th column of the matrix S;

[0038] After obtaining the final similarity score, record the information of the bilingual short sentence pair with the highest score, as well as the original text and the corresponding text embedded representation of the current bilingual long sentence pair, and insert them into the MongoDB non-relational database.

[0039] Furthermore, the said Step5 includes:

[0040] Step5.1. Perform word segmentation on the text to be translated, and use the LASER cross-lingual text representation model to obtain the embedded representation of the text to be translated;

[0041] Step5.2. For the text to be translated after word segmentation, sequentially use the TF-IDF algorithm and the BM25 algorithm to filter and screen the matching words in the tree-shaped retrieval structure, and retain the words with a BM25 algorithm score greater than 0.0 and their corresponding long and short sentence pairs as the candidate sentence pairs for the context prompt of the Chinese-Burmese machine translation of the large language model;

[0042] Step5.3. Calculate the cosine similarity between the embedding representation of the text to be translated and the embedding representations of the filtered long and short sentence pairs respectively, and take the average of the calculation results as the cosine similarity score of the final sentence pair. Sort the sentence pairs in descending order of the similarity score. After removing the candidate sentence pairs with scores less than 0.3, use the remaining long and short sentence pairs as the context hint sentence pairs for the translation task.

[0043] Further, the Step6 includes:

[0044] Step6.1. Use the bge-reranker-v2-m3 model to reorder the finally obtained context hint sentence pairs, and sort the hint sentence pairs in descending order of the cosine similarity score;

[0045] Step6.2. Determine the final number of hint sentence pairs according to the number of hint examples, and input the hint sentence pairs and the sentence to be translated into the large language model to obtain the final translation result.

[0046] The beneficial effects of the present invention are:

[0047] 1. By constructing a corpus retrieval tree and combining TF-IDF, BM25, LASER embedding, and Frobenius norm regularization techniques, the present invention improves the organization efficiency and utilization effect of parallel corpora, providing a new solution idea for low-resource language translation;

[0048] 2. By introducing a structured parallel corpus retrieval mechanism and combining neural network representation and retrieval enhancement techniques, the present invention can organize and utilize existing resources more efficiently; the present invention can not only alleviate the problem of insufficient corpus, but also provide a new idea for low-resource language translation tasks, bringing new possibilities for solving the translation problems of low-resource languages;

[0049] 3. The present invention designs a complete framework from preprocessing to generation, including the structured organization of corpus preprocessing, the multi-algorithm fusion optimization in the retrieval stage, and the reordering strategy of hint sentence pairs in the generation stage, effectively alleviating the impact of scarce parallel corpora on translation performance;

[0050] 4. The present invention provides a general low-resource translation technology framework. The proposed method is not only applicable to Chinese to Southeast Asian languages but also has generalizability and can be applied to the translation tasks of other low-resource languages, providing reference and technical support for the research and application of low-resource machine translation. Description of the Drawings

[0051] Figure 1 It is a flowchart in the present invention. Detailed Embodiments

[0052] Example 1: As Figure 1As shown, a Chinese-Myanmar machine translation method based on language knowledge retrieval, the method includes:

[0053] Step1. Collect a Chinese-to-Myanmar parallel corpus dataset from the OPUS open parallel corpus website;

[0054] Step2: Segment the Chinese part of the Chinese-to-Myanmar parallel corpus and count the word frequencies, remove stop words and words with low word frequencies, and construct the root node of the tree-shaped retrieval structure according to the occurrence frequencies of Chinese words and phrases.

[0055] Further, the Step2 includes:

[0056] Step2.1. Use the jieba word segmentation tool to segment the Chinese part of the Chinese-to-Myanmar parallel corpus and count the word frequencies, sort the counted words and phrases according to the word frequencies to construct a word list, use the Chinese stop word list of Harbin Institute of Technology to clean the Chinese general words in the word list and remove the noise content, and remove the last 5% of the words with low word frequencies to avoid the influence of the word list long-tail effect on the retrieval efficiency of the final tree-shaped retrieval structure; the noise content includes blank lines and special symbols.

[0057] Step2.2. Use the MongoDB non-relational database to save the word list information as the root node of the tree-shaped retrieval structure, and the word list information includes words and phrases, word frequencies, and word indexes.

[0058] Step3. Divide the Chinese-to-Myanmar parallel corpus into long sentence pairs and short sentence pairs according to the sentence length, and the ratio of long sentence pairs to short sentence pairs is 1:3;

[0059] Among them, the short sentence pairs are matched with the words and phrases of the root node with reference to the TF-IDF algorithm (Term Frequency-Inverse Document Frequency) and the BM25 (Best Matching 25) algorithm, establish the connection between the short sentences and the words and phrases, and use the cross-lingual text representation tool to perform embedding representation on the inserted short sentence pairs and save them to complete the construction of the preliminary tree-shaped retrieval structure;

[0060] Further, the Step3 includes:

[0061] Step3.1. Divide the collected Chinese-to-Myanmar parallel corpus into long sentence pairs and short sentence pairs according to the sentence length, and the ratio of long sentence pairs to short sentence pairs is 1:3;

[0062] Step3.2. Filter and score the Chinese part of the short sentence pairs with the words and phrases of the root node in the MongoDB non-relational database as the tree-shaped retrieval structure with reference to the TF-IDF algorithm and the BM25 algorithm in turn, determine the insertion position of the short sentence pairs in the retrieval tree and insert them into the MongoDB non-relational database.

[0063] Step 3.3. Use the LASER cross - language text representation model to perform cross - language embedding representation on the Chinese and corresponding Burmese translations of the inserted short sentence pairs, and update the content of the MongoDB non - relational database;

[0064] Furthermore, in the above Step 3, the TF - IDF algorithm (Term Frequency - Inverse Document Frequency) is a statistical method for measuring the importance of words in a piece of text; the TF - IDF algorithm is used in the construction stage of the tree - shaped retrieval structure of parallel corpora and the final hint retrieval stage. By calculating the importance scores of words in the tree - shaped retrieval structure, the hint sentence pairs most relevant to the sentence pairs to be inserted and the sentences to be translated are screened; the TF - IDF algorithm screens the root node words of the currently saved tree - shaped retrieval structure to which the currently inserted sentence pairs and the sentences to be translated belong by calculating the relationship between the frequency of occurrence of words in a piece of text and the text length;

[0065] The formula definition of the TF - IDF algorithm is:

[0066]

[0067] where, f t,d represents the number of occurrences of word t in document d, N represents the total number of documents, n t is the number of documents containing word t. To improve the retrieval efficiency, the root node words of the tree - shaped retrieval structure are regarded as documents, that is, the value of N is the number of screened words, and the value of n t is 1 or 0. Such a design can speed up the calculation speed and improve the construction and retrieval efficiency of the tree - shaped retrieval structure; TF is the term frequency, representing the frequency of occurrence of a certain word in a document, reflecting the local importance of the word in the sentence; IDF represents the inverse document frequency, which measures the universality of words in the entire corpus, used to reduce the weight of high - frequency general words and increase the weight of words with high distinctiveness;

[0068] The BM25 (Best Matching 25) algorithm is an improved probabilistic retrieval model, widely used in information retrieval and document relevance ranking tasks; the BM25 algorithm is used to score the sentence pairs in the retrieval tree to further optimize the accuracy of sentence matching; the core idea of the BM25 algorithm is to dynamically adjust the weight according to the term frequency, document length, and word distribution. Compared with TF - IDF, it performs better in dealing with the problem of uneven distribution of long and short sentences; the formula definition of the BM25 algorithm is:

[0069]

[0070] Among them, D represents the sentence to be inserted currently, Q represents the words and phrases of the screening result in the previous step, and f t,D represents the word frequency of word t in the sentence to be inserted currently. The value of M is 1, indicating the number of sentences to be inserted currently; n t takes a value of 1 or 0, indicating whether the sentence to be inserted currently contains word t; |D| is the number of words and phrases in the sentence to be inserted currently, and avgdl is the average word and phrase length after word segmentation of the sentence to be inserted currently; k1 and k are adjustment parameters. Among them, the value range of k1 is usually in [0.5, 2]. The larger it is, the more the algorithm depends on word frequency. The value range of k is in [0, 1]. The larger k is, the more the algorithm depends on the fully normalized document length. In the present invention, k1 = 2.0 and k = 0.5.

[0071] Further, the Step 3.3 includes:

[0072] LASER (Language-Agnostic Sentence Representations) is a multilingual sequence-to-sequence sentence embedding model. Use the LASER cross-lingual text representation model to perform text representation on the short sentence pair at the determined insertion position; input the text into the LASER cross-lingual text representation model, first enter the 5-layer bidirectional long short-term memory network BiLSTM (Bi-directional Long Short-Term Memory), and convert Chinese and Burmese into x1, x2....x i1 , (i1 = 1024) embedding representations; then input into the Decoder layer. The Decoder layer is jointly trained in multiple languages. After determining the generated language, the Decoder layer converts the embedding representation input from the upper layer into a [1, 1024]-dimensional sentence embedding representation vector through a loss function (softmax function). Since LASER can map sentence pairs in any language to close positions in a high-dimensional space, the outer product is directly removed as part of the sentence score during the score calculation process; after obtaining the sentence embedding representation of the bilingual text, save the sentence embedding representation of the bilingual text, the original text information, and the word and phrase information determined after screening in the MongoDB non-relational database.

[0073] Step 4: For the processing of long sentence pairs, use the LASER cross-lingual representation tool to perform embedding representation on the sentences, screen the short sentences related to the long sentence in the retrieval tree through the TF-IDF algorithm and the BM25 algorithm, and then calculate the cosine similarity between the short sentences and the long sentence in the source language and the target language to determine the insertion position of the long sentence pair in the tree-shaped retrieval structure;

[0074] Further, the said Step4 includes: processing the divided long sentence pairs, using a cross-lingual text representation tool to perform embedding representation on the Chinese and corresponding Burmese translations in each group of long sentence pairs, screening the root nodes of the tree-shaped retrieval structure and their associated short sentence pairs through the TF-IDF algorithm and the BM25 algorithm, calculating the cross-cosine similarity between the current long sentence pair and the screened short sentence pairs to generate a similarity matrix, and regularizing the similarity matrix through the Frobenius norm to obtain the final similarity score; determining the insertion position of the long sentence pairs in the tree-shaped retrieval structure according to the order of the similarity scores from high to low, thereby completing the insertion of the long sentence pairs into the tree-shaped retrieval structure.

[0075] Further, the specific steps of the said Step4 include:

[0076] Step4.1: Using the LASER cross-lingual text representation model to perform cross-lingual representation on the source language and target language of the long sentence pairs;

[0077] Step4.2: Sequentially using the TF-IDF algorithm and the BM25 algorithm to filter and score the words, and retaining the top 20 words with the highest scores and their affiliated short sentence pairs;

[0078] Step4.3: Using the cosine similarity to sequentially calculate the similarity scores between the source language and target language of the short sentence pairs and the long sentence pairs to obtain a similarity matrix, taking the Frobenius norm of the similarity matrix as the final similarity score of the current long sentence pair to be inserted, determining the position of the long sentence pair in the retrieval tree according to the similarity score from high to low, and inserting it into the MongoDB non-relational database to complete the insertion of the long sentence pairs into the tree-shaped retrieval structure.

[0079] Further, in the said Step4, the cosine similarity can calculate the semantic similarity between two texts after high-dimensional representation; if the semantics between the two texts are similar, the final result approaches 1, otherwise it approaches 0; the cosine similarity algorithm is defined as:

[0080]

[0081] where Sim(X,Y) represents the final score of the text similarity calculation, X and Y respectively represent the high-dimensional embedding vectors of the texts participating in the cosine similarity calculation, X i and Y i respectively represent an element in the vector, and n represents the total number of elements;

[0082] The Frobenius norm is a measure of a matrix, used to measure the square root of the sum of the squares of all elements in the matrix; in the stage of confirming the insertion position of the bilingual long sentence pair in the tree retrieval structure, the LASER cross-lingual text representation model is used to obtain the embedding representation of the bilingual text of the long sentence pair, and the embedding representation of the bilingual text of the long sentence pair is cross-calculated with the bilingual short sentence pairs and the embedding representations of the bilingual short sentence pairs screened from the database to form a similarity matrix S. The similarity matrix S is defined as:

[0083]

[0084] where x s and y s represent the Chinese and Burmese bilingual short sentence pairs that have been inserted into the database, and x l and y l respectively represent the Chinese and Burmese bilingual long sentence pairs whose insertion positions in the tree retrieval structure are to be determined currently;

[0085] The definition of taking the Frobenius norm of the similarity matrix S is:

[0086] [[ID=z0]]

[0087] where Score(S) represents the final score of the sentence pair to be inserted, and ||S|| F represents the calculation of taking the Frobenius norm of the matrix, and s ab is the corresponding sentence pair score after cross-cosine similarity calculation in the ath row and bth column of the matrix S;

[0088] After obtaining the final similarity score, record the information of the bilingual short sentence pair with the highest score, as well as the original text and the corresponding text embedding representation of the current bilingual long sentence pair, and insert them into the MongoDB non-relational database. The similarity calculation method adopted by the present invention can ensure that there is a strong similarity relationship between the bilingual long sentence pair finally inserted into the tree retrieval structure and its associated short sentence pair, and ensure the accuracy and retrieval efficiency of the final retrieval result.

[0089] Step5. Retrieval stage: Segment the text to be translated and sequentially match the root nodes of the tree retrieval structure according to the segmentation results, and screen the relevant candidate trees for the words appearing in the text to be translated; secondly, use the cross-lingual text representation tool to obtain the embedding representation of the text to be translated, and use the embedding representation of the text to be translated to perform cosine similarity calculations with the Chinese and Burmese embedding representations of the long and short sentence pairs in the candidate trees in turn, and take the average value of the cosine similarity scores of each pair of candidate sentence pairs as the final relevance score, and sort them from high to low according to the relevance score as the context prompt sentence pairs for the translation task;

[0090] Further, the said Step5 includes:

[0091] Step5.1. Tokenize the text to be translated and use the LASER cross-lingual text representation model to obtain the embedding representation of the text to be translated;

[0092] Step5.2. For the tokenized text to be translated, use the TF-IDF algorithm and the BM25 algorithm in sequence to filter and select matching words in the tree retrieval structure, and retain the words with a BM25 algorithm score greater than 0.0 and their corresponding long and short sentence pairs as candidate sentence pairs for context prompts in the large language model Chinese-Myanmar machine translation;

[0093] Step5.3. Calculate the cosine similarity between the embedding representation of the text to be translated and the embedding representations of the filtered long and short sentence pairs respectively, and take the average of the calculation results as the final cosine similarity score of the sentence pairs. Sort the sentence pairs in descending order of similarity score. After removing the candidate sentence pairs with a score less than 0.3, the remaining long and short sentence pairs are used as context prompt sentence pairs for the translation task.

[0094] Step6. Before using the large language model for the final translation, use a text re-ranking model to re-rank the retrieved context prompt sentence pairs, and input the context prompt sentence pairs and the text to be translated into the large language model to obtain the final translation result.

[0095] Further, the said Step6 includes:

[0096] Step6.1. Use the bge-reranker-v2-m3 model to re-rank the finally obtained context prompt sentence pairs, and sort the prompt sentence pairs in descending order of cosine similarity score;

[0097] Step6.2. Determine the final number of prompt sentence pairs according to the number of prompt examples, and input the prompt sentence pairs and the sentence to be translated into the large language model to obtain the final translation result.

[0098] To verify the effectiveness of the method of the present invention, the Qwen2.5-7B-Instruct model and the Llama3-8B-Instruct model were selected for final translation in the experimental stage. Both of these models are representatives in the field of current large language models, with powerful text generation and understanding capabilities, and showing significant advantages in low-resource machine translation tasks. Qwen2.5-7B-Instruct is a large model focused on multilingual instruction tasks, performing excellently in multilingual pre-training and instruction fine-tuning, and being able to quickly adapt to translation tasks in different language scenarios. Thanks to its large-scale pre-training on multilingual corpora, Qwen2.5-7B-Instruct has high robustness in dealing with complex language structures and long sentence translations. Llama3-8B-Instruct is another high-performance language model, performing well in dealing with low-resource language tasks. With a larger number of parameters and a deeply optimized instruction tuning strategy, Llama3-8B-Instruct demonstrates good capabilities in translation quality, language fluency, and semantic consistency, especially having obvious advantages in translation scenarios with complex context information. By comparing the translation effects of these two models, the actual performance of the method proposed in this paper in low-resource translation tasks can be comprehensively evaluated, providing an important reference for subsequent research.

[0099] In the experimental data collection stage, to verify the effectiveness of the method of the present invention, the present invention respectively constructed parallel corpora in the Chinese-Vietnamese and Chinese-Myanmar directions. Among them, the scale of the Chinese-Vietnamese parallel corpus reached more than 16 million, and the scale of the Chinese-Myanmar parallel corpus was more than 900,000. These corpora cover a variety of text types, including news reports, educational materials, and web documents, etc., aiming to reflect the diversity and complexity in actual translation scenarios to the greatest extent. To ensure the quality of the parallel corpora, the present invention carried out systematic cleaning and preprocessing on the parallel corpora, respectively filtering parallel corpora with large length differences, too long and too short ones, and at the same time removing non-UTF-8 encoded characters in the corpora to reduce the influence of noise on the experimental results.

[0100] In the experimental stage, the present invention conducts experimental verification based on the FLORES-200 multilingual evaluation dataset. This dataset covers 200 language pairs and provides a standardized benchmark for cross-lingual translation ability assessment. In the experiment, the present invention dynamically selects five groups of cross-lingual parallel examples (5-shots) through the above retrieval enhancement method to construct a prompt template, guiding the large language model to capture the translation rules of low-resource languages. To comprehensively verify the effectiveness of the method of the present invention, a double control group is set up in the experiment: one is the control of the original zero-shot (0-shot) translation ability of the large language model, and the other is to select the pre-trained multilingual machine translation model NLLB-200-distilled-600M as the control baseline. The latter is based on the encoder-decoder architecture and supports direct mutual translation between more than 200 languages through the language adaptive parameter sharing mechanism.

[0101] In the design of the evaluation framework, the experiment uses spBLEU (SentencePiece-BLEU) as the core indicator. This indicator significantly alleviates the evaluation bias problem caused by traditional BLEU indicators based on space or rule-based word segmentation for low-resource languages by uniformly applying the SentencePiece subword segmentation technology.

[0102] Table 1 shows the comparison of spBLEU scores between the method of the present invention and the baseline method in the experimental results

[0103]

[0104] It can be seen from the experimental results in Table 1 that by introducing a retrieval mechanism that combines algorithms such as TF-IDF and BM25 with the Frobenius norm, as well as the cross-lingual semantic embedding of the LASER model, the accuracy and semantic relevance of prompt sentence screening are effectively improved, and finally the overall machine translation performance of the large language model in low-resource languages is improved. In addition, aiming at the characteristics of uneven quality of low-resource corpora, the present invention significantly improves the translation generation quality through the optimization strategy of prompt sentence reordering and the strengthening of context association. The method of the present invention not only provides an efficient and feasible solution for low-resource language translation, but also further expands the application potential of large models in multilingual translation tasks.

[0105] The specific implementation manners of the present invention have been described in detail above with reference to the accompanying drawings. However, the present invention is not limited to the above implementation manners, and various changes can be made without departing from the spirit of the present invention within the scope of knowledge possessed by those of ordinary skill in the art.

Claims

1. A Chinese-Myanmar machine translation method for large models based on language knowledge retrieval, characterized in that: The method includes: Step 1: Collect a Chinese-to-Burmese parallel corpus dataset; Step 2: Perform word segmentation and word frequency statistics on the Chinese part of the Chinese-to-Burmese parallel corpus, remove stop words and words with low word frequencies, and construct the root node of the tree retrieval structure according to the occurrence frequency of Chinese words; Step 3: Divide the Chinese-to-Burmese parallel corpus into long sentence pairs and short sentence pairs according to the sentence length; Among them, the short sentence pairs are matched with the words at the root node with reference to the TF-IDF algorithm and the BM25 algorithm, establish the connection between the short sentences and the words, and use the cross-lingual text representation tool to perform embedding representation on the inserted short sentence pairs and save them to complete the preliminary construction of the tree retrieval structure; Step 4: For the processing of long sentence pairs, use the LASER cross-lingual representation tool to perform embedding representation on the sentences, screen the short sentences related to the long sentences in the retrieval tree through the TF-IDF algorithm and the BM25 algorithm, and then calculate the cosine similarity between the short sentences and the long sentences in the source language and the target language to determine the insertion position of the long sentence pair in the tree retrieval structure; Step 5: In the retrieval stage, perform word segmentation on the text to be translated and sequentially match the root nodes of the tree retrieval structure according to the word segmentation results, and screen the relevant candidate trees for the words that appear in the text to be translated; secondly, use the cross-lingual text representation tool to obtain the embedding representation of the text to be translated, and calculate the cosine similarity between the embedding representation of the text to be translated and the Chinese and Burmese embedding representations of the long and short sentence pairs in the candidate trees in turn. Take the average value of the cosine similarity scores of each pair of candidate sentence pairs as the final relevance score, and sort them from high to low according to the relevance score as the context prompt sentence pairs for the translation task; Step 6: Before using the large language model for the final translation, use the text re-ranking model to re-rank the retrieved context prompt sentence pairs, and input the context prompt sentence pairs and the text to be translated into the large language model to obtain the final translation result.

2. The method for Chinese-Myanmar machine translation of large models based on language knowledge retrieval according to claim 1, characterized in that: The said Step 1 includes: Collect the Chinese-to-Burmese parallel corpus from the OPUS public parallel corpus website.

3. The method for Chinese-Myanmar machine translation of large models based on language knowledge retrieval according to claim 1, wherein: The said Step 2 includes: Step 2.1: Use the jieba word segmentation tool to perform word segmentation and count the word frequencies on the Chinese part of the Chinese-to-Burmese parallel corpus, sort the counted words according to the word frequencies to construct a word list, clean the Chinese general words in the word list and remove the noise content, and remove the last 5% of the words with low word frequencies to avoid the influence of the word list long-tail effect on the retrieval efficiency of the final tree retrieval structure; the noise content includes blank lines and special symbols; Step 2.2: Use the MongoDB non-relational database to save the word list information as the root node of the tree retrieval structure, and the word list information includes words, word frequencies, and word indexes.

4. The method for Chinese-Myanmar machine translation of a large model based on language knowledge retrieval according to claim 1, wherein: The said Step 3 includes: Step 3.1: Divide the collected Chinese-to-Burmese parallel corpus into long sentence pairs and short sentence pairs according to the sentence length, and the ratio of long sentence pairs to short sentence pairs is 1:3; Step3.

2. Filter and score the Chinese part of the short sentence pair and the root node words in the MongoDB non-relational database as the tree-shaped retrieval structure root node words in turn with reference to the TF-IDF algorithm and the BM25 algorithm, determine the insertion position of the short sentence pair in the retrieval tree and insert it into the MongoDB non-relational database; Step3.

3. Use the LASER cross-lingual text representation model to perform cross-lingual embedding representation on the Chinese and corresponding Burmese translations of the inserted short sentence pair, and update the content of the MongoDB non-relational database.

5. The method for Chinese-Myanmar machine translation of a large model based on language knowledge retrieval according to claim 4, wherein: The said Step3.3 includes: Use the LASER cross - language text representation model to represent the short sentence pairs for determining the insertion positions; input the text into the LASER cross - language text representation model, first enter the 5 - layer bidirectional long short - term memory network BiLSTM, and convert Chinese and Burmese into x1, x2....x i1 , (i1 = 1024) embedding representation; then input it into the Decoder layer. After determining the generated language, the Decoder layer converts the embedding representation input from the upper layer into a [1, 1024] - dimensional sentence embedding representation vector through the loss function; after obtaining the sentence embedding representation of the bilingual text, save the sentence embedding representation of the bilingual text, the original text information, and the determined word information after screening in the MongoDB non - relational database.

6. The method for Chinese-Myanmar machine translation of large models based on language knowledge retrieval according to claim 1, characterized in that: The said Step4 includes: Process the divided long sentence pairs, use the cross-lingual text representation tool to perform embedding representation on the Chinese and corresponding Burmese translations in each group of long sentence pairs, screen the root node of the tree-shaped retrieval structure and its associated short sentence pairs through the TF-IDF algorithm and the BM25 algorithm, use the cross-calculation of the cosine similarity between the current long sentence pair and the screened short sentence pairs to generate a similarity matrix, and regularize the similarity matrix through the Frobenius norm to obtain the final similarity score; According to the order of the similarity scores from high to low, determine the insertion position of the long sentence pair in the tree-shaped retrieval structure, thus completing the insertion of the long sentence pair in the tree-shaped retrieval structure.

7. The method for Chinese-Myanmar machine translation of a large model based on language knowledge retrieval according to claim 1, characterized in that: The specific steps of the said Step4 include: Step4.

1. Use the LASER cross-lingual text representation model to perform cross-lingual representation on the source language and target language of the long sentence pair; Step4.

2. Filter and score the words in turn with the TF-IDF algorithm and the BM25 algorithm, and retain the top 20 words with the highest scores and the short sentence pairs to which they belong; Step4.

3. Use the cosine similarity to cross-calculate the similarity scores of the source language and target language in the short sentence pairs and long sentence pairs in turn to obtain a similarity matrix, take the Frobenius norm of the similarity matrix as the final similarity score of the current long sentence pair to be inserted, determine the position of the long sentence pair in the retrieval tree according to the similarity score from high to low, and insert it into the MongoDB non-relational database to complete the insertion of the long sentence pair in the tree-shaped retrieval structure.

8. The method for Chinese-Myanmar machine translation of a large model based on language knowledge retrieval according to claim 1, characterized in that: In the said Step4, the said cosine similarity can calculate the semantic similarity between two texts after high-dimensional representation; if the semantics between two texts are similar, the final result approaches 1, otherwise it approaches 0; the cosine similarity algorithm is defined as: Among them, Sim(X, Y) represents the final score of text similarity calculation, where X and Y respectively represent the high-dimensional embedding vectors of the texts participating in the cosine similarity calculation, and X i and Y i respectively represent an element in the vector, and n represents the total number of elements; In the stage of confirming the insertion position of the bilingual long sentence pair in the tree-shaped retrieval structure, use the LASER cross-lingual text representation model to obtain the embedding representation of the bilingual long sentence pair text, cross-calculate the cosine similarity scores between the embedding representation of the bilingual long sentence pair text and the bilingual short sentence pairs and bilingual short sentence pair embedding representations screened from the database to form a similarity matrix S, and the similarity matrix S is defined as: where x s and y s represent the Chinese and Burmese bilingual short sentence pairs that have been inserted into the database, and x l and y l represent the Chinese and Burmese bilingual long sentence pairs for which the insertion positions in the current tree-shaped retrieval structure are to be determined, respectively; The definition of taking the Frobenius norm of the similarity matrix S is: Among them, Score(S) represents the final score of the sentence pair to be inserted, ||S|| F represents the calculation of the Frobenius norm of the matrix, s ab is the corresponding sentence pair score after cross-cosine similarity calculation for the element in the ath row and bth column of the matrix S; After obtaining the final similarity score, record the information of the bilingual short sentence pair with the highest score and the original text and corresponding text embedding representation of the current bilingual long sentence pair, and insert it into the MongoDB non-relational database.

9. The Chinese-Myanmar machine translation method for large models based on language knowledge retrieval according to claim 1, wherein: The said Step5 includes: Step5.

1. Tokenize the text to be translated and use the LASER cross-lingual text representation model to obtain the embedded representation of the text to be translated; Step5.

2. For the tokenized text to be translated, sequentially use the TF-IDF algorithm and the BM25 algorithm to filter and screen for matching words in the tree retrieval structure, and retain the words with a BM25 algorithm score greater than 0.0 and their corresponding long and short sentence pairs as candidate sentence pairs for context prompting in the large language model Chinese-Myanmar machine translation; Step5.

3. Calculate the cosine similarity between the embedded representation of the text to be translated and the embedded representations of the filtered long and short sentence pairs respectively, take the average of the calculation results as the final cosine similarity score of the sentence pairs, sort them from high to low according to the similarity score, and after removing the candidate sentence pairs with a score less than 0.3, use the remaining long and short sentence pairs as context prompting sentence pairs for the translation task.

10. The method for Myanmar-Chinese machine translation of a large model based on language knowledge retrieval according to claim 1, characterized in that: The said Step6 includes: Step6.

1. Use the bge-reranker-v2-m3 model to re-rank the finally obtained context prompting sentence pairs, and sort the prompting sentence pairs from high to low according to the cosine similarity score; Step6.

2. Determine the final number of prompting sentence pairs according to the number of prompting examples, and input the prompting sentence pairs and the sentence to be translated into the large language model to obtain the final translation result.