Personalized translation engine based on large language model, translation text generation method

By using an auxiliary model with fewer than 1 billion parameters to extract labels and combining it with the main model for translation, the problems of high cost and poor timeliness in existing technologies are solved, enabling the rapid generation of high-quality translation engines and improving translation efficiency and accuracy.

CN117151124BActive Publication Date: 2026-01-09BESTEASY (BEIJING) TRANSLATION CORP
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310994937.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-08-09
Publication Date
2026-01-09
Estimated Expiration
2043-08-09

AI Technical Summary

Technical Problem

Existing technologies for machine translation in specific fields suffer from high costs, poor timeliness, and the need for incremental training with large amounts of data and time, making it impossible to quickly respond to the needs of translation projects.

Method used

Auxiliary models with fewer than 1 billion parameters are used to extract labels. The optimal set of instructions is obtained through pre-training, evaluation and fine-tuning. The results are then combined with the main model for translation, and a personalized translation engine is used to quickly generate high-quality translations.

Benefits of technology

It enables the generation of personalized translation engines in a short time, improving translation efficiency and accuracy, simplifying the development process for algorithm engineers, and reducing costs.

✦ Generated by Eureka AI based on patent content.
Patent Text Reader

Abstract

The application belongs to the technical field of automatic translation of texts, and discloses a personalized translation engine based on a large language model and a method for generating translated texts. The method comprises the following steps: obtaining a file to be translated; obtaining an auxiliary model; extracting a label from the file to be translated by using the auxiliary model, obtaining an optimal instruction set; and taking the optimal instruction set as a carrier of the personalized translation engine. The translation engine generated by the method for generating a personalized translation engine based on a large language model is used to translate a translation project file. The construction of the personalized translation engine is simpler and faster than that of a traditional method, and the overall translation accuracy is effectively improved by using the personalized translation engine for translation.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of automatic translation of text, and in particular to a personalized translation engine based on a large language model and a method for generating translated text. BACKGROUND

[0002] In the field of machine translation, large language models have become the current mainstream research direction due to their cross-domain capabilities and emergent capabilities, but the specific industry application direction is still unclear.

[0003] After professional field data incremental training, neural network machine translation can greatly improve translation quality in the professional field. However, the incremental training method has problems such as high cost, poor timeliness, large amount of required data, and field limitations.

[0004] When translating a specific field, this method requires algorithm engineers to design and develop related algorithms for a long time, and a large amount of training data (several thousand sentence pairs or more for adjacent fields, and more for non-adjacent fields) is required. After a period of training, the trained model can be obtained, which cannot meet the timeliness and low-cost requirements of translation projects.

[0005] Therefore, how to generate a corresponding personalized translation engine in a short time according to an actual translation project requirement is particularly important for improving translation efficiency. SUMMARY

[0006] The embodiments of the present application provide a personalized translation engine based on a large language model and a method for generating translated text.

[0007] The personalized translation engine generation method based on a large language model includes the following steps:

[0008] Obtain a file to be translated;

[0009] Obtain an auxiliary model; wherein the parameter amount of the auxiliary model is less than 1 billion;

[0010] The auxiliary model extracts a label from the file to be translated to obtain an optimal instruction set;

[0011] The optimal instruction set is used as a carrier of the personalized translation engine.

[0012] After obtaining the auxiliary model, the auxiliary model is further trained, and the training includes:

[0013] Pre-training;

[0014] Evaluation fine-tuning; the parameters of the evaluation fine-tuning include translation quality, instruction evaluation, text analysis, conversion of labels to instructions, and automatic post-editing.

[0015] The auxiliary model extracts tags from the file to be translated, and obtaining the optimal instruction set comprises,

[0016] The auxiliary model extracts domain tags from the file to be translated.

[0017] In the memory bank, search and extract historical tags according to the file to be translated.

[0018] In the term bank, search and extract term tags according to the file to be translated.

[0019] The domain tags, historical tags, and term tags are combined to form an instruction set.

[0020] Iterative instruction optimization is performed on the instruction set to obtain an optimal instruction set.

[0021] The iterative instruction optimization on the instruction set to obtain an optimal instruction set comprises:

[0022] Trial translation is performed on the file to be translated.

[0023] The quality of the trial translation is evaluated.

[0024] If there are unqualified tags, modify the unqualified tags in the instruction set.

[0025] If there are no unqualified tags, the current instruction set is taken as the optimal instruction set.

[0026] The personalized translation text generation method based on a large language model uses the translation engine generated by the personalized translation engine generation method based on a large language model to translate the translation project file.

[0027] The translation of the translation project file comprises the following steps:

[0028] Create a translation project, wherein the translation project includes a file to be translated.

[0029] Obtain a main model.

[0030] According to the main model and the translation engine, output the translation text corresponding to the file to be translated.

[0031] The main model is a model pre-trained with world knowledge and fine-tuned for translation ability.

[0032] After outputting the translation text corresponding to the file to be translated according to the main model and the translation engine, obtain translation feedback and optimize the translation engine according to the translation feedback.

[0033] The embodiment of the application provides a personalized translation engine based on a large language model, a translation text generation method, an auxiliary model smaller than a main model is adopted to extract a label of a to-be-translated file to obtain an optimal instruction set, so that the personalized translation engine is generated. Then, the entire to-be-translated file is translated by the main model combined with the translation engine. The above personalized translation engine is simple and fast to construct, and the overall translation accuracy is effectively improved by using the personalized translation engine for translation. DETAILED DESCRIPTION

[0034] To make the objectives, technical solutions and advantages of the embodiments of the application clearer, the technical solutions in the embodiments of the application are described clearly and completely. Obviously, the described embodiments are some embodiments of the application, rather than all the embodiments. Based on the embodiments in the application, all other embodiments obtained by those skilled in the art without creative work fall within the protection scope of the application.

[0035] The embodiment of the application provides a personalized translation engine generation method based on a large language model, which comprises the following steps:

[0036] Obtaining a to-be-translated file;

[0037] Obtaining an auxiliary model; the parameter quantity of the auxiliary model is much smaller than that of the main model, and preferably, the parameter quantity of the auxiliary model is less than 1 billion. Only when the auxiliary model is small enough, the personalized engine generation can be more lightweight.

[0038] Training the auxiliary model, which can be divided into two stages, comprising:

[0039] Pre-training, the auxiliary model can be trained on a large amount of text.

[0040] Evaluating fine-tuning; preset text and corresponding requests for evaluating translation quality are used as parameters for evaluating fine-tuning, specifically including: translation quality, instruction evaluation, text analysis, converting labels into instructions, and automatic post-editing.

[0041] Translation quality: the auxiliary model outputs basic features of the translation quality of the auxiliary model, including translation errors, omissions, translation failures, and normal.

[0042] As follows:

[0043] Please evaluate the following translation quality:

[0044] Original text: Beijing continues to monitor PM2.5

[0045] Translated text: Beijing is controlling PM2.5

[0046] Evaluation: Translation error

[0047] Please evaluate the translation quality as follows:

[0048] Original: Beijing continues to monitor PM2.5

[0049] Translation: Beijing continues to monitor PM2.5

[0050] Evaluation: Translation failure

[0051] Please evaluate the translation quality as follows:

[0052] Original: The team won first place in two major translation projects, including Chinese-English translation

[0053] Translation: The team won first place in two major translation projects, including Chinese-English translation

[0054] Evaluation: Omission {two}

[0055] Instruction evaluation: Auxiliary model R needs to evaluate the degree to which the translation complies with the translation instructions, and return normally when the original translation request does not exist instructions.

[0056] As follows:

[0057] Please evaluate the degree to which the translation complies with the instructions

[0058] Instructions: Oral, adverbial preposition

[0059] Original: The team won first place in two major translation projects, including Chinese-English translation

[0060] Translation: They won first place in two major translation tasks, including Chinese to English

[0061] Evaluation: Omission {adverbial preposition}

[0062] Please evaluate the degree to which the translation complies with the instructions

[0063] Instructions: (empty)

[0064] Original: The team won first place in two major translation projects, including Chinese-English translation

[0065] Translation: They won first place in two major translation tasks, including Chinese to English

[0066] Evaluation: Normal

[0067] Please evaluate the adherence to the translation instructions

[0068] Instructions: Adverbial fronting, past tense

[0069] Original text: There was a row of short houses on the edge of the lake.

[0070] Translation: There was a row of short houses on the edge of the lake.

[0071] Evaluation: Omission {Adverbial fronting}

[0072] Text analysis: Identify the industry, domain, and genre of a given text or pair of bilingual texts.

[0073] As follows:

[0074] Text: The company has been committed to using the power of technology to work together with users to achieve "the joy of human culture." The company's operating system believes that the mobile operating system is responsible for connecting the digital world and the physical world, carrying many technological innovations and expectations.

[0075] Analysis: Industry = Technology, Domain = Technology News Release, Genre = News

[0076] Please analyze the features {Genre} of the following text or text pair:

[0077] Text: This array of methods is passed down from Sunzi and has not changed for thousands of years.

[0078] Analysis: Genre = Classical Novel

[0079] Convert the label to instructions: Make the translation conform to the specified features.

[0080] As follows:

[0081] Original text: You can do tensor parallelism on two GPUs on a machine, but the effect may not be very good because your GPU memory and computing power are very limited.

[0082] Features: Technology domain, "you" as the subject

[0083] Instructions: English proper name duplication, proper name translation must be consistent

[0084] Original text: You can do tensor parallelism on two GPUs on a machine, but the effect may not be very good because your GPU memory and computing power are very limited.

[0085] We can use two GPUs of one node for Tensor Parallelism,but itmight not work,as your GPUs and memory are rather limited in terms ofcomputability.

[0086] Feature: tech field, "we" as subject

[0087] Instruction: proper name duplication, consistent translation of proper names, strong localization

[0088] Post-editing: for continuous optimization of the translation process, become part of the training data that provide auxiliary model translation understanding ability.

[0089] As follows:

[0090] Translate the following Chinese text into English

[0091] Instruction: written

[0092] Original text: They won the first place in two big translation tasks, including Chinese to English

[0093] Translation: They won the first place in two major translation tasks, including Chinese to English translation

[0094] Correction: The team won the first place in two major translation tasks, including Chinese to English translation

[0095] Comment (optional): The original translation is not written enough

[0096] The auxiliary model extracts tags from the file to be translated, and obtains the optimal instruction set; Specifically, it includes:

[0097] Extracting domain tags from the file to be translated: the auxiliary model analyzes the theme, industry, domain, etc. Information of the file to be translated, and extracts keywords about language features at the same time, calls the auxiliary model with different information combinations, and obtains the information and keywords by set, denoted as domain tag Tag1.

[0098] Search for similar sentence pairs in the memory bank according to the file to be translated, and extract historical tags;

[0099] In the terminology database, search for the terms contained in the original text according to the file to be translated, find the corresponding translation words in the field, extract the term tags, and record the set of historical tags and data tags as Tag2.

[0100] Convert Tag1 and Tag2 into translation instructions using the auxiliary model, and store them as an instruction set;

[0101] Iterative instruction optimization is performed on the instruction set to obtain the optimal instruction set. The optimal instruction set is used as a personalized translation engine carrier.

[0102] In instruction translation, the function insTrans(original text, language direction, instruction, *translation) is used; the final translation is stored in *translation. Translation can have various errors, including ignoring instructions; therefore, the function is designed to return a non-zero value when an error occurs, otherwise it returns 0.

[0103] When translating, for example:

[0104] 1. Given input original text S, language direction LP, instruction INS, translation pointer *T; M is the main model, and R is the auxiliary model.

[0105] 2. Let the translation T = M(S, LP, INS);

[0106] 3. Get the evaluation C = R(S, T, INS);

[0107] 4. If C contains translation errors or omissions, return 1;

[0108] 5. If C contains missing instructions, return 2;

[0109] 6. *T = T;

[0110] 7. Return 0;

[0111] The instructions required by the personalized translation engine are likely to be more than one, and requesting multiple instructions together in practice may have the following risks:

[0112] The main model M may miss some instructions (insTrans returns 2);

[0113] Multiple instructions are redundant or conflicting with each other;

[0114] Even if a single instruction is extracted for translation, some instructions will always be ineffective (insTrans returns 2);

[0115] Therefore, the following process is designed to iteratively find the optimal fine-tuning instructions. Try to translate the file to be translated;

[0116] Get the quality evaluation of the trial translation;

[0117] If there is a substandard label, modify the substandard label in the instruction set;

[0118] If there is no substandard label, the current instruction set is the optimal instruction set.

[0119] Specifically, the logic of the iterative instruction optimization (denoted as enhance (text pair, instruction set)) is as follows:

[0120] 1. Combine all possible instructions to get P;

[0121] 2. Instruction-missing level binary tuple list M;

[0122] 3. For each combination p in P:

[0123] Iterative translation is performed to obtain the translation T; if the translation is incorrect, T is empty;

[0124] Get instruction compliance evaluation C = R (S, T, INS);

[0125] Find the truly missing part in C, count the number, and mark it as m;

[0126] Push {p, m} into M;

[0127] 4. Sort M according to the size of each m in ascending order, and take the first one;

[0128] 5. If there is a tie for first place, the one with the highest instruction px diversity is the best;

[0129] 6. That is, if {colloquialization + oral expression, 1} and {colloquialization + multi-use present participle, 1} appear, the latter is better.

[0130] Finally, we get the optimal instruction set carrier of the personality engine Eng.

[0131] In one embodiment, custom instructions can also be added to the optimal instruction set to manually optimize the optimal instruction set.

[0132] The application also discloses a personalized translation text generation method based on a large language model, which uses a translation engine generated by the personalized translation engine generation method based on a large language model to translate a translation project file.

[0133] The translation of the translation project file includes the following steps:

[0134] Create a translation project; wherein the translation project includes a file to be translated;

[0135] Get the main model;

[0136] According to the main model and the translation engine, the translation text corresponding to the to-be-translated file is output; translation feedback is obtained, and the translation engine is optimized according to the translation feedback.

[0137] The main model is a model pre-trained with world knowledge and fine-tuned with translation ability, specifically as follows:

[0138] The main model can be a self-owned model trained by the provider of the machine translation, or can be publicly provided by other service providers; the personal engine acquisition and management method proposed in the present solution is not limited by the main model.

[0139] The following is a solution for training a main model by oneself. If an API of other service providers is selected, it needs to have similar functions, and the following requests are modified accordingly to adapt.

[0140] A single-decoder language model is trained on a large-scale corpus, which predicts the word token of the right position for each position of the input sequence.

[0141] The training is divided into three stages:

[0142] World knowledge pre-training: training on large-scale texts in various fields and languages. The result of this stage is denoted as PT.

[0143] Translation ability fine-tuning: training using translation request texts based on world knowledge pre-training. The translation request text is similar to the following; the result of this stage is denoted as FT.

[0144] Examples:

[0145] Please translate the original text from Chinese to English

[0146] Chinese: Beijing will continue to monitor PM2.5

[0147] English: Beijing continuously monitors the level of PM2.5

[0148] Instruction translation ability fine-tuning: training using expert-constructed instruction translation requests based on FT. The instruction translation request is similar to the following example; the result of this stage is the main model in the subsequent task.

[0149] Type 1: direct translation, denoted as main model (original text, language direction, instruction)

[0150] Please translate the original text from Chinese to English according to the translation instruction

[0151] Instruction: colloquial

[0152] Chinese: The company has been deeply engaged in artificial intelligence for 12 years

[0153] English: We have been working with AI for over a decade

[0154] Instructions: Direct translation

[0155] Instructions: Fuzzy number expression

[0156] Chinese: Our company has dug deeply in the AI for 12years

[0157] English: Our company has dug deeply in the AI for 12years

[0158] Type 2: Continue translation (modify based on existing translation according to instructions); this type is represented as M (original text, original translation, direction, instructions)

[0159] Instructions: Fuzzy number expression

[0160] Instructions: Fuzzy number expression

[0161] Chinese: Our company has dug deeply in the AI for 12years

[0162] Original translation: Our company has dug deeply in the AI for 12years

[0163] English: Our company has dug deeply in the AI for over a decade.

[0164] Finally, it should be noted that: the above examples are only used to illustrate the technical solutions of the present application, but not to limit it; although the present application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that: it can still modify the technical solutions recorded in the foregoing embodiments, or make equivalent replacement for part of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present application.

Claims

1. A method for generating a personalized translation engine based on a large language model, characterized by, The method comprises the following steps: obtaining a to-be-translated file; obtaining an auxiliary model; wherein the parameter quantity of the auxiliary model is less than 1 billion; the auxiliary model extracts labels from the to-be-translated file and obtains an optimal instruction set; the auxiliary model extracts labels from the to-be-translated file and obtains an optimal instruction set, which comprises, the auxiliary model extracts domain labels from the to-be-translated file; searching for and extracting historical labels from the to-be-translated file in a memory bank; searching for and extracting term labels from the to-be-translated file in a term bank; forming an instruction set from the domain labels, the historical labels, and the term labels; iteratively optimizing the instruction set to obtain an optimal instruction set; using the optimal instruction set as a personalized translation engine carrier; translating a translation project file, which comprises the following steps: creating a translation project; wherein the translation project includes a to-be-translated file; obtaining a main model; outputting a translated text corresponding to the to-be-translated file according to the main model and the translation engine. 2.The method of claim 1, wherein, After obtaining the auxiliary model, the method further comprises training the auxiliary model, which comprises: pre-training; evaluation fine-tuning; the parameters of the evaluation fine-tuning include translation quality, instruction evaluation, text analysis, conversion of labels to instructions, and automatic post-editing. 3.The method of claim 2, wherein, the iterative instruction optimization of the instruction set to obtain an optimal instruction set comprises: trial translation of the to-be-translated file; quality evaluation of the trial translation; if there are unqualified labels, modifying the unqualified labels in the instruction set; if there are no unqualified labels, using the current instruction set as the optimal instruction set.

4. A personalized translation text generation method based on a large language model, characterized by, using the translation engine generated by the personalized translation engine generation method based on a large language model according to any one of claims 1-3 to translate a translation project file. 5.The method of claim 4, wherein, The main model is a model pre-trained with world knowledge and fine-tuned for translation ability. 6.The method of claim 5, wherein, After outputting a translated text corresponding to the to-be-translated file according to the main model and the translation engine, the method further comprises obtaining translation feedback and optimizing the translation engine according to the translation feedback.

Citation Information

Patent Citations

  • Text feature extraction method based on machine learning

    CN112686039A

  • Machine translation style migration method and system based on curriculum pre-training

    CN115114940A