Multi-language localization agent based on large language model

Through multilingual localization agents based on large language models, the multilingual translation process is automated, and the high cost and low efficiency problems caused by relying on language experts in the existing technology are solved, and high-quality and low-cost multilingual translation and localization are achieved.

CN120373320APending Publication Date: 2025-07-25BESTEASY (BEIJING) TRANSLATION CORP
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202411023273.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-07-29
Publication Date
2025-07-25

AI Technical Summary

Technical Problem

Existing multilingual translation and localization workflows rely on language experts and have low automation, resulting in high costs and difficult to guarantee translation quality and consistency.

Method used

Multilingual localized agents based on large language models are adopted to realize automated and intelligent translation processes through steps such as project preparation, term extraction and filtering, glossary confirmation and interaction, personalized translation model generation, translation quality evaluation and model selection, preliminary translation and polishing, quality inspection and feedback, translation correction and re-checking, etc., and reduce manual intervention.

Benefits of technology

It improves translation efficiency and accuracy, reduces costs, ensures consistency and high quality of translations, and improves user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120373320A_ABST
    Figure CN120373320A_ABST
Patent Text Reader

Abstract

The invention belongs to the field of multilingual agents, and particularly relates to a multilingual localization agent based on a large language model, which comprises the following steps of: 1, project preparation and feature extraction; 2, terms are extracted and filtered; step 3, term table confirmation and interaction; 4, generating a personalized translation model; 5, translation quality evaluation and model selection; step 6, preliminarily translating and retouching; step 7, quality inspection and feedback; 8, correcting and re-checking the translated text; according to the method, the multi-language localization workflow is established based on the large-language model and the intelligent agents, different intelligent agents are defined through technologies such as cue words and model fine adjustment, and meanwhile, the overall workflow is automated through understanding, reasoning and planning capabilities of the large model; and low-cost and high-quality localization output can be realized, the manual intervention degree is reduced to a brand new level, and the global multi-language information and knowledge circulation efficiency is greatly improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of multilingual agents, and particularly to a multilingual localization agent based on a large language model. Background Art

[0002] The multilingual translation and localization workflow is a relatively complex process, mainly relying on language experts, and introducing automated steps such as machine translation and automated QA in some links. The premise assumption of this workflow is that AI can only handle single tasks, and language experts are required to plan, coordinate, give feedback, and check the entire project. The current two technical solutions are as follows:

[0003] One is to automatically typeset the document (after marking the paragraph format of the document through document parsing technology and sending it to machine translation) and automate machine translation (after the machine translation returns the translation, backfill the translation according to the original format). Users only need to upload the document and can download the document with paragraph-level translation.

[0004] One is a human-machine collaboration mode, that is, after uploading the document, the machine parses and pre-translates it, and language experts use tools to perform tasks in different links such as term extraction (term extraction model), term confirmation, post-editing, and QA inspection (text inspection rules and models), with language experts taking the lead.

[0005] The first technical solution realizes full automation. The disadvantages are that there are only two processes: format parsing and restoration and pre-translation, and there are problems and risks in aspects such as term unification, translation quality, and low error rate of the translation. Because the document only has certain reference reading value, the application scope is limited.

[0006] The second technical solution can ensure the unity of terms, high quality of the translation, and low error rate in line with quality specifications in the output translation document. However, the disadvantage is that the entire process requires the collaboration of language experts, mainly manual work, with high costs and limited delivery speed.

[0007] Therefore, based on practice, we propose a multilingual localization agent based on a large model, which can achieve a low-cost and high-quality localization process. Summary of the Invention

[0008] The multilingual localization agent based on a large language model proposed by the present invention solves the problems of the existing multilingual translation and localization workflow, which is a relatively complex process, mainly relying on language experts, and introducing automated steps such as machine translation and automated QA in some links. The premise assumption of this workflow is that AI can only handle single tasks, and language experts are required to plan, coordinate, give feedback, and check the entire project.

[0009] To achieve the above object, the present invention adopts the following technical solutions:

[0010] A multilingual localization agent based on a large language model, including the following steps:

[0011] Step 1: Project preparation and feature extraction;

[0012] Step 2: Term extraction and filtering;

[0013] Step 3: Glossary confirmation and interaction;

[0014] Step 4: Generation of personalized translation models;

[0015] Step 5: Translation quality evaluation and model selection;

[0016] Step 6: Preliminary translation and polishing;

[0017] Step 7: Quality inspection and feedback;

[0018] Step 8: Translation correction and re-inspection;

[0019] Step 9: Expert intervention and final delivery.

[0020] Preferably, in the above Step 1, project features are extracted according to documents such as document titles, tags, project information, and style guides, and the project workflow and model requirements are output, providing a basis for subsequent steps.

[0021] Preferably, in the above Step 2, term clusters are identified through term extraction models such as TF-IDF and TextRank, single-language document terms are extracted, and low-quality terms are filtered according to information such as term occurrence frequency.

[0022] Preferably, in the above Step 3, after the glossary is output, feedback interaction is carried out with external language experts to confirm the accuracy of the terms. After the language experts modify and confirm, proceed to the next step.

[0023] Preferably, in the above Step 4, a personalized translation model for the project is generated based on the RAG (Retrieval-Augmented Generation) technology and is used as one of the alternative models for the pre-translation process of the project.

[0024] Preferably, in the above Step 5, some document contents are randomly selected, translations are generated through multiple alternative models, and translation quality evaluation metrics such as BLEU, TER, METEOR, and BERT Score are used for scoring, and the model with the highest score is selected as the pre-translation model.

[0025] Preferably, in the sixth step, in combination with the glossary, the optimal pre - translation model is used for preliminary translation, and the polishing model polishes some translations according to requirements such as the project style guide. This link mainly focuses on the diction of the translation in terms of grammar, rhetoric, etc. in the target language, and optimizes the translation in terms of grammar, rhetoric, etc. through the polishing model.

[0026] Preferably, in the seventh step, a QA check is performed on the polished translation, including key contents such as punctuation marks, numbers, dates, etc., and the consistency of the glossary in the text is checked.

[0027] Preferably, in the eighth step, the translation is corrected according to the QA feedback, and the modified sentence segments are sent to the QA check model again for verification.

[0028] Preferably, in the ninth step, usually three rounds of interaction are set. When there are still risk items after three interactions, it is delivered to a language expert for final inspection and processing to generate the final multilingual localization document.

[0029] Compared with the prior art, the beneficial effects of the present invention are as follows:

[0030] Improve efficiency and accuracy: Through the automated process and the collaboration of multiple agents, the efficiency and accuracy of multilingual document translation and localization are greatly improved.

[0031] Ensure consistency: Through the unified management of the glossary and QA checks, the consistency of the translation in terms of term usage, style, etc. is ensured.

[0032] Flexibility and scalability: The system can be customized and extended according to the needs of different projects, supporting multiple languages and domains.

[0033] Reduce costs: Reduce manual intervention and lower the costs in the translation and localization process.

[0034] Enhance user experience: Through high - quality translations and consistent styles, the satisfaction and trust of users in multilingual content are enhanced.

[0035] The present invention constructs a multilingual localization workflow based on large - language models and agents, defines different agents through techniques such as prompts and model fine - tuning, and at the same time automates the overall workflow through the understanding, reasoning, and planning capabilities of the large model. It can achieve low - cost and high - quality localization output, reduce the degree of manual intervention to a new level, and greatly improve the efficiency of global multilingual information and knowledge circulation. Brief Description of the Drawings

[0036] Figure 1 It is a flowchart of the multilingual localization agent based on the large - language model proposed by the present invention. Detailed Embodiments

[0037] The technical solutions in the embodiments of the present invention will be clearly and completely described below. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments.

[0038] Referring Figure 1 , the multi - language localization intelligent agent based on the large - language model includes the following steps:

[0039] Step 1: Project preparation and feature extraction;

[0040] Step 2: Term extraction and filtering;

[0041] Step 3: Glossary confirmation and interaction;

[0042] Step 4: Generation of personalized translation models;

[0043] Step 5: Translation quality assessment and model selection;

[0044] Step 6: Preliminary translation and polishing;

[0045] Step 7: Quality inspection and feedback;

[0046] Step 8: Translation correction and re - inspection;

[0047] Step 9: Expert intervention and final delivery.

[0048] In this embodiment, in the above - mentioned Step 1, according to documents such as document titles, tags, project information, and style guides, project features are extracted, and the project workflow and model requirements are output, providing a basis for subsequent steps;

[0049] To improve efficiency and accuracy, natural language processing technology (NLP) can be further introduced to automatically extract and understand the content of these documents, improving efficiency and accuracy.

[0050] In this embodiment, in the above - mentioned Step 2, term clusters are identified through term extraction models such as TF - IDF and TextRank, and monolingual document terms are extracted. Low - quality terms are filtered according to information such as the frequency of term occurrence;

[0051] To improve the accuracy and professionalism of the glossary, a domain knowledge base or an expert system can be considered to assist in the identification and confirmation of terms.

[0052] In this embodiment, in the above - mentioned Step 3, after the glossary is output, feedback interaction is carried out with external language experts to confirm the accuracy of the terms. After the language experts modify and confirm, the next step is entered;

[0053] To accelerate the confirmation process, an online collaboration platform can be developed to enable language experts to provide real - time feedback and modify the glossary.

[0054] In this embodiment, in step four, a personalized translation model for the project is generated based on the RAG (Retrieval-Augmented Generation) technology and used as one of the alternative models for the pre-translation process of the project;

[0055] To further improve the performance and adaptability of the translation model, more advanced generative AI technologies such as the GPT series of models can be explored.

[0056] In this embodiment, in step five, some document content is randomly selected, translations are generated through multiple alternative models, and translation quality evaluation metrics such as BLEU, TER, METEOR, and BERT Score are used for scoring. The model with the highest score is selected as the pre-translation model;

[0057] To enable it to more accurately reflect the translation quality, machine learning algorithms can be introduced to optimize the scoring model.

[0058] In this embodiment, in step six, in combination with the glossary, the optimal pre-translation model is used for preliminary translation, and the polishing model polishes some translations according to requirements such as the project style guide. This link mainly focuses on the diction of the translation in terms of grammar, rhetoric, etc. in the target language, and optimizes the translation in terms of grammar, rhetoric, etc. through the polishing model;

[0059] The polishing model can further integrate domain-specific language rules and style guides to generate translations that are more in line with the target language habits.

[0060] In this embodiment, in step seven, a QA check is performed on the polished translation, including key contents such as punctuation marks, numbers, dates, etc., and the consistency of the glossary in the text is checked;

[0061] To quickly identify and correct common problems, an automated QA tool can be developed using technologies such as regular expressions and natural language understanding.

[0062] In this embodiment, in step eight, the translation is corrected according to the QA feedback, and the modified sentence segments are sent back to the QA check model for verification;

[0063] To improve the inspection efficiency and accuracy, a continuous learning mechanism can be introduced to enable the QA check model to continuously learn new error patterns and correction methods.

[0064] In this embodiment, in step nine, usually three rounds of interaction are set. When there are still risk items after three rounds of interaction, it is delivered to a language expert for final inspection and processing to generate the final multilingual localization document;

[0065] To improve delivery quality and efficiency, an expert database and an online collaboration mechanism can be established to ensure that experts can respond and handle complex problems in a timely manner.

[0066] The present invention constructs a multilingual localization workflow based on large language models and agents. Different agents are defined through techniques such as prompting and model fine-tuning. At the same time, the overall workflow is automated through the understanding, reasoning, and planning capabilities of the large model, enabling low-cost and high-quality localization output, reducing the degree of manual intervention to a new level, and significantly improving the efficiency of global multilingual information and knowledge transfer. Thus, agents based on large models achieve applications such as project planning, calling external models and tools, and obtaining feedback, which are the best practices for applying large language models to multilingual localization. By aggregating project resources such as documents to be translated and localized, style guides, term bases, and memory banks, the overall project is analyzed and planned, and different processes and tasks are defined as professional agents through the large model, such as term extraction and unification agents, model comparison and recommendation agents, polishing agents, quality inspection agents, etc. A complete translation and localization workflow is constructed through the invocation between agents and with external tools and models.

[0067] As described above, the above are only the preferred specific embodiments of the present invention, but the protection scope of the present invention is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present invention, according to the technical solution and inventive concept of the present invention, makes equivalent substitutions or changes, and should be covered by the protection scope of the present invention.

Claims

1. A multilingual localization agent based on a large language model, characterized in that, It includes the following steps: Step 1: Project preparation and feature extraction; Step 2: Term extraction and filtering; Step 3: Glossary confirmation and interaction; Step 4: Generation of personalized translation model; Step 5: Translation quality assessment and model selection; Step 6: Preliminary translation and polishing; Step 7: Quality inspection and feedback; Step 8: Translation correction and re-inspection; Step 9: Expert intervention and final delivery.

2. The multilingual localization agent based on a large language model according to claim 1, characterized in that, In the said Step 1, project features are extracted according to documents such as document titles, tags, project information, style guides, etc., and the project workflow and model requirements are output, providing a basis for subsequent steps.

3. The multilingual localization agent based on a large language model according to claim 1, characterized in that, In the said Step 2, term clusters are identified through term extraction models such as TF-IDF and TextRank, monolingual document terms are extracted, and low-quality terms are filtered according to information such as the frequency of term occurrence.

4. The multilingual localization intelligent agent based on the large language model according to claim 1, wherein In the said Step 3, after the glossary is output, feedback interaction is carried out with external language experts to confirm the accuracy of the terms. After the language experts modify and confirm, it proceeds to the next step.

5. The multilingual localization agent based on a large language model according to claim 1, wherein In the said Step 4, a personalized translation model for the project is generated based on the RAG (Retrieval-Augmented Generation) technology and serves as one of the alternative models for the pre-translation process of this project.

6. The multilingual localization agent based on the large language model according to claim 1, wherein In the said Step 5, some document contents are randomly selected, translations are generated through multiple alternative models, and translation quality assessment metrics such as BLEU, TER, METEOR, and BERT Score are used for scoring. The model with the highest score is selected as the pre-translation model.

7. The multi-language localization intelligent agent based on the large language model according to claim 1, wherein In the said Step 6, in combination with the glossary, the optimal pre-translation model is used for preliminary translation. The polishing model polishes some translations according to requirements such as the project style guide. This link mainly focuses on the diction of the translations in aspects such as grammar and rhetoric in the target language, and optimizes the translations in terms of grammar, rhetoric, etc. through the polishing model.

8. The multilingual localization agent based on a large language model according to claim 1, characterized in that, In the said Step 7, the polished translations are checked by the QA model, including key contents such as punctuation marks, numbers, dates, etc., and the consistency of the glossary in the text is checked.

9. The multilingual localization agent based on a large language model according to claim 1, characterized in that, In the said Step 8, the translations are corrected according to the QA feedback, and the modified sentence segments are sent to the QA inspection model again for verification.

10. The multilingual localization agent based on a large language model according to claim 1, wherein In the said Step 9, usually three rounds of interaction are set. When there are still risk items after three interactions, it is delivered to language experts for final inspection and processing to generate the final multilingual localization document.