Text translation method and system and electronic equipment

By combining corpus, term translation library and untranslated item list library to obtain multiple translation references, and using large language models for translation, the problems of inaccurate translation of professional terms and inconsistent translation style in the existing technology are solved, and a higher quality translation effect is achieved.

CN120068894APending Publication Date: 2025-05-30AUTEL INTELLIGENT TECHNOLOGY CORP LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510218165.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-26
Publication Date
2025-05-30

AI Technical Summary

Technical Problem

Existing machine translation models cannot obtain ideal translation results when processing professional terms, and the translated text does not have a specific translation style, resulting in low translation quality.

Method used

By obtaining the text to be translated and obtaining multiple translation references based on the corpus, term translation library and untranslated item list library, the prompt words are determined and input into the pre-trained large language model to obtain the final translation result.

Benefits of technology

Improve the accuracy and quality of translations, especially when dealing with professional terms, you can obtain more accurate translations, and by obtaining a list of untranslated items, avoiding translations of specific entries, making the translation results more in line with the original context.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120068894A_ABST
    Figure CN120068894A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of language translation, and discloses a text translation method and system and electronic device.The method comprises the steps that a to-be-translated text is obtained, a first translation result corresponding to the to-be-translated text is obtained based on a corpus, a second translation result corresponding to the to-be-translated text is obtained based on a term translation library, and the to-be-translated text is translated based on an untranslated item list library; and obtaining an untranslated item list corresponding to the to-be-translated text, determining a cue word based on the to-be-translated text, the first translation result, the second translation result and the untranslated item list, and inputting the cue word into a pre-trained large language model to obtain a final translation result of the to-be-translated text. The translation result is obtained based on the corpus, the term translation library and the untranslated item list library, the translation accuracy can be improved, and the translation quality is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The embodiments of the present application relate to the technical field of language translation, and in particular, to a text translation method, system, and electronic device. Background Art

[0002] With the development of the globalization process, all industries are facing the challenge of translating a large number of professional materials into multiple languages to meet the expanding needs of the international market.

[0003] In the automotive diagnosis industry, translation materials with high professionalism and accuracy are required. The traditional translation method is based on a machine translation model, such as a neural network-based machine translation model. However, the above models cannot obtain ideal translation results when dealing with professional terms, and the translated text does not have a specific translation style, resulting in low translation quality. Summary of the Invention

[0004] To solve the above technical problems, the embodiments of the present application provide a text translation method, system, and electronic device, which can improve the accuracy of translation to improve the translation quality.

[0005] To solve the above technical problems, the embodiments of the present application provide the following technical solutions:

[0006] In a first aspect, the embodiments of the present application provide a text translation method, which includes:

[0007] Obtain the text to be translated;

[0008] Based on a corpus, obtain a first translation result corresponding to the text to be translated, where the corpus is used to store multiple historical texts and the historical translation texts corresponding to the historical texts;

[0009] Based on a term translation library, obtain a second translation result corresponding to the text to be translated, where the term translation library is used to store multiple terms and the term translation results corresponding to the terms;

[0010] Based on a non-translation item list library, obtain a non-translation item list corresponding to the text to be translated, where the non-translation item list library is used to store multiple non-translation entries;

[0011] Based on the text to be translated, the first translation result, the second translation result, and the non-translation item list, determine a prompt word;

[0012] Input the prompt word into a pre-trained large language model to obtain the final translation result of the text to be translated, where the pre-trained large language model is used to translate text.

[0013] In some embodiments, before obtaining the first translation result corresponding to the text to be translated, the method further includes:

[0014] Preprocess multiple historical texts and historical translation texts, specifically including:

[0015] Establish a mapping relationship between the historical texts and the historical translation texts, such that each historical text corresponds one-to-one to a historical translation text in one language.

[0016] In some embodiments, based on a corpus, obtaining the first translation result corresponding to the text to be translated includes:

[0017] Calculate the first cosine similarity between the text to be translated and multiple historical texts;

[0018] Determine at least one historical text and the corresponding historical translation text with the first cosine similarity greater than the similarity threshold as the first translation result corresponding to the text to be translated.

[0019] In some embodiments, based on a term translation library, obtaining the second translation result corresponding to the text to be translated includes:

[0020] Perform word segmentation on the text to be translated to obtain a word segmentation result, where the word segmentation result includes several specific phrases;

[0021] Calculate the second cosine similarity between each specific phrase and multiple terms;

[0022] Based on the term corresponding to the specific phrase with the highest second cosine similarity, determine the term translation result corresponding to the term, and use the specific phrase, the term, and the term translation result as the second translation result.

[0023] In some embodiments, based on a non-translation item list library, obtaining the non-translation item list corresponding to the text to be translated includes:

[0024] Match the word segmentation result with multiple non-translation entries to obtain a matching result;

[0025] Based on the matching result, obtain the non-translation item list corresponding to the text to be translated, where the non-translation item list includes several non-translation words.

[0026] In some embodiments, the prompt words include a first prompt word, a second prompt word, and a third prompt word. Determining the prompt words based on the text to be translated, the first translation result, the second translation result, and the non-translation item list includes:

[0027] Based on the text to be translated and the first translation result, determine the first prompt word, where the first prompt word includes a translation instruction word, and the translation instruction word is used to indicate the text translated by a pre-trained large language model and the format of the text translation;

[0028] Based on the text to be translated and the second translation result, determine the second prompt word, where the second prompt word includes a term comparison word, and the term comparison word is used to indicate the terms in the text to be translated that need to be translated for the pre-trained large language model;

[0029] Based on the text to be translated and the list of non-translated items, determine the third prompt word, where the third prompt word includes a non-translated item annotation word, and the non-translated item annotation word is used to indicate the words in the text to be translated that do not need to be translated for the pre-trained large language model.

[0030] In some embodiments, the method further includes:

[0031] Based on the text to be translated and the final translation result, update the corpus;

[0032] Based on the updated corpus, retrain the pre-trained large language model.

[0033] In some embodiments, based on the text to be translated and the final translation result, updating the corpus includes:

[0034] At every preset time interval, add multiple texts to be translated and the final translation results to the corpus to update the corpus, where each text to be translated corresponds to one final translation result.

[0035] In a second aspect, an embodiment of the present application provides a text translation system, and the system includes:

[0036] A pre-trained large language model, configured to translate the text to be translated to output the final translation result of the text to be translated;

[0037] A corpus, configured to store historical texts and historical translation texts corresponding to the historical texts;

[0038] A term translation library, configured to store terms and term translation results corresponding to the terms;

[0039] A non-translated item list library, configured to store multiple non-translated entries.

[0040] In a third aspect, an embodiment of the present application provides an electronic device, including:

[0041] At least one processor, and

[0042] A memory communicatively connected to the at least one processor, where

[0043] The memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the text translation method as in the first aspect.

[0044] The beneficial effects of the embodiments of the present application are as follows: Different from the prior art, the embodiments of the present application provide a text translation method, which includes: obtaining the text to be translated, obtaining the first translation result corresponding to the text to be translated based on a corpus, where the corpus is used to store multiple historical texts and historical translation texts, obtaining the second translation result corresponding to the text to be translated based on a term translation library, where the term translation library is used to store multiple terms and the term translation results corresponding to the terms one by one, obtaining the non-translation item list corresponding to the text to be translated based on a non-translation item list library, where the non-translation item list library is used to store multiple non-translated entries, determining a prompt word based on the text to be translated, the first translation result, the second translation result, and the non-translation item list, and inputting the prompt word into a pre-trained large language model to obtain the final translation result of the text to be translated, where the pre-trained large language model is used to translate text.

[0045] By integrating the corpus, the term translation library, and the non-translation item list library, the present application provides multi-faceted translation references for the model to obtain accurate translations of professional terms, improve the professionalism of the translation results, and avoid translating specific entries by obtaining the non-translation item list, making the translation results more in line with the original context and improving the translation quality. BRIEF DESCRIPTION OF THE DRAWINGS

[0046] One or more embodiments are exemplarily illustrated by the corresponding drawings. These exemplary illustrations do not limit the embodiments. Elements with the same reference numerals in the drawings are represented as similar elements. Unless otherwise stated, the drawings in the figures do not constitute a scale limitation.

[0047] Figure 1 is a schematic flowchart of a text translation method provided by an embodiment of the present application;

[0048] Figure 2 is a schematic flowchart of a preprocessing of historical texts and historical translation texts provided by an embodiment of the present application;

[0049] Figure 3 is Figure 1 a detailed flowchart of step S102 in

[0050] Figure 4 is Figure 1 a detailed flowchart of step S103 in

[0051] Figure 5 is Figure 1 a detailed flowchart of step S104 in

[0052] Figure 6 is Figure 1 a detailed flowchart of step S105 in

[0053] Figure 7 It is a schematic flow chart for retraining a pre-trained large language model provided by an embodiment of the present application;

[0054] Figure 8 is Figure 7 a refined flow chart of step S701 in

[0055] Figure 9 It is a schematic structural diagram of a text translation system provided by an embodiment of the present application;

[0056] Figure 10 It is a schematic structural diagram of a text translation device provided by an embodiment of the present application;

[0057] Figure 11 It is a schematic structural diagram of an electronic device provided by an embodiment of the present application.

[0058] Explanation of the reference numerals in the drawings:

[0059] Label Name Label Name 900 Text translation system 1002 Historical translation result acquisition unit 901 List of items not to be translated library 1003 Term translation result acquisition unit 902 Term translation library 1004 Unit for obtaining the list of items not to be translated 903 Corpus 1005 Prompt acquisition unit 904 Prompt engineering 1006 Final translation result acquisition unit 905 Pre-trained large language model 1100 Electronic device 1000 Text translation device 1101 Processor 1001 Text acquisition unit 1102 Memory Detailed implementation manners

[0060] To make the objectives, technical solutions and advantages of the embodiments of the present application clearer, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments in the present application without creative efforts shall fall within the protection scope of the present application.

[0061] It should be noted that if there is no conflict, the various features in the embodiments of the present application can be combined with each other, and all are within the protection scope of the present application. In addition, the terms "first", "second", etc. adopted in the present application do not limit the data, but are only used to divide the same items or similar items with basically the same functions and effects.

[0062] Before introducing the embodiments of the present application, a brief introduction to the text translation methods known to the inventors of the present application will be given first, so as to facilitate the subsequent understanding of the embodiments of the present application.

[0063] Currently, there are mainly three translation modes in the automotive diagnosis industry: Manual Translation, Google Translate, and Machine Translation Model.

[0064] The above three methods have the following disadvantages: (1) Manual translation requires a large amount of time, increasing the time cost; (2) Google Translate mainly performs translation through machine translation based on translation rules and a large amount of statistical data. Since its translation rules cannot cover all possible language phenomena and translation situations, and the statistical data is not diverse enough, the model cannot learn enough translation knowledge, resulting in a decline in the quality of the translation results; (3) Machine translation models mainly rely on deep learning technology to learn the mapping relationship between languages, but lack sufficient knowledge in the professional field, leading to a decline in translation quality.

[0065] To address the above problems, the present application provides a text translation method. By obtaining the text to be translated, and based on a corpus, obtaining a first translation result corresponding to the text to be translated, based on a term translation library, obtaining a second translation result corresponding to the text to be translated, based on a do-not-translate list library, obtaining a do-not-translate list corresponding to the text to be translated, and based on the text to be translated, the first translation result, the second translation result, and the do-not-translate list, determining a prompt word, and inputting the prompt word into a pre-trained large language model to obtain the final translation result of the text to be translated. The present application obtains translation results based on a corpus, a term translation library, and a do-not-translate list library, which can improve the accuracy of translation and thus improve the quality of translation.

[0066] Before elaborating on the present application in detail, the nouns and terms involved in the embodiments of the present application are explained. The nouns and terms involved in the embodiments of the present application are applicable to the following explanations:

[0067] Large Language Model (LLM): A deep learning-based language model generated by training with a large amount of text data. Among them, large language models include GPT-4 (Generative Pre-trained Transformer 4), Qwen 2, Llama 3.

[0068] Semantic Search: A search technology that uses a vector database to perform semantic analysis and matching on the input text to provide accurate results.

[0069] Corpus: Specific industry text data collected and stored, which is used to train and evaluate machine learning models.

[0070] Do-not-translate List: A list containing specific words and phrases that do not need to be translated.

[0071] The technical solution of the present application is specifically described below in conjunction with the accompanying drawings of the specification:

[0072] Please refer toFigure 1 , Figure 1 is a schematic flowchart of a text translation method provided by an embodiment of the present application;

[0073] Among them, the text translation method is applied to an electronic device. Specifically, the execution subject of the text translation method is one or at least two processors of the electronic device.

[0074] As Figure 1 shown, the text translation method includes:

[0075] Step S101: Obtain the text to be translated;

[0076] Among them, the text to be translated is a text segment or a document. Among them, the format of the document includes Word, PDF, Excel, PowerPoint, etc., and the document includes text content.

[0077] It can be understood that since there are multiple languages, the text to be translated is one of the multiple languages, such as Chinese.

[0078] Step S102: Based on the corpus, obtain the first translation result corresponding to the text to be translated, where the corpus is used to store multiple historical texts and the historical translation texts corresponding to the historical texts;

[0079] Among them, the corpus is used to store multiple historical texts and the historical translation texts corresponding to the historical texts.

[0080] In the embodiment of the present application, before obtaining the first translation result corresponding to the text to be translated, preprocess the multiple historical texts and the historical translation texts in the corpus to reorganize the historical texts and the historical translation texts so that each historical text corresponds to a historical translation text in one language.

[0081] It can be understood that each language corresponds to a corpus. When it is necessary to translate a certain language into another language, the translation result is obtained based on the corpus corresponding to the language to be translated. For example, when the text to be translated is Chinese and it is necessary to translate Chinese into English, the translation result is obtained from the corpus corresponding to English.

[0082] Please refer to Figure 2 , Figure 2 is a schematic flowchart of a method for preprocessing historical texts and historical translation texts provided by an embodiment of the present application;

[0083] As Figure 2 shown, the preprocessing of the historical texts and the historical translation texts includes:

[0084] Step S201: Establish a mapping relationship between historical texts and historical translation texts, such that each historical text corresponds one-to-one with a historical translation text in a specific language;

[0085] Specifically, establishing a mapping relationship between historical texts and historical translation texts means establishing a one-to-one correspondence between each historical text and a historical translation text in a specific language, so that a historical text can be mapped to a historical translation text in a specific language. For example, there are two historical texts A and B, and two historical translation texts C and D, where D is the historical translation text of B. Then map A to C and B to D.

[0086] In the embodiments of the present application, it is also necessary to check the mapping relationship between historical texts and historical translation texts to determine whether there are incorrect mapping relationships. For example, if the historical translation text of historical text A is B, but in the corpus, historical text A corresponds to historical translation text D, then the mapping relationship needs to be corrected.

[0087] In the embodiments of the present application, by establishing a mapping relationship between historical texts and historical translation texts, during the actual translation process, the translation result of the translation text can be directly obtained based on this mapping relationship.

[0088] Specifically, based on the corpus, obtain the first translation result corresponding to the text to be translated, that is, based on the corpus, calculate the first cosine similarity between the text to be translated and historical texts to determine the historical texts similar to the text to be translated, and directly obtain the historical texts similar to the text to be translated and the historical translation texts corresponding to the historical texts as the first translation result based on the mapping relationship between historical texts and historical translation texts stored in the corpus. Among them, the first translation result includes historical texts and the historical translation texts corresponding to the historical texts. The specific steps for calculating the first cosine similarity refer to Figure 3 .

[0089] Please refer to Figure 3 , Figure 3 is Figure 1 a detailed process schematic diagram of step S102 in

[0090] As Figure 3 shown, this step S102 includes:

[0091] Step S121: Calculate the first cosine similarity between the text to be translated and multiple historical texts;

[0092] In the embodiments of the present application, the first cosine similarity is used to measure the similarity between the text to be translated and historical texts, such as the similarity in content, structure, style, or context between the text to be translated and historical texts. For example, the text to be translated is "communicating", and the historical text similar to the text to be translated in the corpus is "connecting".

[0093] Specifically, the text to be translated and multiple historical texts are converted into vectors, and the first cosine similarity between the vector of the text to be translated and the vectors of multiple historical texts is calculated. The formula is as follows:

[0094] cos(θ) = (X * Y) / (‖X‖ * ‖Y‖)

[0095] Where cos(θ) is the first cosine similarity, X is the vector of the text to be translated, Y is the vector of the text to be translated, X * Y is the dot product of X and Y, ‖X‖ is the norm (length) of X, and ‖Y‖ is the norm (length) of Y.

[0096] Among them, the value range of the first cosine similarity is [-1, 1]. The closer the first cosine similarity is to 1, the more similar the two texts are.

[0097] Step S122: Determine whether the first cosine similarity is less than the similarity threshold;

[0098] Specifically, compare the size between the first cosine similarity and the similarity threshold to determine whether the first cosine similarity is less than the similarity threshold. If the first cosine similarity is greater than or equal to the similarity threshold, jump to step S123; if the first cosine similarity is less than the similarity threshold, jump to step S124.

[0099] Among them, the similarity threshold is adjusted according to specific circumstances. For example, the similarity threshold is 0.8.

[0100] Step S123: Use the historical texts with the first cosine similarity greater than the similarity threshold and the corresponding historical translation texts as the first translation result corresponding to the text to be translated;

[0101] Specifically, when the first cosine similarity is greater than or equal to the similarity threshold, the historical text corresponding to the first cosine similarity greater than the similarity threshold is used as the text similar to the text to be translated, and the historical translation text corresponding to the historical text similar to the text to be translated is obtained from the corpus. The historical text corresponding to the first cosine similarity greater than the similarity threshold and the historical translation text corresponding to the historical text are used as the first translation result corresponding to the text to be translated.

[0102] In the embodiments of the present application, when there are multiple historical texts similar to the text to be translated, the multiple historical texts similar to the text to be translated are sorted according to the first cosine similarity from large to small, and several historical texts with the highest first cosine similarity are obtained as the texts similar to the text to be translated. For example, when there are 10 historical texts similar to the text to be translated, 3 historical texts with the highest first cosine similarity are obtained as the texts similar to the text to be translated.

[0103] Step S124: Determine that the historical text corresponding to the current first cosine similarity and the corresponding historical translation text are not the first translation result corresponding to the text to be translated;

[0104] Specifically, when the first cosine similarity is less than the similarity threshold, it is determined that the historical text corresponding to the current first cosine similarity and the historical translation text corresponding to the historical text are not the first translation result corresponding to the text to be translated, that is, the historical translation text corresponding to the historical text corresponding to the current first cosine similarity is not similar to the text to be translated.

[0105] Step S103: Based on the term translation library, obtain the second translation result corresponding to the text to be translated, where the term translation library is used to store multiple terms and the term translation results corresponding to the terms;

[0106] In the embodiments of the present application, the term translation library is used to store multiple terms and the term translation results corresponding to the terms, where each term corresponds to a term translation result.

[0107] Specifically, extract a specific phrase from the text to be translated, and calculate the second cosine similarity between the specific phrase and the terms in the term translation library to determine the terms identical to the specific phrase, and based on the relationship between the terms and the corresponding term translation results, determine the terms identical to the specific phrase and the term translation results corresponding to the terms as the second translation result corresponding to the text to be translated, where the second translation result includes the terms identical to the specific phrase and the term translation results corresponding to the terms.

[0108] Among them, the specific phrase refers to a professional term in a certain field. For example, in the automotive diagnosis industry, professional terms include engine control unit, anti-lock braking system, and so on.

[0109] Please refer to Figure 4 , Figure 4 is Figure 1 the detailed process schematic diagram of step S103 in

[0110] As Figure 4 shown, this step S103 includes:

[0111] Step S131: Perform word segmentation on the text to be translated to obtain a word segmentation result;

[0112] Specifically, use a word segmentation tool to perform word segmentation on the text to be translated to obtain a word segmentation result, where the word segmentation result includes specific phrases, and the word segmentation tools include tools such as Jieba, SnowNLP, PkuSeg, THULAC, and HanLP.

[0113] Step S132: Calculate the second cosine similarity between each specific phrase and multiple terms;

[0114] In the embodiments of the present application, the second cosine similarity is used to measure the similarity between specific phrases and terms.

[0115] Specifically, based on the calculation formula of the first cosine similarity, the second cosine similarity between each specific phrase and multiple terms is calculated.

[0116] Step S133: Determine the term translation result corresponding to the term based on the term corresponding to the specific phrase with the highest second cosine similarity, and use the specific phrase, the term, and the term translation result as the second translation result;

[0117] Specifically, sort the second cosine similarities corresponding to each specific phrase, and take the term with the highest value of the second cosine similarity as the term most similar to the specific phrase. Based on the one-to-one correspondence between the term and the term translation result, obtain the term translation result corresponding to the term from the term translation library, and use the specific phrase, the term corresponding to the specific phrase, and the term translation result corresponding to the term in the text to be translated as the second translation result. Among them, the second translation result includes the specific phrase, the term most similar to the specific phrase, and the term translation result corresponding to the term.

[0118] Step S104: Obtain the non-translation item list corresponding to the text to be translated based on the non-translation item list library, where the non-translation item list library is used to store multiple non-translated entries;

[0119] In the embodiments of the present application, the non-translation item list library is used to store multiple non-translated entries, where the non-translated entries include proper nouns, measurement units, etc. For example, the measurement unit is km (kilometer).

[0120] Specifically, search in the non-translation item list library to see if there is an entry identical to the word segmentation result. If there is an identical entry, add the word segmentation result to the non-translation item list. Among them, the non-translation item list includes the non-translated words in the word segmentation result, and the non-translated words include measurement units, etc.

[0121] Please refer to Figure 5 , Figure 5 is Figure 1 a detailed flowchart of step S104 in

[0122] As Figure 5 shown, this step S104 includes:

[0123] Step S141: Match the word segmentation result with multiple non-translated entries to obtain a matching result;

[0124] Among them, the word segmentation result includes multiple words.

[0125] Specifically, each word in the word segmentation result is matched one by one with multiple non-translation entries in the non-translation item list library to obtain a matching result, where the matching result includes a first matching result and a second matching result. The first matching result is that the word in the word segmentation result is the same as the non-translation entry in the non-translation item list library, and the second matching result is that the word in the word segmentation result is different from the non-translation entry in the non-translation item list library.

[0126] Step S142: Based on the matching result, obtain a non-translation item list corresponding to the text to be translated, where the non-translation item list includes several non-translated words;

[0127] Specifically, when the matching result is the first matching result, add the words in the word segmentation result that are the same as the non-translation entry to the non-translation item list. The non-translation item list includes non-translated words, and the non-translated words are the words in the word segmentation result that are the same as the non-translation entry.

[0128] In the embodiment of the present application, when the matching result is the second matching result, it is determined that the word in the word segmentation result is different from the non-translation entry, and there is no need to add the word that is different from the non-translation entry to the non-translation item list.

[0129] In the embodiment of the present application, the non-translation item list library, the corpus, and the term translation library can be replaced with a dictionary database. During the translation process, the translation result of the text to be translated is obtained based on the dictionary database, where the content of the dictionary database is updated by the cloud, and the content of the dictionary database can be updated in real time, enabling the model to learn the latest translation rules and enabling the model to be adjusted in a timely manner according to the updated content in the dictionary database.

[0130] Step S105: Determine a prompt word based on the text to be translated, the first translation result, the second translation result, and the non-translation item list;

[0131] Specifically, input the text to be translated, the first translation result, the second translation result, and the non-translation item list into the prompt engineering to obtain a prompt word, where the prompt word is used to instruct a pre-trained large language model to translate the text to be translated to output a translation result that meets expectations.

[0132] In the embodiment of the present application, the prompt word includes a first prompt word, a second prompt word, and a third prompt word. For the specific steps of determining the first prompt word, the second prompt word, and the third prompt word, please refer to Figure 6 .

[0133] Among them, the prompt engineering (Prompt Engineering) is used to generate and optimize the prompt word (Prompt). The prompt engineering includes methods such as template design, context addition, and example provision.

[0134] Among them, the large language model (LLM) is a language model based on deep learning. The pre-trained large language model is a large language model trained with a large amount of text data. The large language model is used to translate text to output a translation result. The large language model includes but is not limited to GPT-4 (Generative Pre-trained Transformer 4), Qwen, Llama3 (Large Language Model 3).

[0135] Please refer to Figure 6 , Figure 6 is Figure 1 a detailed process schematic diagram of step S105 in

[0136] As Figure 6 shown, this step S105 includes:

[0137] Step S151: Determine a first prompt word based on the text to be translated and the first translation result;

[0138] In the embodiments of the present application, the first prompt word is a prompt word determined based on the text to be translated and the first translation result.

[0139] Among them, the first translation result includes historical texts similar to the text to be translated and the corresponding historical translation texts of the historical texts.

[0140] Specifically, input the text to be translated and the first translation result into the prompt engineering to generate a translation instruction word, that is, determine the first prompt word. The first prompt word includes the translation instruction word. The translation instruction word is used to indicate how the large language model translates the text. The translation prompt word includes the style, tone, specific format, etc. of the translation. For example, the translation prompt word instructs the large language model to maintain the format of the historical text or use a specific translation style, etc. Among them, the specific format is the text format output by the large language model, such as the json format.

[0141] Step S152: Determine a second prompt word based on the text to be translated and the second translation result;

[0142] Among them, the second translation result includes specific phrases, terms, and the corresponding term translation results of the terms.

[0143] Specifically, input the text to be translated and the second translation result into the prompt engineering to determine the term comparison word, that is, determine the second prompt word. The second prompt word includes the term comparison word. The term comparison word is used to indicate the terms in the text to be translated that need to be translated by the large language model. The term comparison word includes the terms and the corresponding term translation results of the terms.

[0144] Step S153: Determine the third prompt word based on the text to be translated and the non-translation item list;

[0145] Among them, the non-translation item list includes words that are not to be translated.

[0146] Specifically, input the text to be translated and the non-translation item list into the prompt engineering to determine the non-translation item annotation words, that is, determine the third prompt word. The third prompt word includes the non-translation item annotation words, and the non-translation item annotation words are used to indicate which words or phrases in the text to be translated do not need to be translated. The non-translation item annotation words include specific symbols, tags or descriptions.

[0147] In the embodiments of the present application, generating prompt words through prompt engineering can guide the large language model to output translation texts in a specific format.

[0148] Step S106: Input the prompt word into the pre-trained large language model to obtain the final translation result of the text to be translated;

[0149] Specifically, the prompt word includes the first prompt word, the second prompt word, and the third prompt word. Input the first prompt word, the second prompt word, and the third prompt word into the pre-trained large language model to output the final translation result of the text to be translated through the pre-trained large language model.

[0150] In the embodiments of the present application, the final translation result also needs to be reviewed to ensure the accuracy of the final translation result. Specifically, the translation result is corrected through a crowd-sourced review platform (Crowd-sourced Review) to improve the accuracy of the translation result. Among them, the crowd-sourced review platform refers to an online platform that utilizes the power of the public, that is, the contributions of numerous users or professionals, to review, correct, and improve the accuracy of specific content (such as the final translation result). On the crowd-sourced review platform, the task publisher can submit the content to be reviewed, and users or experts on the crowd-sourced review platform can review these contents, provide modification suggestions or directly make corrections based on their knowledge and experience, thereby helping to improve the overall quality and accuracy of the translated content.

[0151] In the embodiments of the present application, semantic search is performed in the corpus, the term translation library, and the non-translation item list library to obtain the first translation result, the second translation result, and the third translation result of the text to be translated, and input the obtained translation results into the prompt engineering to output the prompt word, and input the prompt word into the pre-trained large language model to obtain the final translation result, which can improve the translation quality of the text to be translated.

[0152] Please refer to Figure 7 , Figure 7 which is a schematic flowchart of a process for retraining a pre-trained large language model provided by the embodiments of the present application;

[0153] As shown Figure 7 in the figure, retraining a pre-trained large language model includes:

[0154] Step S701: Update the corpus based on the text to be translated and the final translation result;

[0155] In the embodiments of the present application, a large amount of historical translation data is stored in the corpus. During the translation process, the historical translation data provides a unified translation standard for the text to be translated. Since new vocabulary, expressions, and grammatical structures often change, it is necessary to regularly update the historical translation data in the corpus to improve translation quality and efficiency.

[0156] Specifically, regularly obtain the text to be translated and the corresponding final translation result to update the corpus. For the specific update process, please refer to Figure 8 .

[0157] Please refer to Figure 8 , Figure 8 which Figure 7 is a schematic diagram of the refined process of step S701 in

[0158] As shown Figure 8 in the figure, this step S701 includes:

[0159] Step S711: At every preset time interval, add multiple texts to be translated and the final translation results to the corpus to update the corpus;

[0160] Specifically, at every preset time interval, add multiple texts to be translated and the final translation results to the corpus to update the corpus. For example, update once every three months.

[0161] In the embodiments of the present application, updating the corpus also includes batch update and dynamic updating.

[0162] Among them, batch update refers to updating multiple texts to be translated and the final translation results to the corpus in the form of a centralized batch. For example, when the number of multiple texts to be translated and the final translation results accumulates to 100,000, then update these 100,000 texts to be translated and the final translation results to the corpus. By means of batch update, the overhead of frequent updates can be reduced to save time costs.

[0163] Among them, dynamic update refers to real-time update. For example, every time a translation task is executed to obtain the text to be translated and the final translation result, directly update the text to be translated and the final translation result to the corpus. By dynamically updating the corpus, the accuracy and timeliness of translation can be improved.

[0164] Step S702: Retrain the pre-trained large language model based on the updated corpus;

[0165] Specifically, obtain multiple translation texts in the updated corpus and the corresponding translation results of the translation texts, and use the multiple translation texts and the corresponding translation results of the translation texts as training data to retrain the pre-trained large language model to improve the translation quality and efficiency of the large language model.

[0166] In the embodiment of the present application, the large language model can also be updated by means of incremental fine-tuning. Incremental fine-tuning means that after obtaining the training data, a part of the parameters or structure of the large language model is adjusted through the training data. In actual operation, it can be determined through experiments and verification which parameters or structures of the large language model have the greatest impact on the current translation task. For example, different parts of the model are updated, and the performance of the model on the validation set is observed, and the parameters of the best-performing part are selected for update. The model can be updated more frequently by means of incremental fine-tuning.

[0167] In the embodiment of the present application, a confidence information-based hybrid translation system can also be used. This system provides more accurate and fluent translation results by dynamically selecting the optimal translation path and the correction mechanism for low-confidence paragraphs according to the translation confidence information. The confidence can be calculated based on multiple factors, such as translation probability, language model score, semantic similarity, etc.

[0168] In the embodiment of the present application, the translation results of the large language model and the machine translation system can be scored through the confidence information-based hybrid translation system to determine one of the translation results as the final translation result.

[0169] Specifically, selecting the optimal translation path according to the translation confidence information includes: setting one or more confidence thresholds to distinguish high-confidence and low-confidence translation results, calculating the confidence of the translation result of the large language model and the confidence of the translation result of the machine translation system respectively, and selecting the translation result with a confidence greater than the confidence threshold. If the confidence of the translation result of the large language model is the same as the confidence of the translation result of the translation system, the translation result of the large language model is preferentially selected as the final translation result. Confidence = (1 - significance level) × 100%, where the significance level is obtained through experiments.

[0170] In the embodiments of the present application, the translation text can also be translated by a multi-stage translation system (Multi-stage Translation System) to obtain a translation result. The multi-stage translation system uses multiple dedicated small language models to perform translation in stages. The first stage performs basic translation, the second stage processes professional terms, and the third stage adjusts the language style. Through staged processing, each stage can focus on specific problems and achieve high-quality translation.

[0171] In the embodiments of the present application, a multi-agent system can also be established through a multi-agent translation system (Multi-agent Translation System) to simulate different agents to process the translation task. One agent focuses on technical term translation, another agent focuses on language style adjustment, and the third agent performs syntactic correction. Through agent collaboration, a high-quality translation task can be jointly completed.

[0172] Please refer to Figure 9 , Figure 9 which is a schematic structural diagram of the text translation system provided by the embodiments of the present application;

[0173] As Figure 9 shown, the text translation system 900 includes a non-translation item list library 901, a term translation library 902, a corpus 903, a prompt engineering 904, and a pre-trained large language model 905.

[0174] The non-translation item list library 901 is used to store multiple non-translated entries.

[0175] The term translation library 902 is used to store terms and the corresponding term translation results.

[0176] The corpus 903 is used to store historical texts and the corresponding historical translation texts.

[0177] The prompt engineering 904 is used to obtain the prompt corresponding to the text to be translated.

[0178] The pre-trained large language model 905 is used to translate the text to be translated to output the final translation result of the text to be translated.

[0179] Specifically, after obtaining the text to be translated, retrieve the non-translation item list from the non-translation item list library 901 to obtain the non-translation item list, retrieve the term translation result from the term translation library 902 to obtain the term translation result, perform semantic translation on the corpus 903 to obtain a similar text similar to the text to be translated and the translation result corresponding to the similar text, input the text to be translated, the non-translation item list, the term translation result, and the translation result corresponding to the similar text into the prompt engineering 904. The prompt engineering 904 outputs the prompt corresponding to the text to be translated, and input the prompt into the pre-trained large language model 905. The pre-trained large language model 905 translates the text to be translated according to the prompt to output the final translation result corresponding to the text to be translated.

[0180] In the embodiment of the present application, by combining the non-translation item list library 901, the term translation library 902, and the corpus 903, obtain the non-translation item list corresponding to the text to be translated, the term translation result, the similar text similar to the text to be translated, and the translation result corresponding to the similar text, so as to further obtain the prompt, and input the prompt into the pre-trained large language model 905 to obtain the final translation result, which can ensure the accuracy of the expression of professional terms based on a specific term library.

[0181] Please refer to Figure 10 , Figure 10 which is the structural schematic diagram of the text translation device provided by the embodiment of the present application;

[0182] As Figure 10 shown, the text translation device 1000 includes a text acquisition unit 1001, a historical translation result acquisition unit 1002, a term translation result acquisition unit 1003, a non-translation item list acquisition unit 1004, a prompt acquisition unit 1005, and a final translation result acquisition unit 1006.

[0183] The text acquisition unit 1001 is used to acquire the text to be translated;

[0184] The historical translation result acquisition unit 1002 is used to acquire the first translation result corresponding to the text to be translated based on the corpus, where the corpus is used to store multiple historical texts and the historical translation texts corresponding to the historical texts;

[0185] The term translation result acquisition unit 1003 is used to acquire the second translation result corresponding to the text to be translated based on the term translation library, where the term translation library is used to store multiple terms and the term translation results corresponding to the terms;

[0186] The non-translation item list acquisition unit 1004 is used to acquire the non-translation item list corresponding to the text to be translated based on the non-translation item list library, where the non-translation item list library is used to store multiple non-translated entries;

[0187] The prompt word acquisition unit 1005 is configured to determine a prompt word based on the text to be translated, the first translation result, the second translation result, and the non-translation item list;

[0188] The final translation result acquisition unit 1006 is configured to input the prompt word into a pre-trained large language model to obtain the final translation result of the text to be translated, where the pre-trained large language model is used for text translation.

[0189] In an embodiment of the present application, the text translation device 1000 may be a software module. The software module includes several instructions, which are stored in a memory. The processor can access the memory and call the instructions for execution to complete the text translation method of the above various embodiments.

[0190] In an embodiment of the present application, the text translation device may also be built by hardware devices. For example, the text translation device may be built by one or more than two chips, and each chip can work in coordination with each other to complete the text translation method described in the above various embodiments. For another example, the text translation device may also be built by various logic devices, such as being built by a general-purpose processor, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), a single-chip microcomputer, an ARM (Acorn RISC Machine), or other programmable logic devices, discrete gate or transistor logic, discrete hardware components, or any combination of these components.

[0191] The text translation device in the embodiment of the present application may be a device, or a component, an integrated circuit, or a chip in a terminal. The device may be a mobile electronic device or a non-mobile electronic device. Exemplarily, the mobile electronic device may be a mobile phone, a tablet computer, a laptop computer, a handheld computer, a vehicle-mounted electronic device, a wearable device, an ultra-mobile personal computer (UMPC), a netbook, or a personal digital assistant (PDA), etc. The non-mobile electronic device may be a server, a network attached storage (NAS), a personal computer (PC), a television (TV), a teller machine, or a self-service machine, etc. The embodiment of the present application does not make a specific limitation.

[0192] The text translation device in the embodiment of the present application may be a device with an operating system. The operating system may be an Android operating system, an iOS operating system, or other possible operating systems. The embodiment of the present application does not make a specific limitation.

[0193] The text translation devices provided in the embodiments of this application can all achieve Figure 1 each process implemented by the method embodiments. To avoid repetition, they will not be elaborated here.

[0194] It should be noted that the above device can execute the text translation method provided in the embodiments of this application and has corresponding functional modules and beneficial effects for executing the method. For technical details not described in detail in the device embodiments, reference can be made to the text translation method provided in the embodiments of this application.

[0195] In the embodiments of this application, by combining the corpus, term translation library, and non-translation item list library in the text translation device to obtain the translation result, the accuracy of translation can be improved, thereby improving the quality of translation.

[0196] Please refer to Figure 11 , Figure 11 which is a schematic structural diagram of an electronic device provided in the embodiments of this application;

[0197] As Figure 11 shown, the electronic device 1100 includes one or more processors 1101 and a memory 1102. Among them, Figure 11 one processor 1101 is taken as an example in

[0198] The processor 1101 and the memory 1102 can be connected through a bus or other means, Figure 11 and taking the connection through the bus as an example in

[0199] The processor is configured to execute the text translation method in any embodiment of this application. The method includes:

[0200] Obtain the text to be translated, and based on the corpus, obtain the first translation result corresponding to the text to be translated. Among them, the corpus is used to store multiple historical texts and the historical translation texts corresponding to the historical texts. Based on the term translation library, obtain the second translation result corresponding to the text to be translated. Among them, the term translation library is used to store multiple terms and the term translation results corresponding to the terms. Based on the non-translation item list library, obtain the non-translation item list corresponding to the text to be translated. Among them, the non-translation item list library is used to store multiple non-translation entries. Based on the text to be translated, the first translation result, the second translation result, and the non-translation item list, determine the prompt word, and input the prompt word into the pre-trained large language model to obtain the final translation result of the text to be translated. Among them, the pre-trained large language model is used to translate text.

[0201] The memory 1102 serves as a non-volatile computer-readable storage medium and can be used to store non-volatile software programs, non-volatile computer-executable programs, and modules, such as the program instructions / modules corresponding to the text translation method in the embodiments of the present invention. The processor 1101 executes various functional applications and data processing of the electronic device by running the non-volatile software programs, instructions, and modules stored in the memory 1102, that is, implements the text translation method in the above method embodiments.

[0202] The memory 1102 may include high-speed random access memory and may also include non-volatile memory, such as at least one magnetic disk storage device, a flash memory device, or other non-volatile solid-state storage devices. In some embodiments, the memory 1102 optionally includes a memory remotely disposed relative to the processor 1101. Examples of the above networks include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.

[0203] One or more modules are stored in the memory 1102 and, when executed by one or more processors 1101, execute the text translation method in any of the above method embodiments. For example, execute each of the steps described above Figure 1 as shown.

[0204] The embodiments of the present application also provide a computer program product. The computer program product includes one or more program codes, and the program codes are stored in a non-volatile computer-readable storage medium. The processor of the electronic device reads the program codes from the non-volatile computer-readable storage medium, and the processor executes the program codes to complete the steps of the text translation method provided in the above embodiments.

[0205] Through the description of the above embodiments, those of ordinary skill in the art can understand that all or part of the steps of implementing the above embodiments can be completed by hardware, or can be completed by hardware related to program codes. The program can be stored in a non-volatile computer-readable storage medium. The non-volatile computer-readable storage medium mentioned above can be a read-only memory, a magnetic disk, an optical disc, or the like.

[0206] Through the description of the above embodiments, those of ordinary skill in the art can clearly understand that each embodiment can be implemented by means of software plus a general hardware platform, and of course, it can also be implemented by hardware. Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through a computer program, and the program can be stored in a non-volatile computer-readable storage medium. When the program is executed, it can include the processes of the embodiments of the above methods. Among them, the non-volatile computer-readable storage medium can be a magnetic disk, an optical disk, a read-only memory (ROM), or a random access memory (RAM), etc.

[0207] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present application, rather than to limit them; under the idea of the present application, the technical features in the above embodiments or different embodiments can also be combined, and the steps can be implemented in any order, and there are many other changes in different aspects of the present application as described above. For the sake of brevity, they are not provided in detail; although the present application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements on some of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present application.

Claims

1. A text translation method, characterized in that: The method comprises: Get the text to be translated; Based on a corpus, obtaining a first translation result corresponding to the text to be translated, wherein the corpus is used to store a plurality of historical texts and historical translation texts corresponding to the historical texts; Based on a term translation library, obtaining a second translation result corresponding to the text to be translated, wherein the term translation library is used to store a plurality of terms and term translation results corresponding to the terms; Based on an untranslatable item list library, obtaining an untranslatable item list corresponding to the text to be translated, wherein the untranslatable item list library is used to store a plurality of untranslatable items; Determine a prompt word based on the text to be translated, the first translation result, the second translation result, and the list of untranslatable items; The prompt word is input into a pre-trained large language model to obtain a final translation result of the text to be translated, wherein the pre-trained large language model is used to translate the text.

2. The method according to claim 1, characterized in that: Before obtaining the first translation result corresponding to the text to be translated, the method further includes: Preprocessing the plurality of historical texts and historical translation texts includes: A mapping relationship between the historical texts and the historical translation texts is established so that each of the historical texts corresponds one-to-one to a historical translation text in one language.

3. The method according to claim 1, characterized in that The obtaining, based on the corpus, a first translation result corresponding to the text to be translated includes: Calculating the first cosine similarity between the text to be translated and the plurality of historical texts; At least one of the historical texts and the corresponding historical translation text whose first cosine similarity is greater than a similarity threshold is determined as a first translation result corresponding to the text to be translated.

4. The method according to claim 1, characterized in that Based on the terminology translation library, obtaining a second translation result corresponding to the text to be translated includes: Performing word segmentation processing on the text to be translated to obtain a word segmentation result, wherein the word segmentation result includes a plurality of specific phrases; Calculating a second cosine similarity between each of the specific phrases and the plurality of terms; Based on the term corresponding to the specific phrase with the highest second cosine similarity, a term translation result corresponding to the term is determined, and the specific phrase, the term and the term translation result are used as the second translation result.

5. The method according to claim 4, characterized in that Based on the untranslatable item list library, obtaining the untranslatable item list corresponding to the text to be translated includes: Matching the word segmentation result with the plurality of untranslated entries to obtain a matching result; Based on the matching result, a list of untranslatable items corresponding to the text to be translated is obtained, wherein the list of untranslatable items includes a number of untranslatable words.

6. The method according to claim 5, characterized in that The prompt words include a first prompt word, a second prompt word, and a third prompt word. The determining of the prompt words based on the text to be translated, the first translation result, the second translation result, and the list of untranslatable items includes: Determine a first prompt word based on the text to be translated and the first translation result, wherein the first prompt word includes a translation indicator word, and the translation indicator word is used to indicate the text translated by the pre-trained large language model and the format of the text translation; Determine a second prompt word based on the text to be translated and the second translation result, wherein the second prompt word includes a terminology control word, and the terminology control word is used to indicate a term to be translated in the text to be translated of the pre-trained large language model; Based on the text to be translated and the list of untranslated items, a third prompt word is determined, wherein the third prompt word includes an untranslated item marking word, and the untranslated item marking word is used to indicate words that do not need to be translated in the text to be translated by the pre-trained large language model.

7. The method according to claim 1, characterized in that The method further comprises: Based on the text to be translated and the final translation result, updating the corpus; The pre-trained large language model is retrained based on the updated corpus.

8. The method according to claim 7, characterized in that The updating of the corpus based on the text to be translated and the final translation result comprises: At intervals of a preset time period, a plurality of the texts to be translated and the final translation results are added to the corpus to update the corpus, wherein each of the texts to be translated corresponds to a final translation result.

9. A text translation system, characterized in that: The system comprises: A pre-trained large language model is used to translate the text to be translated to output a final translation result of the text to be translated; Corpus, used to store historical texts and historical translation texts corresponding to historical texts; A term translation library, used to store terms and term translation results corresponding to the terms; The untranslatable item list library is used to store multiple untranslatable items.

10. An electronic device, characterized in that: include: at least one processor, and a memory communicatively coupled to the at least one processor, wherein: The memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can perform the text translation method according to any one of claims 1 to 8.