Text translation method and device, electronic equipment and storage medium

By filtering out the reference original text and translated text with the highest similarity to the text to be translated in the large language model, combining the dual recall and preset strategy, the problem of poor performance of the large language model in complex translation tasks is solved, and more efficient and accurate translation results are achieved.

CN120297293APending Publication Date: 2025-07-11BEIJING REALAI TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510506081.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-22
Publication Date
2025-07-11

AI Technical Summary

Technical Problem

The existing large language model has poor translation results when facing complex and highly professional translation tasks, and the existing methods consume a lot of computing resources and are not conducive to rapid information updates.

Method used

By performing dual recalls based on the title of the text to be translated, the recall title collection is selected, and the recall weight is determined. The target reference original text and large language model are selected based on the preset strategy, and the target large language model is used for translation.

Benefits of technology

It improves the accuracy and efficiency of translation results, reduces computing resource consumption, supports rapid updates and adapts to translation needs in different fields.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120297293A_ABST
    Figure CN120297293A_ABST
Patent Text Reader

Abstract

According to the text translation method and device, the electronic equipment and the storage medium, compared with the prior art, the reference original text and the reference translated text which have the highest similarity with the to-be-translated text are screened out from an existing translated text database comprising the original text and the translated text through twice screening of the title content and the text content; while the screening range is gradually reduced, double-path recall is used for screening during screening to improve the accuracy and the coverage range of a screening result, and meanwhile, a target large language model with a good translation effect on a to-be-translated text is further screened out from a plurality of large language models by using a reference original text and a reference translation; and the to-be-translated text is translated by utilizing the screened target large language model, so that the accuracy of a translation result can be effectively improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of translation technology, and particularly to a text translation method, device, electronic device and storage medium. Background Art

[0002] With the development of machine translation technology, traditional rule-based and statistical methods have gradually been replaced by neural machine translation (NMT) due to their limitations in dealing with complex grammar and context associations. In recent years, with the rise of deep learning and large-scale pre-trained language models (such as BERT, GPT series), significant progress has been made in the field of NLP. Large language models (LLMs) can capture rich semantic information through pre-training on massive texts. Fine-tuning on specific tasks enables these models to better understand the context in translation tasks and generate more context-appropriate translations. LLMs can not only handle translations of a single language pair but also support multi-language and multi-modal translation requirements. This ability stems from the extensive and diverse corpus the model has been exposed to during training, making it more generalizable when dealing with different languages and input forms.

[0003] Although large language models have relatively good translation effects when dealing with data they have seen during translation training, in actual translation tasks, new corpora emerge, and large models have basically not seen them, so the translation effects are relatively poor. In such cases, we usually prepare some similar datasets and fine-tune the model to make it adapt to different fields. However, such methods consume a large amount of computing resources and are not conducive to rapid information update. At the same time, in actual translation tasks, due to the diversity in writing styles and word usage, as well as the limitations of the large model itself and the uncertainty of answer return, when facing translation tasks with complex business definitions and strong professionalism, the translation effects of large models may not be satisfactory. Summary of the Invention

[0004] Embodiments of this application provide a text translation method, device, electronic device and storage medium, aiming to solve the problem of poor translation effects of existing text translation methods.

[0005] In a first aspect, the text translation method provided by this application includes: performing dual-channel recall in a preset translation text database based on the title of the text to be translated to obtain a set of recalled titles, where the set of recalled titles includes multiple recalled titles;

[0006] Determining the recall weight corresponding to each recalled title in the set of recalled titles to obtain multiple recall weights;

[0007] Based on the multiple recall weights and the preset title screening strategy, at least one recalled title that meets the preset title screening strategy is screened out from the set of recalled titles;

[0008] Determine the reference original texts corresponding to the at least one recalled title to obtain at least one reference original text;

[0009] Based on the at least one reference original text and the preset original text screening strategy, a target reference original text that meets the preset original text screening strategy is screened out from the at least one reference original text;

[0010] Determine the target reference translation corresponding to the target reference original text, and based on the target reference translation, determine a target large language model that meets the preset large language model screening strategy among the preset multiple large language models;

[0011] Based on the target large language model, translate the text to be translated to obtain a translation result.

[0012] In some possible embodiments, the dual-channel recall is performed in the preset translation text database based on the title of the text to be translated to obtain a set of recalled titles, including:

[0013] Determine the target title of the text to be translated, and vectorize the target title to obtain a title vector;

[0014] Obtain a preset translation text database, where the translation text database includes translation originals and corresponding translation translations;

[0015] Based on the title vector, perform semantic recall and keyword recall in the translation text database respectively to obtain a set of recalled titles, and the set of recalled titles includes multiple recalled titles.

[0016] In some possible embodiments, the set of recalled titles includes a target recalled title;

[0017] The determination of the recall weights corresponding to each recalled title in the set of recalled titles to obtain multiple recall weights includes:

[0018] Determine the recall score of the target recalled title during the recall process;

[0019] Respectively determine whether the target recalled title and the target title are titles of the same series;

[0020] If the target recalled title and the target title are titles of the same series, determine that the title score of the target recalled title is the first title score;

[0021] If the target recall title and the target title are not titles of the same series, determine the title similarity between the target recall title and the target title;

[0022] Based on the title similarity, determine the title score of the target recall title as the second title score;

[0023] Based on the recall score, the first title score, and the second title score, determine the target recall weight corresponding to the target recall title, so as to obtain the recall weight corresponding to each recall title in the recall title set.

[0024] In some possible embodiments, the screening of the target reference text that meets the preset source text screening strategy from the at least one reference source text based on the at least one reference source text and the preset source text screening strategy includes:

[0025] Calculate the text similarity between the at least one reference source text and the text to be translated to obtain at least one text similarity; reorder the at least one text similarity in descending order to obtain a sorted text similarity sequence;

[0026] Screen out the first text similarity in the text similarity sequence whose text similarity is greater than the preset text similarity threshold, and determine the target reference source text corresponding to the first text similarity;

[0027] Or screen out the second text similarity in the text similarity sequence that is ranked before the preset rank, and determine the target reference source text corresponding to the second text similarity.

[0028] In some possible embodiments, the translating the text to be translated based on the target large language model to obtain a translation result includes:

[0029] Obtain the entity objects in the text to be translated, and determine the relationship knowledge of the text to be translated based on the entity objects; determine the target idiom in the text to be translated, and determine the idiom translation text corresponding to the target idiom to obtain an idiom translation pair; combine the relationship knowledge of the text to be translated, the idiom translation pair, and the target large language model to translate the text to be translated to obtain a translation result.

[0030] In some possible embodiments, the method further includes: creating a single-article database corresponding to the text to be translated;

[0031] According to the text to be translated and the corresponding translation result, determine a new translation pair corresponding to the text to be translated;

[0032] Store the translation result corresponding to the text to be translated, the relationship knowledge, and the new translation pair in the single-article database; correct and confirm the content in the single-article database, and update the confirmed single-article database to the translation text database.

[0033] In a second aspect, the text translation device provided by the present application includes:

[0034] A title recall module for performing dual-channel recall in a preset translation text database based on the title of the text to be translated to obtain a set of recalled titles, where the set of recalled titles includes multiple recalled titles;

[0035] A recall weight calculation module for determining the recall weight corresponding to each recalled title in the set of recalled titles to obtain multiple recall weights;

[0036] A recalled title screening module for screening at least one recalled title that meets the preset title screening strategy from the set of recalled titles based on the multiple recall weights and the preset title screening strategy;

[0037] An original text determination module for determining the reference original text corresponding to the at least one recalled title to obtain at least one reference original text;

[0038] An original text screening module for screening a target reference original text that meets the preset original text screening strategy from the at least one reference original text based on the at least one reference original text and the preset original text screening strategy;

[0039] A large language model screening module for determining the target reference translation corresponding to the target reference original text, and determining a target large language model that meets the preset large language model screening strategy from a preset multiple large language models based on the target reference translation;

[0040] A translation module for translating the text to be translated based on the target large language model to obtain a translation result.

[0041] In a third aspect, the present application provides an electronic device, including a memory and a processor, the memory stores a computer program, and the processor is used to run the computer program in the memory to implement the steps in the text translation method provided by the present application.

[0042] In a fourth aspect, the computer-readable storage medium provided by the present application stores multiple instructions, and these instructions are suitable for being loaded by a processor to implement the steps in the text translation method provided by the present application.

[0043] In a fifth aspect, the computer program product provided by the present application includes a computer program or instructions, and when the computer program or instructions are executed by a processor, the steps in the text translation method provided by the present application are implemented.

[0044] The present application provides a text translation method, apparatus, electronic device, and storage medium. Compared with the prior art, the present application uses two rounds of screening through the title and the body content to screen out the reference original text and the reference translation text with the highest similarity to the text to be translated in the existing translation text database including the original text and the translation. While gradually narrowing the screening scope, double-channel recall is also used during the screening to improve the accuracy and coverage of the screening results. At the same time, further screening is carried out among multiple large language models using the reference original text and the reference translation text to select the target large language model with better translation effect for the text to be translated. Using the selected target large language model to translate the text to be translated can effectively improve the accuracy of the translation result. Brief Description of the Drawings

[0045] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the accompanying drawings required for the description of the embodiments. Obviously, the accompanying drawings in the following description are only some embodiments of the present application. For those skilled in the art, without creative efforts, other accompanying drawings can also be obtained based on these drawings.

[0046] Figure 1 is the architecture diagram of the text translation system provided by the embodiment of the present application;

[0047] Figure 2 is the flowchart of an embodiment of the text translation method provided by the embodiment of the present application;

[0048] Figure 3 is the flowchart of an embodiment for determining multiple recall weights provided by the embodiment of the present application;

[0049] Figure 4 is the flowchart of an embodiment for screening the target reference original text provided by the embodiment of the present application;

[0050] Figure 5 is the flowchart of an embodiment for translating the text to be translated using a large language model provided by the embodiment of the present application;

[0051] Figure 6 is the flowchart of an embodiment for updating the translation text database provided by the embodiment of the present application;

[0052] Figure 7 is the complete flowchart of an embodiment of the text translation method provided by the present application;

[0053] Figure 8 is the structural diagram of the text translation apparatus provided by the embodiment of the present application

[0054] Figure 9Schematic structural diagram of the electronic device provided by the embodiment of the present application;

[0055] Figure 10 It is a schematic structural diagram of an embodiment of the terminal device provided in the embodiment of the present application;

[0056] Figure 11 It is a schematic structural diagram of the server provided by the embodiment of the present application. Detailed implementation manners

[0057] It should be noted that the principle of the present application is illustrated by being implemented in a suitable computing environment. The following description is based on the specific embodiments of the present application illustrated, and it should not be regarded as limiting other specific embodiments of the present application not detailed herein. In the following description of the present application, reference is made to "some embodiments", which describe a subset of all possible embodiments. However, it can be understood that "some embodiments" can be the same subset or different subsets of all possible embodiments, and can be combined with each other without conflict.

[0058] In the following description of the present application, the terms "first / second / third" are only used to distinguish similar objects and do not represent a specific order for the objects. It can be understood that "first / second / third" can be interchanged with a specific order or sequence when allowed, so that the embodiments of the present application described herein can be implemented in an order other than that illustrated or described herein. Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the technical field to which the present application belongs. The terms used herein are only for the purpose of describing the embodiments of the present application and are not intended to limit the present application.

[0059] In order to improve the effect of text translation based on large language models, the embodiments of the present application provide a text translation method, apparatus, electronic device and computer-readable storage medium. Among them, the text translation method can be executed by the text translation apparatus, or by an electronic device integrated with the text translation apparatus. The technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative efforts belong to the scope of protection of the present application.

[0060] Please refer to Figure 1 , which is the architecture diagram of the text translation system provided by the present application. As Figure 1 shown, the text translation system based on large language models includes an electronic device 100, and the text translation apparatus provided by the present application is integrated in the electronic device 100.

[0061] Among them, the electronic device 100 can be any device configured with a processor and having processing capabilities, such as mobile electronic devices with processors like smartphones, tablets, personal digital assistants, laptops, smart speakers, etc., or fixed electronic devices with processors like desktop computers, TVs, servers, industrial devices, etc.

[0062] In addition, as Figure 1 shown, the text translation system based on the large language model may further include a memory 200 for storing adversarial defense samples.

[0063] In the embodiments of the present application, the memory 200 may be a cloud memory. Cloud storage is a new concept extended and developed based on the concept of cloud computing. A distributed cloud storage system (hereinafter referred to as the storage system) refers to a storage system that combines a large number of different types of storage devices (storage devices are also called storage nodes) in the network through functions such as cluster applications, grid technology, and distributed file systems, and collaborates through application software or application interfaces to jointly provide data storage and service access functions to the outside world.

[0064] Currently, the storage method of the storage system is as follows: Create a logical volume. When creating a logical volume, physical storage space is allocated for each logical volume. This physical storage space may be composed of the disks of a certain storage device or several storage devices. The client stores data on a certain logical volume, that is, stores the data on the file system. The file system divides the data into many parts, and each part is an object. The object not only contains the data but also contains additional information such as data identifiers (ID, ID entity). The file system writes each object into the physical storage space of the logical volume respectively, and the file system records the storage location information of each object. Thus, when the client requests to access the data, the file system can enable the client to access the data according to the storage location information of each object.

[0065] The process of the storage system allocating physical storage space for a logical volume is specifically as follows: According to the capacity estimation of the objects stored in the logical volume (this estimation often has a large margin relative to the actual capacity of the objects to be stored) and the group of the redundant array of independent disks (RAID), the physical storage space is pre-divided into stripes. A logical volume can be understood as a stripe, thereby allocating physical storage space for the logical volume.

[0066] It should be noted that Figure 1The scenario schematic diagram of the large language model-based text translation system shown is merely an example. The large language model-based text translation system and scenario described in the embodiments of this application are for more clearly explaining the technical solutions of the embodiments of this application, and do not constitute a limitation on the technical solutions provided by the embodiments of this application. Those of ordinary skill in the art know that with the evolution of the large language model-based text translation system and the emergence of new business scenarios, the technical solutions provided by the embodiments of this application are equally applicable to similar technical problems.

[0067] The solution provided by the embodiments of this application involves technologies such as Artificial Intelligence (AI), Computer Vision (CV), and Machine Learning (ML). The specific description is as follows through the following embodiments:

[0068] Among them, AI is to use a digital computer or a machine controlled by a digital computer to simulate, extend, and expand human intelligence, a theory, method, technology, and application system that perceives the environment, acquires knowledge, and uses knowledge to obtain the best results. In other words, artificial intelligence is a comprehensive technology in computer science. It attempts to understand the essence of intelligence and produce a new intelligent machine that can react in a way similar to human intelligence. Artificial intelligence is also the study of the design principles and implementation methods of various intelligent machines, enabling the machine to have the functions of perception, reasoning, and decision-making.

[0069] AI technology is an interdisciplinary subject with a wide range of fields, including both hardware-level technologies and software-level technologies. The basic technologies of artificial intelligence generally include technologies such as sensors, dedicated artificial intelligence chips, cloud computing, distributed storage, big data processing technologies, operation / interaction systems, and mechatronics. The software technologies of artificial intelligence mainly include several major directions such as computer vision technology, speech processing technology, natural language processing technology, and machine learning / deep learning.

[0070] CV is a science that studies how to make machines "see". Further, it refers to using cameras and computers to replace human eyes to perform machine vision such as target recognition, tracking, and measurement on targets, and further perform graphic processing to make the computer process into an image that is more suitable for human eyes to observe or be transmitted to instruments for detection. As a scientific discipline, computer vision studies related theories and technologies and attempts to build an artificial intelligence system that can obtain information from images or multi-dimensional data. Computer vision technology usually includes technologies such as adversarial perturbation generation, image recognition, image semantic understanding, image retrieval, OCR, video processing, video semantic understanding, video content / behavior recognition, three-dimensional object reconstruction, 3D technology, virtual reality, augmented reality, simultaneous localization and mapping, and also includes common biometric recognition technologies such as face recognition and fingerprint recognition.

[0071] The text translation method, apparatus, electronic device, and storage medium provided in this application will be described in detail below. It should be noted that the serial numbers of the following embodiments do not limit the preferred order of the embodiments.

[0072] Please refer to Figure 2 , Figure 2 which is a schematic flowchart of an embodiment of the text translation method provided in an embodiment of this application. As Figure 2 shown, the process of the text translation method provided in this application is as follows:

[0073] 201. Perform two-way recall on the text to be translated in a preset translation text database to obtain a set of recalled titles, where the set of recalled titles includes multiple recalled titles.

[0074] The text translation method provided in this application mainly uses a large language model with good translation effects to translate the text to be translated; and before that, it is necessary to screen out a large language model with good translation effects from multiple large language models. For different types of texts to be translated or translation scenarios, the translation effects of different large language models are different. Therefore, the methods for screening out a large language model with good translation effects are also different. This application determines the relationship between the text to be translated and the existing original text with a translation, and then determines a large language model with good translation effects for the current text to be translated by referring to the existing translated original text and translation.

[0075] Specifically, this application first needs to determine the relationship between the text to be translated and the existing original text with a translation. This application mainly determines the relationship through a recall method. In some embodiments, a recall can be performed based on the current text to be translated to determine books or articles that are relatively close or similar in content to the current text to be translated. The books or articles related to the text to be translated can assist in the translation of the text to be translated. For example, existing translations can be referred to for translating certain specific names of people and places in the text to be translated to be consistent with the existing translations. For this application, it is first necessary to obtain a preset translation text database. The translation text database includes the original text of the article and the corresponding translation text, and usually needs to include the title of the article, specific chapter titles, paragraph titles, etc., and the corresponding translation text. For example, taking the original text of the article as "Harry Potter and the Philosopher's Stone", the translation text database includes the original text of the book "Harry Potter and the Philosopher's Stone" (including the book title, each chapter title, each paragraph title, etc.), and at least one translation text different from the original text corresponding to the original text. For example, the translation text database includes the English original text of "Harry Potter and the Philosopher's Stone", as well as translation texts in multiple versions such as Chinese, Japanese, and Korean. It should be noted that in the actual preset translation text database, there are usually the original texts of multiple books or articles and translation texts in different versions. This is only an example here and does not represent a limitation on the content in the translation text database.

[0076] After obtaining the translation text database, the data in the translation text database can be recalled based on the text to be translated to determine the existing original text and the corresponding translation text that are relevant to the text to be translated. For this application, the title can be used for recall first, that is, the title is used to search in the translation text database to determine books or articles that are relevant to the text to be translated. This is because for actual books or articles, there are usually a large number of fixed nouns in the title, and the title has a high degree of condensation and can well summarize the content of the book and article. Using the title for recall can perform a preliminary screening of the data in the translation text database and reduce the amount of data that needs to be obtained or calculated subsequently. For example, taking the title of the text to be translated as "Harry Potter and the Deathly Hallows" as an example, which mainly includes two keywords, "Harry Potter" and "Deathly Hallows". Recalling based on these two keywords can more likely recall books or articles that are relevant to the text to be translated. For example, two books, "Harry Potter and the Philosopher's Stone" and "Harry Potter and the Chamber of Secrets", can be recalled. Subsequently, usually only the content of the two books, "Harry Potter and the Philosopher's Stone" and "Harry Potter and the Chamber of Secrets", needs to be processed, and there is no need to judge other books in the translation text database, which can greatly reduce the amount of data processing.

[0077] In some embodiments of the present application, a dual-channel recall is performed in a preset translation text database based on the title of the text to be translated, and a set of recalled titles can be obtained, which may include: determining the target title of the text to be translated and vectorizing the target title to obtain a title vector; obtaining the preset translation text database, where the translation text database includes the original translation and the corresponding translated text; performing semantic recall and keyword recall in the translation text database based on the title vector to obtain a set of recalled titles, and the set of recalled titles includes multiple recalled titles.

[0078] Specifically, when the present application actually uses the title for recall, a dual-channel recall needs to be performed, that is, multiple recall strategies are used for recall. This can retrieve and screen the title from multiple perspectives, avoid information omission caused by single-channel recall, and can comprehensively sort the results of the dual-channel recall. By integrating different scoring mechanisms, the accuracy of the final result can be improved, ensuring that the most relevant books are ranked at the front and guaranteeing the accuracy of subsequent translations. In a specific embodiment, semantic recall and keyword recall can be respectively performed in the translation text database based on the title vector corresponding to the target title of the text to be translated. Among them, keyword recall can quickly and efficiently screen out the content that matches the keyword, while semantic recall can query the content that is fuzzy, colloquial, or has different expressions but similar semantics, which can improve the accuracy and coverage of the recall, has a good generalization effect, and can provide more comprehensive search results; therefore, semantic recall and keyword recall can be combined to ensure the accuracy and coverage of the recall results. Taking the target title "Harry Potter and the Deathly Hallows" as an example, the result of keyword recall can be "Harry Potter and the Philosopher's Stone", and the result of using semantic recall can be "Harry Potter and the Philosopher's Stone". It should be noted that the recalled results are also the titles of books or articles, and there is no need to recall the specific text content of the books or articles here.

[0079] 202. Determine the recall weight corresponding to each recalled title in the set of recalled titles to obtain multiple recall weights.

[0080] 203. Based on the multiple recall weights and the preset title screening strategy, screen out at least one recalled title that meets the preset title screening strategy from the set of recalled titles.

[0081] Based on dual-channel recall, a recall title set including multiple title recall results can be obtained. Generally speaking, the number of multiple title recall results is still relatively large. The multiple title recall results can be further screened to further reduce the number of titles for which the main text content of the book or article needs to be determined subsequently, thereby reducing the subsequent computation amount. In the embodiments of the present application, the recall weight corresponding to each recall title can be determined, and then at least one recall title meeting the requirements can be screened out from the multiple recall titles in the recall title set according to a preset title screening strategy. The specific method for screening at least one recall title will be described in subsequent embodiments and is not limited herein.

[0082] 204. Determine the reference original texts corresponding to at least one recall title to obtain at least one reference original text.

[0083] 205. Based on at least one reference original text and a preset original text screening strategy, screen out the target reference original texts meeting the preset original text screening strategy from the at least one reference original text.

[0084] In the foregoing embodiments, at least one recall title having a correlation with the title of the text to be translated is determined according to the title of the text to be translated. However, since the title is usually short and contains less information, it is also necessary to determine the main text content in the preset translation text database that has a correlation with the main text content of the text to be translated according to the specific main text content corresponding to the title. Specifically, the reference original texts corresponding to at least one recall title can be determined respectively to obtain at least one reference original text, and then the target reference original texts meeting the preset original text screening strategy can be screened out from the at least one reference original text by using the preset original text screening strategy; the target reference original text is one or several original texts having the highest correlation with the text to be translated.

[0085] 206. Determine the target reference translation corresponding to the target reference original text, and determine the target large language model meeting the preset large language model screening strategy from the preset multiple large language models based on the target reference translation.

[0086] 207. Based on the target large language model, translate the text to be translated to obtain a translation result.

[0087] The main purpose of this application is to screen out the target large language model with better translation effect for the text to be translated. The method for determining the target large language model is mainly to judge which large language model has a better translation effect on the target reference original text. Since there is a high correlation between the target reference original text and the text to be translated, if the target large language model has a good translation effect on the target reference original text, it can also be determined that the target large language model has a good translation effect on the text to be translated. Therefore, this application can determine the target reference translation corresponding to the target reference original text, and determine the target large language model among multiple preset large language models based on the target reference translation. Then, the target large language model is used to translate the text to be translated to obtain the translation result.

[0088] This application provides a text translation method. Compared with the prior art, this application uses two screenings through the title and the content of the text body to screen out the reference original text and reference translation with the highest similarity to the text to be translated in the existing translation text database including the original text and the translation. While gradually narrowing the screening scope, it also uses dual-channel recall for screening during the screening to improve the accuracy and coverage of the screening results. At the same time, it further screens out the target large language model with better translation effect on the text to be translated among multiple large language models using the reference original text and reference translation. Using the screened target large language model to translate the text to be translated can effectively improve the accuracy of the translation result.

[0089] As Figure 3 shown, it is a schematic flowchart of an embodiment for determining multiple recall weights provided by an embodiment of this application, which may include:

[0090] 301. Determine the recall score of the target recall title during the recall process.

[0091] The set of recall titles provided by this application includes multiple recall titles. Taking any one of the recall titles as the target recall title as an example; during the actual dual-channel recall, different recall results correspond to different recall scores. For example, the recall titles are "Harry Potter and the Philosopher's Stone" and "JK Rowling - The Creative Journey of Harry Potter"; among them, "Harry Potter and the Philosopher's Stone" has a relatively high similarity with the target title "Harry Potter and the Deathly Hallows", usually corresponding to a relatively high recall score, while "JK Rowling - The Creative Journey of Harry Potter" is also an article related to Harry Potter, but has a lower similarity with "Harry Potter and the Deathly Hallows", so it corresponds to a lower recall score. This application can determine the recall score corresponding to each recall title during the actual dual-channel recall process. The specific method can refer to the prior art and will not be limited here.

[0092] 302. Respectively judge whether the target recall title and the target title are of the same series of titles.

[0093] 303. If the target recall title and the target title are of the same series, determine the title score of the target recall title as the first title score.

[0094] The aforementioned recall score can be determined during the actual recall process. This application can also additionally add a title score to further determine the relevance between the recalled title and the target title of the text to be translated. Specifically, it can first be determined separately whether the target recall title and the target title are of the same series. If the target recall title and the target title are of the same series, determine the title score of the target recall title as the first title score, and the first title score is usually a relatively high score. It should be noted that when determining whether the target recall title and the target title belong to the same series, it is mainly determined by whether the author corresponding to the target recall title and the author corresponding to the target title are the same person; this is because the same author usually writes using similar nouns or grammar, and there is a high similarity between the books or articles written by the same author. Therefore, if it is determined that the target title and the target recall title are written by the same author, it can be considered that there is a high similarity between the text to be translated and the content of the text corresponding to the target recall title, which helps the subsequent determination of the large language model.

[0095] 304. If the target recall title and the target title are not of the same series, determine the title similarity between the target recall title and the target title.

[0096] 305. Determine the title score of the target recall title as the second title score based on the title similarity.

[0097] If it is determined that the target recall title and the target title are not of the same series, the title similarity between the target recall title and the target title can be further determined, and the title score of the target recall title can be further determined based on the title similarity. Specifically, in the actual translation text database, there are usually also books or articles written by multiple other authors. These authors usually choose the same names for special nouns when writing, making the titles of some books or articles have a certain similarity to the target title; therefore, the similarity between the target recall title and the target title can be determined, and the title score can be determined based on the similarity. Since the target recall title in this application is a dual-channel recall mainly obtained through semantic recall and keyword recall, when calculating the similarity, the similarity can be calculated based on keywords and semantics; the specific calculation process can refer to the prior art and will not be limited here.

[0098] After calculating the title similarity between the target recall title and the target title, the second title score corresponding to each title similarity can be determined based on a preset title similarity-title score relationship. For example, when the title similarity is between 60% and 70%, the second title score is 30 points; when the title similarity is between 20% and 30%, the second title score can be 8 points. Based on the preset title similarity-title score relationship, the second title score corresponding to each recall title that does not belong to the same series as the target title can be determined; in other embodiments, other methods can also be used to determine the second title score, which is not limited here.

[0099] It should be noted that the target recall title usually only corresponds to the first title score or the second title score, and the target recall title will not have both the first title score and the second title score at the same time.

[0100] 306. Based on the recall score, the first title score, and the second title score, determine the target recall weight corresponding to the target recall title, so as to obtain the recall weight corresponding to each recall title in the recall title set.

[0101] After obtaining the recall score, the first title score, and the second title score corresponding to the target recall title, the target recall weight of the target recall title can be further obtained. In some embodiments, the target recall weight can be obtained by combining the recall score and the first title score, or by combining the recall score and the second title score. In a specific embodiment, the recall score and the first title score can be fused by inverse sorting to obtain the recall weight, or the recall score and the second title score can be fused by inverse sorting to obtain the recall weight. The specific process of determining the recall weight using inverse fusion sorting can refer to the prior art and is not limited in this application.

[0102] In the foregoing embodiments, taking any retrieved title in the retrieved title set as the target retrieved title, the retrieval weight corresponding to each retrieved title can be obtained; after obtaining multiple retrieval weights, at least one retrieved title that meets the preset title screening strategy needs to be screened out from the multiple retrieval weights. In this application, the preset title screening strategy mainly screens out some retrieved titles with higher retrieval weights, or screens out some retrieved titles whose retrieval weights are greater than the retrieval weight threshold. This is because the retrieval weight in this application basically represents the correlation between the retrieved title and the target title. Generally, the higher the retrieval weight, the higher the correlation between the retrieved title and the target title; and the higher the correlation of the title, it also means that the text content of the retrieved title has a higher correlation with the text to be translated. When the correlation is higher, it is easier to determine a large language model with better translation effect for the text to be translated based on the retrieved title and its corresponding text content. Therefore, in this application, at least one retrieved title with a higher retrieval weight can be screened out, or at least one retrieved title with a higher retrieval weight can be screened out. Generally speaking, a certain number of retrieved titles will be screened out, rather than only one retrieved title being screened out.

[0103] After screening out at least one retrieved title, the reference original text corresponding to at least one retrieved title can be further determined to obtain at least one reference original text; then, further screening is performed according to the reference original text to screen out the target reference original text with a high correlation between the text content and the text to be translated. As Figure 4 shown, it is a schematic flowchart of an embodiment for screening the target reference original text provided by the embodiment of this application, which may include:

[0104] 401. Calculate the text similarity between at least one reference original text and the text to be translated to obtain at least one text similarity.

[0105] 402. Re-sort at least one text similarity in descending order to obtain a sorted text similarity sequence.

[0106] 403. Screen out the first text similarity in the text similarity sequence whose text similarity is greater than the preset text similarity threshold, and determine the target reference original text corresponding to the first text similarity.

[0107] 404. Or screen out the second text similarity ranked before the preset rank in the text similarity sequence, and determine the target reference original text corresponding to the second text similarity.

[0108] Specifically, the purpose of screening out the target reference original text from at least one reference original text using the original text screening strategy is to determine the original text with a relatively high correlation with the text to be translated from the perspective of specific text content. Therefore, the similarity between at least one reference original text and the text to be translated can be determined by calculating the text similarity, and the correlation between the two can be determined based on the similarity. Specifically, the text similarity between at least one reference original text and the text to be translated can be calculated respectively to obtain at least one text similarity, and then the at least one text similarity can be re-sorted in a certain order, such as from large to small, to obtain a sorted text similarity sequence. Based on the sorted text similarity sequence, the first text similarity with a text similarity greater than the preset text similarity can be screened out, and the target reference original text corresponding to the first text similarity can be determined. In some other embodiments, the second text similarity ranked before the preset ranking can also be screened out from the sorted text similarity sequence, and the target reference original text corresponding to the second text similarity can be determined.

[0109] In the embodiments of the present application, when screening the target reference original text, usually only one target reference original text is screened out, and the only target reference original text is determined as the target reference original text with the highest correlation or similarity with the text to be translated. Therefore, the text similarity with the highest similarity can be determined from the sorted text similarity sequence, and the reference original text corresponding to the text similarity with the highest similarity is used as the target reference original text.

[0110] In the foregoing embodiments, through two screenings of the title and the text content, the target reference original text with the highest similarity or the highest correlation with the text to be translated is determined. Then, the target large language model with the best translation effect can be screened out using the reference original text. Specifically, determining the target reference translation corresponding to the target reference original text, and determining the target large language model that meets the preset large language screening strategy among the preset multiple large language models based on the target reference translation may include the following steps: determining the target reference translation corresponding to the target reference original text; using the preset multiple large language models to translate the target reference original text to obtain multiple initial translation results; calculating the translation text similarity between the multiple initial translation results and the target reference translation respectively to obtain multiple translation text similarities; and determining the target large language model among the multiple large language models according to the multiple translation text similarities.

[0111] Since the two screenings of this application are both screenings of the content in the preset translation text database, and the translation text database includes the original translation text and the translated text; after determining the target reference original text, the target reference translation corresponding to the target reference original text can be further determined. Then, multiple large language models are used to translate the target reference original text to obtain multiple initial translation results, and then the translation text similarity between the initial translation results of different large language models and the already translated target reference translation is judged, and the translation effect of the large language model is determined according to the translation text similarity. Specifically, multiple translation text similarities can be re-sorted to obtain a re-sorted sequence of translation text similarities, and the translation text similarity with the highest translation text similarity is selected from it, and the large language model corresponding to the translation text similarity with the highest translation text similarity is selected from it, which is the target large language model. The highest translation text similarity means that the translation result of this large language model is closest to the existing target reference translation, so it can be determined that the translation effect of this large language model is the best.

[0112] When actually translating the text to be translated, other knowledge or content related to the text to be translated can also be combined to assist the translation to improve the accuracy of the translation. Specifically, as Figure 5 shown, it is a schematic flowchart of an embodiment of using a large language model to translate the text to be translated provided by an embodiment of this application, including:

[0113] 501. Obtain the entity objects in the text to be translated, and determine the relationship knowledge of the text to be translated based on the entity objects.

[0114] Specifically, in the actual translation process, it is usually necessary to combine specific context content to improve the accuracy of translation. At the same time, it is also necessary to make certain adjustments to the translated text to make the writing of the translated text smooth. Therefore, the entity objects in the text to be translated can be obtained, the relationship knowledge of the text to be translated can be determined based on the entity objects, and the relationship knowledge can be combined to assist in translation. Among them, the entity objects can be specific physical objects, such as people, places, organizations, etc., or abstract concepts such as physics, economics, etc.; each entity object has a unique identifier and a series of attributes used to describe its characteristics. The relationship knowledge is the link connecting different entities and represents various relationships or connections between different entities. Obtaining the entity objects and relationship knowledge in the text to be translated can effectively combine the connections between context contents during the translation process, thereby improving the translation effect. In some embodiments, Zhipu mapping technology can be used to extract the entity objects in the text to be translated, such as extracting the names of people, places, etc. in the text to be translated, and retrieving the relationships between multiple names and the relationships between names and places through the knowledge graph, so as to obtain the relationship knowledge of the text to be translated. In a specific embodiment, the relationship knowledge of the entity object can be determined according to the entity object of the country. For example, the relationship knowledge of the entity object includes country-province / nation / terrain, and then different provinces can be extended from the province, different nations can be extended from the nation, and different terrains can be extended from the terrain, etc.; the relationship knowledge of the entity object can be obtained based on the entity objects in the text to be translated. And when the entity objects are different, the relationship knowledge is also different; therefore, different entity objects can be selected according to actual needs to obtain different relationship knowledge.

[0115] 502. Determine the target idiom in the text to be translated, and determine the idiom translation text corresponding to the target idiom to obtain an idiom translation pair.

[0116] In the actual translation process, not only can the relational knowledge of the text to be translated be obtained to assist the translation, but also the fixed idioms in the text to be translated can be determined, and the existing translation of the fixed idioms can be determined to assist the translation of the text to be translated, thereby reducing the amount of translation required in the actual translation process. Among them, idioms are usually phrases that are often used together and have a specific form. The corresponding translations of these idioms are usually determined, and no additional translation is required. Therefore, an idiom translation pair including idioms and corresponding translations can be obtained. For example, there are corresponding translations for certain idioms or famous sayings or aphorisms. By determining the target idioms in the text to be translated and determining the idiom translation text corresponding to the target idiom, the actual amount of translation required in the text to be translated can be effectively reduced, and the idiom translation in the translated text after the text to be translated can be ensured to be consistent with the existing translation. For example, "Learn until you are old" can be fixedly translated as: "It is never too old to learn", and the two form a complete idiom translation pair. It should be noted that before determining the idiom translation pair, it is first necessary to determine the idiom in the text to be translated, and the idiom needs to be determined in combination with the context. Specifically, the idiom can be determined in combination with the relationship knowledge described in the above embodiment, and then the corresponding translation of the idiom can be searched in the existing idiom knowledge base. Among them, the idiom knowledge base includes multiple idioms and their different versions of translations, for example, including a certain idiom and its corresponding English translation, Japanese translation, Korean translation, etc.

[0117] 503. Combining the relational knowledge of the text to be translated, the idiom translation pair and the target large language model, the text to be translated is translated to obtain a translation result.

[0118] After obtaining the relational knowledge and idiom translation pairs of the text to be translated, the large language model can be used to translate the text to be translated, and the translation after the model translation is adjusted in combination with the relational knowledge and idiom translation pairs of the text to be translated to obtain the final translation result. The process of translating using the target large language model can refer to the existing technology and is not limited here.

[0119] After the text to be translated is translated and the translation result is obtained, the translation text database can be updated according to the translated translation result. Figure 6 As shown, a schematic diagram of an embodiment of the process of updating the translation text database provided in the embodiment of the present application may include:

[0120] 601. Create a single article database corresponding to the text to be translated.

[0121] 602. Determine a new translation pair corresponding to the text to be translated according to the text to be translated and the corresponding translation result.

[0122] 603. Store the translation result, relationship knowledge, and new translation pairs corresponding to the text to be translated in the single-article database.

[0123] 604. Check and confirm the content in the single-article database, and update the confirmed single-article database to the translation text database.

[0124] Specifically, after obtaining the translation result corresponding to the text to be translated, a single-article database corresponding to the text to be translated can be created. This database includes the text to be translated and its corresponding translation result. Moreover, new translation pairs in the translation result can be further determined. For example, specific translations for newly emerged personal names or place names are used to obtain new translation pairs. Store the translation result, relationship knowledge, and new translation pairs corresponding to the text to be translated in the single-article database. The translated text of the text to be translated that has been completed can be used to assist in the translation of subsequent untranslated texts. When the entire text to be translated is completely translated, the content in the single-article database can be manually checked and confirmed to ensure that the content in the single-article database is correct. Then, the confirmed single-article database can be updated to the translation text database to obtain a new translation text database. When translating a new text to be translated subsequently, it can be recalled from the new translation text database. That is, the translation text database can be continuously iterated to improve the translation effect of the thickness.

[0125] As Figure 7 shown, it is a schematic diagram of the complete process of an embodiment of the text translation method provided by this application. In Figure 7In the illustrated embodiment, for the passage or sentence to be translated, entity objects such as names of people or places can be extracted, and knowledge graph technology can be used to retrieve the relationships between people or the relationships between people and places to obtain relationship knowledge. At the same time, idioms in the passage or sentence to be translated can be determined to obtain the corresponding translations of the idioms in the idiom database, resulting in idiom translation pairs. On this basis, through screening by the title and the content of the text, the reference original text with the highest similarity to the passage or sentence to be translated and the corresponding reference translation are determined, and the target large language model with the best translation effect is determined among multiple large language models. The translation is performed by combining the obtained relationship knowledge, idiom translation pairs, and the target large language model to obtain the translation result. Among them, the title of the passage or sentence to be translated can be vectorized first to obtain the title vector, and semantic recall and dual-path recall are performed in the existing translation text database to obtain multiple recalled titles; at the same time as obtaining multiple recalled titles, the recall weight of each recalled title needs to be determined to screen out some recalled titles that meet the requirements according to the recall weight. After screening out some recalled titles that meet the requirements, the corresponding text content of these recalled titles can be further determined, and re-ranked by combining the title and the text content to screen out the only reference original text and the corresponding reference translation with the highest similarity or relevance to the passage or sentence to be translated among multiple articles or books. At this time, multiple large language models can be called to translate the reference original text to obtain the translation results of model one, model two... model N; the multiple translation results are compared with the reference translation, and the model X with the highest similarity to the reference translation is selected from them, and then the passage or sentence to be translated is translated by combining the relationship knowledge, idiom translation pairs, and model X to obtain the translation result. In Figure 7 it, a single-article database corresponding to the passage or sentence to be translated can also be created, and the passage or sentence to be translated, its corresponding translation result, relationship knowledge, and newly determined translation pairs are stored in the single-article database; when the translation of the passage or sentence to be translated is completed, after manually confirming that the content in the single-article database is correct, the single-article database is imported into the public library to update the public database, where the public database can include a translation text database and an idiom database, the translation text database includes the original text and the corresponding translation, and the idiom database includes idioms and their corresponding translations.

[0126] In a specific embodiment, when translating derivative works of the same IP, even if there are multiple types of derivative works such as novels, comics, animations, movies, and TV dramas, the names of characters, place names, character relationships, character information, etc. in these derivative works are usually consistent; therefore, the existing translation results of the original IP can be referred to, and based on the existing translations, other derivative works of the IP can be assisted in translation, which can effectively improve the translation efficiency and accuracy. In other embodiments, when translating a sequel of the same IP, the names of characters, place names, character relationships, character information, etc. in novels of the same series are related, and combining the existing translation results can better translate the sequel and maintain the unity of the translation style.

[0127] To facilitate the better implementation of the text translation method provided by the embodiments of the present application, the embodiments of the present application also provide a text translation device based on the above large language model. The meanings of the nouns are the same as those in the above text translation method, and the specific implementation details can be referred to the descriptions in the above method embodiments.

[0128] Please refer to Figure 8 , Figure 8 which is the structural schematic diagram of the text translation device provided by the embodiments of the present application.

[0129] The title recall module 801 is used to perform two-way recall in a preset translation text database based on the title of the text to be translated, and obtain a set of recalled titles, where the set of recalled titles includes multiple recalled titles.

[0130] The recall weight calculation module 802 is used to determine the recall weight corresponding to each recalled title in the set of recalled titles, and obtain a plurality of recall weights.

[0131] The recalled title screening module 803 is used to screen at least one recalled title that meets the preset title screening strategy from the set of recalled titles based on the plurality of recall weights and the preset title screening strategy.

[0132] The original text determination module 804 is used to determine the reference original text corresponding to at least one recalled title, and obtain at least one reference original text.

[0133] The original text screening module 805 is used to screen the target reference original text that meets the preset original text screening strategy from at least one reference original text based on at least one reference original text and the preset original text screening strategy.

[0134] The large language model screening module 806 is used to determine the target reference translation corresponding to the target reference original text, and determine the target large language model that meets the preset large language model screening strategy from the preset multiple large language models based on the target reference translation.

[0135] A translation module 807 is configured to translate the text to be translated based on a target large language model to obtain a translation result.

[0136] This application provides a text translation device. Compared with the prior art, this application uses two rounds of screening through the title and the text content to screen out the reference original text and the reference translation text with the highest similarity to the text to be translated in the existing translation text database including the original text and the translation. While gradually narrowing the screening scope, it also uses dual-channel recall during screening to improve the accuracy and coverage of the screening results. At the same time, it further screens out a target large language model with better translation effect for the text to be translated from multiple large language models using the reference original text and the reference translation text. Using the screened target large language model to translate the text to be translated can effectively improve the accuracy of the translation result.

[0137] In some embodiments of this application, the title recall module 801 may specifically be configured to: determine the target title of the text to be translated, and vectorize the target title to obtain a title vector; obtain a preset translation text database, where the translation text database includes the original translation text and the corresponding translation; perform semantic recall and keyword recall in the translation text database based on the title vector to obtain a set of recalled titles, and the set of recalled titles includes multiple recalled titles.

[0138] In some embodiments of this application, the set of recalled titles includes a target recalled title; the recall weight calculation module 802 may specifically be configured to: determine the recall score of the target recalled title during the recall process; respectively determine whether the target recalled title and the target title are titles of the same series;

[0139] If the target recalled title and the target title are titles of the same series, determine the title score of the target recalled title as the first title score; if the target recalled title and the target title are not titles of the same series, determine the title similarity between the target recalled title and the target title; determine the title score of the target recalled title as the second title score based on the title similarity;

[0140] Based on the recall score, the first title score, and the second title score, determine the target recall weight corresponding to the target recalled title, so as to obtain the recall weight corresponding to each recalled title in the set of recalled titles.

[0141] In some embodiments of the present application, the original text screening module 805 may specifically be configured to: calculate the text similarity between at least one reference original text and the text to be translated to obtain at least one text similarity; reorder the at least one text similarity in descending order to obtain a sorted text similarity sequence; screen out the first text similarity in the text similarity sequence whose text similarity is greater than a preset text similarity threshold, and determine the target reference original text corresponding to the first text similarity; or screen out the second text similarity in the text similarity sequence that is ranked before a preset rank, and determine the target reference original text corresponding to the second text similarity.

[0142] In some embodiments of the present application, the large language model screening module 806 may specifically be configured to: determine the target reference translation corresponding to the target reference original text; use a preset plurality of large language models to translate the target reference original text to obtain a plurality of initial translation results; calculate the translation text similarity between each of the plurality of initial translation results and the target reference translation respectively to obtain a plurality of translation text similarities; and determine the target large language model among the plurality of large language models according to the plurality of translation text similarities.

[0143] In some embodiments of the present application, the translation module 807 may specifically be configured to: obtain the entity objects in the text to be translated, and determine the relationship knowledge of the text to be translated based on the entity objects; determine the target idiom in the text to be translated, and determine the idiom translation text corresponding to the target idiom to obtain an idiom translation pair; and combine the relationship knowledge of the text to be translated, the idiom translation pair, and the target large language model to translate the text to be translated to obtain a translation result.

[0144] The translation device based on a large language model provided by the present application further includes a database update module, which is specifically configured to: create a single-article database corresponding to the text to be translated; determine a new translation pair corresponding to the text to be translated according to the text to be translated and the corresponding translation result; store the translation result, relationship knowledge, and new translation pair corresponding to the text to be translated into the single-article database; correct and confirm the content in the single-article database, and update the confirmed single-article database to the translation text database.

[0145] Please refer to Figure 9 , Figure 9 which is a schematic structural diagram of the electronic device provided in the embodiments of the present application.

[0146] The electronic device may include components such as a processor 901 with one or more processing cores, a memory 902 with one or more computer-readable storage media, a power supply 903, and an input unit 904. Those skilled in the art can understand that Figure 9The electronic device structure shown does not constitute a limitation on the electronic device, and may include more or fewer components than shown, or combine certain components, or have different component arrangements. Among them:

[0147] The processor 901 is the control center of the electronic device, connecting various parts of the entire electronic device through various interfaces and lines. By running or executing software programs and / or modules stored in the memory 902, and by calling the data stored in the memory 902, it executes various functions of the electronic device and processes data. Optionally, the processor 901 may include one or more processing cores; optionally, the processor 901 may integrate an application processor and a modem processor. Among them, the application processor mainly processes the operating system, user interface, application programs, etc., and the modem processor mainly processes wireless communication. It can be understood that the above-mentioned modem processor may not be integrated into the processor 901 either.

[0148] The memory 902 can be used to store software programs and modules. The processor 901 executes various functional applications and data processing by running the software programs and modules stored in the memory 902. The memory 902 mainly includes a program storage area and a data storage area. Among them, the program storage area can store the operating system, application programs required for at least one function (such as the sound playback function, image playback function, etc.); the data storage area can store data created according to the use of the electronic device. In addition, the memory 902 may include high-speed random access memory, and may also include non-volatile memory, such as at least one magnetic disk storage device, flash memory device, or other volatile solid-state storage devices. Correspondingly, the memory 902 may also include a memory controller to provide the processor 901 with access to the memory 902.

[0149] The electronic device also includes a power supply 903 that powers each component. Optionally, the power supply 903 can be logically connected to the processor 901 through a power management system, so as to realize functions such as management of charging, discharging, and power consumption management through the power management system. The power supply 903 may also include any components such as one or more DC or AC power supplies, a recharge system, a power failure detection circuit, a power converter or inverter, and a power status indicator.

[0150] The electronic device may also include an input unit 904, which can be used to receive input digital or character information, and generate keyboard, mouse, joystick, optical or trackball signal inputs related to user settings and function controls.

[0151] Although not shown, the electronic device may further include a display unit, an image acquisition component, etc., which will not be elaborated herein. Specifically, in this embodiment, the processor 901 in the electronic device will load the executable code corresponding to one or more computer programs into the memory 902 according to the following instructions, and the processor 901 will execute the steps in the text translation method provided in this application, such as:

[0152] Perform a two-way recall in a preset translation text database based on the title of the text to be translated to obtain a set of recalled titles, where the set of recalled titles includes multiple recalled titles; determine the recall weights corresponding to each recalled title in the set of recalled titles to obtain multiple recall weights; based on the multiple recall weights and a preset title screening strategy, screen out at least one recalled title that meets the preset title screening strategy from the set of recalled titles; determine the reference original texts corresponding to the at least one recalled title to obtain at least one reference original text;

[0153] Based on the at least one reference original text and a preset original text screening strategy, screen out the target reference original text that meets the preset original text screening strategy from the at least one reference original text; determine the target reference translation corresponding to the target reference original text, and based on the target reference translation, determine the target large language model that meets the preset large language model screening strategy from a preset multiple large language models; based on the target large language model, translate the text to be translated to obtain a translation result.

[0154] Those of ordinary skill in the art can understand that all or part of the steps in the various methods of the above embodiments can be completed by instructions or by controlling related hardware through instructions. The instructions can be stored in a computer-readable storage medium and loaded and executed by a processor.

[0155] Therefore, an embodiment of this application provides a computer-readable storage medium, on which a computer program is stored. The computer program is loaded by a processor to execute the steps in any text translation method provided in the embodiments of this application.

[0156] For the specific implementation of the above operations, reference may be made to the previous embodiments, which will not be elaborated herein.

[0157] Among them, the computer-readable storage medium may include: read-only memory (ROM, Read Only Memory), random access memory (RAM, Random Access Memory), magnetic disk or optical disk, etc. Since the computer program stored in the computer-readable storage medium can execute the steps in the text translation method provided in the embodiments of this application, the beneficial effects that can be achieved by the text translation method provided in the embodiments of this application can be realized. For details, see the previous embodiments, which will not be elaborated herein.

[0158] When the computing device in the embodiments of the present application is a terminal device, the embodiments of the present application further provide a terminal device. As Figure 10 shown, for the sake of convenience of description, only the parts related to the embodiments of the present application are shown. For the specific technical details not disclosed, please refer to the method part of the embodiments of the present application. The terminal device may be any terminal device including a mobile phone, a tablet computer, a personal digital assistant (PDA), a point of sales (POS) device, an in-vehicle computer, etc. Taking the terminal device as a mobile phone as an example:

[0159] Figure 10 The block diagram of a part of the structure of the mobile phone related to the terminal device provided by the embodiments of the present application is shown. Referring to Figure 10 , the mobile phone includes: a radio frequency (RF) circuit 1010, a memory 1020, an input unit 1030, a display unit 1040, a sensor 1050, an audio circuit 1060, a wireless fidelity (WiFi) module 1070, a processor 1080, and a power supply 1090 and other components. Those skilled in the art can understand that Figure 10 the structure of the mobile phone shown in

[0160] does not constitute a limitation on the mobile phone, and may include more or fewer components than shown, or combine certain components, or have different component arrangements.

[0160] Next, the various components of the mobile phone will be specifically introduced in conjunction with Figure 10 :

[0161] The RF circuit 1010 can be used for receiving and transmitting information or signals during communication. Specifically, after receiving the downlink information from the base station, it is sent to the processor 1080 for processing. Additionally, the uplink data designed is sent to the base station. Generally, the RF circuit 1010 includes, but is not limited to, an antenna, at least one amplifier, a transceiver, a coupler, a low noise amplifier (LNA), a duplexer, etc. In addition, the RF circuit 1010 can also communicate with the network and other devices through wireless communication. The above wireless communication can use any communication standard or protocol, including but not limited to the Global System of Mobile communication (GSM), General Packet Radio Service (GPRS), Code Division Multiple Access (CDMA), Wideband Code Division Multiple Access (WCDMA), Long Term Evolution (LTE), email, Short Messaging Service (SMS), etc.

[0162] The memory 1020 can be used to store software programs and modules. The processor 1080 executes various functional applications and data processing of the mobile phone by running the software programs and modules stored in the memory 1020. The memory 1020 mainly includes a program storage area and a data storage area. Among them, the program storage area can store an operating system, application programs required for at least one function (such as a sound playback function, an image playback function, etc.); the data storage area can store data created according to the use of the mobile phone (such as audio data, a phone book, etc.). In addition, the memory 1020 can include a high-speed random access memory and can also include a non-volatile memory, such as at least one magnetic disk storage device, a flash memory device, or other non-volatile solid-state storage devices.

[0163] The input unit 1030 can be used to receive input digital or character information, and generate key signal inputs related to the user settings and function controls of the mobile phone. Specifically, the input unit 1030 can include a touch panel 1031 and other input devices 1032. The touch panel 1031, also known as a touch screen, can collect touch operations of the user thereon or nearby (such as operations of the user using any suitable object or accessory such as a finger or a stylus on or near the touch panel 1031), and drive corresponding connection devices according to a preset program. Optionally, the touch panel 1031 can include two parts: a touch detection device and a touch controller. Among them, the touch detection device detects the touch position of the user, detects the signal brought by the touch operation, and transmits the signal to the touch controller; the touch controller receives the touch information from the touch detection device, converts it into contact coordinates, and then sends it to the processor 1080, and can receive and execute commands sent by the processor 1080. In addition, the touch panel 1031 can be implemented in multiple types such as resistive, capacitive, infrared, and surface acoustic wave. In addition to the touch panel 1031, the input unit 1030 can also include other input devices 1032. Specifically, the other input devices 1032 can include, but are not limited to, one or more of a physical keyboard, function keys (such as volume control keys, power on / off keys, etc.), a trackball, a mouse, a joystick, etc.

[0164] The display unit 1040 can be used to display information input by the user or information provided to the user and various menus of the mobile phone. The display unit 1040 can include a display panel 1041. Optionally, the display panel 1041 can be configured in the form of a liquid crystal display (LCD), an organic light-emitting diode (OLED), etc. Further, the touch panel 1031 can cover the display panel 1041. When the touch panel 1031 detects a touch operation thereon or nearby, it is transmitted to the processor 1080 to determine the type of touch event. Subsequently, the processor 1080 provides a corresponding visual output on the display panel 1041 according to the type of touch event. Although in Figure 9 the touch panel 1031 and the display panel 1041 are implemented as two independent components to realize the input and input functions of the mobile phone, in some embodiments, the touch panel 1031 and the display panel 1041 can be integrated to realize the input and output functions of the mobile phone.

[0165] The mobile phone may further include at least one sensor 1050, such as a light sensor, a motion sensor, and other sensors. Specifically, the light sensor may include an ambient light sensor and a proximity sensor. Among them, the ambient light sensor can adjust the brightness of the display panel 1041 according to the brightness of the ambient light, and the proximity sensor can turn off the display panel 1041 and / or the backlight when the mobile phone is moved to the ear. As a kind of motion sensor, the accelerometer sensor can detect the magnitude of acceleration in various directions (generally three axes), and can detect the magnitude and direction of gravity when stationary, and can be used for applications that identify the posture of the mobile phone (such as horizontal and vertical screen switching, related games, magnetometer posture calibration), vibration recognition related functions (such as pedometer, tapping), etc.; as for other sensors that the mobile phone can also be configured with, such as gyroscopes, barometers, hygrometers, thermometers, infrared sensors, etc., they will not be elaborated here.

[0166] The audio circuit 1060, the speaker 1061, and the microphone 1062 can provide an audio interface between the user and the mobile phone. The audio circuit 1060 can transmit the electrical signal converted from the received audio data to the speaker 1061, and the speaker 1061 converts it into a sound signal for output; on the other hand, the microphone 1062 converts the collected sound signal into an electrical signal, which is received by the audio circuit 1060 and converted into audio data, and then the audio data is output to the processor 1080 for processing, and then sent to another mobile phone, for example, through the RF circuit 1010, or the audio data is output to the memory 1020 for further processing.

[0167] Wi-Fi belongs to short-distance wireless transmission technology. The mobile phone can help users send and receive emails, browse the web, and access streaming media through the Wi-Fi module 1070, which provides users with wireless broadband Internet access. Although Figure 10 the Wi-Fi module 1070 is shown, it can be understood that it does not belong to the essential composition of the mobile phone and can be omitted completely within the scope of not changing the essence of the invention according to needs.

[0168] The processor 1080 is the control center of the mobile phone, connecting various parts of the entire mobile phone through various interfaces and lines. By running or executing software programs and / or modules stored in the memory 1020, and calling data stored in the memory 1020, it executes various functions of the mobile phone and processes data, thereby monitoring the mobile phone as a whole. Optionally, the processor 1080 may include one or more processing units; optionally, the processor 1080 may integrate an application processor and a modem processor. Among them, the application processor mainly processes the operating system, user interface, and application programs, etc., and the modem processor mainly processes wireless communication. It can be understood that the above-mentioned modem processor may not be integrated into the processor 1080 either.

[0169] The mobile phone further includes a power source 1090 (such as a battery) for supplying power to each component. Optionally, the power source can be logically connected to the processor 1080 through a power management system, so as to manage functions such as charging, discharging, and power consumption management through the power management system. Although not shown, the mobile phone may further include a camera, a Bluetooth module, etc., which will not be elaborated here.

[0170] In the embodiment of the present application, the processor 1080 included in the mobile phone is further configured to control the execution of the text translation method flow executed by the above text translation device.

[0171] The embodiment of the present application further provides a server. Please refer to Figure 11 , Figure 11 FIG. is a schematic structural diagram of a server provided by the embodiment of the present application. The server 1100 may vary greatly due to different configurations or performances, and may include one or more central processing units (English full name: central processing units, English abbreviation: CPU) 1122 (for example, one or more processors) and a memory 1132, and one or more storage media 1130 (for example, one or more mass storage devices) for storing application programs 1142 or data 1144. Among them, the memory 1132 and the storage media 1130 may be transient storage or persistent storage. The program stored in the storage media 1130 may include one or more modules (not shown in the figure), and each module may include a series of instruction operations on the server. Further, the central processing unit 1122 may be configured to communicate with the storage media 1130 and execute a series of instruction operations in the storage media 1130 on the server 1100.

[0172] The server 1100 may further include one or more power sources 1126, one or more wired or wireless network interfaces 1150, one or more input / output interfaces 1158, and / or one or more operating systems 1141, such as Windows Server, Mac OS X, Unix, Linux, FreeBSD, and so on.

[0173] The steps in the text translation method in the above embodiment may be based on the Figure 11 structure of the server 1100 shown. For example, the central processing unit 1122 executes the following operations by calling the instructions in the memory 1132:

[0174] Perform two-way recall in a preset translation text database based on the title of the text to be translated to obtain a set of recalled titles, where the set of recalled titles includes multiple recalled titles; determine the recall weights corresponding to each recalled title in the set of recalled titles to obtain multiple recall weights; based on the multiple recall weights and a preset title screening strategy, screen out at least one recalled title that meets the preset title screening strategy from the set of recalled titles; determine the reference original texts corresponding to the at least one recalled title to obtain at least one reference original text; based on the at least one reference original text and a preset original text screening strategy, screen out the target reference original text that meets the preset original text screening strategy from the at least one reference original text; determine the target reference translation corresponding to the target reference original text, and based on the target reference translation, determine the target large language model that meets the preset large language model screening strategy among the preset multiple large language models; based on the target large language model, translate the text to be translated to obtain a translation result.

[0175] In the above embodiments, the descriptions of the various embodiments have their own emphases. For the parts not detailed in a certain embodiment, reference may be made to the relevant descriptions of other embodiments.

[0176] Those skilled in the art can clearly understand that for the convenience and brevity of description, the specific working processes of the systems, devices, and modules described above can refer to the corresponding processes in the foregoing method embodiments and will not be elaborated herein.

[0177] In the several embodiments provided by the embodiments of the present application, it should be understood that the disclosed systems, devices, and methods can be implemented in other ways. For example, the device embodiments described above are only illustrative. For example, the division of the modules is only a logical function division, and there can be other division methods in actual implementation. For example, multiple modules or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed couplings or direct couplings or communication connections to each other can be indirect couplings or communication connections through some interfaces, devices, or modules, and can be in electrical, mechanical, or other forms.

[0178] The modules described as separate components may or may not be physically separated, and the components displayed as modules may or may not be physical modules, that is, they can be located in one place or distributed to multiple network modules. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0179] In addition, in each embodiment of the present application, each functional module can be integrated into one processing module, or each module can exist physically alone, or two or more modules can be integrated into one module. The above-mentioned integrated modules can be implemented in the form of hardware or in the form of software functional modules. If the integrated module is implemented in the form of a software functional module and sold or used as an independent product, it can be stored in a computer-readable storage medium.

[0180] In the above embodiments, it can be implemented in whole or in part by software, hardware, firmware, or any combination thereof. When implemented using software, it can be implemented in whole or in part in the form of a computer program product.

[0181] The computer program product includes one or more computer instructions. When the computer program is loaded and executed on a computer, it generates all or part of the processes or functions described in the embodiments of the present application. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions can be stored in a computer-readable storage medium, or transmitted from one computer-readable storage medium to another computer-readable storage medium. For example, the computer instructions can be transmitted from a website, computer, server, or data center to another website, computer, server, or data center by wire (such as coaxial cable, optical fiber, digital subscriber line (DSL)) or wireless (such as infrared, wireless, microwave, etc.). The computer-readable storage medium can be any available medium that a computer can store, or a data storage device such as a server or data center that includes one or more integrated available media. The available medium can be a magnetic medium (such as a floppy disk, hard disk, magnetic tape), an optical medium (such as a DVD), or a semiconductor medium (such as a solid-state drive (SSD)).

[0182] The technical solutions provided in the embodiments of the present application have been introduced in detail above. In the embodiments of the present application, specific examples are used to elaborate on the principles and implementation manners of the embodiments of the present application. The descriptions of the above embodiments are only used to help understand the method and its core idea of the embodiments of the present application; at the same time, for those of ordinary skill in the art, according to the idea of the embodiments of the present application, there will be changes in the specific implementation manners and application scopes. In summary, the content of this specification should not be construed as a limitation to the embodiments of the present application.

[0183] It should be noted that when the above embodiments of the present application are applied to specific products or technologies and involve relevant user data, user permission or consent needs to be obtained, and the collection, use, and processing of relevant data need to comply with relevant laws, regulations, and standards of relevant countries and regions.

Claims

1. A text translation method, characterized in that, including: Performing two-way recall in a preset translation text database based on the title of the text to be translated to obtain a set of recalled titles, where the set of recalled titles includes multiple recalled titles; Determining the recall weights corresponding to each recalled title in the set of recalled titles to obtain multiple recall weights; Based on the multiple recall weights and a preset title screening strategy, screening at least one recalled title that meets the preset title screening strategy from the set of recalled titles; Determining the reference original texts corresponding to the at least one recalled title to obtain at least one reference original text; Based on the at least one reference original text and a preset original text screening strategy, screening a target reference original text that meets the preset original text screening strategy from the at least one reference original text; Determining the target reference translation corresponding to the target reference original text, and based on the target reference translation, determining a target large language model that meets the preset large language model screening strategy among a preset plurality of large language models; Based on the target large language model, translating the text to be translated to obtain a translation result.

2. The text translation method according to claim 1, characterized in that The performing two-way recall in a preset translation text database based on the title of the text to be translated to obtain a set of recalled titles includes: Determining the target title of the text to be translated and vectorizing the target title to obtain a title vector; Obtaining a preset translation text database, where the translation text database includes translation original texts and corresponding translation translations; Performing semantic recall and keyword recall in the translation text database based on the title vector to obtain a set of recalled titles, where the set of recalled titles includes multiple recalled titles.

3. The text translation method according to claim 2, characterized in that, The set of recalled titles includes a target recalled title; The determining the recall weights corresponding to each recalled title in the set of recalled titles to obtain multiple recall weights includes: Determining the recall score of the target recalled title during the recall process; Respectively determining whether the target recalled title and the target title are titles of the same series; If the target recalled title and the target title are titles of the same series, determining the title score of the target recalled title as the first title score; If the target recalled title and the target title are not titles of the same series, determining the title similarity between the target recalled title and the target title; Based on the title similarity, determining the title score of the target recalled title as the second title score; Based on the recall score, the first title score, and the second title score, determining the target recall weight corresponding to the target recalled title to obtain the recall weights corresponding to each recalled title in the set of recalled titles.

4. The text translation method according to claim 2, characterized in that The screening a target reference original text that meets the preset original text screening strategy from the at least one reference original text based on the at least one reference original text and a preset original text screening strategy includes: Calculating the text similarity between the at least one reference original text and the text to be translated to obtain at least one text similarity; Re-sorting the at least one text similarity in descending order to obtain a sorted text similarity sequence; Filter out the first text similarity in the text similarity sequence where the text similarity is greater than a preset text similarity threshold, and determine the target reference original text corresponding to the first text similarity; Or filter out the second text similarity in the text similarity sequence that is ranked before a preset ranking, and determine the target reference original text corresponding to the second text similarity.

5. The text translation method according to claim 2, characterized in that Determining the target reference translation corresponding to the target reference original text, and determining a target large language model that meets the preset large language model screening strategy among a preset plurality of large language models, includes: Determine the target reference translation corresponding to the target reference original text; Use a preset plurality of large language models to translate the target reference original text to obtain a plurality of initial translation results; Calculate the translation text similarity between each of the plurality of initial translation results and the target reference translation respectively to obtain a plurality of translation text similarities; Determine a target large language model among the plurality of large language models according to the plurality of translation text similarities.

6. The text translation method according to claim 2, characterized in that Based on the target large language model, translating the text to be translated to obtain a translation result, includes: Obtain the entity objects in the text to be translated, and determine the relationship knowledge of the text to be translated based on the entity objects; Determine the target idiom in the text to be translated, and determine the idiom translation text corresponding to the target idiom to obtain an idiom translation pair; Combine the relationship knowledge, idiom translation pair of the text to be translated and the target large language model to translate the text to be translated to obtain a translation result.

7. The text translation method according to claim 6, characterized in that, The method further includes: Create a single-article database corresponding to the text to be translated; Determine a new translation pair corresponding to the text to be translated according to the text to be translated and the corresponding translation result; Store the translation result, the relationship knowledge and the new translation pair corresponding to the text to be translated into the single-article database; Check and confirm the content in the single-article database, and update the confirmed single-article database to the translation text database.

8. A text translation device, characterized in that, The device includes: A title recall module, configured to perform dual-channel recall in a preset translation text database based on the title of the text to be translated to obtain a set of recalled titles, and the set of recalled titles includes a plurality of recalled titles; A recall weight calculation module, configured to determine the recall weight corresponding to each recalled title in the set of recalled titles to obtain a plurality of recall weights; A recalled title screening module, configured to screen out at least one recalled title that meets the preset title screening strategy from the set of recalled titles based on the plurality of recall weights and the preset title screening strategy; An original text determination module, configured to determine the reference original text corresponding to the at least one recalled title to obtain at least one reference original text; An original text screening module, configured to screen out a target reference original text that meets the preset original text screening strategy from the at least one reference original text based on the at least one reference original text and the preset original text screening strategy; The large language model screening module is used to determine the target reference translation corresponding to the target reference original text, and determine the target large language model that meets the preset large language model screening strategy among a preset plurality of large language models; The translation module is used to translate the text to be translated based on the target large language model to obtain a translation result.

9. An electronic device, characterized in that, It includes a memory and a processor. The memory stores a computer program, and the processor is used to run the computer program in the memory to execute the steps in the text translation method according to any one of claims 1 to 7.

10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores multiple instructions, and the instructions are suitable for being loaded by the processor to execute the steps in the text translation method according to any one of claims 1 to 7.