Text translation methods, methods for acquiring text translation models, apparatus, devices, and computer programs

The method and device enhance text translation accuracy by determining probabilities and reliability through text feature analysis and model updating, addressing the challenge of improving translation precision.

JP7870354B2Active Publication Date: 2026-06-04TENCENT TECHNOLOGY (SHENZHEN) CO LTD

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
TENCENT TECHNOLOGY (SHENZHEN) CO LTD
Filing Date
2023-06-19
Publication Date
2026-06-04

AI Technical Summary

Technical Problem

Improving the accuracy of text translation is a pressing technical challenge in various scenarios.

Method used

A method and device for determining probabilities and reliability of text translation using text features and data pairs, and updating a text translation model based on confidence and match analysis.

Benefits of technology

Enhances the accuracy of text translation by leveraging probability and reliability assessments to improve the precision of language conversion.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007870354000024
    Figure 0007870354000024
  • Figure 0007870354000025
    Figure 0007870354000025
  • Figure 0007870354000026
    Figure 0007870354000026
Patent Text Reader

Abstract

This application discloses a method for translating text, a method for obtaining a text translation model, an apparatus, a device and a medium, which includes the steps of: determining at least one first probability based on a first text feature (201); obtaining at least one target data pair that matches the first text feature (202); determining a confidence and a match degree of the at least one target data pair (203); determining at least one second probability based on the confidence and the match degree of the at least one target data pair (204); and determining a translation text corresponding to the first text based on the at least one first probability and the at least one second probability (205).
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] [Cross - Reference to Related Applications] This application claims the benefit of priority to Chinese Patent Application No. 202211049110.8, filed on August 30, 2022, with the title "Text Translation Method, Method for Obtaining a Text Translation Model, Apparatus, Device, and Medium", the entire content of which is incorporated herein by reference.

[0002] [Technical Field] Embodiments of this application relate to the field of computer technology, and particularly to text translation methods, methods for obtaining text translation models, apparatuses, devices, and media.

Background Art

[0003] With the development of computer technology, text translation has come to be widely used in various scenarios. By performing text translation, text can be translated from one language to another. Improving the accuracy of text translation has become a technical problem that needs to be solved urgently.

Summary of the Invention

Means for Solving the Problems

[0004] Embodiments of this application provide a text translation method, a method for obtaining a text translation model, an apparatus, a device, and a storage medium that can improve the accuracy of text translation. The technical means are as follows.

[0005] According to one aspect, embodiments of this application provide a text translation method applied to a computer device. The method includes Determining at least one first probability based on the first text feature, wherein the first text feature is a text feature of a first text, the first text is a text in a first language, and the at least one first probability is used to indicate the probability that the first text is translated into each candidate text among at least one candidate text, and all of the at least one candidate text are texts in a second language; Obtaining at least one target data pair that matches the first text feature, wherein any one target data pair includes one second text feature and a standard translation text of one second text, the second text feature is a text feature of the second text, the second text is a text in the first language, and the standard translation text is a text in the second language; Determining the reliability and matching degree of the at least one target data pair, wherein the reliability of any one target data pair is used to indicate the degree of reliability of the any one target data pair, and the matching degree of the any one target data pair is used to indicate the similarity between the second text feature and the first text feature in the any one target data pair; Determining at least one second probability based on the reliability and matching degree of the at least one target data pair, wherein the at least one second probability is used to indicate the probability that the first text is translated into each standard translation text in the at least one target data pair; Determining a translation text corresponding to the first text based on the at least one first probability and the at least one second probability.

[0006] According to another aspect, a method for obtaining a text translation model applied to a computer device is provided. The method includes: A step of obtaining a first sample text, a first standard translation text, and an initial text translation model, wherein the first sample text is a text in a first language, and the first standard translation text is a text obtained by translating the first sample text into a second language. The steps include processing a first sample text feature using the initial text translation model to obtain at least one first sample probability, wherein the first sample text feature is a text feature of the first sample text, and the at least one first sample probability is used to indicate the probability that the first sample text is translated into each of at least one candidate texts, and each of the at least one candidate texts is a text in the second language, A step of obtaining at least one sample data pair that matches the first sample text feature, wherein any one sample data pair includes one second sample text feature and one second standard translation text, the second sample text feature is a text feature of the second sample text, the second sample text is text in the first language, and the second standard translation text is text obtained by translating the second sample text into the second language, A step of determining the confidence and match of at least one sample data pair, wherein the confidence of any one sample data pair is used to indicate the degree of reliability of the any one sample data pair, and the match of any one sample data pair is used to indicate the similarity between a second sample text feature and a first sample text feature in any one sample data pair. A step of determining at least one second sample probability based on the confidence and match of the at least one sample data pair, wherein the at least one second sample probability is used to indicate the probability that the first sample text is translated into each second standard translation text in the at least one sample data pair, A step of determining a predicted translation text corresponding to the first sample text based on the at least one first sample probability and the at least one second sample probability, The process includes the step of updating the initial text translation model based on the difference between the predicted translation text and the first standard translation text to obtain a target text translation model.

[0007] Another aspect is the provision of a text translation device to be placed on a computer device. The device is A decision module configured to perform the step of determining at least one first probability based on a first text feature, wherein the first text feature is a text feature of a first text, the first text is a text in a first language, and the at least one first probability is used to indicate the probability that the first text is translated into each of at least one candidate texts, and each of the at least one candidate texts is a text in a second language, An acquisition module configured to perform the step of acquiring at least one target data pair that matches the first text feature, wherein any one target data pair includes one second text feature and one standard translation text of the second text, the second text feature is a text feature of the second text, the second text is text in the first language, and the standard translation text is text in the second language, The decision module is configured to perform a step of determining the confidence and match of at least one target data pair, wherein the confidence of any one target data pair is used to indicate the degree of reliability of any one target data pair, and the match of any one target data pair is used to indicate the similarity between a second text feature in any one target data pair and a first text feature, The decision module is configured to perform a step of determining at least one second probability based on the confidence and match of the at least one target data pair, wherein the at least one second probability is used to indicate the probability that the first text is translated into each standard translation text in the at least one target data pair, The system includes a decision module configured to further perform the step of determining a translated text corresponding to the first text based on the at least one first probability and the at least one second probability.

[0008] In other respects, the present invention provides a device for acquiring text translation models to be placed on a computer device. An acquisition module configured to perform the steps of acquiring a first sample text, a first standard translation text, and an initial text translation model, wherein the first sample text is text in a first language, and the first standard translation text is text obtained by translating the first sample text into a second language, A decision module configured to perform the steps of processing a first sample text feature using the initial text translation model and obtaining at least one first sample probability, wherein the first sample text feature is a text feature of the first sample text, and the at least one first sample probability is used to indicate the probability that the first sample text is translated into each of at least one candidate texts, and each of the at least one candidate texts is a text in the second language, The acquisition module is configured to perform the step of acquiring at least one sample data pair that matches the first sample text feature, wherein any one of the sample data pairs includes one second sample text feature and one second standard translation text, the second sample text feature is a text feature of the second sample text, the second sample text is text in the first language, and the second standard translation text is text in which the second sample text has been translated into the second language, The decision module is configured to perform a step of determining the confidence and match of at least one sample data pair, wherein the confidence of any one sample data pair is used to indicate the degree of reliability of the any one sample data pair, and the match of any one sample data pair is used to indicate the similarity between a second sample text feature and a first sample text feature in any one sample data pair. The decision module is configured to perform a step of determining at least one second sample probability based on the confidence and match of the at least one sample data pair, wherein the at least one second sample probability is used to indicate the probability that the first sample text is translated into each second standard translation text in the at least one sample data pair, The decision module is configured to further perform the step of determining a predicted translation text corresponding to the first sample text based on the at least one first sample probability and the at least one second sample probability, The system includes an update module configured to perform the step of updating the initial text translation model and obtaining a target text translation model based on the difference between the predicted translation text and the first standard translation text.

[0009] In other words, the present invention provides a computer device including a processor and memory. The memory stores at least one computer program, and the computer device is made to implement the text translation method or method for obtaining a text translation model described above by loading and executing the at least one computer program by the processor.

[0010] In other words, the present invention further provides a computer-readable storage medium in which at least one computer program is stored. The at least one computer program is loaded and executed by a processor to cause the computer to implement any of the above-described text translation methods or methods for obtaining text translation models.

[0011] In another aspect, the present invention further provides a computer program product including a computer program or computer command. The computer program or computer command is loaded and executed by a processor to cause a computer to implement the text translation method or method for obtaining a text translation model described in any of the above. [Brief explanation of the drawing]

[0012] [Figure 1] This is a schematic diagram of an implementation environment according to the embodiment of this application. [Figure 2] This is a flowchart of a text translation method according to an embodiment of this application. [Figure 3] This is a schematic diagram of a confidence-based text translation model according to an embodiment of this application. [Figure 4] This is a flowchart illustrating the method for obtaining a text translation model according to the embodiment of this application. [Figure 5] This is a schematic diagram for constructing a noisy data pair according to the embodiment of this application. [Figure 6]This is a schematic diagram for obtaining a sample data pair according to the embodiment of this application. [Figure 7] This is a schematic diagram of a text translation device according to an embodiment of this application. [Figure 8] This is a schematic diagram of a text translation model acquisition device according to an embodiment of this application. [Figure 9] This is a schematic diagram of the server according to the embodiment of this application. [Figure 10] This is a schematic diagram of the terminal according to the embodiment of this application. [Modes for carrying out the invention]

[0013] To further clarify the purpose, technical methods, and advantages of this application, embodiments of this application will be described in more detail below with reference to the attached drawings.

[0014] In some embodiments, the text translation method and method for obtaining a text translation model according to the embodiments of this application may be applied to a variety of scenarios, including but not limited to cloud technology, artificial intelligence, smart transportation, and driver assistance.

[0015] Artificial intelligence (AI) refers to the theories, methods, technologies, and application systems that use digital computers or machines controlled by digital computers to mimic and extend human intelligence, enabling them to perceive their environment, acquire knowledge, and utilize that knowledge to achieve optimal results. In other words, AI is considered a comprehensive technology within computer science, aiming to understand the essence of intelligence and create new intelligent devices that can react in a manner similar to human intelligence. Essentially, AI is the study of the design principles and implementation methods of various intelligent devices, and the technology to give machines the functions of perception, reasoning, and decision-making.

[0016] Artificial intelligence (AI) technology is considered a comprehensive field of study, applicable to a wide range of areas, encompassing both hardware and software-level technologies. Fundamental AI technologies typically include sensors, dedicated AI chips, cloud computing, distributed storage, big data processing technologies, operating / interactive systems, and mechatronics. AI software technologies primarily encompass several areas, including computer vision, speech processing, natural language processing, machine learning / deep learning, autonomous driving, and smart transportation.

[0017] Natural Language Processing (NLP) is considered an important direction in the fields of computer science and artificial intelligence. Various theories and methods are being explored to achieve effective communication between humans and computers using natural language. Natural Language Processing is a science that integrates linguistics, computer science, and mathematics. Therefore, research in this field is closely related to linguistics, as it deals with natural language, that is, the language humans use every day. Natural Language Processing techniques typically include text processing, semantic analysis of words, machine translation, robot quizzes, and knowledge graphs.

[0018] Machine learning (ML) is a discipline encompassing a wide range of fields, including probability theory, statistics, approximation theory, convex analysis, and algorithmic complexity theory. It focuses on enabling computers to mimic or replicate human learning behavior, acquire new knowledge and skills, and reconstruct existing knowledge systems to continuously improve their performance. Machine learning is considered the core of artificial intelligence, a fundamental method for giving computers intelligence, and is applied across various fields of AI. Machine learning and deep learning typically include techniques such as artificial neural networks, trust networks, reinforcement learning, transfer learning, inductive learning, and teaching-learning.

[0019] With the ongoing research and advancement of artificial intelligence (AI) technology, its research and application are progressing in many fields, including general smart homes, smart wearable devices, virtual assistants, smart speakers, smart marketing, autonomous driving, drones, robots, smart healthcare, smart customer service, connected cars, and smart transportation. As the technology develops, AI technology is expected to be applied in even more fields and provide increasingly important value.

[0020] Figure 1 shows a schematic diagram of an implementation environment according to the embodiment of this application. The implementation environment includes a terminal 11 and a server 12.

[0021] The text translation method according to the embodiment of this application may be performed by terminal 11, by server 12, or jointly by terminal 11 and server 12, but is not limited to the embodiments of this application. When the text translation method according to the embodiment of this application is jointly performed by terminal 11 and server 12, server 12 may be responsible for the primary computing tasks and terminal 11 for the secondary computing tasks; or server 12 may be responsible for the secondary computing tasks and terminal 11 for the primary computing tasks; or collaborative computing may be performed between server 12 and terminal 11 using a distributed computing architecture.

[0022] The method for obtaining a text translation model according to the embodiment of this application may be performed by terminal 11, by server 12, or jointly by terminal 11 and server 12, but is not limited to the embodiments of this application. When the text translation method according to the embodiment of this application is jointly performed by terminal 11 and server 12, server 12 may be responsible for the primary computing tasks and terminal 11 for the secondary computing tasks; or server 12 may be responsible for the secondary computing tasks and terminal 11 for the primary computing tasks; or collaborative computing may be performed between server 12 and terminal 11 using a distributed computing architecture.

[0023] The device that performs the text translation method and the device that performs the method for obtaining the text translation model may be the same or different, but are not limited to the embodiments of this application.

[0024] In some embodiments, terminal 11 is any electronic product capable of human-computer interaction with a user via one or more methods such as a keyboard, touch panel, touchscreen, remote control, voice interactive, or handwriting input device. Examples include PCs (Personal Computers), mobile phones, smartphones, PDAs (Personal Digital Assistants), wearable devices, PPCs (Pocket PCs), tablets, smart sensor devices, smart TVs, smart speakers, smart voice interactive devices, smart home appliances, in-car terminals, VR (Virtual Reality) devices, and AR (Augmented Reality) devices. Server 12 may be a single server, a server cluster consisting of multiple servers, or a cloud computing service center. Terminal 11 and server 12 are connected to each other via a wired or wireless network.

[0025] Those skilled in the art will understand that the terminal 11 and server 12 described above are merely examples, and that other existing or potentially emerging terminals or servers, if applicable to this application, are included within the scope of protection of this application and should be included hereby by reference.

[0026] The method according to the embodiment of this application is applicable to a variety of situations.

[0027] For example, in an online translation scenario, the server trains an initial text translation model using the method for acquiring a text translation model according to the embodiment of this application, and deploys the trained target text translation model within the server. The terminal logs into the translation application with a user ID. The server provides services for the translation application. The terminal uses the translation application to send a first text in a first language to be translated to the server. The server receives this first text and, using the text translation method according to the embodiment of this application with the target text translation model, obtains a translated text by translating this first text into a second language, and sends the translated text to the terminal. The terminal uses the translation application to receive and display this translated text. Here, the first language and the second language are different languages. In some embodiments, the first language may also be called the source language, and the second language may also be called the target language.

[0028] Furthermore, for example, in a face-to-face conversation scene, the server trains an initial text translation model using the method for acquiring a text translation model according to the embodiment of this application, and deploys the trained target text translation model within the server. The terminal logs into the translation application with a user ID. The server provides services for the translation application. The terminal uses the translation application to collect audio data belonging to a first language from any speaker, converts the audio data into a first text belonging to the first language, and sends the first text to be translated to the server using the translation application. The server receives the first text and, using the text translation method according to the embodiment of this application with the target text translation model, obtains a translated text that has the same meaning as the first text and has been translated into a second language, and sends that translated text to the terminal. The terminal uses the translation application to receive the translated text, converts the translated text into audio data belonging to a second language so that the played audio data can be heard by the speaker corresponding to the terminal, and plays the converted audio data, thereby realizing a simultaneous interpretation effect and enabling communication between two speakers who speak different languages.

[0029] Embodiments of this application provide a text translation method applicable to the implementation environment shown in Figure 1 above. This text translation method is performed by a computer device. This computer device may be a terminal 11 or a server 12, but is not limited to the embodiments of this application. As shown in Figure 2, the text translation method according to the embodiment of this application includes the following steps 201 to 205.

[0030] Step 201 determines at least one first probability based on the first text feature.

[0031] Here, the first text feature is a text feature of the first text, and the first text is a text in the first language to be translated, but in embodiments of this application, the type of the first language is not limited. Exemplaryly, the first language may be Chinese, English, or the like. The first text may contain one or more characters. The length of the characters contained in the first text may be determined according to experience or actual translation requirements. For example, if the first language is Chinese, the first text may contain one Chinese character or multiple Chinese characters, where multiple Chinese characters may constitute one word or one sentence. At least one first probability is used to indicate the probability that the first text is translated into each of at least one candidate text. In other words, the first probability corresponding to any one of the candidate texts is used to indicate the probability that the first text is translated into any one of the candidate texts.

[0032] The method for obtaining the first text by a computer device may include receiving the first text uploaded by a user on the computer device, converting audio in a first language uploaded by a user into text on the computer device to obtain the first text, or extracting the first text from a web page on the computer device, and is not limited to the embodiments of this application.

[0033] In some embodiments, the method for obtaining the first text by a computer device may further involve the computer device extracting the first text from a target text. The target text refers to the text containing the first text. For example, if the target text is a sentence to be translated, the process of translating the sentence is achieved by sequentially translating each word in the sentence; therefore, the first text is a single word in the target text that is to be translated.

[0034] After acquiring the first text, the server needs to extract features from the first text to obtain the first text features. Based on the obtained first text features, the server can then determine the first probability for each candidate text in the second language. The first text features are used to characterize the first text. In the embodiments of this application, the form of the first text features is not limited, as long as it facilitates identification and processing by a computer device. For example, the form of the first text features may be a vector, a matrix, or the like.

[0035] In some embodiments, the process for extracting features from a first text and obtaining first text features may include the steps of encoding the first text and obtaining encoded features, and decoding the encoded features and obtaining first text features.

[0036] When the server determines a first probability for each candidate text based on a first text feature, each candidate text is in a second language, where the second language is the language of the translation text to be obtained. Unlike the first language, the type of the second language can be flexibly set according to the translation needs, but is not limited to the embodiments of this application. For example, if a translation from Chinese to English is required, the first language is Chinese and the second language is English.

[0037] Each candidate text can be set based on experience or flexibly adjusted according to the application scenario. For example, each candidate text may include text extracted from sentences in the second language whose frequency of occurrence is greater than a frequency threshold, or text extracted from a text library in the second language, etc.

[0038] The first probability of any one candidate text refers to the probability that the translated text of the first text, determined based on the first text features, is one of those candidate texts. Exemplarily, the first probability of any one candidate text being a single number between 0 and 1. Exemplarily, the sum of the first probabilities of each candidate text being a single number may be 1. Exemplarily, the first probabilities of each candidate text being a single number are represented using a bar graph, where each candidate text is represented by a single bar. The height of the bar corresponding to any one candidate text is used to indicate the first probability of that candidate text being a single number.

[0039] In some embodiments, step 201 may be achieved by calling a target text translation model. That is, the target text translation model is called and a first probability corresponding to each of several candidate texts is determined based on a first text feature. The target text translation model is a model for translating text in a first language to text in a second language. In embodiments of this application, the structure of the target text translation model is not limited, as long as it can achieve text translation.

[0040] In some embodiments, the target text translation model includes a first translation submodel, a second translation submodel, and a third translation submodel. The first translation submodel is used to extract features from the text to be translated and to predict a first probability corresponding to each of several candidate texts based on the extracted features. The second translation submodel is used to search for matching data pairs according to the features extracted by the first translation submodel and to determine a second probability corresponding to each of several standard translation texts in each searched data pair. The third translation submodel is used to determine the translation text corresponding to the first text according to the first probability determined by the first translation submodel and the second probability determined by the second translation submodel.

[0041] If the structure of the target text translation model is as described above, the implementation process for calling the target text translation model and determining a first probability corresponding to each of the multiple candidate texts based on a first text feature refers to calling a first translation submodel of the target text translation model and determining a first probability corresponding to each of the multiple candidate texts based on the first text feature. In the embodiments of this application, the type of the first translation submodel is not limited and may be any model having feature extraction and probability determination functions. For example, the first translation submodel may be an NMT (Neural Machine Translation) model, an RNN (Recurrent Neural Network) model, or any other model.

[0042] In the embodiments of this application, an example is described in which the first translation submodel is an NMT model. The NMT model has an encoder-decoder framework. After the first text is input to the first translation submodel, the encoder in the first translation submodel encodes the first text and obtains encoded features. Then, the obtained encoded features are input to the decoding layer in the decoder and decoded to obtain first text features. The prediction layer in the encoder determines a first probability corresponding to each of a plurality of candidate texts according to the first text features. In some embodiments, the NMT model may be a model based on a Transformer structure.

[0043] Step 202 retrieves at least one target data pair that matches the first text feature. Each target data pair contains one second text feature and one standard translation text of the second text.

[0044] Here, the second text feature is the text feature of the second text, the second text is the text of the first language, and the standard translation text is the text of the second language.

[0045] In some embodiments, the server retrieves at least one target data pair from a data pair library that matches a first text feature based on the first text feature. This data pair library contains at least one data pair, and any one of these data pairs contains one second text feature and one standard translation text corresponding to the second text feature, where the second text feature is a feature obtained by performing feature extraction on the second text, and the standard translation text corresponding to the second text feature is the exact translation text of this second text.

[0046] A target data pair is a data pair that matches a first text feature in the data pair library. The number of target data pairs to be acquired may be determined based on experience or flexibly adjusted depending on the application scene, but is not limited to the embodiments of this application. For example, the number of target data pairs may be four, eight, or other.

[0047] In some embodiments, the implementation process for obtaining at least one target data pair matching a first text feature from a data pair library includes determining the match degree of each data pair in the data pair library, and designating data pairs whose match degree satisfies the match condition as at least one target data pair matching the first text feature. The match degree of any one data pair is used to indicate the similarity between the second text feature and the first text feature in that data pair. Exemplaryly, the match degree of any one data pair may show a positive correlation with the similarity between the second text feature and the first text feature in that data pair, in which case a higher similarity corresponds to a higher match degree. Alternatively, it may show a negative correlation with the similarity between the second text feature and the first text feature in that data pair, in which case a lower similarity corresponds to a higher match degree.

[0048] In some embodiments, the match score of any one data pair may show a negative correlation with the similarity between the second text feature and the first text feature in any one data pair. For example, the distance between the second text feature and the first text feature in any one data pair may be the match score of any one data pair. In embodiments of this application, the method for calculating the distance between two text features is not limited, and may include calculating the L2 distance (also called the Euclidean distance) between two text features, calculating the cosine distance between two text features, or calculating the L1 distance (also called the Manhattan distance) between two text features.

[0049] In some embodiments, the match score of any one data pair can show a positive correlation with the similarity between the second text feature and the first text feature in any one data pair. For example, the similarity between the second text feature and the first text feature in any one data pair can be the match score of any one data pair. Similarity is used to represent the similarity between the second text feature and the first text feature. The embodiments of this application are not limited to calculating the similarity between two text features, such as calculating the cosine similarity between two text features or calculating the Pearson similarity between two text features.

[0050] In some embodiments, a data pair whose match score satisfies the matching condition refers to a data pair where the second text feature and the first text feature have a high degree of similarity. The criteria for a match score to satisfy the matching condition can be flexibly adjusted depending on how the match score is calculated. The following two cases can be considered.

[0051] In Case 1, if the match degree of any one data pair refers to the distance between the second text feature and the first text feature in any one data pair, then the data pairs whose match degree satisfies the match condition may refer to data pairs whose match degree is less than the distance threshold, or it may refer to the K (K is an integer greater than or equal to 1) data pairs with the smallest match degrees among all the match degrees, where K is the number of target data pairs to be retrieved. The distance threshold is set based on experience or can be flexibly adjusted according to the application scene.

[0052] In Case 2, if the match degree of any one data pair refers to the similarity between the second text feature and the first text feature in any one data pair, then a data pair whose match degree satisfies the matching condition may refer to a data pair whose match degree is greater than the similarity threshold, or it may refer to the K (where K is an integer greater than or equal to 1) data pairs with the highest match degrees among all the match degrees. The similarity threshold is set based on experience or can be flexibly adjusted depending on the application scenario.

[0053] Before obtaining at least one target data pair from the data pair library that matches the first text feature, the data pair library must be built first. Exemplaryly, the process of building a data pair library includes the steps of obtaining multiple second texts and extracting features from each of the multiple second texts to obtain multiple second text features. Since each second text corresponds to one second text feature, all the second text features corresponding to each second text and the corresponding standard translation texts are made into a single data pair, so that each data pair contains one second text feature and one standard translation text.

[0054] Here, the second text may be extracted from a sample text containing this second text, and the sample text is a text in the first language. The standard translation text corresponding to the second text may be extracted from the standard translation text corresponding to the sample text, and the sample text is a text that has a standard translation text. The standard translation text corresponding to the sample text may be obtained by a professional translator translating the sample text, and this standard translation text is a text in the second language. The sample text and the standard translation text corresponding to the sample text express the same meaning using different languages. Exemplaryly, one sample text and the standard translation text corresponding to this one sample text can constitute one sample example, and multiple sample examples can constitute a sample set. In some embodiments, a sample example may also be called a training example, and a sample set may also be called a training set.

[0055] In some embodiments, the process of extracting second text features from a second text can be achieved by calling a text feature extraction model. In embodiments of this application, the type of text feature extraction model is not limited; for example, a text feature extraction model can refer to a sub-model for extracting text features within an NMT model. The method for extracting second text features is the same in principle as the method for extracting first text features in step 201, and therefore will not be described further here.

[0056] In the process of constructing a data pair library, a text feature extraction model (for example, a submodel for extracting text features within an NMT model) is used to extract features from the second text in all sample examples within the sample set, obtaining multiple second text features. These second text features and the corresponding standard translation texts are recorded and stored as data pairs within the data pair library. Exemplary, the second text features may also be called representations generated by a decoder corresponding to the second text, and the standard translation texts corresponding to the second text features may also be called exact translation texts corresponding to the second text features.

[0057] In some embodiments, each data pair can be represented as a key-value pair, where the second text feature in each data pair is the key and the standard translated text in each data pair is the value.

[0058] In some embodiments, given a sample set {(x,y)} (where (x,y) represents a sample example, x represents a sample text, and y represents a standard translation text corresponding to the sample text), a data pair library D may be constructed based on the following equation (1).

[0059]

number

[0060] In some embodiments, when the number of at least one target data pair is K (where K is an integer greater than or equal to 1), the k-th (where k is any one integer between 1 and K) data pair among these K target data pairs can be represented as (h k , v k ). Here, h k represents the second text feature in the k-th data pair. v k represents the standard translation text in the k-th data pair.

[0061] In some embodiments, step 202 may be achieved by calling a target text translation model. That is, the target text translation model can be called to obtain at least one target data pair that matches the first text feature. Exemplary, if the structure of the target text translation model is the structure described in step 201, calling the target text translation model to obtain at least one target data pair that matches the first text feature may mean calling a second translation submodel of the target text translation model to obtain at least one target data pair that matches the first text feature. Exemplary, since the second translation submodel includes a data pair lookup network for searching for matching data pairs from a data pair library, the process of obtaining at least one target data pair that matches the first text feature may be achieved via the data pair lookup network in the second translation submodel. Exemplary, the data pair lookup network may be a simple feedforward neural network or other more complex networks.

[0062] Step 203 determines the confidence and match of at least one target data pair. The confidence of any one target data pair is used to indicate the degree of confidence of any one target data pair, and the match of any one target data pair is used to indicate the similarity between the second text feature and the first text feature in any one target data pair.

[0063] Here, the method for determining the match of at least one target data pair was described in step 202 and will not be described further here. In the method according to the embodiment of this application, after obtaining at least one target data pair that matches the first text feature, it is also necessary to determine the confidence level of at least one target data pair. The confidence level of any one target data pair is used to indicate the degree of reliability of any one target data pair. Optionally, the confidence level of any one target data pair shows a positive correlation with the degree of reliability of any one target data pair, i.e., the higher the confidence level of any one target data pair, the higher the degree of reliability of that one target data pair. By considering the confidence level of at least one target data pair, the determined second probability can be made more reliable, thereby further improving the accuracy of the translated text corresponding to the first text.

[0064] In some embodiments, step 203 can be achieved by calling a target text translation model. That is, the target text translation model can be called to determine the confidence level of at least one target data pair. Optionally, if the structure of the target text translation model is the structure described in step 201, calling the target text translation model to determine the confidence level of at least one target data pair can mean calling a second translation submodel of the target text translation model to determine the confidence level of at least one target data pair. Optionally, the second translation submodel further includes a probability distribution prediction network in addition to the data pair lookup network in step 202. The process for determining the confidence level of at least one target data pair can be achieved via the probability distribution prediction network in the second translation submodel.

[0065] The principle for determining the confidence level of each target data pair among at least one target data pair is the same. In embodiments of this application, the process for determining the confidence level of any one of the target data pairs will be described as an example. In some embodiments, the implementation process for determining the confidence level of any one target data pair includes the steps of: determining at least one third probability, i.e., each candidate text corresponds to a third probability, based on a second text feature in any one target data pair, the at least one third probability being used to indicate the probability that a second text in any one target data pair is translated into each candidate text; in other words, the third probability corresponding to any one candidate text is used to indicate the probability that a second text corresponding to any one candidate text is translated into any one candidate text; determining a fourth probability based on at least one third probability, the fourth probability being used to indicate the probability that a second text corresponding to any one target data pair is translated into a standard translation text in any one target data pair; and determining the confidence level of any one target data pair based on the fourth probability.

[0066] The principle for determining the third probability for each candidate text based on the second text feature is the same as the principle for determining the first probability for each candidate text based on the first text feature, so it will not be explained further here. The probability that the second text corresponding to any one target data pair is translated into any one candidate text is called the third probability. Based on the third probability for each of the multiple candidate texts, the probability that the second text is translated into the standard translation text in any one target data pair is determined, and this is called the fourth probability.

[0067] In some embodiments, a process for determining the probability that a second text is translated into a standard translation text in any one of the target data pairs, based on a third probability corresponding to each of a plurality of candidate texts, i.e., a process for determining a fourth probability based on at least one third probability, is such that if the third probability corresponding to each of the plurality of candidate texts includes the corresponding third probability of the standard translation text in any one of the target data pairs, then it is shown that the standard translation text is one of the candidate texts, and in this case the standard translation text in any one of the target data pairs The process includes the steps of setting the corresponding third probability as the fourth probability, i.e., the probability that the second text is translated into a standard translation text in any one of the target data pairs, and if the corresponding third probability for each candidate text does not include the corresponding third probability for the standard translation text in any one of the target data pairs, it is shown that the standard translation text in any one of the target data pairs is not one of the candidate texts, in which case the first number is set as the fourth probability, i.e., the probability that the second text is translated into a standard translation text in any one of the target data pairs. In this process, the first number is a number less than or equal to the minimum of the corresponding third probabilities for each candidate text, for example, if the numerical range of each third probability is 0 to 1, the first number may be 0. If the standard translation text in any one target data pair is one of the candidate texts, the higher the probability that the second text will be translated into the standard translation text in any one target data pair, the higher the probability that the standard translation text can be predicted based on the second text features, i.e., the higher the confidence level of that one target data pair.

[0068] In some embodiments, the process for determining the confidence level of any one target data pair based on the probability that a second text is translated into a standard translation text in any one target data pair, i.e., a fourth probability, includes the steps of transforming the probability that a second text is translated into a standard translation text in any one target data pair, i.e., transforming the fourth probability, and taking the numerical value obtained after the transformation as the confidence level of any one target data pair. For example, taking the example of determining the confidence level of at least one target data pair using a probability distribution prediction network in a second translation submodel, the probability that a second text is translated into a standard translation text in any one target data pair, i.e., the fourth probability, is input to the probability distribution prediction network, the fourth probability is transformed through the probability distribution prediction network, and the numerical value output from the probability distribution prediction network is taken as the confidence level of any one target data pair. The process of using a probability distribution prediction network to convert the probability that a second text is translated into a standard translation text in any one of the target data pairs is an internal computation process of the probability distribution prediction network, and is therefore not limited to the embodiments of this application. It is sufficient to ensure that the output confidence score shows a positive correlation with the probability that a second text is translated into a standard translation text in any one of the target data pairs.

[0069] In some embodiments, a probability distribution prediction network is used to predict the k-th target data pair (h) (where k is any integer from 1 to K). k ,v k The process for transforming the fourth probability determined based on ) can be expressed using the following equation (2).

[0070]

number

[0071] In some embodiments, the realization process for determining the confidence level of any one target data pair based on the probability that a second text is translated into a standard translation text in any one target data pair, i.e., a fourth probability, includes the steps of determining a fifth probability based on a first probability corresponding to each of a plurality of candidate texts, i.e., at least one first probability, wherein the fifth probability is used to indicate the probability that the first text is translated into a standard translation text in any one target data pair, and determining the confidence level of any one target data pair based on the probability that the second text is translated into a standard translation text in any one target data pair and the probability that the first text is translated into a standard translation text in any one target data pair, i.e., determining the confidence level of any one target data pair based on the fourth probability and the fifth probability.

[0072] In some embodiments, the realization process for determining the probability that a first text is translated into a standard translation text in any one of the target data pairs, based on a first probability corresponding to each of a plurality of candidate texts, i.e., the process for determining a fifth probability based on at least one first probability, is such that if the first probability corresponding to each candidate text includes a first probability corresponding to a standard translation text in any one of the target data pairs, then it is shown that the standard translation text in any one of the target data pairs is one of the candidate texts, and in this case, in any one of the target data pairs The process includes the steps of: setting the first probability corresponding to the standard translation text in a given data pair to the fifth probability, i.e., the probability that the first text is translated to the standard translation text in any one of the target data pairs; and if the first probability corresponding to each of the multiple candidate texts does not include the first probability corresponding to the standard translation text in any one of the target data pairs, it is shown that the standard translation text in any one of the target data pairs is not one of the candidate texts, in which case the second numerical value is set to the probability that the first text is translated to the standard translation text in any one of the target data pairs. In this process, the second numerical value is a number less than or equal to the minimum value of the first probabilities corresponding to each of the multiple candidate texts. For example, if the numerical range of each first probability is 0 to 1, the second numerical value may be 0. When the standard translation text in any one target data pair is one of several candidate texts, the higher the probability that the first text will be translated into the standard translation text in any one target data pair, the higher the probability that the standard translation text can be predicted based on the first text features.The high similarity between the first text feature and the second text feature in any one of the target data pairs indicates to some extent that the higher the probability of predicting the standard translation text in any one of the target data pairs based on the first text feature prediction, the higher the probability of predicting the standard translation text in any one of the target data pairs based on the second text feature in any one of the target data pairs, thus indicating to some extent that the reliability of any one of the target data pairs is higher.

[0073] In some embodiments, the implementation process for determining the confidence level of any one target data pair based on the probability that the second text is translated into a standard translation text in any one target data pair and the probability that the first text is translated into a standard translation text in any one target data pair, i.e., the implementation process for determining the confidence level of a target data pair based on the fourth probability and the fifth probability, involves inputting the probability that the second text is translated into a standard translation text in any one target data pair and the probability that the first text is translated into a standard translation text in any one target data pair into an input probability distribution prediction network, using the probability distribution prediction network to convert the probability that the second text is translated into a standard translation text in any one target data pair and the probability that the first text is translated into a standard translation text in any one target data pair, and taking the numerical value output from the probability distribution prediction network as the confidence level of any one target data pair. In other words, the fourth and fifth probabilities are input into the input probability distribution prediction network, the probability distribution prediction network is used to transform the fourth and fifth probabilities, and the numerical value output from the probability distribution prediction network is used as the confidence level for one of the target data pairs. The process of transforming the fourth and fifth probabilities using the probability distribution prediction network is an internal calculation process of the probability distribution prediction network, and is therefore not limited to the embodiments of this application. It is sufficient to ensure that the output confidence level shows a positive correlation with the fourth and fifth probabilities.

[0074] In some embodiments, the confidence level of at least one target data pair may be determined according to equation (3) below.

[0075]

number

number

number

number

[0076] In some embodiments, a method for determining the confidence level of any one target data pair may further include the steps of: determining the probability that a first text is translated into a standard translation text in any one target data pair, based on a first probability corresponding to each of a plurality of candidate texts; and transforming the probability that the first text is translated into a standard translation text in any one target data pair, with the resulting numerical value being the confidence level of any one target data pair. That is, the method may include the steps of: determining a fifth probability based on at least one first probability, the fifth probability being used to indicate the probability that the first text is translated into a standard translation text in any one target data pair; and transforming the fifth probability, with the resulting numerical value being the confidence level of any one target data pair.

[0077] In some embodiments, taking as an example the determination of the confidence level of at least one target data pair using a probability distribution prediction network in a second translation submodel, the probability that the first text is translated into a standard translation text in any one of the target data pairs is input to the probability distribution prediction network, the probability distribution prediction network is used to transform the probability that the first text is translated into a standard translation text in any one of the target data pairs, and the numerical value output from the probability distribution prediction network is taken as the confidence level of any one of the target data pairs. In other words, a fifth probability is input to the probability distribution prediction network, the probability distribution prediction network is used to transform the fifth probability, and the numerical value output from the probability distribution prediction network is taken as the confidence level of any one of the target data pairs. The process of using the probability distribution prediction network to transform the probability that the first text is translated into a standard translation text in any one of the target data pairs is an internal calculation process of the probability distribution prediction network, and is not limited to the embodiments of this application. It is sufficient to ensure that the output confidence level shows a positive correlation with the probability that the first text is translated into a standard translation text in any one of the target data pairs.

[0078] For example, using a probability distribution prediction network, we can predict the k-th target data (h) (where k is any integer from 1 to K). k ,v k The process for transforming the fifth probability determined based on ) can be expressed using the following equation (4).

[0079]

number

number

number

[0080] In step 204, at least one second probability is determined based on the confidence and match of at least one target data pair.

[0081] Here, at least one second probability is used to indicate the probability that the first text is translated into each standard translation text in at least one target data pair. The second probability corresponding to any one standard translation text is used to indicate the probability that the first text is translated into any one of the standard translation texts. Each standard translation text must be a unique translation text. For example, if two of the ten target data pairs found have the same standard translation text, then there are nine standard translation texts. To calculate the second probability corresponding to each standard translation text, simply add up the probabilities for each of the same standard translation texts.

[0082] In some embodiments, step 204 may be achieved by calling a target text translation model. That is, the target text translation model can determine a second probability corresponding to each of the standard translated texts in at least one target data pair, based on the confidence and match of at least one target data pair. Exemplary, assuming the structure of the target text translation model is the same as described in step 201, calling the target text translation model and determining the confidence of at least one target data pair may mean calling a second translation submodel of the target text translation model and determining the confidence of at least one target data pair. Exemplary, the second translation submodel further includes a probability distribution prediction network in addition to the data pair lookup network associated with step 202. The process for determining a second probability corresponding to each of the standard translated texts in at least one target data pair can be achieved via the probability distribution prediction network in the second translation submodel. For example, since the process of determining the second probability considers not only the degree of match but also the confidence level, the probability distribution prediction network can be considered a distribution calibration (DC) network compared to a network that determines the second probability by considering only the degree of match.

[0083] In some embodiments, determining a second probability corresponding to each of the standard translation texts in at least one target data pair based on the confidence and match of at least one target data pair includes the steps of: standardizing the match of a first data pair for any one of a plurality of standard translation texts to obtain a standardized match, wherein the first data pair is a data pair containing any one of the standard translation texts in at least one target data pair; correcting the standardized match using the confidence of the first data pair to obtain a corrected match; and determining a second probability corresponding to any one of the standard translation texts, i.e., determining a second probability corresponding to any one of the standard translation texts based on the corrected match, wherein the corrected match and the second probability show a positive correlation.

[0084] A first data pair is a data pair that contains a standard translation text from at least one of the target data pairs. There may be one or more first data pairs. Each first data pair has a match degree and a confidence degree. Standardizing the match degrees of the first data pairs to obtain standardized match degrees means standardizing the match degree of each first data pair so that each first data pair obtains its corresponding standardized match degree. Correcting the standardized match degrees using the confidence degrees of the first data pairs to obtain corrected match degrees means correcting the standardized match degree of each first data using the confidence degrees of each first data pair so that each first data pair obtains its corresponding corrected match degree.

[0085] After obtaining the match degree of a first data pair, standardizing that match degree can yield a standardized match degree, thereby improving the predictability of the match degree of the first data pair. Taking one first data pair as an example, in some embodiments, the method for standardizing the match degree of the first data pair may be to standardize the match degree of the first data pair using hyperparameters. Before standardizing the match degree of the first data pair using hyperparameters, it is necessary to determine the size of the hyperparameters. Here, the numerical values ​​of the hyperparameters may be set empirically or flexibly adjusted according to the target data, but are not limited to the embodiments of this application.

[0086] In embodiments of this application, the hyperparameters are described as being dynamically determined according to the target data pairs. The process for determining the hyperparameters includes the step of determining the hyperparameters based on at least one of the following pieces of information: a quantitative index for each target data pair and the degree of match for each target data pair. Here, the quantitative index for any one of the target data pairs is the number of target data pairs that are not ranked lower than any one of the target data pairs after each target data pair has been sorted according to reference order.

[0087] In embodiments of this application, the determination of hyperparameters relates to one of two sets of data, which include a quantitative index for each target data pair and the degree of match for each target data pair. Here, the quantitative index for one of the target data pairs is the number of target data pairs that are not ranked lower than any one of the target data pairs after each target data pair has been sorted according to reference order. The reference order may be determined empirically or can be flexibly adjusted according to the application scene. For example, different target data pairs may be assigned different numbers, and the reference order may refer to the order in descending order of number or the order in ascending order of number. After each target data pair has been sorted according to reference order, each target data pair has its own position, and the quantitative index for that one of the target data pairs is the number of non-overlapping standard translated texts among each target data pair that are not ranked lower than any one of the target data pairs.

[0088] For example, suppose there are three target data pairs searched, data pair 1, data pair 2, and data pair 3, and the standard translation text for data pair 1 and data pair 2 is M1, and the standard translation text for data pair 3 is M2. Assuming that after sorting according to the reference order, data pair 1, data pair 2, and data pair 3 are positioned from top to bottom, then the quantitative indicator for data pair 1 is 1, the quantitative indicator for data pair 2 is 2, and the quantitative indicator for data pair 3 is 2.

[0089] Hyperparameters may be determined based solely on the quantitative indicators of each target data pair, solely on the degree of match of each target data pair, or based on both the quantitative indicators and the degree of match of each target data pair. The numerical values ​​of the hyperparameters can be obtained by inputting at least one of the quantitative indicators and the degree of match of each target data pair into a probability distribution prediction network and performing calculations.

[0090] Taking the example of determining hyperparameters based on the quantitative indicators of each target data pair and the degree of match between each target data pair, the hyperparameters can be calculated according to the following equation (5).

[0091]

number

[0092] In some embodiments, the method for standardizing the match degree of a first data pair using hyperparameters may be the ratio of the match degree of the first data pair to the hyperparameter, or the product of the match degree of the first data pair to the hyperparameter, but is not limited to the embodiments of this application.

[0093] After obtaining the standardized match score corresponding to the first data pair, the standardized match score is corrected using the confidence score of the first data pair to obtain the corrected match score. This corrected match score is a match score that matches the degree of confidence of the first data pair. Exemplaryly, the method of correcting the standardized match score using the confidence score of the first data pair may depend on the specific circumstances of the match score of the first data pair. For example, if the match score of the first data pair shows a positive correlation with the similarity between the second text feature and the first text feature in the first data pair, the corrected match score can be the sum of the confidence score and the standardized match score of the first data pair. On the other hand, if the match score of the first data pair shows a negative correlation with the similarity between the second text feature and the first text feature in the first data pair, the corrected match score can be the difference between the confidence score and the standardized match score of the first data pair.

[0094] In some embodiments, if there is one first data pair, the second probability corresponding to any one of the standard translation texts is the probability that the corrected match score determined according to that one first data pair shows a positive correlation. On the other hand, if there are multiple first data pairs, the sum of the corrected match scores determined according to the multiple first data pairs is calculated, and the second probability corresponding to any one of the standard translation texts is the probability that the calculated sum of match scores shows a positive correlation.

[0095] For example, use one of the standard translation texts v k Let's take this as an example. The second probability corresponding to any one of the standard translation texts can be calculated according to equation (6) below.

[0096]

number

number

number

number

number

number

[0097] By referring to the method for obtaining the second probability corresponding to any one of the standard translation texts, we can determine the corresponding second probability for each standard translation text.

[0098] In the embodiments of this application, the order in which the first probability corresponding to each of the multiple candidate texts and the second probability corresponding to each of the multiple standard translation texts are determined is not limited and can be flexibly set according to actual needs. After determining the first probability corresponding to each candidate text and the second probability corresponding to each standard translation text, step 205 is performed.

[0099] In step 205, the translated text corresponding to the first text is determined based on at least one first probability and at least one second probability.

[0100] Here, at least one first probability is a first probability corresponding to each of the multiple candidate texts, and at least one second probability is a second probability corresponding to each of the multiple standard translation texts. Accordingly, step 205 may be expressed as determining the translation text corresponding to the first text based on the first probability corresponding to each of the multiple candidate texts and the second probability corresponding to each of the standard translation texts.

[0101] The translated text corresponding to the first text refers to a translation of the first text into the second language that is appropriate to that text. The translated text corresponding to the first text is determined by comprehensively considering the first probability corresponding to each of the multiple candidate texts and the second probability corresponding to each of the multiple standard translated texts. The process of determining the translated text corresponding to the first text involves a wealth of information, which contributes to ensuring the reliability of the translated text corresponding to the first text. Furthermore, since the second probability corresponding to each standard translated text is determined by comprehensively considering the degree of match and reliability of the target data pair, the information taken into consideration is also rich, and the determined second probability matches the degree of reliability of the target data pair, increasing the reliability of the second probability and thereby contributing to further enhancing the reliability of the translated text corresponding to the first text.

[0102] In some embodiments, the process for determining a translation text corresponding to a first text, based on a first probability for each candidate text and a second probability for each standard translation text, includes the steps of: determining a first approximate probability distribution based on the first probability for each candidate text; determining a second probability distribution based on the second probability for each standard translation text; fusing the first and second probability distributions to obtain a fused probability distribution, wherein the fused probability distribution includes the translation probability for each target text, and each target text includes each candidate text and each standard translation text; and selecting the target text with the highest translation probability among the target texts as the translation text. In other words, the steps include determining a first probability distribution based on at least one first probability, determining a second probability distribution based on at least one second probability, fusing the first and second probability distributions to obtain a fused probability distribution, wherein the fused probability distribution includes the translation probability of each target text, and each target text includes each candidate text and each standard translation text, and selecting the target text with the highest translation probability among the target texts as the translation text.

[0103] Here, the first probability distribution includes a first probability for each candidate text, and the second probability distribution includes a second probability for each standard translation text. In the embodiments of this application, the method of fusing the obtained first and second probability distributions is not limited; it is sufficient to obtain a fusing probability distribution in which each target text includes a corresponding translation probability. Here, each target text includes each candidate text and each standard translation text. In other words, each target text is a non-overlapping text from each candidate text and each standard translation text. Exemplarily, a translation text corresponding to the first text can be obtained by fusing the first and second probability distributions using a weight interpolation method.

[0104] In some embodiments, merging a first probability distribution and a second probability distribution to obtain a merged probability distribution includes the steps of determining a first importance and a second importance, wherein the first importance is used to indicate the importance of the first probability distribution in obtaining the translated text, and the second importance is used to indicate the importance of the second probability distribution in obtaining the translated text; determining a target parameter based on the first importance and the second importance, wherein the target parameter is also called a normalization parameter; transforming the first importance based on the target parameter to obtain a first weight; transforming the second importance based on the target parameter to obtain a second weight; and merging the first probability distribution and the second probability distribution based on the first weight and the second weight to obtain a merged probability distribution.

[0105] In some embodiments, step 205 may be achieved by calling a target text translation model. That is, the target text translation model can determine the translation text corresponding to the first text based on a first probability that each candidate text corresponds to and a second probability that each standard translation text corresponds to. In other words, the target text translation model determines the translation text corresponding to the first text based on at least one first probability and at least one second probability. If the structure of the target text translation model is as described in step 201, then step 205 may be achieved by calling a third translation submodel of the target text translation model. Exemplarily, the third translation submodel may include a weight prediction network (WP) for predicting first and second weights, and a fusion network for fusing the first and second probability distributions according to the first and second weights.

[0106] In some embodiments, the first importance is calculated by the weighted prediction network from at least one of the following pieces of information: the probability that each standard translation text can be predicted based on a first text feature, the probability that the standard translation text in each target data pair can be predicted based on a second text feature in each target data pair, and the first probability to which each candidate text corresponds. In other words, the first importance is calculated by the weighted prediction network from at least one of the following pieces of information: at least one fifth probability, at least one fourth probability, and at least one first probability.

[0107] In some embodiments, the first importance can be calculated according to equation (7) below, for example, by taking the probability that the weight prediction network can predict each standard translation text based on a first text feature, the probability that the standard translation text in each target data pair can be predicted based on a second text feature in each target data pair, and the corresponding first probability for each candidate text.

[0108]

number

number

number

[0109] In some embodiments, the second importance is determined by a weighted prediction network according to at least one piece of information: the quantitative index of each target data pair and the degree of match between each target data pair. Optionally, taking the case where the second importance is determined by a weighted prediction network according to the quantitative index of each target data pair and the degree of match between each target data pair, the second importance can be calculated according to equation (8) below.

[0110]

number

[0111] After calculating the first and second importance values, a normalization parameter is determined based on the first and second importance values. This normalization parameter is the parameter used as the basis for converting the first and second importance values, and the sum of the first and second weights obtained by converting the first and second importance values ​​according to the normalization parameter is 1. Optionally, the second weight can be calculated according to equation (9) below.

[0112]

number

[0113] In some embodiments, the sum of the first importance and the second importance may be used as the target parameter, the ratio of the first importance to the target parameter may be used as the first weight, and the ratio of the second importance to the target parameter may be used as the second weight.

[0114] In some embodiments, the fusion probability distribution can be calculated based on the first weight and the second weight according to the following equation (10).

[0115]

number

[0116] In some embodiments, the fusion probability distribution obtained according to equation (10) above includes the corresponding translation probabilities for multiple target texts. Among the multiple target texts, the text with the highest translation probability is determined, and that text is set as the translation text corresponding to the first text.

[0117] Figure 3 is a schematic diagram of a confidence-based text translation model. Using an NMT translation model as an example, Figure 3 shows the translation process from inputting a first text to outputting a translated text corresponding to the first text, including steps 201 to 205 described above. In Figure 3, 301 is the first text to be translated input into the model, and this first text is Chinese text; 302 is the NMT translation model; 303 is the first text feature; 304 is the first probability distribution; 305 is the data pair library; 306 is at least one target data pair retrieved according to the first text feature; 307 is the second probability distribution; 308 is the fused probability distribution obtained by fusing the first and second probability distributions; and 309 is the translated text corresponding to the outputted first text.

[0118] According to the technical method of the embodiment of this application, in the process of determining the second probability, in addition to the degree of match between the second text feature and the first text feature in the target data pair, the reliability of the target data pair is also taken into consideration, thus providing a wealth of information. Furthermore, since the reliability of the target data pair is used to evaluate the degree of reliability of the target data pair, considering the reliability of the target data pair can increase the reliability of the second probability and further improve the accuracy of text translation.

[0119] Embodiments of this application provide a method for obtaining a text translation model applicable to the implementation environment shown in Figure 1 above. This method for obtaining a text translation model is performed by a computer device. This computer device may be a terminal 11 or a server 12, but is not limited to the embodiments of this application. As shown in Figure 4, the method for obtaining a text translation model according to the embodiment of this application includes the following steps 401 to 407.

[0120] Step 401 involves obtaining the first sample text, the first standard translation text, and the initial text translation model in the first language.

[0121] Here, the first sample text is the text in the first language, and the first standard translation text is the text obtained by translating the first sample text into the second language.

[0122] In some embodiments, where translation from Chinese to English is required, the first language is Chinese and the second language is English. The first sample text is a text with a standard translation. In embodiments of this application, the first standard translation text is the standard translation text of the first sample text. To facilitate the provision of training data to the initial text translation model training process using the standard translation text corresponding to the first sample text, the language of the standard translation text corresponding to the first sample text is the same as the language of the translated text that needs to be output using the initial text translation model. Since the standard translation text corresponds to the first sample text, the process for training the initial text translation model using the first sample text is a supervised training process.

[0123] Furthermore, the first sample text is the text used as the basis for training the text translation model once. The number of first sample texts may be one or more, but is not limited to the embodiments of this application. In the embodiments of this application, we will explain using the example of having one first sample text. The method for obtaining the first sample text can be seen by referring to the relevant process in step 201 in the embodiment shown in Figure 2, and will not be explained further here.

[0124] In step 402, the initial text translation model processes the first sample text features to obtain at least one first sample probability.

[0125] Here, the first sample text feature is the text feature of the first sample text. At least one first sample probability is used to indicate the probability that the first sample text is translated into each of the at least one candidate texts. The first sample text corresponding to any one of the candidate texts is used to indicate the probability that the first sample text is translated into any one of the candidate texts. At least one of the candidate texts is a text in the second language.

[0126] The process for realizing step 402 can be described by referring to step 201 in the embodiment shown in Figure 2, so it will not be explained further here.

[0127] In step 403, obtain at least one sample data pair that matches the first sample text feature. Any one of the sample data pairs will contain the second sample text feature and the second standard translation text.

[0128] Here, the second sample text feature is the text feature of the second sample text. The second sample text is the text in the first language. The second standard translation text is the text obtained by translating the second sample text into the second language.

[0129] In some embodiments, obtaining at least one sample data pair that matches a first sample text feature includes the steps of searching for at least one initial data pair that matches a first sample text feature in a data pair library, wherein any one initial data pair includes one third sample text feature and one second standard translation text, where the third sample text feature is a text feature of the second sample text, the second sample text is text in a first language, and the second standard translation text is text obtained by translating the second sample text into a second language; and determining at least one sample data pair based on at least one initial data pair.

[0130] In some embodiments, a method for determining at least one sample data pair based on at least one initial data pair is to make at least one sample data pair the at least one initial data pair. In such cases, the third sample text feature of the second sample text in the initial data pair is directly used as the second sample text feature of the second sample text in the sample data pair.

[0131] In some embodiments, a method for determining at least one sample data pair based on at least one initial data pair includes the steps of interfering with at least one initial data pair according to interference probabilities to obtain an interfered data pair, and determining at least one sample data pair based on the interfered data pair.

[0132] Since the data pair library and the first sample text may not be a perfect match, and at least one of the searched sample data pairs may not contain the first standard translation text, the model can be made more robust during the training phase by adding perturbation to at least one initial data pair (i.e., interfering with at least one initial data pair), thereby improving the accuracy of the model's translation results.

[0133] Exemplary examples, the interference probability may be set empirically. Exemplary examples, the interference probability may be determined according to the number of updates corresponding to the initial text translation model. Exemplary examples, the interference probability shows a negative correlation with the number of updates corresponding to the initial text translation model. For example, the ratio of the number of updates corresponding to the initial text translation model to the rate of decrease of the interference probability is determined, a numerical value that shows a negative correlation with that ratio is determined, and the product of that numerical value and the initial interference probability is taken as the interference probability. The initial interference probability and the rate of decrease of the interference probability may be set empirically or flexibly adjusted according to the application scene, but are not limited to the embodiments of this application.

[0134] For example, the interference probability can be calculated according to the following formula (11).

[0135] α=α0*exp(-step / β) ···(11) Here, α0 is the initial interference probability. β is the rate at which the interference probability decreases. step is the number of updates corresponding to the initial text translation model. α is the interference probability. According to equation (11) above, it was found that the larger the number of updates corresponding to the initial text translation model, the smaller the interference probability α becomes.

[0136] In some embodiments, interfering with at least one initial data pair according to the interference probability means interfering with at least one initial data pair with a probability of interference probability and not interfering with at least one initial data pair with a probability of (1 - interference probability).

[0137] In some embodiments, where the interference probability includes a first interference probability, interfering at least one initial data pair according to the interference probability to obtain an interfered data pair includes adding noise features to a third sample text feature in each initial data pair according to the first interference probability to obtain an interfered data pair. In such cases, a method for determining at least one sample data pair based on the interfered data pair is to make the interfered data pair at least one sample data pair. The first interference probability is the probability of performing the interference scheme of adding noise features to a third sample text feature in each data pair.

[0138] To address the problem that the data pair library and the first sample text may not be a perfect match, a noise feature can be added to the third sample text feature of at least one retrieved initial data pair to construct a noisy data pair. A second sample text feature in the noisy data pair can be constructed according to equation (12) below.

[0139]

number

[0140] If the data pair library and the first sample text do not perfectly match, the retrieved initial data pair cannot efficiently assist in completing the model training. Therefore, by adding noise features to the third sample text features of at least one initial data pair and shifting the second sample text features from the third sample text features of the initial data pair, the data pair library and the first sample text can be better matched. Note that the second standard translation text in each initial data pair remains unchanged during this process.

[0141] Figure 5 is a schematic diagram of constructing a noisy data pair. In Figure 5, 501 is at least one initial data pair retrieved from the data pair library, 502 is the added noise feature, and 503 is the sample data pair constructed by adding the noise feature.

[0142] In some embodiments, if the interference probability includes a second interference probability, then interfering at least one initial data pair according to the interference probability to obtain an interfered data pair includes removing at least one initial data pair that does not satisfy the matching condition according to the second interference probability to obtain an interfered data pair. In such cases, a method for determining at least one sample data pair based on the interfered data pair includes the steps of constructing a reference data pair based on a first sample text feature and a first standard translation text, wherein the number of reference data pairs is the same as the number of removed initial data pairs, and determining at least one sample data pair based on the interfered data pair and the reference data pair. The second interference probability is the probability of performing the interference scheme to remove at least one initial data pair that does not satisfy the matching condition. Exemplaryly, the second interference probability may be the same as the first interference probability or may be different from the first interference probability.

[0143] In some embodiments, if the second standard translation text does not include the first standard translation text, a reference data pair can be constructed based on the first sample text feature and the first standard translation text to ensure that at least one sample data pair includes the first standard translation text. Optionally, constructing a reference data pair based on the first sample text feature and the first standard translation text may mean directly constructing the reference data pair based on the first sample text feature and the first standard translation text, or it may mean adding noise features to the first sample text feature and constructing the reference data pair based on the sample text feature obtained by adding the noise features and the first standard translation text.

[0144] In some embodiments, determining at least one sample data pair based on the interfered data pair and the reference data pair means using both the interfered data pair and the reference data pair as the sample data pair. In the process of using the interfered data pair as the sample data pair, a third sample text feature in the interfered data pair is used as the second sample text feature in the sample data pair, and a second standard translation text in the interfered data pair is used as the second standard translation text in the sample data pair. In the process of using the reference data pair as the sample data pair, a first sample text feature in the sample data pair, or a sample text feature obtained by adding a noise feature to the first sample text feature, is used as the second sample text feature in the sample data pair, and a first standard translation text in the reference data pair is used as the second standard translation text in the sample data pair.

[0145] Figure 6 is a schematic diagram for obtaining sample data pairs. In Figure 6, 601 is at least one initial data pair retrieved within the data pair library, 602 is a reference data pair constructed based on a first sample text feature and a first standard translation text, and 603 is at least one determined sample data pair.

[0146] In some embodiments, removing an initial data pair that does not satisfy the match condition according to a second interference probability may mean removing the initial data pair from which the distance between the third sample text feature and the first sample text feature is furthest. As shown in Figure 6, removing the initial data pair from which the distance between the third sample text feature and the first sample text feature is furthest is illustrated. Optionally, not satisfying the match condition may mean that the distance between the third sample text feature and the first sample text feature is greater than a distance threshold. This distance threshold may be set empirically or flexibly adjusted according to the actual situation, but is not limited to the embodiments of this application.

[0147] In some embodiments, if the interference probability includes a first interference probability and a second interference probability, interfering at least one initial data pair according to the interference probability to obtain an interfered data pair includes the steps of: adding noise features to a third sample text feature in each initial data pair according to the first interference probability to obtain an intermediate data pair; and removing initial data pairs from the intermediate data pair that do not satisfy the matching condition according to the second interference probability to obtain an interfered data pair. In such cases, a method for determining at least one sample data pair based on the interfered data pair includes the steps of: constructing a reference data pair based on a first sample text feature and a first standard translation text, wherein the number of reference data pairs is the same as the number of initial data pairs removed; and determining at least one sample data pair based on the interfered data pair and the reference data pair.

[0148] In some embodiments, if the interference probability includes a first interference probability and a second interference probability, interfering at least one initial data pair according to the interference probability to obtain an interfered data pair includes the steps of: removing at least one initial data pair that does not satisfy the matching condition according to the second interference probability to obtain an intermediate data pair; and adding noise features to a third sample text feature in the intermediate data pair according to the first interference probability to obtain an interfered data pair. In such cases, a method for determining at least one sample data pair based on the interfered data pair includes the steps of: constructing a reference data pair based on a first sample text feature and a first standard translation text, wherein the number of reference data pairs is the same as the number of removed initial data pairs; and determining at least one sample data pair based on the interfered data pair and the reference data pair.

[0149] In some embodiments, unlike the process of obtaining the translated text of the first text using the text translation model shown in Figure 3, the process of training the initial text translation model involves adding a certain level of interference to the initial data pairs retrieved within the data pair library, and then determining the confidence level and translated text on the data pairs with added interference. This significantly improves the robustness of the model and makes it resistant to noise interference.

[0150] Step 404 determines the confidence and match scores for at least one sample data pair. The confidence score of any one sample data pair is used to indicate the degree of reliability of any one sample data pair. The match score of any one sample data pair is used to indicate the similarity between the second sample text feature and the first sample text feature in any one sample data pair.

[0151] The process for realizing step 404 can be described by referring to step 203 in the embodiment shown in Figure 2, and therefore will not be explained further here.

[0152] In step 405, at least one second sample probability is determined based on the confidence and match of at least one sample data pair.

[0153] Here, at least one second sample probability is used to indicate the probability that the first sample text is translated into each second standard translation text in at least one sample data pair. A second sample probability corresponding to any one of the second standard translation texts is used to indicate the probability that the first sample text is translated into any one of the second standard translation texts.

[0154] The process for realizing step 405 can be described by referring to step 204 in the embodiment shown in Figure 2, and therefore will not be explained further here.

[0155] In step 406, the predicted translated text corresponding to the first sample text is determined based on at least one first sample probability and at least one second sample probability.

[0156] The process for realizing step 406 can be described by referring to step 205 in the embodiment shown in Figure 2, and therefore will not be explained further here.

[0157] In step 407, the initial text translation model is updated based on the difference between the predicted translation text and the first standard translation text to obtain the target text translation model.

[0158] In some embodiments, an outcome loss is obtained based on a predicted translation text corresponding to a first sample text and a first standard translation text, and the outcome loss is used to show the difference between the predicted translation text corresponding to the first sample text and the first standard translation text. The outcome loss is used to update the model parameters of the initial text translation model to obtain a target text translation model.

[0159] After obtaining a predicted translation text corresponding to a first sample text, the result loss is obtained based on the predicted translation text corresponding to the first sample text and the first standard translation text. In embodiments of this application, the method for obtaining the result loss based on the predicted translation text corresponding to the first sample text and the first standard translation text is not limited. Exemplaryly, the result loss may be the cross-entropy loss or mean squared error loss between the predicted translation text corresponding to the first sample text and the first standard translation text.

[0160] After obtaining the result loss, the model parameters of the initial text translation model are updated using the result loss. Updating the model parameters of the initial text translation model using the result loss may mean updating all of the model parameters of the initial text translation model using the result loss, or it may mean updating only a portion of the model parameters of the initial text translation model (for example, other model parameters excluding the model parameters of the first translation submodel), but is not limited to the embodiments of this application.

[0161] After updating the model parameters of the initial text translation model using the result loss, a trained text translation model is obtained. It is then determined whether the trained text translation model meets the training termination condition. If the trained text translation model meets the training termination condition, it is set as the target text translation model. On the other hand, if the trained text translation model does not meet the training termination condition, the trained text translation model is continuously updated by referring to steps 401 to 407 until a text translation model that meets the training termination condition is obtained and that model is set as the target text translation model.

[0162] The conditions for completing training may be determined by experience or flexibly adjusted depending on the application scene, but are not limited to the embodiments of this application. Exemplary, a trained text translation model completing training includes, but is not limited to, one of the following: the number of model parameter updates performed at the time the trained text translation model is acquired has reached a threshold; the result loss at the time the trained text translation model is acquired is less than a loss threshold; or the result loss at the time the trained text translation model is acquired has converged.

[0163] In the technical method according to the embodiments of this application, the interference probability is dynamically determined based on the number of updates corresponding to the initial text translation model, making the addition of the interference probability more rational. Furthermore, by interfering at least one initial data pair according to the interference probability and obtaining the interfered data pair, it is possible to some extent to resolve the issue that the data pair library and the first sample text do not perfectly match, and that the first standard translation text is not included in the at least one sample data pair found, thereby further improving the accuracy of the model's translation results.

[0164] In the technical method according to the embodiments of this application, in the process of determining the second sample probability, the reliability of the sample data pair is further considered in addition to the degree of match between the second sample text features and the first sample text features in the sample data pair, thus providing a wealth of information to consider. Furthermore, since the reliability of the sample data pair is used to evaluate the reliability of the sample data pair, considering the reliability of the sample data pair can improve the reliability of the second sample probability, further improve the accuracy of the preliminary translated text, improve the efficiency of model acquisition and the reliability of the acquired model, and further improve the accuracy of text translation using the model.

[0165] The text translation method according to the embodiment of this application can be considered as text translation based on the k-Nearest-Neighbor Machine Translation (kNN-MT) method. The kNN-MT method is considered an important research direction in neural machine translation tasks. Such methods assist in the generation of translations by searching for useful key-value pairs from a constructed data pair library, and there is no need to update the NMT model in this process. However, the retrieved potential noise samples can severely disrupt the performance of the model. Therefore, in order to improve the robustness of the model, the embodiment of this application proposes a reliability-based robust k-nearest-neighbor machine translation model. Specifically, while the reliability of the NMT model itself was not considered in conventional methods, the embodiment of this application introduces an NMT reliability, distribution correction network, and weight prediction network to optimize the distribution of predictions by the k-nearest-neighbor method and the weights of distribution interpolation. Furthermore, a robust training method has been added to the training process, including the addition of two types of interference to the search results, which further improves the model's ability to resist noisy search results.

[0166] The embodiments of this application, compared to conventional k-nearest neighbor machine translation models, add confidence information of an NMT model to the model structure and optimize the prediction of the k-nearest neighbor distribution and interpolated weights through two networks (a distribution correction network and a weight prediction network). By considering the confidence information of the NMT model, the model can better balance the weights of the k-nearest neighbor distribution and the distribution predicted by NMT, and can avoid the degradation of model performance caused by excessive weights of the noisy k-nearest neighbor distribution. Furthermore, since two types of interference are applied during the training process, the robustness of the model can be increased while further avoiding the effects of noise during the model training process.

[0167] Referring to Figure 7, an embodiment of the present application provides a text translation device. This device is A decision module 701 is configured to perform the step of determining at least one first probability based on a first text feature, wherein the first text feature is a text feature of a first text, the first text is a text in a first language, and at least one first probability is used to indicate the probability that the first text is translated into each of at least one candidate texts, and at least one candidate text is a text in a second language. Acquisition module 702 is configured to perform the step of acquiring at least one target data pair that matches a first text feature, wherein any one of the target data pairs includes one second text feature and one standard translation text of the second text, the second text feature is a text feature of the second text, the second text is text in the first language, and the standard translation text is text in the second language. A decision module 701 is configured to perform a step of determining the confidence and match of at least one target data pair, wherein the confidence of any one target data pair is used to indicate the degree of confidence of any one target data pair, and the match of any one target data pair is used to indicate the similarity between a second text feature and a first text feature in any one target data pair. A decision module 701 is configured to perform a step of determining at least one second probability based on the confidence and match of at least one target data pair, wherein the at least one second probability is used to indicate the probability that a first text is translated into each standard translation text in at least one target data pair, The system includes a decision module 701 configured to further perform the step of determining a translated text corresponding to a first text based on at least one first probability and at least one second probability.

[0168] In some embodiments, the decision module 701 is configured to perform the steps of: determining at least one third probability for any one of at least one target data pair based on a second text feature in any one of the target data pairs, wherein at least one third probability is used to indicate the probability that a second text corresponding to any one of the target data pairs is translated into each of the candidate texts; determining a fourth probability based on at least one third probability, wherein the fourth probability is used to indicate the probability that a second text corresponding to any one of the target data pairs is translated into a standard translation text in any one of the target data pairs; and determining the confidence level of any one of the target data pairs based on the fourth probability.

[0169] In some embodiments, the decision module 701 is configured to perform the steps of: determining a fifth probability based on at least one first probability, wherein the fifth probability is used to indicate the probability that the first text is translated into a standard translation text in any one of the target data pairs; and determining the confidence level of any one of the target data pairs based on the fourth and fifth probabilities.

[0170] In some embodiments, the decision module 701 is configured to perform the following steps: for any one of the standard translation texts, standardize the match of a first data pair to obtain a standardized match, wherein the first data pair is a data pair that includes any one of at least one target data pair of standard translation texts; correct the standardized match using the confidence of the first data pair to obtain a corrected match; and determine a second probability corresponding to any one of the standard translation texts based on the corrected match, wherein the corrected match shows a positive correlation with the second probability.

[0171] In some embodiments, the decision module 701 is configured to perform the steps of determining hyperparameters based on at least one piece of information, namely a quantitative index for each target data pair and the degree of match for each target data pair, wherein the quantitative index for any one target data pair is the number of target data pairs that are not ranked lower than any one target data pair after each target data pair has been sorted according to reference order, and the ratio of the degree of match for the first data pair to the hyperparameter is set to a standardized degree of match.

[0172] In some embodiments, the decision module 701 is configured to perform the steps of: determining a first probability distribution based on at least one first probability; determining a second probability distribution based on at least one second probability; fusing the first and second probability distributions to obtain a fused probability distribution, wherein the fused probability distribution includes the translation probability of each target text, and each target text includes each candidate text and each standard translation text; and selecting the target text with the highest translation probability among the target texts as the translation text.

[0173] In some embodiments, the decision module 701 is configured to perform the steps of: determining a first importance and a second importance, wherein the first importance is used to indicate the importance of a first probability distribution in obtaining the translation text, and the second importance is used to indicate the importance of a second probability distribution in obtaining the translation text; determining a target parameter based on the first importance and the second importance; transforming the first importance based on the target parameter to obtain a first weight; transforming the second importance based on the target parameter to obtain a second weight; and fusing the first probability distribution and the second probability distribution based on the first weight and the second weight to obtain a fusing probability distribution.

[0174] In some embodiments, the text translation method is implemented by a target text translation model for translating text in a first language to text in a second language.

[0175] According to the technical method of the embodiment of this application, in the process of determining the second probability, in addition to the degree of match between the second text feature and the first text feature in the target data pair, the reliability of the target data pair is also taken into consideration, thus providing a wealth of information. Furthermore, since the reliability of the target data pair is used to evaluate the degree of reliability of the target data pair, considering the reliability of the target data pair can increase the reliability of the second probability and further improve the accuracy of text translation.

[0176] Referring to Figure 8, an embodiment of this application provides a device for acquiring a text translation model. This device is An acquisition module 801 is configured to perform the steps of acquiring a first sample text, a first standard translation text, and an initial text translation model, wherein the first sample text is text in a first language, and the first standard translation text is text obtained by translating the first sample text into a second language. A decision module 802 is configured to perform the steps of processing a first sample text feature using an initial text translation model to obtain at least one first sample probability, wherein the first sample text feature is a text feature of the first sample text, and the at least one first sample probability is used to indicate the probability that the first sample text is translated into each of at least one candidate texts, and at least one candidate text is a text in a second language. Acquisition module 801 is configured to perform the step of acquiring at least one sample data pair that matches a first sample text feature, wherein any one of the sample data pairs includes one second sample text feature and one second standard translation text, the second sample text feature is a text feature of the second sample text, the second sample text is text in the first language, and the second standard translation text is text in which the second sample text has been translated into the second language. A decision module 802 is configured to perform a step of determining the confidence and match of at least one sample data pair, wherein the confidence of any one sample data pair is used to indicate the degree of reliability of any one sample data pair, and the match of any one sample data pair is used to indicate the similarity between a second sample text feature and a first sample text feature in any one sample data pair. A decision module 802 is configured to perform a step of determining at least one second sample probability based on the confidence and match of at least one sample data pair, wherein the at least one second sample probability is used to indicate the probability that a first sample text is translated into each second standard translation text in the at least one sample data pair, A decision module 802 is configured to further perform the step of determining a predicted translated text corresponding to a first sample text based on at least one first sample probability and at least one second sample probability, The system includes an update module 803 configured to perform the steps of updating the initial text translation model and obtaining a target text translation model based on the difference between the predicted translation text and a first standard translation text.

[0177] In some embodiments, the acquisition module 801 is A step of searching a data pair library for at least one initial data pair that matches a first sample text feature, wherein any one initial data pair includes one third sample text feature and one second standard translation text, the third sample text feature being a text feature of the second sample text, the second sample text being text in a first language, and the second standard translation text being text obtained by translating the second sample text into a second language. The steps include interfering at least one initial data pair according to the interference probability to obtain the interfered data pair, The system is configured to perform the steps of determining at least one sample data pair based on the interferometric data pair.

[0178] In some embodiments, the interference probability is determined according to the number of updates of the initial text translation model, and the interference probability shows a negative correlation with the number of updates of the initial text translation model.

[0179] In some embodiments, if the interference probability includes a first interference probability indicating the probability of performing the noise addition as an interference scheme, the acquisition module 801 is configured to perform the steps of adding noise features to a third sample text feature in each initial data pair according to the first interference probability to obtain an interfered data pair, and making the interfered data pair at least one sample data pair.

[0180] In some embodiments, if the interference probability includes a second interference probability indicating the probability of performing the removal of initial data pairs as an interference method, the acquisition module 801 is configured to perform the steps of: removing at least one initial data pair that does not satisfy the matching condition according to the second interference probability to obtain an interfered data pair; constructing a reference data pair based on a first sample text feature and a first standard translated text, wherein the number of reference data pairs is the same as the number of removed initial data pairs; and determining at least one sample data pair based on the interfered data pair and the reference data pair.

[0181] In the technical method according to the embodiments of this application, in the process of determining the second sample probability, the reliability of the sample data pair is further considered in addition to the degree of match between the second sample text features and the first sample text features in the sample data pair, thus providing a wealth of information to consider. Furthermore, since the reliability of the sample data pair is used to evaluate the reliability of the sample data pair, considering the reliability of the sample data pair can improve the reliability of the second sample probability, further improve the accuracy of the preliminary translated text, improve the efficiency of model acquisition and the reliability of the acquired model, and further improve the accuracy of text translation using the model.

[0182] In the above-described embodiment, the device was explained using only the division of each of the above-described functional modules as an example of how it realizes its functions. However, in actual application, the above-described functions can be made to be performed by different functional modules as needed, that is, the internal structure of the device can be divided into different functional modules to complete all or part of the functions described above. Furthermore, the embodiments of the device and method according to the above-described embodiment belong to the same concept, and the specific implementation process can be found in the embodiment of the method, but will not be explained further here.

[0183] In some embodiments, a computer device including a processor and memory is further provided. At least one computer program is stored in this memory. This at least one computer program is loaded and executed by one or more processors to enable the computer device to implement the text translation method or method for obtaining a text translation model described in any one of the above paragraphs. This computer device may be a server or a terminal, but is not limited to the embodiments of this application. Next, the configurations of a server and a terminal will be described, respectively.

[0184] Figure 9 is a schematic diagram of a server according to an embodiment of this application. This server may vary considerably depending on its configuration or performance, and may include one or more processors (Central Processing Units, CPUs) 901 and one or more memories 902. Here, at least one computer program is stored in the one or more memories 902. This at least one computer program is loaded and executed by the one or more processors 901 to enable the server to implement the text translation method or the method for obtaining a text translation model according to each of the method embodiments described above. Of course, this server may have components such as wired or wireless network interfaces, keyboards, and input / output interfaces for convenient input / output. This server may also include other components for implementing device functions, but these will not be described further here.

[0185] Figure 10 is a schematic diagram of a terminal according to an embodiment of this application. This terminal may be a PC, mobile phone, smartphone, PDA, wearable device, PPC, tablet, smart sensor device, smart TV, smart speaker, smart voice interaction device, smart home appliance, in-car terminal, VR device, or AR device. The terminal may also be called a user device, mobile terminal, laptop terminal, desktop terminal, etc.

[0186] Typically, the terminal includes a processor 1501 and memory 1502.

[0187] The processor 1501 may include one or more processing cores, such as a 4-core processor or an 8-core processor. The processor 1501 may be implemented using at least one hardware form from among DSP (Digital Signal Processing), FPGA (Field-Programmable Gate Array), and PLA (Programmable Logic Array). The processor 1501 may include a main processor and a coprocessor. The main processor is a processor for processing data that is active and is also called a CPU (Central Processing Unit). The coprocessor is a low-power processor for processing data that is idle. In some embodiments, the processor 1501 may integrate a GPU (Graphics Processing Unit) for rendering and painting display content required for the display. In some embodiments, the processor 1501 may further include an AI (Artificial Intelligence) processor for processing computing operations related to machine learning.

[0188] The memory 1502 may include one or more computer-readable storage media, which may be non-temporary. The memory 1502 may include high-speed random-access memory and non-volatile memory, such as one or more disk storage devices or flash memory storage devices. In some embodiments, the non-temporary computer-readable storage media in the memory 1502 are used to store at least one command executed by the processor 1501 to implement a text translation method or a method for obtaining a text translation model according to an embodiment of the method of this application.

[0189] In some embodiments, the terminal optionally includes a peripheral device interface 1503 and at least one peripheral device. The processor 1501, memory 1502, and peripheral device interface 1503 may be connected via a bus or signal lines. Each peripheral device may be connected to the peripheral device interface 1503 via a bus, signal lines, or circuit board. Specifically, the peripheral device includes at least one of a display 1505 and a power supply 1508.

[0190] The peripheral device interface 1503 may be used to connect at least one I / O (Input / Output) related peripheral device to the processor 1501 and the memory 1502. In some embodiments, the processor 1501, memory 1502, and peripheral device interface 1503 are integrated on the same chip or circuit board. In some other embodiments, one or two of the processor 1501, memory 1502, and peripheral device interface 1503 may be implemented on separate chips or circuit boards, but are not limited to this embodiment.

[0191] The display 1505 is used to display a UI (User Interface). This UI may include graphics, text, icons, videos, and any combination thereof. If the display 1505 is a touch display, the display 1505 also has the ability to collect touch signals on or over the surface of the display 1505. These touch signals can be input to the processor 1501 as control signals for processing. In this case, the display 1505 may also be used to provide virtual buttons and / or virtual keyboards, also called soft buttons and / or soft keyboards. In some embodiments, the display 1505 may be a single unit located on the front panel of the terminal. In other embodiments, the display 1505 may be at least two units separately located on different surfaces of the terminal, or it may be a folded design. In another embodiment, the display 1505 may be a flexible display located on a curved or folded surface of the terminal. Furthermore, the display 1505 may also be configured as a non-rectangular, irregular shape, i.e., a shaped screen. The display 1505 can be fabricated using materials such as LCD (Liquid Crystal Display) and OLED (Organic Light-Emitting Diode).

[0192] Power supply 1508 is used to supply power to each component within the terminal. Power supply 1508 may be AC ​​power, DC power, a disposable battery, or a rechargeable battery. If power supply 1508 includes a rechargeable battery, this rechargeable battery may support wired or wireless charging. This rechargeable battery may also be used to support fast charging technology.

[0193] Those skilled in the art will understand that the structure shown in Figure 10 does not constitute a limitation on the terminal, and that it may include more or fewer components than shown, or that some components may be combined, or that different component arrangements may be adopted.

[0194] In some embodiments, a computer-readable storage medium is further provided, which stores at least one computer program. This at least one computer program is loaded and executed by the processor of the computer device, causing the computer to implement the text translation method or method for obtaining a text translation model described in any one of the above paragraphs.

[0195] In some embodiments, the computer-readable storage medium described above may be read-only memory (ROM), random access memory (RAM), compact disc read-only memory (CD-ROM), magnetic tape, floppy disk, and optical data storage device.

[0196] In some embodiments, a computer program product including a computer program or computer command is further provided. This computer program or computer command is loaded and executed by a processor to cause a computer to implement the text translation method or method for obtaining a text translation model described in any of the above.

[0197] The foregoing are merely exemplary embodiments of this application and do not limit it. Any modifications, equivalent substitutions, or improvements, within the scope of the principles of this application, should be included within the scope of protection.

Claims

1. A text translation method applicable to computer devices, A step of determining at least one first probability based on a first text feature, wherein the first text feature is a text feature of a first text, the first text is a text in a first language, and the at least one first probability is used to indicate the probability that the first text is translated into each of at least one candidate texts, and each of the at least one candidate texts is a text in a second language. A step of obtaining at least one target data pair that matches the first text feature, wherein any one target data pair includes one second text feature and one standard translation text of the second text, the second text feature is a text feature of the second text, the second text is a text in the first language, and the standard translation text is a text in the second language. A step of determining the confidence and match of at least one target data pair, wherein the confidence of any one target data pair is used to indicate the degree of reliability of the any one target data pair, and the match of any one target data pair is used to indicate the similarity between a second text feature and a first text feature in any one target data pair. A step of determining at least one second probability based on the confidence and match of the at least one target data pair, wherein the at least one second probability is used to indicate the probability that the first text is translated into each standard translation text in the at least one target data pair, A text translation method comprising the step of determining a translated text corresponding to a first text based on the at least one first probability and the at least one second probability.

2. The step of determining the confidence level of at least one target data pair is: A step of determining at least one third probability for any one of the at least one target data pair based on a second text feature in the one target data pair, wherein the at least one third probability is used to indicate the probability that the second text corresponding to any one target data pair is translated into each of the candidate texts, A step of determining a fourth probability based on at least one third probability, wherein the fourth probability is used to indicate the probability that a second text corresponding to any one of the target data pairs is translated into a standard translation text in any one of the target data pairs, The text translation method according to claim 1, comprising the step of determining the confidence level of any one of the target data pairs based on the fourth probability.

3. The step of determining the confidence level of any one of the target data pairs based on the fourth probability is: A step of determining a fifth probability based on at least one first probability, wherein the fifth probability is used to indicate the probability that the first text is translated into a standard translation text in any one of the target data pairs, The text translation method according to claim 2, comprising the step of determining the confidence level of any one of the target data pairs based on the fourth probability and the fifth probability.

4. The step of determining at least one second probability based on the confidence and match of the at least one target data pair is: A step of standardizing the match degree of a first data pair for any one of the aforementioned standard translation texts, and obtaining the standardized match degree, wherein the first data pair is a data pair that includes any one of the aforementioned standard translation texts from the at least one target data pair, The steps include correcting the standardized match score using the confidence score of the first data pair to obtain the corrected match score, A text translation method according to claim 1, comprising the step of determining a second probability corresponding to any one of the standard translation texts based on the corrected match score, wherein the corrected match score shows a positive correlation with the second probability.

5. The step of standardizing the match degree of the first data pair and obtaining the standardized match degree is: A step of determining hyperparameters based on at least one piece of information, namely a quantitative index for each of the at least one target data pair and the degree of match for each of the target data pairs, wherein the quantitative index for any one of the target data pairs is the number of target data pairs that are not ranked lower than any one of the target data pairs after the target data pairs have been sorted according to the reference order, The text translation method according to claim 4, comprising the step of setting the ratio of the match degree of the first data pair to the hyperparameter to be the standardized match degree.

6. The step of determining a translated text corresponding to the first text based on the at least one first probability and the at least one second probability is: A step of determining a first probability distribution based on the aforementioned at least one first probability, The steps include determining a second probability distribution based on the aforementioned at least one second probability, A step of fusing the first probability distribution and the second probability distribution to obtain a fused probability distribution, wherein the fused probability distribution includes the translation probability of each target text, and each target text includes each candidate text and each standard translation text. The text translation method according to claim 1, comprising the step of selecting the target text with the highest translation probability among the aforementioned target texts as the translation text.

7. The step of fusing the first probability distribution and the second probability distribution to obtain a fused probability distribution is: A step of determining a first importance and a second importance, wherein the first importance is used to indicate the importance of the first probability distribution in obtaining the translated text, and the second importance is used to indicate the importance of the second probability distribution in obtaining the translated text, A step of determining the target parameter based on the first importance and the second importance, The steps include: transforming the first importance based on the target parameter and obtaining a first weight; The steps include: transforming the second importance based on the target parameter and obtaining a second weight; A text translation method according to claim 6, comprising the step of fusing the first probability distribution and the second probability distribution based on the first weight and the second weight to obtain a fused probability distribution.

8. The text translation method according to claim 1, wherein the text translation method is implemented by a target text translation model for translating text in a first language to text in a second language.

9. A method for obtaining a text translation model applicable to a computer device, A step of obtaining a first sample text, a first standard translation text, and an initial text translation model, wherein the first sample text is a text in a first language, and the first standard translation text is a text obtained by translating the first sample text into a second language. The initial text translation model processes a first sample text feature to obtain at least one first sample probability, wherein the first sample text feature is a text feature of the first sample text, and the at least one first sample probability is used to indicate the probability that the first sample text is translated into each of at least one candidate texts, and each of the at least one candidate texts is a text in the second language. A step of obtaining at least one sample data pair that matches the first sample text feature, wherein any one sample data pair includes one second sample text feature and one second standard translation text, the second sample text feature is a text feature of the second sample text, the second sample text is text in the first language, and the second standard translation text is text obtained by translating the second sample text into the second language, A step of determining the confidence and match of at least one sample data pair, wherein the confidence of any one sample data pair is used to indicate the degree of reliability of the any one sample data pair, and the match of any one sample data pair is used to indicate the similarity between a second sample text feature and a first sample text feature in any one sample data pair. A step of determining at least one second sample probability based on the confidence and match of the at least one sample data pair, wherein the at least one second sample probability is used to indicate the probability that the first sample text is translated into each second standard translation text in the at least one sample data pair, A step of determining a predicted translation text corresponding to the first sample text based on the at least one first sample probability and the at least one second sample probability, A method for obtaining a text translation model, comprising the steps of: updating the initial text translation model based on the difference between the predicted translation text and the first standard translation text to obtain a target text translation model.

10. The step of obtaining at least one sample data pair that matches the first sample text feature is: A step of searching a data pair library for at least one initial data pair that matches the first sample text feature, wherein any one initial data pair includes one third sample text feature and one second standard translation text, the third sample text feature is a text feature of the second sample text, the second sample text is text in the first language, and the second standard translation text is text obtained by translating the second sample text into the second language. A step of interfering with at least one initial data pair according to interference probabilities to obtain an interfered data pair, wherein the interference probabilities include at least one of a first interference probability indicating the probability of performing noise addition as the interference method and a second interference probability indicating the probability of performing deletion of an initial data pair as the interference method. A method for obtaining a text translation model according to claim 9, comprising the step of determining the at least one sample data pair based on the interfered data pair.

11. The method for obtaining a text translation model according to claim 10, wherein the interference probability is determined according to the number of updates of the initial text translation model, and the interference probability shows a negative correlation with the number of updates of the initial text translation model.

12. When the interference probability includes the first interference probability, the step of interfering with at least one initial data pair according to the interference probability to obtain the interfered data pair is: The process includes the step of adding noise features to the third sample text features in each initial data pair according to the first interference probability, thereby obtaining the interfered data pair. The step of determining the at least one sample data pair based on the interference-processed data pair is: A method for obtaining a text translation model according to claim 10, comprising the step of making the interference-processed data pair into the at least one sample data pair.

13. When the interference probability includes the second interference probability, The step of interfering with at least one initial data pair according to the interference probability to obtain an interfered data pair is: The process includes the step of removing initial data pairs that do not satisfy the matching condition from among the at least one initial data pair according to the second interference probability, thereby obtaining an interference-processed data pair. The step of determining the at least one sample data pair based on the interference-processed data pair is: A step of constructing reference data pairs based on the first sample text features and the first standard translated text, wherein the number of reference data pairs is the same as the number of deleted initial data pairs, A method for obtaining a text translation model according to claim 10, comprising the step of determining the at least one sample data pair based on the interference-processed data pair and the reference data pair.

14. When the interference probability includes the first interference probability and the second interference probability, the step of interfering with at least one initial data pair according to the interference probability and obtaining the interfered data pair is: The steps include: removing initial data pairs that do not satisfy the matching condition from among the at least one initial data pair according to the second interference probability, and obtaining an intermediate data pair; The process includes the steps of adding noise features to the third sample text features in the intermediate data pair according to the first interference probability, thereby obtaining the interfered data pair. The step of determining the at least one sample data pair based on the interference-processed data pair is: A step of constructing reference data pairs based on the first sample text features and the first standard translated text, wherein the number of reference data pairs is the same as the number of deleted initial data pairs, A method for obtaining a text translation model according to claim 10, comprising the step of determining the at least one sample data pair based on the interference-processed data pair and the reference data pair.

15. A text translation device that is placed on a computer device, A decision module configured to perform the step of determining at least one first probability based on a first text feature, wherein the first text feature is a text feature of a first text, the first text is a text in a first language, and the at least one first probability is used to indicate the probability that the first text is translated into each of at least one candidate texts, and each of the at least one candidate texts is a text in a second language, A retrieval module configured to perform the step of retrieving at least one target data pair that matches the first text feature, wherein any one target data pair includes one second text feature and one standard translation text of the second text, the second text feature is a text feature of the second text, the second text is text in the first language, and the standard translation text is text in the second language, The decision module is configured to perform a step of determining the confidence and match of at least one target data pair, wherein the confidence of any one target data pair is used to indicate the degree of reliability of any one target data pair, and the match of any one target data pair is used to indicate the similarity between a second text feature and a first text feature in any one target data pair. The decision module is configured to perform a step of determining at least one second probability based on the confidence and match of the at least one target data pair, wherein the at least one second probability is used to indicate the probability that the first text is translated into each standard translation text in the at least one target data pair. A text translation device wherein the decision module is configured to further perform the step of determining a translated text corresponding to the first text based on the at least one first probability and the at least one second probability.

16. A device for acquiring text translation models to be placed on a computer device, An acquisition module configured to perform the steps of acquiring a first sample text, a first standard translation text, and an initial text translation model, wherein the first sample text is text in a first language, and the first standard translation text is text obtained by translating the first sample text into a second language, The initial text translation model includes a decision module configured to perform the steps of processing a first sample text feature and obtaining at least one first sample probability, wherein the first sample text feature is a text feature of the first sample text, and the at least one first sample probability is used to indicate the probability that the first sample text is translated into each of at least one candidate texts, where each of the at least one candidate texts is a text in the second language, The acquisition module is configured to perform the steps of acquiring at least one sample data pair that matches the first sample text feature, wherein any one sample data pair includes one second sample text feature and one second standard translation text, the second sample text feature is a text feature of the second sample text, the second sample text is text in the first language, and the second standard translation text is text obtained by translating the second sample text into the second language. The decision module is configured to perform a step of determining the confidence and match of at least one sample data pair, wherein the confidence of any one sample data pair is used to indicate the degree of reliability of any one sample data pair, and the match of any one sample data pair is used to indicate the similarity between a second sample text feature and a first sample text feature in any one sample data pair. The decision module is configured to perform the steps of determining at least one second sample probability based on the confidence and match of the at least one sample data pair, wherein the at least one second sample probability is used to indicate the probability that the first sample text is translated into each second standard translation text in the at least one sample data pair. The decision module is configured to further perform the step of determining a predicted translated text corresponding to the first sample text based on the at least one first sample probability and the at least one second sample probability. The text translation model acquisition device further includes an update module configured to perform the step of updating the initial text translation model and obtaining a target text translation model based on the difference between the predicted translation text and the first standard translation text.

17. A computer device including a processor and memory, A computer device wherein at least one computer program is stored in the memory, and the computer device implements the text translation method according to any one of claims 1 to 8, or the method for obtaining a text translation model according to any one of claims 9 to 14, by loading and executing the at least one computer program by the processor.

18. A computer program that, when loaded and executed by a processor, causes a computer to implement the text translation method described in any one of claims 1 to 8, or the method for obtaining a text translation model described in any one of claims 9 to 14.