Text translation method, method for obtaining a text translation model, apparatus, device, and computer program
The text translation method addresses the challenge of improving translation accuracy by determining probabilities based on text features, confidence, and match levels, and updating the translation model accordingly, resulting in enhanced translation precision and model robustness.
Patent Information
- Application Number
- JP2024564506
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2022-08-30
- Filing Date
- 2023-06-19
- Publication Date
- 2025-06-05
- Estimated Expiration
- 2043-06-19
AI Technical Summary
Existing text translation methods struggle with improving the accuracy of translations, particularly in handling variations in text features and reliability of translation data pairs.
A method for text translation that determines probabilities based on text features, confidence levels, and match levels of target data pairs, and updates a text translation model using these probabilities to enhance translation accuracy.
The method significantly improves the accuracy of text translations by considering the reliability and similarity of data pairs, leading to more precise translations and a robust text translation model.
Smart Images

Figure 2025517293000001_ABST
Abstract
Description
[Technical field]
[0001] [CROSS REFERENCE TO RELATED APPLICATIONS] This application claims the benefit of priority to Chinese Patent Application No. 202211049110.8, filed on August 30, 2022, entitled "Text translation method, method, apparatus, device and medium for obtaining text translation model," the entire contents of which are incorporated herein by reference.
[0002] [Technical field] TECHNICAL FIELD The embodiments of the present application relate to the field of computer technology, and in particular to a method for translating text, a method, an apparatus, a device and a medium for obtaining a text translation model. [Background technology]
[0003] With the development of computer technology, text translation has become widely used in various scenes. Text translation allows you to translate text from one language to another. Improving the accuracy of text translation has become a technical problem that is urgently needed to be solved. Summary of the Invention [Means for solving the problem]
[0004] The embodiments of the present application provide a text translation method, a method for obtaining a text translation model, an apparatus, a device and a storage medium, which can improve the accuracy of text translation. The technical solutions are as follows:
[0005] According to one aspect, an embodiment of the present application provides a method for text translation applied to a computing device, the method comprising: determining at least one first probability based on first text features, the first text features being text features of a first text, the first text being a text in a first language, the at least one first probability being used to indicate a probability that the first text will be translated into each of at least one candidate text, each of the at least one candidate text being a text in a second language; obtaining at least one target data pair matching the first text features, each target data pair including a second text feature and a standard translation of the second text, the second text feature being a text feature of the second text, the second text being a text in the first language, and the standard translation being a text in the second language; determining a confidence level and a match level of the at least one target data pair, the confidence level of any one target data pair being used to indicate a degree of reliability of the any one target data pair, and the match level of any one target data pair being used to indicate a degree of similarity between a second text feature and the first text feature in the any one target data pair; determining at least one second probability based on the confidence and match of the at least one target data pair, the at least one second probability being used to indicate a probability that the first text will be translated into each standard translation text in the at least one target data pair; determining a translated text corresponding to the first text based on the at least one first probability and the at least one second probability.
[0006] According to another aspect, there is provided a method for obtaining a text translation model for application to a computing device, the method comprising: obtaining a first sample text, a first standard translation text, and an initial text translation model, the first sample text being a text in a first language and the first standard translation text being a text translated from the first sample text into a second language; processing first sample text features with the initial text translation model to obtain at least one first sample probability, the first sample text features being text features of the first sample text, the at least one first sample probability being used to indicate a probability that the first sample text will be translated into each of at least one candidate text, each of the at least one candidate text being a text in the second language; obtaining at least one sample data pair matching the first sample text features, each sample data pair including a second sample text feature and a second standard translated text, the second sample text features being text features of the second sample text, the second sample text being text in the first language, and the second standard translated text being text translated from the second sample text into a second language; determining a confidence level and a match level of the at least one sample data pair, the confidence level of any one sample data pair being used to indicate a degree of reliability of the one sample data pair, and the match level of any one sample data pair being used to indicate a degree of similarity between the second sample text feature and the first sample text feature in the one sample data pair; determining at least one second sample probability based on the confidence and match of the at least one sample data pair, the at least one second sample probability being used to indicate a probability that the first sample text is translated into each second standard translation text in the at least one sample data pair; determining a predicted translated text corresponding to the first sample text based on the at least one first sample probability and the at least one second sample probability; updating the initial text translation model based on a difference between the predicted translated text and the first standard translated text to obtain a target text translation model.
[0007] According to another aspect, there is provided a text translation apparatus arranged on a computing device, the apparatus comprising: a determination module configured to perform a step of determining at least one first probability based on first text features, the first text features being text features of a first text, the first text being a text in a first language, the at least one first probability being used to indicate a probability that the first text will be translated into each of at least one candidate text, each of the at least one candidate text being a text in a second language; an acquisition module configured to perform the steps of acquiring at least one target data pair matching the first text feature, where any one target data pair includes a second text feature and a standard translation of the second text, the second text feature being a text feature of the second text, the second text being a text in the first language, and the standard translation being a text in the second language; the determination module further configured to perform a step of determining a confidence and a match of the at least one target data pair, where the confidence of any one target data pair is used to indicate a degree of reliability of the any one target data pair, and the match of any one target data pair is used to indicate a similarity between a second text feature and the first text feature in the any one target data pair; the determination module configured to further perform a step of determining at least one second probability based on the confidence and match of the at least one target data pair, the at least one second probability being used to indicate a probability that the first text will be translated into each standard translation text in the at least one target data pair; and the determination module configured to further perform the step of determining a translated text corresponding to the first text based on the at least one first probability and the at least one second probability.
[0008] According to another aspect, there is provided an apparatus for obtaining a text translation model located on a computing device, the apparatus comprising: an acquisition module configured to perform the steps of acquiring a first sample text, a first standard translation text, and an initial text translation model, wherein the first sample text is text in a first language and the first standard translation text is text obtained by translating the first sample text into a second language; a decision module configured to process first sample text features with the initial text translation model to obtain at least one first sample probability, the first sample text features being text features of the first sample text, the at least one first sample probability being used to indicate a probability that the first sample text will be translated into each of at least one candidate text, each of the at least one candidate text being a text in the second language; the acquiring module further configured to perform the steps of acquiring at least one sample data pair matching the first sample text features, each sample data pair including a second sample text feature and a second standard translated text, the second sample text feature being a text feature of the second sample text, the second sample text being text in the first language, and the second standard translated text being text of the second sample text translated into a second language; the determination module further configured to perform a step of determining a confidence and a match of the at least one sample data pair, where the confidence of any one sample data pair is used to indicate a degree of reliability of the any one sample data pair, and the match of any one sample data pair is used to indicate a similarity between the second sample text feature and the first sample text feature in the any one sample data pair; the determination module further configured to perform a step of determining at least one second sample probability based on the confidence and match of the at least one sample data pair, the at least one second sample probability being used to indicate a probability that the first sample text is translated into each second standard translation text in the at least one sample data pair; the determination module configured to further perform the step of determining a predictively translated text corresponding to the first sample text based on the at least one first sample probability and the at least one second sample probability; and an update module configured to perform the steps of updating the initial text translation model based on a difference between the predicted translated text and the first standard translated text to obtain a target text translation model.
[0009] According to another aspect, there is provided a computer device including a processor and a memory, wherein the memory stores at least one computer program, and the at least one computer program is loaded and executed by the processor to cause the computer device to perform any of the above-described methods for translating text or for obtaining a text translation model.
[0010] According to another aspect, there is further provided a computer-readable storage medium having stored thereon at least one computer program, the at least one computer program being configured to cause a computer to perform any of the above-described methods for translating text or for obtaining a text translation model, when loaded and executed by a processor.
[0011] According to another aspect, there is further provided a computer program product including a computer program or computer commands that, when loaded and executed by a processor, cause the computer to perform any of the above-described methods for translating text or for obtaining a text translation model. [Brief description of the drawings]
[0012] [Figure 1] FIG. 1 is a schematic diagram of an implementation environment according to an embodiment of the present application. [Diagram 2] 1 is a flowchart of a text translation method according to an embodiment of the present application. [Diagram 3] FIG. 1 is a schematic diagram of a confidence-based text translation model according to an embodiment of the present application; [Figure 4] 1 is a flowchart of a method for obtaining a text translation model according to an embodiment of the present application. [Diagram 5] FIG. 1 is a schematic diagram for constructing noisy data pairs according to an embodiment of the present application; [Figure 6]FIG. 2 is a schematic diagram for obtaining sample data pairs according to an embodiment of the present application. [Figure 7] 1 is a schematic diagram of a text translation device according to an embodiment of the present application; [Figure 8] FIG. 1 is a schematic diagram of an apparatus for obtaining a text translation model according to an embodiment of the present application. [Figure 9] FIG. 2 is a schematic configuration diagram of a server according to an embodiment of the present application. [Figure 10] 1 is a schematic configuration diagram of a terminal according to an embodiment of the present application. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS
[0013] In order to make the objectives, technical solutions and advantages of the present application clearer, the following describes the embodiments of the present application in more detail with reference to the accompanying drawings.
[0014] In some embodiments, the text translation method and the method for obtaining a text translation model according to the embodiments of the present application may be applied to various scenarios, including but not limited to cloud technology, artificial intelligence, smart transportation, driving assistance, etc.
[0015] Artificial Intelligence (AI) refers to theories, methods, technologies, and application systems that use digital computers or machines controlled by digital computers to imitate and extend human intelligence, to sense the environment, acquire knowledge, and use that knowledge to obtain optimal results. In other words, AI is a comprehensive technology of computer science that aims to understand the nature of intelligence and create new intelligent machines that can react in a manner similar to human intelligence. In other words, AI is a technology that studies the design principles and realization methods of various intelligent machines and gives machines the functions of sensing, reasoning, and decision-making.
[0016] AI technology is considered to be one of the comprehensive disciplines, and is applied to a wide range of fields, including both hardware and software level technologies. The basic technologies of AI generally include sensors, dedicated AI chips, cloud computing, distributed storage, big data processing technology, operating / interactive systems, electromechanical integration and other technologies. AI software technology mainly includes computer vision technology, voice processing technology, natural language processing technology, and several directions such as machine learning / deep learning, autonomous driving, and smart transportation.
[0017] Natural Language Processing (NLP) is an important direction in the fields of computer science and artificial intelligence. Various theories and methods are being considered to realize effective communication between humans and computers using natural language. Natural Language Processing is a science that combines linguistics, computer science, and mathematics into one. Therefore, research in this field is closely related to linguistics research, as it relates to natural language, i.e., language that humans use on a daily basis. Natural Language Processing techniques usually include technologies such as text processing, word semantic analysis, machine translation, robot quizzes, and knowledge graphs.
[0018] Machine learning (ML) is a discipline that covers a wide range of fields, including probability theory, statistics, approximation theory, convex analysis, and algorithmic complexity theory, and is dedicated to enabling computers to imitate or realize human learning behavior, acquire new knowledge and skills, and reconstruct existing knowledge systems to continually improve their performance. Machine learning is positioned as the core of artificial intelligence and is the fundamental method of endowing computers with intelligence, and is applied across various fields of artificial intelligence. Machine learning and deep learning usually include techniques such as artificial neural networks, trust networks, reinforcement learning, transfer learning, inductive learning, and instructional learning.
[0019] With the research and progress of artificial intelligence technology, artificial intelligence technology is being researched and applied in many fields, such as general smart home, smart wearable devices, virtual assistants, smart speakers, smart marketing, unmanned driving, autonomous driving, drones, robots, smart medical care, smart customer service, connected cars, autonomous driving, smart transportation, etc. With the development of technology, artificial intelligence technology is expected to be applied in more fields and play more and more important value.
[0020] 1 shows a schematic diagram of an implementation environment according to an embodiment of the present application. The implementation environment includes a terminal 11 and a server 12.
[0021] The text translation method according to the embodiment of the present application may be executed by the terminal 11, may be executed by the server 12, or may be jointly executed by the terminal 11 and the server 12, but is not limited to the embodiment of the present application. When the text translation method according to the embodiment of the present application is jointly executed by the terminal 11 and the server 12, the server 12 is responsible for the main computing task, and the terminal 11 is responsible for the secondary computing task; alternatively, the server 12 is responsible for the secondary computing task, and the terminal 11 is responsible for the main computing task; alternatively, a distributed computing architecture may be used between the server 12 and the terminal 11 to perform collaborative computing.
[0022] The method for obtaining a text translation model according to an embodiment of the present application may be executed by the terminal 11, may be executed by the server 12, or may be jointly executed by the terminal 11 and the server 12, but is not limited to the embodiment of the present application. When the text translation method according to an embodiment of the present application is jointly executed by the terminal 11 and the server 12, the server 12 is responsible for the main computing task, and the terminal 11 is responsible for the secondary computing task; or the server 12 is responsible for the secondary computing task, and the terminal 11 is responsible for the main computing task; or the server 12 and the terminal 11 may use a distributed computing architecture to perform collaborative computing.
[0023] The device that performs the text translation method and the device that performs the text translation model acquisition method may be the same or different, but this is not limited in the embodiments of the present application.
[0024] In some embodiments, the terminal 11 is any electronic product capable of performing human-computer interaction with a user through one or more methods such as a keyboard, a touch panel, a touch screen, a remote control, a voice interactive or handwriting input device, such as a PC (Personal Computer), a mobile phone, a smartphone, a PDA (Personal Digital Assistant), a wearable device, a PPC (Pocket PC), a tablet, a smart sensor device, a smart TV, a smart speaker, a smart voice interactive device, a smart home appliance, an in-vehicle terminal, a VR (Virtual Reality) device, an AR (Augmented Reality) device, etc. The server 12 may be a single server, a server cluster consisting of multiple servers, or a cloud computing service center. The terminal 11 and the server 12 are communicatively connected via a wired or wireless network.
[0025] It should be understood by those skilled in the art that the above-mentioned terminal 11 and server 12 are merely examples, and other existing or future terminals or servers that may appear in the future shall be included within the scope of protection of the present application if they are applicable to the present application, and shall be incorporated herein by way of reference.
[0026] The method according to the embodiment of the present application can be applied in various scenarios.
[0027] For example, in an online translation scene, a server uses the method for obtaining a text translation model according to an embodiment of the present application to train an initial text translation model, and deploys the trained target text translation model in the server. A terminal logs into a translation application with a user ID. The server provides services for the translation application. The terminal uses the translation application to send a first text in a first language to be translated to the server. The server receives the first text, and uses the text translation method according to an embodiment of the present application to obtain a translation text by translating the first text into a second language according to the target text translation model, and sends the translation text to the terminal. The terminal uses the translation application to receive and display the translation text. Here, the first language and the second language are different languages. In some embodiments, the first language may also be called a source language, and the second language may also be called a target language.
[0028] Further, for example, in a face-to-face conversation scene, the server uses the method for acquiring a text translation model according to an embodiment of the present application to train an initial text translation model, and deploys the trained target text translation model in the server. The terminal logs into a translation application with a user ID. The server provides services for the translation application. The terminal uses the translation application to collect speech data belonging to a first language by any speaker, converts the speech data into a first text belonging to the first language, and uses the translation application to send the first text to be translated to the server. The server receives the first text, and uses the text translation method according to an embodiment of the present application according to the target text translation model to obtain a translation text that has the same meaning as the first text and is translated into a second language, and sends the translation text to the terminal. The terminal uses the translation application to receive the translation text, converts the translation text into speech data belonging to a second language so that the speaker corresponding to the terminal can hear the played speech data, and plays the converted speech data, thereby realizing a simultaneous interpretation effect and enabling communication between two speakers who speak different languages.
[0029] The embodiment of the present application provides a text translation method applicable to the implementation environment shown in Fig. 1 above. The text translation method is executed by a computer device. The computer device may be a terminal 11 or a server 12, but is not limited to this in the embodiment of the present application. As shown in Fig. 2, the text translation method according to the embodiment of the present application includes the following steps 201 to 205.
[0030] In step 201, at least one first probability is determined based on a first text feature.
[0031] Here, the first text feature is a text feature of the first text, and the first text is a text of the first language to be translated, but in the embodiment of the present application, the type of the first language is not limited. Exemplarily, the first language may be Chinese, English, etc. The first text may include one or more characters. The length of the characters included in the first text may be determined according to experience or actual translation requirements. For example, when the first language is Chinese, the first text may include one Chinese character, or may include multiple Chinese characters, where the multiple Chinese characters may constitute one word or one sentence. At least one first probability is used to indicate the probability that the first text is translated into each candidate text among the at least one candidate text. In other words, the first probability corresponding to any one of the candidate texts is used to indicate the probability that the first text is translated into any one of the candidate texts.
[0032] The method of acquiring the first text by the computing device may be to receive the first text uploaded by a user at the computing device, or to obtain the first text by converting speech in a first language uploaded by a user into text at the computing device, or to extract the first text from a web page at the computing device, and is not limited to this in the embodiments of the present application.
[0033] In some embodiments, the method for obtaining the first text by the computer device may further include extracting the first text from the target text by the computer device. The target text refers to a text that includes the first text. For example, if the target text is a sentence to be translated, the process for translating the sentence to be translated is realized by translating each word in the sentence one by one, so the first text is one word to be translated in the target text.
[0034] After obtaining the first text, the server needs to extract features from the first text to obtain first text features. Then, the server can determine a first probability of each candidate text in the second language based on the obtained first text features. The first text features are used to characterize the first text. In the embodiment of the present application, the aspect of the first text features is not limited, and it is sufficient if it can be easily identified and processed by a computer device. For example, the aspect of the first text features may be a vector, a matrix, etc.
[0035] In some embodiments, the process for extracting features from the first text to obtain first text features may include encoding the first text to obtain the encoded features, and decoding the encoded features to obtain the first text features.
[0036] When the server determines the first probability of each candidate text based on the first text feature, each candidate text is a text in a second language, and the second language is a corresponding language of the translation text to be obtained. The second language is different from the first language, and the type of the second language can be flexibly set according to the translation needs, but is not limited in the embodiment of the present application. For example, when a translation from Chinese to English is required, the first language is Chinese and the second language is English.
[0037] Each candidate text may be set by experience or flexibly adjusted according to the application scene. Exemplarily, each candidate text may include text extracted from a sentence in the second language and having a frequency of occurrence greater than a frequency threshold, or may include text extracted from a text library in the second language, etc.
[0038] The first probability of any one of the candidate texts means a probability that the translation text of the first text determined based on the first text feature is the one of the candidate texts. Exemplarily, the first probability to which any one of the candidate texts corresponds is a numerical value between 0 and 1. Exemplarily, the sum of the first probabilities to which each of the candidate texts corresponds may be 1. Exemplarily, the first probability to which each of the candidate texts corresponds is represented using a histogram. In the histogram, one column to which each of the candidate texts corresponds is included. The height of the column to which each of the candidate texts corresponds is used to indicate the first probability to which the one of the candidate texts corresponds.
[0039] In some embodiments, step 201 may be realized by invoking a target text translation model. That is, the target text translation model is invoked to determine a first probability corresponding to each of a plurality of candidate texts based on a first text feature. The target text translation model is a model for translating a text in a first language into a text in a second language. In the embodiments of the present application, the structure of the target text translation model is not limited as long as it can realize text translation.
[0040] In some embodiments, the target text translation model includes a first translation sub-model, a second translation sub-model, and a third translation sub-model. Among them, the first translation sub-model is used to extract features from the text to be translated and predict a first probability corresponding to each of a plurality of candidate texts based on the extracted features. The second translation sub-model is used to search for a data pair matching the features extracted by the first translation sub-model and to determine a second probability corresponding to each of a plurality of standard translation texts for each searched data pair according to the searched data pairs. The third translation sub-model is used to determine a translation text corresponding to the first text according to the first probability determined by the first translation sub-model and the second probability determined by the second translation sub-model.
[0041] When the structure of the target text translation model is the above-mentioned structure, the realization process of calling the target text translation model and determining a first probability corresponding to each of the multiple candidate texts based on the first text feature refers to calling a first translation sub-model of the target text translation model and determining a first probability corresponding to each of the multiple candidate texts based on the first text feature. In the embodiment of the present application, the type of the first translation sub-model is not limited, and it may be one having the function of feature extraction and probability determination. Exemplarily, the first translation sub-model may be an NMT (Neural Machine Translation) model, an RNN (Recurrent Neural Network) model, or other models.
[0042] In the embodiment of the present application, an example will be described in which the first translation sub-model is an NMT model. The NMT model is an encoder-decoder framework. After inputting a first text into the first translation sub-model, the encoder in the first translation sub-model encodes the first text to obtain an encoding feature. The obtained encoding feature is then input to a decoding layer in the decoder for decoding to obtain a first text feature. The prediction layer in the encoder determines a first probability corresponding to each of a plurality of candidate texts according to the first text feature. In some embodiments, the NMT model may be a model based on a Transformer structure.
[0043] In step 202, at least one target data pair that matches the first text feature is obtained, where any one target data pair includes a second text feature and a standard translation of the second text.
[0044] Here, the second text features are text features of a second text, the second text is a text in a first language, and the standard translation text is a text in a second language.
[0045] In some embodiments, the server retrieves, based on the first text feature, at least one target data pair matching the first text feature from a data pair library, the data pair library including at least one data pair, and any one data pair in the data pair library includes a second text feature and a standard translation text corresponding to the second text feature, where the second text feature is a feature obtained by performing feature extraction on the second text, and the standard translation text corresponding to the second text feature is an exact translation of the second text.
[0046] The target data pair is a data pair that matches the first text feature in the data pair library. The number of target data pairs to be obtained can be set according to experience or flexibly adjusted according to application scenes, but is not limited in the embodiment of the present application. For example, the number of target data pairs can be 4, 8, etc.
[0047] In some embodiments, the implementation process for obtaining at least one target data pair matching the first text feature from the data pair library includes: determining a match degree of each data pair in the data pair library, and taking a data pair whose match degree satisfies a match condition as at least one target data pair matching the first text feature. The match degree of any one data pair is used to indicate a similarity degree between the second text feature and the first text feature in the any one data pair. Exemplarily, the match degree of any one data pair may be positively correlated with a similarity degree between the second text feature and the first text feature in the any one data pair, in which case the higher the similarity degree, the higher the match degree. Also, the match degree of any one data pair may be negatively correlated with a similarity degree between the second text feature and the first text feature in the any one data pair, in which case the lower the similarity degree, the higher the match degree.
[0048] In some embodiments, the match degree of any one of the data pairs can show a negative correlation with the similarity degree between the second text feature and the first text feature in any one of the data pairs. For example, the distance between the second text feature and the first text feature in any one of the data pairs can be the match degree of any one of the data pairs. In the embodiment of the present application, the method of calculating the distance between two text features is not limited, such as calculating the L2 distance (also called Euclidean distance) between two text features, calculating the cosine distance between two text features, and calculating the L1 distance (also called Manhattan distance) between two text features.
[0049] In some embodiments, the match degree of any one of the data pairs can show a positive correlation with the similarity degree between the second text feature and the first text feature in any one of the data pairs. For example, the similarity degree between the second text feature and the first text feature in any one of the data pairs can be the match degree of any one of the data pairs. The similarity degree is used to represent the similarity degree between the second text feature and the first text feature. In the embodiment of the present application, the method of calculating the similarity degree between two text features is not limited, such as calculating the cosine similarity degree between two text features, calculating the Pearson similarity degree between two text features, etc.
[0050] In some embodiments, the data pair whose match degree meets the match condition refers to a data pair whose second text feature and the first text feature have high similarity. The match degree that meets the match condition can be flexibly adjusted according to the calculation method of the match degree. The following two cases can be referred to.
[0051] In case 1, when the match degree of any one data pair refers to the distance between the second text feature and the first text feature in any one data pair, the data pair whose match degree satisfies the match condition may refer to a data pair whose match degree is less than a distance threshold, or may refer to K (K is an integer equal to or greater than 1) data pairs whose match degrees are the smallest among all match degrees, where K is the number of target data pairs to be acquired. The distance threshold is set according to experience or flexibly adjusted according to application scenarios.
[0052] In case 2, when the match degree of any one data pair refers to the similarity between the second text feature and the first text feature in any one data pair, the data pair whose match degree satisfies the match condition may refer to a data pair whose match degree is greater than a similarity threshold, or may refer to K (K is an integer equal to or greater than 1) data pairs whose match degrees are the greatest among all match degrees. The similarity threshold is set based on experience or flexibly adjusted according to the application scene.
[0053] Before obtaining at least one target data pair that matches the first text feature from the data pair library, the data pair library needs to be constructed first. Illustratively, the process for constructing the data pair library includes the steps of obtaining a plurality of second texts and extracting features from each of the plurality of second texts to obtain a plurality of second text features. Since each second text corresponds to one second text feature, the second text feature corresponding to each second text and the standard translation text corresponding to each second text are all combined into one data pair, so that one data pair includes one second text feature and one standard translation text.
[0054] Here, the second text may be extracted from a sample text including the second text, where the sample text is a text in a first language. The standard translation text corresponding to the second text may be extracted from a standard translation text corresponding to the sample text, where the sample text is a text having a standard translation text. The standard translation text corresponding to the sample text may be obtained by translating the sample text by a professional translator, where the standard translation text is a text in a second language. The sample text and the standard translation text corresponding to the sample text express the same meaning using different languages. Illustratively, one sample text and a standard translation text corresponding to the one sample text may constitute one sample example, and a sample set may be constituted by a plurality of sample examples. In some embodiments, the sample examples may also be referred to as training examples, and the sample set may also be referred to as a training set.
[0055] In some embodiments, the process for extracting the second text feature of the second text can be realized by calling a text feature extraction model. In the embodiments of the present application, the type of the text feature extraction model is not limited, for example, the text feature extraction model can refer to a partial model for extracting text features in the NMT model. The method for extracting the second text feature is the same in principle as the method for extracting the first text feature in step 201, so it will not be further described here.
[0056] In the process of constructing the data pair library, a text feature extraction model (e.g., a submodel for extracting text features in an NMT model) is used to extract features from the second text in all sample examples in the sample set to obtain a plurality of second text features, and the second text features and standard translated texts corresponding to the second text features are recorded, and both are stored as data pairs in the data pair library. Illustratively, the second text features may also be referred to as a representation generated by a decoder according to the second text, and the standard translated text corresponding to the second text features may also be referred to as an accurate translated text corresponding to the second text features.
[0057] In some embodiments, each data pair can be represented as a key-value pair, where the second text feature in each data pair is the key and the standard translated text in each data pair is the value.
[0058] In some embodiments, given a sample set {(x,y)} (where (x,y) represents a sample example, x represents a sample text, and y represents a standard translation text corresponding to the sample text), a data pairs library D may be constructed based on formula (1) below:
[0059]
number
[0060] In some embodiments, when the number of at least one target data pair is K (where K is an integer equal to or greater than 1), the kth data pair (where k is an integer between 1 and K) among the K target data pairs is (h k ,v k ), where h k v represents the second text feature in the kth data pair. k represents the standard translation text in the kth data pair.
[0061] In some embodiments, this step 202 can be realized by invoking a target text translation model. That is, the target text translation model can be invoked to obtain at least one target data pair that matches the first text feature. Exemplarily, if the structure of the target text translation model is the structure described in step 201, invoking the target text translation model to obtain at least one target data pair that matches the first text feature can refer to invoking a second translation sub-model of the target text translation model to obtain at least one target data pair that matches the first text feature. Exemplarily, the second translation sub-model includes a data pair search network for searching a matching data pair from a data pair library, so that the process for obtaining at least one target data pair that matches the first text feature can be realized through the data pair search network in the second translation sub-model. Exemplarily, the data pair search network can be a simple feed-forward neural network or other more complex networks.
[0062] In step 203, a confidence and a match of at least one target data pair are determined, where the confidence of any one target data pair is used to indicate a degree of reliability of any one target data pair, and the match of any one target data pair is used to indicate a degree of similarity between the second text feature and the first text feature in any one target data pair.
[0063] Here, the method of determining the match degree of at least one target data pair has been described in step 202, and will not be further described here. In the method according to the embodiment of the present application, after obtaining at least one target data pair matching the first text feature, it is also necessary to determine the confidence of at least one target data pair respectively. The confidence of any one target data pair is used to indicate the degree of confidence of any one target data pair. Optionally, the confidence of any one target data pair is positively correlated with the degree of confidence of any one target data pair, i.e., the higher the confidence of any one target data pair, the higher the degree of confidence of that any one target data pair. By taking into account the confidence of at least one target data pair, the determined second probability can be made more reliable, and thus the accuracy of the translated text corresponding to the first text can be further improved.
[0064] In some embodiments, this step 203 can be realized by invoking a target text translation model. That is, the target text translation model can be invoked to determine the confidence of at least one target data pair. Optionally, if the target text translation model has the structure described in step 201, invoking the target text translation model to determine the confidence of at least one target data pair can refer to invoking a second translation sub-model of the target text translation model to determine the confidence of at least one target data pair. Optionally, the second translation sub-model further includes a probability distribution prediction network in addition to the data pair search network in step 202. The process for determining the confidence of at least one target data pair can be realized via a probability distribution prediction network in the second translation sub-model.
[0065] The principle of determining the reliability of each of the at least one target data pair is the same. In the embodiment of the present application, the process for determining the reliability of any one target data pair is taken as an example. In some embodiments, the implementation process for determining the confidence of any one of the target data pairs includes the steps of: determining at least one third probability, i.e., a third probability corresponding to each candidate text, based on second text features in any one of the target data pairs, where the at least one third probability is used to indicate the probability that the second text in any one of the target data pairs is translated into each candidate text, in other words, the third probability corresponding to any one of the candidate texts is used to indicate the probability that the second text corresponding to any one of the candidate texts is translated into any one of the candidate texts; determining a fourth probability based on the at least one third probability, where the fourth probability is used to indicate the probability that the second text corresponding to any one of the target data pairs is translated into a standard translation text in any one of the target data pairs; and determining the confidence of any one of the target data pairs based on the fourth probability.
[0066] The principle of determining the third probability corresponding to each candidate text based on the second text features is the same as the principle of determining the first probability corresponding to each candidate text based on the first text features, so it will not be further described here. The probability that the second text corresponding to any one of the target data pairs is translated into any one of the candidate texts is called the third probability. Based on the third probabilities corresponding to each of the multiple candidate texts, the probability that the second text is translated into the standard translation text in any one of the target data pairs is determined, which is called the fourth probability.
[0067] In some embodiments, the process for determining a probability that the second text is translated into the standard translation text in any one of the target data pairs based on the third probabilities corresponding to each of the plurality of candidate texts, i.e., the process for determining a fourth probability based on at least one third probability, includes determining a probability that the standard translation text is one of the candidate texts when the third probabilities corresponding to each of the plurality of candidate texts include a corresponding third probability of the standard translation text in any one of the target data pairs, and then determining a fourth probability based on the at least one third probability. The method includes a step of setting the corresponding third probability as a fourth probability, i.e., the probability that the second text will be translated into the standard translation text in any one of the target data pairs, and a step of setting the first numerical value as the fourth probability, i.e., the probability that the second text will be translated into the standard translation text in any one of the target data pairs, when the corresponding third probability of each candidate text does not include the corresponding third probability of the standard translation text in any one of the target data pairs, indicating that the standard translation text in any one of the target data pairs is not one of the candidate texts. In this case, the first numerical value is a numerical value equal to or less than the minimum value of the third probabilities corresponding to each candidate text, and for example, when the numerical range of each third probability is 0 to 1, the first numerical value may be 0. It can be shown that, when the standard translation text in one of the target data pairs is used as one of the candidate texts, the higher the probability that the second text is translated into the standard translation text in one of the target data pairs, the higher the probability that the standard translation text can be predicted based on the second text features, i.e., the higher the confidence level of the one of the target data pairs.
[0068] In some embodiments, the process for determining the reliability of any one of the target data pairs based on the probability that the second text is translated into the standard translation text in any one of the target data pairs, i.e., the fourth probability, includes the steps of converting the probability that the second text is translated into the standard translation text in any one of the target data pairs, i.e., converting the fourth probability, and taking the converted numerical value as the reliability of any one of the target data pairs. Illustratively, taking the reliability of at least one target data pair determined using a probability distribution prediction network in the second translation sub-model as an example, the probability that the second text is translated into the standard translation text in any one of the target data pairs, i.e., the fourth probability, is input into the probability distribution prediction network, the fourth probability is converted through the probability distribution prediction network, and the numerical value output from the probability distribution prediction network is taken as the reliability of any one of the target data pairs. The process of converting the probability that the second text will be translated into the standard translation text for any one of the target data pairs using the probability distribution prediction network is an internal calculation process of the probability distribution prediction network, and is not limited in the embodiments of the present application, as long as it can ensure that the output confidence shows a positive correlation with the probability that the second text will be translated into the standard translation text for any one of the target data pairs.
[0069] In some embodiments, a probability distribution prediction network is used to predict the kth (k is an integer between 1 and K) target data pair (h k ,v k The process for converting the fourth probability determined based on ( ) can be expressed using the following equation (2):
[0070]
number
[0071] In some embodiments, the implementation process for determining the confidence of any one of the target data pairs based on the probability that the second text will be translated into the standard translation text in any one of the target data pairs, i.e., the fourth probability, includes the steps of determining a fifth probability based on first probabilities corresponding to each of a plurality of candidate texts, i.e., the at least one first probability, where the fifth probability is used to indicate the probability that the first text will be translated into the standard translation text in any one of the target data pairs, and determining the confidence of any one of the target data pairs based on the probability that the second text will be translated into the standard translation text in any one of the target data pairs and the probability that the first text will be translated into the standard translation text in any one of the target data pairs, i.e., determining the confidence of any one of the target data pairs based on the fourth probability and the fifth probability.
[0072] In some embodiments, the implementation process for determining a probability that the first text is translated into the standard translation text in any one of the target data pairs based on the first probabilities corresponding to each of the plurality of candidate texts, i.e., the process for determining the fifth probability based on at least one first probability, includes determining a probability that the standard translation text in any one of the target data pairs is a candidate text among the respective candidate texts when the first probabilities corresponding to each of the candidate texts include a first probability corresponding to the standard translation text in any one of the target data pairs, and then determining a probability that the standard translation text in any one of the target data pairs is a candidate text among the respective candidate texts. and a step of setting the first probability corresponding to the standard translation text in the target data pair to a fifth probability, i.e., the probability that the first text will be translated into the standard translation text in any one of the target data pairs, and setting the second numerical value as the probability that the first text will be translated into the standard translation text in any one of the target data pairs when the first probabilities corresponding to each of the plurality of candidate texts do not include the first probability corresponding to the standard translation text in any one of the target data pairs, indicating that the standard translation text in the one of the target data pairs is not one of the plurality of candidate texts. In this case, the second numerical value is a numerical value equal to or less than the minimum value of the first probabilities corresponding to each of the plurality of candidate texts. For example, if the numerical range of each of the first probabilities is 0 to 1, the second numerical value may be 0. It can be shown that, when the standard translation text in any one target data pair is one of the candidate texts, the higher the probability that the first text can be translated into the standard translation text in any one target data pair, the higher the probability that the standard translation text can be predicted based on the first text features.Because the similarity between the first text feature and the second text feature in any one of the target data pairs is high, the higher the probability that the standard translated text in any one of the target data pairs can be predicted based on the first text feature prediction, which to some extent indicates that the standard translated text in any one of the target data pairs can be predicted based on the second text feature in the any one of the target data pairs, i.e., the higher the reliability of the any one of the target data pairs.
[0073] In some embodiments, the implementation process for determining the reliability of any one of the target data pairs based on the probability that the second text is translated into the standard translation text in any one of the target data pairs and the probability that the first text is translated into the standard translation text in any one of the target data pairs, i.e., the implementation process for determining the reliability of the target data pair based on the fourth probability and the fifth probability, is to input the probability that the second text is translated into the standard translation text in any one of the target data pairs and the probability that the first text is translated into the standard translation text in any one of the target data pairs into an input probability distribution prediction network, use the probability distribution prediction network to convert the probability that the second text is translated into the standard translation text in any one of the target data pairs and the probability that the first text is translated into the standard translation text in any one of the target data pairs, and use the probability distribution prediction network to determine the reliability of any one of the target data pairs. That is, the fourth probability and the fifth probability are input to an input probability distribution prediction network, the fourth probability and the fifth probability are converted using the probability distribution prediction network, and the numerical value output from the probability distribution prediction network is the reliability of any one of the target data pairs. The process of converting the fourth probability and the fifth probability using the probability distribution prediction network is an internal calculation process of the probability distribution prediction network, and is not limited in the embodiments of the present application, as long as it is ensured that the output reliability shows a positive correlation with the fourth probability and the fifth probability.
[0074] In some embodiments, the confidence level of at least one target data pair may be determined according to equation (3) below.
[0075]
number
number
number
number
[0076] In some embodiments, the method for determining the confidence level of any one of the target data pairs may further include determining a probability that the first text will be translated into a standard translation text in any one of the target data pairs based on the first probabilities corresponding to each of the plurality of candidate texts, and converting the probability that the first text will be translated into a standard translation text in any one of the target data pairs and using the converted numerical value as the confidence level of any one of the target data pairs. That is, the method includes determining a fifth probability based on at least one of the first probabilities, the fifth probability being used to indicate the probability that the first text will be translated into a standard translation text in any one of the target data pairs, and converting the fifth probability and using the converted numerical value as the confidence level of any one of the target data pairs.
[0077] In some embodiments, taking the probability distribution prediction network in the second translation sub-model to determine the reliability of at least one target data pair as an example, the probability that the first text is translated into the standard translation text in any one target data pair is input into the probability distribution prediction network, the probability distribution prediction network is used to convert the probability that the first text is translated into the standard translation text in any one target data pair, and the numerical value output from the probability distribution prediction network is the reliability of any one target data pair. That is, the fifth probability is input into the probability distribution prediction network, the probability distribution prediction network is used to convert the fifth probability, and the numerical value output from the probability distribution prediction network is the reliability of any one target data pair. The process of converting the probability that the first text is translated into the standard translation text in any one target data pair using the probability distribution prediction network is an internal calculation process of the probability distribution prediction network, so it is not limited in the embodiments of the present application, and it is sufficient to ensure that the output reliability shows a positive correlation with the probability that the first text is translated into the standard translation text in any one target data pair.
[0078] For example, a probability distribution prediction network is used to predict the kth (k is an integer from 1 to K) target data (h k ,v k The process for converting the fifth probability determined based on (x,y) can be expressed using the following equation (4):
[0079]
number
number
number
[0080] In step 204, at least one second probability is determined based on the confidence and match of the at least one target data pair.
[0081] Here, at least one second probability is used to indicate the probability that the first text is translated into each standard translation text in at least one target data pair. The second probability corresponding to any one of the standard translation texts is used to indicate the probability that the first text is translated into any one of the standard translation texts. Each standard translation text must be a translation text that does not overlap. For example, if the standard translation texts of two of the ten searched target data pairs are the same, the number of standard translation texts at that time is nine. When calculating the second probability corresponding to each standard translation text, the probabilities corresponding to the same standard translation text may be added.
[0082] In some embodiments, this step 204 can be realized by invoking a target text translation model. That is, the target text translation model can determine a second probability corresponding to each of the standard translation texts in the at least one target data pair based on the confidence and match of the at least one target data pair. Illustratively, if the target text translation model has the structure described in step 201, invoking the target text translation model and determining the confidence of the at least one target data pair can refer to invoking a second translation sub-model of the target text translation model and determining the confidence of the at least one target data pair. Illustratively, the second translation sub-model further includes a probability distribution prediction network in addition to the data pair search network associated with step 202. The process for determining the second probability corresponding to each of the standard translation texts in the at least one target data pair can be realized through the probability distribution prediction network in the second translation sub-model. For example, since not only the degree of match but also the reliability are taken into account in the process of determining the second probability, the probability distribution prediction network can be regarded as a distribution calibration (DC) network for a network that determines the second probability by considering only the degree of match.
[0083] In some embodiments, determining a second probability corresponding to each of the standard translation texts in the at least one target data pair based on the confidence and match degree of the at least one target data pair includes the steps of: standardizing the match degree of a first data pair for any one of the standard translation texts among the plurality of standard translation texts to obtain a standardized match degree, where the first data pair is a data pair including any one of the standard translation texts in the at least one target data pair; correcting the standardized match degree using the confidence of the first data pair to obtain a corrected match degree; and determining a probability that is positively correlated with the corrected match degree as the second probability corresponding to any one of the standard translation texts, i.e., determining the second probability corresponding to any one of the standard translation texts based on the corrected match degree, where the corrected match degree and the second probability are positively correlated.
[0084] The first data pair is a data pair including one standard translation text of at least one target data pair. The number of first data pairs may be one or more. Each first data pair has a match degree and a confidence level. Standardizing the match degree of the first data pair to obtain a standardized match degree refers to standardizing the match degree of each first data pair and obtaining a standardized match degree corresponding to each first data pair. Correcting the standardized match degree using the confidence level of the first data pair to obtain a corrected match degree refers to correcting the standardized match degree corresponding to each first data using the confidence level of each first data pair and obtaining a corrected match degree corresponding to each first data pair.
[0085] After obtaining the match degree of the first data pair, the match degree can be standardized to obtain a standardized match degree, thereby improving the normativeness of the match degree of the first data pair. Taking one first data pair as an example, in some embodiments, the method for standardizing the match degree of the first data pair may be to standardize the match degree of the first data pair using a hyperparameter. Before standardizing the match degree of the first data pair using the hyperparameter, it is necessary to further determine the size of the hyperparameter. Here, the value of the hyperparameter may be set according to experience or flexibly adjusted according to the target data, but is not limited in the embodiment of the present application.
[0086] In the embodiment of the present application, an example is described in which the hyperparameter is dynamically determined according to the target data pair. The process for determining the hyperparameter includes a step of determining the hyperparameter based on at least one information of a quantitative index of each target data pair and a match degree of each target data pair. Here, the quantitative index of any one target data pair is the number of target data pairs that are not positioned lower than any one target data pair after each target data pair is sorted according to the reference order.
[0087] In an embodiment of the present application, the determination of the hyperparameter is related to one of two data including a quantity index of each target data pair and a match degree of each target data pair. Here, the quantity index of any one target data pair is the number of target data pairs that are not ranked lower than any one target data pair after each target data pair is sorted according to the reference order. The reference order can be set by experience or flexibly adjusted according to the application scene. For example, different target data pairs are numbered differently, and the reference order can refer to the order of the numbers being higher or lower. After each target data pair is sorted according to the reference order, each target data pair has its own position, and the number of non-overlapping standard translation texts among each target data pair that are not ranked lower than any one target data pair is the quantity index of any one target data pair.
[0088] For example, suppose the number of searched target data pairs is three, data pair 1, data pair 2, and data pair 3, the standard translation texts for data pair 1 and data pair 2 are both M1, and the standard translation text for data pair 3 is M2. If it is assumed that data pair 1, data pair 2, and data pair 3 are ranked from top to bottom after being sorted according to the reference order, the quantity index of data pair 1 is 1, the quantity index of data pair 2 is 2, and the quantity index of data pair 3 is 2.
[0089] The hyperparameters may be determined based only on the quantitative index of each target data pair, or based only on the match degree of each target data pair, or may be determined based on the quantitative index of each target data pair and the match degree of each target data pair. The numerical value of the hyperparameter can be obtained by inputting at least one of the quantitative index of each target data pair and the match degree of each target data pair into the probability distribution prediction network and performing calculations.
[0090] Taking the example of determining hyperparameters based on the quantitative indicators of each target data pair and the degree of match of each target data pair as an example, the hyperparameters can be calculated according to the following formula (5).
[0091]
number
[0092] In some embodiments, the method of standardizing the degree of match of the first data pair using the hyperparameter may be such that the ratio of the degree of match of the first data pair to the hyperparameter is the standardized degree of match, or the product of the degree of match of the first data pair and the hyperparameter is the standardized degree of match, but is not limited to this in the embodiments of the present application.
[0093] After obtaining the standardized match degree corresponding to the first data pair, the standardized match degree is corrected using the confidence level of the first data pair to obtain a corrected match degree. The corrected match degree is a match degree that matches the degree of confidence of the first data pair. Illustratively, the method of correcting the standardized match degree using the confidence level of the first data pair may be related to the specific situation of the match degree of the first data pair. For example, when the match degree of the first data pair shows a positive correlation with the similarity between the second text feature and the first text feature in the first data pair, the sum of the confidence level of the first data pair and the standardized match degree can be the corrected match degree. On the other hand, when the match degree of the first data pair shows a negative correlation with the similarity between the second text feature and the first text feature in the first data pair, the difference between the confidence level of the first data pair and the standardized match degree can be the corrected match degree.
[0094] In some embodiments, when there is one first data pair, the probability that the corrected match degree determined according to the one first data pair is positively correlated is determined as the second probability corresponding to any one of the standard translation texts. On the other hand, when there is a plurality of first data pairs, the sum of the corrected match degrees determined according to the plurality of first data pairs is calculated, and the second probability corresponding to any one of the standard translation texts is determined as the probability that the calculated sum of the match degrees is positively correlated.
[0095] For example, one of the standard translation texts v k For example, the second probability corresponding to any one of the standard translation texts can be calculated according to the following formula (6).
[0096]
number
number
number
number
number
number
[0097] By referring to the method for obtaining the second probability corresponding to any one of the standard translation texts, the second probability corresponding to each of the standard translation texts can be determined.
[0098] It should be noted that in the embodiment of the present application, the order of determining the first probability corresponding to each of the plurality of candidate texts and the second probability corresponding to each of the plurality of standard translation texts is not limited and can be flexibly set according to actual needs. After determining the first probability corresponding to each of the candidate texts and the second probability corresponding to each of the standard translation texts, step 205 is performed.
[0099] In step 205, a translated text corresponding to the first text is determined based on the at least one first probability and the at least one second probability.
[0100] Here, the at least one first probability is a first probability corresponding to each of the plurality of candidate texts, and the at least one second probability is a second probability corresponding to each of the plurality of standard translation texts. Accordingly, step 205 may be expressed as determining a translation text corresponding to the first text based on the first probability corresponding to each of the plurality of candidate texts and the second probability corresponding to each of the standard translation texts.
[0101] The translated text corresponding to the first text refers to the first text translated into the second language accordingly. The first probability corresponding to each of the plurality of candidate texts and the second probability corresponding to each of the plurality of standard translation texts are comprehensively considered to determine the translated text corresponding to the first text. The information considered in the process of determining the translated text corresponding to the first text is rich, which contributes to ensuring the reliability of the translated text corresponding to the first text. In addition, the second probability corresponding to each of the standard translation texts is determined by comprehensively considering the match degree and reliability of the target data pair, so the information taken into consideration is rich, and the determined second probability matches the reliability degree of the target data pair, which increases the reliability of the second probability, thereby contributing to further increasing the reliability of the translated text corresponding to the first text.
[0102] In some embodiments, the process for determining a translation text corresponding to a first text based on a first probability corresponding to each candidate text and a second probability corresponding to each standard translation text includes the steps of: determining a first rough probability distribution based on the first probability corresponding to each candidate text; determining a second probability distribution based on the second probability corresponding to each standard translation text; fusing the first probability distribution and the second probability distribution to obtain a fused probability distribution, where the fused probability distribution includes translation probabilities corresponding to each target text, each target text including each candidate text and each standard translation text; and determining the target text with the highest translation probability among the target texts as the translation text. That is, the method includes the steps of: determining a first probability distribution based on at least one first probability; determining a second probability distribution based on at least one second probability; fusing the first probability distribution and the second probability distribution to obtain a fused probability distribution, where the fused probability distribution includes a translation probability of each target text, where each target text includes each candidate text and each standard translation text; and selecting the target text with the highest translation probability among the target texts as the translation text.
[0103] Here, the first probability distribution includes a first probability corresponding to each candidate text, and the second probability distribution includes a second probability corresponding to each standard translation text. In the embodiment of the present application, the method of fusing the obtained first probability distribution and the second probability distribution is not limited, as long as a fused probability distribution including the translation probability corresponding to each target text can be obtained. Here, each target text includes each candidate text and each standard translation text. In other words, each target text is a text that does not overlap with each candidate text and each standard translation text. For example, a translation text corresponding to the first text can be obtained by fusing the first probability distribution and the second probability distribution using a method based on weighted interpolation.
[0104] In some embodiments, fusing the first probability distribution and the second probability distribution to obtain a fused probability distribution includes determining a first importance and a second importance, where the first importance is used to indicate an importance of the first probability distribution in obtaining the translation text, and the second importance is used to indicate an importance of the second probability distribution in obtaining the translation text; determining a target parameter based on the first importance and the second importance, where the target parameter is also referred to as a normalization parameter; transforming the first importance based on the target parameter to obtain a first weight; transforming the second importance based on the target parameter to obtain a second weight; and fusing the first probability distribution and the second probability distribution based on the first weight and the second weight to obtain a fused probability distribution.
[0105] In some embodiments, this step 205 can be realized by calling a target text translation model. That is, the target text translation model can determine a translation text corresponding to the first text based on a first probability corresponding to each candidate text and a second probability corresponding to each standard translation text. That is, the target text translation model determines a translation text corresponding to the first text based on at least one first probability and at least one second probability. If the structure of the target text translation model is the structure described in step 201, this step 205 can be realized by calling a third translation sub-model of the target text translation model. Exemplarily, the third translation sub-model can include a weight prediction network (WP) for predicting the first weight and the second weight, and a fusion network for fusing the first probability distribution and the second probability distribution according to the first weight and the second weight.
[0106] In some embodiments, the first importance is calculated by the weight prediction network from at least one of the following information: a probability that each standard translation text may be predicted based on the first text feature, a probability that each standard translation text in each target data pair may be predicted based on the second text feature in each target data pair, and a first probability to which each candidate text corresponds, i.e., the first importance is calculated by the weight prediction network from at least one of the following information: at least one fifth probability, at least one fourth probability, and at least one first probability.
[0107] In some embodiments, taking the first importance as an example, the weighted prediction network calculates the probability that each standard translation text can be predicted based on the first text feature, the probability that each standard translation text in each target data pair can be predicted based on the second text feature in each target data pair, and each candidate text has a corresponding first probability. The first importance can be calculated according to the following formula (7):
[0108]
number
number
number
[0109] In some embodiments, the second importance is determined by the weighted prediction network according to at least one of the following information: the quantity indicator of each target data pair and the match degree of each target data pair. Optionally, for example, the second importance is determined by the weighted prediction network according to the quantity indicator of each target data pair and the match degree of each target data pair, and the second importance can be calculated according to the following formula (8):
[0110]
number
[0111] After calculating the first importance and the second importance, a normalization parameter is determined based on the first importance and the second importance, the normalization parameter is a parameter based on which the first importance and the second importance are converted, and the sum of the first weight and the second weight obtained by converting the first importance and the second importance according to the normalization parameter is 1. Optionally, the second weight can be calculated according to the following formula (9).
[0112]
number
[0113] In some embodiments, the sum of the first importance and the second importance may be taken as the target parameter, and then the ratio of the first importance to the target parameter may be taken as the first weight, and the ratio of the second importance to the target parameter may be taken as the second weight.
[0114] In some embodiments, a fusion probability distribution may be calculated based on the first weight and the second weight according to the following equation (10):
[0115]
number
[0116] In some embodiments, the fusion probability distribution obtained according to the above formula (10) includes translation probabilities corresponding to each of the multiple target texts. Among the multiple target texts, a text with the highest translation probability is determined, and the determined text is the translation text corresponding to the first text.
[0117] FIG. 3 is a schematic diagram of a text translation model based on confidence. FIG. 3 takes an NMT translation model as an example and shows a translation process including the above-mentioned steps 201 to 205, from inputting a first text to outputting a translated text corresponding to the first text by the model. In FIG. 3, 301 is a first text to be translated input to the model, and the first text is a Chinese text, 302 is an NMT translation model, 303 is a first text feature, 304 is a first probability distribution, 305 is a data pair library, 306 is at least one target data pair searched according to the first text feature, 307 is a second probability distribution, 308 is a fusion probability distribution obtained by fusing the first probability distribution and the second probability distribution, and 309 is a translation text corresponding to the outputted first text.
[0118] According to the technical method of the embodiment of the present application, in the process of determining the second probability, in addition to the match degree between the second text feature and the first text feature in the target data pair, the reliability of the target data pair is also taken into consideration, so that a wealth of information is taken into consideration. In addition, the reliability of the target data pair is used to evaluate the degree of reliability of the target data pair, so that by taking the reliability of the target data pair into consideration, the reliability of the second probability can be increased, and the accuracy of the text translation can be further improved.
[0119] The embodiment of the present application provides a method for acquiring a text translation model applicable to the implementation environment shown in Fig. 1 above. The method for acquiring a text translation model is executed by a computer device. The computer device may be a terminal 11 or a server 12, but is not limited to this in the embodiment of the present application. As shown in Fig. 4, the method for acquiring a text translation model according to the embodiment of the present application includes the following steps 401 to 407.
[0120] In step 401, a first sample text in a first language, a first standard translation text, and an initial text translation model are obtained.
[0121] Here, the first sample text is a text in a first language, and the first standard translation text is a text obtained by translating the first sample text into a second language.
[0122] In some embodiments, when a translation from Chinese to English is required, the first language is Chinese and the second language is English. The first sample text is a text with a standard translation. In an embodiment of the present application, the first standard translation text is a standard translation text of the first sample text. In order to facilitate providing supervised data to the training process of the initial text translation model using the standard translation text corresponding to the first sample text, the language of the standard translation text corresponding to the first sample text is the same as the language of the translation text that needs to be output using the initial text translation model. Since the first sample text corresponds to a standard translation text, the process for training the initial text translation model using the first sample text is a supervised training process.
[0123] In addition, the first sample text is a text that is used as a basis for training the text translation model once. The number of the first sample texts may be one or more, but is not limited in the embodiment of the present application. In the embodiment of the present application, the number of the first sample texts is described as one. The method of obtaining the first sample text can refer to the relevant process of step 201 in the embodiment shown in FIG. 2, and will not be described further here.
[0124] In step 402, a first sample text feature is processed by an initial text translation model to obtain at least one first sample probability.
[0125] Here, the first sample text feature is a text feature of the first sample text. The at least one first sample probability is used to indicate a probability that the first sample text is translated into each of the at least one candidate text. The first sample text corresponding to any one of the candidate texts is used to indicate a probability that the first sample text is translated into any one of the candidate texts. Each of the at least one candidate text is a text in a second language.
[0126] The implementation process of step 402 can refer to step 201 in the embodiment shown in FIG. 2, so no further description will be given here.
[0127] In step 403, at least one sample data pair that matches the first sample text feature is obtained, where one sample data pair includes a second sample text feature and a second standard translation text.
[0128] Here, the second sample text features are text features of the second sample text, the second sample text is a text in a first language, and the second standard translation text is a text obtained by translating the second sample text into a second language.
[0129] In some embodiments, obtaining at least one sample data pair matching the first sample text feature includes searching a data pair library for at least one initial data pair matching the first sample text feature, where any one initial data pair includes a third sample text feature and a second standard translated text, the third sample text feature being a text feature of the second sample text, the second sample text being text in a first language, and the second standard translated text being text of the second sample text translated into a second language; and determining the at least one sample data pair based on the at least one initial data pair.
[0130] In some embodiments, the method of determining the at least one sample data pair based on the at least one initial data pair is to use the at least one initial data pair as the at least one sample data pair. In such a case, the third sample text feature of the second sample text in the initial data pair is directly used as the second sample text feature of the second sample text in the sample data pair.
[0131] In some embodiments, a method for determining at least one sample data pair based on at least one initial data pair includes the steps of interference processing the at least one initial data pair according to an interference probability to obtain an interference-processed data pair, and determining at least one sample data pair based on the interference-processed data pair.
[0132] Because the data pair library and the first sample text may not be perfectly matched, and at least one of the retrieved sample data pairs may not contain the first standard translation text, in the model training stage, perturbation can be added to at least one initial data pair (i.e., interference processing is performed on at least one initial data pair) to make the model more robust, and thus the accuracy of the translation result of the model can be improved.
[0133] Exemplarily, the interference probability may be set according to experience. Exemplarily, the interference probability may be determined according to the number of updates corresponding to the initial text translation model. Exemplarily, the interference probability shows a negative correlation with the number of updates corresponding to the initial text translation model. For example, the ratio between the number of updates corresponding to the initial text translation model and the falling speed of the interference probability is determined, a numerical value showing a negative correlation with the ratio is determined, and the product of the numerical value and the initial interference probability is set as the interference probability. The initial interference probability and the falling speed of the interference probability may be set according to experience or flexibly adjusted according to the application scene, but are not limited in the embodiment of the present application.
[0134] For example, the interference probability can be calculated according to the following equation (11).
[0135] α=α 0 *exp(-step / β) (11) Here, α 0 is the initial interference probability; β is the rate of decline of the interference probability; step is the number of updates corresponding to the initial text translation model; and α is the interference probability. According to the above formula (11), it can be seen that the greater the number of updates corresponding to the initial text translation model, the smaller the interference probability α becomes.
[0136] In some embodiments, interfering with at least one initial data pair according to an interference probability refers to processing the at least one initial data pair with interference with a probability of the interference probability, and processing the at least one initial data pair without interference with a probability of (1-interference probability).
[0137] In some embodiments, when the interference probability includes a first interference probability, interference processing at least one initial data pair according to the interference probability to obtain an interference-processed data pair includes adding a noise feature to a third sample text feature in each initial data pair according to the first interference probability to obtain an interference-processed data pair. In such a case, the method of determining the at least one sample data pair based on the interference-processed data pair is to determine the interference-processed data pair as the at least one sample data pair. The first interference probability is a probability of performing an interference scheme of adding a noise feature to a third sample text feature in each data pair.
[0138] To address the problem that the data pair library and the first sample text may not be perfectly matched, a noise feature can be added to the third sample text feature of at least one retrieved initial data pair to construct a noisy data pair. The second sample text feature in the noisy data pair can be constructed according to the following formula (12).
[0139]
number
[0140] If the data pair library and the first sample text are not perfectly matched, the retrieved at least one initial data pair cannot effectively assist the model in completing training, so a noise feature is added to the third sample text feature of the at least one initial data pair, and the second sample text feature is shifted from the third sample text feature of the initial data pair, so that the data pair library and the first sample text can be more closely matched. Note that, in this process, the second standard translation text in each initial data pair is not changed.
[0141] Fig. 5 is a schematic diagram of constructing a noisy data pair, in which 501 is at least one initial data pair searched in a data pair library, 502 is an added noise feature, and 503 is a sample data pair constructed by adding the noise feature.
[0142] In some embodiments, when the interference probability includes a second interference probability, interference processing at least one initial data pair according to the interference probability to obtain an interference-processed data pair includes deleting an initial data pair that does not satisfy a match condition among the at least one initial data pair according to the second interference probability to obtain an interference-processed data pair. In such a case, the method of determining at least one sample data pair based on the interference-processed data pair includes constructing a reference data pair based on a first sample text feature and a first standard translation text, the number of the reference data pairs being the same as the number of the deleted initial data pairs, and determining at least one sample data pair based on the interference-processed data pair and the reference data pair. The second interference probability is a probability of performing an interference scheme of deleting an initial data pair that does not satisfy a match condition among the at least one initial data pair. Exemplarily, the second interference probability may be the same as the first interference probability or may be different from the first interference probability.
[0143] In some embodiments, if the second standard translation text does not include the first standard translation text, a reference data pair may be constructed based on the first sample text features and the first standard translation text to enable at least one sample data pair to include the first standard translation text. Optionally, constructing a reference data pair based on the first sample text features and the first standard translation text may refer to directly constructing a reference data pair based on the first sample text features and the first standard translation text, or may refer to adding noise features to the first sample text features and constructing a reference data pair based on the sample text features obtained by adding the noise features and the first standard translation text.
[0144] In some embodiments, determining at least one sample data pair based on the interference-processed data pair and the reference data pair refers to determining both the interference-processed data pair and the reference data pair as the sample data pair. In the process of determining the interference-processed data pair as the sample data pair, the third sample text feature in the interference-processed data pair is determined as the second sample text feature in the sample data pair, and the second standard translation text in the interference-processed data pair is determined as the second standard translation text in the sample data pair. In the process of determining the reference data pair as the sample data pair, the first sample text feature in the sample data pair or the sample text feature obtained by adding a noise feature to the first sample text feature is determined as the second sample text feature in the sample data pair, and the first standard translation text in the reference data pair is determined as the second standard translation text in the sample data pair.
[0145] Fig. 6 is a schematic diagram for obtaining sample data pairs, in which 601 is at least one initial data pair searched in a data pair library, 602 is a reference data pair constructed based on a first sample text feature and a first standard translation text, and 603 is at least one determined sample data pair.
[0146] In some embodiments, according to the second interference probability, deleting the initial data pair that does not satisfy the match condition among the at least one initial data pair can refer to deleting the initial data pair that has the longest distance between the third sample text feature and the first sample text feature among the at least one initial data pair. As shown in FIG. 6, deleting the initial data pair that has the longest distance between the third sample text feature and the first sample text feature among the at least one initial data pair is illustrated. Optionally, not satisfying the match condition may be that the distance between the third sample text feature and the first sample text feature is greater than a distance threshold. The distance threshold can be set according to experience or flexibly adjusted according to actual circumstances, but is not limited in the embodiment of the present application.
[0147] In some embodiments, when the interference probability includes a first interference probability and a second interference probability, interference processing at least one initial data pair according to the interference probability to obtain interference-processed data pairs includes adding a noise feature to the third sample text feature in each initial data pair according to the first interference probability to obtain intermediate data pairs, and deleting an initial data pair that does not satisfy the match condition among the intermediate data pairs according to the second interference probability to obtain interference-processed data pairs. In such a case, the method for determining at least one sample data pair based on the interference-processed data pairs includes constructing a reference data pair based on the first sample text feature and the first standard translation text, where the number of the reference data pairs is the same as the number of the deleted initial data pairs, and determining at least one sample data pair based on the interference-processed data pair and the reference data pair.
[0148] In some embodiments, when the interference probability includes a first interference probability and a second interference probability, interference processing at least one initial data pair according to the interference probability to obtain an interference-processed data pair includes deleting an initial data pair that does not satisfy a match condition among the at least one initial data pair according to the second interference probability to obtain an intermediate data pair, and adding a noise feature to the third sample text feature in the intermediate data pair according to the first interference probability to obtain an interference-processed data pair. In such a case, the method for determining at least one sample data pair based on the interference-processed data pair includes constructing a reference data pair based on the first sample text feature and the first standard translation text, where the number of the reference data pairs is the same as the number of the deleted initial data pairs, and determining at least one sample data pair based on the interference-processed data pair and the reference data pair.
[0149] In some embodiments, unlike the process of obtaining a translation text of the first text using the text translation model shown in FIG. 3, in the process of training the initial text translation model, a certain interference is added to the initial data pair searched for in the data pair library, and the confidence and the translation text are determined on the data pair with the added interference, thereby greatly improving the robustness of the model and resisting noise interference.
[0150] At step 404, a confidence and a match of at least one sample data pair are determined. The confidence of any one sample data pair is used to indicate a degree of reliability of any one sample data pair. The match of any one sample data pair is used to indicate a degree of similarity between the second sample text feature and the first sample text feature in any one sample data pair.
[0151] The implementation process of step 404 can refer to step 203 in the embodiment shown in FIG. 2, and therefore will not be described further here.
[0152] In step 405, at least one second sample probability is determined based on the confidence and match of the at least one sample data pair.
[0153] Here, at least one second sample probability is used to indicate a probability that the first sample text is translated into each second standard translation text in the at least one sample data pair, and the second sample probability corresponding to any one of the second standard translation texts is used to indicate a probability that the first sample text is translated into any one of the second standard translation texts.
[0154] The implementation process of step 405 can refer to step 204 in the embodiment shown in FIG. 2, and therefore will not be described further here.
[0155] In step 406, predictive translated text corresponding to the first sample text is determined based on the at least one first sample probability and the at least one second sample probability.
[0156] The implementation process of step 406 can refer to step 205 in the embodiment shown in FIG. 2, so no further description will be given here.
[0157] In step 407, the initial text translation model is updated based on the difference between the predicted translated text and the first standard translated text to obtain a target text translation model.
[0158] In some embodiments, a result loss is obtained based on the predictively translated text corresponding to the first sample text and the first standard translated text, and the result loss is used to indicate a difference between the predictively translated text corresponding to the first sample text and the first standard translated text. The result loss is used to update model parameters of the initial text translation model to obtain a target text translation model.
[0159] After obtaining the predictive translation text corresponding to the first sample text, obtain a result loss based on the predictive translation text corresponding to the first sample text and the first standard translation text. In the embodiment of the present application, the method of obtaining the result loss based on the predictive translation text corresponding to the first sample text and the first standard translation text is not limited. Illustratively, the cross entropy loss or the mean squared error loss between the predictive translation text corresponding to the first sample text and the first standard translation text is taken as the result loss.
[0160] After obtaining the result loss, use the result loss to update the model parameters of the initial text translation model. The result loss used to update the model parameters of the initial text translation model may refer to using the result loss to update the entire model parameters of the initial text translation model, or may refer to using the result loss to update part of the model parameters of the initial text translation model (for example, other model parameters except the model parameters of the first translation sub-model), but is not limited in the embodiment of the present application.
[0161] After updating the model parameters of the initial text translation model using the resulting loss, a trained text translation model is obtained. It is determined whether the trained text translation model satisfies the training end condition. If the trained text translation model satisfies the training end condition, the trained text translation model is set as the target text translation model. On the other hand, if the trained text translation model does not satisfy the training end condition, the trained text translation model is continued to be updated by referring to steps 401 to 407 until a text translation model that satisfies the training end condition is obtained and the text translation model that satisfies the training end condition is set as the target text translation model.
[0162] The training end condition may be set by experience or flexibly adjusted according to the application scene, but is not limited in the embodiment of the present application. For example, the training end condition of the trained text translation model includes, but is not limited to, any one of the following: the number of model parameter updates performed at the time of obtaining the trained text translation model reaches a number threshold, the result loss at the time of obtaining the trained text translation model is less than a loss threshold, or the result loss at the time of obtaining the trained text translation model has converged.
[0163] In the technical method according to the embodiment of the present application, the interference probability is dynamically determined based on the update times corresponding to the initial text translation model, so that the addition of the interference probability is more reasonable. In addition, by performing interference processing on at least one initial data pair according to the interference probability and obtaining an interference-processed data pair, the problem that the data pair library and the first sample text are not completely matched, or the first standard translation text is not included in the at least one retrieved sample data pair can be resolved to a certain extent, so that the accuracy of the translation result of the model can be further improved.
[0164] In the technical method according to the embodiment of the present application, in addition to the match degree between the second sample text feature and the first sample text feature in the sample data pair, the reliability of the sample data pair is also taken into consideration in the process of determining the second sample probability, so that the information taken into consideration is rich. In addition, the reliability of the sample data pair is used to evaluate the reliability of the sample data pair, so that the reliability of the second sample probability can be improved by taking the reliability of the sample data pair into consideration, which can further improve the accuracy of the pre-translated text, improve the efficiency of model acquisition and the reliability of the acquired model, and further improve the accuracy of text translation using the model.
[0165] The text translation method according to the embodiment of the present application can be regarded as text translation based on the k-Nearest-Neighbor Machine Translation (kNN-MT) method. The kNN-MT method is considered to be an important research direction in neural machine translation tasks. Such a method assists in the generation of translations by searching for useful key-value pairs from a constructed data pair library, and does not require updating the NMT model in this process. However, the potential noise samples searched for may severely destroy the performance of the model. Therefore, in order to improve the robustness of the model, the embodiment of the present application proposes a robust k-nearest neighbor machine translation model based on reliability. Specifically, while the reliability of the NMT model itself is not considered in the conventional method, the embodiment of the present application introduces NMT reliability, a distribution correction network, and a weight prediction network to optimize the distribution of prediction by the k-nearest neighbor method and the weight of distribution interpolation. In addition, a robust training method, including adding two types of interference to the search result, has been added in the training process, which can further improve the model's ability to resist noise search results.
[0166] Compared with the conventional k-nearest neighbor machine translation model, the embodiment of the present application adds the reliability information of the NMT model to the model structure, and optimizes the prediction of the k-nearest neighbor distribution and the interpolation weight through two networks (distribution correction network and weight prediction network). By considering the reliability information of the NMT model, the model can better balance the weight of the k-nearest neighbor distribution and the NMT predicted distribution, and avoid the performance of the model being reduced due to the weight of the noisy k-nearest neighbor distribution being too large. In addition, two kinds of interference are added during the training process, so that the robustness of the model can be improved while further avoiding the influence of noise during the model training process.
[0167] Referring to FIG. 7, an embodiment of the present application provides a text translation device, which comprises: a determination module 701 configured to perform a step of determining at least one first probability based on first text features, the first text features being text features of a first text, the first text being a text in a first language, the at least one first probability being used to indicate a probability that the first text will be translated into each of at least one candidate text, each of the at least one candidate text being a text in a second language; an acquisition module 702 configured to perform a step of acquiring at least one target data pair matching a first text feature, where any one target data pair includes a second text feature and a standard translation of the second text, the second text feature being a text feature of the second text, the second text being a text in the first language, and the standard translation being a text in the second language; a determination module 701 configured to further perform a step of determining a confidence and a match of at least one target data pair, where the confidence of any one target data pair is used to indicate a degree of reliability of any one target data pair, and the match of any one target data pair is used to indicate a similarity between the second text feature and the first text feature in any one target data pair; A determination module 701 configured to further perform a step of determining at least one second probability based on the confidence and match of the at least one target data pair, the at least one second probability being used to indicate the probability that the first text is translated into each standard translation text in the at least one target data pair; and a determination module 701 configured to further perform the step of determining a translated text corresponding to the first text based on the at least one first probability and the at least one second probability.
[0168] In some embodiments, the determination module 701 is configured to perform, for any one of the at least one target data pairs, a step of determining at least one third probability based on second text features in any one of the target data pairs, where the at least one third probability is used to indicate a probability that the second text corresponding to any one of the target data pairs will be translated into the respective candidate text; a step of determining a fourth probability based on the at least one third probability, where the fourth probability is used to indicate a probability that the second text corresponding to any one of the target data pairs will be translated into a standard translation text in any one of the target data pairs; and a step of determining a confidence level for any one of the target data pairs based on the fourth probability.
[0169] In some embodiments, the determination module 701 is configured to perform the steps of determining a fifth probability based on the at least one first probability, where the fifth probability is used to indicate the probability that the first text will be translated into the standard translation text in any one of the target data pairs, and determining a confidence level for any one of the target data pairs based on the fourth probability and the fifth probability.
[0170] In some embodiments, the determination module 701 is configured to perform the steps of: standardizing the match degree of the first data pair for each of the standard translation texts to obtain a standardized match degree, where the first data pair is a data pair including the standard translation text of any one of the at least one target data pair; correcting the standardized match degree using the confidence of the first data pair to obtain a corrected match degree; and determining a second probability corresponding to the one of the standard translation texts based on the corrected match degree, where the corrected match degree is positively correlated with the second probability.
[0171] In some embodiments, the determination module 701 is configured to perform the steps of determining a hyperparameter based on at least one information of a quantitative indicator of each target data pair and a match degree of each target data pair among at least one target data pair, where the quantitative indicator of any one target data pair is the number of target data pairs that are not ranked lower than any one target data pair after each target data pair is sorted according to a reference order, and determining the ratio of the match degree of the first data pair and the hyperparameter as the standardized match degree.
[0172] In some embodiments, the determination module 701 is configured to perform the steps of determining a first probability distribution based on at least one first probability, determining a second probability distribution based on at least one second probability, fusing the first probability distribution and the second probability distribution to obtain a fused probability distribution, where the fused probability distribution includes a translation probability of each target text, where each target text includes each candidate text and each standard translation text, and determining the target text with the highest translation probability among each target text as the translation text.
[0173] In some embodiments, the determination module 701 is configured to perform the steps of determining a first importance and a second importance, where the first importance is used to indicate the importance of a first probability distribution in obtaining the translation text and the second importance is used to indicate the importance of a second probability distribution in obtaining the translation text; determining a target parameter based on the first importance and the second importance; transforming the first importance based on the target parameter to obtain a first weight; transforming the second importance based on the target parameter to obtain a second weight; and fusing the first probability distribution and the second probability distribution based on the first weight and the second weight to obtain a fused probability distribution.
[0174] In some embodiments, a text translation method is implemented by a target text translation model for translating text in a first language into text in a second language.
[0175] According to the technical method of the embodiment of the present application, in the process of determining the second probability, in addition to the match degree between the second text feature and the first text feature in the target data pair, the reliability of the target data pair is also taken into consideration, so that a wealth of information is taken into consideration. In addition, the reliability of the target data pair is used to evaluate the degree of reliability of the target data pair, so that by taking the reliability of the target data pair into consideration, the reliability of the second probability can be increased, and the accuracy of the text translation can be further improved.
[0176] Referring to FIG. 8, an embodiment of the present application provides an apparatus for obtaining a text translation model, the apparatus comprising: an acquisition module 801 configured to perform the steps of acquiring a first sample text, a first standard translation text, and an initial text translation model, where the first sample text is a text in a first language and the first standard translation text is a text obtained by translating the first sample text into a second language; a determining module 802 configured to perform a step of processing first sample text features with an initial text translation model to obtain at least one first sample probability, the first sample text features being text features of the first sample text, the at least one first sample probability being used to indicate a probability that the first sample text is translated into each of at least one candidate text, each of the at least one candidate text being a text in a second language; an acquisition module 801 configured to further perform the steps of acquiring at least one sample data pair matching the first sample text features, where any one sample data pair includes a second sample text feature and a second standard translated text, the second sample text feature being a text feature of the second sample text, the second sample text being a text in a first language, and the second standard translated text being a text of the second sample text translated into a second language; a determination module 802 configured to further perform a step of determining a confidence and a match of at least one sample data pair, where the confidence of any one sample data pair is used to indicate a degree of reliability of any one sample data pair, and the match of any one sample data pair is used to indicate a similarity between the second sample text feature and the first sample text feature in any one sample data pair; a determination module 802 configured to further perform a step of determining at least one second sample probability based on the confidence and match of the at least one sample data pair, the at least one second sample probability being used to indicate a probability that the first sample text is translated into each second standard translation text in the at least one sample data pair; a determination module 802 configured to further perform the step of determining a predictively translated text corresponding to the first sample text based on the at least one first sample probability and the at least one second sample probability; and an update module 803 configured to perform the steps of updating the initial text translation model based on the difference between the predicted translated text and the first standard translated text to obtain a target text translation model.
[0177] In some embodiments, the acquisition module 801 includes: searching a data pair library for at least one initial data pair matching the first sample text features, any one of the initial data pairs including a third sample text feature and a second standard translated text, the third sample text feature being a text feature of the second sample text, the second sample text being a text in the first language, and the second standard translated text being a text of the second sample text translated into the second language; coherently processing at least one initial data pair according to an interference probability to obtain a coherently processed data pair; and determining at least one sample data pair based on the interference processed data pairs.
[0178] In some embodiments, the interference probability is determined according to the number of updates of the initial text translation model, and the interference probability is negatively correlated with the number of updates of the initial text translation model.
[0179] In some embodiments, when the interference probability includes a first interference probability to indicate the execution probability of adding noise as the interference method, the acquisition module 801 is configured to perform a step of adding a noise feature to the third sample text feature in each initial data pair according to the first interference probability to obtain an interference-processed data pair, and a step of making the interference-processed data pair into at least one sample data pair.
[0180] In some embodiments, when the interference probability includes a second interference probability for indicating the execution probability of deleting the initial data pairs as the interference method, the acquisition module 801 is configured to perform the steps of deleting at least one initial data pair that does not satisfy the match condition according to the second interference probability to obtain an interference-processed data pair; constructing reference data pairs based on the first sample text feature and the first standard translation text, wherein the number of reference data pairs is the same as the number of the deleted initial data pairs; and determining at least one sample data pair based on the interference-processed data pair and the reference data pair.
[0181] In the technical method according to the embodiment of the present application, in addition to the match degree between the second sample text feature and the first sample text feature in the sample data pair, the reliability of the sample data pair is also taken into consideration in the process of determining the second sample probability, so that the information taken into consideration is rich. In addition, the reliability of the sample data pair is used to evaluate the reliability of the sample data pair, so that the reliability of the second sample probability can be improved by taking the reliability of the sample data pair into consideration, which can further improve the accuracy of the pre-translated text, improve the efficiency of model acquisition and the reliability of the acquired model, and further improve the accuracy of text translation using the model.
[0182] It should be noted that, in the above-mentioned embodiment of the device, when the function is realized, only the division of each of the above-mentioned functional modules is taken as an example, but in practical application, the above-mentioned functions can be performed by different functional modules as necessary, that is, the internal structure of the device can be divided into different functional modules to complete all or part of the above-mentioned functions. In addition, the embodiment of the device and the embodiment of the method according to the above-mentioned embodiment belong to the same concept, and the specific implementation process can be referred to the embodiment of the method, which will not be described further here.
[0183] In some embodiments, a computer device including a processor and a memory is further provided. At least one computer program is stored in the memory. The at least one computer program is loaded and executed by one or more processors to cause the computer device to realize the text translation method or the method for obtaining a text translation model described in any one of the above. The computer device may be a server or a terminal, but is not limited to this in the embodiments of the present application. Next, the configurations of the server and the terminal will be described respectively.
[0184] FIG. 9 is a schematic diagram of a server according to an embodiment of the present application. The server may have relatively large differences due to differences in configuration or performance, and may include one or more processors (Central Processing Units, CPU) 901 and one or more memories 902. Here, the one or more memories 902 store at least one computer program. The at least one computer program is loaded and executed by the one or more processors 901, thereby enabling the server to realize the text translation method or the text translation model acquisition method according to each of the above method embodiments. Of course, the server may have components such as a wired or wireless network interface, a keyboard, and an input / output interface for convenient input / output. The server may also include other components for realizing device functions, which will not be described further here.
[0185] 10 is a schematic diagram of a terminal according to an embodiment of the present application. The terminal may be a PC, a mobile phone, a smartphone, a PDA, a wearable device, a PPC, a tablet, a smart sensor device, a smart TV, a smart speaker, a smart voice dialogue device, a smart home appliance, an in-vehicle terminal, a VR device, or an AR device. The terminal may be called a user device, a mobile terminal, a laptop terminal, a desktop terminal, or the like.
[0186] Typically, the terminal includes a processor 1501 and a memory 1502 .
[0187] The processor 1501 may include one or more processing cores, such as a 4-core processor, an 8-core processor, etc. The processor 1501 may be implemented using at least one of the following hardware formats: DSP (Digital Signal Processing), FPGA (Field-Programmable Gate Array), and PLA (Programmable Logic Array). The processor 1501 may include a main processor and a coprocessor. The main processor is a processor for processing data in an active state, and is also called a CPU (Central Processing Unit). The coprocessor is a low-power processor for processing data in a standby state. In some embodiments, the processor 1501 may integrate a GPU (Graphics Processing Unit) for rendering and painting display content required for the display. In some embodiments, the processor 1501 may further include an AI (Artificial Intelligence) processor for processing computing operations related to machine learning.
[0188] The memory 1502 may include one or more computer readable storage media, which may be non-transitory. The memory 1502 may include high speed random access memory and non-volatile memory, such as one or more disk storage devices or flash memory storage devices. In some embodiments, the non-transitory computer readable storage medium in the memory 1502 is used to store at least one command executed by the processor 1501 to realize the text translation method or the method of obtaining a text translation model according to the embodiment of the method in the present application.
[0189] In some embodiments, the terminal optionally includes a peripheral device interface 1503 and at least one peripheral device. The processor 1501, the memory 1502, and the peripheral device interface 1503 may be connected via a bus or signal lines. Each peripheral device may be connected to the peripheral device interface 1503 via a bus, signal lines, or a circuit board. Specifically, the peripheral device includes at least one of a display 1505 and a power supply 1508.
[0190] The peripheral device interface 1503 may be used to connect at least one peripheral device associated with I / O (Input / Output) to the processor 1501 and the memory 1502. In some embodiments, the processor 1501, the memory 1502, and the peripheral device interface 1503 are integrated on the same chip or circuit board. In some other embodiments, any one or two of the processor 1501, the memory 1502, and the peripheral device interface 1503 may be implemented on separate chips or circuit boards, but this embodiment is not limited thereto.
[0191] The display 1505 is used to display a UI (User Interface). This UI can include graphics, text, icons, videos, and any combination thereof. If the display 1505 is a touch display, the display 1505 also has the ability to collect touch signals on or above the surface of the display 1505. The touch signals can be input as control signals to the processor 1501 for processing. At this time, the display 1505 can also be used to provide virtual buttons and / or virtual keyboards, also called soft buttons and / or soft keyboards. In some embodiments, the display 1505 can be one provided on the front panel of the terminal. In another embodiment, the display 1505 can be at least two separately arranged on different surfaces of the terminal or a folded design. In another embodiment, the display 1505 can be a flexible display arranged on a curved or folded surface of the terminal. Furthermore, the display 1505 can also be configured as a non-rectangular irregular shape, i.e., an irregular screen. The display 1505 may be made of materials such as LCD (Liquid Crystal Display), OLED (Organic Light-Emitting Diode), and the like.
[0192] The power source 1508 is used to provide power to each component in the terminal. The power source 1508 may be AC power, DC power, a disposable battery, or a rechargeable battery. If the power source 1508 includes a rechargeable battery, the rechargeable battery may support wired or wireless charging. The rechargeable battery may also be used to support fast charging technology.
[0193] It should be understood by those skilled in the art that the structure shown in FIG. 10 does not constitute a limitation on the terminal, which may include more or fewer components than shown, or may combine some components or employ different component arrangements.
[0194] In some embodiments, there is further provided a computer readable storage medium having stored thereon at least one computer program, which when loaded and executed by a processor of a computing device, causes the computer to perform the method for translating text or the method for obtaining a text translation model according to any one of the above claims.
[0195] In some embodiments, the aforementioned computer readable storage medium may be a Read-Only Memory (ROM), a Random Access Memory (RAM), a Compact Disc Read-Only Memory (CD-ROM), a magnetic tape, a floppy disk, an optical data storage device, and the like.
[0196] In some embodiments, there is further provided a computer program product comprising a computer program or computer commands that, when loaded and executed by a processor, cause the computer to perform any of the methods for translating text or for obtaining a text translation model described above.
[0197] The above are only exemplary embodiments of the present application, and do not limit the present application. Within the scope of the principles of the present application, any modifications, equivalent replacements, improvements, etc. should be included in the protection scope of the present application.
Claims
1. 1. A method for text translation applied to a computing device, comprising: determining at least one first probability based on first text features, the first text features being text features of a first text, the first text being a text in a first language, the at least one first probability being used to indicate a probability that the first text will be translated into each of at least one candidate text, each of the at least one candidate text being a text in a second language; obtaining at least one target data pair matching the first text feature, where any one target data pair includes a second text feature and a standard translation of the second text, the second text feature being a text feature of the second text, the second text being a text in the first language, and the standard translation being a text in the second language; determining a confidence level and a match level of the at least one target data pair, the confidence level of any one target data pair being used to indicate a degree of reliability of the any one target data pair, and the match level of any one target data pair being used to indicate a degree of similarity between a second text feature and the first text feature in the any one target data pair; determining at least one second probability based on the confidence and match of the at least one target data pair, the at least one second probability being used to indicate a probability that the first text will be translated into each standard translation text in the at least one target data pair; determining a translated text corresponding to the first text based on the at least one first probability and the at least one second probability.
2. The step of determining a confidence level of the at least one target data pair comprises: determining at least one third probability for each of the at least one target data pairs based on second text features in the target data pair, the at least one third probability being used to indicate a probability that a second text corresponding to the target data pair is translated into each of the candidate texts; determining a fourth probability based on the at least one third probability, the fourth probability being used to indicate a probability that a second text corresponding to the one of the target data pairs will be translated into a standard translation text for the one of the target data pairs; and determining a confidence level of the one target data pair based on the fourth probability.
3. determining a confidence level of the one target data pair based on the fourth probability, determining a fifth probability based on the at least one first probability, the fifth probability being used to indicate a probability that the first text will be translated into a standard translation text in any one of the target data pairs; and determining a confidence level of the one target data pair based on the fourth probability and the fifth probability.
4. The step of determining at least one second probability based on the confidence and match of the at least one target data pair comprises: A step of standardizing a match degree of a first data pair for any one of the standard translation texts to obtain a standardized match degree, the first data pair being a data pair including the one of the at least one target data pairs; correcting the standardized match measure using a confidence measure of the first data pair to obtain a corrected match measure; 2. The method of claim 1, further comprising: determining a second probability corresponding to the one of the standard translation texts based on the corrected match degree, the corrected match degree exhibiting a positive correlation with the second probability.
5. said step of standardizing the match measure of the first data pair to obtain a standardized match measure comprises: A step of determining a hyperparameter based on at least one information of a quantitative indicator of each of the at least one target data pairs and a degree of match of each of the target data pairs, wherein the quantitative indicator of any one of the target data pairs is the number of target data pairs that are not positioned lower than any one of the target data pairs after the target data pairs are sorted according to a reference order; The text translation method according to claim 4 , further comprising a step of: determining a ratio of the match degree of the first data pair and the hyperparameter as the standardized match degree.
6. The step of determining a translation text corresponding to the first text based on the at least one first probability and the at least one second probability comprises: determining a first probability distribution based on the at least one first probability; determining a second probability distribution based on the at least one second probability; blending the first probability distribution and the second probability distribution to obtain a blended probability distribution, the blended probability distribution including translation probabilities for each target text, the each target text including the each candidate text and the each standard translation text; The text translation method according to claim 1 , further comprising the step of: determining, as the translation text, the target text having the highest translation probability among the target texts.
7. The step of fusing the first probability distribution and the second probability distribution to obtain a fused probability distribution includes: determining a first importance and a second importance, the first importance being used to indicate the importance of the first probability distribution in obtaining the translated text, and the second importance being used to indicate the importance of the second probability distribution in obtaining the translated text; determining a target parameter based on the first importance and the second importance; converting the first importance based on the target parameter to obtain a first weight; converting the second importance based on the target parameter to obtain a second weight; and fusing the first probability distribution and the second probability distribution based on the first weight and the second weight to obtain a fusing probability distribution.
8. The method of claim 1 , wherein the method is implemented by a target text translation model for translating text in the first language into text in the second language.
9. 1. A method for obtaining a text translation model for application to a computing device, comprising: obtaining a first sample text, a first standard translation text, and an initial text translation model, the first sample text being a text in a first language and the first standard translation text being a translation of the first sample text into a second language; processing first sample text features with the initial text translation model to obtain at least one first sample probability, the first sample text features being text features of the first sample text, the at least one first sample probability being used to indicate a probability that the first sample text will be translated into each of at least one candidate text, each of the at least one candidate text being a text in the second language; obtaining at least one sample data pair matching the first sample text features, each sample data pair including a second sample text feature and a second standard translation text, the second sample text feature being a text feature of the second sample text, the second sample text being a text in the first language, and the second standard translation text being a text translated from the second sample text into a second language; determining a confidence level and a match level of the at least one sample data pair, the confidence level of any one sample data pair being used to indicate a degree of reliability of the any one sample data pair, and the match level of any one sample data pair being used to indicate a degree of similarity between a second sample text feature and the first sample text feature in the any one sample data pair; determining at least one second sample probability based on the confidence and match of the at least one sample data pair, the at least one second sample probability being used to indicate a probability that the first sample text is translated into each second standard translation text in the at least one sample data pair; determining a predicted translated text corresponding to the first sample text based on the at least one first sample probability and the at least one second sample probability; updating the initial text translation model based on a difference between the predicted translated text and the first standard translated text to obtain a target text translation model.
10. The step of obtaining at least one sample data pair matching the first sample text feature comprises: searching a data pair library for at least one initial data pair matching the first sample text feature, any one of the initial data pairs including a third sample text feature and a second standard translated text, the third sample text feature being a text feature of the second sample text, the second sample text being a text in the first language, and the second standard translated text being a text translated from the second sample text into a second language; coherently processing the at least one initial data pair according to an interference probability to obtain a coherently processed data pair; and determining the at least one sample data pair based on the interference processed data pairs.
11. The method of claim 10 , wherein the interference probability is determined according to a number of updates of the initial text translation model, and the interference probability is negatively correlated with a number of updates of the initial text translation model.
12. The interference probability includes a first interference probability indicating a probability of performing the interference method of adding noise; The step of coherently processing the at least one initial data pair according to an interference probability to obtain a coherently processed data pair comprises: adding a noise feature to a third sample text feature in each initial data pair according to the first interference probability to obtain the interference processed data pair; The step of determining the at least one sample data pair based on the interference processed data pair comprises: The method for obtaining a text translation model according to claim 10, further comprising the step of: taking the interference processed data pair as the at least one sample data pair.
13. The interference probability includes a second interference probability for indicating a probability of performing the deletion of the initial data pair as an interference method; The step of coherently processing the at least one initial data pair according to an interference probability to obtain a coherently processed data pair comprises: removing an initial data pair that does not satisfy a match condition from among the at least one initial data pair according to the second interference probability to obtain an interference-processed data pair; The step of determining the at least one sample data pair based on the interference processed data pair comprises: constructing reference data pairs based on the first sample text features and the first standard translation text, the number of the reference data pairs being equal to the number of the deleted initial data pairs; and determining the at least one sample data pair based on the interference processed data pairs and the reference data pairs.
14. The interference probability includes a first interference probability indicating a probability of performing an interference method of adding noise, and a second interference probability indicating a probability of performing an interference method of deleting an initial data pair; The step of coherently processing the at least one initial data pair according to an interference probability to obtain a coherently processed data pair comprises: removing an initial data pair that does not satisfy a match condition from among the at least one initial data pair according to the second interference probability to obtain an intermediate data pair; adding noise features to third sample text features in the intermediate data pairs according to the first interference probability to obtain the interference processed data pairs; The step of determining the at least one sample data pair based on the interference processed data pair comprises: constructing reference data pairs based on the first sample text features and the first standard translation text, the number of the reference data pairs being equal to the number of the deleted initial data pairs; and determining the at least one sample data pair based on the interference processed data pairs and the reference data pairs.
15. A text translation apparatus disposed on a computing device, comprising: a determination module configured to perform a step of determining at least one first probability based on first text features, the first text features being text features of a first text, the first text being a text in a first language, the at least one first probability being used to indicate a probability that the first text will be translated into each of at least one candidate text, each of the at least one candidate text being a text in a second language; and an acquisition module configured to perform the steps of acquiring at least one target data pair matching the first text feature, where any one target data pair includes a second text feature and a standard translation of the second text, the second text feature being a text feature of the second text, the second text being a text in the first language, and the standard translation being a text in the second language; the determination module is further configured to perform a step of determining a confidence and a match of the at least one target data pair, the confidence of any one target data pair being used to indicate a degree of reliability of the any one target data pair, and the match of any one target data pair being used to indicate a similarity between a second text feature and the first text feature in the any one target data pair; The determination module is further configured to perform a step of determining at least one second probability based on the confidence and match of the at least one target data pair, the at least one second probability being used to indicate a probability that the first text is translated into each standard translation text in the at least one target data pair; The text translation apparatus, wherein the determination module is further configured to perform the step of determining a translated text corresponding to the first text based on the at least one first probability and the at least one second probability.
16. 1. An apparatus for obtaining a text translation model located on a computing device, comprising: an acquisition module configured to perform the steps of acquiring a first sample text, a first standard translation text, and an initial text translation model, the first sample text being a text in a first language, and the first standard translation text being a text obtained by translating the first sample text into a second language; a decision module configured to process first sample text features with the initial text translation model to obtain at least one first sample probability, the first sample text features being text features of the first sample text, the at least one first sample probability being used to indicate a probability that the first sample text will be translated into each of at least one candidate text, each of the at least one candidate text being a text in the second language; the acquiring module is further configured to acquire at least one sample data pair matching the first sample text features, each sample data pair including a second sample text feature and a second standard translated text, the second sample text feature being a text feature of the second sample text, the second sample text being text in the first language, and the second standard translated text being text translated from the second sample text into a second language; the determination module is further configured to perform a step of determining a confidence and a match of the at least one sample data pair, the confidence of any one sample data pair being used to indicate a degree of reliability of the any one sample data pair, and the match of any one sample data pair being used to indicate a degree of similarity between a second sample text feature and the first sample text feature in the any one sample data pair; The determination module is further configured to perform the steps of determining at least one second sample probability based on the confidence and match of the at least one sample data pair, the at least one second sample probability being used to indicate a probability that the first sample text is translated into each second standard translation text in the at least one sample data pair; the determination module is further configured to determine a predictively translated text corresponding to the first sample text based on the at least one first sample probability and the at least one second sample probability; The apparatus for obtaining a text translation model further includes an update module configured to perform a step of updating the initial text translation model based on a difference between the predicted translated text and the first standard translated text to obtain a target text translation model.
17. A computing device including a processor and a memory, The memory stores at least one computer program, and the at least one computer program is loaded and executed by the processor to cause the computer device to realize the text translation method according to any one of claims 1 to 8 or the method for obtaining a text translation model according to any one of claims 9 to 14.
18. A computer program that, when loaded and executed by a processor, causes the computer to implement the text translation method according to any one of claims 1 to 8 or the method for obtaining a text translation model according to any one of claims 9 to 14.
Citation Information
Patent Citations
Information processing method and device and electronic equipment
CN113591490A
Machine translation method, device, computer device, and storage medium
JP2020535564A