Bilingual Corpus Detection Method, Device, and Computer-Readable Medium

Through the bilingual corpus detection method combined with attention mechanism, the problems of poor accuracy and high cost in the prior art are solved, and the accurate detection of translation errors in the bilingual corpus is achieved, which reduces the misjudgment rate and reduces labor costs.

CN114065777BActive Publication Date: 2025-05-30ALIBABA GROUP HOLDING LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202010762257.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2020-07-31
Publication Date
2025-05-30
Estimated Expiration
2040-07-31

AI Technical Summary

Technical Problem

The existing bilingual corpus detection methods have poor accuracy and high cost, especially because they cannot effectively consider the semantics and word order of words, resulting in a high misjudgment rate and the need for manual labeling of data, which increases labor costs.

Method used

A bilingual corpus detection method is adopted. By obtaining the word information of the first corpus and the target words of the second corpus and their previous and subsequent information, combined with the attention mechanism, the target words are predicted to determine whether they are incorrect.

Benefits of technology

This method can provide semantic and word order support, reduce misjudgment, and eliminate the need for manual annotation of data. It only requires using the translation accurate bilingual corpus as the training set, which can accurately detect translation errors in the bilingual corpus.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114065777B_ABST
    Figure CN114065777B_ABST
Patent Text Reader

Abstract

The present application provides a bilingual corpus detection solution. This solution can take the preceding information, following information of the target word in the second corpus sentence, and the word information of the first corpus as inputs, and combines an attention mechanism. Thus, it can provide semantic and word order support for the target word to be predicted. Even if a word in one corpus does not have a strictly corresponding word in the other corpus or corresponds to multiple words, it is not easy to produce incorrect detection results. At the same time, without relying on manually annotated corpora, only bilingual corpora with relatively accurate translations need to be used as the training set to complete the training of a neural network model adopting the attention mechanism, and then it can accurately detect whether there are words with translation errors in any bilingual corpus.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of information technology, and in particular to a bilingual corpus detection method, device and computer-readable medium. Background Art

[0002] Bilingual corpus, also known as "bilingual parallel sentence pairs", is a kind of text corpus that can be translated into each other. Taking Chinese and English as an example, "It's a nice day today" and "It's a nice day today" are a pair of bilingual corpus. Bilingual corpus is the key training data for machine translation models. Both statistical machine translation (SMT) and neural machine translation (NMT) rely on this type of corpus. In machine translation, the support of multiple languages ​​and the quality of translation in each language direction are closely related to the scale and quality of bilingual corpus.

[0003] There are currently two main ways to detect whether words in bilingual corpora are translated accurately. One way is to build a word alignment model, map the original words and translated words of the bilingual corpora, and then count the mapping results. If a word on the original side cannot be matched with any word on the translated side, then the original word can be considered as missed translation or mistranslation; if a word on the translated side cannot be mapped with any word on the original side, then the word is multiple translation or mistranslation. However, this word alignment method does not consider the semantics and word order of words in the sentence, and it is easy to misjudge some words. For example, for the bilingual corpus "it's fine today" and "the weather is good today", since the English corpus does not contain the word corresponding to "weather", the recognition of "weather" in the Chinese corpus may be judged as a multiple translation error. Moreover, there will be one-to-many problems that cannot be solved in many cases. For example, in the bilingual corpus of "target text is very good" and "the translation is very good", the two English words "target" and "text" actually correspond to a Chinese word "translation", but the word alignment method is to map each word, and it is not easy to match "target text" and "translation", which will lead to detection errors.

[0004] Another method is to manually annotate the wrong words in the bilingual corpus. The error types can be mistranslation, missing translation, multiple translation, etc. Then, the recognition model is trained based on the manually annotated data, and the recognition model is used to identify word errors in the bilingual corpus to be tested. However, the main problem with this method is that it requires manual annotation of data, which has a high labor cost and cannot be applied on a large scale. Summary of the invention

[0005] An object of the present application is to provide a bilingual corpus detection method, device, and computer-readable medium to solve the problems of poor accuracy and high cost in existing detection methods.

[0006] In an embodiment of the present application, a bilingual corpus detection method is provided, and the method includes:

[0007] Obtain the word information of the first corpus, the target word of the second corpus, and its previous and subsequent information;

[0008] Use the word information of the first corpus in combination with the attention mechanism, and the previous information and subsequent information of the target word to predict the target word;

[0009] Determine whether the target word is incorrect according to the prediction result.

[0010] In an embodiment of the present application, a bilingual corpus detection device is further provided, and the device includes:

[0011] A prediction processing module, configured to obtain the word information of the first corpus, the target word of the second corpus, and its previous and subsequent information; use the word information of the first corpus in combination with the attention mechanism, and the previous information and subsequent information of the target word to predict the target word;

[0012] A detection processing module, configured to determine whether the target word is incorrect according to the prediction result.

[0013] Some embodiments of the present application further provide a computing device, where the device includes a memory for storing computer program instructions and a processor for executing the computer program instructions. When the computer program instructions are executed by the processor, the device is triggered to execute the foregoing bilingual corpus detection method.

[0014] Some other embodiments of the present application further provide a computer-readable medium, on which computer program instructions are stored, and the computer-readable instructions can be executed by a processor to implement the bilingual corpus detection method.

[0015] A bilingual corpus detection solution provided by an embodiment of this application first obtains the word information of the first corpus, the target word of the second corpus, and its preceding and following context information, and then uses the word information of the first corpus in combination with the attention mechanism, as well as the preceding context information and the following context information of the target word, to predict the target word. According to the prediction result, it is determined whether the target word is incorrect. Since the preceding context information, the following context information of the target word in the second corpus sentence, and the word information of the first corpus are used as inputs and combined with the attention mechanism, semantic and word order support can be provided for the target word to be predicted. Even if there is no strictly corresponding word in one corpus for a word in another corpus or there are multiple corresponding words, it is not easy to produce incorrect detection results. At the same time, without relying on manually annotated corpora, only bilingual corpora with relatively accurate translations need to be used as the training set to complete the training of a neural network model using the attention mechanism, and it is possible to accurately detect whether any bilingual corpus contains mis-translated words. BRIEF DESCRIPTION OF THE DRAWINGS

[0016] Other features, objects, and advantages of this application will become more apparent by reading the detailed description of the non-limiting embodiments with reference to the following drawings:

[0017] Figure 1 It is a processing flow chart of a bilingual corpus detection method provided by an embodiment of this application;

[0018] Figure 2 It is a schematic diagram of a bilingual corpus and its corresponding feature vectors in an embodiment of this application;

[0019] Figure 3 It is a schematic diagram of the basic principle of obtaining the prediction result of the target word for the target word in an embodiment of this application;

[0020] Figure 4 It is a schematic diagram of the structure of a bilingual corpus detection device provided by an embodiment of this application;

[0021] Figure 5 It is a schematic diagram of the structure of a computing device for implementing bilingual corpus detection provided by an embodiment of this application;

[0022] The same or similar reference numerals in the drawings represent the same or similar components. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0023] The embodiments of the present invention will be described in detail below. The examples of the embodiments are shown in the drawings, where the same or similar reference numerals represent the same or similar elements or elements with the same or similar functions from beginning to end. The embodiments described below with reference to the drawings are exemplary and are only used to explain the present invention and should not be construed as limiting the present invention.

[0024] Those skilled in the art can understand that, unless specifically stated otherwise, the singular forms "a", "an", "the" and "said" used herein may also include the plural forms. It should be further understood that the term "comprising" used in the specification of the present invention means the presence of the stated features, integers, steps, operations, elements and / or components, but does not exclude the presence or addition of one or more other features, integers, steps, operations, elements, components and / or their groups. It should be understood that when we say that an element is "connected" or "coupled" to another element, it can be directly connected or coupled to other elements, or there may also be intermediate elements. In addition, the "connection" or "coupling" used herein may include wireless connection or wireless coupling. The phrase "and / or" used herein includes all or any unit and all combinations of one or more of the associated listed items.

[0025] A bilingual corpus detection method provided by an embodiment of the present application takes the previous information, subsequent information of a target word in a second corpus sentence, and word information of a first corpus as inputs, and combines an attention mechanism. Thus, semantic and word order support can be provided for the target word to be predicted. Even if a word in one corpus does not have a strictly corresponding word in another corpus or corresponds to multiple words, it is not easy to produce incorrect detection results. At the same time, without relying on manually annotated corpora, only bilingual corpora with relatively accurate translations need to be used as the training set to complete the training of a neural network model adopting an attention mechanism, and then it is possible to accurately detect whether there are words with translation errors in any bilingual corpus.

[0026] In an actual scenario, the execution subject of this method can be a user device, a network device, or a device formed by integrating a user device and a network device through a network. In addition, it can also be a program running on the above devices. The user device includes but is not limited to various terminal devices such as computers, mobile phones, and tablet computers; the network device includes but is not limited to being implemented such as a network host, a single network server, a set of multiple network servers, or a computer set based on cloud computing. Here, the cloud consists of a large number of hosts or network servers based on cloud computing (Cloud Computing), where cloud computing is a type of distributed computing and consists of a virtual computer formed by a group of loosely coupled computer sets.

[0027] Figure 1 The processing flow of a bilingual corpus detection method provided by an embodiment of the present application is shown, and at least the following processing steps can be included:

[0028] Step S101, obtain the word information of the first corpus, the target word of the second corpus, and its previous information and subsequent information.

[0029] Among them, the bilingual corpus processed in the embodiments of the present application may include a first corpus and a second corpus. The first corpus and the second corpus correspond to one language respectively, and they are translations of each other. For example, "The weather is nice today" and "It's a nice day today" are bilingual corpora in Chinese and English, and "I agree with that" and "Ich stimme dem zu" are bilingual corpora in English and German. In actual scenarios, some content in the bilingual corpus may be mistranslated due to various factors. For example, for the bilingual corpus "I especially like to eat red apples" and "I like eating green apple and banana", the word "especially" in the Chinese corpus is a missing word, the words "and" and "banana" in the English corpus are extra words, and the words "red" and "green" are mistranslated words. The purpose of the detection method in the embodiments of the present application is to detect these mistranslated words in the bilingual corpus.

[0030] The word information of the first corpus, the target word of the second corpus, and its previous and subsequent information are all information used to represent the text content contained in the corpus. Taking these information as input can provide semantic and word order support for the target word to be predicted, making the prediction result more accurate.

[0031] In some embodiments of the present application, the word information of the first corpus, the target word of the second corpus, and its previous and subsequent information can all be represented in the form of a word vector sequence. Among them, the previous information can be the word vector sequence of the first N words before the target word in the second corpus, where N can be set according to the requirements of the actual scenario. For example, it can be set to a specific value such as 3, 4, 6, etc., or it can be dynamically adjusted according to the position of the target word in the second corpus. For example, it can be the position serial number of the target word minus 1. At this time, when the target word is the 4th word in the second corpus, the word vectors of the first 3 words (i.e., all the words before the target word) can be obtained as the previous information. For the corpus "I like eating green apple and banana", if N is set to 3, when the target word is "apple", its corresponding previous information is the word vectors of "like", "eating", and "green".

[0032] The post-context information is the word vector sequence of the last M words of the target word in the second corpus. Similar to N in the pre-context information, M can also be set according to the needs of the actual scene, for example, it can be set to a specific value such as 2, 3, 6, etc., or it can be dynamically adjusted according to the position of the target word in the second corpus. For example, it can be the total number of words in the second corpus minus the position number of the target word. At this time, when the target word is the 4th word in the second corpus and the total number of words in the second corpus minus 8, the word vectors of the last 4 words (i.e., all words after the target word) can be obtained as post-context information. For the corpus "I like eating green apple and banana", if M is set to 3, when the target word is green, the corresponding post-context information is the word vector of "apple", "apple", and "banana".

[0033] Therefore, when obtaining the target word of the second corpus and its preceding and following context information, the target word of the second corpus and the first N words and the last M words of the target word in the second corpus can be obtained, and then the word vector sequence of the first N words and the word vector sequence of the last M words can be obtained.

[0034] The word information of the first corpus may be a word vector sequence of words contained in the first corpus, so when obtaining the word information of the first corpus, multiple words of the first corpus may be obtained, and then the word vector sequence of the multiple words may be obtained.

[0035] In a process of processing bilingual corpora, one corpus can be set as the first corpus, and the other corpus can be set as the second corpus. Taking the bilingual corpora "I like eating red apples" and "I like eating green apple and banana" as examples, in a process of processing, the Chinese corpus "I like eating red apples" can be set as the first corpus, and the English corpus "I like eating green apple and banana" can be set as the second corpus. Thus, the first corpus information is the feature vector sequence of "I", "especially", "like", "eat", "red", and "apple". After completing a process, the corpus can be swapped, that is, the Chinese corpus "I like eating red apples" can be set as the second corpus, and the English corpus "I like eating green apple and banana" can be set as the first corpus, and processed again.

[0036] In some implementations of the present application, when obtaining the word vector sequence of the multiple words, the multiple words can be first segmented to obtain the word sequence, and then the word sequence can be embedded to generate the word vector sequence of the multiple words.

[0037] Among them, the word vector sequence represents the word vectors of a word sequence in any corpus, and the word sequence represents words arranged in a certain order. In this embodiment, word segmentation processing can be performed on the multiple words to obtain the corresponding word sequence. For example, for the multiple words included in the aforementioned bilingual corpus, they can be considered as sentences before word segmentation. The sentences composed of the words in the two corpora are "The weather is good today" and "It's a nice day today". After performing word segmentation processing on the Chinese corpus, the following word sequence can be obtained: "today / weather / good", and after performing word segmentation processing on the English corpus, the following word sequence can be obtained: "It's / a / nice / day / today".

[0038] The word segmentation algorithm adopted in the embodiments of the present application can be a dictionary-based word segmentation algorithm, such as the forward maximum matching algorithm, the backward maximum matching algorithm, the bidirectional maximum matching algorithm, etc.; it can also be a statistics-based word segmentation algorithm, such as the N-gram algorithm; in addition, it can also be a word annotation-based word segmentation algorithm, such as the word segmentation algorithm based on BPE (byte pair encoding). In an actual scenario, since various languages have different grammar habits, corresponding word segmentation algorithms can be adopted for different languages to make the word segmentation results more accurate.

[0039] In some embodiments of the present application, the first corpus can be cleaned before performing word segmentation processing on the multiple words. After cleaning the corpus, some non-standard or content that affects subsequent processing can be adjusted to make subsequent processing more efficient and accurate. For example, in this embodiment, the cleaning of the corpus can be to unify the case, remove punctuation marks, add start symbols, end symbols, etc.

[0040] The word vector is used to uniquely identify a specific word, and different words have their own different word vectors. For example, the specific form of the word vector can adopt one-hot encoding, converting the word in text form into the word in encoded form. Assuming that there are a total of n words in the dictionary, an n-dimensional one-hot encoding can be created. Taking the aforementioned Chinese corpus "today / weather / good" as an example, the word "today" can be represented as: [1,0,0,...,0], "weather" can be represented as: [0,1,0,...,0], and "good" can be represented as: [0,0,1,...,0]. Thus, the word vector sequence of the Chinese corpus "The weather is good today" can be represented as a 3×n matrix, as shown in Table 1:

[0041] Dimension 1 Dimension 2 Dimension 3 Dimension n Today 1 0 0 …… 0 Weather 0 1 0 …… 0 Good 0 0 1 …… 0

[0042] Table 1

[0043] Because in actual scenarios, one-hot encoding has some problems when performing data processing, such as the data is relatively sparse, each word is represented as an orthogonal vector in the n-dimensional space, and there is no association between each other. This method cannot be used to calculate the similarity between words, resulting in the inability to perform efficient data processing. Therefore, after obtaining the word sequence of the bilingual corpus, the word sequence can be subjected to word embedding (Word embedding) processing to generate a word vector sequence of the bilingual corpus to be detected. In general, the word embedding can reduce the dimension of the n-dimensional vector in the dictionary to m-dimensional, and the association relationship between words can be represented by the spatial relationship between the vectors. For example, in an embodiment of the present application, after the Chinese corpus "The weather is good today" is word embedded, the following can be generated.

[0044] The word vector sequence shown in Table 2:

[0045]

[0046]

[0047] Table 2

[0048] Here, m is smaller than n, thereby reducing the dimension of data processing and improving processing efficiency.

[0049] Step S102: predict the target word by using the word information of the first corpus in combination with the attention mechanism, as well as the preceding context information of the target word and the following context information of the target word.

[0050] The prediction process can adopt a neural network model based on an encoder-decoder framework, which can include two parts of processing. First, the encoder (Encoder) based on the neural network model using the attention mechanism encodes to obtain a feature vector sequence, and then the corresponding decoder (Decoder) decodes the feature vector sequence to obtain the predicted probability of the target word.

[0051] The attention mechanism used in the embodiment of the present application may be a self-attention mechanism. After combining the word information of the first corpus with the attention mechanism, the corresponding feature vector sequence can be encoded. For example, the bilingual corpus and its corresponding feature vectors are as follows: Figure 2 As shown, the content of the first corpus includes the start symbol <s>, word A, word B, ……, word C, and terminator< / s> , where the start symbol <s>and terminator< / s>It is used to indicate the start and end of a sentence and can be regarded as a word during processing. After encoding it with Self-Attention, the corresponding feature vector sequence (H0, H1, H2, …, Hn) can be obtained, where Hi represents the feature vector corresponding to the i-th word.

[0052] If the content of the second corpus includes the start symbol <s>, word a, word b, ……, word c, and terminator< / s> , if the target word is word k. After encoding it in combination with Self-Attention, the previous information can correspond to the feature vector sequence (h0, h1, ..., hk-1), and the subsequent information can correspond to the feature vector sequence (hk+1, hk+2, ..., hn), where hi represents the feature vector corresponding to the i-th word.

[0053] Since the self-attention mechanism is adopted, when predicting the target word k, the relevant feature vector will be given a certain weight, which is related to the semantics of the word corresponding to the feature vector and the target word. Take the aforementioned bilingual corpus "I like eating red apples" and "I like eating green apple and banana" as an example. When the Chinese corpus is the first corpus and the English corpus is the second corpus, if the target word "apple" in the second corpus needs to be predicted, the feature vector corresponding to "apple" in the first corpus will have a higher weight, and when predicting other target words in the second corpus, the weight of the feature vector corresponding to "apple" will be reduced. In this way, the prediction probability of the target word can be made more accurate.

[0054] The feature vector sequence obtained by encoding with the attention mechanism is input into the decoder of the neural network model, and the predicted probability of the target word can be obtained by decoding. Figure 3 The basic principle of obtaining the predicted probability of a target word in an embodiment of the present application is shown. After determining the feature vector sequence (h0, h1, ..., hk-1) corresponding to the preceding information, the feature vector sequence (hk+1, hk+2, ..., hn) corresponding to the succeeding information, and the feature vector sequence (H0, H1, H2, ..., Hn) corresponding to the word information of the first corpus, this information can be used as input to a pre-trained neural network model, which can be decoded by a decoder to output the predicted probability P of the target word.

[0055] The neural network model for prediction can be pre-trained based on the bilingual corpus in the training set before it is needed. The bilingual corpus in the training set is of high quality. Here, high quality means that the meanings expressed by the first corpus and the corresponding second corpus are the same (i.e., the translation is correct), and there are no errors that need to be detected in the solution of this embodiment of the present application, such as over-translation, under-translation, wrong translation, etc. Thus, for the correct target word, the trained neural network model will output a relatively high prediction probability, while for the wrong target word, it will output a relatively low prediction probability.

[0056] Step S103, according to the prediction result, determine whether the target word is incorrect.

[0057] In an actual scenario, if the target word in the second corpus is incorrect (such as wrong translation or over-translation), after inputting the previous context information, the subsequent context information, and the first corpus information into the neural network model, a relatively low prediction probability will be output. If the target word is correct, a relatively high prediction probability will be output. Thus, when determining the incorrect target word, a probability threshold can be preset according to the actual scenario. After obtaining the prediction probability of the target word, compare the prediction probability of the target word with the preset probability threshold. If the prediction probability is lower than the preset probability threshold, determine the target word as an incorrect target word.

[0058] Thus, by traversing all the words in the second corpus, the detection of all the words in the second corpus can be completed, and all the words with translation errors in the second corpus can be identified. For example, if the Chinese corpus "I especially like to eat red apples" is set as the first corpus and the English corpus "I like eating green apple and banana" is used as the second corpus, and the words "I", "like", "eating", "green", "apple", "and", "banana" are used as target words for processing in turn to complete the traversal. If the detected prediction probabilities are as shown in Table 3:

[0059] Target word Prediction probability I 0.87 like 0.91 eating 0.77 green 0.11 apple 0.92 and 0.07 banana 0.05

[0060] Table 3

[0061] After comparing the predicted probabilities of the above target words with a preset probability threshold, it can be determined that "green", "and", and "banana" are below the probability threshold, from which it can be identified that "green", "and", and "banana" are incorrect words. If it is necessary to simultaneously identify the incorrect words in the Chinese corpus "I especially like to eat red apples", the first corpus and the second corpus can be swapped (set the original first corpus as the second corpus and the original second corpus as the first corpus), and then traversed again, so as to identify that "especially" and "red" are incorrect words.

[0062] Thus, the bilingual corpus detection solution provided by the embodiments of the present application can be applied to the scenario of bilingual translation to detect the translation accuracy of existing bilingual corpora, thereby improving the translation quality. And with the maturity of machine translation technology, more and more manufacturers are trying to provide multilingual translation services in scenarios such as commodity, search, comment, and chat. Bilingual corpora are the key training data for machine translation models. Whether it is statistical machine translation or neural network machine translation, they rely on such corpus data. Since this solution provides a solution that can detect the accuracy of bilingual corpora and can effectively improve the translation quality of bilingual corpora, it can be widely applied to various language translation services based on machine translation and has high market value.

[0063] Based on the same inventive concept, an embodiment of the present application also provides a bilingual corpus detection device. The method corresponding to the device is the bilingual corpus detection method in the foregoing embodiment, and the principle of solving problems is similar to that of the method.

[0064] A bilingual corpus detection device provided by an embodiment of the present application can use the previous information, the subsequent information of the target word in the second corpus sentence, and the word information of the first corpus as inputs, and combines an attention mechanism. Thus, semantic and word order support can be provided for the target word to be predicted. Even if there is no strictly corresponding word or there are multiple corresponding words in one corpus for a word in another corpus, it is not easy to produce incorrect detection results. At the same time, without relying on manually annotated corpora, only bilingual corpora with relatively accurate translations need to be used as the training set to complete the training of a neural network model with an attention mechanism, and then it can accurately detect whether any bilingual corpus contains words with translation errors.

[0065] In an actual scenario, the device can be a user device, a network device, or a device formed by integrating a user device and a network device through a network. In addition, it can also be a program running on the above devices. The user device includes, but is not limited to, various terminal devices such as computers, mobile phones, and tablets; the network device includes, but is not limited to, network hosts, single network servers, multiple network server sets, or computer collections based on cloud computing. Here, the cloud consists of a large number of hosts or network servers based on cloud computing. Among them, cloud computing is a type of distributed computing, which is a virtual computer composed of a group of loosely coupled computers.

[0066] Figure 4 The structure of a bilingual corpus detection device provided by an embodiment of the present application is shown, and it can at least include a prediction processing module 410 and a detection processing module 420. Among them, the prediction processing module 410 is used to obtain the word information of the first corpus, the target word of the second corpus, and its previous and subsequent information; use the word information of the first corpus in combination with the attention mechanism, and the previous information and the subsequent information of the target word to predict the target word. The detection processing module 420 is used to determine whether the target word is incorrect according to the prediction result.

[0067] The bilingual corpus processed in the embodiment of the present application can include a first corpus and a second corpus. The first corpus and the second corpus respectively correspond to one language, and they are translations of each other. For example, "The weather is nice today" and "It's a nice day today" are Chinese-English bilingual corpora, and "I agree with that" and "Ich stimme dem zu" are English-German bilingual corpora. In an actual scenario, some content in the bilingual corpus may be translated incorrectly due to various factors. For example, for the bilingual corpus "I especially like eating red apples" and "I like eating green apple and banana", the word "especially" in the Chinese corpus is a missing word, the words "and" and "banana" in the English corpus are extra words, and the words "red" and "green" are mis-translated words. The purpose of the detection method in the embodiment of the present application is to detect these incorrectly translated words in the bilingual corpus.

[0068] The word information of the first corpus, the target word of the second corpus, and its previous and subsequent information are all information used to represent the text content included in the corpus. Taking these information as inputs can provide semantic and word order support for the target word to be predicted, making the prediction result more accurate.

[0069] In some embodiments of the present application, the word information of the first corpus, the target word of the second corpus, and its preceding context information and succeeding context information can all be represented in the form of a word vector sequence. Among them, the preceding context information can be the word vector sequence of the first N words of the target word in the second corpus, where N can be set according to the requirements of the actual scenario. For example, it can be set to a specific value such as 3, 4, 6, etc., or it can be dynamically adjusted according to the position of the target word in the second corpus. For example, it can be the position serial number of the target word minus 1. At this time, when the target word is the 4th word in the second corpus, the word vectors of the first 3 words (i.e., all the words before the target word) can be obtained as the preceding context information. For the corpus "I like eating green apple and banana", if N is set to 3 and the target word is "apple", its corresponding preceding context information is the word vectors of "like", "eating", and "green".

[0070] The succeeding context information is the word vector sequence of the last M words of the target word in the second corpus. Similar to N in the preceding context information, M can also be set according to the requirements of the actual scenario. For example, it can be set to a specific value such as 2, 3, 6, etc., or it can be dynamically adjusted according to the position of the target word in the second corpus. For example, it can be the total number of words in the second corpus minus the position serial number of the target word. At this time, when the target word is the 4th word in the second corpus and the total number of words in the second corpus is 8, the word vectors of the last 4 words (i.e., all the words after the target word) can be obtained as the succeeding context information. For the corpus "I like eating green apple and banana", if M is set to 3 and the target word is "green", its corresponding succeeding context information is the word vectors of "apple", "and", "banana".

[0071] Thus, when obtaining the target word of the second corpus and its preceding context information and succeeding context information, the target word of the second corpus, and the first N words and the last M words of the target word in the second corpus can be obtained, and then the word vector sequence of the first N words and the word vector sequence of the last M words can be obtained.

[0072] The word information of the first corpus can be the word vector sequence of the words included in the first corpus. Thus, when obtaining the word information of the first corpus, multiple words of the first corpus can be obtained, and then the word vector sequence of the multiple words can be obtained.

[0073] In a process of processing bilingual corpora, one corpus can be set as the first corpus, and the other corpus can be set as the second corpus. Taking the bilingual corpora "I like eating red apples" and "I like eating green apple and banana" as examples, in a process of processing, the Chinese corpus "I like eating red apples" can be set as the first corpus, and the English corpus "I like eating green apple and banana" can be set as the second corpus. Thus, the first corpus information is the feature vector sequence of "I", "especially", "like", "eat", "red", and "apple". After completing a process, the corpus can be swapped, that is, the Chinese corpus "I like eating red apples" can be set as the second corpus, and the English corpus "I like eating green apple and banana" can be set as the first corpus, and processed again.

[0074] In some implementations of the present application, when obtaining the word vector sequence of the multiple words, the multiple words can be first segmented to obtain the word sequence, and then the word sequence can be embedded to generate the word vector sequence of the multiple words.

[0075] Among them, the word vector sequence represents the word vector of the word sequence in any corpus, and the word sequence represents the words arranged in a certain order. In this embodiment, the multiple words can be segmented to obtain the corresponding word sequence. For example, the multiple words contained in the aforementioned bilingual corpus can be considered as sentences before word segmentation. The sentences composed of the words contained in the two corpora are "The weather is good today" and "It's a nice day today". After the Chinese corpus is segmented, the following word sequence "Today / weather / good" can be obtained, and after the English corpus is segmented, the following word sequence "It's / a / nice / day / today" can be obtained.

[0076] Among them, the word segmentation algorithm used in the embodiment of the present application can be a word segmentation algorithm based on a dictionary, such as a forward maximum matching algorithm, a reverse maximum matching algorithm, a bidirectional maximum matching algorithm, etc.; it can also be a word segmentation algorithm based on statistics, such as an N-gram algorithm; in addition, it can also be a word segmentation algorithm based on word annotation, such as a word segmentation algorithm based on BPE (byte pair encoding), etc. In actual scenarios, since various languages ​​have different grammatical habits, corresponding word segmentation algorithms can be used for different languages ​​to make the word segmentation results more accurate.

[0077] In some embodiments of the present application, the first corpus may be cleaned before segmenting the multiple words. After the corpus is cleaned, some irregularities or contents that affect subsequent processing may be adjusted to make subsequent processing more efficient and accurate. For example, in this embodiment, the cleaning of the corpus may be to unify uppercase and lowercase letters, remove punctuation marks, add start and end marks, etc.

[0078] The word vector is used to uniquely identify a specific word, and different words have their own different word vectors. For example, the specific form of the word vector can be one-hot encoding, which converts words in text form into words in encoded form. Assuming that there are a total of n words in the dictionary, an n-dimensional one-hot encoding can be created. Taking the aforementioned Chinese corpus "Today / Weather / Good" as an example, the word "Today" can be expressed as: [1,0,0,...,0], "Weather" can be expressed as: [0,1,0,...,0], and "Good" can be expressed as: [0,0,1,...,0]. Therefore, the word vector sequence of the Chinese corpus "Today's weather is good" can be expressed as a 3×n matrix, as shown in Table 1.

[0079] Because in actual scenarios, there are some problems in data processing with unique hot encoding, such as sparse data, each word is represented as an orthogonal vector in n-dimensional space, and there is no association between each other. It is impossible to use this method to calculate the similarity between words, resulting in the inability to perform efficient data processing. Therefore, after obtaining the word sequence of the bilingual corpus, the word sequence can be processed by word embedding to generate a word vector sequence of the bilingual corpus to be detected. In general, the n-dimensional vector in the dictionary can be reduced to m dimensions during word embedding, and the association relationship between words can be represented by the spatial relationship between vectors. For example, in an embodiment of the present application, after the Chinese corpus "The weather is good today" is word embedded, a word vector sequence as shown in Table 2 can be generated. Where m is less than n, thereby reducing the dimension of data processing and improving processing efficiency.

[0080] When the prediction processing module predicts the target word using the word information of the first corpus combined with the attention mechanism, as well as the previous information of the target word and the subsequent information of the target word, the prediction process can adopt a neural network model based on an encoder-decoder framework, which can include two parts of processing. First, an encoder (Encoder) based on a neural network model using an attention mechanism is encoded to obtain a feature vector sequence, and then the feature vector sequence is decoded by a corresponding decoder (Decoder) to obtain the prediction probability of the target word.

[0081] The attention mechanism used in the embodiment of the present application may be a self-attention mechanism. After combining the word information of the first corpus with the attention mechanism, the corresponding feature vector sequence can be encoded. For example, the bilingual corpus and its corresponding feature vectors are as follows: Figure 2 As shown, the content of the first corpus includes the start symbol <s>, word A, word B, ……, word C, and terminator< / s> , where the start symbol <s>and terminator< / s> It is used to indicate the start and end of a sentence and can be regarded as a word during processing. After encoding it with Self-Attention, the corresponding feature vector sequence (H0, H1, H2, …, Hn) can be obtained, where Hi represents the feature vector corresponding to the i-th word.

[0082] If the content of the second corpus includes the start symbol <s>, word a, word b, ……, word c, and terminator< / s> , if the target word is word k. After encoding it in combination with Self-Attention, the previous information can correspond to the feature vector sequence (h0, h1, ..., hk-1), and the subsequent information can correspond to the feature vector sequence (hk+1, hk+2, ..., hn), where hi represents the feature vector corresponding to the i-th word.

[0083] Since the self-attention mechanism is adopted, when predicting the target word k, the relevant feature vector will be given a certain weight, which is related to the semantics of the word corresponding to the feature vector and the target word. Take the aforementioned bilingual corpus "I like eating red apples" and "I like eating green apple and banana" as an example. When the Chinese corpus is the first corpus and the English corpus is the second corpus, if the target word "apple" in the second corpus needs to be predicted, the feature vector corresponding to "apple" in the first corpus will have a higher weight, and when predicting other target words in the second corpus, the weight of the feature vector corresponding to "apple" will be reduced. In this way, the prediction probability of the target word can be made more accurate.

[0084] The feature vector sequence obtained by encoding with the attention mechanism is input into the decoder of the neural network model, and the predicted probability of the target word can be obtained by decoding. Figure 3 The basic principle of obtaining the predicted probability of a target word in an embodiment of the present application is shown. After determining the feature vector sequence (h0, h1, ..., hk-1) corresponding to the preceding information, the feature vector sequence (hk+1, hk+2, ..., hn) corresponding to the succeeding information, and the feature vector sequence (H0, H1, H2, ..., Hn) corresponding to the word information of the first corpus, this information can be used as input to a pre-trained neural network model, which can be decoded by a decoder to output the predicted probability P of the target word.

[0085] The neural network model used for prediction in the prediction processing module can be trained in advance by the training module based on the bilingual corpus in the training set before it is needed. The bilingual corpus in the training set are all high-quality bilingual corpus, wherein high quality means that the first corpus and the corresponding second corpus have the same meaning (i.e., the translation is correct), and there are no errors that need to be detected in the embodiment of the present application, such as multiple translations, missing translations, and wrong translations. Thus, the neural network model obtained by training will output a higher prediction probability for the correct target word, and a lower prediction probability for the wrong target word.

[0086] In actual scenarios, if the target word in the second corpus is wrong (such as mistranslation or multiple translation), after inputting the preceding information, the following information and the first corpus information into the neural network model, a lower prediction probability will be output, and if the target word is correct, a higher prediction probability will be output. Therefore, when determining the wrong target word, a probability threshold can be pre-set according to the actual scenario. After obtaining the probability of the target word, the predicted probability of the target word is compared with the preset probability threshold. If the predicted probability is lower than the preset probability threshold, the target word is determined as an incorrect target word.

[0087] Thus, by traversing all the words in the second corpus, the detection of all the words in the second corpus can be completed, and all the words with translation errors in the second corpus can be identified. For example, if the Chinese corpus "I especially like to eat red apples" is set as the first corpus, the English corpus "I like eating green apple and banana" is used as the second corpus, "I", "like", "eating", "green", "apple", "and", "banana" are processed as target words in turn to complete the traversal. If the prediction probabilities obtained by the detection are shown in Table 3.

[0088] After comparing the predicted probability of the above target words with the preset probability threshold, it can be determined that "green", "and", and "banana" are lower than the probability threshold, so "green", "and", and "banana" can be identified as wrong words. If it is necessary to simultaneously identify the wrong words in the Chinese corpus "I particularly like to eat red apples", the first corpus and the second corpus can be swapped (the original first corpus is set as the second corpus, and the original second corpus is set as the first corpus), and then traversed again, so as to identify "especially" and "red" as wrong words.

[0089] In summary, the bilingual corpus detection solution provided by the embodiments of the present application can use the previous information, subsequent information, and the first corpus information of the target word in the second corpus sentence as the input of the neural network model. Thus, semantic and word order support can be provided for the target word to be predicted. Even if there is no strictly corresponding word or there are multiple corresponding words in one corpus for a word in another corpus, it is not easy to produce incorrect detection results. At the same time, without relying on manually annotated corpora, only bilingual corpora with relatively accurate translations need to be used as the training set to complete the training of the neural network model, and then it is possible to accurately detect whether there are words with translation errors in any bilingual corpus.

[0090] Therefore, the bilingual corpus detection solution provided by the embodiments of the present application can be applied to the scenario of bilingual translation to detect the translation accuracy of existing bilingual corpora, thereby improving the translation quality. And with the maturity of machine translation technology, more and more manufacturers are trying to provide multilingual translation services in scenarios such as commodity, search, review, and chat. Bilingual corpora are the key training data for machine translation models, whether it is statistical machine translation or neural network machine translation, which rely on such corpus data. Since this solution provides a solution that can detect the accuracy of bilingual corpora and can effectively improve the translation quality of bilingual corpora, it can be widely applied to various language translation services based on machine translation and has high market value.

[0091] In addition, a part of the present application can be applied as a computer program product, for example, computer program instructions. When executed by a computer, through the operation of the computer, the methods and / or technical solutions according to the present application can be invoked or provided. The program instructions for invoking the methods of the present application may be stored in a fixed or removable recording medium, and / or transmitted through a data stream in a broadcast or other signal-bearing medium, and / or stored in the working memory of a computer device running according to the program instructions. Here, some embodiments according to the present application include a Figure 5 computing device as shown, which includes one or more memories 510 storing computer-readable instructions and a processor 520 for executing the computer-readable instructions. When the computer-readable instructions are executed by the processor, the device is caused to execute the methods and / or technical solutions based on the foregoing multiple embodiments of the present application.

[0092] In addition, some embodiments of the present application also provide a computer-readable medium, on which computer program instructions are stored, and the computer-readable instructions can be executed by a processor to implement the methods and / or technical solutions of the foregoing multiple embodiments of the present application.

[0093] It should be noted that the present application can be implemented in software and / or a combination of software and hardware, for example, can be implemented using an application specific integrated circuit (ASIC), a general purpose computer or any other similar hardware device. In certain embodiments, the software program of the present application can be executed by a processor to implement the above steps or functions. Similarly, the software program of the present application (including related data structures) can be stored in a computer-readable recording medium, for example, a RAM memory, a magnetic or optical drive or a floppy disk and similar devices. In addition, some steps or functions of the present application can be implemented using hardware, for example, as a circuit that cooperates with a processor to perform each step or function.

[0094] It is obvious to those skilled in the art that the present application is not limited to the details of the above exemplary embodiments, and that the present application can be implemented in other specific forms without departing from the spirit or basic features of the present application. Therefore, from any point of view, the embodiments should be regarded as exemplary and non-restrictive, and the scope of the present application is limited by the attached claims rather than the above description, so it is intended to include all changes that fall within the meaning and scope of the equivalent elements of the claims in the present application. Any figure mark in the claims should not be regarded as limiting the claims involved. In addition, it is obvious that the word "comprising" does not exclude other units or steps, and the singular does not exclude the plural. Multiple units or devices stated in the device claim can also be implemented by one unit or device through software or hardware. The words first, second, etc. are used to indicate names, and do not indicate any particular order.

Claims

1. A bilingual corpus detection method, wherein, the method includes: obtaining the word information of the first corpus, the target word of the second corpus, and its preceding context information and following context information; using the word information of the first corpus in combination with the attention mechanism, as well as the preceding context information and the following context information of the target word, to predict the target word and obtain a prediction result; wherein, the prediction result includes the prediction probability of the target word; determining whether the target word is incorrect according to the prediction result.

2. The method according to claim 1, wherein, the obtaining of the word information of the first corpus, the target word of the second corpus, and its preceding context information and following context information includes: obtaining multiple words of the first corpus and obtaining a word vector sequence of the multiple words; obtaining the target word of the second corpus, as well as the first N words and the last M words of the target word in the second corpus, and obtaining a word vector sequence of the first N words and a word vector sequence of the last M words.

3. The method according to claim 2, wherein, obtaining the word vector sequence of the multiple words includes: performing word segmentation processing on the multiple words to obtain a word sequence; performing word embedding processing on the word sequence to generate the word vector sequence of the multiple words.

4. The method according to claim 3, wherein, before performing word segmentation processing on the multiple words, it further includes: cleaning the first corpus.

5. The method according to claim 1, wherein, the determining whether the target word is incorrect according to the prediction result includes: comparing the prediction probability of the target word with a preset probability threshold, and if the prediction probability is lower than the preset probability threshold, determining that the target word is incorrect.

6. A bilingual corpus detection device, wherein, the device includes: a prediction processing module, configured to obtain the word information of the first corpus, the target word of the second corpus, and its preceding context information and following context information; use the word information of the first corpus in combination with the attention mechanism, as well as the preceding context information and the following context information of the target word, to predict the target word and obtain a prediction result; wherein, the prediction result includes the prediction probability of the target word; a detection processing module, configured to determine whether the target word is incorrect according to the prediction result.

7. The device according to claim 6, wherein, the prediction processing module is configured to obtain multiple words of the first corpus and obtain a word vector sequence of the multiple words; obtain the target word of the second corpus, as well as the first N words and the last M words of the target word in the second corpus, and obtain a word vector sequence of the first N words and a word vector sequence of the last M words.

8. The device according to claim 7, wherein, the prediction processing module is configured to perform word segmentation processing on the multiple words to obtain a word sequence; perform word embedding processing on the word sequence to generate the word vector sequence of the multiple words.

9. The device according to claim 8, wherein, the device further includes: a cleaning module, configured to clean the first corpus before performing word segmentation processing on the multiple words.

10. The device according to claim 6, wherein, the detection and processing module is configured to compare the predicted probability of the target word with a preset probability threshold, and if the predicted probability is lower than the preset probability threshold, determine that the target word is incorrect.

11. A computing device, wherein, the device includes a memory for storing computer program instructions and a processor for executing the computer program instructions, wherein when the computer program instructions are executed by the processor, the device is triggered to execute the method according to any one of claims 1 to 5.

12. A computer-readable medium having computer program instructions stored thereon, the computer program instructions being executable by a processor to implement the method according to any one of claims 1 to 5.

Citation Information

Patent Citations

  • Filtering method and device for corpuses

    CN104750820A

  • Data processing method and device and device for data processing

    CN111160046A